Summary:
The Apple M5 chip features a 16-core Neural Engine, a crucial component often overlooked in standard benchmarks, yet it represents a significant, long-term investment by Apple.
Apple's vertical integration, controlling chip design, memory architecture, OS, and developer frameworks like Core ML, allows for deep optimization, such as unified memory and embedded neural accelerators within the GPU. Since its introduction in 2017 with 2 cores, the Neural Engine has grown to 16 cores, achieving 38+ trillion operations per second in the M5, alongside GPU accelerators. This expansion reflects Apple's strong commitment to on-device AI, prioritizing privacy, cost-effectiveness, and low latency by processing AI workloads directly on the device. The industry trend, mirrored by companies like Qualcomm (Snapdragon X2 Elite with 80 TOPS) and Intel (Lunar Lake with 48 TOPS), also indicates a widespread shift towards integrating dedicated AI hardware. This ongoing investment in local AI compute aims to support future applications that will leverage powerful on-device machine learning capabilities, even as software development catches up to the hardware's potential.