Apple’s Neural Engine has been a core part of on‑device machine learning for several generations. How exactly does the Neural Engine’s architecture differ from the CPU and GPU when handling tasks like image recognition or voice activation? What are the main architectural changes introduced in the latest generation, and how do they impact performance and power efficiency?
How does the Neural Engine differ from CPU/GPU in recent iPhone models?
👁️ 50 görüntüleme💬 1 cevap❤️ 0 beğeni
1 Cevap
The Neural Engine (NE) is basically a matrix‑multiply accelerator built from a massive grid of small, fixed‑function compute units that are wired together for 8‑bit/16‑bit tensor ops. Unlike the CPU’s out‑of‑order scalar cores or the GPU’s SIMD shaders, the NE doesn’t have a deep instruction pipeline or a cache hierarchy – it streams data straight into a systolic array, performs the multiply‑accumulate in a single clock cycle, and writes the result back. This makes it ideal for the dense linear‑algebra workloads you see in image‑classification or voice‑activation models, where the same operation is repeated millions of times on different data.
In the A15/A16 generation Apple moved from a 16‑core NE (A14) to a 16‑core design with doubled MAC throughput per core, plus a new “mixed‑precision” path that can handle both 8‑bit and 16‑bit operands without extra conversion overhead. The newer NE also integrates a dedicated high‑bandwidth SRAM buffer that sits right next to the compute fabric, cutting the number of trips to main memory. That reduces latency and, more importantly, cuts the energy per inference dramatically – Apple claims roughly a 2‑3× improvement in TOPS/W over the previous generation.
Because the NE is isolated from the CPU/GPU power domains, the system can spin it up only when a model is needed and shut it down instantly, keeping the rest of the chip in a low‑power state. The CPU can still handle control‑flow, pre‑ and post‑processing, while the GPU can offload any residual graphics‑oriented ML work (e.g., vision pipelines that need texture sampling). In practice this partitioning means a voice‑assistant request that used to cost a few hundred milliwatts on the CPU/GPU alone now sits in the tens of milliwatts on the NE, and latency drops from ~30 ms to sub‑10 ms on the latest iPhones.