tech
The NPU in your phone keeps improving—why isn’t that making AI better?
Almost every technological innovation of the past several years has been laser-focused on one thing: generative AI. Many of these supposedly revolutionary systems run on big, expensive servers in a data center somewhere, but at the same time, chipmakers are crowing about the power of the neural processing units (NPU) they have brought to consumer devices. Every few months, it’s the same thing: This new NPU is 30 or 40 percent faster than the last one. That’s supposed to let you do something important, but no one really gets around to explaining what that is.

TL;DR
- Generative AI innovations are heavily marketed, with chipmakers emphasizing on-device NPUs, but most AI tools run in the cloud.
- NPUs are specialized components in Systems-on-a-Chip (SoCs) designed for parallel computing, evolving from DSPs.
- Cloud-based AI models have significantly larger context windows and parameter counts than models optimized for edge devices.
- Edge AI offers advantages in user privacy and reliability by processing data locally, avoiding data sharing with cloud servers.
- Despite advancements, cloud AI currently dominates due to its superior performance, accuracy, and resource availability.
- A hybrid approach is common, but many advertised on-device features still rely on cloud processing.
- Increased RAM capacity in devices is being driven by the potential for on-device AI, even if cloud AI is currently more prevalent.