On-Device AI Models Might Be the Next Reason to Upgrade Your Phone
The iPhone 17 runs a 3 billion parameter language model on-device at 30 tokens per second. Obviously, the average consumer has no idea what that sentence means, and Apple hasn’t figured out how to make them care.
I believe that’s about to change. Apple now has complete access to Google’s Gemini model in its own data centers, with the ability to distill it into smaller models built for iPhones and iPads. Knowledge distillation works like this: you take a large model, have it perform tasks with detailed reasoning, then feed those reasoning traces to a smaller model until the student learns to mimic the teacher. The smaller model ends up far more capable than if you’d trained it from scratch on the same data. Apple can now do this with the full Gemini, not just their own in-house models, and the distilled output runs locally. No internet required.
Smartphones haven’t had a real upgrade story in years. The camera is great. The screen is great. The processor was fast enough three generations ago. Battery life has overtaken price as the top purchase driver for the first time. The global replacement cycle has stretched to 3.5 years . People hold onto their phones because nothing about the new one feels different enough. Deloitte’s 2025 TMT Predictions report frames on-device generative AI as the feature that could break this cycle, if the experience delivers on the promise. On-device AI might become the next reason.
In the late 1990s it was megahertz: Intel and AMD raced clock speeds past the point where consumers could distinguish real-world performance differences, but the number on the box still drove purchases. Then it was megapixels. Samsung shipped a 200 MP camera sensor knowing that most phones use 16-to-1 pixel binning to output a 12.5 MP image by default.
Parameters could be next. The iPhone 17’s standard A19 chip has 8GB of RAM. The Pro gets 12GB with faster memory bandwidth, which determines how large a model the phone can run and how quickly. Samsung’s 2026 flagships with the Exynos 2600 hit 80 TOPS on a 2nm process, more than double the prior generation. These are already the numbers in press releases. It’s not hard to imagine an Apple keynote where someone says, with rehearsed enthusiasm, that the iPhone 18 Pro runs a 7 billion parameter model while the standard model is limited to 3 billion.
The difference from previous spec wars is that this one might actually correlate with user experience. Megahertz past a certain threshold didn’t make Word open faster. Megapixels past 12 MP didn’t make photos look better on a phone screen. But a 7 billion parameter model running locally outperforms a 3 billion one on nearly every task. It handles longer documents, follows more complex instructions, holds better conversational context.