The Snapdragon 8 Elite Gen 5, MediaTek Dimensity 9500, and Apple A19 Pro each claim leadership in on-device AI, but comparing their NPUs head-to-head runs into a real problem: none of the three companies published a directly comparable raw TOPS figure this generation, and each is measuring itself against its own predecessor rather than its direct rivals. That's not an accident — it reflects how differently these three chips actually approach on-device AI, and reading past the marketing percentages reveals genuinely different strategic bets, not just three versions of the same NPU.
Why On-Device AI Benchmarking Is Genuinely Harder Than CPU or GPU
CPU and GPU performance has decades of standardized benchmarks — Geekbench, AnTuTu, 3DMark — that let you compare a Snapdragon chip against an Apple chip on the same numeric scale. NPU performance doesn't have an equivalent industry standard yet, partly because the underlying architectures are genuinely different (Qualcomm's Hexagon design, MediaTek's dedicated APU, and Apple's Neural Engine tightly fused with its GPU all process AI workloads differently), and partly because raw TOPS (trillions of operations per second) figures don't reliably predict real-world performance across different architectures the way clock speed once did for CPUs. That's why every company covered here reports its NPU gains as a percentage improvement over its own previous generation, rather than a head-to-head number against its two biggest rivals.
Qualcomm's Hexagon NPU: Betting on Agentic AI
Qualcomm's messaging around the Snapdragon 8 Elite Gen 5's Hexagon NPU centers specifically on "agentic AI" — on-device AI systems capable of completing multi-step tasks rather than just answering single queries. The company reports the new NPU is 37% faster and 16% more power-efficient than the previous generation's Hexagon NPU, positioning it explicitly for generative AI workloads running locally on the device rather than relying on cloud processing. That framing matters strategically: Qualcomm is betting that phone makers and app developers will build more autonomous, task-completing AI features, and wants its NPU positioned as the hardware foundation for that shift specifically, not just faster photo processing or voice recognition.
MediaTek's NPU 990: Optimizing for LLM Token Throughput
MediaTek's Dimensity 9500 takes a more specific technical angle with its NPU 990, which the company reports as twice as fast as its predecessor while cutting peak power consumption by 56%. Where MediaTek's messaging diverges from Qualcomm's is specificity: MediaTek explicitly frames the NPU 990's improvements around large language model token throughput — essentially, how quickly the chip can generate the next word or token when running a language model directly on the device, without a cloud connection. That's a more measurable, LLM-specific claim than Qualcomm's broader "agentic AI" framing, and it signals MediaTek is positioning itself specifically for on-device chatbot and text-generation features rather than a wider range of AI task types.
Apple's 16-Core Neural Engine: Fused With the GPU, Not Separate From It
Apple's approach differs architecturally from both Android chipmakers. The A19 Pro's 16-core Neural Engine is described as tightly integrated with Apple's GPU-based accelerators, fusing AI and graphics processing rather than treating the NPU as a fully separate compute block. Apple's public messaging around the A19 Pro's AI capabilities is notably less specific on standalone NPU performance percentages compared to Qualcomm and MediaTek, instead emphasizing how AI and graphics tasks work together for what the company describes as smoother on-device intelligence — a framing that fits Apple's broader pattern of prioritizing integrated system-level experience over publishing standalone component benchmarks.
What This Actually Means for AI Features You Use
The practical difference shows up less in raw specs and more in which on-device AI features each platform actually ships. Apple Intelligence features on iPhones lean on the tight Neural Engine-GPU integration for tasks like on-device writing tools and image generation. Android flagships running Snapdragon chips increasingly ship with Google's Gemini Nano models running partly on-device, benefiting from Qualcomm's agentic AI-focused Hexagon improvements. MediaTek's LLM-throughput focus on the Dimensity 9500 shows up most directly in how responsive on-device chatbot and text-generation features feel on Vivo, Honor, and other Dimensity-powered flagships, since faster token generation is the specific bottleneck that determines whether an on-device chatbot response feels instant or noticeably laggy.
The Honest Verdict: No Clean Winner Without Real Workload Testing
Without a standardized, independent NPU benchmark that all three chips have been run through, any claim of an outright NPU "winner" is really a claim about which vendor's self-reported metric you trust most, or which specific AI task you care about. Qualcomm's percentage gains emphasize breadth (agentic AI across many task types), MediaTek's emphasize a specific measurable bottleneck (LLM token throughput), and Apple's emphasize system-level integration over standalone component claims at all. That's a genuinely different set of priorities being marketed as three versions of "better AI," and the most honest advice is to judge based on the specific on-device AI feature you actually plan to use daily, rather than any single percentage figure from a spec sheet.
NPU Approach at a Glance
| Chip | NPU | Reported gain vs. predecessor | Strategic focus |
|---|---|---|---|
| Snapdragon 8 Elite Gen 5 | Hexagon NPU | 37% faster, 16% more power-efficient | Agentic AI — multi-step, autonomous on-device tasks |
| Dimensity 9500 | NPU 990 | 2x faster, 56% lower peak power | LLM token throughput — faster on-device chatbot responses |
| Apple A19 Pro | 16-core Neural Engine | Not separately quantified vs. GPU-fused performance | System-level integration between AI and graphics compute |
Frequently Asked Questions
Which chip has the fastest NPU: Snapdragon, Dimensity, or Apple?
There's no single answer, because none of the three companies published a directly comparable raw TOPS figure this generation. Each reports gains as a percentage improvement over its own previous generation rather than a head-to-head benchmark against its rivals, so the honest comparison depends on which specific AI task you care about.
What is "agentic AI" and why does Qualcomm emphasize it?
Agentic AI refers to on-device AI systems capable of completing multi-step tasks autonomously, rather than just answering single queries. Qualcomm's Hexagon NPU messaging for the Snapdragon 8 Elite Gen 5 centers specifically on this capability, positioning the chip as hardware built for more autonomous AI features.
Does MediaTek's Dimensity 9500 actually run AI chatbots faster?
MediaTek specifically frames its NPU 990 improvements around LLM token throughput, meaning how quickly the chip generates each word when running a language model on-device. This is a more measurable claim than general AI performance, and it should show up most directly in how responsive on-device chatbot features feel.
Why doesn't Apple publish specific NPU performance numbers like Qualcomm and MediaTek?
Apple's A19 Pro tightly integrates its 16-core Neural Engine with its GPU-based accelerators rather than treating them as fully separate components, and the company's public messaging emphasizes this system-level integration over standalone NPU benchmarks.
