Skip to main content

Dimensity 9600 Pro Claims Support for 30B On-Device AI Models

MediaTek announced the Dimensity 9600 Pro on September 15, 2026, and says its NPU supports on-device AI models up to 30 billion parameters. It also claims faster language-model prompt processing and better token-generation efficiency than its previous flagship. No shipping phone or independently measured local chat app yet establishes what a 30B model will feel like on this chip. Memory and software support remain decisive.

On this page

The NPU 1090 carries the 30B claim

MediaTek announced the Dimensity 9600 Pro on September 15. Its NPU 1090 is said to support on-device applications using models with up to 30 billion parameters. MediaTek also claims 51 percent higher language-model prefill performance and 55 percent more generated tokens per watt than its previous flagship processor. These are vendor figures. The company has not published a comparable phone-and-app test that would show a user's time to first visible word or sustained chat speed.

Prefill is the work of processing the prompt before generation begins. An improvement there could make a local assistant feel quicker when a conversation or document excerpt is long. Tokens per watt concerns the energy used while generating output. It does not tell us the absolute tokens per second, battery drain during a full session, or how the chip behaves once a phone warms up.

MediaTek pairs the NPU with a second, more efficient unit for always-on tasks. It also lists a 34.5 MB cache and support for LPDDR6 memory and UFS 5.0 storage. Phone makers choose the actual RAM, storage, cooling, and software configuration, so the chip specification alone does not describe a product a reader can buy today. MediaTek expects the first phones using the 9600 Pro or 9600M this quarter.

A 30B ceiling is not a phone recommendation

Thirty billion parameters at four bits per weight require about 15 GB just for the raw weights. A real model also needs scales or other quantization metadata, a key-value cache, runtime buffers, the app, and memory for Android. Some weights may use higher precision. The arithmetic makes clear why “supports up to 30B” cannot be read as “any 30B model will run comfortably on any 9600 Pro phone.”

There is a second gate: the app must have a compatible model artifact and runtime path for this NPU. A checkpoint that runs on a desktop GPU may have unsupported operators, a different quantization layout, or a tokenizer and prompt setup that the phone app has not qualified. If an app falls back to the CPU or GPU, the NPU's claimed efficiency may not apply. The Android model-choice guide explains why memory, format, and a real-device test matter more than a parameter count.

MediaTek's product sheet identifies the NPU 1090 and the Arm Mali G2-Ultra NX GPU, but it does not provide a public compatibility list for popular local chat models. Arm's separate GPU announcement describes neural acceleration inside graphics shader cores. That architecture is related to the chip's on-device story, but a GPU capability and an NPU model-size claim should not be combined into one assumed benchmark.

What a buyer can check when phones arrive

Android Central independently reported the September 15 launch and noted that no 9600 Pro phone was available for testing at publication. Its coverage repeats MediaTek's 30B and efficiency figures as company claims, rather than independent measurements.

Once a phone ships, the useful evidence is a named model and artifact, the app version, the actual execution backend, installed and available RAM, cold-load time, first visible output, sustained generation, and heat after several prompts. An airplane-mode test can confirm that the claimed local workflow does not quietly depend on a hosted model. A Stop-and-retry test catches a different failure: a fast chip cannot help if the app leaves generation stuck.

CuriousLM runs supported AI models locally on your device. Try CuriousLM.