Skip to main content

Local AI models for Android: choose what fits

Choose and compare local AI models for Android using task fit, memory, storage, quantisation, runtime support, license, heat, and device testing.

What you need to know

The right local AI model is the smallest compatible artifact that performs your real tasks reliably. Check the exact file, quantisation, runtime, modality, memory needs, storage headroom, license, and supported prompt format. Then measure first output, answer quality, cancellation, heat, and repeated runs on your own device.

Choose from the task backwards

Start by writing down what the model must do: short chat, rewriting, document questions, multilingual work, image understanding, or code. A model family name does not guarantee that every artifact supports the same inputs, license, context length, or runtime. Verify the exact checkpoint and file offered by the app.

Parameter count is only one signal. Quantisation changes file size and memory behaviour, while architecture and runtime optimisation affect speed. Available storage is not the same as usable memory. The operating system, graphics stack, active applications, context length, and temporary buffers all compete for resources during generation.

Compare models with controlled tests

Use the same device, app version, runtime settings, prompt, and approximate context when comparing two models. Record time to first visible output separately from generation speed, then assess whether each answer follows instructions and handles the intended subject. Repeat prompts after the phone warms up because sustained performance can differ from a single cold run.

A smaller model that starts promptly, stops cleanly, and answers a narrow task well may be more useful than a larger model that strains the device. Keep model cards, licences, and runtime compatibility visible in the decision. Treat universal winner claims cautiously unless they are supported by reproducible tests on hardware comparable to yours.

More local ai models articles

  1. LFM2.5 vs Qwen3.5 for Mobile Local AI

    LFM2.5 1.2B and Qwen3.5 0.8B are compact model families with different architectures, language claims, licenses, and intended uses. There is no honest universal winner: compare the exact quantised artifacts in the same app, on the same phone, with your own prompts.

    · 8 min read
  2. Why Local AI Is Slow on a Phone and How to Diagnose It

    Local AI speed has two separate parts: time before the first visible token and the rate of tokens after output begins. Diagnose them separately.

    · 10 min read
  3. What Is AI Model Quantization for Local AI on a Phone?

    Quantization can make a local model smaller and more practical, but a lower bit count is not a universal speed or quality score. Learn how precision, runtime support, memory, and task testing fit together.

    · 10 min read
  4. ExecuTorch 1.4 Expands Android and Qualcomm On-Device AI Support

    PyTorch's on-device runtime now covers more Android, Qualcomm, Arm, and embedded workflows. The release improves deployment options without promising that every model will run faster on every phone.

    · 5 min read
  5. Meta Muse Glimmer 30B Brings Local AI Agents to Consumer Hardware

    Meta has released a 30B multimodal agent model with official local artifacts. The smallest build still needs about 17 GB before vision, context, and speculative decoding.

    · 5 min read
  6. ONNX Runtime 1.29 Moves Browser AI Towards Native WebGPU

    Microsoft's inference runtime adds WebGPU attention and quantisation support while beginning the move away from WebGL and JSEP.

    · 5 min read
  7. Qwen3.8 27B Hardware Requirements Start Above 30 GB

    Qwen's smaller 27B release is practical for high-memory computers, but the official files and context cache still put it beyond ordinary phones.

    · 5 min read
  8. GLM-5.3 Open Weights Delayed After Cyber Tests

    Z.ai has launched GLM-5.3 through its hosted service while delaying the model weights after its own tests found stronger cyber capabilities.

    · 5 min read
  9. OpenAI Pauses Frontier Training Over Astra Cyber Risk

    OpenAI has slowed some frontier training while it changes security, monitoring, and alignment controls around its unreleased Astra model.

    · 5 min read
  10. Koboldcpp v1.120 Adds DirectIO Loading and Two New Models

    The 29 August Koboldcpp release speeds model loading, unlocks two efficient new MoE models, and brings tool calling to local setups.

    · 3 min read
  11. Anthropic's Fable 5.1 Arrives Two Months After Export Shutdown

    Anthropic's newest frontier models ship with 25 to 45 percent cost cuts and unusually candid benchmark caveats.

    · 4 min read
  12. Microsoft's VibeVoice Streaming ASR Models Go Open Weights

    New VibeVoice-ASR-Streaming weights turn hour-long audio into speaker-tagged, timestamped transcriptions locally in a single pass.

    · 4 min read
  13. Nvidia Confirms Its $13 Billion Acquisition of Hugging Face

    Nvidia's SEC filing confirms the acquisition: $11.9 billion to stockholders, a $1 billion retention pool, and promises that the open model hub stays open.

    · 4 min read
  14. Nvidia PAIR Turns Your Home PCs Into a Local AI Cluster

    PAIR is a free virtual inference router for your home network: it finds idle machines and spreads parallel AI jobs across them, entirely offline.

    · 4 min read
  15. MBZUAI Ships K2 Horizon: Six Fully Open Models From 0.9B to 375B

    K2 Horizon is a six-model open fleet spanning watch-sized to datacenter-sized, and the training data and methodology are public too.

    · 4 min read
  16. Microsoft's Project Zenith Is Windows Tuned for Local AI Development

    Microsoft announced Project Zenith: a ready-to-code Windows setup with local AI in mind, debuting on AMD Ryzen AI Halo mini PCs with 128GB memory.

    · 4 min read
  17. OpenAI Launches GPT-6 Astra With Staged Access and Cyber Limits

    GPT-6 Astra starts with limited organizations before reaching ChatGPT plans, APIs, Azure, and AWS Bedrock, while cybersecurity access stays restricted and monitoring broadens.

    · 6 min read
  18. Claude Formalized Fermat's Last Theorem in Lean in 11 Days

    A multi-agent Claude workflow formalized Fermat's Last Theorem in Lean: 13 million lines, 30,300 theorems, human input limited to priority nudges.

    · 4 min read
  19. Corporate America Is Switching to Open-Weight AI Models to Cut Costs

    A New York Times report says corporate America is hooked on open-source AI: AT&T runs 40 percent of employee queries on open models and wants more.

    · 4 min read
  20. GLM-5.3 Open Weights Ship After the Cyber Safety Delay

    The delayed GLM-5.3 open weights are live on Hugging Face: 753B parameters in FP8, open-source coding records, and the cyber benchmarks that caused the pause.

    · 4 min read
  21. MiniCPM5-2B Packs 131K Context and SOTA Results Into 2.5B Parameters

    A 2.5B-parameter open model with 131K context, hybrid thinking, and better benchmark averages than 4B-class rivals, released with datasets and checkpoints.

    · 4 min read
  22. Mistral Raises €3 Billion to Push Sovereign Open-Weight AI

    Mistral announced a €3 billion Series D led by Samsung at a valuation above €21 billion, betting that sovereign open-weight AI can reach the frontier.

    · 4 min read
  23. ONNX Runtime 1.30 Adds Quantized WebGPU KV Caches

    Microsoft's inference runtime expands browser and Arm64 paths while adding quantized attention caches and model-loading hardening.

    · 4 min read