Models

GGUF models that run on Android.

A practical index of local LLM models for Android phones. Quantized GGUF weights, RAM estimates, and what works offline with LlamaBox.

Compact model modules align with the internal layers of an Android phone.

What is the LlamaBox model hub?

The model hub inside LlamaBox is a curated list of GGUF models that fit Android phones. Each entry includes recommended quantization, estimated RAM use, and whether the model supports vision or TTS. You can download directly in the app or import a GGUF you already have.

Quick-start model picks

Model familySize classBest forApprox. RAMVision
Qwen2.5-Instruct0.5B–3B Q4_K_MGeneral chat, drafting, coding help0.5–2 GBNo
SmolLM2360M–1.7B Q4_K_MFast answers on low-RAM phones0.3–1 GBNo
Gemma 3 4B IT4B Q4_K_MBalanced quality on 8GB+ phones1.5–3 GBYes (with mmproj)
MiniCPM-V / InternVL22B–4B Q4_K_MOn-device image understanding1.5–3 GBYes (with mmproj)

These are starting points, not guarantees. Real performance depends on free RAM, Android version, background apps, and the exact quantization. Verified device measurements will be published as they are recorded.

How to choose a quantization

  • Q4_K_M — best default for mobile. Good quality at roughly half the float16 size.
  • Q5_K_M / Q6_K — slightly better quality if you have RAM headroom.
  • Q3_K_M / Q2_K — smaller but quality drops noticeably; use only when RAM is tight.

Vision models need an mmproj

Vision-capable GGUF models also need a matching multimodal projector file (mmproj). LlamaBox pairs base and mmproj automatically when they share a filename prefix. See GGUF on Android for naming conventions.

Where to get GGUF models

  • Hugging Face — largest collection of quantized models
  • GGUF model catalog — filter by GGUF format
  • In-app LlamaBox model hub (curated for Android RAM limits)

Next: LLM download guide · how to run an LLM on Android · GGUF on Android.

FAQ

What GGUF models work on Android?
Smaller quantized models generally work. Start with 0.5B–3B class Q4_K_M weights. Vision models need a base GGUF plus an mmproj file. LlamaBox detects presets from filename hints.
Can I run Llama 3 on Android?
Yes, if you use a small enough GGUF quantization (for example a 1B–3B parameter variant). Larger 8B+ models usually need more RAM than mid-range phones offer.
What is the best model for offline chat on a phone?
For most users, Qwen2.5 1.5B Q4_K_M or SmolLM2 360M/1.7B Q4_K_M are strong starting points: small, fast, and capable enough for drafting and问答.
How much RAM does a model need?
A good rule of thumb: the GGUF file size plus 200 MB–1 GB runtime/KV overhead. A 1 GB model can run on a 4–6 GB RAM phone if background apps are closed.
Where do I download models?
Use the LlamaBox in-app model hub, or import GGUF files you download from Hugging Face and other model hubs.
Available on Google Play

Try private offline AI on Android.

Install from Google Play, download a compatible model, then chat offline.