GGUF

Run GGUF models on Android.

Builders already quantize for desktops. LlamaBox brings those GGUF files to arm64 phones without a cloud relay.

Compact model modules align with the internal layers of an Android phone.

Why GGUF on mobile

GGUF + llama.cpp is the de facto portable stack for local LLMs. Android arm64 can run the same ecosystem — carefully — with smaller models and honest thread counts.

Recommendations

  • Prefer Q4_K_M to start
  • Match model size to free RAM (leave headroom for OS + UI)
  • Context defaults to 2048 (configurable); vision may auto-raise
  • Vision models need base GGUF + mmproj

Workflow

  1. Install LlamaBox
  2. Download from the in-app hub or import a local GGUF
  3. Load model · chat offline

Deep stack notes: architecture · tutorial: how to run an LLM on Android.

FAQ

What is GGUF?
GGUF is a file format for quantized LLM weights commonly used with llama.cpp. It packs tensors and metadata for efficient local loading.
Which quantization should I use?
Q4_K_M is a strong default balance of quality and size for mobile. Smaller quants save RAM; larger improve quality if you have headroom.
Can I import my own GGUF?
Yes — LlamaBox supports in-app downloads and local import workflows for models that fit device memory.
Available on Google Play

Try private offline AI on Android.

Install from Google Play, download a compatible model, then chat offline.