GGUF
Run GGUF models on Android.
Builders already quantize for desktops. LlamaBox brings those GGUF files to arm64 phones without a cloud relay.
Why GGUF on mobile
GGUF + llama.cpp is the de facto portable stack for local LLMs. Android arm64 can run the same ecosystem — carefully — with smaller models and honest thread counts.
Recommendations
- Prefer Q4_K_M to start
- Match model size to free RAM (leave headroom for OS + UI)
- Context defaults to 2048 (configurable); vision may auto-raise
- Vision models need base GGUF + mmproj
Workflow
- Install LlamaBox
- Download from the in-app hub or import a local GGUF
- Load model · chat offline
Deep stack notes: architecture · tutorial: how to run an LLM on Android.
FAQ
What is GGUF?
GGUF is a file format for quantized LLM weights commonly used with llama.cpp. It packs tensors and metadata for efficient local loading.
Which quantization should I use?
Q4_K_M is a strong default balance of quality and size for mobile. Smaller quants save RAM; larger improve quality if you have headroom.
Can I import my own GGUF?
Yes — LlamaBox supports in-app downloads and local import workflows for models that fit device memory.
Available on Google Play
Try private offline AI on Android.
Install from Google Play, download a compatible model, then chat offline.