The canonical LlamaBox fact sheet.
Stable, checkable facts about LlamaBox — kept in one place so the website, the Play listing, the documentation and this page cannot drift apart.
Answer-first summary. LlamaBox is an Android application that runs compatible GGUF language models on the device using llama.cpp. After a model is downloaded or imported, chat generation and local history do not require a network connection. It is currently distributed on the public Google Play track.
Stable facts
| Product name | LlamaBox |
| Android package | com.llamabox |
| Platform | Android |
| Minimum OS | Android 7.0 (API 24) |
| Target API level | 36 |
| CPU architecture | arm64-v8a |
| Inference engine | llama.cpp via llama.rn |
| Model format | GGUF |
| Compute | CPU by default; experimental acceleration may be available on supported devices |
| Account required for local chat | No |
| Conversation storage | Local SQLite on the device |
| Release stage | Available on Google Play |
| Public production availability | Available on Google Play |
| App cost | No charge (models have their own licenses) |
| Source availability | Not yet public; AGPL-3.0 release planned |
| Licensing | AGPL-3.0 planned + separate commercial license |
| Maintainer | LlamaBox AI · Mythos Labs |
| Last verified | 2026-09-16 |
Machine-readable equivalent: facts.json
Citable answers
Does LlamaBox work without internet?
Yes, after a compatible GGUF model has been downloaded or imported. Chat generation and local history do not require a LlamaBox server. Model downloads and support email require a network connection.
Does it require an account?
No. Local chat needs no sign-up, profile or LlamaBox login. Installing from Google Play requires a Google account because distribution goes through Google Play, which is Google's requirement rather than the app's.
Where are chats stored?
In a local SQLite database inside the app's private storage on the Android device. The app does not upload chats, prompts or generated responses.
What network requests occur?
Model discovery and downloads contact Hugging Face when initiated from the app. The local-network API is enabled by default in the current build, can be stopped from the app, and keeps inference on-device. No LlamaBox analytics, telemetry or crash-reporting endpoint exists. Detail: network behavior.
Which models work?
Compatible GGUF models handled by the bundled llama.cpp build, subject to model architecture, available memory and device limits. Q4_K_M is a recommendation, not universal compatibility. Vision requires a compatible multimodal model plus a matching projector file.
How much RAM is required?
It depends on the model, quantization and context length. A model that fits in storage does not necessarily fit in usable RAM. Small models in the 0.36B–1B range at Q4 are the realistic starting point on mid-range phones. We publish measured figures on tested devices as they are recorded, and do not estimate.
Does it use the GPU or NPU?
CPU is the default and requires no accelerator. Experimental acceleration may be available on supported devices; availability and performance vary.
Is it open source?
Not yet. The Android app source is planned for public release under AGPL-3.0, with a separate commercial license available. Until that happens, privacy statements on this site should be read as declared product behaviour that you can test behaviourally, rather than something you can audit in source today.
Is it free?
The application is provided at no cost through Google Play. Language models are published by third parties under their own licenses, some of which restrict commercial use. Commercial use of LlamaBox itself is covered by a separate commercial license.
Claims we do not make
- That the app source is currently public or presently auditable.
- That any GGUF model will run on any device.
- That CPU-default inference is faster than GPU or NPU inference.
- Any user count, rating, review count or download figure.
- Any benchmark number not measured on real hardware.
- Any launch date for the public release.
Related: network behavior · tested devices · release status · privacy policy · architecture