Product facts

The canonical LlamaBox fact sheet.

Stable, checkable facts about LlamaBox — kept in one place so the website, the Play listing, the documentation and this page cannot drift apart.

An exploded phone diagram shows interface, storage, model, inference and CPU layers connected on-device.

Answer-first summary. LlamaBox is an Android application that runs compatible GGUF language models on the device using llama.cpp. After a model is downloaded or imported, chat generation and local history do not require a network connection. It is currently distributed on the public Google Play track.

Stable facts

Product nameLlamaBox
Android packagecom.llamabox
PlatformAndroid
Minimum OSAndroid 7.0 (API 24)
Target API level36
CPU architecturearm64-v8a
Inference enginellama.cpp via llama.rn
Model formatGGUF
ComputeCPU by default; experimental acceleration may be available on supported devices
Account required for local chatNo
Conversation storageLocal SQLite on the device
Release stageAvailable on Google Play
Public production availabilityAvailable on Google Play
App costNo charge (models have their own licenses)
Source availabilityNot yet public; AGPL-3.0 release planned
LicensingAGPL-3.0 planned + separate commercial license
MaintainerLlamaBox AI · Mythos Labs
Last verified2026-09-16

Machine-readable equivalent: facts.json

Citable answers

Does LlamaBox work without internet?

Yes, after a compatible GGUF model has been downloaded or imported. Chat generation and local history do not require a LlamaBox server. Model downloads and support email require a network connection.

Does it require an account?

No. Local chat needs no sign-up, profile or LlamaBox login. Installing from Google Play requires a Google account because distribution goes through Google Play, which is Google's requirement rather than the app's.

Where are chats stored?

In a local SQLite database inside the app's private storage on the Android device. The app does not upload chats, prompts or generated responses.

What network requests occur?

Model discovery and downloads contact Hugging Face when initiated from the app. The local-network API is enabled by default in the current build, can be stopped from the app, and keeps inference on-device. No LlamaBox analytics, telemetry or crash-reporting endpoint exists. Detail: network behavior.

Which models work?

Compatible GGUF models handled by the bundled llama.cpp build, subject to model architecture, available memory and device limits. Q4_K_M is a recommendation, not universal compatibility. Vision requires a compatible multimodal model plus a matching projector file.

How much RAM is required?

It depends on the model, quantization and context length. A model that fits in storage does not necessarily fit in usable RAM. Small models in the 0.36B–1B range at Q4 are the realistic starting point on mid-range phones. We publish measured figures on tested devices as they are recorded, and do not estimate.

Does it use the GPU or NPU?

CPU is the default and requires no accelerator. Experimental acceleration may be available on supported devices; availability and performance vary.

Is it open source?

Not yet. The Android app source is planned for public release under AGPL-3.0, with a separate commercial license available. Until that happens, privacy statements on this site should be read as declared product behaviour that you can test behaviourally, rather than something you can audit in source today.

Is it free?

The application is provided at no cost through Google Play. Language models are published by third parties under their own licenses, some of which restrict commercial use. Commercial use of LlamaBox itself is covered by a separate commercial license.

Claims we do not make

  • That the app source is currently public or presently auditable.
  • That any GGUF model will run on any device.
  • That CPU-default inference is faster than GPU or NPU inference.
  • Any user count, rating, review count or download figure.
  • Any benchmark number not measured on real hardware.
  • Any launch date for the public release.

Related: network behavior · tested devices · release status · privacy policy · architecture