How LlamaBox works — and why

LlamaBox runs chat inference on your Android phone without a cloud inference service or LlamaBox account. Here is the architecture and its network boundary.

The complete LlamaBox on-device inference path and streamed-token return loop.

LlamaBox on-device inference architecture

Expanded LlamaBox architecture diagram
Verified 16 September 2026 against LlamaBox 1.0.7. The Android app source is not yet public; an AGPL-3.0 source release is planned.

1.What LlamaBox is

A React Native Android app that runs LLMs on-device using llama.cpp via the llama.rn bridge. You download or import a compatible GGUF model, the app memory-maps it locally, and inference runs on your phone. No inference call is sent to a LlamaBox cloud service.

2.High-level architecture

JavaScript layer   →   React Native bridge (JSI)   →   Native (C++)
  UI + Zustand                                    llama.rn → llama.cpp → ggml

Three layers: a TypeScript UI and state layer, a JSI/native boundary carrying completion calls and token callbacks, and a native C++ inference core.

3.The decisions, explained

On-device / offline

Chat inference and local history do not require a LlamaBox server. Model discovery and downloads use the network when initiated, and the local-network API is a separate path for compatible clients. See network behavior for the exact boundary.

llama.cpp + llama.rn

llama.cpp is the most mature, mobile-optimized CPU engine for GGUF, with hand-tuned ARM NEON kernels. llama.rn is its React Native bridge. We get native-class inference without writing our own JNI.

React Native + New Architecture

One TypeScript codebase runs on Android today and can target iOS later. The New Architecture provides the native boundary used for model lifecycle, completion calls and streamed token callbacks.

CPU-first execution

The CPU path is the default because it provides the broadest compatibility. Experimental acceleration can be enabled on supported devices; availability varies and unsuccessful initialization falls back to CPU.

Vision encoder on CPU

Multimodal sessions load a matching projector and run image encoding on CPU. Text generation then continues through the same llama.cpp completion path.

Three stores (Zustand / SQLite / AsyncStorage)

Each does what it is best at: Zustand for reactive in-memory UI state, SQLite (WAL) for durable chat history, AsyncStorage for key/value settings. Keeping them separate is simpler and faster than fighting one store's limits.

Memory-mapped weights, no mlock

GGUF weights are memory-mapped and paged in on demand. Pages are not pinned in RAM, reducing pressure on lower-memory phones at the cost of possible page-fault stalls.

KV cache is not persisted

Resuming an old chat replays its retained history and rebuilds the KV cache. Persisted caches are not reused across model or context changes, favoring predictable state over faster resume.

Context 2048 (4096 for vision)

KV-cache memory grows with context length. Text chat defaults to 2048 tokens; multimodal sessions use at least 4096 and disable context shifting to preserve image embeddings.

Q4_K_M quantization

Q4_K_M is the recommended starting point for model size and compatibility. Actual quality, memory use and performance depend on the model and device.

Source and licensing

The Android app source is not yet public. An AGPL-3.0 source release is planned, a separate commercial license is available, and “LlamaBox” remains a reserved trademark.

4.Inference, in brief

Per output token: embedding lookup → transformer blocks (normalization, Q/K/V projection, RoPE, attention and feed-forward layers) → final projection → sampling → KV-cache append. Performance varies substantially by phone, model, quantization, context and thermal state. Verified measurements will appear on Tested devices.

5.Multimodal vision

Images are picked via expo-image-picker, copied to persistent storage (the picker's temp URIs are unreliable), and sent as an OpenAI-style content array with a file:// image. A model is vision-capable when its filename matches vision keywords and a sibling mmproj file is found.

6.State & persistence

Two Zustand stores (chatStore, modelStore) drive the UI. SQLite stores conversations and messages (with an attachments column for image metadata, migrated idempotently). AsyncStorage holds settings and per-model overrides. Sending a message: build the prompt, prune to fit context (system prompt always kept), insert the user message + a streaming placeholder, stream tokens back on a ~150 ms throttled UI tick, then finalize in SQLite.

7.Memory footprint

Memory use is dominated by mapped model weights and the KV cache, with additional runtime, UI and image allocations. A file fitting in storage does not guarantee it will fit in usable RAM; measured results are published only through Tested devices.

8.Limitations & roadmap

This page is the human-readable reference. A matching Markdown version is available for developer tools and generative engines.
Read the markdown version →