1.What LlamaBox is
A React Native Android app that runs LLMs on-device using llama.cpp via the llama.rn bridge. You download or import a compatible GGUF model, the app memory-maps it locally, and inference runs on your phone. No inference call is sent to a LlamaBox cloud service.
- Platform: Android 7.0+ (arm64). iOS is a future target.
- Framework: React Native 0.81, New Architecture (JSI/TurboModules)
- Inference: llama.rn 0.12.7 wrapping llama.cpp
- Compute: CPU by default (4 threads); experimental acceleration may be available on supported devices
- State: Zustand (UI) · SQLite (history) · AsyncStorage (settings)
2.High-level architecture
JavaScript layer → React Native bridge (JSI) → Native (C++)
UI + Zustand llama.rn → llama.cpp → ggml
Three layers: a TypeScript UI and state layer, a JSI/native boundary carrying completion calls and token callbacks, and a native C++ inference core.
3.The decisions, explained
On-device / offline
Chat inference and local history do not require a LlamaBox server. Model discovery and downloads use the network when initiated, and the local-network API is a separate path for compatible clients. See network behavior for the exact boundary.
llama.cpp + llama.rn
llama.cpp is the most mature, mobile-optimized CPU engine for GGUF, with hand-tuned ARM NEON kernels. llama.rn is its React Native bridge. We get native-class inference without writing our own JNI.
React Native + New Architecture
One TypeScript codebase runs on Android today and can target iOS later. The New Architecture provides the native boundary used for model lifecycle, completion calls and streamed token callbacks.
CPU-first execution
The CPU path is the default because it provides the broadest compatibility. Experimental acceleration can be enabled on supported devices; availability varies and unsuccessful initialization falls back to CPU.
Vision encoder on CPU
Multimodal sessions load a matching projector and run image encoding on CPU. Text generation then continues through the same llama.cpp completion path.
Three stores (Zustand / SQLite / AsyncStorage)
Each does what it is best at: Zustand for reactive in-memory UI state, SQLite (WAL) for durable chat history, AsyncStorage for key/value settings. Keeping them separate is simpler and faster than fighting one store's limits.
Memory-mapped weights, no mlock
GGUF weights are memory-mapped and paged in on demand. Pages are not pinned in RAM, reducing pressure on lower-memory phones at the cost of possible page-fault stalls.
KV cache is not persisted
Resuming an old chat replays its retained history and rebuilds the KV cache. Persisted caches are not reused across model or context changes, favoring predictable state over faster resume.
Context 2048 (4096 for vision)
KV-cache memory grows with context length. Text chat defaults to 2048 tokens; multimodal sessions use at least 4096 and disable context shifting to preserve image embeddings.
Q4_K_M quantization
Q4_K_M is the recommended starting point for model size and compatibility. Actual quality, memory use and performance depend on the model and device.
Source and licensing
The Android app source is not yet public. An AGPL-3.0 source release is planned, a separate commercial license is available, and “LlamaBox” remains a reserved trademark.
4.Inference, in brief
Per output token: embedding lookup → transformer blocks (normalization, Q/K/V projection, RoPE, attention and feed-forward layers) → final projection → sampling → KV-cache append. Performance varies substantially by phone, model, quantization, context and thermal state. Verified measurements will appear on Tested devices.
5.Multimodal vision
Images are picked via expo-image-picker, copied to persistent storage (the picker's temp URIs are unreliable), and sent as an OpenAI-style content array with a file:// image. A model is vision-capable when its filename matches vision keywords and a sibling mmproj file is found.
6.State & persistence
Two Zustand stores (chatStore, modelStore) drive the UI. SQLite stores conversations and messages (with an attachments column for image metadata, migrated idempotently). AsyncStorage holds settings and per-model overrides. Sending a message: build the prompt, prune to fit context (system prompt always kept), insert the user message + a streaming placeholder, stream tokens back on a ~150 ms throttled UI tick, then finalize in SQLite.
7.Memory footprint
Memory use is dominated by mapped model weights and the KV cache, with additional runtime, UI and image allocations. A file fitting in storage does not guarantee it will fit in usable RAM; measured results are published only through Tested devices.
8.Limitations & roadmap
- Today: CPU is the default; acceleration is experimental; KV cache is not persisted; one model is loaded at a time; Android is the supported platform.
- Planned: Public source release and broader platform support. No dates are announced.