# LlamaBox — full architecture reference (for LLMs) This file mirrors /architecture.md. It is the canonical technical reference for LlamaBox, provided in markdown for generative engines and developer tools. # LlamaBox — Architecture This document explains **what LlamaBox is** and **why it is built the way it is**. It is written for developers, contributors, and curious users who want to understand the system beneath the app — not just the feature list. > Verified 16 September 2026 against LlamaBox 1.0.7. The app source is not yet public; an AGPL-3.0 source release is planned. Until then, this document describes the public architecture and its observable behavior. --- ## 1. What LlamaBox is LlamaBox is a **React Native Android application that runs large language models (LLMs) on-device**, with no LlamaBox cloud dependency for inference. You download or import a compatible GGUF model, the app loads it locally, and llama.cpp generates tokens on your phone. CPU execution is the default; experimental acceleration may be available on supported devices. The product goal is simple: **private, offline AI chat.** Chat inference and local history remain on-device. Model discovery, downloads and the local-network API are separate network surfaces described in [Network behavior](/network-behavior.html). ### Key facts | | | |---|---| | **Platform** | Android 7.0+ (arm64-v8a). iOS is a future target. | | **Framework** | React Native 0.81 with the New Architecture (Fabric + JSI/TurboModules) | | **Inference** | `llama.rn` 0.12.7 wrapping `llama.cpp` | | **Model format** | GGUF (Q4_K_M recommended) | | **Compute today** | CPU by default (4 threads); experimental acceleration on supported devices | | **Multimodal** | Vision models via `initMultimodal` — image encoder on CPU | | **State** | Zustand (UI), SQLite (chat history), AsyncStorage (settings) | | **Context window** | 2048 tokens default (512–8192 configurable); auto 4096 for vision | --- ## 2. High-level architecture ``` ┌─────────────────────────────────────────────┐ │ JavaScript Layer │ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │ │ ChatStore│ │ModelStore│ │ Settings │ │ │ │ (Zustand)│ │ (Zustand)│ │(Zustand) │ │ │ └────┬─────┘ └────┬─────┘ └────┬─────┘ │ │ │ │ │ │ │ ┌────▼─────────────▼─────────────▼──────┐ │ │ │ UI Components │ │ │ │ ChatScreen · ChatBubble · ChatInput │ │ │ │ ModelSidebar · SystemMonitor · … │ │ │ └────┬─────────────────────────────────┘ │ │ │ │ │ ┌────▼─────┐ ┌──────────┐ ┌───────────┐ │ │ │ LlmService│ │ Database │ │ Settings │ │ │ │ (llama.rn)│ │ (SQLite) │ │Persistence│ │ │ └────┬─────┘ └────┬─────┘ └────┬─────┘ │ └───────┼─────────────┼─────────────┼────────┘ │ │ │ ┌───────▼─────────────▼─────────────▼────────┐ │ React Native Bridge (JSI) │ └───────┬────────────────────────────────────┘ │ ┌───────▼────────────────────────────────────┐ │ Native Layer (C++) │ │ llama.rn module → llama.cpp → ggml (CPU) │ └─────────────────────────────────────────────┘ ``` **Three layers:** 1. **JavaScript layer** — React Native UI + Zustand stores + service classes. 2. **Bridge** — JSI/TurboModules give near-zero-overhead JS↔native calls (New Architecture). 3. **Native layer** — `llama.rn` ships `llama.cpp` + `ggml` compiled for arm64. Inference runs here. --- ## 3. The "why" — design decisions explained This is the section most people want. Each decision has a short rationale. ### Why on-device / offline? Chat inference and local history do not require a LlamaBox server. Model discovery and downloads use the network when initiated, and the local-network API is a separate path for compatible clients. See [Network behavior](/network-behavior.html) for the exact boundary. ### Why `llama.cpp` + `llama.rn`? `llama.cpp` provides the native GGUF inference engine, including optimized CPU kernels. `llama.rn` is its React Native bridge, exposing model lifecycle, completion, tokenization and multimodal APIs to the app. ### Why React Native + the New Architecture? A single TypeScript codebase runs on Android today and can target iOS later. The New Architecture provides the native boundary used for model lifecycle, completion calls and streamed token callbacks. ### Why CPU-first execution? The CPU path is the default because it provides the broadest compatibility. Experimental acceleration can be enabled on supported devices; availability varies and unsuccessful initialization falls back to CPU. ### Why is the vision encoder forced onto CPU? Multimodal sessions load a matching projector and run image encoding on CPU. Text generation then continues through the same llama.cpp completion path. ### Why Zustand + SQLite + AsyncStorage (three stores)? Each does what it is best at: - **Zustand** — reactive UI state that does not need persistence (in-memory message list, generation flags, theme). Lightweight, no boilerplate. - **SQLite (WAL mode)** — the durable chat history. Indexed by `(conversation_id, position)` for fast message loads and `updated_at` for sidebar sorting. - **AsyncStorage** — key/value settings and per-model overrides, hydrated into Zustand on startup. Mixing these would mean fighting a single store's limits. Keeping them separate is simpler and faster. ### Why memory-mapped weights (`use_mmap: true`) with no `mlock`? GGUF weights are mapped into virtual address space and paged in on demand. Pages are not pinned with `mlock`, reducing pressure on lower-memory phones at the cost of possible page-fault stalls. ### Why is the KV cache not persisted across sessions? When you resume an old conversation, retained history is fed back as a fresh prompt and llama.cpp rebuilds the KV cache. Persisted caches are not reused across model or context changes, favoring predictable state over faster resume. ### Why a singleton model context? Only one model is loaded at a time; loading a new one unloads the previous. This limits concurrent memory pressure on mobile devices. ### Why context 2048 (4096 for vision)? KV-cache memory grows with context length. Text chat defaults to 2048 tokens; multimodal sessions use at least 4096 and disable context shifting to preserve image embeddings. ### Why Q4_K_M? Q4_K_M is the recommended starting point for model size and compatibility. Actual quality, memory use and performance depend on the model and device. ### Source and licensing The Android app source is not yet public. An AGPL-3.0 source release is planned, a separate commercial license is available, and “LlamaBox” remains a reserved trademark. --- ## 4. The inference engine ### Native API surface used | JS function | Native equivalent | Purpose | |---|---|---| | `initLlama()` | `llama_load_model_from_file` | Load GGUF, create context | | `context.completion()` | `llama_decode` loop | Generate tokens | | `context.tokenize()` | `llama_tokenize` | String → token IDs | | `context.clearCache()` | `llama_kv_cache_clear` | Wipe KV cache | | `context.stopCompletion()` | abort callback | Stop generation | | `context.initMultimodal()` | clip/llava encoder init | Load vision projector (mmproj) | | `getBackendDevicesInfo()` | `ggml_backend_dev_get_info` | Query compute devices | | `releaseAllLlama()` | `llama_free` / backend free | Cleanup | ### Model loading parameters | Parameter | Default | Effect | |---|---|---| | `n_ctx` | 2048 (4096 vision) | Max tokens the model can attend to | | `n_threads` | 4 | CPU threads for matrix ops (1–8) | | `n_gpu_layers` | 0 | CPU default; experimental offload when enabled and supported | | `use_mmap` | true | Memory-map weights instead of loading to RAM | | `use_mlock` | false | Do not pin pages in RAM | | `ctx_shift` | true (false for vision) | Slide context window when full | The app surfaces the full `llama.cpp` sampling parameter set — temperature, top_p, top_k, min_p, repeat_penalty, Mirostat, DRY, XTC, typical_p, top_n_sigma, seed, and more — with user-friendly labels ("Creativity", "Memory Length", "Anti-Repetition") in Settings. ### Token generation pipeline (per output token) ``` 1. Embedding lookup (single-threaded) 2. Transformer block × N layers (4 worker threads) - LayerNorm, Q/K/V GEMM, RoPE, attention, FFN, residuals 3. Final LayerNorm + output projection (multi-threaded) 4. Sampling (temperature, top-p/k, repeat penalty) 5. KV cache append (RAM write) ``` Performance varies substantially by phone, model, quantization, context and thermal state. Verified measurements will appear on [Tested devices](/tested-devices.html). --- ## 5. Multimodal vision ``` User picks/captures image → expo-image-picker returns a temp URI → copied to persistent storage (FileSystem.Paths.document/attachments/{uuid}.jpg) → message sent with OpenAI-style content array: { role:'user', content:[ {type:'text',...}, {type:'image_url', image_url:{url:'file://...'}} ] } → native: image → CLIP encoder (CPU) → image embeddings → llama.cpp decode ``` **Image persistence is critical.** `expo-image-picker` returns temp/cache URIs that get cleaned up unpredictably; the native encoder needs a stable `file://` path at generation time. Images are copied to persistent storage before being sent. A model is detected as vision-capable when its filename matches vision keywords (`llava`, `vision`, `mmproj`, `minicpm`, `moondream`, `internvl`, `idefics`, …) **and** a matching `mmproj` file is found alongside it (`{base}.mmproj.gguf`, `mmproj-{base}.gguf`, `{base}_mmproj.gguf`, plus heuristic fallbacks). --- ## 6. State, persistence, and the chat lifecycle ### Zustand stores - **`chatStore`** — active conversation, in-memory messages, generation flags, streaming buffer, generation config, theme, presets. - **`modelStore`** — loaded model, available models, download progress. ### SQLite schema (`llamabox.db`, WAL mode) - `conversations` — id, title, model_name, model_path, system_prompt, context_size, timestamps, message_count. - `messages` — id, conversation_id (FK, cascade delete), content, is_user, is_streaming, position, created_at, `attachments` (JSON image metadata). - Indexes on `(conversation_id, position)` and `updated_at DESC`. The `attachments` column is added via an idempotent `ALTER TABLE … ADD COLUMN` wrapped in try/catch, so existing installs migrate safely without crashing. ### Sending a message (the critical path) 1. Build prompt as `[system, ...history, user]`. 2. Estimate tokens per message; prune oldest user/assistant pairs (system prompt always kept) until the prompt fits `contextSize − maxTokens − 256`. 3. Insert the user message + an empty streaming placeholder into SQLite. 4. Call `LlmService.generateChat`; tokens stream back via a throttled callback (~150 ms UI updates). 5. On completion, finalize the assistant message in SQLite; on error, append the error text. Always release the generation flag. **The KV cache is never restored** — resuming an old chat reprocesses the whole history (see *Why* above). --- ## 7. UI architecture - **Screens** (expo-router): `/` → ChatScreen, `/settings` → SettingsScreen. - **Key components:** `ChatScreen`, `ChatBubble`, `ChatInput`, `ModePill` (animated Chat↔Vision toggle), `ModelSidebar`, `SystemMonitor`, `ModelDownloader`. - **Theming:** custom `useTheme()` hook over Zustand `colorScheme` — light, dark, amoled, system. Color tokens in `src/constants/theme.ts`. - **Animations:** React Native's built-in `Animated` API. --- ## 8. Model management - **Format:** GGUF; Q4_K_M recommended; typical size 500 MB – 4 GB. - **Import:** pick a file → copy to `FileSystem.Paths.document/models/` → tracked in `modelStore`. - **Download:** stream from Hugging Face URLs by `hfRepo` + `hfFile`, with live progress. - **Vision detection:** scanner pairs `.gguf` models with sibling `mmproj` files (see §5). - **One model at a time:** loading a new model unloads the current one (RAM constraint). --- ## 9. Memory footprint Memory use is dominated by mapped model weights and the KV cache, with additional runtime, UI and image allocations. A file fitting in storage does not guarantee it will fit in usable RAM. OOM during `initLlama` is caught and surfaced with guidance to use a smaller model or context. --- ## 10. Current limitations & roadmap **Limitations today** - CPU is the default; acceleration is experimental and device-dependent. - KV cache is not persisted; resuming long chats is slower to first token. - One model loaded at a time. - Android-first; iOS not yet built. **Roadmap (tentative)** - Public source release and broader platform support. No dates are announced. --- ## 11. Licensing & trademark The Android app source is not yet public. An **AGPL-3.0 source release is planned**, and a separate commercial license is available; contact `work.aalhad@gmail.com`. "LlamaBox" is a **reserved trademark** of the project. The name may not be used to endorse or promote derived works without written permission. See `LICENSE` and `CONTRIBUTING.md` (CLA) for details. --- *Last verified: 2026-09-16 against LlamaBox 1.0.7. Public app availability and public source availability are separate.* ## 12. Public site map (marketing + SEO) | URL | Purpose | |-----|---------| | https://llamabox-ai.vercel.app/ | Landing | | https://play.google.com/store/apps/details?id=com.llamabox | Google Play install | | https://llamabox-ai.vercel.app/download.html | Download / install status | | https://llamabox-ai.vercel.app/guides.html | Guides hub | | https://llamabox-ai.vercel.app/offline-ai-android.html | Offline AI Android | | https://llamabox-ai.vercel.app/on-device-llm.html | On-device AI / on-device LLM | | https://llamabox-ai.vercel.app/private-chatgpt-alternative.html | Private ChatGPT alternative | | https://llamabox-ai.vercel.app/private-ai-android.html | Private AI for Android | | https://llamabox-ai.vercel.app/chatgpt-android-offline.html | Offline ChatGPT alternative | | https://llamabox-ai.vercel.app/gguf-android.html | GGUF on Android | | https://llamabox-ai.vercel.app/llm-download.html | LLM download | | https://llamabox-ai.vercel.app/models.html | GGUF model hub for Android | | https://llamabox-ai.vercel.app/what-is-an-llm.html | What is an LLM? | | https://llamabox-ai.vercel.app/free-llms.html | Free LLMs for local use | | https://llamabox-ai.vercel.app/how-to-run-llm-on-android.html | Tutorial | | https://llamabox-ai.vercel.app/vs-chatgpt.html | Comparison | | https://llamabox-ai.vercel.app/vs-ollama.html | Ollama Android alternative / comparison | | https://llamabox-ai.vercel.app/vs-pocketpal.html | Comparison | | https://llamabox-ai.vercel.app/vs-mlc-llm.html | Comparison | | https://llamabox-ai.vercel.app/best-local-llm-apps-android.html | Roundup | | https://llamabox-ai.vercel.app/enterprise.html | Enterprise | | https://llamabox-ai.vercel.app/commercial-license.html | Licensing | | https://llamabox-ai.vercel.app/partners.html | Partnerships | | https://llamabox-ai.vercel.app/investors.html | Investors | | https://llamabox-ai.vercel.app/press-kit.html | Press kit | | https://llamabox-ai.vercel.app/blog/ | Blog index | | https://llamabox-ai.vercel.app/blog/2026-07-30-phone-faster-than-pc.html | Blog post | | https://llamabox-ai.vercel.app/blog/2026-07-30-llamabox-vs-pocketpal.html | Blog post | | https://llamabox-ai.vercel.app/blog/2026-07-30-llamabox-vs-mlc-llm.html | Blog post | | https://llamabox-ai.vercel.app/blog/2026-07-30-best-private-ai-chat-android.html | Blog post | | https://llamabox-ai.vercel.app/blog/2026-07-30-what-is-offline-ai-chat.html | Blog post | | https://llamabox-ai.vercel.app/blog/2026-07-30-why-cpu-only.html | Blog post | | https://llamabox-ai.vercel.app/architecture.html | Architecture (HTML) | | https://llamabox-ai.vercel.app/llms.txt | Short LLM brief | *Last verified: 2026-09-16 against LlamaBox 1.0.7.* ## Visual guide library Ten visual guides with readable HTML and original three-page PDFs: https://llamabox-ai.vercel.app/resources/ - Meet LlamaBox: offline AI chat on Android: https://llamabox-ai.vercel.app/resources/app-introduction/ - Set up AI chat before you go offline: https://llamabox-ai.vercel.app/resources/offline-setup/ - Choose a GGUF model for your Android task: https://llamabox-ai.vercel.app/resources/choose-gguf-model/ - Local AI: storage and RAM are different: https://llamabox-ai.vercel.app/resources/storage-and-memory/ - Give a local model a clear system prompt: https://llamabox-ai.vercel.app/resources/system-prompts/ - Tune generation settings one change at a time: https://llamabox-ai.vercel.app/resources/generation-and-context/ - Understand local chat and network boundaries: https://llamabox-ai.vercel.app/resources/privacy-and-history/ - Use a short passage for offline study practice: https://llamabox-ai.vercel.app/resources/study-practice/ - Review a draft with local AI on Android: https://llamabox-ai.vercel.app/resources/writing-drafts/ - Your first local AI chat: a short checklist: https://llamabox-ai.vercel.app/resources/first-chat-checklist/