# LlamaBox — Architecture

This document explains **what LlamaBox is** and **why it is built the way it is**. It is written for developers, contributors, and curious users who want to understand the system beneath the app — not just the feature list.

> Verified 16 September 2026 against LlamaBox 1.0.7. The app source is not yet public; an AGPL-3.0 source release is planned. Until then, this document describes the public architecture and its observable behavior.

---

## 1. What LlamaBox is

LlamaBox is a **React Native Android application that runs large language models (LLMs) on-device**, with no LlamaBox cloud dependency for inference. You download or import a compatible GGUF model, the app loads it locally, and llama.cpp generates tokens on your phone. CPU execution is the default; experimental acceleration may be available on supported devices.

The product goal is simple: **private, offline AI chat.** Chat inference and local history remain on-device. Model discovery, downloads and the local-network API are separate network surfaces described in [Network behavior](/network-behavior.html).

### Key facts

| | |
|---|---|
| **Platform** | Android 7.0+ (arm64-v8a). iOS is a future target. |
| **Framework** | React Native 0.81 with the New Architecture (Fabric + JSI/TurboModules) |
| **Inference** | `llama.rn` 0.12.7 wrapping `llama.cpp` |
| **Model format** | GGUF (Q4_K_M recommended) |
| **Compute today** | CPU by default (4 threads); experimental acceleration on supported devices |
| **Multimodal** | Vision models via `initMultimodal` — image encoder on CPU |
| **State** | Zustand (UI), SQLite (chat history), AsyncStorage (settings) |
| **Context window** | 2048 tokens default (512–8192 configurable); auto 4096 for vision |

---

## 2. High-level architecture

```
┌─────────────────────────────────────────────┐
│              JavaScript Layer                │
│  ┌──────────┐  ┌──────────┐  ┌──────────┐   │
│  │ ChatStore│  │ModelStore│  │ Settings │   │
│  │ (Zustand)│  │ (Zustand)│  │(Zustand) │   │
│  └────┬─────┘  └────┬─────┘  └────┬─────┘   │
│       │             │             │          │
│  ┌────▼─────────────▼─────────────▼──────┐  │
│  │           UI Components               │  │
│  │  ChatScreen · ChatBubble · ChatInput  │  │
│  │  ModelSidebar · SystemMonitor · …      │  │
│  └────┬─────────────────────────────────┘  │
│       │                                      │
│  ┌────▼─────┐  ┌──────────┐  ┌───────────┐  │
│  │ LlmService│  │ Database │  │ Settings  │  │
│  │ (llama.rn)│  │ (SQLite) │  │Persistence│  │
│  └────┬─────┘  └────┬─────┘  └────┬─────┘  │
└───────┼─────────────┼─────────────┼────────┘
        │             │             │
┌───────▼─────────────▼─────────────▼────────┐
│        React Native Bridge (JSI)            │
└───────┬────────────────────────────────────┘
        │
┌───────▼────────────────────────────────────┐
│              Native Layer (C++)             │
│   llama.rn module → llama.cpp → ggml (CPU)  │
└─────────────────────────────────────────────┘
```

**Three layers:**
1. **JavaScript layer** — React Native UI + Zustand stores + service classes.
2. **Bridge** — JSI/TurboModules give near-zero-overhead JS↔native calls (New Architecture).
3. **Native layer** — `llama.rn` ships `llama.cpp` + `ggml` compiled for arm64. Inference runs here.

---

## 3. The "why" — design decisions explained

This is the section most people want. Each decision has a short rationale.

### Why on-device / offline?
Chat inference and local history do not require a LlamaBox server. Model discovery and downloads use the network when initiated, and the local-network API is a separate path for compatible clients. See [Network behavior](/network-behavior.html) for the exact boundary.

### Why `llama.cpp` + `llama.rn`?
`llama.cpp` provides the native GGUF inference engine, including optimized CPU kernels. `llama.rn` is its React Native bridge, exposing model lifecycle, completion, tokenization and multimodal APIs to the app.

### Why React Native + the New Architecture?
A single TypeScript codebase runs on Android today and can target iOS later. The New Architecture provides the native boundary used for model lifecycle, completion calls and streamed token callbacks.

### Why CPU-first execution?
The CPU path is the default because it provides the broadest compatibility. Experimental acceleration can be enabled on supported devices; availability varies and unsuccessful initialization falls back to CPU.

### Why is the vision encoder forced onto CPU?
Multimodal sessions load a matching projector and run image encoding on CPU. Text generation then continues through the same llama.cpp completion path.

### Why Zustand + SQLite + AsyncStorage (three stores)?
Each does what it is best at:
- **Zustand** — reactive UI state that does not need persistence (in-memory message list, generation flags, theme). Lightweight, no boilerplate.
- **SQLite (WAL mode)** — the durable chat history. Indexed by `(conversation_id, position)` for fast message loads and `updated_at` for sidebar sorting.
- **AsyncStorage** — key/value settings and per-model overrides, hydrated into Zustand on startup.

Mixing these would mean fighting a single store's limits. Keeping them separate is simpler and faster.

### Why memory-mapped weights (`use_mmap: true`) with no `mlock`?
GGUF weights are mapped into virtual address space and paged in on demand. Pages are not pinned with `mlock`, reducing pressure on lower-memory phones at the cost of possible page-fault stalls.

### Why is the KV cache not persisted across sessions?
When you resume an old conversation, retained history is fed back as a fresh prompt and llama.cpp rebuilds the KV cache. Persisted caches are not reused across model or context changes, favoring predictable state over faster resume.

### Why a singleton model context?
Only one model is loaded at a time; loading a new one unloads the previous. This limits concurrent memory pressure on mobile devices.

### Why context 2048 (4096 for vision)?
KV-cache memory grows with context length. Text chat defaults to 2048 tokens; multimodal sessions use at least 4096 and disable context shifting to preserve image embeddings.

### Why Q4_K_M?
Q4_K_M is the recommended starting point for model size and compatibility. Actual quality, memory use and performance depend on the model and device.

### Source and licensing
The Android app source is not yet public. An AGPL-3.0 source release is planned, a separate commercial license is available, and “LlamaBox” remains a reserved trademark.

---

## 4. The inference engine

### Native API surface used

| JS function | Native equivalent | Purpose |
|---|---|---|
| `initLlama()` | `llama_load_model_from_file` | Load GGUF, create context |
| `context.completion()` | `llama_decode` loop | Generate tokens |
| `context.tokenize()` | `llama_tokenize` | String → token IDs |
| `context.clearCache()` | `llama_kv_cache_clear` | Wipe KV cache |
| `context.stopCompletion()` | abort callback | Stop generation |
| `context.initMultimodal()` | clip/llava encoder init | Load vision projector (mmproj) |
| `getBackendDevicesInfo()` | `ggml_backend_dev_get_info` | Query compute devices |
| `releaseAllLlama()` | `llama_free` / backend free | Cleanup |

### Model loading parameters

| Parameter | Default | Effect |
|---|---|---|
| `n_ctx` | 2048 (4096 vision) | Max tokens the model can attend to |
| `n_threads` | 4 | CPU threads for matrix ops (1–8) |
| `n_gpu_layers` | 0 | CPU default; experimental offload when enabled and supported |
| `use_mmap` | true | Memory-map weights instead of loading to RAM |
| `use_mlock` | false | Do not pin pages in RAM |
| `ctx_shift` | true (false for vision) | Slide context window when full |

The app surfaces the full `llama.cpp` sampling parameter set — temperature, top_p, top_k, min_p, repeat_penalty, Mirostat, DRY, XTC, typical_p, top_n_sigma, seed, and more — with user-friendly labels ("Creativity", "Memory Length", "Anti-Repetition") in Settings.

### Token generation pipeline (per output token)

```
1. Embedding lookup            (single-threaded)
2. Transformer block × N layers (4 worker threads)
     - LayerNorm, Q/K/V GEMM, RoPE, attention, FFN, residuals
3. Final LayerNorm + output projection (multi-threaded)
4. Sampling                     (temperature, top-p/k, repeat penalty)
5. KV cache append              (RAM write)
```

Performance varies substantially by phone, model, quantization, context and thermal state. Verified measurements will appear on [Tested devices](/tested-devices.html).

---

## 5. Multimodal vision

```
User picks/captures image
   → expo-image-picker returns a temp URI
   → copied to persistent storage (FileSystem.Paths.document/attachments/{uuid}.jpg)
   → message sent with OpenAI-style content array:
       { role:'user', content:[ {type:'text',...}, {type:'image_url', image_url:{url:'file://...'}} ] }
   → native: image → CLIP encoder (CPU) → image embeddings → llama.cpp decode
```

**Image persistence is critical.** `expo-image-picker` returns temp/cache URIs that get cleaned up unpredictably; the native encoder needs a stable `file://` path at generation time. Images are copied to persistent storage before being sent.

A model is detected as vision-capable when its filename matches vision keywords (`llava`, `vision`, `mmproj`, `minicpm`, `moondream`, `internvl`, `idefics`, …) **and** a matching `mmproj` file is found alongside it (`{base}.mmproj.gguf`, `mmproj-{base}.gguf`, `{base}_mmproj.gguf`, plus heuristic fallbacks).

---

## 6. State, persistence, and the chat lifecycle

### Zustand stores
- **`chatStore`** — active conversation, in-memory messages, generation flags, streaming buffer, generation config, theme, presets.
- **`modelStore`** — loaded model, available models, download progress.

### SQLite schema (`llamabox.db`, WAL mode)
- `conversations` — id, title, model_name, model_path, system_prompt, context_size, timestamps, message_count.
- `messages` — id, conversation_id (FK, cascade delete), content, is_user, is_streaming, position, created_at, `attachments` (JSON image metadata).
- Indexes on `(conversation_id, position)` and `updated_at DESC`.

The `attachments` column is added via an idempotent `ALTER TABLE … ADD COLUMN` wrapped in try/catch, so existing installs migrate safely without crashing.

### Sending a message (the critical path)
1. Build prompt as `[system, ...history, user]`.
2. Estimate tokens per message; prune oldest user/assistant pairs (system prompt always kept) until the prompt fits `contextSize − maxTokens − 256`.
3. Insert the user message + an empty streaming placeholder into SQLite.
4. Call `LlmService.generateChat`; tokens stream back via a throttled callback (~150 ms UI updates).
5. On completion, finalize the assistant message in SQLite; on error, append the error text. Always release the generation flag.

**The KV cache is never restored** — resuming an old chat reprocesses the whole history (see *Why* above).

---

## 7. UI architecture

- **Screens** (expo-router): `/` → ChatScreen, `/settings` → SettingsScreen.
- **Key components:** `ChatScreen`, `ChatBubble`, `ChatInput`, `ModePill` (animated Chat↔Vision toggle), `ModelSidebar`, `SystemMonitor`, `ModelDownloader`.
- **Theming:** custom `useTheme()` hook over Zustand `colorScheme` — light, dark, amoled, system. Color tokens in `src/constants/theme.ts`.
- **Animations:** React Native's built-in `Animated` API.

---

## 8. Model management

- **Format:** GGUF; Q4_K_M recommended; typical size 500 MB – 4 GB.
- **Import:** pick a file → copy to `FileSystem.Paths.document/models/` → tracked in `modelStore`.
- **Download:** stream from Hugging Face URLs by `hfRepo` + `hfFile`, with live progress.
- **Vision detection:** scanner pairs `.gguf` models with sibling `mmproj` files (see §5).
- **One model at a time:** loading a new model unloads the current one (RAM constraint).

---

## 9. Memory footprint

Memory use is dominated by mapped model weights and the KV cache, with additional runtime, UI and image allocations. A file fitting in storage does not guarantee it will fit in usable RAM. OOM during `initLlama` is caught and surfaced with guidance to use a smaller model or context.

---

## 10. Current limitations & roadmap

**Limitations today**
- CPU is the default; acceleration is experimental and device-dependent.
- KV cache is not persisted; resuming long chats is slower to first token.
- One model loaded at a time.
- Android-first; iOS not yet built.

**Roadmap (tentative)**
- Public source release and broader platform support. No dates are announced.

---

## 11. Licensing & trademark

The Android app source is not yet public. An **AGPL-3.0 source release is planned**, and a separate commercial license is available; contact `work.aalhad@gmail.com`.

"LlamaBox" is a **reserved trademark** of the project. The name may not be used to endorse or promote derived works without written permission. See `LICENSE` and `CONTRIBUTING.md` (CLA) for details.

---

*Last verified: 2026-09-16 against LlamaBox 1.0.7. Public app availability and public source availability are separate.*
