Blog · 2026-07-28
Why we keep the vision encoder on CPU
Engineering note: multimodal image encoding stays on CPU in LlamaBox so the UI remains responsive on every Android device.
Multimodal models tempt many mobile AI apps to reach for GPU acceleration. LlamaBox is CPU by default; experimental acceleration may be available on supported devices.
Today LlamaBox runs text generation on CPU threads and keeps the vision encoder on CPU so the interface stays responsive under memory pressure. A frozen UI is a product bug, not a benchmark flex.
CPU-first scope
Keeping inference on CPU removes driver fragmentation and lets LlamaBox target the widest range of Android devices. Read architecture and llms.txt.
Available on Google Play
Try private offline AI on Android.
Install from Google Play, download a compatible model, then chat offline.