This phone, this model, this result.
Real measurements on real Android hardware, published with methodology and checksums — or an empty table until they exist.
Answer-first summary. No verified device benchmarks are published yet. The data structure, schema and methodology are defined below, and rows will appear here as real measurements are recorded during device testing. We would rather show an empty table than an invented one.
Measured results
| Device | Android | RAM | Model | Quantization | Load time | Generation | Result |
|---|---|---|---|---|---|---|---|
|
No verified measurements published yet. Benchmark runs are being collected during device testing. |
|||||||
Schema
Each published row corresponds to one object in benchmarks.json:
| Field | Meaning |
|---|---|
device | Marketing name and model number |
soc | System-on-chip identifier |
androidVersion | Android release and API level |
ramGb | Total physical RAM |
model | Model name and parameter count |
quantization | GGUF quantization, e.g. Q4_K_M |
modelSha256 | Checksum of the exact file tested |
contextLength | Context window used for the run |
threads | Inference thread count |
loadTimeSec | Seconds from load start to ready |
timeToFirstTokenSec | Latency to the first generated token |
promptTokPerSec | Prompt-processing throughput |
generationTokPerSec | Generation throughput |
peakMemoryMb | Peak resident memory observed |
thermalNotes | Throttling or heat observations |
result | ok · slow · unstable · failed-oom · failed-unsupported |
appVersion | LlamaBox version under test |
testDate | ISO date of the run |
Methodology
- Install the Google Play build and record the exact app version.
- Download the model and record its SHA-256 so the file is unambiguous.
- Reboot, then let the device settle to ambient temperature before measuring.
- Load the model and record load time and peak memory.
- Run a fixed prompt set and record time to first token, prompt throughput and generation throughput.
- Note any thermal throttling, background termination or out-of-memory event.
- Record the outcome verbatim, including failures. Failed runs are the most useful rows in the table.
Why this table matters more than adjectives
"Fast on-device AI" is unfalsifiable. "This phone, this model, this app version, this load time, this token rate" is checkable, reproducible and useful when deciding whether your device can run a given model. Read it alongside the known limitations.
Caveats that apply to every row
- Results vary by CPU, memory bandwidth, thermal conditions, context length and model architecture.
- A model fitting in storage does not guarantee it will fit in usable RAM.
- Q4_K_M is a recommendation, not universal compatibility.
- Vision requires a compatible multimodal model and a matching projector.
- Android may reclaim memory and interrupt generation under pressure, independently of LlamaBox.
Contribute a measurement
Users can send a run to work.aalhad@gmail.com with the schema fields above. Submissions are published only with the contributor's agreement, and device details are published without any personal identifier. Benchmark submission is never automatic — the app does not upload measurements on its own.
Related: performance overview · model compatibility · product facts · architecture