Tested devices

This phone, this model, this result.

Real measurements on real Android hardware, published with methodology and checksums — or an empty table until they exist.

Four phones are paired with differently sized model objects in a neutral compatibility matrix.

Answer-first summary. No verified device benchmarks are published yet. The data structure, schema and methodology are defined below, and rows will appear here as real measurements are recorded during device testing. We would rather show an empty table than an invented one.

Measured results

Verified LlamaBox device and model benchmark measurements
Device Android RAM Model Quantization Load time Generation Result
No verified measurements published yet.
Benchmark runs are being collected during device testing.

Schema

Each published row corresponds to one object in benchmarks.json:

FieldMeaning
deviceMarketing name and model number
socSystem-on-chip identifier
androidVersionAndroid release and API level
ramGbTotal physical RAM
modelModel name and parameter count
quantizationGGUF quantization, e.g. Q4_K_M
modelSha256Checksum of the exact file tested
contextLengthContext window used for the run
threadsInference thread count
loadTimeSecSeconds from load start to ready
timeToFirstTokenSecLatency to the first generated token
promptTokPerSecPrompt-processing throughput
generationTokPerSecGeneration throughput
peakMemoryMbPeak resident memory observed
thermalNotesThrottling or heat observations
resultok · slow · unstable · failed-oom · failed-unsupported
appVersionLlamaBox version under test
testDateISO date of the run

Methodology

  1. Install the Google Play build and record the exact app version.
  2. Download the model and record its SHA-256 so the file is unambiguous.
  3. Reboot, then let the device settle to ambient temperature before measuring.
  4. Load the model and record load time and peak memory.
  5. Run a fixed prompt set and record time to first token, prompt throughput and generation throughput.
  6. Note any thermal throttling, background termination or out-of-memory event.
  7. Record the outcome verbatim, including failures. Failed runs are the most useful rows in the table.

Why this table matters more than adjectives

"Fast on-device AI" is unfalsifiable. "This phone, this model, this app version, this load time, this token rate" is checkable, reproducible and useful when deciding whether your device can run a given model. Read it alongside the known limitations.

Caveats that apply to every row

  • Results vary by CPU, memory bandwidth, thermal conditions, context length and model architecture.
  • A model fitting in storage does not guarantee it will fit in usable RAM.
  • Q4_K_M is a recommendation, not universal compatibility.
  • Vision requires a compatible multimodal model and a matching projector.
  • Android may reclaim memory and interrupt generation under pressure, independently of LlamaBox.

Contribute a measurement

Users can send a run to work.aalhad@gmail.com with the schema fields above. Submissions are published only with the contributor's agreement, and device details are published without any personal identifier. Benchmark submission is never automatic — the app does not upload measurements on its own.

Related: performance overview · model compatibility · product facts · architecture