About this project
edge-llm-bench is a neutral, reproducible benchmark harness for local LLM engines running on real devices, designed to run continuously rather than as one-off measurement campaigns. It targets macOS, iOS and Android hardware.
Engines covered include LiteRT-LM, llama.cpp, MLX, Apple Core AI and Cactus, with every version pin recorded in environment.lock.json. Devices referenced include Mac Studio (M4 Max), iPhone 17 Pro, Pixel 8a and Galaxy S26. All results share one schema (schema/result.v1.json), one accumulation layer and one leaderboard format.
The repository ships the pipeline, every raw capture record and machine-readable regression verdicts, but deliberately does not publish cross-runtime standings. Users render their own leaderboard locally from the shipped raw data using build_summary.py and render_leaderboard.py, producing a gitignored LEADERBOARD.md.
Core commands include release-watch (compares upstream releases against pinned versions), matrix (runs a cells file of benchmark configurations) and regress (compares an engine version against a baseline). Adding a model is a single line in a cells file.
Every run emits schema-v1 JSON per record plus its raw console log under results/raw/. Derived CSVs live in results/summary/, and regression verdicts persist as JSON under results/regression-reports/. CI keeps these consistent.
Setup paths differ by platform. The Mac lane uses Homebrew, a bootstrap script for vendored engines and a CLI build. The Android lane uses prebuilt engine binaries from the releases page, avoiding bazel/NDK builds. The iPhone lane requires Xcode signing and GUI-only memory entitlements, making it the heaviest setup.
Fairness and honesty mechanisms are a central feature. Each row records its recipe, including quantization and engine pin, because a faster number under a different recipe represents a different deployment profile. Wide trial spread causes a cell to be marked UNRELIABLE rather than scored. Failed runs, crashes and OOMs remain in the table with their reasons. Cross-session deltas are normalized through session anchors. Mismatched budgets or modes refuse to score. Mac and iPhone captures are automatically quarantined and retried once when a run starts hot or shows wide spread.
Coverage is tracked as measured (stored captures exist), wired (builds at pinned version but no captures yet) or n/a with a stated reason. LiteRT-LM is measured across all four devices; llama.cpp is measured on Android devices; MLX and Core AI are Apple-only; Cactus has no Mac arm and Android support is planned. Disclosed gaps include Android LiteRT NPU being Early Access only, Android llama.cpp using the official CPU release binary, and Android v1 lacking a warm regime and TTFT measurement.
An endurance task (endurance-chat-30m) measures sustained behavior rather than single-turn speed: one engine process, one conversation with accumulating KV cache, a fixed 12-prompt script and a native 256-token per-turn cap. Each turn records decode rate, KV occupancy, memory, thermal state and degeneracy. A crash at turn 37 preserves turns 1-36 as evidence. Sessions derive four verdicts: decode decay, memory slope, degeneracy onset and completion. Lanes exist for Mac and Android (Galaxy S26), currently LiteRT-LM only.
The repository includes a seed baseline from a 2026-08-17 LiteRT-LM v0.15.0 to v0.16.0 regression run on Pixel 8a and Mac Studio M4 Max. It was carved out of apple-silicon-llm-bench, which remains the measurement archive. Licensed under MIT.
Comments
0 Rating appears after 10 ratings
Sign in to join the discussion.