Every benchmark on this site, including the ones we lose. Each links to the artifact the number was read from, and each says what was not measured.
Atome LM EDGE vs TensorFlow Lite for Microcontrollers
The most carefully run comparison here: UCI HAR, 2,947 held-out windows from 30 subjects, the same 6,567-parameter network through both engines, 5 seeds, leak-free subject-wise split.
Language model: routed-ternary vs a plain float transformer
TinyStories, 3,000 steps, single seed. A direction rather than a result, and labelled that way.
| Regime | Atome LM | Vanilla FP32 | Verdict |
|---|---|---|---|
| 60K, parameter-fair | 6.31 ppl | 8.12 ppl | Atome −22% |
| 60K, flash-fair | 6.31 ppl | 13.10 ppl | Atome −52% |
| 944K, parameter-fair | 2.87 ppl | 2.54 ppl | Atome +11% — loses |
| 944K, on disk | 271 KB | 3.7 MB | Atome 20× smaller |
Single seed, multi-seed run pending. The reversal at 944K is the honest headline: this architecture is a bet on the sub-1M regime and it loses outside it.
Verification, not comparison
| Check | Result |
|---|---|
| Max |Δ|, Python → C99 → Cortex-M3 | 3.7e-7 |
| Multi-token parity | 48/48 at 60K, 16/16 at 944K |
| C99 engine on all held-out EDGE windows | 0.9206, 99.19% agreement with float |
| Ternary-32 C engine vs reference | 2947/2947 identical |
Against the wider field
| Project | Smallest | Weights | Runs on |
|---|---|---|---|
| BitNet b1.58 | 700M – 3B | ternary | server / phone |
| llama2.c (Stories260K) | 260K – 110M | FP32 / Q8 | MCU class possible |
| TinyMaix, esp32-llm | ~15M | Q8 / FP32 | ESP32-S3 with PSRAM |
| Atome LM | 60K · 944K | ternary, zero-heap | Cortex-M3 under QEMU, bit-exact; physical ESP32-WROOM-32 |
The claim is narrower than "best tiny LM" and more checkable: ternary weights, a zero-heap pure-C99 engine, and bit-exact parity verified under emulation. Each of those is falsifiable from the repository.
What is not benchmarked
- Energy. Never measured on hardware, on any board.
- X-CUBE-AI itself. Never run — CMSIS-NN, the library it generates, is the measured stand-in.
- Cortex-M33. The build faults under both emulators used here.
- A second dataset for the EDGE line. Everything is UCI HAR.