⚡️ Tiny model, real reasoning. 6,616 tok/sec.
Greg Diamos published a paper on an outrageously small reasoning model trained almost entirely on synthetic data curated by large models. Training loss never stopped falling across a 4.91B-token run.
His read: this is basically distillation. And it works because that kind of corpus didn't exist last time anyone took small models seriously.

huggingface.co
paper.pdf · gdiamos/amx-reasoning-v1-instruct at main
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1September 8, 2026 250 1