Capturing data…
Capturing data…
MON, 12 OCT · 83 ITEMS
r/LocalLLaMA, Hacker News · AI & Technology
is a newly released open speech-to-text model from Cactus Compute that fits in a single 16.9 MB file (about 55M parameters, 36M active, CQ2-bit quantized). It is designed to run offline on ordinary CPUs, transcribes seven languages, and can be paired with Cactus Compute's engine to turn audio directly into tool-call requests. The practical claim is that you can do on-device speech recognition without a GPU or cloud connection, at the cost of some accuracy relative to much larger models.
The model's existence, size, and parameter count are consistently reported across multiple independent hands-on write-ups (Reddit, dev.to, Medium) and appear grounded in Cactus Compute's own release, so the core size/capability claim is solid. Quality claims are more contested: one Hacker News comment reports noticeably worse transcriptions than a much larger Qwen-ASR model on 170 messages, so 'runs on CPU' is credible but 'accurate enough' likely depends on the use case.
Genuinely new: this coverage appeared in the October 8-11, 2026 window, shortly after the model's release.