Cactus Compute's Whistle is a 16.9 MB on-device speech-recognition model released on 2026-10-02. It runs on the CPU in the same C++ engine as Needle, transcribes up to 30 seconds in seven languages, and can take audio into the tool-calling path in one binary. The benchmark results below are vendor-reported and should be tested on target hardware.
Key facts
- Release and authors: Cactus Compute announced Whistle on 2026-10-02; the post names Jakub Mroz and Henry Ndubuaku.
- Footprint: Cactus Compute reports one 16.9 MB
.cactfile with 55 million total parameters, 36 million active parameters, and CQ2bit quantization. - Speed: Cactus Compute's Apple M4 Pro measurements report 11.1 ms to first token on a 10-second clip and 1,319 tokens per second after decoding begins.
- Languages: Transcribes English, German, French, Spanish, Italian, Dutch, and Polish with automatic language detection.
- Audio limit: It accepts 16 kHz mono float samples or WAV files for up to 30 seconds, or 480,000 samples, in one pass; longer audio is refused.
- On-device tasks: It returns a transcription, word timestamps with start/end/probability values at 80 ms resolution, or frame-level speech embeddings.
- Shared runtime: The
needleC++ binary can load Whistle beside Needle and connect the resulting audio understanding to tool calls.
How does Cactus Whistle process audio inside the Needle engine?
Cactus Compute's design reuses the runtime and several building blocks from Needle. Audio processing runs across five stages:
- Log-mel front end: 16 kHz mono audio is sliced into 25 ms windows with a 10 ms hop into 80 log-mel bins (250–3,500 Hz). Thirty seconds produces 3,000 frames normalized per channel.
- Convolutional stem: A 1D convolutional stem (128 channels, kernel 9) halves temporal resolution three times, yielding 375 frames at one per 80 ms.
- Audio encoder: Eight non-causal Simple Attention blocks process all 375 frames at once using four multi-lane hyper-connection (mHC) residual lanes and Monarch Hadamard MLPs, matching Needle's encoder blocks.
- Decoder with gated cross-attention: Eight Laddered Simple Attention blocks (width 512, 8 query heads to 2 KV heads, 48-dim Q/K, 64-dim V, 3-tap causal convolution) query the encoder via gated cross-attention:
x ← x + σ(g) · softmax(q̂ K̂ᵀ / √d) V. Key and value projections are computed once per clip and cached across all search beams. - Beam search and keyword biasing: Five beams decode using length-normalized log probability over an 8,192 text piece vocabulary plus seven language tokens. An Aho-Corasick automaton walks the beams to boost probabilities of specified keywords.
In the launch post Whistle: Speech to Text in 16.9 MB, Cactus Compute notes that "The transcription stays inside the engine, so audio in and tool calls out is one call." When running the unified CLI, both models load into memory together:
needle --model needle3.cact --model whistle.cact --tools tools.json --audio clip.wavThe third invocation answers the transcript against the tools and returns one JSON object containing tool calls and speech fields. The example includes audio_text, so the caller can receive the text; the important boundary is that it does not have to shuttle that text between separate processes:
{
"function_calls": [
{
"name": "set_lights",
"arguments": {
"room": "kitchen",
"on": false
}
}
],
"confidence": 0.94,
"audio_text": "turn off the kitchen lights",
"audio_language": "en"
}Before beam search runs, Whistle calculates audio loudness across the clip. Audio below the silence threshold returns an empty transcript without invoking decoder search passes.
Whistle is a local deployment choice. Inside Gemini Spark: Tasks, Schedules, and Chrome Auto Browse describes a different browser split. What changed in the Codex CLI refresh of September 2026 covers voice support in a desktop coding CLI, while Whistle puts recognition in the edge runtime.
How does Cactus Whistle compare to Whisper base and Moonshine tiny v2?
Cactus Compute's benchmark page reports the following vendor measurements on an Apple M4 Pro CPU. Each model used its official runtime and defaults; the latency and decode rows use 10 seconds of audio.
Metric | Whistle | Whisper base | Moonshine tiny v2 |
|---|---|---|---|
Model size | 16.9 MB | 145.3 MB | 41.9 MB |
Time to first token | 11.1 ms | 73.2 ms | 22.8 ms |
Decode speed | 1,319 tokens/s | 266 tokens/s | 262 tokens/s |
Whisper base pads every input to 30 seconds, so its reported first-token time stays flat across clip lengths. Whistle's reported time tracks clip duration: 5.9 ms at 5 seconds, 11.1 ms at 10 seconds, and 36.3 ms at 30 seconds.
The same rendered chart reports these vendor WER values. A dash means the chart says that model's authors did not publish a value; Whisper's AMI number is AMI-IHM, a different subset from the AMI result shown for the other models.
Dataset | Whistle WER | Whisper base WER | Moonshine tiny v2 WER |
|---|---|---|---|
LibriSpeech test-clean | 4.31% | 4.90% | 4.52% |
LibriSpeech test-other | 10.49% | 11.00% | 11.71% |
SPGISpeech | 7.65% | — | 7.70% |
Earnings-22 | 19.01% | — | 21.25% |
AMI | 26.07% | 21.50% (AMI-IHM) | 22.77% |
TED-LIUM | 7.61% | 5.00% | 5.64% |
FLEURS average | 21.40% | 24.50% | — |
Multilingual LibriSpeech (MLS) | 24.90% | 23.10% | — |
The chart shows Whistle ahead of Whisper base on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22, and FLEURS. Whisper base is ahead on TED-LIUM, AMI, and MLS. These are vendor-reported comparisons, not an independent reproduction.
Cactus Compute's guide also reports keyword-biasing results on 2,611 LibriSpeech test-clean utterances. With 100 keywords, biased word error rate (B-WER) falls from 18.43% to 4.46%, while overall WER falls from 4.73% to 2.93%. The feature is useful for names, rooms, devices, stations, product codes, and other terms that a general vocabulary may miss. The same decoder ladder exposes trained depths from 2 to 8 through --audio-depth, so a deployment can trade accuracy against decode speed without swapping model files.
What hardware targets and licenses support Cactus Whistle?
The weights and runtime source are open under Apache-2.0.
The model files are available in `Cactus-Compute/whistle`; the C++ engine and bindings are in `cactus-compute/needle`.
Cactus Compute says its precompiled engine supports 17 targets, including macOS, Linux, Android, iOS, watchOS, Windows on ARM, RISC-V, MIPS, WebAssembly, and WASI.
Python integration installs via PyPI:
pip install cactus-needleMicrophone recording and audio resampling require pip install cactus-needle[mic], adding soxr and sounddevice:
import needle
result = needle.transcribe("clip.wav")
print(result["text"])RuntimeWire reported on 2026-10-03 that the seven-language scope and 30-second single-pass cap are deployment trade-offs that warrant testing on the target acoustic environment. The r/LocalLLaMA launch thread repeats the headline benchmark claims; its comments ask about noise, French accuracy, longer clips, and comparisons with other production ASR models. Those comments are reactions, not independent benchmark results.
Sources
- Whistle: Speech to Text in 16.9 MB (release, architecture, runtime, and rendered benchmark chart; read 2026-10-06)
- Getting the Most out of Whistle (input limits, keyword biasing, and decoder ladder; read 2026-10-06)
- cactus-compute/needle on GitHub (runtime, platform targets, install path, and license; read 2026-10-06)
- Cactus-Compute/whistle on Hugging Face (model files, languages, and Apache-2.0 metadata; read 2026-10-06)
- Cactus Compute releases a 16.9MB speech model for local CPUs (independent coverage; read 2026-10-06)
- Whistle: speech to text in a 16.9MB file on r/LocalLLaMA (launch-thread claims and reactions; read 2026-10-06)
- Inside Gemini Spark: Tasks, Schedules, and Chrome Auto Browse (related post; read 2026-10-06)
- What changed in the Codex CLI refresh of September 2026 (related post; read 2026-10-06)
Last verified: 2026-10-06.