Whistle speech-to-text model fits in 16.9 MB and runs on the CPU
The authors have released Whistle, a speech recognition model aimed at mobiles, wearables, robots, smart home, automotive and microcontrollers. It is one 16.9 MB file, runs on the CPU with no dependencies, and loads into the same C++ engine as Needle, using the same container and the same quantisation.
Whistle does three jobs on the device. Transcription takes 16 kHz mono audio, up to 30 seconds in one pass, in English, German, French, Spanish, Italian, Dutch and Polish; the language is detected unless you name it. Word timestamps give every word with its start, end and probability, aligned from the decoder's attention. Speech embedding returns the encoder output, one row per 80 ms frame, without decoding a transcript.
The front end frames audio at a 25 ms window and a 10 ms hop into 80 log-mel bins, band-limited to 250-3500 Hz and normalised per channel. Thirty seconds is 3,000 frames; a convolutional stem of 128 channels and kernel 9 halves the count three times, leaving 375 frames at one per 80 ms. The encoder is eight Simple Attention blocks with four mHC residual lanes and a Monarch Hadamard MLP in place of the feed-forward network, the same blocks Needle uses. Its attention is not causal, so a frame at 3 s can attend to a frame at 12 s.
The decoder has eight Laddered Simple Attention blocks at width 512, with 8 query heads to 2 KV heads, 48-dimensional queries and keys, 64-dimensional values, a 3-tap causal convolution on Q, K and V, and engram lookups at layers 3 and 7 over 18,432 slots. That is Needle's block list with a different layer count. The speech-specific addition is a gated cross attention in each decoder layer, with a gate learned per layer. Its K and V projections are computed once when the clip arrives (375 frames across 8 layers) and held for the whole decode, so five beams cost five short transcript caches rather than five passes over the audio.
Decoding uses five beams scored by length-normalised log probability. Keyword biasing walks an Aho-Corasick automaton over phrases the caller passes in and lifts their log probability as the automaton advances. Transcripts are capped at 320 tokens. The vocabulary is 8,192 text pieces plus seven language tokens, one per language, so the detected language comes out as a token. Every decoder depth from 2 layers up was trained as a separate model, and --audio-depth picks one at load time; the encoder is never sliced, so all eight blocks run at every depth. On silence, the engine measures the clip's loudness range first and, below a threshold, returns an empty transcript and an empty language without entering the beam search.
On benchmarks, the authors say Whistle is ahead of Whisper base on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average. Whisper base is ahead on TED-LIUM, AMI and the MLS average, at 145.3 MB against Whistle's 16.9 MB. Each model ran on its official runtime at its defaults: Whistle's C++ engine at 5 beams, openai-whisper, and moonshine-voice non-streaming over whole audio. Whisper pads every input to 30 seconds, so its time to first token is flat across clip lengths, while Whistle's tracks the clip: 5.9 ms at 5 seconds, 11.1 ms at 10 seconds and 36.3 ms at 30 seconds. Word error rates use the Whisper normalizers. Whistle's were measured over 86,174 utterances; Whisper's and Moonshine's are the figures their authors published, from the multilingual checkpoints rather than the English-only ones. The authors state that no test audio appears in Whistle's training or validation data, verified by comparing audio checksums and speaker IDs across every reported test set.
The needle_load call reads whichever model a .cact file holds, so the same binary can do speech, text, or both. needle_complete can take a clip directly: the engine transcribes it, answers the transcript against your tools, and returns one JSON object with the calls and the speech fields, the speech ones prefixed audio_. A 16 kHz WAV or raw samples need nothing beyond the base install; other sample rates and microphone capture need the [mic] extra, which adds soxr and sounddevice. Each call returns the text, the language, the milliseconds to first token and the decoder's tokens per second. word_timestamps=True adds per-word times and probability, keywords=[...] biases the search toward given phrases, and language="de" forces a language. needle.Whistle() exposes the same model as an object for embed(audio). Two CLI commands exist: needle whistle playground transcribes from the microphone in the terminal, and needle whistle compare runs one clip through Whistle, Whisper and Moonshine side by side with timings.
The engine ships prebuilt for seventeen targets, from macOS and Linux through Android, iOS, watchOS, Windows on ARM, RISC-V, MIPS, the browser and a WASI component. Each folder holds a needle binary, libneedle.a and needle.h. The speech C API is needle_load, needle_transcribe and needle_embed, and the engine reads no environment variables. Weights are on Hugging Face, the engine and platform folders are in Cactus-Compute/needle3, and the source is on GitHub.
Key facts
- Whistle is a speech recognition model in one 16.9 MB file that runs on the CPU with no dependencies, in the same C++ engine as Needle.
- It transcribes up to 30 seconds of 16 kHz mono audio per pass in seven languages (English, German, French, Spanish, Italian, Dutch, Polish), and also returns word timestamps and speech embeddings (one row per 80 ms frame).
- The authors say Whistle beats Whisper base on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average; Whisper base leads on TED-LIUM, AMI and the MLS average and is 145.3 MB against 16.9 MB.
- Time to first token for Whistle is 5.9 ms at 5 seconds of audio, 11.1 ms at 10 seconds and 36.3 ms at 30 seconds, while Whisper's stays flat because it pads input to 30 seconds.
- The engine ships prebuilt for seventeen targets, including Android, iOS, watchOS, RISC-V, the browser and a WASI component.
Why it matters
Speech recognition that fits in 16.9 MB and runs on the CPU with no dependencies puts transcription within reach of devices as small as microcontrollers, wearables and robots, according to the authors. The comparison point they give is Whisper base at 145.3 MB. Whistle shares its C++ engine, container, quantisation and encoder blocks with Needle, so one binary can load speech models, text models, or both, and a single call can transcribe a clip and answer the transcript against your tools.
Who it affects
The release targets developers building for mobiles, wearables, robots, smart home, automotive and microcontrollers who want speech handled on the device. Developers already using the Needle engine can load Whistle through the same needle_load call. Teams that need word-level timing or frame-level speech embeddings without a transcript get both from the same model. The seven supported languages are English, German, French, Spanish, Italian, Dutch and Polish.
How to use it
A 16 kHz WAV or raw samples work with the base install; other sample rates and microphone capture need the [mic] extra (soxr and sounddevice). Each call returns the text, the language, milliseconds to first token and decoder tokens per second. Pass word_timestamps=True for per-word start, end and probability, keywords=["Siobhan", "Krzysztof"] to raise the log probability of specific phrases, and language="de" to force a language instead of detecting it. needle.Whistle() gives the model as an object for embed(audio). In the terminal, needle whistle playground transcribes from the microphone, and needle whistle compare runs one clip through Whistle, Whisper and Moonshine side by side. Weights are on Hugging Face and the engine is in Cactus-Compute/needle3. No licence or price is stated.
How solid is it
This is a first-party release post with architectural detail, and the authors describe the test setup: each model on its official runtime at defaults, Whisper normalizers for scoring, and a leakage check comparing audio checksums and speaker IDs across every reported test set. Whistle's word error rates were measured by the authors over 86,174 utterances. Whisper's and Moonshine's rates are the figures their authors published, from the multilingual checkpoints, not numbers measured in the same run. The text gives no actual word error rate values for Whistle, Whisper base or Moonshine; only which model is ahead on which dataset. Independent reproduction is not mentioned.
Risks and caveats
Whisper base is ahead on TED-LIUM, AMI and the MLS average, so Whistle's lead is not across the board. Whistle's comparative standing against Moonshine on specific datasets is not stated. A single pass handles up to 30 seconds of audio, transcripts are capped at 320 tokens, and only seven languages are covered. The hardware on which the timings were measured is not stated, and no latency figures are given for Whisper or Moonshine in this text. The silence loudness threshold value is not given, and which of the seventeen targets were actually benchmarked is not stated.
“It is one 16.9 MB file, runs on the CPU with no dependencies, and loads into the same C++ engine as Needle, from the same container and the same quantisation.”
— Whistle release post