ENCODERWaveform16 kHz mono, up to 30 s = 480,000 samplesLog-mel80 bins, 25 ms window, 10 ms hop, 250-3500 Hz, normalised per channel → 3,000 framesConvolutional stem128 channels, kernel 9, three halvings → 375 frames, one per 80 msSimple Attention blocks × 8shared with Needleself-attention over every frame at once, 4 mHC lanes, Monarch Hadamard MLP in place of the FFNDECODERCross memoryK and V projected once per clip, 375 frames × 8 layers, shared by every beamLaddered Simple Attention blocks × 8shared with NeedleGQA 8q : 2kv, 48 qk / 64 v, 3-tap causal conv, engram at layers 3 and 7, width 512Gated cross attentionevery decoder layer reads the encoder: x ← x + σ(g) · softmax(q̂ K̂ᵀ/√d) VBeam search × 5length-normalised log prob, keyword bias by Aho-Corasick automatonTranscript8,192 text pieces + 7 language tokens, up to 320 of them, word times from the decoder’s own attention
Blocks marked shared run Needle’s code, not a copy of it. --audio-depth selects decoder layers; the encoder always runs all eight.
The front end. 16 kHz mono audio is framed at a 25 ms window and a 10 ms hop into 80 log-mel bins, band-limited to 250-3500 Hz and normalised per channel. Thirty seconds is 3,000 frames. A convolutional stem of 128 channels and kernel 9 halves that count three times, leaving 375 frames at one per 80 ms. Every stage after this runs at that rate, and embed returns one row per frame.
The encoder. Eight Simple Attention blocks: four mHC residual lanes and a Monarch Hadamard MLP in place of the feed-forward network, the same blocks Needle uses. The attention is not causal. A frame at 3 s attends to a frame at 12 s.
The decoder. Eight Laddered Simple Attention blocks at width 512, 8 query heads to 2 KV heads, 48-dimensional queries and keys, 64-dimensional values, a 3-tap causal convolution on Q, K and V, and engram lookups at layers 3 and 7 over 18,432 slots. That is Needle’s block list with a different layer count.
The speech-specific part is one addition per layer. Each decoder layer reads the encoder through a gated cross attention, x ← x + σ(g) · softmax(q̂ K̂ᵀ/√d) V, with a gate learned per layer and K and V taken from the clip. Those projections run once when the clip arrives, 375 frames across 8 layers, and are then held for the whole decode. Five beams therefore cost five short transcript caches, not five passes over the audio.
Decoding. Five beams scored by length-normalised log probability. Keyword biasing walks an Aho-Corasick automaton over the phrases you pass in, alongside the beams, and lifts their log probability as the automaton advances. The transcript is capped at 320 tokens. The vocabulary is 8,192 text pieces plus seven language tokens, one per language, so the detected language is emitted as a token rather than returned out of band.
The ladder is on the decoder. Every depth from 2 layers up was trained as a model of its own, and --audio-depth selects one at load time. The encoder is never sliced: all eight blocks run at every depth.
Silence. The engine measures the clip’s loudness range before the decoder starts. Below the threshold it returns an empty transcript and an empty language, and never enters the beam search.
WhistleWhisper baseMoonshine tiny v2
Word error rate, lower is better. A missing bar is a benchmark that model’s authors never Whistle is ahead on LibriSpeech test-clean and test-other, on SPGISpeech, on Earnings-22 and on the FLEURS average. Whisper base is ahead on TED-LIUM, on AMI and on the MLS average, at 145.3 MB against 16.9.
Each model ran on its official runtime at its defaults: Whistle’s C++ engine at 5 beams, openai-whisper, and moonshine-voice non-streaming over whole audio. Time to first token is audio in to first token. Decode is tokens divided by the wall time after it, so the encoder is not counted twice. Whisper pads every input to 30 seconds, so its time to first token is flat across clip lengths. Whistle’s tracks the clip: 5.9 ms at 5 seconds, 11.1 ms at 10, 36.3 ms at 30.
Word error rates are scored with the Whisper normalizers. Whistle’s are measured over 86,174 utterances. Whisper’s and Moonshine’s are the figures their authors
needle_load reads whichever model a .cact file holds, so the same binary does speech, text, or both:
needle --model whistle.cact --audio clip.wav
needle --model needle3.cact --tools tools.json --prompt "turn off the kitchen lights"
needle --model needle3.cact --model whistle.cact --tools tools.json --audio clip.wav
On the third line needle_complete takes the clip directly. The engine transcribes it, answers the transcript against your tools, and returns one JSON object with the calls and the speech fields, the speech ones prefixed audio_. No transcript is handled by the caller.
{"function_calls":[{"name":"set_lights","arguments":{"room":"kitchen","on":false}}],
"confidence":0.94,
"audio_text":"turn off the kitchen lights",
"audio_language":"en"}
Get started
pip install cactus-needle
import needle
print(needle.transcribe("clip.wav")["text"])
# turn off the kitchen lights
A 16 kHz WAV or raw samples need nothing beyond the base install. Other sample rates and microphone capture need the [mic] extra, which adds soxr and sounddevice.
Every call returns the text, the language, the milliseconds to the first token and the decoder’s tokens per second after it. word_timestamps=True adds each word with its times and probability. keywords=["Siobhan", "Krzysztof"] raises the log probability of those phrases during the search. language="de" forces the language instead of detecting it. needle.Whistle() is the same model as an object, for embed(audio) or to hold one tuned .cact.