Whistle: Speech To Text In 16.9 MB
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Before you orderOffer from Amazon

Get privacy and security gear delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Cactus Compute released Whistle, a 16.9 MB speech-recognition model designed to run on-device using the company’s C++ engine. The company reports support for seven languages and strong results on several speech benchmarks, while acknowledging that competing models perform better on some datasets and that the comparisons use different runtimes and test conditions.

Cactus Compute released Whistle on October 2, describing it as a 16.9 MB speech-recognition model that runs on a CPU without external dependencies. The company says Whistle transcribes speech in seven languages on-device, a design intended for mobile devices, wearables, robots, smart-home products, vehicles and microcontrollers where network access or larger models may be impractical.

Whistle supports English, German, French, Spanish, Italian, Dutch and Polish. According to Cactus Compute, it processes mono audio sampled at 16 kHz, in clips of up to 30 seconds, and detects the language automatically unless a user specifies it. The company’s in-browser demonstration downloads the model on the first press of the microphone button; it says the audio stays on the device. That description concerns the demo and model’s local processing, not an independent audit of data handling across every product that may use Whistle.

The release includes more than transcript text. Cactus Compute says Whistle can return word-level timestamps and probabilities, and can generate speech embeddings without decoding a transcript. Its encoder reduces the audio to one representation per 80 milliseconds. The decoder uses five-beam search, with an optional keyword-biasing feature, and the transcript is capped at 320 tokens. The company says a silence check can return an empty transcript without running beam search.

Cactus Compute’s report gives CPU test results for 10 seconds of audio on an Apple M4 Pro. It reports 11.1 milliseconds to the first token and a decoding rate of 1,319 tokens per second for Whistle, compared with 73.2 milliseconds and 266 tokens per second for Whisper base, and 22.8 milliseconds and 262 tokens per second for Moonshine tiny v2. The company specifies that these figures use each model’s official runtime at its defaults; its decode-rate calculation excludes the encoder time. The figures are vendor-reported, not an independently reproduced benchmark.

At a glance
announcementWhen: Announced October 2, 2026
The developmentCactus Compute announced Whistle, a compact speech-to-text model that it says runs locally on CPUs and supports seven languages.

Local Transcription on Smaller Devices

A 16.9 MB model that runs on a CPU could make speech recognition easier to include in products with limited storage or without a reliable connection. Local processing may also reduce the need to send recorded speech to a remote service, though actual privacy depends on how a product is implemented and what other data it collects. Cactus Compute says the Whistle demonstration keeps audio on-device; the release does not establish the practices of future third-party integrations.

The reported latency and decoding numbers point to a possible practical benefit for short, interactive voice features. But model size and speed do not settle whether Whistle is suitable for a particular device or use case. Recognition accuracy, supported languages, clip length and runtime behavior matter too, and the company’s own results show that performance varies by dataset. Developers will need to test it against their audio and hardware rather than treating the headline figures as universal guarantees.

Amazon

on-device speech recognition software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Whistle Uses Needle’s Engine

Cactus Compute says Whistle uses the same C++ engine as Needle, its existing model, and shares several model blocks with it. The company describes an eight-block audio encoder and an eight-block decoder, with gated cross-attention connecting the two. It says users can select a decoder depth when loading the model, while all eight encoder blocks still run. This shared implementation is presented as a way to use the same engine and quantisation approach across both models.

The report compares Whistle with Whisper base and Moonshine tiny v2 on size, speed and word error rate. Cactus Compute says Whistle is smaller than both models in its comparison: 16.9 MB, against 145.3 MB for Whisper base and 41.9 MB for Moonshine tiny v2. On word error rate, the company reports Whistle ahead on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average. It reports Whisper base ahead on TED-LIUM, AMI and the MLS average. Moonshine tiny v2 is English-only, and the report notes that the models do not have published results for every benchmark. Whisper’s AMI figure uses the AMI-IHM subset, which is not the same subset used for the other two models.

““one 16.9 MB file, runs on the CPU with no dependencies””

— Cactus Compute

Amazon

multilingual voice transcription device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Benchmark Limits and Open Questions

The published results are reported by the model’s developer; the release does not provide independent replication. Comparisons also involve separate official runtimes and default settings, so the speed results reflect those particular configurations rather than a controlled test of identical software conditions. The report identifies differences in benchmark coverage and AMI subsets, which limit direct comparisons across some scores.

The release does not specify performance across a range of devices beyond the stated M4 Pro CPU test, nor does it give enough detail here to establish power consumption, memory use during inference or accuracy in noisy real-world conditions. It also remains unclear how the seven-language results vary by accent and recording quality. The company’s benchmark summary indicates that accuracy depends on the dataset, with different models leading on different tests.

Amazon

portable speech-to-text app

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Tests Will Clarify Use

The next useful evidence will be independent evaluations of Whistle’s accuracy and speed, especially on lower-powered devices and recordings outside benchmark datasets. Developers considering the model can compare its word-error rates on their own audio, check how its language detection behaves, and measure its latency and resource use on intended hardware. Cactus Compute’s release does not announce a schedule for further benchmarks or a broader product rollout.

Amazon

voice recognition microcontroller

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Whistle?

Whistle is a speech-recognition model released by Cactus Compute. The company says it runs on CPUs through its C++ engine and occupies a 16.9 MB model file.

Which languages does it support?

Cactus Compute lists seven languages: English, German, French, Spanish, Italian, Dutch and Polish. It says the language is detected automatically unless the user specifies one.

Does Whistle send audio to the cloud?

The company says Whistle processes audio on-device and that audio in its browser demonstration does not leave the device. That statement describes the model and demo; privacy practices in products built with it depend on their implementation.

Is Whistle more accurate than Whisper?

Not across every test in Cactus Compute’s report. The company reports Whistle ahead on several listed datasets, while Whisper base leads on TED-LIUM, AMI and the MLS average. The report also notes gaps and differences in benchmark coverage.

How were the speed figures measured?

Cactus Compute reports results for 10 seconds of audio on an Apple M4 Pro CPU, using each model’s official runtime at its default settings. The company says its decoding-rate figure excludes encoder time; the results have not been independently verified in the supplied report.

Source: hn

COLUMBUS DAY / I

Columbus Day / Indigenous Peoples' Day Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
OpenAI Poached The Latest Fields Medal Winner: Who Is ByteDance’s Newly Launched Scientist Program Targeting? – 36 Kr

OpenAI Poached The Latest Fields Medal Winner: Who Is ByteDance’s Newly Launched Scientist Program Targeting? – 36 Kr

OpenAI reportedly recruits the latest Fields Medal recipient, while ByteDance launches a new scientist program, highlighting fierce US-China AI talent competition.
AI’s Management Gap Appears After The Right Answer

AI’s Management Gap Appears After The Right Answer

A recent experiment reveals AI models can diagnose and respond accurately but struggle to complete trustworthy, actionable work under real-world pressures.
One Markdown File, Publish-ready For Every Platform

One Markdown File, Publish-ready For Every Platform

A web tool now enables creators to convert a single markdown file into platform-ready formats, streamlining content distribution for independent creators.
13 Best Guides to AI-Powered Marketing Automation Tools for Smarter Campaigns in 2026

13 Best Guides to AI-Powered Marketing Automation Tools for Smarter Campaigns in 2026

Discover the 13 best books and guides on AI-driven marketing automation, helping marketers choose strategies and workflows for smarter campaigns.