Whistle: Speech To Text In 16.9 MB
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Cactus Compute released Whistle, a 16.9 MB speech-recognition model designed to run on a device CPU without dependencies. The company reports support for seven languages and benchmark advantages on several datasets, but results vary by test and the comparisons use each model’s official runtime and defaults.

Cactus Compute announced Whistle on October 2, describing it as a 16.9 MB speech-recognition model that runs on a device’s CPU without external dependencies. The company says the model transcribes speech in seven languages and is built for uses such as phones, wearables, robots, smart-home devices and cars, where a smaller local model could reduce reliance on cloud processing.

Whistle handles 16 kHz mono audio clips up to 30 seconds in English, German, French, Spanish, Italian, Dutch and Polish. Cactus says it detects the language automatically unless the user specifies one. The model also provides word-level timestamps and probabilities, and can return speech embeddings without generating a transcript. Its browser demo downloads the model on first use; the company says audio remains on the device.

The model runs in the same C++ engine and container as Cactus Compute’s Needle model, using the same quantisation, according to the company. Whistle’s encoder and decoder each contain eight attention blocks, with an audio-specific cross-attention mechanism connecting them. Users can select decoder depth at load time, although the full eight-block encoder still runs at every selected depth. For decoding, the system uses five beams and can apply keyword biasing to phrases supplied by the user.

Cactus reports a 11.1-millisecond time to first token for ten seconds of audio on an Apple M4 Pro CPU. In the same test, the company reports 1,319 decoded tokens per second, compared with 266 for Whisper base and 262 for Moonshine tiny v2. These are vendor-reported measurements from each model’s official runtime at its defaults; the company says Whistle used five beams, and that decode speed excludes the encoder stage after measuring time to first token.

At a glance
announcementWhen: Announced October 2, 2026
The developmentCactus Compute announced Whistle, a compact speech-to-text model that runs locally on CPU and shares an engine with its Needle model.

Local Transcription on Smaller Devices

A 16.9 MB model could make speech recognition practical on devices with limited storage or compute, including wearables and embedded products. Running transcription locally may also help applications avoid sending audio to a remote service, though the privacy benefit depends on how the surrounding software handles recordings and results. Cactus says the browser demo keeps audio on-device, but that statement describes its demo and does not establish how every future integration will operate.

The reported speed and size figures are relevant to developers weighing local recognition against larger models. Cactus lists Whisper base at 145.3 MB and Moonshine tiny v2 at 41.9 MB, while reporting Whistle’s lower time to first token in its test. Those measurements are not a full assessment of real-world battery use, accuracy across accents, or performance on different processors. The company’s benchmark chart shows Whistle ahead on some speech datasets, but not all, so the results do not establish a universal accuracy lead.

Whistle’s ability to provide timestamps and speech embeddings also points to uses beyond displaying a transcript, such as aligning spoken words with media or supplying audio representations to another on-device system. Cactus says Whistle can load beside Needle, allowing a single binary to pass from audio transcription toward tool calls. The practical value of that setup will depend on integration details and the needs of each application.

Amazon

on-device speech to text app

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Whistle Fits With Needle

Cactus presents Whistle as part of a family of models sharing a common CPU engine. The release description says the blocks marked as shared use Needle’s code rather than a separate copy. The speech model adds an encoder for audio and gated cross-attention in the decoder, which lets transcript generation use encoded audio features. The company says audio projections are calculated once per clip and reused during beam search.

The release compares Whistle with Whisper base and Moonshine tiny v2 on model size, speed and word error rate. Cactus notes that coverage differs across the accuracy tests: Moonshine is English-only, while Whisper has no published results in the comparison for SPGISpeech, Earnings-22 or AMI cleaned. It also says Whisper’s AMI result uses AMI-IHM, a different subset from the AMI data used for the other models. On the reported datasets, Whistle leads on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average; Whisper base leads on TED-LIUM, AMI and the MLS average.

“It is one 16.9 MB file, runs on the CPU with no dependencies, and loads into the same C++ engine as Needle.”

— Cactus Compute, in its release

Amazon

small speech recognition model for mobile

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Benchmark Coverage and Device Results

The published material does not specify how Whistle performs across a broad range of phones, wearables or embedded processors. Its headline latency and token-throughput figures come from a ten-second audio test on an Apple M4 Pro CPU; results on lower-powered chips, longer sessions and real-world background noise are not included in the supplied figures. Cactus also does not provide power-consumption measurements in the report.

The reported accuracy comparisons have limits. Some models lack results for particular datasets, and the AMI figures do not all use the same subset. The release supplies comparative word error rates but does not establish that Whistle is more accurate across every language or use case. Independent replication of the performance figures, as well as detailed results across accents and noisy environments, is not included in the source material.

It is also not clear from the release what licensing terms apply, how developers can obtain the model outside the demo, or whether the stated 16.9 MB file size covers every component needed for deployment. Cactus describes the model as open, but the provided report does not spell out those distribution details.

Amazon

offline speech transcription device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Developer Access and Broader Testing

The next practical test is whether developers can run Whistle reliably on the range of devices named by Cactus, rather than only on the benchmarked Apple CPU. Broader evaluations could clarify accuracy, latency and power use across different processors, languages, accents and recording conditions. The release does not name a date for additional results or a follow-up product milestone.

Developers and prospective users will also need clear information on model access, licensing and deployment requirements before adopting Whistle in commercial or open-source projects. Until those details and independent tests are available, the October 2 announcement establishes the model’s stated design and vendor-reported results, but leaves its wider performance and adoption uncertain.

Amazon

multi-language speech to text software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Whistle?

Whistle is a speech-recognition model released by Cactus Compute. The company says it runs on a CPU and is packaged as a 16.9 MB file.

Which languages does it support?

Cactus lists English, German, French, Spanish, Italian, Dutch and Polish. The model detects the language unless the user names it.

Does Whistle send audio to the cloud?

Cactus says audio in its browser demo stays on the device. The release does not establish how audio would be handled in every third-party integration.

Is Whistle faster or more accurate than Whisper?

Cactus reports faster time to first token and higher decode throughput than Whisper base in its ten-second Apple M4 Pro test. Its accuracy comparison is mixed: Whistle leads on some listed datasets, while Whisper base leads on others. These are vendor-reported results, not proof of an advantage in every setting.

What remains unknown about the release?

The supplied report does not give broad results across consumer devices, battery-use measurements, independent benchmark replication or detailed licensing and distribution terms.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Introducing ChatGPT For Teens: AI That Combines Education, Safety, And Innovation

OpenAI announced ChatGPT for Teens, a new AI service aimed at supporting learning with safety protections, though details on safeguards and availability remain unclear.

The European Bet: How Mistral, Aleph Alpha, and Black Forest Labs Are Playing a Different Game

European AI vendors Mistral, Aleph Alpha, and Black Forest Labs are positioning for the EU AI Act’s enforcement, emphasizing compliance, sovereignty, and open models.

ChannelHelm – Drop a video. Get a publishing kit.

ChannelHelm introduces a new tool that automates video asset creation, enabling creators to generate comprehensive publishing kits from a single video file or link.

Capability or Control: The European Enterprise AI Playbook for the AI Act Era

How European companies navigate the AI Act: balancing capability, control, licensing, and jurisdiction for compliance and resilience.