Whistle: Speech To Text In 16.9 MB
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Cactus Compute announced Whistle, a 16.9 MB speech-recognition model designed to run locally on CPUs and support seven languages. The company reports competitive word-error results on several benchmarks and an 11.1 ms time to first token for a 10-second clip on an Apple M4 Pro; independent verification and broader device results are not provided in the source.

Cactus Compute announced Whistle, a 16.9 MB speech-recognition model that the company says can transcribe speech directly on CPUs without software dependencies. The model supports English, German, French, Spanish, Italian, Dutch and Polish, and is intended for devices including phones, wearables, robots, smart-home systems and vehicles.

According to Cactus Compute’s October 2 announcement, Whistle processes 16 kHz mono audio and can handle clips of up to 30 seconds in one pass. It can detect the spoken language automatically or accept a language choice. Its reported outputs include a transcript, word-level timestamps and probabilities, and speech embeddings: encoder representations produced without generating a transcript.

The company says the model runs locally, so audio stays on the device when the model is used in its browser demonstration after the download. The first use downloads the 16.9 MB model file. Cactus says Whistle shares the C++ engine, container and quantisation used by its Needle model. It describes the speech encoder and decoder as using components shared with Needle, with gated cross-attention added to connect speech features to text generation.

For decoding, Whistle uses five-beam search, according to the report. Developers can provide keywords for biasing, while an audio-depth setting selects among decoder configurations. Cactus says the audio encoder still runs all eight of its blocks at every setting. The company also describes a silence check that can return an empty transcript without starting beam search when audio falls below its loudness threshold.

At a glance
announcementWhen: Announced October 2, 2026
The developmentCactus Compute released Whistle, a compact speech-to-text model that it says runs on-device in its existing C++ engine.

Small Models for Local Speech

Whistle’s main pitch is the combination of a small download and speech processing on the device. Local inference can be useful where an internet connection is unavailable, where response time matters, or where developers want to keep recorded audio from being sent to a remote service. The announcement describes possible use in mobile, wearable, automotive and embedded products, though it does not establish that the model has been deployed in those settings.

The reported size and speed could also make speech recognition practical on hardware with limited resources. Cactus reports an 11.1 millisecond time to first token for a 10-second audio clip on an Apple M4 Pro CPU, along with 1,319 decoded tokens per second. Those figures are the company’s benchmark results, not a guarantee of performance on other processors or in applications with different audio and latency requirements.

Model size alone does not determine whether a speech system is useful. Accuracy across accents, noisy recordings and real-world speech, as well as memory use, power consumption and licensing, will affect adoption. The supplied report compares Whistle with larger systems and says it scores better on some test sets and worse on others, rather than establishing a universal accuracy advantage.

Amazon

on-device speech to text software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Cactus Benchmarked Whistle

Cactus published word-error-rate comparisons with Whisper base and Moonshine tiny v2. The report says Whistle had lower word error rates on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average. It says Whisper base scored better on TED-LIUM, AMI and the MLS average. Moonshine tiny v2 is described as English-only, and some benchmark cells are absent because the compared models’ authors did not publish those results.

The report flags a limitation in the AMI comparison: Whisper’s listed result uses the AMI-IHM subset, while the other models’ results use a different AMI subset. It also compares model sizes—16.9 MB for Whistle, 145.3 MB for Whisper base and 41.9 MB for Moonshine tiny v2—and reports first-token times of 11.1, 73.2 and 22.8 milliseconds, respectively.

For the speed test, Cactus says each model ran on its official runtime with default settings using 10 seconds of audio on an Apple M4 Pro CPU. Whistle used five beams; Whisper ran through openai-whisper; Moonshine used moonshine-voice in non-streaming mode over the full clip. The company defines first-token time as audio input to first token and calculates decode speed after that point. It notes that Whisper pads every input to 30 seconds, while Whistle’s first-token time changes with clip length.

““It is one 16.9 MB file, runs on the CPU with no dependencies, and loads into the same C++ engine as Needle.””

— Cactus Compute, in its Whistle announcement

Amazon

offline speech recognition app

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Results Still Pending

The supplied announcement does not include an independent evaluation of Whistle’s accuracy, latency or resource use. Its benchmarks describe particular datasets, runtimes, hardware and settings; results may differ on other processors, with other audio lengths, or in noisy environments. The source does not give a complete account of memory requirements during inference, power draw, supported hardware beyond its stated CPU target, or model licensing terms.

The privacy description is also specific to the browser demo: the first press downloads the model, and Cactus says audio remains on the device during transcription. The report does not detail data handling for every possible integration or product built with the model. It is also unclear from the supplied material when model weights, developer documentation or production-ready packages will be available, and whether the stated multilingual performance is comparable across all seven supported languages.

Amazon

multilingual speech transcription device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Availability and Device Testing

The Whistle browser demonstration described by Cactus offers a way to try transcription, subject to downloading the 16.9 MB model. The company’s post presents the release for use in mobile, wearable, robotic, smart-home, automotive and microcontroller applications, but it does not provide a product-by-product rollout schedule.

The next useful evidence will be independent testing across devices and languages, along with practical information about licensing, memory and power use. Developers will also need to establish whether Whistle’s stated 30-second clip limit and accuracy profile fit their applications. Cactus has not specified further milestones in the supplied announcement.

Amazon

small speech recognition model

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Whistle?

Whistle is a speech-recognition model released by Cactus Compute. The company says its 16.9 MB model runs on a CPU and transcribes speech in seven languages.

Which languages does Whistle support?

Cactus lists English, German, French, Spanish, Italian, Dutch and Polish. The model can detect a language automatically or use a language specified by the user.

Does Whistle send audio to a server?

Cactus says audio in its browser demo stays on the device after the model is downloaded. The source does not describe data handling for every possible third-party integration.

How fast is Whistle?

Cactus reports an 11.1 ms time to first token for 10 seconds of audio on an Apple M4 Pro CPU. The company’s result is tied to that hardware and test setup; performance on other devices is not established by the supplied report.

Is Whistle more accurate than Whisper?

Not across every benchmark reported. Cactus says Whistle had lower word error rates on several named datasets, while Whisper base scored better on TED-LIUM, AMI and the MLS average. The company notes that the AMI results use different subsets, and the figures have not been independently verified in the supplied material.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How A Dedicated Logistics Workspace Supports Nonprofit Civic Initiatives

A new dedicated logistics workspace prototype supports nonprofit organizers of citizens’ assemblies, addressing manual workflows and scaling civic engagement efforts.

Virtual Museum Tour Guide

Discover captivating virtual museum tours that transport you around the world, but wait until you see the tips to enhance your experience even more!

Unlocking The Power Of AI Search To Better Enjoy Everyday Life

Google outlines five ways its Search AI Mode can assist with offline activities, including finding classes, shopping, and event tickets, with some functions requiring app connections.

Group Activities for Art Appreciation

Appreciate art like never before through engaging group activities that spark creativity and collaboration—discover innovative ideas to enhance your experience!