A transcript can tell a system that someone said, “Yes, that’s me.” It cannot by itself tell whether the voice was synthesized or what acoustic signals surrounded the words. Modulate’s $25 million financing is a bet that voice AI needs a second intelligence layer built from the audio itself.1
Modulate raised $25 million to expand audio-native AI that analyzes signals such as tone, conversational behavior and synthetic speech alongside words. The company says it has analyzed more than 600 million hours of audio and reports 98.9% deepfake-detection accuracy on public benchmark data. Those figures do not establish the same accuracy in live adversarial calls.1, 2

Modulate’s company-reported scale
Speech-to-text throws away information on purpose
Transcription reduces an audio stream to words. That is useful for search, summarization and language reasoning. It also discards characteristics of the signal that can carry information about speaker behavior, synthetic generation and conversational dynamics.1
| Layer | Examples | Useful for |
|---|---|---|
| Transcript | Words, entities, stated intent | Search, summarization and language reasoning |
| Audio-native signals | Acoustic patterns, tone, synthetic-speech cues, turn-taking | Deepfake detection, moderation and richer call analysis |
The security use case makes the distinction obvious
A fraudster can generate words that look ordinary in a transcript. Deepfake detection instead examines characteristics of the audio signal. Performance can therefore change with microphones, compression, languages and attack methods that a clean benchmark may not reproduce.1
Voice agents create a supervision market too
As companies deploy AI agents into phone calls, the same audio layer can analyze both sides of the interaction. That can help flag unusual conversations without pretending the transcript captures every relevant signal.1
Questions behind any voice-AI accuracy claim
- Was the result measured on a public benchmark or live production calls?
- Which languages, codecs and recording conditions were represented?
- Were the synthetic voices known to the detector or genuinely novel attacks?
- What false-positive rate accompanies the headline accuracy number?
Voice AI is often described as speech-to-text followed by a language model. Modulate’s thesis is that the sound itself deserves a model too. The financing shows capital moving into that layer, while real-world robustness remains the harder proof.
Sources and methodology
Sources checked September 29, 2026. Dates and periods for individual figures are stated beside them.
- Modulate: $25 million financing announcement ↗Accessed 2026-09-29
- FTC: voice cloning challenge and consumer harms ↗Accessed 2026-09-29
Scope and assumptions
The 600M+ audio-hours figure and 98.9% benchmark accuracy are company-reported.
Public benchmark performance may not transfer to novel deepfakes, noisy channels or live adversarial calls.
Continue reading
1B Monthly Users—and 63% Talk to Gemini Out Loud →
Unit 42 Says New CVEs Can Be Weaponized in 15 Minutes. Cybersecurity Is Becoming AI vs. AI →