A transcript can tell a system that someone said, “Yes, that’s me.” It cannot by itself tell whether the voice was synthesized or what acoustic signals surrounded the words. Modulate’s $25 million financing is a bet that voice AI needs a second intelligence layer built from the audio itself.1

IN BRIEF

Modulate raised $25 million to expand audio-native AI that analyzes signals such as tone, conversational behavior and synthetic speech alongside words. The company says it has analyzed more than 600 million hours of audio and reports 98.9% deepfake-detection accuracy on public benchmark data. Those figures do not establish the same accuracy in live adversarial calls.1, 2

Modulate’s company-reported scale. New financing: $25M — Financing announced in September 2026.. Audio hours analyzed: 600M+ — Company-reported cumulative audio analyzed across its systems.. Deepfake benchmark accuracy: 98.9% — Company-reported result on public benchmark data, not a guarantee for live calls.. Values and their context are also available as HTML below.
Modulate’s company-reported scale. Values and their context are also available as HTML below.1

Modulate’s company-reported scale

$25M
New financing1

Financing announced in September 2026.

600M+
Audio hours analyzed1

Company-reported cumulative audio analyzed across its systems.

98.9%
Deepfake benchmark accuracy1

Company-reported result on public benchmark data, not a guarantee for live calls.

Speech-to-text throws away information on purpose

Transcription reduces an audio stream to words. That is useful for search, summarization and language reasoning. It also discards characteristics of the signal that can carry information about speaker behavior, synthetic generation and conversational dynamics.1

Two layers of a voice interaction1
LayerExamplesUseful for
TranscriptWords, entities, stated intentSearch, summarization and language reasoning
Audio-native signalsAcoustic patterns, tone, synthetic-speech cues, turn-takingDeepfake detection, moderation and richer call analysis

The security use case makes the distinction obvious

A fraudster can generate words that look ordinary in a transcript. Deepfake detection instead examines characteristics of the audio signal. Performance can therefore change with microphones, compression, languages and attack methods that a clean benchmark may not reproduce.1

Voice agents create a supervision market too

As companies deploy AI agents into phone calls, the same audio layer can analyze both sides of the interaction. That can help flag unusual conversations without pretending the transcript captures every relevant signal.1

Questions behind any voice-AI accuracy claim

  • Was the result measured on a public benchmark or live production calls?
  • Which languages, codecs and recording conditions were represented?
  • Were the synthetic voices known to the detector or genuinely novel attacks?
  • What false-positive rate accompanies the headline accuracy number?

Voice AI is often described as speech-to-text followed by a language model. Modulate’s thesis is that the sound itself deserves a model too. The financing shows capital moving into that layer, while real-world robustness remains the harder proof.

Sources and methodology

Sources checked September 29, 2026. Dates and periods for individual figures are stated beside them.

  1. Modulate: $25 million financing announcement ↗Accessed 2026-09-29
  2. FTC: voice cloning challenge and consumer harms ↗Accessed 2026-09-29
Scope and assumptions

The 600M+ audio-hours figure and 98.9% benchmark accuracy are company-reported.

Public benchmark performance may not transfer to novel deepfakes, noisy channels or live adversarial calls.

Continue reading

1B Monthly Users—and 63% Talk to Gemini Out Loud →

Unit 42 Says New CVEs Can Be Weaponized in 15 Minutes. Cybersecurity Is Becoming AI vs. AI →

86% Agreement Can Still Miss Half the AI Failures →