$25M for Modulate: The Next Voice AI Breakthrough Is Listening
Story
Modulateʼs $25 million financing, its benchmark-leading models, and why audio-native intelligence is essential infrastructure.
Sep 29, 2026

Long before the generative AI boom, we led Modulate’s pre-seed and seed financing with a clear conviction: voice is more than an interface for AI; instead, it’s a foundational layer of AI.
Today, we’re proud to support the company’s next chapter: $25 million in new funding led by Steve Jurvetson of Future Ventures, with participation from Hyperplane and Lakestar. The financing brings Modulate’s total funding to $60 million.
Modulate’s audio-native models have reached #1 on public transcription and voice deepfake detection benchmarks. Those results are part of a broader ambition: building a platform that understands conversations beyond the words being spoken.
For us, these milestones reinforce the thesis behind our original investment:
The opportunity in voice AI is not just making machines speak. It is understanding emotion, tone, and intent in a conversation.
That is the layer Modulate has spent years developing.
A transcript is not the conversation
Consider a simple sentence: “Sure, that sounds just fine.”
Depending on how they are delivered, those words might express enthusiasm, resignation, irritation, or sarcasm. A useful response depends on hearing the difference.
The same distinction applies to entire conversations. Did a customer agree, or did they simply give up? Is an agent helping, or interrupting the person it is supposed to serve? Does a suspicious call contain synthetic speech, manipulative behavior, or both?
A transcript is valuable. But reducing an interaction to text ignores the signals needed to answer those questions.
Modulate’s flagship Velma platform analyzes audio for signals including emotion, tone, intent, emphasis, synthetic speech, and conversational behavior. Those signals can be combined to identify higher-level events like potential fraud, customer dissatisfaction, harassment, and voice-agent failures. Velma can operate in real time, enabling intervention during a conversation, not just afterward.
The strategic question is not simply whether an AI model accepts audio. It is how reliably, efficiently, and usefully it understands the conversation.
What the #1 results say
Benchmark results make that ambition tangible. Three comparisons are especially useful to unpack.
1. English transcription: competing at the frontier
In a comparison among ten providers, Modulate’s vfast ranked first with a 4.43% average word error rate, ahead of the models at Microsoft, ElevenLabs, NVIDIA, and OpenAI. All ten scores use the same se public English test sets.

Figure 1. Hugging Face Open ASR, August 2026: ten selected providers, not the overall top ten. These archived results predate the August 21 update; lower average word error rate is better.
Word error rate measures substitutions, deletions, and insertions relative to the reference transcript.
These results suggest that Modulate can compete at the frontier of speech recognition while building a broader audio-understanding platform.
2. Voice deepfake detection: another leading result
Speech DF Arena tests systems across 14 evaluation datasets. In the comparison below, Modulate’s average equal error rate is 1.104% compared to 2.113% for Hiya and
2.570% for Resemble AI, which is almost 50% lower than the next-best average EER shown.

Figure 2. Speech DF Arena, August 2026 entry cohort: ten systems ordered by average equal error rate; lower is better. Published scores are used to reconstruct the cohort, not a frozen August score archive.
The strategic distinction matters. Transcription asks what was said. Deepfake detection asks whether the voice was synthetically generated. Neither, on its own, tells you whether a conversation is fraudulent: legitimate applications use synthetic voice and human callers can commit fraud.
We believe trusted voice applications will need to combine these capabilities with an understanding of behavior and context. That combination is much more powerful than a standalone detector.
3. Beyond English: accuracy and latency together
The Mandarin leaderboard on Sierra’s μ-Bench shows Modulate’s multilingual vfas model in first place among eight entries at 27.38% utterance error rate versus 29.52% for Google Chirp 3.

Figure 3. Sierra μ-Bench, August 2026: all eight entries with Chinese / Mandarin (zh-CN) results in the archived leaderboard. Lower utterance error rate is better.
The same archive reports P95 batch latency of 658 milliseconds for Modulate and 1,196 milliseconds for Google Chirp 3. These batch-processing measurements capture the end-to-end response times of a live voice agent.
The conclusion from these results is a practical lesson for founders: evaluate quality and latency together in the languages and conditions your customers use. A generic claim of “good transcription” is not enough.
Why the architecture matters
The company’s approach is built around a proprietary Ensemble Listening Model, ELM. Rather than relying on a single massive foundation model for every task, Velma’s architecture selects and combines more than 100 specialized audio models.
Think of the difference between asking one generalist to explain everything and coordinating specialists around the questions that matter. Different models contribute different signals; orchestration turns those signals into a useful understanding of the interaction.
We like this architecture because it creates room to improve individual capabilities, combine them for new applications, and match the work performed to the specific task. An ensemble is not a moat on its own. The value is the specialists, orchestration, evaluation discipline, and economics of running the system.
Economics matter. Modulate’s announced price for batch transcription is $0. per audio hour. At that base rate, one million hours costs $30,000 for transcription alone, not for the full Velma platform or every additional capability.
For a founder building a voice product, inference cost is critical. It helps determine whether a capability can be used throughout the product or only in a small subset of interactions.
Our thesis is that winners in voice infrastructure will need strong accuracy and viable economics.
Accumulated learning, not just accumulated hours
Modulate has processed more than 600 million hours of audio, with more than 10 million hours analyzed each month.
The company’s work in gaming is central to that story. Activision publicly announced its use of Modulate’s ToxMod for Call of Duty voice moderation in 2023. Modulate has also published case studies involving Grand Theft Auto and Rainbow Six Siege.
The Call of Duty results illustrate the power of its solution. Activision reported a combined 67% reduction in repeat voice-chat offenders across Modern Warfare III and Warzone after rolling out improved voice-chat enforcement. Over 80% of players who had received a voice-chat enforcement action since launch had not reoffended.
One finding from the joint Modulate–Activision case study speaks directly to our thesis: only about 23% of player-generated reports contained actionable evidence of Code of Conduct violation. A report is a signal, not a substitute for understanding what happened. You have to be listening.
Gaming is a demanding proving ground for voice intelligence. Consider the environment: overlapping speakers, variable microphones, background noise, slang, emotionally charged exchanges, and people who may deliberately try to evade moderation. Understanding the words is only part of the problem. Distinguishing harmless banter from harmful behavior requires context.
The moat we underwrite is the combination of specialized models, real-world operating experience, evaluation judgment, and integration into consequential workflows.
A competitor can reproduce an impressive demo more quickly than it can reproduce years of experience with the situations where a system fails, the false alarms users do not tolerate, and the evidence customers need before acting.
Where customer permissions allow learning from data and feedback, that experience can support a virtuous cycle: better evaluation exposes weaknesses, improved specialists address them, and stronger products create opportunities for further deployment. This represents a potential compounding advantage.
What this means for the voice AI stack
Our thesis is that speech generation and conversation understanding are complementary layers.
A voice agent needs to produce an answer. The business deploying it also needs to know whether it understood the customer, interrupted at the wrong moment, missed a warning sign, or should have handed the conversation to a person.
A security team needs to examine synthetic speech alongside suspicious behavior. Social platforms need to detect harm without treating every heated exchange as a violation. Different applications can draw on the same underlying audio capabilities.
We believe that creates an opportunity for a shared understanding layer: infrastructure developers can integrate rather than rebuilding specialized audio intelligence for their product.
For the founders in our portfolio, the implication is practical. Evaluate not just how convincingly your application speaks, but how well it recognizes risks and the need to intervene. The most polished voice is not necessarily the most reliable one.
For our LPs, Modulate illustrates the kind of AI advantage we find compelling: technical depth developed in a demanding application with the potential to reach a much broader market.
The next chapter
The new financing will support Modulate’s research, product, and engineering work, developer ecosystem, and partnerships as it expands its models, APIs, and deployment options.
We are proud to have led the pre-seed and seed financings, grateful to continue supporting the company, and excited to welcome Steve Jurvetson and Future Ventures to this next chapter.
Congratulations to Mike Pappas, Carter Huffman, Graham Gullans, and the entire Modulate team.
The future of voice AI will not be defined only by how well machines speak. It will also be defined by how well they listen.
We believe Modulate is building a foundational part of that future.