The funding brings Wispr Flow’s total capital raised to $361 million, a figure that signals how aggressively venture backers are positioning themselves in the AI voice-to-text category. Menlo Ventures led the round as a long-time investor, with existing backers Notable Capital, NEA, Neo Ventures, 8VC, and MVP Ventures participating alongside new investors Acrew, Forerunner, and Goodwater.
Canto Model Targets Real-World Conditions
Alongside the funding, Wispr Flow previewed its proprietary speech model, called Canto. The company built Canto to handle the environments where users actually deploy dictation tools — settings with background noise, wind, heavy accents, or music playing. Under those conditions, Wispr Flow says error rates drop from more than 30% of words to between 5% and 10%.
“We built this model for where people actually use Flow,” CEO Tanay Kothari said. “In the hardest conditions, with background noise, wind, heavy accents or music, error rates fall from more than 30% of words to somewhere between 5 and 10%.”
Why Investors Are Betting Big On Voice AI
The $2 billion valuation reflects broader momentum in AI voice-to-text, a category that has moved from niche accessibility tooling into mainstream productivity workflows. Speech recognition has improved dramatically with large language model architectures, and the use cases have expanded well beyond medical dictation and courtroom transcription. Enterprise users now expect voice input to work reliably across noisy offices, mobile environments, and multilingual settings.
Wispr Flow’s pitch hinges on closing the gap between controlled-environment accuracy and real-world performance. Most speech recognition systems perform well in quiet rooms with clear audio. The harder problem — and the one Wispr Flow is positioning Canto to solve — is maintaining accuracy when conditions degrade. A reduction from 30% word error rates to single digits in adversarial conditions would represent a meaningful step change in usability.
The Competitive Landscape
Wispr Flow is not alone in chasing this market. Established players like Microsoft-owned Nuance, Google’s speech APIs, and OpenAI’s Whisper model have set baselines for accuracy. Startups including AssemblyAI, Deepgram, and Speechmatics are also building specialized speech recognition infrastructure. What differentiates Wispr Flow is its end-user product focus — Flow is a consumer and prosumer dictation application rather than a developer API, which positions the company closer to the productivity software market than the infrastructure layer.
The mix of investors in this round also tells a story. Menlo Ventures leading as a repeat backer suggests confidence in the team’s execution. The addition of Forerunner, known for consumer technology investments, and Acrew, which focuses on enterprise and infrastructure, indicates Wispr Flow is drawing interest from firms that see applicability across both consumer and business markets.
What Happens Next
With $361 million in total funding and a $2 billion valuation, Wispr Flow now faces the pressure of justifying that price tag through product traction and revenue growth. The Canto model remains in preview, meaning the company will need to ship it to production and demonstrate that the claimed accuracy improvements hold up at scale across diverse user bases and languages.
Watch for Wispr Flow to expand Canto’s language coverage, deepen enterprise integrations with productivity suites, and potentially move upmarket into verticals like healthcare and legal where dictation accuracy is mission-critical. The broader AI voice-to-text category will likely see further consolidation as well-funded startups compete with hyperscaler offerings, and Wispr Flow’s capital position gives it runway to be an acquirer rather than an acquisition target in the near term.
— David Kim, technology desk, AXO News