Speechmatics and LiveKit integrate to close voice agent accuracy gap
Speechmatics and LiveKit have announced a native integration that makes Speechmatics' Linden speech-to-text model available directly through LiveKit Inference, the San Francisco company's unified model-serving platform. Developers building voice agents on LiveKit can now select Linden via the UI or a single line of code, with no separate API key, account or billing relationship required.
Linden is described by Speechmatics as its first model purpose-built for voice agent use cases, with particular attention to the conditions that cause transcription to fail in production: strong accents, non-native speakers, low-quality telephone audio, and alphanumeric strings such as card and account numbers. The model supports more than 55 languages. LiveKit Inference, which launched in October 2025, consolidates STT, LLM and TTS model access under a single API key and handles routing, billing, rate-limit management and turn detection.
The production accuracy problem
The integration targets a well-documented failure mode in deployed voice agents: demo performance rarely reflects what happens on a live contact-centre line or a poor mobile connection. When a speech-to-text layer mis-transcribes a key entity, such as a policy number or a postcode, every downstream step in the agent pipeline inherits that error. The compounding effect tends to erode user trust faster than outright silence.
Ricardo Herreros-Symons, chief revenue officer at Speechmatics, said the gap between a trusted agent and an exhausting one comes down to whether the system understands what was actually said. "Linden not only excels at 55+ languages, but also transcribing non-native speakers, dialects, and accents. LiveKit built the infrastructure that made those deployments possible," he added.
Norwegian conversational AI firm boost.ai, which already runs both vendors in production for European enterprise customers, also provided a named endorsement. Rasmus Hauch, CTO at boost.ai, cited low-latency real-time models, fast iteration on edge cases including alphanumerics and end-of-utterance detection, and breadth of European language coverage as the factors that matter in live deployments. The boost.ai infrastructure is noted in the release as also powering xAI Grok and Salesforce Agentforce, though the nature of those relationships is not elaborated upon.
Market context and competitive positioning
The enterprise voice agent market is consolidating around a small number of orchestration platforms that abstract away model-level complexity, a pattern already established in the LLM layer by services such as AWS Bedrock and Azure OpenAI Service. LiveKit Inference is pursuing a comparable position for real-time, multimodal agent pipelines, where low-latency STT is often the most latency-sensitive component.
Speechmatics competes in the speech recognition space with established names including Deepgram, AssemblyAI and the speech APIs from the major hyperscalers. Its differentiation claim, centred on accent and language breadth rather than raw throughput benchmarks, is consistent with its historical positioning in regulated verticals such as financial services, government and media, where transcript accuracy on rare vocabulary carries legal or compliance weight.
David Zhao, co-founder and CTO of LiveKit, highlighted speaker diarisation as an immediate capability gain for developers, noting it is particularly relevant for physical AI and multi-speaker environments. LiveKit frames the integration as part of a broader strategy to let builders swap models without rewriting their stack.
The two companies did not disclose commercial terms, revenue-share arrangements or a combined customer count. Speechmatics is headquartered in Cambridge, UK; LiveKit in San Francisco. Both companies serve enterprise customers in regulated sectors where accuracy on edge-case inputs is a procurement criterion rather than a differentiator to be taken on trust.