
Linden by Speechmatics
Overview
- Deliver near-instant voice agent responses with sub-350ms median end-of-speech finalization, so LLMs draft replies before callers notice a pause.
- Eliminate misheard names, postcodes, and account numbers with purpose-built alphanumeric handling that stops the model from "correcting" spoken codes into wrong ones.
- Know exactly who said what on every call with real-time speaker diarization that tags each word to the correct speaker, with no cap on speakers per call.
- Recognize returning callers instantly with Speaker ID, which identifies previously-heard voices at no extra cost.
- Serve global callers accurately across 55+ languages, with particular strength in global English accents, Arabic with code-switching, Nordic, Japanese, and Central/Eastern European speech.
- Get company names, product names, and industry jargon right the first time by adding up to 1,000 custom vocabulary words.
- Cut integration time to a config toggle with pre-built support for LiveKit Inference, LiveKit Agents, Pipecat, Vapi, and Jambonz, plus a Python SDK and direct WebSocket API for custom pipelines.
- Meet data residency requirements and reduce latency by pinning traffic to US, Europe, or Australia regional endpoints.
- Start production voice agents on a budget with pricing from $0.30/hour, volume discounts to $0.15–0.16/hour, and $100 in free credit for new users.
- Hand pre-structured, punctuated, turn-based transcripts straight to your LLM without post-processing, thanks to built-in turn detection and intelligent segmentation.
- Trust benchmarked performance: Linden sits on Pipecat's Pareto frontier at 369ms median time-to-final-segment and 1.05% pooled semantic word error rate—the only model both faster and more accurate than competitors.
Pros & Cons
Pros
- Sub-350ms latency without sacrificing accuracy
- Speaker diarization and Speaker ID included at no extra cost
- Purpose-built for voice-agent failure points (names, numbers, addresses, short one-word turns)
- Supports 55+ languages with strong accent handling
- Custom vocabulary support for domain-specific terms
- Independently benchmarked on Pipecat's Pareto frontier, fastest AND most accurate among 23 models tested
- Multiple integration paths: LiveKit, Pipecat, Vapi, Jambonz, SDK, or raw API
- Lower published price than Deepgram Flux and AssemblyAI Universal-3.5
- Backed by 15+ years of speech recognition work at Speechmatics
- $100 free credit to start
- ISO 27001, SOC 2, HIPAA, and GDPR compliant
Cons
- SaaS-only, no on-device or local deployment for Agent STT specifically
- Turn detection currently based on speaker pauses; smart turn detection is "coming soon"
- Fully multilingual mid-sentence language switching not yet available (arriving with Linden 2)
- Pricing is usage-based (per hour), which can be harder to predict than flat subscription pricing
- Requires technical integration (API/SDK/framework), not a no-code end-user tool
Reviews
Rate this tool
Loading reviews...
❓ Frequently Asked Questions
A speech-to-text API built for production voice agents, powered by Linden, Speechmatics' latest model. It returns transcription, structured conversational signals, speaker attribution, and custom vocabulary support over a single connection.
End-of-speech finalization is under 350ms mean, and around 440ms at P95, in Speechmatics' internal measurements. On Pipecat's independent benchmark, it lands at 369ms median with 1.05% pooled semantic word error rate.
55+ languages at launch, including bilingual and quadlingual models, with particular strength in global English accents, Arabic (including Saudi and Egyptian accents with code-switching), Nordic languages, Japanese, and Central and Eastern European speech. Fully multilingual mid-conversation language switching arrives with Linden 2.
Agent STT provides speech and segment events to drive turn logic, currently based on speaker pauses. Smart turn detection is coming soon, contact Speechmatics for early access.
$0.30/hour, with volume discounts bringing it down to $0.15-0.16/hour as usage scales.
LiveKit Inference and LiveKit Agents, Pipecat, Vapi, and Jambonz, plus a Python SDK and a direct WebSocket API for custom pipelines or unsupported platforms.
Not detailed in the material provided, check Speechmatics' documentation on regional endpoints and data residency for specifics, they do support pinning to US, Europe, or Australia regions.
Yes, up to 1,000 custom words can be added for company names, product names, or domain-specific jargon.
Agent STT is purpose-built for the turn-based, low-latency needs of voice agents, speaker diarization, turn detection, and structured segments arrive by default, whereas general-purpose transcription APIs are optimized for accuracy on a fuller range of use cases without the same agent-specific latency and structuring guarantees.
Pricing
Pricing model
Paid
Paid options from
$0.30/unit
Billing frequency
Pay-as-you-go
















