Liden by Speechmatics
Overview
- Cut caller-perceived silence to near zero with end-of-speech finalization landing under 350ms median, giving your LLM a head start on every reply
- Eliminate wrong names, postcodes, and account numbers with purpose-built handling for alphanumeric strings that stops the model from 'correcting' spoken codes into plausible-but-wrong values
- Know exactly who said what on every call with speaker diarization included by default, tagging every word to the correct speaker in real time with no cap on speaker count
- Recognize returning callers by voice with Speaker ID, letting your agent identify previously-heard voices without extra cost or setup
- Get transcripts your LLM can consume immediately with turn-based, punctuated, speaker-labelled segments delivered pre-structured over a single streaming connection
- Serve global callers without accent drop-off, covering 55+ languages with particular strength in global English accents, Arabic with Saudi and Egyptian code-switching, Nordic languages, Japanese, and Central and Eastern European speech
- Stop mis-transcribing company names, product names, and industry jargon by loading up to 1,000 custom words into the model's vocabulary
- Ship faster on infrastructure you already run with pre-built integrations for LiveKit Inference, LiveKit Agents, Pipecat, Vapi, and Jambonz, plus a Python SDK and direct WebSocket API for custom pipelines
- Meet data residency requirements by pinning traffic to US, Europe, or Australia regional endpoints while reducing latency for local callers
- Scale from prototype to production on $0.30/hour pricing that drops to $0.15–0.16/hour at volume, starting with $100 in free credit
Pros & Cons
Pros
- Sub-350ms latency without sacrificing accuracy
- Speaker diarization and Speaker ID included at no extra cost
- Purpose-built for voice-agent failure points (names, numbers, addresses, short one-word turns)
- Supports 55+ languages with strong accent handling
- Custom vocabulary support for domain-specific terms
- Independently benchmarked on Pipecat's Pareto frontier, fastest AND most accurate among 23 models tested
- Multiple integration paths: LiveKit, Pipecat, Vapi, Jambonz, SDK, or raw API
- Lower published price than Deepgram Flux and AssemblyAI Universal-3.5
- Backed by 15+ years of speech recognition work at Speechmatics
- $100 free credit to start
- ISO 27001, SOC 2, HIPAA, and GDPR compliant
Cons
- SaaS-only, no on-device or local deployment for Agent STT specifically
- Turn detection currently based on speaker pauses; smart turn detection is "coming soon"
- Fully multilingual mid-sentence language switching not yet available (arriving with Linden 2)
- Pricing is usage-based (per hour), which can be harder to predict than flat subscription pricing
- Requires technical integration (API/SDK/framework), not a no-code end-user tool
Reviews
Rate this tool
Loading reviews...
❓ Frequently Asked Questions
A speech-to-text API built for production voice agents, powered by Linden, Speechmatics' latest model. It returns transcription, structured conversational signals, speaker attribution, and custom vocabulary support over a single connection.
End-of-speech finalization is under 350ms mean, and around 440ms at P95, in Speechmatics' internal measurements. On Pipecat's independent benchmark, it lands at 369ms median with 1.05% pooled semantic word error rate.
55+ languages at launch, including bilingual and quadlingual models, with particular strength in global English accents, Arabic (including Saudi and Egyptian accents with code-switching), Nordic languages, Japanese, and Central and Eastern European speech. Fully multilingual mid-conversation language switching arrives with Linden 2.
Agent STT provides speech and segment events to drive turn logic, currently based on speaker pauses. Smart turn detection is coming soon, contact Speechmatics for early access.
$0.30/hour, with volume discounts bringing it down to $0.15-0.16/hour as usage scales.
LiveKit Inference and LiveKit Agents, Pipecat, Vapi, and Jambonz, plus a Python SDK and a direct WebSocket API for custom pipelines or unsupported platforms.
Not detailed in the material provided, check Speechmatics' documentation on regional endpoints and data residency for specifics, they do support pinning to US, Europe, or Australia regions.
Yes, up to 1,000 custom words can be added for company names, product names, or domain-specific jargon.
Agent STT is purpose-built for the turn-based, low-latency needs of voice agents, speaker diarization, turn detection, and structured segments arrive by default, whereas general-purpose transcription APIs are optimized for accuracy on a fuller range of use cases without the same agent-specific latency and structuring guarantees.
Pricing
Pricing model
Paid
Paid options from
$0.30/unit
Billing frequency
Pay-as-you-go








