Skip to content
Woyce Technologies
AboutTeamCareersContactStart a project →

Hugging Face's speech-to-speech Explained: A Modular Open Voice Stack

Hugging Face's speech-to-speech is an open-source, OpenAI Realtime-compatible voice agent pipeline with swappable VAD, STT, LLM, and TTS components — production-tested as the conversation backend for thousands of Reachy Mini robots.

Hugging Face's speech-to-speech Explained: A Modular Open Voice Stack — Woyce Technologies

Loading repository details…

——

Building a voice agent usually means choosing between an easy, fully-hosted realtime API you don't control, or stitching together separate speech, language, and synthesis models yourself. Hugging Face's speech-to-speech is built to close that gap: an open-source pipeline that speaks the OpenAI Realtime API protocol so any client built for it works unmodified, while every component underneath — the transcription model, the language model, the voice synthesizer — is swappable, self-hostable, and can run fully offline.

That gap matters because the choice tends to get locked in early. A team that builds on a hosted realtime API gets low latency fast, but inherits per-minute pricing, a fixed set of voices, and no way to keep audio on its own hardware. A team that wires its own cascade gets control, but spends weeks on turn detection, interruption handling, and session plumbing before the agent feels natural. This explainer walks through what Hugging Face speech-to-speech actually is (and what its name does not mean), why the Realtime protocol compatibility is the real product decision, how each stage can be swapped, how a modular cascade compares with a unified audio model, and what it takes to get a server running locally or behind your own gateway.

Three voice agent build paths: a hosted realtime API with lock-in, a DIY cascade that costs weeks of plumbing, and HF speech-to-speech with a compatible protocol and swappable stages.

Despite the Name, It's a Modular Pipeline, Not a Single-Model Architecture

This is worth clarifying directly, because the naming is easy to misread if you already know the term: "speech-to-speech" in the voice AI field increasingly refers to a specific architecture — a single model that processes audio in and produces audio out directly, with no intermediate text transcript. Hugging Face's project, despite the name, is the traditional four-stage cascade: Voice Activity Detection, then Speech-to-Text, then an LLM, then Text-to-Speech, with each stage running in its own thread connected by queues. That's not a criticism — the modular approach has real advantages the single-model approach doesn't (every component is independently swappable, debuggable, and upgradable) — but it's a meaningfully different architecture than what the term "speech-to-speech" usually implies elsewhere, and worth knowing before you assume this is a true unified audio model.

OpenAI Realtime Compatibility Is the Actual Product Decision

The pipeline exposes itself as an OpenAI Realtime-compatible WebSocket (and, as of the current release, WebRTC) API. That single design choice is what makes the project genuinely practical rather than just another voice pipeline demo: any client already built against OpenAI's Realtime protocol can point at a self-hosted speech-to-speech server instead, without rewriting the client. The project's own demo shows exactly that — an OpenAI Realtime client switching its endpoint from hosted OpenAI to a self-hosted server live, with no other changes. For a team that's already built a product against OpenAI's Realtime API and wants a self-hosted or fully local fallback path, that compatibility is the whole reason to look at this project first.

Every Stage Is Genuinely Swappable

The four-stage cascade — VAD, STT, LLM, TTS — has multiple interchangeable backends at every stage, selected via CLI flags, with a stated focus on models available through Transformers and the Hugging Face Hub. The default install alone covers a complete realtime path: Parakeet TDT for speech-to-text, an OpenAI-compatible API slot for the language model, and Qwen3-TTS for speech output. Because the LLM slot speaks OpenAI-compatible protocols specifically, it can point at a hosted provider, at Hugging Face's own Inference Providers, or at a local vLLM or llama.cpp server — meaning the exact same pipeline code supports everything from a fully hosted setup to a fully local, fully offline one, just by changing which endpoint the LLM slot points at.

Smart Turn-Taking Is the Detail That Actually Matters for Feel

A voice agent's most noticeable quality problem usually isn't transcription accuracy — it's turn-taking. Cutting a user off mid-thought, or leaving an awkward silence after they've clearly finished speaking, breaks the conversational feel faster than almost anything else. The current release enables a quantized Smart Turn model by default specifically to distinguish a completed turn from a mid-thought pause, running alongside speculative STT and LLM work rather than waiting for silence alone (the traditional Silero-only approach, still available as a fallback). This is exactly the kind of unglamorous latency-and-timing engineering that separates a voice agent that feels natural from one that technically works but feels robotic.

Smart Turn flow: Silero VAD detects a pause, a quantized model scores whether the turn is complete while STT and LLM start speculatively, then processing starts or waits briefly.

It's Not a Research Toy — It Runs in Production

The project states directly that this pipeline runs in production as the conversation backend for thousands of Reachy Mini robots — a real, shipped hardware product, not just a GitHub demo. That's a meaningful credibility signal for anyone evaluating whether an open-source voice pipeline is actually production-hardened: the recent release notes include exactly the kind of reliability work you'd expect from something running at real scale — fixing stuck pipeline units during session teardown, preventing stale teardown signals from releasing the wrong session, and locking around Metal command-buffer operations on Apple Silicon to prevent crashes when speech recognition overlaps other GPU work.

An unmodified OpenAI Realtime client talks to the server over WebSocket or WebRTC, behind which run four swappable stages: VAD, STT, LLM and TTS, each in its own thread connected by queues.

Getting a Server Running

The quickstart is genuinely three commands: pip install speech-to-speech, set an OPENAI_API_KEY, then speech-to-speech serve. That boots an OpenAI Realtime-compatible server at ws://localhost:8765/v1/realtime using Parakeet TDT for STT and Qwen3-TTS for speech output by default, with the LLM slot pointed at whatever OpenAI-compatible endpoint you configured. A second terminal running speech-to-speech talk --url ws://127.0.0.1:8765/v1/realtime connects the packaged microphone/speaker client, or speech-to-speech local composes both in one process over loopback for a single-command trial. Swapping in a fully local LLM is a matter of pointing the same server at a different endpoint — running llama-server with a GGUF model locally and passing its URL via --responses_api_base_url gets you an entirely offline pipeline with no code changes, just different CLI flags. Requires Python 3.10+, and Docker Compose support (with the NVIDIA Container Toolkit for GPU access) starts both a llama.cpp server and the Realtime server together for a two-port, fully self-hosted setup out of the box.

Backend Choices Go Deeper Than the Big Four

Beyond the default STT, LLM, and TTS choices, each stage has more backends available as pip extras than the core pitch suggests. STT options include Whisper through Transformers, Faster Whisper, Lightning Whisper MLX (Apple Silicon specifically), MLX Audio Whisper, and Paraformer through FunASR for Chinese-oriented transcription. TTS options extend to Kokoro-82M, Pocket TTS from Kyutai Labs (with streaming voice cloning across eight preset voices), ChatTTS for English and Chinese, and MMS TTS for broad multilingual coverage. That range matters for language coverage specifically: Parakeet TDT's default language set covers 25 European languages, so a team building for Chinese, Japanese, or another language outside that set needs to deliberately pick a different STT and TTS pairing rather than assume the defaults handle it — the pipeline supports both single-language operation and automatic per-utterance language detection via --language auto.

Benefits of Hugging Face Speech-to-Speech

The project's value comes less from any single model and more from how the pieces are put together.

No Client Rewrite to Leave a Hosted API

Because the server speaks the OpenAI Realtime protocol over WebSocket and WebRTC, an existing client can switch to a self-hosted backend by changing its endpoint. That lowers the cost of trying self-hosting, of running a fallback when the hosted API is unavailable, or of moving specific customers onto infrastructure you control. The decision about where audio is processed becomes reversible rather than a rebuild.

Control Over Every Component

Each stage of the cascade can be replaced through CLI flags. If transcription is weak for an accent, swap the STT model. If the voice doesn't suit the brand, change the TTS backend. If a better LLM appears, point the LLM slot at it. Teams can upgrade the weakest part of the stack without touching the rest, which keeps improvement incremental and low-risk.

Audio That Can Stay on Your Hardware

With local STT, a local LLM server, and local TTS, the whole conversation can run without sending audio or text to a cloud provider. For products where recordings are sensitive, or where deployments must work without internet access, that is a capability a hosted realtime API cannot offer. It also gives teams a clear option when customers ask where their voice data goes.

Visibility Into What the Agent Heard and Said

A cascade produces a transcript and an LLM response at each turn. Engineers can see whether a bad answer came from misheard speech or poor reasoning, apply text-level guardrails before anything is spoken, and log conversations for review. Debugging a voice agent becomes much closer to debugging a text chatbot.

Natural Turn-Taking Out of the Box

The default Smart Turn model distinguishes a finished thought from a pause and starts speculative work while it decides. Teams get conversational timing that would otherwise take weeks of tuning to build, and they can still fall back to silence-based detection when needed. That matters because poor timing is the flaw users notice first.

Speech-to-Speech Use Cases

These are the deployments the pipeline's design suits best, based on what it supports today.

Self-Hosted Fallback for Realtime Voice Products

A product already built on OpenAI's Realtime API depends on one provider's availability, pricing, and voice options. Running a speech-to-speech server alongside gives a second path the same client can use. Traffic can shift during outages, for cost-sensitive customers, or for regions with data requirements. The product keeps one client codebase while gaining a backend it controls.

Offline and On-Premises Voice Assistants

Factories, hospitals, field equipment, and secure facilities may not allow audio to leave the site or may lack reliable connectivity. A fully local stack with llama.cpp or vLLM serving the language model runs entirely on local hardware, including Apple Silicon via the MLX-based backends. Users get a voice interface without any cloud dependency, at the cost of sizing hardware for all three models. The Docker Compose setup with llama.cpp is a quick way to trial this.

Robots and Physical Devices

Hardware products need a conversation backend that handles sessions reliably over long periods. The pipeline's production use behind Reachy Mini robots shows this pattern in practice, with release notes focused on session teardown and GPU stability. Device makers can adopt the same approach and swap models to match their hardware budget.

Multilingual Voice Agents

Products serving Chinese, Japanese, or mixed-language users can pair STT backends such as Paraformer or Whisper variants with multilingual TTS options and per-utterance language detection. Teams choose the combination that fits their audience rather than accepting one provider's language list. Test with native speakers before launch.

Prototyping and Comparing Voice Models

Research and product teams evaluating which STT, LLM, or TTS model to use can hold the rest of the pipeline constant and swap one stage at a time. That makes comparisons fair and fast, and the same pipeline can then go to production with the winning combination. Because each experiment changes one flag rather than a codebase, teams can run many comparisons cheaply and keep a record of which combination performed best on their own recordings, accents, and vocabulary.

Modular Cascade vs Unified Speech-to-Speech Model

Because the name invites confusion, it helps to put the two architectures side by side. Neither is strictly better; they optimize for different things.

FactorModular cascade (HF speech-to-speech)Unified audio-in, audio-out model
ComponentsVAD, STT, LLM, TTS as separate modelsOne model handles listening and speaking
SwappabilityAny stage can be replaced via CLI flagsWhole model is replaced at once
DebuggabilityText transcript and LLM output visible at each stepIntermediate reasoning is mostly opaque
Latency profileSum of stages, reduced by streaming and speculative workPotentially lower, since no text hand-offs
Prosody and emotionTTS rebuilds tone from text, so some nuance is lostCan carry tone from input audio to output
Self-hostingMix local and hosted parts freelyDepends on whether open weights exist
Language coverageSet by the STT/TTS pairing you chooseSet by the model's training data

The practical reading: if you need to inspect transcripts, enforce text-level guardrails, or swap a single weak component, the cascade wins. If your product lives or dies on expressive, emotionally aware back-and-forth and you can accept a single vendor's model, a unified approach may fit better. Our broader comparison of speech-to-speech voice agents covers the unified side in more depth.

Common Speech-to-Speech Deployment Mistakes

The pipeline is easy to start and easy to deploy badly. These are the errors most likely to cause trouble.

Exposing the LLM Proxy to the Internet

The optional remote-LLM proxy doesn't authenticate or throttle requests, and the project says so explicitly. Running it on a public address hands anyone who finds it free access to your LLM endpoint and its bill. Keep it on a trusted network or behind an authenticated gateway, and treat that placement as a launch requirement rather than a later hardening task.

Assuming the Defaults Cover Your Language

Parakeet TDT's default language set is European. Teams building for other languages sometimes deploy the defaults, see poor transcription, and blame the whole pipeline. The fix is choosing a different STT and TTS pairing, which the project supports, but it has to be a deliberate decision made before testing with real users.

Expecting Unified-Model Behaviour From a Cascade

The name suggests a single audio model, and teams sometimes expect the prosody and emotional carry-through such models aim for. A cascade rebuilds tone from text, so some nuance from the user's voice is lost. Judge the pipeline on what a cascade does well, such as visible transcripts and swappable parts, rather than against expectations borrowed from a different architecture.

Undersizing Hardware for a Fully Local Stack

Running STT, an LLM, and TTS locally means holding all three models in memory and serving them fast enough for conversation. A machine that handles each model alone may struggle with all three, producing long pauses that make the agent feel broken. Size hardware against the combined load and test under realistic concurrency.

Skipping Turn-Taking Tests

Teams often benchmark transcription accuracy and model quality but never test interruptions, long pauses, or users who think aloud. Those are the moments users notice most. Script conversations that include hesitation and barge-in, and confirm that Smart Turn settings behave well for your audience before launch.

Speech-to-Speech Best Practices

  • If you already have a client built for OpenAI's Realtime API, this is close to a drop-in self-hosted or local alternative — the protocol compatibility is specifically designed to make that switch require minimal client-side changes.
  • The LLM proxy feature needs deliberate network placement. The project is explicit that its optional remote-LLM proxy doesn't authenticate or throttle requests by design, so any standalone deployment using it needs to sit behind a trusted network or an authenticated gateway — not exposed directly to the internet.
  • For fully offline deployments, the combination of local STT, a local LLM server (llama.cpp or vLLM), and local TTS gives a genuinely complete, non-cloud-dependent voice stack — worth evaluating specifically against self-hosting cost versus a hosted realtime API.
  • Don't assume "speech-to-speech" in the name means single-model latency characteristics. The modular cascade has different latency and quality trade-offs than a true unified audio model — evaluate it on its own architecture's merits, not assumptions carried over from the newer single-model approach.
  • Pick STT and TTS backends for your languages first. The default STT covers European languages; if your users speak Chinese, Japanese, or another language outside that set, choose a matching STT and TTS pairing before tuning anything else, and test --language auto with real mixed-language speech if you need it.
  • Measure turn latency on the target hardware. Time the gap between the end of a user's speech and the first audio from the agent on the machine you will actually deploy, with the models you will actually use. Numbers from a different GPU or a smaller model don't transfer.
  • Pin versions and read release notes before upgrading. The project ships session-handling and stability fixes regularly; upgrade deliberately and re-run your latency and interruption tests afterwards.
  • Check every model licence before commercial use. The pipeline is Apache 2.0, but each STT, LLM, and TTS model carries its own terms.

Practical Takeaway

Hugging Face's speech-to-speech is a well-engineered, production-tested answer to a specific need: an open, self-hostable, protocol-compatible alternative to a hosted realtime voice API, built with the swappable-components philosophy that makes Hugging Face's own ecosystem useful in the first place. For teams building voice-first products that want a self-hosted or fully local option without abandoning OpenAI Realtime API compatibility, it's a strong starting point — just go in clear-eyed about the modular architecture versus the newer single-model speech-to-speech approach, since the name alone doesn't tell you which one you're getting.

Teams building voice AI products — hosted, self-hosted, or fully local — can get hands-on architecture help from Woyce Technologies.

FAQ

What is Hugging Face's speech-to-speech?

It's an open-source, modular voice agent pipeline — Voice Activity Detection, Speech-to-Text, an LLM, and Text-to-Speech — exposed through an OpenAI Realtime API-compatible WebSocket and WebRTC interface, with every component independently swappable. In practice that means you can keep a client built for OpenAI's realtime protocol and run the speech recognition, language model, and voice synthesis on hardware and models you choose, from fully hosted to fully local.

Is this a true single-model speech-to-speech architecture?

No, despite the name. It's a four-stage cascade pipeline, not a unified model that processes audio directly in and out. That's a different architecture than what "speech-to-speech" typically refers to elsewhere in voice AI, though it offers different trade-offs — namely, independently swappable and debuggable components. Each stage runs in its own thread connected by queues, so you can upgrade one model without touching the rest.

Can I run Hugging Face's speech-to-speech fully offline?

Yes — with local STT (like Parakeet TDT), a local LLM server (llama.cpp or vLLM), and local TTS (like Qwen3-TTS), the full pipeline can run without any cloud dependency. The trade-off is hardware: you need enough GPU or Apple Silicon memory to hold all three models at once, and a smaller local LLM will usually give weaker answers than a large hosted one. Test response quality and latency on your target machine before committing.

Does it work with existing OpenAI Realtime API clients?

Yes — it implements the OpenAI Realtime API protocol over both WebSocket and WebRTC, so a client already built against that protocol can typically point at a self-hosted server with minimal changes. Usually that means swapping the endpoint URL. Check any less common events or session options your client relies on, since compatibility covers the core conversation flow rather than guaranteeing every hosted-only feature.

Is this project just a demo, or does it run in production?

It runs in production as the conversation backend for thousands of Reachy Mini robots, a real shipped hardware product — not just a research demo. The release notes reflect that: fixes for stuck pipeline units during session teardown, stale teardown signals, and GPU command-buffer crashes on Apple Silicon. Your own deployment still needs load testing, monitoring, and an authenticated gateway in front of it.

Is Hugging Face's speech-to-speech free to use?

Yes, it's Apache 2.0 licensed and open source, installable via pip, with the underlying models it uses (Parakeet TDT, Qwen3-TTS, and others) available through the Hugging Face Hub. The pipeline itself costs nothing, but running it does: hosted LLM endpoints bill per token, and self-hosting means paying for GPUs and operations. Each model also carries its own license, so check those before commercial use.

What languages does it support?

It depends entirely on which STT and TTS backends you pick rather than the pipeline itself. The default Parakeet TDT covers 25 European languages; swapping to Whisper variants or Paraformer extends that to broader multilingual or Chinese-oriented coverage respectively, and Qwen3-TTS handles multiple output languages with automatic language selection by default.

Does it support WebRTC, or only WebSocket?

Both. WebSocket has been supported from early on; WebRTC support (installed via the webrtc pip extra) was added for SDP negotiation and RTP audio, and both transports share the same underlying pipeline pool, event dispatch, and interruption behavior. Because both follow the OpenAI Realtime protocol, the choice comes down to what your existing client already speaks, and switching transports should not change how the agent behaves in conversation.

Can it feed audio straight to an audio-capable LLM instead of transcribing first?

Yes, with --stt none and the Chat Completions backend pointed at a model that explicitly accepts audio input, such as OpenAI's audio-capable models. This bypass isn't available on the Responses API backend, and the CLI's default model doesn't accept audio, so this mode requires deliberately picking a compatible model and backend combination.

How is turn-taking handled without cutting the user off mid-sentence?

A quantized Smart Turn model, enabled by default, scores whether Silero VAD's detected pause is a completed turn or a mid-thought gap, running speculative STT and LLM work in the background while it decides — complete turns start processing immediately, incomplete ones wait briefly before starting and are held longer before their output is released.

Conclusion

The problem Hugging Face speech-to-speech addresses is lock-in: teams want realtime voice that feels natural, but don't want every audio stream, voice choice, and pricing decision tied to one hosted API. Its answer is pragmatic rather than novel. Keep the OpenAI Realtime protocol on the outside so existing clients keep working, and make every stage on the inside replaceable.

The key insights are that protocol compatibility is the main reason to adopt it, that Smart Turn detection does more for perceived quality than any single model swap, and that the backend menu is wide enough to cover offline, multilingual, and Apple Silicon deployments. The production use behind Reachy Mini is a reasonable signal that the session handling has been tested in real use.

The caveats are just as important. This is a cascade, not a unified audio model, so expect its latency and prosody trade-offs. The optional LLM proxy has no authentication by design and must sit behind a trusted gateway. And a fully local stack is only as good as the smallest model you can afford to run.

A sensible next step is to stand up the default server, point an existing Realtime client at it, and measure turn latency on your own hardware. If you want help designing a self-hosted or hybrid pipeline, explore our voice AI development services.

WT

Woyce Technologies

AI & Engineering Team · Woyce

Woyce Technologies builds AI chatbots, LLM integrations, voice AI, and full-stack web applications for businesses in the US, UK, Europe & APAC. Based in Rajkot, Gujarat.

READY TO BUILD?

Let's build something
that actually works.

Tell us about your project. We'll be honest about whether we're the right fit — and if we are, we move fast.