Build a LiveKit voice agent with Speechify TTS

LiveKit's Python plugin adds Speechify TTS to an AgentSession in one line. It streams each sentence: 82 ms median first audio from US East on 28 Sep 2026.

Engineering and research · SpeechifyAI
6 min read

You add Speechify TTS to a LiveKit voice agent with one line: install LiveKit’s Python plugin and pass speechify.TTS(...) as the tts of your AgentSession. The plugin streams, so audio starts while the rest of the sentence is still being synthesized. On 28 September 2026 we measured 82 ms from handing it a sentence to its first audio frame, at the median, from a server in US East.

This guide wires a complete agent with Deepgram for speech-to-text, OpenAI for the language model and speechify.TTS(voice_id="dominic_32", model="simba-3.2") for speech, then shows you how to measure first audio from where your own agent runs. It is Python only, because LiveKit publishes the Speechify plugin for its Python framework. You need a Speechify API key, a Deepgram key and an OpenAI key. Get a Speechify key here if you do not have one yet.

How the plugin streams

A LiveKit voice agent is a loop: the user speaks, STT turns it into text, the LLM writes a reply, and TTS speaks it. Each stage is a provider object on the AgentSession, and LiveKit runs the conversation around them.

LiveKit AgentSession

Speechify is the TTS provider in the voice pipeline

LiveKit handles the realtime room and turn loop. The Python agent passes Deepgram, OpenAI, and Speechify providers into one session.

  1. STTDeepgram listensUser audio becomes text inside the LiveKit session.deepgram.STT(model="nova-3")
  2. LLMOpenAI repliesThe agent turns the transcript into a short conversational response.openai.LLM(...)
  3. TTSSpeechify speaksThe official LiveKit plugin calls Speechify with the voice and model you choose.speechify.TTS(...)

The Speechify key stays in the Python environment as SPEECHIFY_API_KEY. The browser or room participant never receives it.

The Speechify stage splits the LLM’s reply into sentences as the tokens arrive. Each sentence becomes one request to Speechify’s /v1/audio/stream/with-timestamps endpoint, and the plugin hands LiveKit each audio chunk the moment it lands, with the word timestamps for that chunk riding on the same stream. So the first word plays after the first complete sentence and one streaming request, not after the whole reply.

Two things follow from that shape. First, the TTS number below is only one part of what a caller waits for: the turn also includes end-of-speech detection, the transcript, and the LLM’s first sentence. Second, the word timestamps are free: the plugin reports them to LiveKit, so use_tts_aligned_transcript=True on the session syncs the transcript to the audio word by word. Word-level timestamps in a LiveKit voice agent builds captions on them.

Install the Python packages

Create a fresh Python environment and install LiveKit Agents, the Speechify plugin, and the STT and LLM plugins this agent uses:

python3 -m venv .venv
source .venv/bin/activate
pip install "livekit-agents[codecs]>=1.8.3" \
  "livekit-plugins-speechify>=1.8.3" \
  "livekit-plugins-deepgram>=1.8.3" \
  "livekit-plugins-openai>=1.8.3" \
  python-dotenv

livekit-plugins-speechify is authored and published by LiveKit. On 28 September 2026 that command resolved livekit-agents 1.8.3 and livekit-plugins-speechify 1.8.3, and every snippet in this guide was checked against those versions.

Configure the environment

Put your keys in .env:

SPEECHIFY_API_KEY=your_speechify_api_key
DEEPGRAM_API_KEY=your_deepgram_api_key
OPENAI_API_KEY=your_openai_api_key

# Only needed when you run against a real LiveKit room (dev or start):
LIVEKIT_URL=wss://your-project.livekit.cloud
LIVEKIT_API_KEY=your_livekit_api_key
LIVEKIT_API_SECRET=your_livekit_api_secret

The plugin reads SPEECHIFY_API_KEY unless you pass api_key= to speechify.TTS(...) yourself. Keep it server-side in the Python agent, and never hand it to a browser participant.

Wire Speechify into AgentSession

Here is the full agent, also in the runnable livekit-agent-speechify-python demo:

from dotenv import load_dotenv

from livekit import agents
from livekit.agents import Agent, AgentServer, AgentSession
from livekit.plugins import deepgram, openai, speechify

load_dotenv()


class Assistant(Agent):
    def __init__(self) -> None:
        super().__init__(
            instructions=(
                "You are a helpful voice assistant speaking with a Speechify voice. "
                "Keep replies short, clear, and conversational."
            )
        )


server = AgentServer()


@server.rtc_session(agent_name="speechify-tts-demo")
async def entrypoint(ctx: agents.JobContext) -> None:
    session = AgentSession(
        stt=deepgram.STT(model="nova-3"),
        llm=openai.LLM(model="gpt-4o-mini"),
        tts=speechify.TTS(voice_id="dominic_32", model="simba-3.2"),
    )

    await session.start(room=ctx.room, agent=Assistant())
    await session.generate_reply(
        instructions=(
            "Greet the user and mention that your voice "
            "is powered by Speechify TTS."
        )
    )


if __name__ == "__main__":
    agents.cli.run_app(server)

Save it as agent.py. The tts= line is the whole integration. simba-3.2 is the streaming model for English; for German, Spanish, French, Italian or Brazilian Portuguese, pass a voice for that language with model="simba-3.0".

Measure first audio before you open a room

The number a caller feels depends on where your agent runs, so measure it there before you tune anything else. This script needs only SPEECHIFY_API_KEY, no LiveKit room and no microphone. It hands the plugin one sentence at a time, the way AgentSession does, and times the first audio frame that comes back:

import asyncio
import statistics
import time

from dotenv import load_dotenv
from livekit.plugins import speechify


async def first_audio_ms(tts: speechify.TTS, text: str) -> float:
    stream = tts.stream()
    start = time.perf_counter()
    stream.push_text(text)
    stream.end_input()
    first = None
    async for _ in stream:
        if first is None:
            first = (time.perf_counter() - start) * 1000
    await stream.aclose()
    return first


async def main() -> None:
    load_dotenv()
    tts = speechify.TTS(voice_id="dominic_32", model="simba-3.2")
    await first_audio_ms(tts, "Warming up.")  # not counted

    times = []
    for n in range(20):
        text = f"Your order {1000 + n} shipped today and arrives Thursday."
        times.append(await first_audio_ms(tts, text))

    deciles = statistics.quantiles(times, n=10)
    print(f"first audio p50 {deciles[4]:.0f} ms, p90 {deciles[8]:.0f} ms")
    await tts.aclose()


asyncio.run(main())

Run from our development server in Zurich on 28 September 2026, it printed:

first audio p50 169 ms, p90 185 ms

Zurich is a long way from the API, which serves from US East, so most of that is the Atlantic. Run it from the region your agent is deployed in.

What we measured

We ran the same measurement from two Google Cloud regions on 28 September 2026, with livekit-plugins-speechify 1.8.3, simba-3.2 and dominic_32. Every request was a different sentence of about 90 characters, each cell is 25 requests, and the cells ran interleaved so that time of day hit them alike. “First reply” is a new TTS object, as in the first turn of a new call; “later turn” reuses one. The last column is a bare POST /v1/audio/stream on an open connection, the floor for anything built on the API.

First audio, p50 / p90First replyLater turnBare API
us-east4, Virginia82 / 106 ms79 / 104 ms63 / 71 ms
europe-west6, Zurich187 / 202 ms193 / 203 ms162 / 180 ms
Time from handing the plugin a sentence to its first audio frame, 25 requests per cell, 28 September 2026.

From US East the plugin’s first audio is under 110 ms even at the 90th percentile. Most of the 15 to 30 ms between the plugin and the bare API is connection setup, which version 1.8.3 pays on each sentence.

One thing to avoid: tts.synthesize() in 1.8.3 makes one non-streamed request for the whole text, so its first audio waited 950 ms at the median from us-east4 for the same sentences. AgentSession calls stream(), so an agent never takes that path; only call synthesize() yourself for audio you render ahead of time.

Run the LiveKit agent

With .env filled in, start local console mode:

python agent.py console

Console mode runs against your microphone and speakers with no room, which keeps the loop local while you tune prompts, voice choice and interruption behavior.

To connect the agent to a real LiveKit project, use development mode:

python agent.py dev

For production, use LiveKit’s deployment flow or run python agent.py start. At that point you are operating a realtime service: rotate keys, watch model errors, and measure the whole turn from the caller’s last word to your agent’s first audible word.

Choosing the Speechify voice

This demo uses dominic_32 with simba-3.2. simba-3.2 serves every English voice in the catalogue, cloned voices included, and the _32 voices are its built-in examples. Pass another voice_id whenever you want a different speaker, and check the voice’s models list on the /v1/voices endpoint to confirm the pairing.

For how the first-audio numbers compare across providers, read how to choose a low-latency TTS API for voice agents in 2026. For the streaming API underneath the plugin, read Streaming TTS in Python with Speechify. If you use Deepgram’s Voice Agent instead of LiveKit, Add a better voice to Deepgram’s Voice Agent with Speechify is the same idea on that stack. The Speechify docs and LiveKit’s Speechify plugin guide cover the rest of the surface.

FAQ

Common questions

Does the LiveKit Speechify plugin stream audio?
Yes. With simba-3.2 or simba-3.0, the plugin's stream() sends each sentence to Speechify's /v1/audio/stream/with-timestamps endpoint and hands LiveKit audio frames as they arrive, with word timestamps on the same stream. The agent's reply is split into sentences first, so it is one streaming request per sentence.
How fast is first audio through the LiveKit Speechify plugin?
On 28 September 2026 we measured 82 ms at the median and 106 ms at the 90th percentile from a server in Google Cloud us-east4, from handing the plugin a sentence to the first audio frame out of livekit-plugins-speechify 1.8.3 with simba-3.2, over 25 requests. From Zurich it was 187 ms at the median, because the API serves from US East.
Is there a Node.js LiveKit Speechify plugin?
Not today. LiveKit lists Speechify TTS for its Python Agents framework only, so this guide uses the PyPI package livekit-plugins-speechify, which LiveKit publishes.
Which environment variable does the LiveKit Speechify plugin read?
SPEECHIFY_API_KEY, unless you pass api_key= to speechify.TTS(...) yourself. Keep the key server-side in the Python agent and never hand it to a browser participant.
Which keys do I need to run the full LiveKit agent?
SPEECHIFY_API_KEY for TTS, DEEPGRAM_API_KEY for speech-to-text and OPENAI_API_KEY for the LLM. LIVEKIT_URL, LIVEKIT_API_KEY and LIVEKIT_API_SECRET are only needed when you run against a real LiveKit room with the dev or start commands.

Privacy preferences

Choose what we may store on this device. You can change this at any time from the footer.

Strictly necessary

Sign-in, security, load balancing, and remembering your privacy choices. These cannot be switched off.

Always on

Analytics

How the site is used in aggregate - which pages get read, where people get stuck - so we can improve it.

Marketing

Measures which campaigns bring people here, and lets us show relevant ads on other platforms.