Skip to content
ShipVoice

Talker and Thinker: answer fast, reason slow

A fast model answers every turn while a slower reasoning model plans alongside it.

Mahimai Raja 2 min read What runs beside the call
A small fast red node above a larger slow ring, representing a fast talker model and a slower thinker model.
A fast talker model answers the caller. It passes the hard question down to a slow thinker model, which sends a plan back up for the next turn.
A fast talker model answers the caller. It passes the hard question down to a slow thinker model, which sends a plan back up for the next turn. Download the Excalidraw file

Reasoning models are good at the hard parts of a call: rescheduling three appointments around a caller’s constraints, troubleshooting a boiler, negotiating a payment plan. They are also slow. Four seconds of silence on a phone line feels like the call dropped.

The Talker and Thinker pattern splits the job between two models. A fast model, the Talker, answers every turn so the caller always hears something within a second. When a question needs real reasoning, it hands that question to a slower model, the Thinker, and keeps the conversation going. When the Thinker’s plan is ready, it lands in the Talker’s context and shapes what it says next.

Use it when

The answer needs real reasoning but the caller cannot be left in silence. Scheduling under constraints, diagnosis, quoting a job with several variables, anything where a fast model alone would guess.

Avoid it when

Answers are simple lookups. “What time do you close?” does not need a reasoning model, and running one anyway is pure cost.

Thinker or Worker?

This pattern is easy to confuse with Background Worker. The difference: a Thinker shapes what the agent says. A Worker does a job, like fetching a quote, and reports a result.

Latency

Time to first word stays at fast-model speed. The considered answer lands one turn later. Design the Talker’s instructions around that: it should acknowledge, ask a useful clarifying question while it waits, and never improvise the answer the Thinker is working on.

Build it on LiveKit

LiveKit calls this subagent delegation. The Talker gets a tool that sends the question to a reasoning model. Calling ctx.update() at the start is what makes the tool non-blocking: the Talker voices an acknowledgement and the conversation continues while the Thinker works.

from livekit.agents import Agent, RunContext, function_tool, inference
from livekit.agents.llm import ToolFlag

thinker = inference.LLM(model="google/gemini-3.1-pro")  # any strong reasoning model


class Scheduler(Agent):
    def __init__(self) -> None:
        super().__init__(
            instructions=(
                "Answer simple questions yourself. For rescheduling with constraints, "
                "call plan_schedule, say in one sentence that you are checking, and keep "
                "chatting. Never guess the plan yourself while it is running."
            ),
        )

    @function_tool(flags=ToolFlag.CANCELLABLE)
    async def plan_schedule(self, ctx: RunContext, constraints: str) -> str:
        """Work out a schedule that satisfies the caller's constraints."""
        await ctx.update("Started working out the schedule.")
        return await plan_with(thinker, constraints, await calendar.week())  # your prompt + one LLM call

CANCELLABLE lets the Talker drop the plan if the caller moves on, so abandoned reasoning stops costing tokens. LiveKit’s subagent delegation guide has the full example. The pattern’s name comes from Google DeepMind’s 2024 Talker-Reasoner paper.

How it combines

Put a Thinker beside the agent that needs it, not beside every agent. In a Supervisor, that is usually one task: the scheduling task gets a Thinker, the address task does not. Add model routing as a capability so the Talker itself stays on the fastest model you can get away with.

Common questions

How do I use a reasoning model in a voice agent without long silences?

Let a fast model keep talking and send the hard question to the reasoning model through a non-blocking tool. In LiveKit this is subagent delegation: the tool calls ctx.update() early so the conversation continues.

What is the difference between a Thinker and a Background Worker?

A Thinker shapes what the agent says next. A Background Worker does a job, like fetching a quote, and reports a result.

Where does the name Talker and Thinker come from?

It follows Google DeepMind's 2024 Talker-Reasoner paper, which splits an agent into a fast conversational part and a slower planning part.

Sources and further reading

Mahimai Raja

Mahimai builds production voice agents on LiveKit for small businesses and maintains the open-source livekit-starter. He is building ShipVoice, the voice AI platform for agencies. Find him on X or at [email protected].