Reasoning models are good at the hard parts of a call: rescheduling three appointments around a caller’s constraints, troubleshooting a boiler, negotiating a payment plan. They are also slow. Four seconds of silence on a phone line feels like the call dropped.
The Talker and Thinker pattern splits the job between two models. A fast model, the Talker, answers every turn so the caller always hears something within a second. When a question needs real reasoning, it hands that question to a slower model, the Thinker, and keeps the conversation going. When the Thinker’s plan is ready, it lands in the Talker’s context and shapes what it says next.
Use it when
The answer needs real reasoning but the caller cannot be left in silence. Scheduling under constraints, diagnosis, quoting a job with several variables, anything where a fast model alone would guess.
Avoid it when
Answers are simple lookups. “What time do you close?” does not need a reasoning model, and running one anyway is pure cost.
Thinker or Worker?
This pattern is easy to confuse with Background Worker. The difference: a Thinker shapes what the agent says. A Worker does a job, like fetching a quote, and reports a result.
Latency
Time to first word stays at fast-model speed. The considered answer lands one turn later. Design the Talker’s instructions around that: it should acknowledge, ask a useful clarifying question while it waits, and never improvise the answer the Thinker is working on.
Build it on LiveKit
LiveKit calls this subagent delegation. The Talker gets a tool that sends the question to a reasoning model. Calling ctx.update() at the start is what makes the tool non-blocking: the Talker voices an acknowledgement and the conversation continues while the Thinker works.
from livekit.agents import Agent, RunContext, function_tool, inference
from livekit.agents.llm import ToolFlag
thinker = inference.LLM(model="google/gemini-3.1-pro") # any strong reasoning model
class Scheduler(Agent):
def __init__(self) -> None:
super().__init__(
instructions=(
"Answer simple questions yourself. For rescheduling with constraints, "
"call plan_schedule, say in one sentence that you are checking, and keep "
"chatting. Never guess the plan yourself while it is running."
),
)
@function_tool(flags=ToolFlag.CANCELLABLE)
async def plan_schedule(self, ctx: RunContext, constraints: str) -> str:
"""Work out a schedule that satisfies the caller's constraints."""
await ctx.update("Started working out the schedule.")
return await plan_with(thinker, constraints, await calendar.week()) # your prompt + one LLM call
CANCELLABLE lets the Talker drop the plan if the caller moves on, so abandoned reasoning stops costing tokens. LiveKit’s subagent delegation guide has the full example. The pattern’s name comes from Google DeepMind’s 2024 Talker-Reasoner paper.
How it combines
Put a Thinker beside the agent that needs it, not beside every agent. In a Supervisor, that is usually one task: the scheduling task gets a Thinker, the address task does not. Add model routing as a capability so the Talker itself stays on the fastest model you can get away with.
Common questions
How do I use a reasoning model in a voice agent without long silences?
Let a fast model keep talking and send the hard question to the reasoning model through a non-blocking tool. In LiveKit this is subagent delegation: the tool calls ctx.update() early so the conversation continues.
What is the difference between a Thinker and a Background Worker?
A Thinker shapes what the agent says next. A Background Worker does a job, like fetching a quote, and reports a result.
Where does the name Talker and Thinker come from?
It follows Google DeepMind's 2024 Talker-Reasoner paper, which splits an agent into a fast conversational part and a slower planning part.
Sources and further reading
Mahimai Raja
Mahimai builds production voice agents on LiveKit for small businesses and maintains the open-source livekit-starter. He is building ShipVoice, the voice AI platform for agencies. Find him on X or at [email protected].