The agent on the phone is busy. It has to listen, answer within a second, and call the right tool. It is not well placed to notice that the caller has sounded more annoyed with every turn, or that it forgot to offer the dessert special three turns ago.
An Observer is a second model that does nothing but watch. It reads the transcript as the call happens, and every few turns it decides whether the speaking agent needs a nudge. If it does, it slips a short hint into the agent’s context. It never speaks to the caller.
Use it when
Quality matters across long calls. Sales, support, collections, anything where frustration, drift from the script, a missed upsell or a policy slip is costly. The Observer is the pattern that turns “the agent is fine on average” into “the agent catches the bad call while it is still happening”.
For agencies it is also a reporting tool. The flags an Observer raises, like frustration detected or upsell offered, are exactly the numbers a client wants to see on a monthly report.
Avoid it when
Calls are short and scripted. On a thirty-second hours-and-directions call there is nothing for it to catch, and it is pure cost.
Latency
None on the turn. The Observer runs off to the side and never sits between the caller and the speaking agent. It does cost an extra model call every few turns, so run it on a small, cheap model and only every N caller turns.
Build it on LiveKit
Listen to the session’s conversation events, review the transcript in the background, and update the agent’s instructions when the review finds something.
import asyncio
from livekit.agents import AgentSession, ConversationItemAddedEvent
REVIEW_EVERY = 3
def attach_observer(session: AgentSession, agent, base_instructions: str) -> None:
turns = 0
async def review() -> None:
hint = await observer_llm.review(session.history) # small, cheap model
if hint:
await agent.update_instructions(f"{base_instructions}\n\nCoach note: {hint}")
@session.on("conversation_item_added")
def on_item(ev: ConversationItemAddedEvent) -> None:
nonlocal turns
if ev.item.role == "user":
turns += 1
if turns % REVIEW_EVERY == 0:
asyncio.create_task(review())
ShipVoice Pro ships an observer that checks every three caller turns. Keep hints short and specific (“the caller asked twice about parking, answer it directly”), and replace the previous hint rather than stacking them.
How it combines
An Observer works next to any holding pattern. It is most valuable next to a Supervisor or Relay, where no single task or phase sees the whole call. Its flags feed naturally into the wrap-up half of Bookends, and a high frustration score is a good trigger for a Warm Handover.
Common questions
Does an observer model add latency to a voice agent?
No. It runs beside the conversation and never sits in the turn. It does add the cost of an extra model call every few turns, so use a small model.
How does an observer change what the agent says?
It updates the speaking agent's instructions with a short hint, such as "the caller asked twice about parking, answer it directly". Replace the previous hint instead of stacking them.
How often should the observer run?
Every few caller turns is enough. ShipVoice Pro checks every three caller turns.
Sources and further reading
Mahimai Raja
Mahimai builds production voice agents on LiveKit for small businesses and maintains the open-source livekit-starter. He is building ShipVoice, the voice AI platform for agencies. Find him on X or at [email protected].