The Transcript Looked Fine. The Call Wasn't. — Debugging Voice Agents, Arize

Read the talk

The Transcript Looked Fine. The Call Wasn't. — Debugging Voice Agents

Selected presentation frame from The Transcript Looked Fine. The Call Wasn't. — Debugging Voice Agents, Arize at 184 secondsOpen full source frame
The displayed refund exchange moves from order forty to the customer’s correction, “No, fourteen.”

Fuad Ali shows how a refund call can appear to recover in text while still failing in audio and execution—and how linked traces, audio evaluations and replayed experiments make those failures easier to find and fix.

From a talk by Fuad Ali

At a glance

Ideas worth remembering

  • A correction in the transcript does not prove that the action changed: the refund example still processed order forty instead of fourteen.

  • Common audio conventions make provider events queryable; session IDs connect those events into a timeline of speech and tool activity.

  • Evaluate audible behavior and executed outcomes together. Tone, response latency, interruptions and task success answer different questions.

  • Attach evaluations to relevant spans so failures can be filtered, inspected and used to trigger investigations.

  • The proposed self-healing workflow tests development fixes by replaying failed traces and gives a human reviewer comparison evidence before approval.

A readable transcript can hide a broken conversation

Voice agents can fail in ways that their transcripts barely register. A response arrives late. The assistant talks over a correction. A number sounds close enough to another number that the wrong action follows. Fuad Ali, a product manager at Arize AI working on experimentation, opens with this debugging problem: voice needs observations that preserve how a conversation unfolded, alongside what its participants said.

The stakes extend from customer support to drive-through ordering. Ali jokes about getting six hundred cheeseburgers at McDonald’s at two in the morning, but the underlying problem is ordinary: spoken requests lead to actions, and a plausible response does not establish that the action was right. Latency, turn-taking and transcription drift each create a separate way for that interaction to break.

The central example is an AI-generated refund call. The customer asks for order fourteen. The agent announces a refund for order forty. The customer corrects it: “No, fourteen.” A final confirmation follows. Read as consecutive lines, the exchange invites a reassuring interpretation: the agent heard the correction and completed the intended refund.

The call’s timing undermines that interpretation. There are 2.4 seconds of dead air before the first response, the agent talks over the caller, and fourteen is misheard. Its flat, robotic tone adds another audible problem. Text preserves the words of the correction without preserving whether the agent gave that correction room to change the interaction. The refund example therefore starts with a useful distinction: an apparent conversational recovery is something to investigate, rather than proof of a successful action.

0:120:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

Keep audio, transcript and execution together

Debugging the call requires three views in the same place: playable audio, the transcript and the execution trace. The session must include multiple user and assistant turns, so a pause or interruption can be understood in context. Arize AX’s session view brings these observations together with latency and evaluation results.

Each observation answers a different question:

  • Audio and turns: Who spoke, how did they sound, and where did speech overlap or stop?
  • Execution spans: What inputs, outputs and tool calls occurred during the interaction?
  • Metrics and evaluations: How long did the response take, what did audio tokens cost, and which interruption or quality checks failed?

The links between these observations matter as much as their presence. A list of events leaves the engineer to reconstruct the conversation; a linked trace lets the engineer inspect a particular turn alongside the work that produced it.

This carries familiar agent-development practices into voice. Coding agents already have tool calls, handoffs, sub-agents, harnesses and evaluations. Voice adds the need to preserve audible behavior and timing within that same debugging workflow. An interruption is part of the interaction’s behavior, even when every sentence looks reasonable on its own.

3:404:10
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

3:40 · section reference included

Normalize provider events, then stitch them into a session

OpenInference supplies the conventions underneath this view. Its open-source, OpenTelemetry-compliant semantic conventions give generative-AI activity a common representation. For real-time audio, the engineering work includes mapping WebSocket events into recognizable concepts: session lifecycle, audio input, transcripts, conversation items, output, response lifecycle and token counts.

A shared schema makes events from providers such as OpenAI Realtime and Google Gemini Live queryable in the same way. Instead of writing a different parser for each vendor’s event stream, the instrumenter normalizes the supported provider events. That helps both during an incident—when filtering and sorting must find the relevant interaction among many traces—and during a provider change, when engineers need to keep checking the same system behavior. The conventions can also send observations to another backend, including Grafana; this tracing approach does not require Arize’s analysis interface.

Normalization alone does not reconstruct a call. A session ID stitches turns and WebSocket events into one timeline, with tags identifying user and assistant activity. Playing audio directly from that trace makes it possible to inspect who spoke when, where the agent paused, where either participant was cut off, and when a tool ran.

How do vendor-specific events become a conversation an engineer can inspect? The diagram separates two jobs: common conventions make events comparable, while session identity connects them into a call. The resulting view preserves both the conversation timeline and its execution relationships.

A weather-agent preview illustrates the practical payoff. The session view places a weather tool call alongside the audio, so a tool failure can be investigated at the point where it affected the conversation. The engineer can follow the session’s trace tree rather than reconcile disconnected records from different systems.

How it fits togetherFrom provider events to a replayable session

Real-time events from OpenAI Realtime or Google Gemini Live.

Semantic conventions normalize the events; session identity connects turns, audio and tool activity for inspection.

5:406:10
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:40 · section reference included

Check the action, and evaluate the audio itself

Returning to the refund call, the session view changes the diagnosis. The first turn exposes the 2.4-second wait for audio. The third turn exposes the assistant speaking over the caller. The tool call establishes the consequential outcome: the refund was processed for order forty, rather than fourteen. The customer’s correction existed in the conversation, but the executed action still targeted the wrong order.

These observations separate failures that a single success label would blur. The call had a responsiveness problem, a turn-taking problem and an incorrect transaction. Scrubbing the audio beside the tool activity lets an engineer locate each one. It also prevents the agent’s final confirmation from serving as the only evidence that the task succeeded.

Audio evaluations turn those observations into repeatable checks:

  • Tone and sentiment: Classify emotion from the actual audio, including positive, neutral or negative tone. Words alone can miss audible frustration.
  • Response latency: Check time to first audio against the service-level expectations of a customer-service interaction.
  • Conversation accuracy: Flag interruptions and transcription drift, including cases where spoken content and its textual representation diverge.
  • Task success: Check whether the interaction accomplished the requested task, rather than relying on a fluent confirmation.

Audio-capable models can judge the underlying recording. A transcript-only LLM judge has no direct access to the pauses, overlap or vocal delivery that these checks need. Access to audio supplies the missing evidence; the talk does not establish the accuracy of the resulting sentiment classifiers.

Some checks need evidence beyond either audio or text. An “agent as a judge” can use tools and third-party systems to gather evaluation context. Ali’s example is a role-based access-control check: did the agent access the system it was supposed to have access to? This makes evaluation an investigation with its own tool use, rather than only a model’s judgment of a conversation.

8:409:10
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:40 · section reference included

Attach scores to the spans that need investigation

Arize’s skills provide a natural-language route to configuring an audio evaluator. The example request asks Claude to create a sentiment evaluation that catches frustrated tone on audio spans. Underneath that request, the CLI queries spans, analyzes traces, filters for audio-specific spans, creates and deploys the evaluator, and applies it to the system’s traces and spans.

The operational work continues after evaluator creation. Incoming interactions need scoring, and engineers need to query the results at scale. Ali frames querying billions of spans with sub-second latency and running live evaluations as reasons to use a platform rather than build the whole workflow yourself. These are the scale requirements he describes, not a benchmark demonstrated in the recording.

Putting evaluation outputs on the relevant spans keeps a score beside the audio and execution that explain it. Engineers can filter directly for failures and inspect them without searching a separate evaluation system. That placement also gives monitors something concrete to act on: a failed evaluation can trigger an investigation agent to examine misrepresented tool calls, uncovered edge cases or a missing tool.

11:3912:10
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:39 · section reference included

Replay failed interactions against a proposed fix

The ending expands debugging into an observe, evaluate, improve loop. Observations preserve the interaction; evaluations attach judgments to it; fixes change the system; another experiment checks whether the change helped. Arize’s newly released agent experiments are described as “Postman with tracing”: a way to exercise an endpoint while retaining the trace needed to understand its behavior.

Ali then sketches a future workflow for an SRE agent. Live audio is traced and scored. The agent investigates failures, proposes a fix, stands up the changed endpoint in development and replays the failed traces against it. A fresh trace checks whether latency decreased. This complete self-healing workflow is presented as a future vision, rather than a demonstrated autonomous repair in the recording.

The hypothetical repair has a specific cause: tool results were too large, so the agent received too much information and its context became overloaded. The proposed change truncates tool results. Replaying the failed interactions against the development endpoint would then test the claimed latency improvement, and a pull request would present the problem, change and report card for human approval or rejection. This is a separate example from the refund call; it does not establish what caused that call’s 2.4-second delay.

What turns a proposed repair into something a reviewer can assess? The cycle below keeps replay between implementation and approval. The important relationship is that the original failure supplies the test case, while a new trace supplies evidence about the changed endpoint.

The immediate starting point is smaller: add the OpenInference instrumenter, trace audio, check time to first audio and interruptions, and evaluate whether calls accomplish their intended tasks. Tone and latency evaluations are useful first checks. Ali closes by asking for the same engineering rigor applied to coding agents and agents that execute financial actions—even when the voice agent is only taking a cheeseburger order. The task may sound mundane; the system still needs to hear the request and execute the right action.

How it fits togetherThe proposed repair loop keeps failed traces as test cases

Capture incoming audio interactions and score them.

In Ali’s future workflow, investigation leads to a development fix, replay produces comparison evidence, and a person reviews the resulting pull request.

13:3914:09
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:39 · section reference included

Resources

Read the complete timestamped transcript
  1. 0:12

    Folks, thanks for coming in out. I know it's lunchtime, so I know a lot of you folks are trying to eat as well. Appreciate you making the time. And for everyone listening online, thanks for tuning in as well. Uh, my name's Fuad. I'm a product manager here at Arize. Uh, I've been leading up kind of the experimentation side of the house, and one thing that I've been really interested in is voice agents. Uh, quick show of hands here, maybe online too, who here is building a voice agent? Anyone? Got... Okay, we got a couple. Uh, I

  2. 0:42

    think it's a really cool space. I'll be talking a little bit about voice traces, session views of voices, voice evals, and some of the stuff that's new in Arize. But I do wanna kinda mention this isn't a product pitch. This is, you know, really on why voice agents are invisible and how to see them. You can use Arize to do that. You can use another platform. But what I want you guys to take away from this is some of the failure modes and how you can actually start to see that. So we all know

  3. 1:12

    voice agents are growing really, really fast. Have a few tweets up, uh, kinda proving this. GPT Real-Time just dropped. That was literally May seventh. Uh, I have a screenshot from a tweet an hour ago from the LiveKit CEO talking about some of this. Uh, and Vercel on June twenty-ninth, which is only a couple days ago, um, on voice agents in Vercel. So as you can tell, uh, space is exploding fast, and that is gonna start to matter because as this space grows

  4. 1:42

    really fast, it is also one of the hardest to debug categories. So first off, it's exploding. Second off, it's pretty fragile. There's a lot of failure modes you can't really see. Um, and, uh, things like latency, turn-taking, and transcription drift become really, really important. Uh, and the log lies. If you're just looking at transcripts, I'll play a little audio file for you guys in a second, uh, it doesn't really capture what's actually going on under the hood. And just to give you guys some ideas

  5. 2:12

    of use cases, um, there's a lot of them from customer support, I think, which are the more boring ones, down through drive-through ordering. If you've been to White Castle, they have a, a bot called Linda now. Uh, Bojangles has, uh, Bo Linda as well. There's-- McDonald's is doing a lot of these as well, which I think is pretty cool. And personally, I'd be pretty pissed off if I was at McDonald's at two AM and I got six hundred cheeseburgers. So I think I'll show you a little bit about the failure modes and why it actually matters. So,

  6. 2:42

    uh, this is an example where the transcript looks fine, but the call wasn't. And if you read the text logs, you know, user says, "I need a refund for order fourteen." Agent says, "Sure. Start a refund for order forty." User says, "No, fourteen." And it looks like the agent corrected itself, but this wasn't actually the case, and I'm gonna, I'm gonna play the call. It's AI generated. Let's see if it works. I need a refund for order fourteen.

  7. 3:11

    Sure. I've started a refund for order forty. No, fourteen. Your refund is confirmed. So if you notice in the transcript, if you could hear the audio there, uh, it wasn't, you know, the, the confir-confirmation on the, on the forty wasn't actually done by the agent. It kind of interrupted. Uh, and there was a big, uh, kinda latency spike, uh, after the user asked for the initial refund. Um, so w-what actually happened was that two and-- two point four seconds of dead air

  8. 3:40

    before the agent actually responded. The agent talking over the caller. Fourteen misheard, and obviously, the flat robotic tone is, is not a great look either. Um, so these are all audio-specific issues. If you just read this from an LLM output, you would have no idea anything is going on. So what you need to do is you need to see the conversation beyond just the transcript. Um, so that means the audio, the transcript, and the trace all in one place. The session as a whole, so

  9. 4:10

    multiple turns from the user back and forth in the same spa-- in the same trace view. And then metrics on those. So what was the latency like? What was the time to first audio like? Were there any interruptions? Sentiment analysis, et cetera. And these are some of the features we've just released at the beginning of this month in Arize AX. Uh, and really the point of this is the same workflows that you're running now to help your coding agents, uh, which are, you know, pretty complex processes now with harnesses, evals,

  10. 4:40

    uh, handoffs and sub-agents, tool calls. You wanna start to bring some of those best practices into your voice agents as well. So first off, you gotta trace it. You gotta see each individual span. Um, you gotta see the live audio and play it in line from your trace. Really matters to have all this data in the same place so you can start to debug things. Uh, you wanna see inputs and outputs. You wanna see the transcript obviously, but you also want to see, you know, specific metrics like I mentioned, like time to first token, uh, audio token costs,

  11. 5:10

    interruption events. And then finally, you wanna see linkage between these. You don't wanna just see these as a list. You actually wanna see some of the linkage between things, right?

  12. 5:20

    So Arize gives you that one trace view where you can actually play the audio in line, view the transcript, view latency in line, see the conversation turn between the user, uh, and the assistant, um, and see your evals in line as well. And this gives you a really, really powerful tool to debug.

  13. 5:40

    Under the hood, what this looks like is we've actually mapped, uh, the semantic conventions for audio onto OpenInference. So if you guys aren't familiar, Arize was first to market with OpenInference in twenty twenty-three. Uh, it is the first set of semantic conventions that is completely open source that maps, um, kind of what is happening with GenAI to an OTel compliant set of conventions. Uh, and I say this to say it's open source, it's free, and it maps pretty much any of the major foundational models or model

  14. 6:10

    providers. You guys can go rip this and send it to a Grafana backend, completely free to reuse. Not pitching the product here, just explaining some of the real engineering work it takes to map things like WebSocket events to- An easily understandable set of semantic conventions. So you can start to see things like the session lifecycle, the audio input, the transcript, conversation items, the output, uh, and then the response lifecycle and token counts as well. And so, yeah, real engineering work went behind this. It's all open source, uh, and

  15. 6:40

    we've mapped all the foundational models for you. What this allows you to do is you can now be provider-agnostic and have one queryable schema. Uh, so those semantic conventions normalize sort of these disparate model providers, whether it's OpenAI, Realtime, or Google with Gemini Live, into a set of schema that you can query. And if you've ever tried to debug at two AM, uh, off of a customer call, you know how important it is to be able to query

  16. 7:10

    across thousands, millions, even billions of traces that are coming in. You know how important it is to actually get to the part that you need to see, to filter that, to sort by that. All that becomes incredibly important. Um, and, you know, as new models drop, you also want to be able to swap in new providers and make sure that your systems are still functioning. So no bespoke parsing per vendor. It's one auto instrumenter. It captures every single thing that I just mentioned.

  17. 7:40

    And then finally, you want to be able to replay the whole conversation. So I pl- I played the audio for you guys in the beginning of this, but the session ID is essentially what stitches every turn into the timeline, the full call with every WebSocket event that's submitted, uh, with some tags on whether it's the user or the assistant. Um, you can actually play the audio in the UI directly from the trace. No external tooling, no exports, just play it straight from the trace. That allows you to see who spoke when, when the agent paused, where

  18. 8:10

    it got cut off, where the user got cut off, what tools the agent actually went and called, uh, when it did this. And so this really starts the paradigm shift, which is you're not just reading transcripts, you're listening to audio, you're really understanding the failure cases from there with the trace open right next to it. And, um, I'll just show you kind of q- quick sneak preview of what this looks like in Arize. As you can see here, this is an example of like, um, you know, an agent that can call to get, uh, a

  19. 8:40

    specific weather event. You can actually see when the agent kicks off that tool call in line with the rest of the audio. Uh, and this allows you to debug failures of that tool call while you're looking at the session audio. So you're no longer looking at disparate events from distant systems, but you're actually looking at the entire trace tree for everything that happened in that session. So that real session, few issues that we saw, right? Turn one, "Hi, I need help with the order, a refund for order fourteen." We saw, you know, kind of

  20. 9:10

    that time to first audio of two point four seconds of dead air, which really impacts the user experience. We saw in turn three that the agent actually spoke over the caller's interaction, uh, which is not a good look for any assistant, especially one that's trying to be helpful. Um, and then ultimately, if you look at the tool call, you'll be able to also understand that the refund was processed for order forty and not fourteen, um, which is a huge, huge mistake. And actually seeing that live in the tool call is

  21. 9:40

    really, really important. And every single one of these things would be invisible in a text log. But with the session view, you can actually scrub the audio and, and find those failure surfaces within the tools, within, uh, what the agent is doing as well. Uh, the next thing I want to talk about real quick is evals that are specific to audio. So I kind of mentioned a few of them. Um, I think tonality and sentiment analysis is really important. You want to be able to classify user emotion based on the actual audio. Uh, transcript can't do this for

  22. 10:10

    you. It's very, very easy to look at a transcript and see that it, it, you know, s- looks fine, but tonality really matters. Um, positive, neutral, negative, not just a guess. Latency is really, really important in audio, and it's important in a few use cases. Models are getting better. There's, you know, Jetpack architectures, et cetera. But time to first audio SLAs are still really important for customer service use cases. Um, and you want to be able to flag things that are breaking that. You also want to flag things like I mentioned, like

  23. 10:40

    interruptions in accuracy, transcription drift, uh, and the task success. Um, and you want to be able to run these evals directly against the audio with evals that are based on audio models. Um, so, you know, LLM as a judge evals on a transcript are only so good, but actually evaling the underlying audio is really important. And so, um, we have kind of the tooling available in Arize to be able to do that, and we're supporting all the major audio platforms as well as they come out with new models. Um, so again,

  24. 11:10

    evals specific to audio really matter. They're not just transcription-based evals. Um, and as you start to scale up your LLM as a judge evals or even move to agent as a judge that are doing like complex tasks, like for example, doing RBAC checks and making sure the agent actually accessed the system it was supposed to have access to. Um, being able to run these agent as a judge workflows in Arize starts to become really important. And we're kind of seeing the industry move away from, you know, just an LLM as a judge analyzing a

  25. 11:39

    transcript or even analyzing audio, but actually accessing tools, accessing third-party systems, uh, and hydrating some of that evaluation information inside of it. So really easy to get started with evals in Arize through skills. Uh, it is literally as simple as, "Hey Claude, create an audio eval for me for sentiment analysis using GPT Audio and catch frustrated user tone on those audio spans." And Arize will kind of take care of all the rest for you. The CLI under the hood,

  26. 12:10

    um, is able to query spans, do some trace analysis, filter for the spans that are audio specific, actually create that eval, deploy it, uh, and put it against traces and spans in your actual system. Um, and this is part of why Arize exists. Yes, of course, you could do this yourself. Uh, but when you're trying to query, you know, billions of spans, uh, with sub-second query latency, it becomes really hard. And also deploying these live evals as they come in, so you're scoring them, um, while users are actually

  27. 12:40

    interacting with your platform is really important. And then the final thing I'll mention is- You also want these evals on the spans themselves, uh, because then you start to be able to do really complex querying and filtering across your data. So you're not just looking at a failure and looking up an external system. You actually have the eval scores in line with the exact audio that you're looking at. Um, and that's why some of this is really powerful in Arize. Being able to see those eval outputs directly on the spans that are relevant, uh, helps you kind of get a

  28. 13:09

    full understanding of what is going wrong with, uh, the user interaction. Uh, and then you can kick off workflows from that. So evals and monitors attached to those evals can kick off investigation agents that automatically start to dive into, you know, what tool calls am I misrepresenting? Um, where is my agent failing? Is it capturing, you know, maybe a product surface area that I d- hadn't coded for? Is there a new edge case I need to code for in my agent? Maybe I need to expose a new tool. All

  29. 13:39

    those become, um, possible once you have the results directly on the spans. And so I'll, I'll just kind of flag and, and end with this thought. The same loop that you're doing to improve some of these tools like coding agents, we now need to do for voice. And the loop really follows observe, evaluate, and improve. And then from that loop, you can start to get some really, really cool continuous improvement loops within the Arize platform or within another

  30. 14:09

    platform if you choose to use it as well. Um, and the idea is once you observe and you have the data available to you, then you evaluate and you have the scores in line with your observations, then you can start to kick off fixes, whether they're fixes that you make or whether they're fixes that an agent is automatically making on behalf of you, and then you can start to verify and validate those fixes. So we just released agent experiments as well in Arize. Agent experiments, you can think of it kind of like Postman with tracing. So

  31. 14:39

    imagine a world where you have audio streaming in from your users. It's automatically traced. You have live evaluations in the platform that are automatically scoring those interactions. And then you have SRE agents, so not an actual engineer that you hire, but an SRE agent that's looking through the traces, understanding the failure modes, hypothesizing a fix, and then actually standing up a fix for you in dev, replaying those failed traces for you against that fixed endpoint, tracing that to validate

  32. 15:09

    that the latency actually decreases. And you wake up in the morning, you look at a PR, the PR has a description of a problem. "Hey, I found out why the latency was high. Turns out we were not truncating our tool results, and so we were passing in way too much information to the agent. Context got overloaded. I introduced a truncation algorithm. I shipped it. I actually stood it up in dev. I tested against the traces that failed. Here's a report card. Here's how latency decreased." And you wake up in the morning, and all

  33. 15:39

    this is ready for you, and you just click approve or deny on the PR. That is the future that I think we want to see at Arize, where self-healing software and continuous improvement are realities for everyone who's building agents. And I think making that possible for voice is really, really important. So start seeing your agents. Add the OpenInference instrumenter. Open source, free to use, OTel compliant. You can send it to any back end. Um, so use Arize, don't use Arize, start tracing your audio.

  34. 16:09

    The second part of this is actually starting to do that analysis, and we would recommend Arize to do that. Sign up for a free account. Uh, I'll have a code up at the end of my presentation for you guys to sign up for, uh, a free year of Pro. Um, but start checking time to first audio. Start checking your interruptions. And start evaluating, um, how these interactions are actually going and whether they're accomplishing your use cases. Whether it's a tone or latency eval, um, try using the Arize skills to actually create that eval. Uh, we have some docs set up for you if you

  35. 16:39

    want to follow along, a bunch of examples and cookbooks ready, uh, and bring voice into the same rigor as the rest of your stack. You wouldn't be okay with this happening at the coding agent level. You wouldn't be okay with it happening, uh, for an agent that's able to execute stocks on Robinhood. You shouldn't be okay with it, uh, though the stakes might seem a little lower, for a McDonald's agent that's ordering you a cheeseburger either. Um, and so I'll end with that. Please, please keep building. Please connect with me also on Twitter. I still call it

  36. 17:09

    Twitter, not X. Uh, on Twitter or LinkedIn. Would love to kind of talk and hear about some of the use cases you guys are building towards so we can add more evals to our library and make sure this is useful for you. And if you found this talk useful and you want to start investigating your voice agents, there is a free code up on the screen. It gives you a year of Arize Pro for free. Uh, it's AIEWF2026. Um, and you can get started and start tracing your voice agents. So thank you very much. My name is Fuad Ali. I'll be available for questions for a couple

  37. 17:39

    minutes after, but really appreciate you guys listening. Cheers.