Act, Confirm, or Stop? Smarter behavior for AI assistants, wearables & robots — Amit Desai, Roku
Read the talk
Act, Confirm, or Stop? Choosing Assistant Behavior Under Uncertainty
Amit Desai explains how assistants can reduce the effort users spend recovering from errors by optimizing when to act, ask for confirmation, or stop—even when recognition accuracy stays unchanged.
From a talk by Amit Desai
At a glance
Ideas worth remembering
Interpretation accuracy and behavior under uncertainty are separate controls. The example keeps 790 of 1,000 hypotheses correct while improving which ones the assistant acts on.
Choose confidence thresholds by minimizing weighted outcome costs. Desai's illustrative recovery costs are 10 seconds for wrong playback, four for stopping, two for affirming a confirmation, and six for correcting it.
The example selects a 43 percent act-or-stop threshold rather than the intuitive 65 percent. With confirmation, the screen distinguishes CURRENT boundaries of 30 and 60 percent from OPTIMIZED boundaries of 41 and 49 percent. The optimum depends on the costs and confidence distributions.
The recorded screen resolves the numeric ambiguity: 1,464 is the current confirmation configuration's total cost; 1,260 is the optimized total. The recap displays a reduction from 1.274 to 1.260 OUCH points per turn with optimized confirmation, at unchanged accuracy. Segment 341's retained caption, “employed that then we would go to 1464.”, ambiguously associates the current cost with optimization; the visual evidence resolves the configuration, while the cause of the spoken/caption discrepancy remains uncertain.
Costs must reflect the interface and action. Television choices selected with a remote can reduce confirmation effort, while a wrong channel launch can disrupt the current state and increase recovery cost. Desai proposes carrying this principle into learned real-time decisions.
Voice errors become more consequential when assistants take actions
After a brief greeting, Amit Desai introduces his experience building voice interfaces across Alexa, Roku, and his own startups. His perspective combines voice user interface design with technical approaches: understanding how people interact with a system matters alongside understanding how the system interprets speech. He sees that combination as especially relevant when new technology changes the human interface.
Voice offers a natural way to interact, but it remains error prone. Desai argues that the consequences of errors grow as assistants move from giving answers to taking physical or digital actions. A robot throwing a watch away with the trash illustrates the difference: the mistake is much more consequential than playing the wrong song. Improving the interpretation of a request therefore addresses only part of the user experience.
He separates two ways to improve satisfaction. The first is increasing technical accuracy. The second is changing the system's decision under uncertainty: what it does with a hypothesis that might be wrong. Desai treats this decision as a separate control that teams can improve without changing recognition accuracy. He introduces a simple smart speaker example to develop the method before considering other devices.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Hold accuracy at 79 percent
The example speaker accepts music requests and immediately plays a song. Each interpretation is either correct or incorrect. Desai proposes collecting 1,000 spoken requests, observing their inputs and outputs, and labeling the results. In this illustrative dataset, 790 requests produce the correct song and 210 produce the wrong song, giving 79 percent accuracy.
A conventional improvement effort would try to reduce those 210 errors. In a cascaded voice system, errors can arise in wake word detection, automatic speech recognition, natural language understanding, intent classification, entity extraction, or voice activity detection. Work at any of those layers may improve end-to-end accuracy. Desai instead freezes the example at 79 percent and asks how much better the experience can become through decisions made after the system forms its interpretation.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Add a stop decision, then choose when to use it
The original system always acts: a request leads directly to playback. Desai adds the option to reject the hypothesis and stop, asking the user to repeat instead of playing a potentially wrong song. This creates a quantitative design question. Having a stop behavior is easy to describe; deciding exactly when to invoke it requires a rule.
For the example, each of the 1,000 labeled hypotheses receives a confidence score between zero and one. Desai assumes that this score is reasonably calibrated and plots separate confidence distributions for correct and incorrect hypotheses. This is a deliberate simplification: a cascaded system may produce multiple confidence scores across its layers, rather than one score that summarizes the entire interpretation.
A threshold t divides the decisions: stop when confidence c is below t; otherwise play. Moving t changes which correct and incorrect hypotheses reach playback. Desai offers 65 percent as a plausible intuitive choice, then challenges that intuition. A confidence value that feels sufficiently reassuring does not yet say how much effort the resulting mistakes will impose on users.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Measure the effort of recovering from each outcome
The threshold produces two kinds of undesirable outcome. Stopping delays the song the user wants, while acting on an incorrect hypothesis plays the wrong song. These outcomes are not equally costly. In Desai's example, a request intended for Prince produces a song by Chris Brown. The user must listen long enough to recognize the mistake, speak over the music to stop playback, and then request the desired song again. A rejection lets the user retry without first undoing playback.
Desai quantifies this difference with a heuristic: the additional seconds needed to return to success, meaning playback of the intended song. He assigns 10 seconds to a wrong song and four seconds to a stop, including repeating the request and the extra latency. These are illustrative assignments rather than established measurements. He leaves open how best to quantify relative badness, but makes the chosen proxy concrete enough to optimize.
Threshold selection now becomes cost minimization. For a candidate t, count the wrong hypotheses that are acted on and all hypotheses that are stopped. The total is C(t) = 10 × wrong acts(t) + 4 × stops(t). A correct immediate playback adds no recovery cost in this formulation. Raising the threshold can prevent wrong playback, but it also rejects more requests, including requests the system understood correctly. The weighted sum expresses that tradeoff. Desai calls the approach the Outcome User Cost Heuristic, or OUCH: the goal is to minimize the pain of the interaction.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Optimize the threshold instead of guessing
Desai presents an interactive graph of the 1,000 requests, showing how total cost changes with the threshold. With immediate playback for every request, the total cost is 2,100, or 2.1 OUCH points per turn. Introducing a stop threshold at the intuitive 65 percent reduces the total to 1,904, about 1.9 per turn. That choice helps, but it does not minimize the objective.
The demonstration identifies 43 percent as the cost-minimizing threshold for these data and assigned costs. Desai moves the control to that point and reports approximately 1.27 OUCH points per turn. He emphasizes that the system's interpretation accuracy has not changed. The improvement comes from deciding which hypotheses to execute and which to reject. The optimum belongs to this cost function and confidence distribution; it is an example of a decision method rather than a general confidence requirement for assistants.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Give uncertain requests a confirmation region
Desai next adds confirmation. The speaker states its proposed song before playing, allowing the user to affirm or correct the interpretation. This creates three behaviors separated by two thresholds: stop at low confidence, confirm in the middle, and act at high confidence. The optimization must now choose both boundaries rather than one.
Confirmation also imposes effort, even when the proposed interpretation is correct. Listening and saying yes delays playback; rejecting the proposal and restating the request takes longer. Desai assigns two seconds to affirmation and six seconds to correction. Keeping the previous costs for wrong action and stopping, the objective becomes C(t1, t2) = 10 × wrong acts + 4 × stops + 2 × affirmations + 6 × corrections. Each count depends on where the two thresholds place the labeled requests.
The graph becomes a two-dimensional heat map of cost across threshold pairs. The recorded screen at 920 seconds explicitly separates CURRENT thresholds of 30 percent and 60 percent, with total cost 1,464, from OPTIMIZED thresholds of 41 percent and 49 percent, with total cost 1,260. At 928 seconds, the current configuration's displayed breakdown shows 601 acts: 596 right and five wrong, costing 50; 277 confirmations: 184 yes and 93 no, costing 926; and 122 stops, costing 488. Those displayed costs total 1,464 and belong to the current 30 percent/60 percent configuration. The original caption for segment 341 remains “employed that then we would go to 1464.” Following the captions naming the optimal boundaries, that wording ambiguously attaches 1,464 to the optimized result. The screen establishes which configuration each number describes, but without an audio check the discrepancy cannot be attributed specifically to transcription error or to the spoken statement.
Desai then considers increasing the wrong-action cost to 20 because recovery could be more irritating or take longer. That change moves the optimum. Both the relative costs of outcomes and the confidence distributions determine the preferred behavior; adding confirmation does not remove the need to evaluate those inputs.
The recap slide at 980 seconds displays average costs of 2.100 OUCH points per turn for always acting, 1.904 with the guessed stop threshold, 1.274 with optimized act-or-stop behavior, and 1.260 with optimized act-stop-confirm behavior. The original captions summarize this as “2.1 act and stop 1.9 then” and “1.27 then 1.26.” The slide supplies the more precise values and independently confirms a modest additional reduction from 1.274 to 1.260 when confirmation is optimized. The demonstrated current total of 1,464 and the optimized total of 1,260 therefore describe different threshold settings, rather than conflicting results for one setting. Accuracy remains unchanged throughout this simplified example.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Adapt decisions to the interface and the action
Desai distinguishes the simplified threshold exercise from his proposal for real systems. He expects a learned decision model operating in real time, rather than only one or two thresholds selected offline. The principle remains choosing behavior according to user outcome cost. He proposes that analogous decisions can apply across voice interfaces, although the devices and interactions introduce differences.
A television assistant makes those differences concrete. For a request to open a channel, confirmation can display multiple choices, including ABC News Live, instead of speaking one proposed interpretation and waiting for a verbal response. The user can select a choice with the remote control. Desai argues that this visual interaction can make confirmation less painful, changing its assigned cost and therefore the decision the system should prefer.
The wrong action can become more costly at the same time. Launching the wrong channel may kick the user out of the current state, creating more recovery work. A different modality therefore changes both the available confirmation behavior and the consequences of acting incorrectly. The method carries over by changing its variables and costs to reflect the actual interaction, rather than assuming the smart speaker's assignments apply unchanged.
Desai closes by returning to assistants that take physical and digital actions, including making phone calls and sending emails. He argues that relying on accuracy improvements alone becomes increasingly difficult as those actions raise the consequences of mistakes. Smarter conversational behavior under uncertainty provides another way to reach an acceptable experience. His final objective is to minimize OUCH—the pain users experience—because an interface that remains frustrating can become a bottleneck even while other parts of the technology improve. He ends by thanking the audience and offering to stay for questions.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
Hi everyone. How's it going? Hey
- 0:15
Patricia, how are you?
- 0:17
>> Uh so last presentation of the day, so
- 0:20
let's make it count. Um,
- 0:23
all right. Let's, uh, let me start with
- 0:25
a little bit of background on myself.
- 0:28
And, um, my background, I'm a voice
- 0:33
subject matter expert. I've been working
- 0:35
in voice AI for a long time across
- 0:37
different surfaces, devices, and um,
- 0:42
both at Alexa, at at Roku, at my own
- 0:46
startups, you know, in the app store.
- 0:49
And my perspective is a little different
- 0:52
from a lot of other voice AI
- 0:55
practitioners. I think it's a
- 0:58
combination of um a deep um voice user
- 1:02
interface expertise and intu intuition
- 1:05
mixed in with new technical approaches
- 1:09
uh that I think can produce really
- 1:11
magical experiences. So I think it's
- 1:12
both sides and I think that's especially
- 1:15
true in this new area that we're in with
- 1:17
frontier tech where the human interface
- 1:20
is basically being redefined. So let me
- 1:24
start with uh I'll just blast through
- 1:26
the first couple of slides then get to
- 1:28
the premise. I think everybody knows
- 1:30
that voice has incredible potential.
- 1:32
There's the power of voice I think
- 1:35
across everywhere. It's the most natural
- 1:37
interface. Humans love talking. And uh
- 1:40
the problem is the other half is the
- 1:43
pain of voice. So it's the power and the
- 1:44
pain. Voice is errorprone. And I think
- 1:49
those errors are going to continue for a
- 1:52
while. And I think the cost or
- 1:54
consequence of those errors is going to
- 1:56
grow, especially as we go fromational
- 1:59
AI bots to embodied AI where rather than
- 2:03
just giving answers that might be
- 2:05
erroneous,
- 2:07
we're going to have AI systems take
- 2:09
physical actions or digital actions
- 2:12
where, you know, if the robot throws
- 2:15
your watch out with the trash, it's a
- 2:17
lot worse than playing the wrong song.
- 2:19
So I do think that a new approach is
- 2:23
definitely needed and here's the TLDDR
- 2:26
of the premise we're going to walk
- 2:28
through today. Um there are two ways to
- 2:31
improve customer or user satisfaction of
- 2:35
a voice AI assistant and that is by
- 2:38
increasing accuracy which people know
- 2:39
about I mean technically accuracy and
- 2:43
the other is a different knob that we
- 2:46
have that we are not using adequately
- 2:48
and I'll call that a system decision
- 2:51
which we will define which is orthogonal
- 2:54
which is different from accuracy and I
- 2:57
believe This approach which I have used
- 3:01
in several different environments and
- 3:04
seen some success I think is a promising
- 3:07
area that we should consider developing.
- 3:10
Um let me walk through this with a
- 3:13
simple smart speaker example and we'll
- 3:16
go step by step with this approach but
- 3:19
it is a scalable approach that I think
- 3:22
uh can apply across different surfaces
- 3:24
and devices. So let's get started. So
- 3:27
suppose we all you know are making a
- 3:30
smart speaker coincidentally called uh
- 3:33
Alexa and Alexa is very simple. It just
- 3:37
allows you to you know ask for music and
- 3:40
it'll play a song and of course it will
- 3:43
play either the song you wanted or a
- 3:45
different song. So it'll be right or
- 3:47
it'll be wrong. This isn't that
- 3:49
different from what you've seen out
- 3:51
there. Um now let's to first talk about
- 3:54
accuracy. Accuracy. Let's say we define
- 3:57
it as we you know take a thousand spoken
- 3:59
requests. We observe the input and the
- 4:02
output. We label it and we look at this.
- 4:05
This is the map of a thousand points and
- 4:08
79% of the time 790 dots here were
- 4:12
actually the correct song. This is let's
- 4:14
say human annotated 20% 21% wrong song.
- 4:19
So that's the accuracy. Now, like I
- 4:22
said, knob one is to spend a lot of time
- 4:25
working on improving the accuracy, you
- 4:28
know, um, percentage point by percentage
- 4:30
point at any layer in the stack.
- 4:33
There's, if it's a cascaded system, you
- 4:35
know, there's a perhaps a wakeword layer
- 4:38
and a speech ASR layer and a NLU layer
- 4:41
which might have intent classification,
- 4:44
entity extraction, a lot of different
- 4:45
layers, VAD, etc. And any of those can
- 4:47
contribute to errors. So we spent time
- 4:50
we might be able to reduce that 210 to a
- 4:52
smaller number that is I think a known
- 4:56
area that we're tackling but I think
- 4:59
knob 2 which is what I was talking about
- 5:01
is what we'll go through here which is
- 5:03
keeping the accuracy exactly the same.
- 5:06
So 79% what could we do in conditions of
- 5:10
uncertainty to improve user satisfaction
- 5:13
apparent and I I think we can do a lot.
- 5:15
So let's start first with the original
- 5:17
system is just acting like I said user
- 5:20
says something system plays a song it's
- 5:21
either the right song or the wrong song
- 5:24
immediately I think just common sense
- 5:26
tells us that we could introduce at
- 5:28
least one system behavior to stop or
- 5:30
rather to reject the hypothesis and do
- 5:33
nothing. So uh there is now one more
- 5:36
option to decide the system may decide
- 5:38
and say sorry I didn't get that or sorry
- 5:41
could you repeat that? Uh the challenge
- 5:43
of course is how how when do we decide
- 5:47
to stop and I mean quantitatively. Um
- 5:51
here's one approach to kind of
- 5:53
visualizing this because if we don't
- 5:55
we'll just take probably some swag like
- 5:57
some guesstimate and I'll prove that if
- 5:59
we just took a guesstimate we would end
- 6:02
up with a worse situation than a more
- 6:04
rigorous approach. So let's just assume
- 6:06
I took those thousand data points and
- 6:08
like I said they've been annotated and
- 6:10
we assign a confidence score a single
- 6:13
confidence score to the hypothesis that
- 6:16
was generated by the system you know
- 6:18
between zero and one and let's say it's
- 6:19
reasonably calibrated. This is a
- 6:21
simplification of if it's a cascaded
- 6:23
system there are multiple layers and
- 6:24
multiple you know confidence scores but
- 6:26
let's just assume that for now. Whoops.
- 6:28
So we're going to have 790 points 200
- 6:31
that are correct 210 wrong. Each one has
- 6:34
a confidence score and we're going to
- 6:36
plot it, you know, plot the
- 6:37
distributions. Uh on the x-axis, I've
- 6:40
just converted from 0ero to one to
- 6:42
percentages. And the question is how do
- 6:46
we choose a threshold t such that
- 6:49
whatever that percentage is um to the
- 6:52
left of it meaning if when the system um
- 6:55
forms a hypothesis if the confidence
- 6:58
score c is less than that t stop and say
- 7:01
sorry otherwise play question is how do
- 7:04
we choose a t so far everything I'm
- 7:06
saying is fairly common sensical but
- 7:08
this is where um intuition will fail us
- 7:12
we might say something like okay I don't
- 7:14
know let's do 65%. It seems you know gut
- 7:17
feeling like okay it's kind of confident
- 7:19
that's probably when we should speak. Um
- 7:22
now here's where we start coming out
- 7:24
with some sophistication.
- 7:26
Any tea we choose is producing bad
- 7:29
outcomes. Bad in the in two fields. One
- 7:33
is obviously on the left side anytime
- 7:35
you stop it's bad. The user doesn't want
- 7:38
it to stop. He wants to they want to
- 7:40
hear their song. The other bad is if you
- 7:43
do play a wrong song, of course that's
- 7:46
bad as well. So these are two two kinds
- 7:48
of bad outcomes. But here's the
- 7:51
important part. Now I've like elaborated
- 7:53
on the um tree diagram on the right hand
- 7:56
side. The bad outcomes are not equally
- 8:00
bad. They're not the same thing from a
- 8:02
user perspective. And obviously let's
- 8:05
let's think about it. If the wrong song
- 8:07
plays, you said play kiss and it starts
- 8:11
playing kiss by Chris Brown instead of
- 8:14
the one by Prince. That's going to be um
- 8:18
the highest user cost. Now I'm defining
- 8:21
user cost from the user's perspective.
- 8:23
First I have to like hear music and
- 8:25
realize that is not Prince. Then I have
- 8:27
to shout over my Alexa and um you know
- 8:31
get it to stop and then I have to
- 8:33
re-request. All of that is a lot of
- 8:35
effort. that is definitely a worse
- 8:37
outcome than the system stopping and
- 8:39
saying sorry I didn't understand that
- 8:42
however we should go further and try to
- 8:45
quantify that relative badness and there
- 8:47
many ways to do it and I think this is
- 8:49
an area to be explored for now let's
- 8:51
just consider this a heristic of if that
- 8:55
outcome happens how many more seconds
- 8:58
additional seconds will it take for the
- 8:59
user to get back to success which is to
- 9:01
play the song they wanted kiss by Prince
- 9:04
and I'm I just put down some numbers
- 9:06
here. Let's say in the case of a bad
- 9:08
song, it's 10 seconds if you add up all
- 9:10
the things I got to do. And if it's a I
- 9:14
didn't understand you, it's 4 seconds
- 9:15
because that's how long it would take
- 9:16
you to respe and and the extra latency.
- 9:19
And now here's where we can start
- 9:23
utilizing that. If we go back to our
- 9:25
distribution curve on trying to find out
- 9:27
where is T. Now we've basically turned
- 9:31
this into a problem of minimizing a cost
- 9:34
function. It's a user cost function. It
- 9:36
is the number of bad acts wherever that
- 9:38
whatever the t causes times 10 because
- 9:40
that was a unit cost we gave plus the
- 9:43
number of stops times four because
- 9:45
that's the the unit cost we gave. By the
- 9:48
way, one thing I should have elaborated
- 9:51
because I work in voice and we like
- 9:52
language and we like puns. So this whole
- 9:55
thing is called an outcome user cost
- 9:57
heruristic. So that spells the word ouch
- 10:00
and that is some expression of pain.
- 10:03
Yes, we are you know language nerds. So
- 10:06
these kinds of things amuse us. Um so
- 10:08
now let's consider that is the cost
- 10:10
function is to minimize the ouch. And
- 10:12
now um that let's see if uh I'm going to
- 10:16
bring up a tool. Let's see if this
- 10:18
works.
- 10:20
Where I have actually gotten or with one
- 10:23
of my coding assistants gotten uh an
- 10:26
interactive
- 10:29
um graph where we have actually plotted
- 10:32
those thousand points and as we vary the
- 10:36
threshold t you can see that the total
- 10:40
user cost here which is that function of
- 10:43
you know x * y + a * b actually changes.
- 10:46
So let's in the very beginning when we
- 10:49
said the system was just playing
- 10:52
the the cost across those thousand
- 10:54
points was 2100 or divided by a,000 is
- 10:57
2.1 ouch points per turn. Then we said
- 11:01
okay let's insert a stop behavior and
- 11:04
let's like wing it and say 65%. That's
- 11:07
when I want the threshold. If we brought
- 11:10
this up to 65 yeah that's better. Now
- 11:13
it's 1904 or 1.9 per turn, but it's not
- 11:16
optimal. As it turns out, if we do
- 11:19
actually um ask for the AI to solve the
- 11:23
uh the problem across this curve, it
- 11:25
turns out 43%. So I'll drag it now to 43
- 11:30
is in fact
- 11:34
the optimal
- 11:36
optimal point of t. This minimizes the
- 11:39
cost function. You can see it's the
- 11:41
lowest point on this graph down here to
- 11:43
1
- 11:44
27. So effectively we haven't changed
- 11:48
the accuracy at all. The system is not
- 11:50
any smarter in that sense. But with some
- 11:52
clever system behavior, conversational
- 11:55
behavior is what we'd call it and some
- 11:56
optimization and a cost function called
- 11:59
ouch. Um we have from the user's
- 12:02
perspective produced a more satisfactory
- 12:06
assistant. And this is not a trivial you
- 12:08
know accomplishment. Okay. Now, let me
- 12:10
go back to this. [clears throat] Let me
- 12:12
see if I can get this. Oh, great. Okay,
- 12:16
let's continue this. Let's continue this
- 12:19
with by now adding one more behavior.
- 12:22
Let's call it the confirm behavior. So,
- 12:23
there was play obviously, then stop,
- 12:26
confirm. Confirm is basically the system
- 12:28
after you said something saying uh kiss
- 12:32
play kiss by Prince or maybe play kiss
- 12:35
by Chris Brown. And uh you know the user
- 12:38
can either confirm like affirm it or
- 12:40
they can correct it. It is a different
- 12:42
kind of behavior and again this is kind
- 12:44
of how humans behave. Um that's
- 12:46
obviously the inspiration. Now if we go
- 12:49
back to our problem of optimization,
- 12:52
we have a third obviously um option
- 12:55
which is to confirm. And so this would
- 12:58
translate to two thresholds
- 13:00
um two thresholds which are separating
- 13:03
the distribution into three spaces of
- 13:07
stop, confirm and uh act. And the
- 13:12
question is now where are these T's? and
- 13:15
we have now given up on guesstimating
- 13:16
because we know it doesn't work. So
- 13:18
we're going to be a lot smarter and go
- 13:21
back to the concept of user outcome cost
- 13:25
and then you know use it go look for
- 13:27
some optimization in that graph. So
- 13:29
let's uh define what are the what are
- 13:32
all the possible bad outcomes that t1
- 13:34
and t2 um make for. So good you can see
- 13:38
my cursor. So uh of course any stops are
- 13:42
still bad. Then in the middle are
- 13:45
confirmations. Confirmations are bad
- 13:47
because they slow the user down. There
- 13:49
is a confirmation outcome called confirm
- 13:52
yes where they just affirmed it by
- 13:54
saying yeah or no where they had to
- 13:56
correct it. And going back to our
- 13:59
formula these outcomes are not equally
- 14:02
bad. And in fact, nobody will, I think,
- 14:06
argue here from a user's perspective.
- 14:08
Affirming, just saying yes is obviously
- 14:11
less painful than saying no and then
- 14:13
having to restate whatever it is that
- 14:15
you wanted in the first place. So now we
- 14:17
I've assigned values of two or six. And
- 14:19
again, I said it was a heristic. This
- 14:21
would be roughly the amount of time it
- 14:23
would take for the extra for the user to
- 14:25
get to the song they want. Saying
- 14:27
listening and then saying yes is like
- 14:28
two seconds. Um and then now
- 14:34
uh we restate the cost function for this
- 14:37
you know added behavior as this number
- 14:40
of you know bad type one times unit cost
- 14:43
bad type plus bad type two times unit
- 14:45
cost etc. And now we try to minimize
- 14:49
this user cost function and minimize the
- 14:52
ouch.
- 14:53
Yes, I'm going to keep doing that pun.
- 14:56
Um let's go back. So this is now the
- 15:00
interactive graph but
- 15:03
with
- 15:06
um the cost values the unit costs here
- 15:08
10264
- 15:10
and uh you know we're just going to ask
- 15:13
the AI to tell us here's the heat map
- 15:16
because it's now two dimensions saying
- 15:18
that the optimal values are 41 for the
- 15:22
the T1 and 49 for the T2 and if we
- 15:27
employed that then we would go to 1464.
- 15:31
Uh, by the way, whatever numbers I put
- 15:33
in here, like let's say I thought wrong
- 15:36
act was 20. It's really irritating and
- 15:39
painful and takes way longer to actually
- 15:42
correct it when you hear a wrong song.
- 15:44
That would change you know all these
- 15:45
numbers uh and the optim optimal point.
- 15:48
So again it is about how what is the
- 15:50
relative badness of these outcomes also
- 15:52
of course the distribution curves
- 15:54
naturally. Uh let's go back here. Okay.
- 15:58
So, um I'm gonna
- 16:02
speed up a little bit. Uh let's go back
- 16:06
here.
- 16:07
Presentation mode. Okay. So, what have
- 16:11
we shown that if we did the super naive
- 16:13
approach, it's 2.1 act and stop 1.9 then
- 16:19
1.27 then 1.26. We are able to bring
- 16:22
this with every added layer of
- 16:25
sophistication, adding more behaviors,
- 16:27
being smart about outcome, uh, user cost
- 16:30
and optimizing. Um, we have made a
- 16:33
tremendous difference without changing
- 16:35
the accuracy at all. Um, this was a
- 16:38
super simplified example. In real
- 16:40
systems, you're not going to have
- 16:42
obviously some offline decision
- 16:44
threshold or two. It's going to be a
- 16:46
real time, you know, learned decision
- 16:48
model. But the principle is the same.
- 16:50
And I believe this is uh scalable across
- 16:54
all voice AI surfaces. Obviously this is
- 16:56
a smart speaker but if we go across any
- 17:01
of these surfaces you will find the
- 17:02
equivalence. If we um we will find the
- 17:07
analogies with some differences but the
- 17:09
spirit and the I think the the gain will
- 17:13
be similar. So just for example in the
- 17:16
TV AI assistant space if you employ it
- 17:20
here it's going to you're going to have
- 17:22
the same thing when users express
- 17:24
intents like on TV it's you know open a
- 17:27
channel that's one of the most common
- 17:29
obviously um requests on a TV voice
- 17:32
assistant same thing you're going to
- 17:34
find you'll have exactly the same
- 17:35
approach but the difference will be
- 17:38
maybe in the the assignments of the user
- 17:42
outcomes because the UI and the
- 17:43
modalities are different when you have a
- 17:46
TV you have a multimodal interface where
- 17:48
choices can be shown. So instead of you
- 17:51
know asking did you mean ABC you know uh
- 17:55
news live by speech that you will the
- 17:59
system would display choices and not
- 18:01
just one it show ABC News live this that
- 18:03
would be the confirm step and if it's
- 18:05
visual and you can use your remote
- 18:07
control to select something it's less
- 18:10
pain so you would change some of these
- 18:12
values or if in fact launching the
- 18:15
channel would kick you out of your
- 18:16
current state then it would go in the
- 18:18
other direction than cost of you know a
- 18:21
bad act would go much higher. So it's
- 18:24
the same concept but in this new
- 18:26
modalities
- 18:27
um variables can change, values can
- 18:30
change, arguments can change but the
- 18:32
premise still holds and you can improve
- 18:35
from the user's perspective because
- 18:37
we're all about you know making humans
- 18:39
happy. Um you can make them happier and
- 18:44
this as I said in conclusion can be
- 18:46
applied across all surfaces. I did say
- 18:49
at the very beginning, just to recap for
- 18:52
us, that voice is great when it works,
- 18:55
bad when it doesn't. And as we get into
- 18:58
embodied AI, where these AI assistants
- 19:01
are taking actions, physical or even
- 19:04
digital, like making a phone call or
- 19:06
sending an email, it is getting more and
- 19:09
more difficult just to rely on accuracy
- 19:12
to improve user satisfaction. I believe
- 19:15
there's a whole knob the second knob
- 19:17
called smarter conversational behavior
- 19:19
under uncertainty
- 19:21
and um if we actually exploit that we
- 19:25
can uh very much help these AI systems
- 19:30
reach a acceptable user experience
- 19:34
otherwise I think this will continue to
- 19:36
be a bottleneck like a lot of things
- 19:38
will get better but if the voice
- 19:40
interface as experienced by user does
- 19:43
not improve it is going to be a a a
- 19:46
choke point. And um if you just remember
- 19:50
one word or two words from this whole um
- 19:54
presentation, it would be to minimize
- 19:57
the ouch of the experience. Um so thank
- 20:00
you. I'll stick around for questions if
- 20:03
you guys got any. Thanks a lot.
- 20:06
[applause]
- 20:21
>> [music]