How long can your skills be before your agent forgets what you told it? — Laurie Voss, Arize AI
Read the talk
How many instructions can an agent follow? Capacity, failure modes, and verification
Laurie Voss’s IFScale experiments show a large increase in simultaneous keyword compliance, but also expose why longer skills files still need careful evaluation.
From a talk by Laurie Voss
At a glance
Ideas worth remembering
IFScale measures exact-word inclusion, giving a countable proxy for simultaneous constraints. It does not prove that equally large skills files will support reliable reasoning or conflict resolution.
Voss reproduced deterioration around 200–300 requirements in older models and reports newer high-accuracy boundaries nearer 2,000–5,000. The approximately tenfold gain is model-dependent, not a universal instruction allowance.
Failure detection must account for omitted requirements, safety refusals, exhausted thinking budgets, and partial completion. A polished opening can conceal a failed task.
More capacity can reduce compression and agent-handoff complexity, but longer prompts still carry cost and latency, and equivalent wording or ordering can change reliability.
Revisit older prompt-size assumptions and evaluate actual outputs. Demonstrated capacity answers whether a model can satisfy many constraints under tested conditions; verification asks whether this response did.
A small instruction ceiling would constrain the whole agent
Laurie Voss opens with a practical question: how much can you put in a skills file before the model stops tracking its instructions? A skills file can accumulate rules about tone, formatting, edge cases, and conditional behavior very quickly. Its length matters because an agent that quietly drops requirements can produce a plausible answer while failing the task it was actually given.
The investigation began after Voss heard an aside at an AI Engineer conference in Miami: an agent could follow about 200 instructions before it started forgetting them, according to a figure from 2025. That sounded like a severe constraint. A conditional rule such as doing Y when the user says X, a requirement to include a section on Z, and a prohibition against phrase W each count as separate instructions. A useful skills file could therefore exceed 200 requirements without seeming unusually elaborate.
Voss separates two questions that users often experience together: can the model handle all the rules, and did this particular output actually satisfy them? Looking at a polished response does not necessarily answer either question. His experiment first investigates where the capacity limit came from and whether it moved, then returns to the problem of recognizing compliance in real work.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
IFScale turns compliance into a count
The 200-instruction figure came from IFScale, a benchmark that asks a model to write a business report containing specified exact words. For example, the report must include customer and revenue. Each required word is treated as one instruction. After generation, the evaluator counts how many of those words appear in the response.
The experiment varies density, denoted n: the number of rules presented at once. Accuracy is the percentage of those rules the output satisfies. This gives the test a directly countable target. An output can sound like a convincing business report and still score poorly if it omits required words; fluency does not substitute for compliance.
Keyword inclusion is a proxy for more useful instructions. Voss argues that requiring revenue has the same broad shape as requiring a pricing section: both are discrete, named constraints the agent must remember while composing an answer. A prohibition against a phrase is another named constraint, although this benchmark directly tests inclusion. His reasoning is that a model that cannot track simple word requirements will probably struggle more with complicated rules. He therefore treats the keyword result as an optimistic ceiling, rather than a demonstrated capacity for arbitrary skills-file instructions.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Reproducing the older decline
Before testing newer models, Voss reran the original benchmark to establish a baseline. The original study used 10 models, but only three remained accessible through APIs when he conducted his replication. He tested those surviving models rather than selecting a fresh comparison group. He also notes that one of the three was retired after he published the research, illustrating how model availability can limit later replication.
He reports that the replicated accuracy curves matched the original paper within its noise boundary. Accuracy began deteriorating around 200–300 requirements, and at 500 rules the models were losing roughly 30%, 40%, or 50% of them. The earlier warning therefore had an experimental basis: for the older models tested, adding more simultaneous constraints substantially reduced the fraction satisfied.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The newer models outgrow the original test
Voss then applied the same prompt and words to newer models from OpenAI, Anthropic, Google, and DeepSeek. All scored 100% on the original test. That result established success within the benchmark’s existing range, but could not locate a new limit: the test topped out at 500 required words, and none of these models broke there.
To expose failure, he expanded the number of words the report had to contain: from 500 to 1,000, then 2,000, and onward to a 10,000-word vocabulary. The expanded results use a logarithmic horizontal axis, so a curve that looks like an abrupt cliff can represent deterioration across a substantial increase in requirements. The important comparison is how far the models retain high accuracy before that deterioration begins.
His headline is an approximately tenfold improvement over about 12 months. Where the older models deteriorated around 200–300 instructions, he describes newer boundaries nearer 2,000, and up to 5,000 for the strongest performance. Those are model-dependent benchmark results, rather than a single limit shared by every agent. What surprised him was the scale of the gain on a practical capability that could otherwise feel like an incremental model improvement.
The engineering implication is to revisit assumptions about prompt and skills-file size. Voss emphasizes that the benchmark was only about a year old and that model releases were already moving beyond the lineup he tested. A design built around an older capacity estimate may impose unnecessary constraints on a newer model; the estimate needs to be reconsidered for the model actually being used.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Failure can mean omission or refusal
The expanded experiment revealed more than a higher capacity limit. Older failures were straightforward to score: the model produced an answer and omitted some required words. The newer models did not all fail that way. DeepSeek retained the familiar pattern, beginning to forget requirements around 750 rules and dropping nearly half by 2,000. Voss considers that behavior comparatively predictable because the missing requirements are easy to count.
Claude repeatedly refused to complete the test at the API level. Voss attributes these refusals to safety classification: combinations of randomly selected words could make the request appear dangerous. He gives anthrax and cyanide as an example of vocabulary that could look concerning in combination. With thousands of random words in an instruction file, the test could encounter this refusal behavior before it measured the model’s ability to retain the requirements.
To obtain cooperative Claude results, Voss ran the vocabulary through an OpenAI safety filter and removed words it flagged. Claude then performed well. That intervention matters to interpretation: its successful performance followed a change to the vocabulary, while the refusals exposed a separate operational limit. Voss also warns that safety-sensitive or dual-purpose subject matter, including medical content, may trigger refusals at much smaller instruction counts. He presents this as a concern about content and classification, rather than ordinary forgetting.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Verification inside the model can consume the answer budget
Gemini performed strongly through 5,000 instructions, according to Voss, but beyond that it could exhaust its budget while trying to check compliance. He describes the model spending its thinking tokens ensuring that it followed all the instructions, leaving little or no capacity to produce the requested report. This failure appears as absent or inadequate output, rather than a normal report with a gradually increasing number of omissions.
His illustrative budget is 10,000 tokens, with 9,500 spent on thinking before a short response. The example mixes token and word quantities, so it should not be read as an exact allocation formula. The supported mechanism is that checking consumes the available generation budget, leaving too little room for a useful answer. That can incur substantial expense even when the final response fails to deliver the report.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A convincing beginning can conceal an abandoned task
Voss identifies GPT as the strongest performer in his test, reporting 99% accuracy through 5,000 rules. When pushed further, however, it sometimes began writing the business report and then declined to continue because it judged the request unreasonable. The report’s exact stopping point is unclear in his account, but the behavior is consistent: useful-looking prose arrives before the model abandons the remaining requirements.
He acknowledges the tension in the task itself. A coherent business report on no particular subject that must contain 5,000 random words is an unreasonable writing assignment. The model’s objection may be understandable, but the partial report still fails the benchmark because it misses most of the required keywords. Judging the assignment and satisfying its explicit constraints are different behaviors.
The danger is detectability. An immediate Claude refusal is obvious, while GPT’s answer can look successful until the reader reaches the ending. Across the four models, Voss observed omission, API refusal, exhaustion of the thinking budget, and abandonment after partial completion. A system that checks only whether some text was returned would not distinguish these outcomes adequately; recognizing failure requires attention to the model’s behavior and the actual completeness of its output.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
More capacity changes the reason to keep prompts short
Voss argues that the older capacity limit encouraged aggressive compression: keep each skills file under 200 instructions, then refer to subskills and additional files. Higher capacity reduces the need for that elaborate structure merely to fit the rules. For a use case with 100 or 300 specific requirements, he suggests putting them directly in the prompt instead of distributing them solely to avoid the earlier ceiling.
His larger example is a style guide with 2,000 named constraints, covering brand rules and legal disclaimers. Distributing such a guide across specialized agents creates another dependency: the agents must hand work to one another cleanly. If a single model can handle the required constraints, consolidation may remove handoff complexity. This is the architectural opportunity he draws from increased capacity.
Capacity does not make long prompts free. Including 10,000 different instructions produces an enormous prompt, and Voss expects greater cost and latency. His proposed decision therefore shifts from whether the model can attempt the task to whether the additional instructions justify their expense and delay. This is a tradeoff about prompt size, not a claim that all 10,000 instructions will reliably be satisfied.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Tracking constraints does not establish clear reasoning
Voss explicitly limits the claim: including random words in a fake business report is evidence that a long skills file might work, not proof that it does. He reports failure points ranging from 750 to more than 9,000 requirements, depending on the model and behavior being measured. The benchmark also does not measure whether the model reasons clearly over a giant prompt or resolves conflicts between its instructions.
He brings in reported context-rot research across 18 models, describing accuracy losses of 30–50% on long inputs before the context-window limit is reached. He also reports a surprising comparison in which coherent, well-structured text was more susceptible than shuffled instructions. Voss does not explain why that happened and says he would need to read the report. The narrower lesson is that fitting input into the available window does not guarantee effective use of it; this account does not establish random ordering as a general solution.
He returns to the distinction between loud and concealed failure. An API refusal signals that the task did not complete. A confident, polished report may conceal abandonment later in the response. Accepting a large rule set and returning convincing opening paragraphs therefore cannot establish success. Voss argues that the whole output must be checked, rather than trusted on the strength of its beginning.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Cheap experiments reveal a harder reliability problem
Voss reports spending $29 on about 2,300 calls across seven models. His point is that a focused empirical investigation can be inexpensive enough to run without a large research budget. The figure describes this study’s reported expense; it does not establish the cost of evaluating every production task.
For real applications with difficult tasks, he recommends monitoring outputs through evaluations. His suggested approach is to use another LLM to judge whether the output went wrong, and he connects that approach to Arize AI’s work. This is his proposed way to detect failures that do not arrive as explicit API refusals. He does not demonstrate here that the judging model is itself infallible or specify how such a judge should be validated.
A further study he describes tested 46 models and found that strong benchmark performance could coexist with unreliable instruction following. Rewording the same instruction could radically change compliance. Voss also warns that rearranging the same 2,000 instructions can make performance substantially worse. The best wording and ordering remain unresolved research questions in his account, so a successful formulation does not establish that equivalent formulations will behave the same way.
Voss welcomes the emergence of additional benchmarks aimed at many real, messy constraints. Their purpose is to investigate the broader question his random-word experiment only approximates: how well models follow numerous requirements at once when those requirements resemble actual work. Greater measured capacity has made that question more pressing, while leaving reliability as a separate problem.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The remaining task is to verify the result
Voss closes by describing a shift from compression to verification. Previously, writing a skill meant fitting the requirements into a small enough package that the model would not lose track. He argues that the expanded capacity has largely removed that pressure for the scale of instructions he discusses. The harder question is now whether the model actually did what it was asked.
His recommendation is to check the output every time, using evaluations as one would use tests for other code. A better prompt alone does not establish that a particular answer complied. He also urges engineers to revisit assumptions made six months earlier about prompt size and instruction capacity, because the approximately tenfold movement he observed shows how quickly those assumptions can become outdated.
The presentation ends with an offer of the experiment’s code and data, a brief invitation to a watch party, and thanks to the audience. The substantive closing claim remains that accommodating more instructions and verifying their execution are different engineering responsibilities.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
All right. Hello everybody.
- 0:15
Thank you for coming to this
- 0:17
delightfully nerdy talk. Uh this talk
- 0:20
has a really long title. Uh so let me
- 0:22
give you the short version up front. You
- 0:24
write skills files uh and stuff them
- 0:26
full of instructions. At some point the
- 0:28
model stops keeping track of all of
- 0:30
them. The question is where is that
- 0:32
point? At what point have you put too
- 0:34
many instructions in your skills files?
- 0:36
Uh and the answer has changed a lot in
- 0:39
the last year. I'm Lori. I'm head of
- 0:42
developer relations at Arise AI. Uh in a
- 0:44
former life, I co-founded npm Inc. So
- 0:46
some of you may know me from the days of
- 0:47
JavaScript. These days I spend a lot of
- 0:49
time thinking about AI and how to test
- 0:51
it.
- 0:53
Uh a few months ago I was at AI engineer
- 0:55
in Miami which was a good conference. Uh
- 0:57
and I was watching a talk by Dexter
- 0:58
Horthy. Uh it was a good talk. It was
- 1:01
not about this topic at all. Uh but
- 1:03
while he was giving that talk he
- 1:04
mentioned as an aside uh that an agent
- 1:07
can follow up to about 200 instructions
- 1:10
uh before it starts forgetting those
- 1:12
instructions. Uh and he then he moved on
- 1:15
in his talk and it was entirely an
- 1:16
aside. Uh and he he mentioned that that
- 1:19
figure is from 2025 so things might be
- 1:21
better now. Um, and I stopped listening
- 1:24
for a second because I was like, 200
- 1:26
instructions. Uh, is not very many
- 1:30
instructions at all. Right? A decent
- 1:31
skills file blows past 200 instructions
- 1:33
almost immediately. Um, if the user says
- 1:36
X, do Y, always include a section on Z,
- 1:39
never use the phrase W, every one of
- 1:41
those is a separate instruction. Uh, and
- 1:43
if the model quietly stops tracking them
- 1:44
after 200, that's a really hard ceiling
- 1:47
on the complexity of what you can build.
- 1:49
Uh, so I wanted to know where he got
- 1:51
that number first. Uh, and I wanted to
- 1:53
know if it was true. Uh, so you know the
- 1:56
feeling that I'm talking about. You
- 1:57
write this big beautiful skills file,
- 1:58
pages of rules, edge cases, tone,
- 2:00
formatting, you hand it to the agent, it
- 2:02
does the thing. Uh, and you look at the
- 2:04
output and go, did it actually pay
- 2:06
attention? Did it actually follow all of
- 2:08
these rules or did it just sort of, you
- 2:10
know, do what it felt like and sort of
- 2:12
give me a close simulacum of what I was
- 2:14
expecting? Um, you can't really tell or
- 2:18
can you? More on that later. Um, and so
- 2:21
you live with this lowgrade anxiety
- 2:23
every time you hit run. Um, and that
- 2:25
feeling is what this research is about
- 2:27
and what we're trying to find out if we
- 2:29
can avoid.
- 2:30
So here's my promise for your next 18
- 2:32
minutes. Uh, I'm going to show you where
- 2:34
that 200 number came from, whether it's
- 2:36
still true, and what the real number is
- 2:38
today. Uh, because it moved by an order
- 2:40
of magnitude. Uh, and then we're going
- 2:42
to talk about what that means for you to
- 2:45
take away. uh how long your skills and
- 2:47
prompts can actually be and what that
- 2:50
should what changes you should make to
- 2:52
your workflow as a result.
- 2:54
So the 200 number isn't folklore. Uh it
- 2:57
comes from a real benchmark called
- 2:58
IFScale uh from a paper uh by this guy
- 3:02
whose name I'm going to mess up
- 3:03
Jeroslowitch uh and co-authors last
- 3:06
year. And the test is beautifully
- 3:08
simple. Uh here's how if scale works. If
- 3:12
you ask the model to write a business
- 3:13
report uh and you give it a list of
- 3:15
specific words that it has to include
- 3:17
exactly in the report, include the exact
- 3:20
word customer, include the exact word
- 3:21
revenue, and so on for as many words as
- 3:23
you want. Each of those is an
- 3:25
instruction that it has to follow. Uh
- 3:27
and then you count how many of those
- 3:29
exact words showed up. Um
- 3:32
so because the test is so simple, you
- 3:34
only have to keep two numbers in your
- 3:36
head. One is density, which we call n.
- 3:38
That is how many rules we're talking
- 3:40
about at once. And the second is
- 3:41
accuracy, which is the percentage of
- 3:43
those rules uh that it was able to
- 3:45
actually follow. Uh now you might say uh
- 3:49
that including random words in a report
- 3:50
is not the same as following real
- 3:52
instructions and fair enough and we're
- 3:54
going to talk about that. Um but the
- 3:56
keywords are a proxy. Uh include the
- 3:59
word revenue is the same shape of task
- 4:01
as include a section on pricing, right?
- 4:03
Or never use this phrase. It is a
- 4:04
discrete named constraint that you've
- 4:06
told the agent that it has to follow.
- 4:08
Um, if a model can't track 200 words in
- 4:11
one prompt, it's definitely going to
- 4:12
struggle with 200 more complicated
- 4:14
instructions. Uh, so uh, if anything,
- 4:19
it's going to do worse. So, this number
- 4:20
is a ceiling. Uh, this number is as high
- 4:23
as you can go. If you give it more
- 4:24
complicated instructions, the number is
- 4:26
probably going to get lower. And 200 is
- 4:28
a really low ceiling. Um so before
- 4:31
chasing new models you have to do good
- 4:33
science which means that you have to
- 4:34
replicate the uh old result and make
- 4:36
sure uh that the 200 ceiling is real. So
- 4:40
I reran the original benchmark. Um the
- 4:43
original paper tested a whole batch of
- 4:44
models uh and models uh live and die
- 4:47
really fast. So uh by the time I got
- 4:49
around to doing this testing only three
- 4:51
of the models in the original set of 10
- 4:53
models that they used were still
- 4:54
available via any kind of API. Uh so
- 4:57
they were GPT 4.1, Claude Sonnet 4, and
- 4:59
Gemini 2.5 Pro. Those were models that
- 5:02
were available 12 months ago that are
- 5:03
still available now. Um and that is why
- 5:06
we tested those three because they were
- 5:07
what was left. Um and since I first
- 5:10
published this research a couple of
- 5:12
weeks ago, uh one of those three models
- 5:14
has been retired. So this was the last
- 5:15
possible time that I could have run this
- 5:17
test. Um so of that lineup, we're
- 5:20
already down to two. So don't get
- 5:21
attached to your models. Um here is the
- 5:23
results that we got replicating the
- 5:25
original if scale finding. Uh that is
- 5:27
accuracy on the vertical axis. So it
- 5:30
starts at 100% and begins to fall off.
- 5:32
Uh and then the number of rules uh going
- 5:34
up along the bottom on log scale. So
- 5:36
every time it gets halfway across it has
- 5:38
doubled uh the number of rules that it's
- 5:40
dealing with. Um so by 500 rules you're
- 5:44
losing 30 40 50% of them. Uh our curves
- 5:47
matched the results in the original
- 5:49
paper within the noise boundary. So the
- 5:50
finding was real. uh a year ago
- 5:53
somewhere around 200 to 300 rules
- 5:55
frontier models started falling apart.
- 5:57
That is a really low ceiling. Uh so that
- 6:01
is our baseline and now comes the fun
- 6:03
part where we took the exact same test
- 6:04
and pointed it at the current frontier
- 6:07
or rather what the current frontier was
- 6:09
when I ran this test. So I ran GPT 5.5,
- 6:12
Claude Opus 4.7 because 4.8 came out a
- 6:15
week after I ran this test. Uh Gemini
- 6:17
3.1 Pro and Deepseek V4 Pro. So, I gave
- 6:20
them the same prompt, the same words,
- 6:22
the same everything. And I immediately
- 6:24
ran into a problem, which is that they
- 6:26
aced it. They all scored 100%
- 6:29
immediately on this test. Absolutely no
- 6:31
bugs. Uh,
- 6:34
so we'd built a test to find the ceiling
- 6:36
and the models had walked straight
- 6:37
through the ceiling without noticing
- 6:38
that the ceiling was there. Um, and that
- 6:40
was a problem because the benchmark was
- 6:42
written to top out at 500 words. So, I
- 6:44
had to change the benchmark in order to
- 6:45
be able to find the new ceiling. So, I
- 6:47
moved the goalposts. I gave it more
- 6:49
words to include. I doubled uh it from
- 6:51
500 to a,000. I doubled it again from
- 6:53
a,000 to 2,000. And I kept doing that
- 6:55
until I hit a 10,000word vocabulary. And
- 6:58
that is where I began to find the
- 6:59
ceiling of what models can do these
- 7:02
days. Um, so let me put up the this is
- 7:05
the money slide. This is the results.
- 7:07
Remember log scale on the on the uh
- 7:11
x-axis there. So it's going from 500 to
- 7:13
1,000 to 5,000 to 10,000. Uh so it looks
- 7:16
like that scale is falling off of a
- 7:18
cliff and it's actually happening over
- 7:19
like a thousand numbers. Um
- 7:22
but uh look how far to the right these
- 7:25
new curves get before they bend. A year
- 7:26
ago they were falling over at 200 to 300
- 7:29
instructions and now depending on the
- 7:30
model the boundary is closer to 2,000.
- 7:32
And for the best of them it is up to
- 7:34
5,000 instructions before they begin to
- 7:37
fall off a cliff. So in about 12 months
- 7:39
frontier models got close to 10 times
- 7:41
better at following instructions
- 7:43
simultaneously. That is the headline
- 7:45
fighting and there is a lot of nuance
- 7:47
that we need to get into. Um the
- 7:50
capacity to track 2,000 named
- 7:51
constraints in a single prompt is there.
- 7:54
Um and that's really interesting because
- 7:56
I think uh I don't know if everybody
- 7:58
else feels this way but like it feel it
- 8:01
felt to me like the the jump from you
- 8:04
know GPT 5.1 to GPT 5.5 was kind of
- 8:06
incremental, right? It didn't feel like
- 8:08
we'd got 10 times better. But this is a
- 8:11
test that really matters to uh a very
- 8:14
practical thing like how long can my
- 8:16
skills file be? Uh and in the course of
- 8:18
a year we got 10 times better. Uh and
- 8:22
that thing that gets me is that this
- 8:23
benchmark is barely a a year old. A year
- 8:25
later 500 is a rounding error. Uh and
- 8:28
this keeps moving under my feet. I
- 8:29
tested 4.7 uh opus 4.7. Opus 4.8 is even
- 8:33
better. Um so this chart is a little out
- 8:36
of date already which is kind of the
- 8:37
whole point. If you set your engineering
- 8:39
assumptions about how skills files
- 8:41
should work, about how prompt how long
- 8:43
your prompt can be, and you did that
- 8:45
more than about six months ago, you are
- 8:47
incorrect now, and you should probably
- 8:49
be re-engineering how you do stuff. Uh,
- 8:52
but there is more to this story uh
- 8:54
because the way that the a models failed
- 8:58
uh changed dramatically. Uh, and the way
- 9:01
that they failed is very important. This
- 9:03
part was a completely unexpected finding
- 9:06
when I started running the experiment.
- 9:08
Uh, and it totally messed up my test to
- 9:09
start with because uh, the old failure
- 9:12
mode was boring. They would just forget
- 9:14
instructions and I could measure how
- 9:15
many instructions they had remembered or
- 9:17
forgotten. Uh, but the new ones fall
- 9:19
apart in their own weird extremely
- 9:21
onbrand way. Uh, so let me introduce you
- 9:24
to how these four models fail. Uh,
- 9:27
Deepseek 4 is a traditional model. It
- 9:29
just forgets things. It doesn't have any
- 9:31
drama. um it starts forgetting
- 9:34
instructions around 750 rules and by
- 9:36
2000 it's dropping nearly half of them.
- 9:38
Uh so it just forgets which frankly is
- 9:40
the failure mode that I trust most
- 9:42
because it's predictable. It's very easy
- 9:43
to measure. Uh and the other models were
- 9:46
not nearly as cooperative. Uh Opus 4.7
- 9:51
uh would decide repeatedly that the test
- 9:53
was dangerous. Uh and what it would do
- 9:55
is it would refuse at the API level to
- 9:58
complete the test. I didn't know that
- 10:00
there was an API response that you could
- 10:01
get from Claude where it was like, "No,
- 10:03
I could do this, but I'm not going to."
- 10:06
Uh, but that's absolutely an API level
- 10:09
response that Claude supports because
- 10:10
they care so much about safety. Uh, and
- 10:12
I started getting those all of the time.
- 10:15
Uh, and the reason that was happening is
- 10:16
because Claude has a very sensitive
- 10:18
safety classifier. Uh, and if you put in
- 10:20
certain combinations of words like say
- 10:22
anthrax and cyanide, it decides that the
- 10:24
whole request is dangerous and it bails
- 10:26
out. Uh, and if you remember what my
- 10:28
test does, my test is throwing uh 5 to
- 10:31
10,000 random words into uh into an
- 10:34
instruction file. And so my my randomly
- 10:37
selected words contained all sorts of
- 10:38
things that looked dangerous in
- 10:40
combination to the safety filter. And so
- 10:41
it kept bailing saying that I was asking
- 10:43
it to, you know, make a bomb or
- 10:44
something. Um,
- 10:47
so, uh, we had to for to get Claude to
- 10:52
cooperate, I had to take all of my words
- 10:54
and run them through OpenAI safety
- 10:55
filter and filter out all of the naughty
- 10:57
looking words so that it could get to
- 10:58
anywhere. Once I given it that, Claude
- 11:01
did really well. Uh so but the failure
- 11:03
mode is that Claude is more likely to
- 11:05
decide what you're doing is dangerous
- 11:07
very early on uh at you know even two or
- 11:10
300 instructions if what you're doing uh
- 11:13
is you know contains anything to do with
- 11:15
medical advice because medical things
- 11:16
often are dual purpose. They can be
- 11:17
dangerous. They can be safe. Um so uh
- 11:21
the third failure mode was Gemini 3.1
- 11:23
Pro. Gemini is rock solid all the way
- 11:26
out uh to 5,000 instructions. It does
- 11:28
extremely well. um genuinely one of the
- 11:31
best on the chart. Uh and then past that
- 11:33
it gets weird. Um it doesn't forget the
- 11:36
instructions, it gets overwhelmed by the
- 11:39
instructions. What it tries to do is it
- 11:42
uh it uses thinking tokens to make sure
- 11:44
that it is following all of the
- 11:45
instructions at once. And when the
- 11:47
number of instructions gets really high,
- 11:48
it uses all of its thinking tokens. It
- 11:50
uses its entire token budget thinking.
- 11:53
And then it doesn't give any output.
- 11:55
It's it gets to like nine, you know, if
- 11:57
you've given it 10,000 tokens worth,
- 11:59
it'll get 9,500 tokens worth of thinking
- 12:02
and then give you a 500word response
- 12:04
which doesn't contain any of the tokens.
- 12:06
Uh so it thinks itself into a corner and
- 12:09
runs out of room to actually answer,
- 12:10
which is very expensive, uh and totally
- 12:13
unhelpful, which is kind of on brand,
- 12:15
isn't it? Um [snorts]
- 12:18
uh
- 12:19
which, you know, I would never say that
- 12:20
out loud. Uh, and finally comes the
- 12:24
winner, which is uh, GPT 5.5. GPT 5.5 is
- 12:27
the best of the lot. 99% accuracy all
- 12:29
the way out to 5,000 rules. Um, but if
- 12:32
you push it far enough, it is by far the
- 12:34
weirdest of the bunch. Uh, because it
- 12:36
doesn't refuse outright. It doesn't
- 12:38
silently forget. Instead, what it does
- 12:39
is it gets frustrated and tells you that
- 12:42
the test is stupid.
- 12:44
Uh, it starts the report. It gets a few
- 12:47
like that's the thing. It doesn't start
- 12:49
out just saying no. It starts the
- 12:51
report, it starts writing the report,
- 12:52
and like 500 words into the report, it's
- 12:54
like, "No, this is dumb. I'm not going
- 12:56
to do this." And then it politely tells
- 12:57
you, "This is dumb. I'm not going to do
- 12:59
this anymore." Uh, that is the actual
- 13:01
response that it gave me, but that that
- 13:02
was like 5,000 words into the into this
- 13:04
business report that I told it to
- 13:06
generate. Um, so it's not wrong, right?
- 13:10
I was asking for a coherent business
- 13:12
report that on no particular subject
- 13:14
that contains 5,000 random words. You're
- 13:17
right, Gemini DPT. this is a a stupid
- 13:20
thing to ask for. Um,
- 13:23
which is a deeply unreasonable request
- 13:25
and GPT called this out on it. Um, but
- 13:27
it still counts as a failure in the test
- 13:29
because the half-finish report that it
- 13:30
gives you is missing most of the
- 13:32
keywords and it is also the hardest one
- 13:34
to detect because claude bails
- 13:36
immediately. Claude says, "No, I'm not
- 13:38
going to do this." Uh, Deepseek does its
- 13:40
best. Uh, but GPT does what looks like a
- 13:44
good job unless you read all the way to
- 13:46
the end of the report where it says,
- 13:47
"No, actually I'm going to bail because
- 13:48
this is stupid." Um,
- 13:51
so if you step back and look at the four
- 13:53
together, Deep Sea quietly forgets,
- 13:54
Claude gets scared and refuses, Gemini
- 13:56
overthinks itself into silence, and GPT
- 13:59
5.5 finishes half of the job and tells
- 14:01
you that the rest of it is beneath it.
- 14:03
Um, and the point was the point isn't
- 14:06
which one of these is funniest, although
- 14:08
it is genuinely a little funny. Uh the
- 14:10
point is that did it follow my
- 14:12
instructions no longer has one failure
- 14:14
mode. It has four different ways that it
- 14:16
can fail and you can't recognize that
- 14:18
failure unless you know which model
- 14:20
you're dealing with and what its m what
- 14:22
its pattern of failure is going to be.
- 14:24
Uh so the models get 10 got 10x better.
- 14:27
They fail in funny ways. Why should you
- 14:29
care when you uh get back to your desk?
- 14:32
Because three things have changed to
- 14:33
your workflow. The first is that a year
- 14:36
ago, the smart move was to keep every
- 14:38
skills file very very short. Uh under
- 14:41
200 instructions, then point off to
- 14:42
subsklls and a whole like you know
- 14:45
byzantine labyrinth of uh additional
- 14:48
skills files and subfiles and things
- 14:50
like that. Uh and you mo you were
- 14:52
compressing your your instructions to
- 14:54
fit into a very small available space
- 14:56
and you don't need to do that anymore.
- 14:58
Your skills files can be very long. Um,
- 15:01
number two is that if your use case
- 15:03
needs a 100 specific rules or 300, you
- 15:05
can just put them all in the prompt. Uh,
- 15:09
you don't have to lie awake wondering
- 15:10
whether which ones the model silently
- 15:12
ignored. Um, and if you've been thinking
- 15:15
uh about uh your own lived experience of
- 15:18
using models, uh you probably recognize
- 15:21
this. you've discovered that you've got
- 15:22
less worried about how long your your
- 15:24
prompt is going to get uh because the
- 15:26
models have genuinely got 10 times
- 15:28
better at following your prompts. Um
- 15:32
2,000 named constraints is an entire
- 15:34
style guide, right? Like it's it's every
- 15:36
brand rule, every legal disclaimer. Uh a
- 15:38
year ago, you'd have had to shard that
- 15:40
across a dozen specialized agents and
- 15:42
hope that your specialized agents are
- 15:43
hand are are handing off to each each
- 15:45
other cleanly. But now you can ignore
- 15:48
that. Um but the third thing is the big
- 15:50
one. The question used to be can the
- 15:52
model even do this? And the answer is
- 15:54
now firmly yes. Well reasonably firmly.
- 15:57
Uh is it worth the cost is the new
- 16:00
question because you can include 10,000
- 16:03
words of of sorry 10,000 different
- 16:05
instructions into your prompt. But that
- 16:06
is going to be an enormous prompt. It's
- 16:08
going to be a very expensive prompt.
- 16:09
It's going to be a very slow prompt. So
- 16:11
what used to be a hard wall that you
- 16:12
would run against has now become a soft
- 16:14
trade-off of is it worth me adding all
- 16:16
of these extra instructions if it's
- 16:18
going to give me more cost and more
- 16:19
latency.
- 16:21
Uh and now some caveats uh to head off
- 16:25
the Q&A. Um first and important first
- 16:28
and most important I mentioned this
- 16:29
earlier this is a proxy task including
- 16:31
random words uh in a in a fake business
- 16:34
report um is evidence that long skills
- 16:37
file works. It is not the same as proof
- 16:39
that a long skills file works. Um, also
- 16:43
the models hit the wall at wildly
- 16:44
different points anywhere from 750 to
- 16:46
9,000 plus. So you have to pick your
- 16:48
model very carefully. Uh, what our test
- 16:53
doesn't do is measure whether the model
- 16:55
reasoned clearly over a giant prompt. So
- 16:59
uh, the good news is since I did my
- 17:00
research several weeks ago, uh, a whole
- 17:02
bunch of people have piled in on this.
- 17:04
Um and now there's good research uh
- 17:07
actual scientists have got involved and
- 17:09
done uh Chroma's has done context rot
- 17:12
work uh across 18 models showing that
- 17:15
accuracy on long inputs can fall 30 to
- 17:18
50% well before you hit the context
- 17:20
window limit. Uh and the weird part of
- 17:23
their finding was that uh coherent well
- 17:26
ststructured text is more likely to hit
- 17:28
that failure mode uh than if you just
- 17:30
put your instructions into a random
- 17:31
order and shuffle them in. Uh, I don't
- 17:35
know why that's the case. I'd have to
- 17:36
read their report. Um, so the model can
- 17:40
track 2,000, 5,000, possibly 10,000
- 17:42
instructions, but it's not necessarily
- 17:44
going to uh reason clearly over them.
- 17:47
It's not necessarily if those if those
- 17:49
instructions conflict, if there is
- 17:50
tension between them, it's not
- 17:52
necessarily going to get that right. Um,
- 17:55
and then there's the other one I
- 17:56
mentioned briefly. Uh, collude's
- 17:59
refusals are annoying, but they are
- 18:00
loud. You get an error, you know it
- 18:02
failed. Uh GPT's polite half-finish
- 18:04
report is much more dangerous because it
- 18:06
looks like a real answer. Uh you have to
- 18:08
read the whole thing to notice that it
- 18:09
gave up quietly halfway, which means
- 18:11
that you can't trust the output. It
- 18:14
means you have to read the output every
- 18:16
single time to make sure whether or not
- 18:17
it's working. Uh so the model will
- 18:20
accept your 2,00 rules and it will hand
- 18:22
you back something that looks at least
- 18:23
to begin with confident and polished but
- 18:25
could be bailing out halfway through.
- 18:28
Um,
- 18:30
so, uh, as an aside, people always ask
- 18:33
me, "How much did all this cost me?" It
- 18:34
cost me $29 to run all of these queries.
- 18:37
2,37
- 18:39
2,300 calls across seven models, uh,
- 18:42
came to $29. Uh, it turns out novel
- 18:44
research doesn't cost very much. Um,
- 18:48
and this is the part of the talk where I
- 18:50
was saying that you have to check this
- 18:51
stuff in production because you can't
- 18:53
trust that your model isn't going to
- 18:55
silently fail. Uh, so you knew I was
- 18:58
going to mention evals eventually
- 18:59
because I work at Arise and this is
- 19:00
where I do that. Um, but there are
- 19:02
plenty of plugs for Arise. So I'm just
- 19:04
going to say one true thing which is
- 19:06
that if you are building a real AI
- 19:07
application and you are giving it
- 19:09
genuinely tricky tasks, you are going to
- 19:11
run into one or more of these failure
- 19:12
modes with a frontier model. Uh, and
- 19:15
unless it's Claude telling you just to
- 19:16
off at the API level, the only way
- 19:19
to know that something went wrong is
- 19:21
monitoring your outputs with another
- 19:22
LLM. That is an eval. And that is what
- 19:24
Arise does. And I'll leave it at that.
- 19:27
Uh, I already mentioned that there's
- 19:29
been new research since we did our own.
- 19:31
Here's another important one. A paper
- 19:32
landed testing 46 models called
- 19:34
revisiting the reliability of language
- 19:36
models in instruction falling, which you
- 19:38
can bet made my ears perk up after I did
- 19:40
that research myself. Uh, and they found
- 19:42
something uncomfortable, which is that a
- 19:44
model can ace a benchmark like ours and
- 19:46
still be wildly unreliable. because if
- 19:48
you reword the same instruction in a
- 19:51
slightly different way, it can make a
- 19:52
radical difference to how well uh it
- 19:55
follows those instructions. So the model
- 19:57
can follow 2,000 instructions and it can
- 19:59
do it really well. But if you put the
- 20:01
same instructions, the same 2,000
- 20:03
instructions in a different order, it
- 20:05
can suddenly make the model much worse
- 20:07
at following those instructions. And how
- 20:09
ex how exactly to do that? what is the
- 20:12
correct order of instructions to give
- 20:14
your model such that it follows them
- 20:15
perfectly as opposed to getting confused
- 20:17
is still research that is being done. So
- 20:20
capacity went up but reliability is
- 20:23
still a problem. Um and then this is
- 20:26
just a little brag because uh I was
- 20:29
happy about it like I'm not a scientist.
- 20:30
I did some research and then a whole
- 20:32
bunch of other actual scientists piled
- 20:34
in uh and did real science on the same
- 20:36
question. There's now a whole bunch of
- 20:37
benchmarks that have shown up uh to
- 20:39
measure this same question. Firebench,
- 20:41
CCR bench, Guidebench uh are all trying
- 20:43
to measure the same thing. How well
- 20:45
models follow a lot of real messy
- 20:47
constraints at once. Uh and now the
- 20:49
whole field is looking at it. So if you
- 20:51
want better science than my, you know,
- 20:53
10,000 random words, uh the real science
- 20:56
exists now. Uh so that gets me to where
- 21:00
I will leave you. A year ago, the hard
- 21:02
part of writing a skill was fitting
- 21:03
everything in without the model losing
- 21:04
the plot. That was a compression
- 21:06
problem, and the compression problem is
- 21:08
gone. uh the model will hold your 2,000
- 21:10
instructions just fine. The new hard
- 21:12
part is knowing whether it actually did
- 21:14
what you said and that is a verification
- 21:16
problem. Uh a verification problem
- 21:18
doesn't get solved by writing a better
- 21:20
prompt. It gets solved by checking the
- 21:21
output every time uh the same way that
- 21:24
you would test any other code, which is
- 21:25
to say an eval. The ceiling moved by 10x
- 21:28
in one year. Uh so go back and check the
- 21:31
assumptions that you made six months ago
- 21:33
about how big your prompts should be,
- 21:35
how big your uh instructions can get. uh
- 21:38
because they might already be wrong.
- 21:39
Boom. Be wrong. So that is the talk. If
- 21:42
you want uh all of the code and all of
- 21:44
the data, uh it is at this GitHub URL.
- 21:47
Uh and this other QR code is uh
- 21:50
something marketing made me insert. We
- 21:52
are having a World Cup watch party
- 21:54
tonight at 5:00 p.m. Uh you can come to
- 21:56
our party. That link is to the Luma that
- 21:58
will get you into the get into get you
- 22:00
into the party. Uh I hope this talk has
- 22:03
given you some novel information or at
- 22:05
least a couple of laughs. And thank you
- 22:06
so much for your time and attention.