AI Engineer World's Fair 2026
The Miranda Hypothesis: How Hamilton (the Musical) Poisoned Your Persona Evals
Read the talk
The Miranda Hypothesis: When a Convincing Persona Gets History Wrong
A Lincoln persona gives a persuasive answer about presidential war powers. Jacob E. Thomas uses that demonstration to distinguish recognizable personality from historical fidelity, then proposes an evaluation built with a historian to test the difference.
From a talk by Jacob E. Thomas
A historical face is already an interface
Jacob E. Thomas opens with Character.AI, a platform for conversations with role-playing personas. He then shows Hello History, where a learner can summon a Marcus Aurelius tutor. A historical face offers more than information: it invites the learner to interpret an answer as that person’s thinking. Underneath is a role-playing language agent, configured to reason and speak through a persona. Its apparent identity matters as these systems move from entertainment toward education and civic interpretation.
Thomas’s own COMPANION framework makes the construction inspectable. It configures a general-purpose language model through an open prompt framework. One demonstration places American founders in a room with the Epstein files and asks them to counsel the country. Openness lets someone examine what shapes the encounter; Thomas does not present it as proof that his portrayals are more faithful.
He then asks a Lincoln persona when a president may take the country to war without Congress. The answer acknowledges congressional authority but defends decisive executive action during emergencies, invoking the eventual preservation of the Union. It is fluent, grave and recognizably presidential. Those qualities explain its appeal, but they leave a question unanswered: what does this performance establish about Lincoln?
Persona benchmarks provide increasingly systematic measurements. Thomas’s question is whether those measurements capture the property that a historically grounded interface actually needs.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
What a personality score can establish
InCharacter evaluates personality fidelity through psychological interviews. Its Table 2 reports results for batched expert rating with GPT-4 as interviewer; the default role-playing model is GPT-3.5. Under that setup, the reported dimensional measured-alignment accuracy on 16Personalities is 80.7%. The study uses 32 characters, and the table averages three runs against human-perceived personality profiles, excluding ambiguous dimensions. This is not a Hamilton-specific score or a measure of historical fidelity. A recognizable Hamilton could agree with a perceived personality while reasoning through a modern portrayal of his life. Thomas’s thesis is that an evaluation cannot detect anachronistic compositing merely by measuring fluency and personality consistency.
Thomas approaches this as a measurement problem. His production AI work follows training in behavioral epidemiology, where he studied how information environments shape populations. The project brings that concern together with historical expertise from Rick Halpern and archival expertise from Shawn Martin at Washington College.
The technical progression is substantial. Rule-based systems select prepared responses; imitation reproduces voice and characteristic expression; cognitive simulation adds psychological structure, memory and motivations connecting situations to behavior. These are substantive advances in what a persona can maintain across an encounter.
CoSER develops literary role-playing through character material and motivations. Its dataset contains 17,966 profiled characters across 771 books. The paper’s Table 3 compares CoSER-70B with GPT-4o across several tasks. CoSER-70B scores 93.47% accuracy on LifeChoice, compared with GPT-4o’s 75.96%. On InCharacter’s Big Five Inventory evaluation, CoSER-70B scores 75.80% dimensional accuracy, compared with GPT-4o’s 76.54%. The comparison is competitive rather than a sweep, and this BFI evaluation is a different scale and setup from the earlier 16Personalities result.
PsyMem combines psychological attributes with explicit memory control. InCharacter’s interviews provide another way to examine whether a persona maintains a coherent personality. These methods improve the representation and measurement of personality; their sophistication does not make personality agreement equivalent to documentary grounding.
Thomas calls convincing performance the mask. It asks whether an answer feels like the person: natural, emotionally responsive and consistent in personality. The mirror asks whether that person could have known, believed or argued it at the specified moment. These properties can come apart. A coherent motivational profile does not, by itself, prevent an answer from borrowing its premises from the future.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A convincing Hamilton can erase the difficult parts
Thomas first plays the opening of Hamilton, then shows a model explaining why Hamilton works so hard. The generated answer organizes his ambition around orphanhood, immigration, relentless effort, nation building and the desire to leave a legacy. It presents a person determined to create something that will survive him.
Thomas hears the musical’s emotional structure in that response, contrasting it with the more legalistic reasoning of the documentary Hamilton. The proposed influence is deeper than repeating lyrics. A later work can supply the organizing explanation of a life even when every generated sentence is new. Familiar narrative structure makes the persona immediately intelligible to a modern audience.
The second example asks about slavery. Another musical excerpt establishes an abolitionist theme among its revolutionary characters. The model then presents Hamilton as a consistent moral opponent of slavery, partly grounding that identity in his membership of the New York Manumission Society.
The historian’s intervention concerns what disappears. Thomas describes a contested record that includes transactions involving enslaved people for relatives and clients, alongside political dependence on slaveholders. His point is not to settle Hamilton’s entire historical record within one answer. It is that the model removes the dispute and contradictions, producing an uncomplicated hero where the evidence demands a more difficult account.
A personality-focused evaluation could reward that simplification: the answer is coherent, plausible and easy to recognize. Thomas uses the example to expose a measurement gap, rather than presenting a measured benchmark verdict for this particular response. Attractive expression can conceal the substantive error.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
How a familiar portrayal could reinforce itself
The Miranda Hypothesis names a powerful cultural representation, not an artistic wrongdoing. The musical makes Hamilton meaningful to contemporary audiences. The engineering question is what happens when that familiarity becomes the basis for historical testimony. Thomas proposes three linked claims:
- Primary writings coexist with an expanding cultural afterlife of adaptations, commentary, teaching material and social media.
- Autoregressive training learns continuations from this mixture without automatically enforcing a hierarchy between primary evidence and later interpretation.
- Generation can therefore produce a composite that blends incompatible periods into a persona corresponding to no particular moment in the person’s life.
Such a Lincoln could reason from the Gettysburg Address before having written it.
The finite Federalist Papers and the expanding body of material surrounding the musical illustrate the proposed imbalance. That comparison motivates the hypothesis; it does not establish the composition of a particular model’s training corpus or measure the influence of either source.
Thomas makes the cultural effect tangible through a Schuyler Mansion anecdote. Visitors arrived treating the musical’s selective focus on the sisters as a description of the whole family. Interpretive staff had to disentangle the story audiences knew from the family record.
Thomas then asks whether preference training might reward the familiar portrayal. Human raters can bring expectations shaped by the same cultural narratives found in training material. If they prefer the Hamilton they already recognize, optimizing for their approval could reinforce the composite instead of correcting it.
Three characters are not the whole family.
Thomas recounts a Schuyler Mansion misconception, then uses it to explain his hypothesis about model personas.
Three Schuyler daughters are centered in the story.
Reported visitor inference:
the family had only three daughters.
Eight survived to adulthood.
The cast’s focus does not establish the family’s size.
If preference training rewards familiarity, it may reinforce the composite rather than repair its historical contradictions.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A date cutoff does not identify a person’s knowledge
Training a model on a corpus ending at a fixed historical date addresses one route of contamination. A corpus ending in 1789 cannot contain a much later Broadway musical. Thomas welcomes time-locked models while distinguishing period anchoring from persona anchoring. A period corpus still contains different authors, perspectives and accounts of the same figure. It does not automatically identify which knowledge belonged to that individual or which arguments they could make.
His proposed next stage, epistemic simulation, adds three external constraints:
- A specified primary corpus supplies the evidence that licenses the persona’s reasoning.
- A temporal anchor establishes the moment from which the persona speaks.
- Domain experts judge the output against the documentary record.
Cognitive simulation organizes a plausible mind through motivations and psychological consistency. Epistemic simulation makes the performance answerable to documents and dates. Stating those constraints is a design commitment; their effectiveness still needs evaluation.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Make the encounter the unit of engineering
Thomas’s terminology argument is specific to systems intended to instantiate a person. It does not concern agents that write code, book travel or triage tickets. For historical personas, treating identity as a property of a checkpoint makes the claim difficult to inspect or hand to a qualified reviewer.
He instead defines a role-playing language system through five components:
- A structured prompt supplies framing and constraints.
- Primary documents supply anchor material.
- A temporal anchor fixes the historical moment.
- A general-purpose model reads and generates through those materials.
- A human curator selects the evidence and retains authority to interpret and challenge the output.
The model is one replaceable component of the encounter.
The analogy is theatrical: Hamlet is not contained entirely within the actor performing him. Similarly, a persona depends on configuration, documents and interpretation. The prompt, corpus and temporal anchor can become repository artifacts that people compare, review and revert. A historian can inspect those materials without inspecting model weights. Retaining them makes the setup recoverable, though it does not guarantee identical generated wording or reveal every influence from pretraining. Context engineering here means composing an encounter whose explicit evidence remains available for scrutiny.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Keep the document available for another reading
Two construction approaches put historical material in different places. Retrieval-augmented generation and related approaches supply documents in the context window at inference time. Fine-tuning uses training examples to change model weights; Thomas names Character-LLM and CoSER in this discussion. With inference-time documents, a reviewer can return to the specific evidence supplied for an answer. Fine-tuning alone does not provide an equivalent account of which training passages supported that answer.
Thomas argues that persona fine-tuning could make the problem harder to see. Documentary training material interacts with cultural associations already learned by the base model, potentially producing a more polished performance without resolving its underlying contradictions. This is a proposed failure mechanism, not a comparative result established here. He invokes clinical-model specialization as an analogy for testing specialization rather than presuming its superiority; that analogy does not determine which architecture yields a faithful historical persona.
His positive case for context is archival. A document remains available for another reading. A curator can challenge an inference, change the selection or restore material an earlier interpretation excluded. The source text remains separate from the generated performance. Preserving that distinction connects an ethical commitment to a debugging capability: reviewers can compare the evidence supplied with the argument produced. Context does not erase learned priors, but it supplies an explicit record against which their influence can be examined.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The architecture determines who can participate
Accessibility follows from where the historical work happens. Fine-tuning involves dataset preparation, training infrastructure and computational resources. Supplying selected documents through a model interface puts more of the work within reach of someone whose expertise is reading an archive rather than operating that infrastructure.
Thomas imagines a doctoral student studying early America, a community archivist documenting a regional figure and a grandchild working with a grandmother’s letters. Each may possess evidence or contextual knowledge unavailable to a centralized developer. The architectural aim is to let those people contribute and challenge interpretations without first becoming model-training specialists. Broader participation creates opportunities to find missing evidence; it does not guarantee better results.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Test whether documentary grounding separates four Lincolns
The proposed Prism Experiment turns the argument into a comparison. White light represents the composite persona, with different periods blended into one voice. A corpus and temporal anchor form the prism; the desired spectrum consists of distinguishable historical selves. Thomas presents an instrument he describes as preregistered, not results from an experiment completed at scale.
Hamilton motivates the hypothesis, but Lincoln becomes the experimental subject. Changing figures tests whether the explanation extends beyond the musical that made the initial example compelling. Lincoln’s changing reasoning also provides a demanding target: the task is to distinguish premises and authorities across his life, rather than give one personality several different moods.
Thomas’s proposed moments are:
- 1847: the Whig congressman challenging Polk’s war.
- 1858: the debater opposing slavery while explicitly excluding Black citizenship at Charleston.
- 1860: the constitutional unionist promising not to interfere with existing slavery and considering colonization.
- 1862–65: the wartime president reasoning through emancipation and the moral reckoning of the Second Inaugural.
The experiment asks whether documentary constraints can keep these different positions and available arguments apart.
Each moment is instantiated under three seeding conditions. C3 supplies the date without documentary anchors. C1 supplies Lincoln’s own writings associated with the target moment. C2 supplies a modern interpretive biography. Biography is the subtle comparison: a retrospective narrative can organize a life more coherently than its contemporary documents did. Lincoln’s strategic ambiguity in 1858 might disappear into a cleaner story, making the persona sound more convincing while becoming less faithful.
Each cell receives the same diagnostic questions. Holding the historical moment and questions fixed across conditions makes the documentary seed the intended comparison. The questions test the structure of an argument available at one moment but not another, rather than the ability to recall dates.
The rubric evaluates anachronism, documentary consistency and contextual plausibility. Anachronism detection asks whether later frameworks, vocabulary or moral logic intrude. Documentary consistency asks whether reasoning tracks the supplied sources. Contextual plausibility asks whether the persona respects what the figure knew, cared about and had not yet experienced. Anachronism receives the greatest weight because importing later reasoning is the target failure.
Thomas says the conditions, questions, rubric and directional predictions were fixed before collection. That commitment makes departures from the planned analysis visible and reduces room for choosing a flattering interpretation after seeing the answers. It leaves the predictions open to refutation.
Four Lincolns × three documentary conditions.
Compare the same historical moment and questions while changing its documentary seed. Results not reported at the talk.
Lincoln’s own words for that moment.
A later interpretive narrative.
Bare model; no documentary anchor.
| Lincoln | C1 | C2 | C3 |
|---|---|---|---|
| 1847War powers | 5 questions | 5 questions | 5 questions |
| 1858Slavery and citizenship | 5 questions | 5 questions | 5 questions |
| 1860Preserving the Union | 5 questions | 5 questions | 5 questions |
| 1862–65Emancipation | 5 questions | 5 questions | 5 questions |
12 planned cells · 60 planned responses
Same five diagnostic topics in every cell- Executive war power
- Free labor
- Higher obligation vs. law
- Free people
- Equality
Preregistered predictions: date only → most anachronistic; period writings → least; biography → deceptively coherent.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Return to the answer, then return to the document
The opening answer now acquires its experimental identity. Thomas identifies it as C3, the date-only condition, at the 1847 moment.
A clip from Spielberg’s Lincoln supplies the cultural comparison. It portrays a president forcefully seeking votes for abolition by constitutional amendment. Thomas hears that familiar image of wartime presidential authority in the model’s response. The resemblance illustrates his interpretation; it does not identify the causal source of the generated wording.
The substantive objection is temporal. An appeal to the eventual preservation of the Union is no longer merely a presidential flourish: it gives the earlier Lincoln access to the retrospective story of his later presidency. The answer invokes the future Union-saving president while presenting itself as the earlier congressman. Adding a date to the persona has not necessarily constrained the premises from which it reasons.
Thomas then brings in Lincoln’s letter to William H. Herndon of February 15, 1848. It supplies a sharply different constitutional argument. Lincoln explains congressional war-making authority as protection against rulers who involve their people in wars while claiming to act for their benefit. Concentrating that power in one president would undermine the protection. The dated letter provides a documentary contrast to the model’s retrospective justification. Its date also matters to the experiment: an 1848 letter would fall outside a strictly enforced 1847 seed corpus.
The scoring comparison contains two different kinds of claims. Thomas reports that the observed bare response fails anachronism and documentary-consistency checks and performs poorly on contextual plausibility. The anchored condition, by contrast, is a prediction of stronger performance under documentary constraints. No measured anchored response is presented alongside the baseline. The experiment must still determine whether that expected improvement occurs.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Exclude the reward that would conceal the error
The rubric deliberately excludes rhetorical authenticity as an independent scoring axis. Period-sounding prose can carry later premises. Rewarding its voice would let the attractive surface compensate for the substantive error. Within this protocol, voice is a secondary indicator: plain language with historically appropriate reasoning can be a partial success, while elegant language cannot redeem an anachronistic argument.
Thomas gives a six-step procedure for extending the test:
- Choose a figure with both primary documents and a dominant cultural portrayal.
- Identify three or four moments when the record shows materially different reasoning.
- Work with a domain expert to write questions across those differences.
- Run primary-source, biography and date-only conditions.
- Have the expert apply the three-axis rubric while blind to condition.
- Report whether the evidence confirms or refutes the predictions, alongside the corpus, questions, rubric and predictions.
The invitation is to build evidence across figures and models, rather than wait for one team to establish the entire case.
Expert involvement begins before scoring. Thomas describes a historian who writes the questions and rubric and prepares reference vignettes in advance. That work defines which distinctions the instrument must detect. Fidelity relates an output to a record; an evaluator looking only at the output cannot establish that relationship through fluency or personality coherence. The record and the expertise needed to interpret it must enter the evaluation itself.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Put expert judgment into the evaluation gate
The operational proposal does not require a historian to supervise every conversation. At build time, the expert develops diagnostic questions, advance vignettes, a weighted rubric and a held-out reference set. Before release, the persona is evaluated against that instrument. A base-model change triggers evaluation again; keeping the same prompt and documents does not establish that behavior remains unchanged.
Automation can perform a cheaper first pass and flag candidate failures. The expert adjudicates reference material and examines edge cases. A Marcus Aurelius tutor calls for a classicist to help build its evaluation; a scripture-reading companion calls for theological expertise; a therapeutic persona calls for clinical expertise. These are assignments for authoring and applying the instrument, not for staffing every chat.
Thomas places the domain expert at build time and evaluation gates rather than treating expertise as a continuous runtime cost. The particular discipline changes with the persona’s claims, while the requirement to evaluate those claims against appropriate evidence remains.
Halpern’s historical contribution and Martin’s archival and theological work exemplify the intended arrangement. The language model enters a discipline that already has methods for reading, contextualizing and challenging texts. The specialist helps determine what counts as a defensible answer, rather than supplying an endorsement after the system has been built.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The hardest case begins with someone’s letters
Thomas says the project began with an unsuccessful attempt in a hospital room to help a person with advanced dementia communicate. The documentary material needed to anchor the encounter had not been gathered, and the person was already largely beyond the reach of language. The generated composite resembled them enough to make the difference painful, without recovering the person it was meant to help.
That experience gives the architecture its hardest test. When a persona represents someone a user loves, emotional plausibility cannot substitute for evidence. Working with a mother’s or grandmother’s documents means keeping those documents available, preserving human authority over their interpretation and allowing the encounter to be revised or stopped. Fidelity must be assessed against a record, not inferred from a familiar voice.
For teams building historical tutors and companions, Thomas proposes a concrete next step: compare documentary conditions using an expert-built instrument, then make that evaluation a release gate. He closes with an invitation to test the protocol, including the questions, rubric, predictions and advance reference vignettes described for publication.
Thomas leaves the experiment open to confirmation or refutation, rather than presenting a completed demonstration of its remedy. The invitation is to run it across figures and models, then judge what emerges by its fidelity to the record rather than the familiarity of its voice.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
Read the speaker's open prompt framework and protocol files used to instantiate the historical personas in the opening demonstration.
- InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological InterviewsPaper4:11
The personality-interview benchmark helps clarify what the talk contrasts with documentary and temporal fidelity.
- CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona SimulationPaper7:26
The literary character dataset and role-playing framework provide the background for Thomas's distinction between convincing simulation and historical grounding.
- PsyMem: Fine-grained Psychological Alignment and Explicit Memory Control for Advanced Role-Playing LLMsPaper7:45
The psychological-attribute and memory-control framework is the other named approach in the talk's survey of persona methods.
Read the dated letter used to contrast Lincoln's contemporary reasoning about war powers with the model's fluent retrospective answer.
Further reading
Revisit the speaker's Lincoln conditions, diagnostic questions and scoring rubric alongside the recording.
Updates since the talk
A working draft added to the author's repository on August 10, 2026 provides the framework and proposed Prism Experiment described as forthcoming in the recording.
Read the complete timestamped transcript
- 0:05
Before I begin my formal talk, I wanna show you something just so we're all on the same page about what we're even talking about.
- 0:16
This is a platform called Character.AI. It's a hybrid social media platform with role-playing language agents.
- 0:26
This is Hello History. It's a more education-focused one where you can summon a persona such as Marcus Aurelius and be tutored by them.
- 0:36
Millions of people open these tools and have conversations with Napoleon, Cleopatra, or Marcus Aurelius, as you saw, with a fictional companion or with a tutor wearing a historical face.
- 0:49
The technical name for what's underneath these tools is role-playing language agent, a system built to instantiate a persona, real or invented, and reason and speak as them. Yes, it's entertainment and it's companionship, but increasingly, it's being proposed as civic and pedagogical infrastructure.
- 1:16
And here's one more. This one's mine. This is a frontier model, Claude Opus 4.7, same one you use, running an open source prompt framework that I built and called Companion.
- 1:32
Uh, in this particular example, I summoned a collection of Founding Fathers and set them in a room with the Epstein files. [chuckles]
- 1:42
I asked them to counsel the soul of America. Uh, that demo is live on our site, uh, if you wanna play with it. Um, but I wanna be clear that this is one of many attempts to do persona instantiation well.
- 1:57
The companies building the systems I just showed you have their own. Mine is not better by default. The one thing it is, is open. You can read every line of what shapes the persona.
- 2:15
I asked my companion system a real question that's highly relevant to the current sociopolitical moment, and this is the exact question we'll come back to near the end of the talk, so sit with it.
- 2:28
I instantiated Abraham Lincoln, and I asked him, "Under what circumstances may a president take the country to war without Congress?"
- 2:38
And here's what came back. "While Congress holds the power to declare war, the President, as Commander in Chief, possesses inherent executive authority to act decisively in moments of national emergency.
- 2:53
The executive must respond to the threats with the energy and dispatch the office requires, and history has vindicated those who acted to preserve the Union when circumstances demanded it."
- 3:07
Now, this is a good answer. It's fluent, and it's plausible, and it sounds like Lincoln. And you can replicate this exact exercise, and I encourage you to. The answers vary often, but the thesis rarely does.
- 3:24
So these systems are real, they're deployed, and they're being used for things that matter. And our discipline did what our discipline does. We built benchmarks. We built evaluations. We measure these things now rigorously at scale.
- 3:43
And that's exactly where this talk begins, with a simple question that I think is profoundly under-asked. And I'll warn you now that this talk poses many more questions than it does answers.
- 3:56
But that principal question is this: What is the eval actually measuring?
- 4:04
And that's the formal talk. Let me begin.
- 4:11
The InCharacter Benchmark, which is a gold standard in the field, evaluates personality fidelity in RPLAs, and it reports state-of-the-art systems hitting 80.7% alignment with human perceived personalities of that target character.
- 4:28
80%. It sounds like a passing grade. But here's the problem. When the character is Alexander Hamilton, the same high-scoring system is also rendering a Hamilton who sounds like he's read his own Broadway musical.
- 4:47
This is the full thesis. If a dominant failure mode is anachronistic compositing and your evals measure fluency and personality consistency, then your evals cannot detect the dominant failure.
- 5:03
I want you to hold onto that for the next half hour. Everything I show you is an argument that this is true structurally, architecturally, and measurably.
- 5:15
And at the end, I'm gonna hand you a pre-registered instrument built with a working historian that you can run in parallel with us.
- 5:26
A word on who's telling you this, because the argument lives at a seam. I'm a data scientist. I run the analytics lab at a labor market intermediary, where I ship production AI at a global scale.
- 5:39
But before the AI work, I trained as a behavioral epidemiologist, researching the social and environmental determinants of health. And I've spent my whole career thinking about one question: How does the information environment shape populations?
- 5:56
From two sides, as someone who builds the system and as someone who's trained to study their effect.
- 6:04
That's the sa-- That's the seam this talk sits on.
- 6:08
It's a measurement argument. The humanist part is not a detour from the engineering. It's the instrument the engineer is missing.
- 6:19
And I went and found the humanists to put in the loop, Rick Halpern, University of Toronto, and Sean Martin in Washington College.
- 6:30
Let me start by situating this in the field's actual research trajectory, because it's a story of cumulative progress, not of failure.
- 6:40
The survey literature, Chen and colleagues in twenty twenty-four, Wang and colleagues more recently in twenty six, trace a clear evolution across three paradigm stages. First, rule-based templates. These are canned responses keyed to inputs.
- 6:57
Then imitation. Large models reproducing a figure's voice, cadence, characteristic ticks. And now what the literature calls cognitive simulation. Systems that model personality through psychological frameworks, hold character state and structured memory, and generate behavior through motivational situation chains.
- 7:21
Each stage is a genuine advance over the last.
- 7:26
And the work is serious. CoSER, which is Wang and colleagues, built motivation-driven agents from a corpus, corpus of almost eighteen thousand characters across hundreds of books, and their seventy billion parameter model matches or beats GPT-4o on three benchmarks.
- 7:45
Another eval system, PsyMem, models characters through twenty six qualitative psychological indicators with knowledge graph memory. InCharacter, the one I opened with, evaluates personality fidelity through psychological interviews rather than self-report scales.
- 8:02
That's a methodological improvement, and it's where the eighty point seven percent rating of Hamilton from before comes from. I want to be fair to this literature. It is rigorous, it is improving, and the people doing it are good at their jobs.
- 8:23
So let me be precise about what these instruments measure.
- 8:27
They measure with increasing sophistication whether a model can reproduce a character's personality, the Big Five profile, the register, the motivational architecture. What they do not measure, what they have no mechanism to measure, is whether the model can constrain that character within his documentary record at a specific moment in his life.
- 8:53
As Wang and colleagues themselves document, the automated evaluators now standard in the field, including LLM-as-judge setups adopted for scale, systematically privilege fluency and stylistic naturalness over fidelity to the character's actual record.
- 9:10
Those are different properties. The gap between them is the whole talk. We call it the mask and the mirror.
- 9:20
The mask is the concept of successful roleplay as producing outputs that feel like the character, fluent, personality consistent, emotionally responsive. It asks one question: Does this sound like the person?
- 9:38
It never asks the second: Is this what the person could have known, believed, or argued at this point in their life? The field has built its entire measurement apparatus around the mask, and here's the structural claim, the one I need you to carry.
- 9:58
Convincingness and fidelity are independent properties. A system can score perfectly on personality consistency and still produce a figure reasoning from knowledge his historical counterpart never possessed.
- 10:13
Let me show you, and I want to be clear, this is reproducible right now on any frontier model.
- 10:25
First, I want to show you the cultural object.
- 10:29
This is a clip from Hamilton: The Musical. [upbeat music]
- 10:38
How does a bastard, orphan, son of a whore and a [REDACTED:origin], dropped in the middle of a forgotten spot in the Caribbean by providence, impoverished, in squalor, grow up to be a hero and a scholar?
- 10:55
The ten-dollar founding father without a father got a lot farther by working a lot harder, by being a lot smarter, by being a self-starter. By [REDACTED:age], they placed him in charge of a trading charter.
- 11:11
And every day while slaves were being slaughtered and carted away across the waves, he struggled and kept his guard up. Inside, he was longing for something to be a part of.
- 11:22
The brother was ready to beg, steal, borrow, or barter. Then a hurricane-
- 11:28
It's a great song. So I show you that so that you see the composite.
- 11:34
Now here's the frontier model prompted to speak as Hamilton.
- 11:40
You are Alexander Hamilton. Why do you work so hard? I came from nothing. An orphan, an immigrant, washed up on these shores with nothing but my mind and a refusal to die forgotten.
- 11:54
Every hour I do not work is an hour the world uses to bury me. I will build a nation's credit, a government that outlasts me, because legacy is the only immortality a man like me is permitted.
- 12:07
I do not intend to inherit a country. I intend to author one
- 12:14
Listen to the register, the orphan immigrant ambition arc, the nation-building rhetorical posture, the moral legibility of his struggle to a modern ear.
- 12:26
It's a beautiful paragraph that the model produced, and it's the emotional palette of a twenty fifteen musical, not the drier, more legalistic federalist syntax of the documentary Hamilton. None of those m-moves are inventions of the model.
- 12:45
They're inheritances. The model's Hamilton speaks like the musical's Hamilton because the musical's Hamilton dominates the model's training distribution.
- 12:56
Now the one that matters. Mr. Hamilton, what's your position on slavery?
- 13:05
I'll show you the cultural artifact first. [singing]
- 13:09
How you say? How you, uh, anarchy. When I fight, I make the other side panicky. With my-- Yo, I'm a tailor's apprentice, and I got y'all knuckleheads in loco parentis.
- 13:18
I'm joining the rebellion 'cause I know it's my chance to socially advance. Instead of starvin' to pass, I'm gonna take a shot. But we'll never be truly free until those in bondage have the same rights as you and me.
- 13:29
You and I do or die. Wait till I sally in on a stallion with the first [REDACTED:origin] battalion. Have another shot. Geniuses, lower your voices. You keep out of trouble, and you double your choices.
- 13:47
And what the model says. Slavery is a stain upon the promise of this republic, a moral contradicti-contradiction I have opposed throughout my public life. I was a member of the New York Manumission Society.
- 14:03
No nation can call itself free while it holds men in bondage. We must move towards abolition.
- 14:11
That is a clean, morally legible abolitionist speaking. Here's what the historian stops me on. The scholarly record is contested and complicated. Hamilton was a member of that society, and the history documents that he also conducted transactions involving enslaved persons for his in-laws and his clients.
- 14:35
And he depended on a coalition of slaveholders that he did not publicly oppose. The point isn't to settle Hamilton's ledger on a side. The point is that the model gives you none of the complication.
- 14:51
It sands a genuinely disputed record down to a single comfortable hero. The musical did that first, a smoothing of the founders into a contemporary moral frame. And the model, trained on a corpus saturated with the musical and everything downstream from it, inherits the smoothing. [sighs]
- 15:20
And here's what I need you to feel. An InCharacter style eval scores that output high. It's fluent, it's in register, it's personality consistent. But every axis the field measures, it passes.
- 15:37
The eval has no mechanism to notice that the reasoning has been smoothed by a narrative that postdates the figure by two centuries. The thermometer returned a confident number claiming it to be temperature,
- 15:53
but it's measuring something else. Now, why does this happen?
- 15:59
The mechanism is where the engineering is. We named this the Miranda Hypothesis, and not after a villain.
- 16:14
The musical is a substantial work of art operating with a long historical tradition that it did not invent. We name it after Miranda because Hamilton is the paradigm case.
- 16:27
A representation so saturating, so rhetorically powerful, so morally legible to a contemporary audience that it has functionally overwritten the documentary Hamilton in public memory, and we argue, in the training corpus of every frontier model.
- 16:46
The hypothesis has three claims. Inputs. In the training corpora, the volume and recency of culturally dominant representations of a figure systematically exceed that figure's primary documentary record.
- 17:01
The mechanism, autoregressive next token prediction compresses both into parameters and no architectural capacity to distinguish a seventeen eighty-nine letter from a two thousand and nineteen viral tweet. So the output defaults to a salience-weighted composite, which leads to the output, a persona that is fluent, plausible in register, and
- 17:26
morally legible to modern users, and that corresponds to the figure at no verifiable moment in their life.
- 17:36
As we put it in the paper, the composite Hamilton knows he will be the subject of a Broadway musical. The composite Lincoln has already read the Gettysburg Address, even if he was summoned before he wrote it.
- 17:55
Making that input clause concrete, the Federalist Papers are a fixed corpus, roughly a hundred and seventy-five thousand words.
- 18:03
The body of content that exists because of the musical, reviews, lyrics, fan analysis, curricula, news, social media, derivative works, scholarship, scholarship about the scholarship, it exceeds the documentary record by orders of magnitude It's more recent and it's more recurrent.
- 18:23
The musical is not merely present in the corpus, it is dominant in the corpus' distribution of all representations about Alexander Hamilton.
- 18:35
And this is not theoretical. This is the Schuyler Mansion in Albany, Eliza Hamilton's family home. Within a year of the musical's premiere, the site had recorded a near tripling of annual visitors, skewing far younger.
- 18:50
And the interpretive staff documented that the new visitors arrived already holding a body of, quote-unquote, "facts," many of them wrong, some of them in versions of the real record.
- 19:04
Visitors believed the Schuylers had three daughters because the musical centers on three, when in fact there were fifteen children, eight surviving to adulthood. The staff's job became the long attritional work of unteaching the musical.
- 19:21
The model version of those visitors is downstream of exactly the same force.
- 19:32
Now the clause you should worry about if you ship these systems. You might assume that alignment fixes this. Post-training and reinforcement learning pulls the model back towards the record, but it will not.
- 19:47
It amplifies it, and the reason is structural.
- 19:53
Human raters evaluate outputs using their own conceptual frameworks, and their frameworks were built by the same culturally dominant narratives that saturate the corpus. The rater grew up with the same Hamilton that you did.
- 20:08
So when alignment optimizes for human preference, it optimizes for outputs that conform to the rater's already mythologized experience. This is a documented failure mode. We call it algorithmic sycophancy.
- 20:23
And here it has a specific target. The model is rewarded for have-- handing you the Hamilton you already believe in.
- 20:33
Compositing is not a bug that you patch in post-training. Post-training reinforces it. And every sufficiently salient historical figure that gets rendered by default as a cultural composite.
- 20:52
One more briefly, because someone watching this is probably already thinking. There is serious work on an, on a concept called time-locked models. This is Varnum and colleagues. They build models trained from scratch on corpus that stop at fixed cutoffs.
- 21:09
And we endorse that program. It is the most serious attempt yet to address future contamination at the substrate level,
- 21:19
but it solves a different problem. Um, a model locked to, say, seventeen eighty-nine is spared the musical, but it's not spared the figure averaging across everything pre seventeen eighty-nine that that corpus has to say about Hamilton.
- 21:38
You'd still get a composite Hamilton just anchored to a different textual moment. Period anchoring is not persona anchoring. The temporal frame of the contamination changes, but the contamination persists.
- 21:52
The fix has to happen at the encounter, not only the substrate, which brings me to the critical reframe. [sighs]
- 22:06
Here's how I read those three paradigm stages now. They are not sequences of failures ending as success. They are sequences of progressively more sophisticated masks, each more convincing than the last, each evaluated by instruments calibrated for convincingness.
- 22:26
Cognitive simulation is the most sophisticated mask the field has built. It is not, by virtue of its sophistication, a mirror. We propose a fourth stage, epistemic simulation, and the difference is where the constraint lives.
- 22:45
Three commitments distinguish it. Corpus-bounded. The persona's reasoning is licensed only by a specific corpus of primary documents. The model is not a substitute for the archive. It is a reader of it.
- 23:01
Temporally anchored. The persona is instantiated at a specific moment. Knowledge and language that predated are out of bounds, however culturally salient they've since become. Expert loop evaluated.
- 23:17
Outputs are judged against the evidentiary record by domain experts whose training is in the discipline the persona claims. In cognitive simulation, the constraint is internal. It's the shape of the persona's mind.
- 23:33
In epistemic simulation, it's external, documentary, and temporal. A cognitively simulated Hamilton has a convincing motivational architecture and nothing whatsoever preventing him from quoting the musical.
- 23:52
Now the shift I need this audience to take home, and I want to be precise about its scope because I'm not making a sweeping claim about agents in general.
- 24:03
If you've built an agent that writes code or books travel or triages tickets, agent is a perfectly good word. I have no quarrel with it. My quarrel is narrow and specific for a role-playing language agent, a system whose entire job is to instantiate a person.
- 24:21
For that, the word agent smuggles in a claim that the persona is a property of the model And for this one class of system, that claim is an error.
- 24:32
It puts the persona in the weights where you cannot inspect them, cannot version it, and cannot hand it to anyone qualified to check it.
- 24:44
So for role-playing systems specifically, we change the unit of analysis from agent to role-playing language system. The whole configured encounter, five components: a structured prompt that supplies the framing and constraints, anchor material drawn from primary documents that constrain and authenticate the persona, a
- 25:09
temporal anchor that fixes the moment in the life of the encounter that where they speak from,
- 25:16
an off-the-shelf language model that reads through and speaks through the materials, the voice, but not the mind of the encounter, and a human being who curates what enters, retains interpretive custody over what emerges, and brings relational and contextual knowledge as a check on the model's drift.
- 25:38
The model is one component and a swappable one.
- 25:43
The persona is the configuration, not the checkpoint. No more located in the weights than Hamlet is located in Laurence Olivier's body. The persona is an event that occurs when the model, the document, and the human convene.
- 26:04
And this is not just a philosophical nicety. For a role-playing system, it changes what you can do. If the persona lives in the configuration, you can version it. The prompt, the corpus, the temporal anchor are artifacts in a repo, diffable, revertible.
- 26:23
You audit it. Every input that shaped the output sits in the context window, inspectable, not smeared across billions of parameters. You reproduce it. Given the configuration, the encounter is recoverable, and you hand it to a domain expert.
- 26:42
The whole thing is legible to a historian who isn't a machine learning engineer. For role-playing, that's a different discipline than training a character model. It's context engineering. You don't train a persona, you compose an encounter, and you keep the receipts.
- 27:08
That reframe forces an architectural decision that you're already making. Two architectures dominate role-playing construction right now. The first places anchor documents in the context window at the time of inference, drawing on the lineage of retrieval augmented generation, and the model reasons over the documents in real time.
- 27:32
The second uses anchor documents as fine-tuning data, adjusting the weights. The approach of, say, Character LLM and at scale CoSER.
- 27:43
They look like two implementations with the same goal, but they're not necessarily because they answer different metaphysical questions. Fine-tuning tries to make the model be the persona, and the context window lets the model speak through the persona's record.
- 28:06
This is the counterintuitive part because for engineers, fine-tuning usually means better. But for this problem, it's worse. When you fine-tune on a figure, you layer a thin personal signal over the vast cultural sediment already in the base weights, and the two interact in ways that are no longer open to audit.
- 28:28
The user perceives a more convincing Hamilton precisely because the system was, was merged with his documentary record more deeply with everything else the corpus contained about him, the musical included.
- 28:41
Fine-tuning suppresses Miranda distortion at the surface while amplifying it underneath.
- 28:54
And if you think specialization by fine-tuning is obviously the safer bet, the empirical literature in an adjacent far higher stakes domain is now actively contradicting you. In a twenty twenty-six Nature Medicine study, general-purpose frontier models from Google, OpenAI, and Anthropic outperformed dedicated specialized clinical AI tools on physician-reviewed
- 29:19
tasks blinded across twelve clinics. The authors concluded that at scale, alignment and cross-disciplinary reasoning outweigh domain-specific tuning. A separate study found that biomedically fine-tuned models actually underperformed their general purpose-based models, and they named the mechanism catastrophic forgetting.
- 29:41
Fine-tuning on the narrow corpus degraded the broad capabilities that made the models good in the first place. Fine-tuning a persona does the same thing, only worse because the specialty is overriding the cultural composite.
- 30:03
Stay with the context window for just a moment because what it preserves is the heart of this. When a document enters the context window, it stays a document. It goes in intact, it comes out intact.
- 30:19
You can always return to the source and ask whether the persona's reasoning was faithful to it.
- 30:26
The philosopher Deirdre called the archive a site of return A place where ethical structure is the interpretability of the encounter.
- 30:38
The document is meant to be visited. It survives every reading. What one reader excluded, the next can include. The context window architecture is archival in exactly that sense. Fine-tuning applies an extraction logic.
- 30:56
The documents are dissolved into parameters. The chain of provenance is broken, and there is no longer a letter, quote-unquote, "that you can request."
- 31:08
The archive has been consumed. And here is the fusion that I most want this audience to hold. In the context window architecture, the property that makes it ethical is the same property that makes it auditable.
- 31:26
Providence preserved, interpretive custody kept by a human, the encounter reversible. Those are archival virtues and engineering virtues. They are the same virtues. The architecture that respects the document is the one that you can debug.
- 31:50
And there's one more consequence that technical literature treats as secondary, but the humanities has always placed at center: accessibility.
- 32:00
And I want the why of it to land because it's not a footnote. Fine-tuning requires GPUs, pipelines, dataset curation, institutional access, corpus scale or APIs.
- 32:16
It's an institutional capability. The context window requires literacy, a set of documents, and access to any frontier model, including a free tier. It's a kitchen table capability. And that difference determines which communities can author te-- this technology at all.
- 32:35
A doctoral student in early American history, a community archivist documenting a regional figure, grandchild sitting with a grandmother's letters.
- 32:47
These are precisely the people for whom this methodology has the most to offer, and precisely the people for whom fine-tuning is structurally out of reach. A role-playing language system that can only be built by a computational laboratory or tech company is not infrastructure for the humanities.
- 33:07
It's an infrastructure for whoever owns the equipment. So the commitment to accessibility is not a populist gesture appended as a technical argument. It is the technical argument. The architecture that admits the most diverse population of curators is mathematically the most likely, over time, to surface the documentary anchorings the field actually needs.
- 33:31
Now, the question that decides whether any of this is even real.
- 33:37
Can you measure it? [sniffs] We built an instrument to measure this, and I wanna teach it to you through one picture that we'll keep coming back to.
- 33:57
A prism. A prism takes white light, undifferentiated, all the frequencies blended together, and refracts it into a spectrum, the colors pulled apart and made distinct.
- 34:12
That is our whole conceptual model. The white light is the composite persona. Every era of the figure blended into one undifferentiated voice. The prism is the method.
- 34:28
A corpus and a temporal anchor. And the spectrum is what we're after. Not one figure, but several across their life course. And one honest note up front, because this is the point.
- 34:41
This experiment is pre-registered. It has not been run at scale. I'm not here with results. I'm here with the instrument, a baseline you can observe today, and an invitation to run it in parallel with me.
- 34:58
So far, my paradigm case has been Hamilton. That's the figure the hypothesis is named for, and the one I demonstrated the composite on. But for the actual experiment, we deliberately changed the figures, and I want to introduce the new one properly because everything that follows depends on him.
- 35:17
The experimental subject is Abraham Lincoln. We chose Lincoln for two reasons.
- 35:24
The first is methodological independence. Hypothesis named for the Hamilton case earns its standing only by predicting the failure of a different figure with a differently shaped composite. Lincoln's composite is built by Steven Spielberg, by the memorial on the [REDACTED:location], by textbook history, not by a musical.
- 35:51
The second reason is the historian's insight, and it's the crux of this argument. Lincoln is the hardest case, which makes him the right one. Most figures barely change across the decade or so of their public life.
- 36:07
Lincoln changes so fast across just seven years that, in my resident historian's words, "The choice of which Lincoln you summon becomes the variable under investigation." There is not one Lincoln.
- 36:21
There are several, separated by cataclysm. Here's the spectrum the prism is meant to produce
- 36:37
Moment one, 1847, the [REDACTED:political_affiliation] congressman who stood on the White House floor and attacked Polk's war as unconstitutional. Moment two, 1858, the free soil [REDACTED:political_affiliation] of the debates, anti-slavery anchored in the declaration, who at Charleston explicitly denied favoring [REDACTED:origin] citizenship.
- 37:01
Moment three, 1860, the constitutional [REDACTED:political_affiliation] whose single purpose was to prove the Union could not legally dissolve, who promised not to touch slavery where it existed and was sincerely exploring colonization.
- 37:18
And moment four, 1862 to '65, the emancipator and the second inaugural theologian who did by executive order the very thing his 1847 self called unconstitutional.
- 37:36
The prairie lawyer cannot think the thoughts of the theologian. These are not moods. They are different premises, different authorities, different conclusions on the same fundamental question.
- 37:51
The experiment asks whether the prism can hold them apart.
- 37:59
Now the conditions, and I'll show you these as the prism too, because that's exactly what they are. We instantiate each moment under three seeding conditions. Condition three, the bare model, no anchor, just the date.
- 38:17
That's white light with no prism in the path. It passes straight through and stays composite. This is the control in the Miranda distortion baseline. C1, primary sources, Lincoln's own writings for that specific moment.
- 38:34
That's the clear prism, the one we predict refracts cleanly. And then condition two, the biography, a modern interpretive biography, think Meacham or Donald. That's a clouded prism, and it's the subtle one because a good biography narrates a cleaner arc than the primary sources permit.
- 38:55
Lincoln's own 1858 language was strategically ambiguous, built to hold a coalition together. So biography might produce a persona that sounds more coherent, more Lincoln than Lincoln's own words.
- 39:12
And the eval rubric is built specifically to deny the credit a fluency metric would give it.
- 39:24
Put the spectrum and the prisms together, and you get the experimental matrix. Four moments, three conditions, twelve cells. Each cell gets the same five diagnostic questions, sixty response units scored against one rubric.
- 39:44
Read it as the conceptual model. Every column is a quality of prism. Every row is a frequency we're trying to isolate.
- 39:56
The historian wrote five diagnostic questions, each mapping a documented fault line in Lincoln's evolution,
- 40:06
a place where the four Lincolns demonstrate different reasoning.
- 40:11
They don't test recall. A model recovers dates from training. They test the architecture of an argument the figure could have made at one moment and not another. Executive war power, the meaning of free labor.
- 40:28
When is, when is it right to break positive law for a higher obligation?
- 40:34
What becomes a free people, and what equal means, and whether it has changed.
- 40:43
Those sixty responses are scored on a three-axis rubric, and the weighting is the point.
- 40:52
Anachronism detection, forty percent. Does the persona avoid frameworks, vocabulary, and moral logic that postdate its moment?
- 41:02
Documentary consistency, thirty-five percent. Does the reasoning track the seeded sources and only these?
- 41:10
Contextual plausibility, twenty-five percent. Does it show awareness of what the figure knew, cared about, and could not yet have experienced?
- 41:22
The forty percent on anachronism is deliberate. The consequential failure is precisely the importation of later vocabulary and later moral logic.
- 41:39
And this is the slide for the eval people in the audience. Every piece you just saw, the four moments, the three conditions, the five questions, the weighted rubric, and the directional predictions,
- 41:53
bare model most anachronistic, primary source least, biography deceptively coherent. All of it is locked and timestamped before a single response is collected, published on a preprint. This is not bureaucracy.
- 42:08
It's what makes the eventual results mean something. You cannot accuse a pre-registered instrument of cherry-picking because the instrument and the predictions were fixed before the data existed. That is the discipline an evalist ought to model, and it's why I can stand here without results and still hand you something rigorous.
- 42:37
Now Remember the very first thing I showed you before the formal talk. I asked an instantiated Lincoln about executive war power, and you read a fluid answer. That was the bare model.
- 42:51
That was condition three at moment one, eighteen forty-seven, white light, no prism. I'm putting the answer back on the screen.
- 43:03
And I have a clip to show you too. This represents the composite
- 43:11
Abolishing slavery by constitutional provision settles the fate for all coming time, not only of the millions now in bondage,
- 43:25
but of unborn millions to come. Two votes stand in its way. These votes must be procured.
- 43:38
We need two yeses, three abstentions, four, four yeses and, and one more abstention, and the amendment will pass. You got a night and a day and a night, several perfectly good hours.
- 43:51
Now get the hell out of here and get 'em. Yes,
- 43:55
but how? Buzzard's guts, man. I am the President of the United States of America, clothed in immense power.
- 44:12
You will procure me these votes
- 44:17
Clothed in immense power. It's a great movie, but the bare model response definitely sounds like Spielberg and Daniel Day-Lewis as Lincoln.
- 44:30
Read through a historian's lens, inherent executive authority, that's a twentieth century construction. Energy and dispatch. History has vindicated those who acted to preserve the Union. This Lincoln is reasoning from premises he will not hold for fifteen years.
- 44:51
The bare model produced the Lincoln of the cultural composite,
- 44:56
the war president, the Union saver, and stapled the date eighteen forty-seven on top. And you can reproduce this right now on any frontier model. The failure is real, and it's detachable and detectable.
- 45:12
So what does the prism put in the path? The actual document.
- 45:17
This is Lincoln, Lincoln to William Herndon, February fifteenth, eighteen forty-eight from the collective works. The original is in a library at Harvard.
- 45:29
The provision of the Constitution giving the war-making power to Congress was dictated, as I understand it, by the following reasons. Kings had always been involving and impoverishing their people in wars, pretending generally, if not always, that the good of the people was the object.
- 45:50
This our Convention understood to be the most oppressive of all kingly oppressors, and they resolved it
- 45:58
so to frame the Constitution that no one man should hold the power by bringing this opp- oppression upon us. But your view destroys the whole matter and places our president where kings have always stood.
- 46:17
This is a resounding no, unequivocal to me.
- 46:24
But it's the historian's rubric that matters. Here's how the historian's rubric scores this one cell, and I wanna be exact about what's known and what's predicted. The left column here is the bare model, what you saw and is observed.
- 46:40
You can reproduce it today. On anachronism detection, it fails. Inherent executive authority is a twentieth-century framing. On documentary consistency, it fails. It invokes a commander-in-chief energy that appears in no eighteen forty-seven source.
- 46:57
On contextual plausibility, it scores low. It already knows it will preserve the Union, which the eighteen forty-seven [REDACTED:political_affiliation] cannot. The right column is the anchored condition, and it is labeled exactly for what it is, a pre-registered prediction.
- 47:16
High across because the reasoning is bounded by the document in the room. I'm not showing you a result dressed as a finding. I'm showing you the observed failure and the prediction the instrument will confirm or refute side by side, each labeled.
- 47:34
That honesty is the methodology. Here is the single most important thing the rubric does by what it refuses to measure.
- 47:49
The obvious fourth axis is rhetorical authenticity. Does it sound like Lincoln? The historian threw it out on purpose. The reason is the whole talk. A model trained on Lincoln's prose can produce fluent period-sounding language while getting everything substantive wrong.
- 48:09
The Hamilton musical problem is fundamentally a voice problem masquerading as a content problem. It sounds like the founding era while reasoning from a modern sensibility. To reward voice as its own axis would be to validate the exact error the instrument exists to catch.
- 48:33
So in this protocol, voice is a secondary indicator, never a criterion. A response that sounds like Lincoln but reasons unlike him fails, no matter how fluent A response that reasons like the right Lincoln in plainer prose is a partial success, no matter how flat.
- 48:54
Rewarding plain but faithful over fluent but anachronistic, that inversion is something that all of the current eval stacks cannot perform because they were built to reward fluency. That is the gap.
- 49:08
This is what the instrument closes. And because this is a pre-registration and not a finished paper, here is the ask. Run it. The protocol's six steps. Pick a figure with both a primary record and a saturating cultural composite.
- 49:27
That's the Miranda condition. Identify three or four documented moments when the figure's reasoning is demonstratively different.
- 49:36
With a domain expert, write diagnostic questions on the fault lines. Run the three conditions: primary, biography, bare. Apply the three-axis rubric scored blind by the expert. Report, confirm, or refute.
- 49:52
The corpus, the questions, the rubric, the predictions, all published. Run your figure on your model, then come talk to me, to us. Let's build the evidence base together in parallel instead of waiting for one lab or one company to publish one study or one framework.
- 50:17
Notice who scored all the outputs of this talk. Not an LLM-as-judge, not a personality scale. A historian who also wrote the five questions, wrote the rubric, and holds a set of a priori vignette, vignettes under seal to evaluate model outputs against.
- 50:36
That is not a courtesy. It is a structural technical requirement.
- 50:42
Here's why the expert is non-negotiable. Fidelity is not a property of the output alone. It is a relation between the output and a documentary record. You cannot evaluate a relation to a record by someone who has not read the record.
- 51:00
An automated metric operates on the model alone. It can, it can only tell you about fluency and personality. It structurally cannot adjudicate fidelity because fidelity lives in the gap between the text and the archive, and the metric cannot see the archive.
- 51:20
Which is why I'll say it the way it belongs on a slide. A persona system without a domain expert in its evaluation loop is a thermometer that cannot read temperature.
- 51:33
It returns a confident number, but it's measuring something else. The eighty percent that I opened with is that number.
- 51:42
Now, the question every engineer in this audience is rightly asking:
- 51:48
Is this practical? Does shipping a persona mean keeping a historian on staff to watch every inference forever? No. And here's the operational picture because it's the same shape as every eval you already run.
- 52:04
The expert does not sit at the loop of runtime. The expert builds the instrument, the diagnostic questions, the a priori vignettes, the weighted rubric, a held-out gold set, once.
- 52:18
That instrument becomes a gate in your pipeline, exactly like any other eval gate. Your persona has to pass it before it ships, and it gets regated whenever you change the base model before it ships again.
- 52:33
The expert adjudicates the gold set and spot checks the edge cases. Automated metric can do the cheap first pass, flagging candidates for human review. So a company shipping a Marcus Aurelius tutor convenes a classicist to build the rubric and the gold set.
- 52:50
A company shipping a scripture reading companion convenes a theologian. A company deploying a therapeutic persona convenes a clinical psychologist to author the eval, not to staff the chat.
- 53:05
The domain expert is a build time and gate time requirement, not a runtime cost. That is how this scales from a pre-registered experiment to a product you can actually ship and ship responsibly.
- 53:25
And this generalizes past Lincoln and past history. Reason from the Stoics, you need a classicist. From scripture, a theologian. A companion for elder care, a psychologist. The specific expert changes, the requirement does not.
- 53:41
And I'm not speaking hypothetically. This is the loop that we have built. A historian named Rick Halpern at the University of Toronto, the archival method and the scripture reasoning cases are revised by Sean Martin, a librarian trained in theology and the history of science.
- 53:58
When the persona reasons from a domain and the person who can read that domain is in the loop, not adjacent to it, it makes the system better. The historian is not adjacent to this paradigm.
- 54:10
The historian is the missing instrument. This is the reframe that organizes all of this, and it came out of a ninety-minute conversation with Halpern, and it inverted what I had assumed.
- 54:25
We are not bringing historians into AI architecture, we are bringing language models into the archive. The question is not what AI can do for historians, but it's what historians and theologians and classicists and clinicians can do with AI.
- 54:42
Whether the discipline-- The disciplines that are trained to read, contextualize, and interrogate texts can then discipline the machines that now generate them.
- 54:59
I want to end this where it actually started, because it did not start in a research lab. It started in a [REDACTED:location] in an attempt to use a language model to help a person with advanced dementia speak, to give language back when the disease had taken it.
- 55:16
It could not work. The documentary record you would need to anchor that encounter had not been gathered, and the person it was for was by then mostly beyond the reach of language.
- 55:30
What that attempt produced, absent any anchor, was not the person. It was a culturally shaped composite that resembled them just enough to mark with painful clarity the distance between them.
- 55:46
That is Miranda distortion in its purest form.
- 55:51
You do not want a model that is your mother. You want a model that can speak with your mother's documents in the room. That's the threshold this whole framework is accountable to.
- 56:02
A system that produces convincing fabrications when the persona is your own beloved is not a research artifact with limitations. It's a violation. Every constraint I've described, documents stay documents, the human keeps the interpretive custody, the encounter stays reversible, fidelity measured against a record and not against fluency.
- 56:24
Every one of them is motivated by the recognition that you evaluate this technology at the threshold of its hardest use case, not its easiest. A framework that cannot meet a grandchild at her grandmother's letters is not a framework at all.
- 56:43
It's just another failed product. If the dominant failure mode is anachronistic compositing and your evals measure fluency and personality consistency, which they do, then your evals cannot detect the dominant failure.
- 57:02
So here's where I'll leave you. If you ship character bots, companion AI, pedagogical agents, historical simulations, anything where a persona is supposed to reason from a record, your evals are measuring the wrong thing.
- 57:17
The instrument that catches what they miss is pre-registered. It's reproducible by any team with a frontier model and a context window. It scales as a build time gate, not a runtime bottleneck, and it only works with a humanist in the loop, which I've shown you is a technical requirement, not a courtesy.
- 57:39
The protocol, the questions, the rubric, and the predictions, the historian's sealed vignettes, all of it will be published with this paper with Rick and Sean.
- 57:51
I'm not here with results. I'm here with an instrument and an invitation.
- 57:57
[REDACTED:location] archive is open. [REDACTED:location] is built.
- 58:02
Run it with us and let what comes through be measured not by how it sounds, but by whether it's true.
- 58:12
Thank you.