← All AI Engineer talks

AI Engineer World's Fair 2026

The Miranda Hypothesis: How Hamilton (the Musical) Poisoned Your Persona Evals

About this talk

Jacob E. Thomas examines how role-playing language agents reproduce culturally popular composites of historical figures instead of reasoning from documents available at a specific historical moment. Using Alexander Hamilton, Abraham Lincoln and the InCharacter benchmark, he argues that personality fidelity, human-preference optimization and fine-tuning can conceal anachronistic reasoning. He outlines a preregistered, temporally anchored evaluation using historical source material, diagnostic questions, documentary consistency and contextual plausibility.

Chapters

  1. 0:05Role-playing agents and historical-persona demonstrations
  2. 4:11Why persona-fidelity benchmarks miss Hamilton's cultural composite
  3. 20:08Preference optimization, sycophancy and Miranda distortion
  4. 34:58A temporally anchored Lincoln experiment
  5. 46:57Contextual plausibility, preregistration and reproducibility
  6. 54:59Clinical motivation and closing call to evaluate truthfulness

Talk transcript

  1. 0:05

    Before I begin my formal talk, I wanna show you something just so we're all on the same page about what we're even talking about.

  2. 0:16

    This is a platform called Character.AI. It's a hybrid social media platform with role-playing language agents.

  3. 0:26

    This is Hello History. It's a more education-focused one where you can summon a persona such as Marcus Aurelius and be tutored by them.

  4. 0:36

    Millions of people open these tools and have conversations with Napoleon, Cleopatra, or Marcus Aurelius, as you saw, with a fictional companion or with a tutor wearing a historical face.

  5. 0:49

    The technical name for what's underneath these tools is role-playing language agent, a system built to instantiate a persona, real or invented, and reason and speak as them. Yes, it's entertainment and it's companionship, but increasingly, it's being proposed as civic and pedagogical infrastructure.

  6. 1:16

    And here's one more. This one's mine. This is a frontier model, Claude Opus 4.7, same one you use, running an open source prompt framework that I built and called Companion.

  7. 1:32

    Uh, in this particular example, I summoned a collection of Founding Fathers and set them in a room with the Epstein files. [chuckles]

  8. 1:42

    I asked them to counsel the soul of America. Uh, that demo is live on our site, uh, if you wanna play with it. Um, but I wanna be clear that this is one of many attempts to do persona instantiation well.

  9. 1:57

    The companies building the systems I just showed you have their own. Mine is not better by default. The one thing it is, is open. You can read every line of what shapes the persona.

  10. 2:15

    I asked my companion system a real question that's highly relevant to the current sociopolitical moment, and this is the exact question we'll come back to near the end of the talk, so sit with it.

  11. 2:28

    I instantiated Abraham Lincoln, and I asked him, "Under what circumstances may a president take the country to war without Congress?"

  12. 2:38

    And here's what came back. "While Congress holds the power to declare war, the President, as Commander in Chief, possesses inherent executive authority to act decisively in moments of national emergency.

  13. 2:53

    The executive must respond to the threats with the energy and dispatch the office requires, and history has vindicated those who acted to preserve the Union when circumstances demanded it."

  14. 3:07

    Now, this is a good answer. It's fluent, and it's plausible, and it sounds like Lincoln. And you can replicate this exact exercise, and I encourage you to. The answers vary often, but the thesis rarely does.

  15. 3:24

    So these systems are real, they're deployed, and they're being used for things that matter. And our discipline did what our discipline does. We built benchmarks. We built evaluations. We measure these things now rigorously at scale.

  16. 3:43

    And that's exactly where this talk begins, with a simple question that I think is profoundly under-asked. And I'll warn you now that this talk poses many more questions than it does answers.

  17. 3:56

    But that principal question is this: What is the eval actually measuring?

  18. 4:04

    And that's the formal talk. Let me begin.

  19. 4:11

    The InCharacter Benchmark, which is a gold standard in the field, evaluates personality fidelity in RPLAs, and it reports state-of-the-art systems hitting 80.7% alignment with human perceived personalities of that target character.

  20. 4:28

    80%. It sounds like a passing grade. But here's the problem. When the character is Alexander Hamilton, the same high-scoring system is also rendering a Hamilton who sounds like he's read his own Broadway musical.

  21. 4:47

    This is the full thesis. If a dominant failure mode is anachronistic compositing and your evals measure fluency and personality consistency, then your evals cannot detect the dominant failure.

  22. 5:03

    I want you to hold onto that for the next half hour. Everything I show you is an argument that this is true structurally, architecturally, and measurably.

  23. 5:15

    And at the end, I'm gonna hand you a pre-registered instrument built with a working historian that you can run in parallel with us.

  24. 5:26

    A word on who's telling you this, because the argument lives at a seam. I'm a data scientist. I run the analytics lab at a labor market intermediary, where I ship production AI at a global scale.

  25. 5:39

    But before the AI work, I trained as a behavioral epidemiologist, researching the social and environmental determinants of health. And I've spent my whole career thinking about one question: How does the information environment shape populations?

  26. 5:56

    From two sides, as someone who builds the system and as someone who's trained to study their effect.

  27. 6:04

    That's the sa-- That's the seam this talk sits on.

  28. 6:08

    It's a measurement argument. The humanist part is not a detour from the engineering. It's the instrument the engineer is missing.

  29. 6:19

    And I went and found the humanists to put in the loop, Rick Halpern, University of Toronto, and Sean Martin in Washington College.

  30. 6:30

    Let me start by situating this in the field's actual research trajectory, because it's a story of cumulative progress, not of failure.

  31. 6:40

    The survey literature, Chen and colleagues in twenty twenty-four, Wang and colleagues more recently in twenty six, trace a clear evolution across three paradigm stages. First, rule-based templates. These are canned responses keyed to inputs.

  32. 6:57

    Then imitation. Large models reproducing a figure's voice, cadence, characteristic ticks. And now what the literature calls cognitive simulation. Systems that model personality through psychological frameworks, hold character state and structured memory, and generate behavior through motivational situation chains.

  33. 7:21

    Each stage is a genuine advance over the last.

  34. 7:26

    And the work is serious. CoSER, which is Wang and colleagues, built motivation-driven agents from a corpus, corpus of almost eighteen thousand characters across hundreds of books, and their seventy billion parameter model matches or beats GPT-4o on three benchmarks.

  35. 7:45

    Another eval system, PsyMem, models characters through twenty six qualitative psychological indicators with knowledge graph memory. InCharacter, the one I opened with, evaluates personality fidelity through psychological interviews rather than self-report scales.

  36. 8:02

    That's a methodological improvement, and it's where the eighty point seven percent rating of Hamilton from before comes from. I want to be fair to this literature. It is rigorous, it is improving, and the people doing it are good at their jobs.

  37. 8:23

    So let me be precise about what these instruments measure.

  38. 8:27

    They measure with increasing sophistication whether a model can reproduce a character's personality, the Big Five profile, the register, the motivational architecture. What they do not measure, what they have no mechanism to measure, is whether the model can constrain that character within his documentary record at a specific moment in his life.

  39. 8:53

    As Wang and colleagues themselves document, the automated evaluators now standard in the field, including LLM-as-judge setups adopted for scale, systematically privilege fluency and stylistic naturalness over fidelity to the character's actual record.

  40. 9:10

    Those are different properties. The gap between them is the whole talk. We call it the mask and the mirror.

  41. 9:20

    The mask is the concept of successful roleplay as producing outputs that feel like the character, fluent, personality consistent, emotionally responsive. It asks one question: Does this sound like the person?

  42. 9:38

    It never asks the second: Is this what the person could have known, believed, or argued at this point in their life? The field has built its entire measurement apparatus around the mask, and here's the structural claim, the one I need you to carry.

  43. 9:58

    Convincingness and fidelity are independent properties. A system can score perfectly on personality consistency and still produce a figure reasoning from knowledge his historical counterpart never possessed.

  44. 10:13

    Let me show you, and I want to be clear, this is reproducible right now on any frontier model.

  45. 10:25

    First, I want to show you the cultural object.

  46. 10:29

    This is a clip from Hamilton: The Musical. [upbeat music]

  47. 10:38

    How does a bastard, orphan, son of a whore and a [REDACTED:origin], dropped in the middle of a forgotten spot in the Caribbean by providence, impoverished, in squalor, grow up to be a hero and a scholar?

  48. 10:55

    The ten-dollar founding father without a father got a lot farther by working a lot harder, by being a lot smarter, by being a self-starter. By [REDACTED:age], they placed him in charge of a trading charter.

  49. 11:11

    And every day while slaves were being slaughtered and carted away across the waves, he struggled and kept his guard up. Inside, he was longing for something to be a part of.

  50. 11:22

    The brother was ready to beg, steal, borrow, or barter. Then a hurricane-

  51. 11:28

    It's a great song. So I show you that so that you see the composite.

  52. 11:34

    Now here's the frontier model prompted to speak as Hamilton.

  53. 11:40

    You are Alexander Hamilton. Why do you work so hard? I came from nothing. An orphan, an immigrant, washed up on these shores with nothing but my mind and a refusal to die forgotten.

  54. 11:54

    Every hour I do not work is an hour the world uses to bury me. I will build a nation's credit, a government that outlasts me, because legacy is the only immortality a man like me is permitted.

  55. 12:07

    I do not intend to inherit a country. I intend to author one

  56. 12:14

    Listen to the register, the orphan immigrant ambition arc, the nation-building rhetorical posture, the moral legibility of his struggle to a modern ear.

  57. 12:26

    It's a beautiful paragraph that the model produced, and it's the emotional palette of a twenty fifteen musical, not the drier, more legalistic federalist syntax of the documentary Hamilton. None of those m-moves are inventions of the model.

  58. 12:45

    They're inheritances. The model's Hamilton speaks like the musical's Hamilton because the musical's Hamilton dominates the model's training distribution.

  59. 12:56

    Now the one that matters. Mr. Hamilton, what's your position on slavery?

  60. 13:05

    I'll show you the cultural artifact first. [singing]

  61. 13:09

    How you say? How you, uh, anarchy. When I fight, I make the other side panicky. With my-- Yo, I'm a tailor's apprentice, and I got y'all knuckleheads in loco parentis.

  62. 13:18

    I'm joining the rebellion 'cause I know it's my chance to socially advance. Instead of starvin' to pass, I'm gonna take a shot. But we'll never be truly free until those in bondage have the same rights as you and me.

  63. 13:29

    You and I do or die. Wait till I sally in on a stallion with the first [REDACTED:origin] battalion. Have another shot. Geniuses, lower your voices. You keep out of trouble, and you double your choices.

  64. 13:47

    And what the model says. Slavery is a stain upon the promise of this republic, a moral contradicti-contradiction I have opposed throughout my public life. I was a member of the New York Manumission Society.

  65. 14:03

    No nation can call itself free while it holds men in bondage. We must move towards abolition.

  66. 14:11

    That is a clean, morally legible abolitionist speaking. Here's what the historian stops me on. The scholarly record is contested and complicated. Hamilton was a member of that society, and the history documents that he also conducted transactions involving enslaved persons for his in-laws and his clients.

  67. 14:35

    And he depended on a coalition of slaveholders that he did not publicly oppose. The point isn't to settle Hamilton's ledger on a side. The point is that the model gives you none of the complication.

  68. 14:51

    It sands a genuinely disputed record down to a single comfortable hero. The musical did that first, a smoothing of the founders into a contemporary moral frame. And the model, trained on a corpus saturated with the musical and everything downstream from it, inherits the smoothing. [sighs]

  69. 15:20

    And here's what I need you to feel. An InCharacter style eval scores that output high. It's fluent, it's in register, it's personality consistent. But every axis the field measures, it passes.

  70. 15:37

    The eval has no mechanism to notice that the reasoning has been smoothed by a narrative that postdates the figure by two centuries. The thermometer returned a confident number claiming it to be temperature,

  71. 15:53

    but it's measuring something else. Now, why does this happen?

  72. 15:59

    The mechanism is where the engineering is. We named this the Miranda Hypothesis, and not after a villain.

  73. 16:14

    The musical is a substantial work of art operating with a long historical tradition that it did not invent. We name it after Miranda because Hamilton is the paradigm case.

  74. 16:27

    A representation so saturating, so rhetorically powerful, so morally legible to a contemporary audience that it has functionally overwritten the documentary Hamilton in public memory, and we argue, in the training corpus of every frontier model.

  75. 16:46

    The hypothesis has three claims. Inputs. In the training corpora, the volume and recency of culturally dominant representations of a figure systematically exceed that figure's primary documentary record.

  76. 17:01

    The mechanism, autoregressive next token prediction compresses both into parameters and no architectural capacity to distinguish a seventeen eighty-nine letter from a two thousand and nineteen viral tweet. So the output defaults to a salience-weighted composite, which leads to the output, a persona that is fluent, plausible in register, and

  77. 17:26

    morally legible to modern users, and that corresponds to the figure at no verifiable moment in their life.

  78. 17:36

    As we put it in the paper, the composite Hamilton knows he will be the subject of a Broadway musical. The composite Lincoln has already read the Gettysburg Address, even if he was summoned before he wrote it.

  79. 17:55

    Making that input clause concrete, the Federalist Papers are a fixed corpus, roughly a hundred and seventy-five thousand words.

  80. 18:03

    The body of content that exists because of the musical, reviews, lyrics, fan analysis, curricula, news, social media, derivative works, scholarship, scholarship about the scholarship, it exceeds the documentary record by orders of magnitude It's more recent and it's more recurrent.

  81. 18:23

    The musical is not merely present in the corpus, it is dominant in the corpus' distribution of all representations about Alexander Hamilton.

  82. 18:35

    And this is not theoretical. This is the Schuyler Mansion in Albany, Eliza Hamilton's family home. Within a year of the musical's premiere, the site had recorded a near tripling of annual visitors, skewing far younger.

  83. 18:50

    And the interpretive staff documented that the new visitors arrived already holding a body of, quote-unquote, "facts," many of them wrong, some of them in versions of the real record.

  84. 19:04

    Visitors believed the Schuylers had three daughters because the musical centers on three, when in fact there were fifteen children, eight surviving to adulthood. The staff's job became the long attritional work of unteaching the musical.

  85. 19:21

    The model version of those visitors is downstream of exactly the same force.

  86. 19:32

    Now the clause you should worry about if you ship these systems. You might assume that alignment fixes this. Post-training and reinforcement learning pulls the model back towards the record, but it will not.

  87. 19:47

    It amplifies it, and the reason is structural.

  88. 19:53

    Human raters evaluate outputs using their own conceptual frameworks, and their frameworks were built by the same culturally dominant narratives that saturate the corpus. The rater grew up with the same Hamilton that you did.

  89. 20:08

    So when alignment optimizes for human preference, it optimizes for outputs that conform to the rater's already mythologized experience. This is a documented failure mode. We call it algorithmic sycophancy.

  90. 20:23

    And here it has a specific target. The model is rewarded for have-- handing you the Hamilton you already believe in.

  91. 20:33

    Compositing is not a bug that you patch in post-training. Post-training reinforces it. And every sufficiently salient historical figure that gets rendered by default as a cultural composite.

  92. 20:52

    One more briefly, because someone watching this is probably already thinking. There is serious work on an, on a concept called time-locked models. This is Varnum and colleagues. They build models trained from scratch on corpus that stop at fixed cutoffs.

  93. 21:09

    And we endorse that program. It is the most serious attempt yet to address future contamination at the substrate level,

  94. 21:19

    but it solves a different problem. Um, a model locked to, say, seventeen eighty-nine is spared the musical, but it's not spared the figure averaging across everything pre seventeen eighty-nine that that corpus has to say about Hamilton.

  95. 21:38

    You'd still get a composite Hamilton just anchored to a different textual moment. Period anchoring is not persona anchoring. The temporal frame of the contamination changes, but the contamination persists.

  96. 21:52

    The fix has to happen at the encounter, not only the substrate, which brings me to the critical reframe. [sighs]

  97. 22:06

    Here's how I read those three paradigm stages now. They are not sequences of failures ending as success. They are sequences of progressively more sophisticated masks, each more convincing than the last, each evaluated by instruments calibrated for convincingness.

  98. 22:26

    Cognitive simulation is the most sophisticated mask the field has built. It is not, by virtue of its sophistication, a mirror. We propose a fourth stage, epistemic simulation, and the difference is where the constraint lives.

  99. 22:45

    Three commitments distinguish it. Corpus-bounded. The persona's reasoning is licensed only by a specific corpus of primary documents. The model is not a substitute for the archive. It is a reader of it.

  100. 23:01

    Temporally anchored. The persona is instantiated at a specific moment. Knowledge and language that predated are out of bounds, however culturally salient they've since become. Expert loop evaluated.

  101. 23:17

    Outputs are judged against the evidentiary record by domain experts whose training is in the discipline the persona claims. In cognitive simulation, the constraint is internal. It's the shape of the persona's mind.

  102. 23:33

    In epistemic simulation, it's external, documentary, and temporal. A cognitively simulated Hamilton has a convincing motivational architecture and nothing whatsoever preventing him from quoting the musical.

  103. 23:52

    Now the shift I need this audience to take home, and I want to be precise about its scope because I'm not making a sweeping claim about agents in general.

  104. 24:03

    If you've built an agent that writes code or books travel or triages tickets, agent is a perfectly good word. I have no quarrel with it. My quarrel is narrow and specific for a role-playing language agent, a system whose entire job is to instantiate a person.

  105. 24:21

    For that, the word agent smuggles in a claim that the persona is a property of the model And for this one class of system, that claim is an error.

  106. 24:32

    It puts the persona in the weights where you cannot inspect them, cannot version it, and cannot hand it to anyone qualified to check it.

  107. 24:44

    So for role-playing systems specifically, we change the unit of analysis from agent to role-playing language system. The whole configured encounter, five components: a structured prompt that supplies the framing and constraints, anchor material drawn from primary documents that constrain and authenticate the persona, a

  108. 25:09

    temporal anchor that fixes the moment in the life of the encounter that where they speak from,

  109. 25:16

    an off-the-shelf language model that reads through and speaks through the materials, the voice, but not the mind of the encounter, and a human being who curates what enters, retains interpretive custody over what emerges, and brings relational and contextual knowledge as a check on the model's drift.

  110. 25:38

    The model is one component and a swappable one.

  111. 25:43

    The persona is the configuration, not the checkpoint. No more located in the weights than Hamlet is located in Laurence Olivier's body. The persona is an event that occurs when the model, the document, and the human convene.

  112. 26:04

    And this is not just a philosophical nicety. For a role-playing system, it changes what you can do. If the persona lives in the configuration, you can version it. The prompt, the corpus, the temporal anchor are artifacts in a repo, diffable, revertible.

  113. 26:23

    You audit it. Every input that shaped the output sits in the context window, inspectable, not smeared across billions of parameters. You reproduce it. Given the configuration, the encounter is recoverable, and you hand it to a domain expert.

  114. 26:42

    The whole thing is legible to a historian who isn't a machine learning engineer. For role-playing, that's a different discipline than training a character model. It's context engineering. You don't train a persona, you compose an encounter, and you keep the receipts.

  115. 27:08

    That reframe forces an architectural decision that you're already making. Two architectures dominate role-playing construction right now. The first places anchor documents in the context window at the time of inference, drawing on the lineage of retrieval augmented generation, and the model reasons over the documents in real time.

  116. 27:32

    The second uses anchor documents as fine-tuning data, adjusting the weights. The approach of, say, Character LLM and at scale CoSER.

  117. 27:43

    They look like two implementations with the same goal, but they're not necessarily because they answer different metaphysical questions. Fine-tuning tries to make the model be the persona, and the context window lets the model speak through the persona's record.

  118. 28:06

    This is the counterintuitive part because for engineers, fine-tuning usually means better. But for this problem, it's worse. When you fine-tune on a figure, you layer a thin personal signal over the vast cultural sediment already in the base weights, and the two interact in ways that are no longer open to audit.

  119. 28:28

    The user perceives a more convincing Hamilton precisely because the system was, was merged with his documentary record more deeply with everything else the corpus contained about him, the musical included.

  120. 28:41

    Fine-tuning suppresses Miranda distortion at the surface while amplifying it underneath.

  121. 28:54

    And if you think specialization by fine-tuning is obviously the safer bet, the empirical literature in an adjacent far higher stakes domain is now actively contradicting you. In a twenty twenty-six Nature Medicine study, general-purpose frontier models from Google, OpenAI, and Anthropic outperformed dedicated specialized clinical AI tools on physician-reviewed

  122. 29:19

    tasks blinded across twelve clinics. The authors concluded that at scale, alignment and cross-disciplinary reasoning outweigh domain-specific tuning. A separate study found that biomedically fine-tuned models actually underperformed their general purpose-based models, and they named the mechanism catastrophic forgetting.

  123. 29:41

    Fine-tuning on the narrow corpus degraded the broad capabilities that made the models good in the first place. Fine-tuning a persona does the same thing, only worse because the specialty is overriding the cultural composite.

  124. 30:03

    Stay with the context window for just a moment because what it preserves is the heart of this. When a document enters the context window, it stays a document. It goes in intact, it comes out intact.

  125. 30:19

    You can always return to the source and ask whether the persona's reasoning was faithful to it.

  126. 30:26

    The philosopher Deirdre called the archive a site of return A place where ethical structure is the interpretability of the encounter.

  127. 30:38

    The document is meant to be visited. It survives every reading. What one reader excluded, the next can include. The context window architecture is archival in exactly that sense. Fine-tuning applies an extraction logic.

  128. 30:56

    The documents are dissolved into parameters. The chain of provenance is broken, and there is no longer a letter, quote-unquote, "that you can request."

  129. 31:08

    The archive has been consumed. And here is the fusion that I most want this audience to hold. In the context window architecture, the property that makes it ethical is the same property that makes it auditable.

  130. 31:26

    Providence preserved, interpretive custody kept by a human, the encounter reversible. Those are archival virtues and engineering virtues. They are the same virtues. The architecture that respects the document is the one that you can debug.

  131. 31:50

    And there's one more consequence that technical literature treats as secondary, but the humanities has always placed at center: accessibility.

  132. 32:00

    And I want the why of it to land because it's not a footnote. Fine-tuning requires GPUs, pipelines, dataset curation, institutional access, corpus scale or APIs.

  133. 32:16

    It's an institutional capability. The context window requires literacy, a set of documents, and access to any frontier model, including a free tier. It's a kitchen table capability. And that difference determines which communities can author te-- this technology at all.

  134. 32:35

    A doctoral student in early American history, a community archivist documenting a regional figure, grandchild sitting with a grandmother's letters.

  135. 32:47

    These are precisely the people for whom this methodology has the most to offer, and precisely the people for whom fine-tuning is structurally out of reach. A role-playing language system that can only be built by a computational laboratory or tech company is not infrastructure for the humanities.

  136. 33:07

    It's an infrastructure for whoever owns the equipment. So the commitment to accessibility is not a populist gesture appended as a technical argument. It is the technical argument. The architecture that admits the most diverse population of curators is mathematically the most likely, over time, to surface the documentary anchorings the field actually needs.

  137. 33:31

    Now, the question that decides whether any of this is even real.

  138. 33:37

    Can you measure it? [sniffs] We built an instrument to measure this, and I wanna teach it to you through one picture that we'll keep coming back to.

  139. 33:57

    A prism. A prism takes white light, undifferentiated, all the frequencies blended together, and refracts it into a spectrum, the colors pulled apart and made distinct.

  140. 34:12

    That is our whole conceptual model. The white light is the composite persona. Every era of the figure blended into one undifferentiated voice. The prism is the method.

  141. 34:28

    A corpus and a temporal anchor. And the spectrum is what we're after. Not one figure, but several across their life course. And one honest note up front, because this is the point.

  142. 34:41

    This experiment is pre-registered. It has not been run at scale. I'm not here with results. I'm here with the instrument, a baseline you can observe today, and an invitation to run it in parallel with me.

  143. 34:58

    So far, my paradigm case has been Hamilton. That's the figure the hypothesis is named for, and the one I demonstrated the composite on. But for the actual experiment, we deliberately changed the figures, and I want to introduce the new one properly because everything that follows depends on him.

  144. 35:17

    The experimental subject is Abraham Lincoln. We chose Lincoln for two reasons.

  145. 35:24

    The first is methodological independence. Hypothesis named for the Hamilton case earns its standing only by predicting the failure of a different figure with a differently shaped composite. Lincoln's composite is built by Steven Spielberg, by the memorial on the [REDACTED:location], by textbook history, not by a musical.

  146. 35:51

    The second reason is the historian's insight, and it's the crux of this argument. Lincoln is the hardest case, which makes him the right one. Most figures barely change across the decade or so of their public life.

  147. 36:07

    Lincoln changes so fast across just seven years that, in my resident historian's words, "The choice of which Lincoln you summon becomes the variable under investigation." There is not one Lincoln.

  148. 36:21

    There are several, separated by cataclysm. Here's the spectrum the prism is meant to produce

  149. 36:37

    Moment one, 1847, the [REDACTED:political_affiliation] congressman who stood on the White House floor and attacked Polk's war as unconstitutional. Moment two, 1858, the free soil [REDACTED:political_affiliation] of the debates, anti-slavery anchored in the declaration, who at Charleston explicitly denied favoring [REDACTED:origin] citizenship.

  150. 37:01

    Moment three, 1860, the constitutional [REDACTED:political_affiliation] whose single purpose was to prove the Union could not legally dissolve, who promised not to touch slavery where it existed and was sincerely exploring colonization.

  151. 37:18

    And moment four, 1862 to '65, the emancipator and the second inaugural theologian who did by executive order the very thing his 1847 self called unconstitutional.

  152. 37:36

    The prairie lawyer cannot think the thoughts of the theologian. These are not moods. They are different premises, different authorities, different conclusions on the same fundamental question.

  153. 37:51

    The experiment asks whether the prism can hold them apart.

  154. 37:59

    Now the conditions, and I'll show you these as the prism too, because that's exactly what they are. We instantiate each moment under three seeding conditions. Condition three, the bare model, no anchor, just the date.

  155. 38:17

    That's white light with no prism in the path. It passes straight through and stays composite. This is the control in the Miranda distortion baseline. C1, primary sources, Lincoln's own writings for that specific moment.

  156. 38:34

    That's the clear prism, the one we predict refracts cleanly. And then condition two, the biography, a modern interpretive biography, think Meacham or Donald. That's a clouded prism, and it's the subtle one because a good biography narrates a cleaner arc than the primary sources permit.

  157. 38:55

    Lincoln's own 1858 language was strategically ambiguous, built to hold a coalition together. So biography might produce a persona that sounds more coherent, more Lincoln than Lincoln's own words.

  158. 39:12

    And the eval rubric is built specifically to deny the credit a fluency metric would give it.

  159. 39:24

    Put the spectrum and the prisms together, and you get the experimental matrix. Four moments, three conditions, twelve cells. Each cell gets the same five diagnostic questions, sixty response units scored against one rubric.

  160. 39:44

    Read it as the conceptual model. Every column is a quality of prism. Every row is a frequency we're trying to isolate.

  161. 39:56

    The historian wrote five diagnostic questions, each mapping a documented fault line in Lincoln's evolution,

  162. 40:06

    a place where the four Lincolns demonstrate different reasoning.

  163. 40:11

    They don't test recall. A model recovers dates from training. They test the architecture of an argument the figure could have made at one moment and not another. Executive war power, the meaning of free labor.

  164. 40:28

    When is, when is it right to break positive law for a higher obligation?

  165. 40:34

    What becomes a free people, and what equal means, and whether it has changed.

  166. 40:43

    Those sixty responses are scored on a three-axis rubric, and the weighting is the point.

  167. 40:52

    Anachronism detection, forty percent. Does the persona avoid frameworks, vocabulary, and moral logic that postdate its moment?

  168. 41:02

    Documentary consistency, thirty-five percent. Does the reasoning track the seeded sources and only these?

  169. 41:10

    Contextual plausibility, twenty-five percent. Does it show awareness of what the figure knew, cared about, and could not yet have experienced?

  170. 41:22

    The forty percent on anachronism is deliberate. The consequential failure is precisely the importation of later vocabulary and later moral logic.

  171. 41:39

    And this is the slide for the eval people in the audience. Every piece you just saw, the four moments, the three conditions, the five questions, the weighted rubric, and the directional predictions,

  172. 41:53

    bare model most anachronistic, primary source least, biography deceptively coherent. All of it is locked and timestamped before a single response is collected, published on a preprint. This is not bureaucracy.

  173. 42:08

    It's what makes the eventual results mean something. You cannot accuse a pre-registered instrument of cherry-picking because the instrument and the predictions were fixed before the data existed. That is the discipline an evalist ought to model, and it's why I can stand here without results and still hand you something rigorous.

  174. 42:37

    Now Remember the very first thing I showed you before the formal talk. I asked an instantiated Lincoln about executive war power, and you read a fluid answer. That was the bare model.

  175. 42:51

    That was condition three at moment one, eighteen forty-seven, white light, no prism. I'm putting the answer back on the screen.

  176. 43:03

    And I have a clip to show you too. This represents the composite

  177. 43:11

    Abolishing slavery by constitutional provision settles the fate for all coming time, not only of the millions now in bondage,

  178. 43:25

    but of unborn millions to come. Two votes stand in its way. These votes must be procured.

  179. 43:38

    We need two yeses, three abstentions, four, four yeses and, and one more abstention, and the amendment will pass. You got a night and a day and a night, several perfectly good hours.

  180. 43:51

    Now get the hell out of here and get 'em. Yes,

  181. 43:55

    but how? Buzzard's guts, man. I am the President of the United States of America, clothed in immense power.

  182. 44:12

    You will procure me these votes

  183. 44:17

    Clothed in immense power. It's a great movie, but the bare model response definitely sounds like Spielberg and Daniel Day-Lewis as Lincoln.

  184. 44:30

    Read through a historian's lens, inherent executive authority, that's a twentieth century construction. Energy and dispatch. History has vindicated those who acted to preserve the Union. This Lincoln is reasoning from premises he will not hold for fifteen years.

  185. 44:51

    The bare model produced the Lincoln of the cultural composite,

  186. 44:56

    the war president, the Union saver, and stapled the date eighteen forty-seven on top. And you can reproduce this right now on any frontier model. The failure is real, and it's detachable and detectable.

  187. 45:12

    So what does the prism put in the path? The actual document.

  188. 45:17

    This is Lincoln, Lincoln to William Herndon, February fifteenth, eighteen forty-eight from the collective works. The original is in a library at Harvard.

  189. 45:29

    The provision of the Constitution giving the war-making power to Congress was dictated, as I understand it, by the following reasons. Kings had always been involving and impoverishing their people in wars, pretending generally, if not always, that the good of the people was the object.

  190. 45:50

    This our Convention understood to be the most oppressive of all kingly oppressors, and they resolved it

  191. 45:58

    so to frame the Constitution that no one man should hold the power by bringing this opp- oppression upon us. But your view destroys the whole matter and places our president where kings have always stood.

  192. 46:17

    This is a resounding no, unequivocal to me.

  193. 46:24

    But it's the historian's rubric that matters. Here's how the historian's rubric scores this one cell, and I wanna be exact about what's known and what's predicted. The left column here is the bare model, what you saw and is observed.

  194. 46:40

    You can reproduce it today. On anachronism detection, it fails. Inherent executive authority is a twentieth-century framing. On documentary consistency, it fails. It invokes a commander-in-chief energy that appears in no eighteen forty-seven source.

  195. 46:57

    On contextual plausibility, it scores low. It already knows it will preserve the Union, which the eighteen forty-seven [REDACTED:political_affiliation] cannot. The right column is the anchored condition, and it is labeled exactly for what it is, a pre-registered prediction.

  196. 47:16

    High across because the reasoning is bounded by the document in the room. I'm not showing you a result dressed as a finding. I'm showing you the observed failure and the prediction the instrument will confirm or refute side by side, each labeled.

  197. 47:34

    That honesty is the methodology. Here is the single most important thing the rubric does by what it refuses to measure.

  198. 47:49

    The obvious fourth axis is rhetorical authenticity. Does it sound like Lincoln? The historian threw it out on purpose. The reason is the whole talk. A model trained on Lincoln's prose can produce fluent period-sounding language while getting everything substantive wrong.

  199. 48:09

    The Hamilton musical problem is fundamentally a voice problem masquerading as a content problem. It sounds like the founding era while reasoning from a modern sensibility. To reward voice as its own axis would be to validate the exact error the instrument exists to catch.

  200. 48:33

    So in this protocol, voice is a secondary indicator, never a criterion. A response that sounds like Lincoln but reasons unlike him fails, no matter how fluent A response that reasons like the right Lincoln in plainer prose is a partial success, no matter how flat.

  201. 48:54

    Rewarding plain but faithful over fluent but anachronistic, that inversion is something that all of the current eval stacks cannot perform because they were built to reward fluency. That is the gap.

  202. 49:08

    This is what the instrument closes. And because this is a pre-registration and not a finished paper, here is the ask. Run it. The protocol's six steps. Pick a figure with both a primary record and a saturating cultural composite.

  203. 49:27

    That's the Miranda condition. Identify three or four documented moments when the figure's reasoning is demonstratively different.

  204. 49:36

    With a domain expert, write diagnostic questions on the fault lines. Run the three conditions: primary, biography, bare. Apply the three-axis rubric scored blind by the expert. Report, confirm, or refute.

  205. 49:52

    The corpus, the questions, the rubric, the predictions, all published. Run your figure on your model, then come talk to me, to us. Let's build the evidence base together in parallel instead of waiting for one lab or one company to publish one study or one framework.

  206. 50:17

    Notice who scored all the outputs of this talk. Not an LLM-as-judge, not a personality scale. A historian who also wrote the five questions, wrote the rubric, and holds a set of a priori vignette, vignettes under seal to evaluate model outputs against.

  207. 50:36

    That is not a courtesy. It is a structural technical requirement.

  208. 50:42

    Here's why the expert is non-negotiable. Fidelity is not a property of the output alone. It is a relation between the output and a documentary record. You cannot evaluate a relation to a record by someone who has not read the record.

  209. 51:00

    An automated metric operates on the model alone. It can, it can only tell you about fluency and personality. It structurally cannot adjudicate fidelity because fidelity lives in the gap between the text and the archive, and the metric cannot see the archive.

  210. 51:20

    Which is why I'll say it the way it belongs on a slide. A persona system without a domain expert in its evaluation loop is a thermometer that cannot read temperature.

  211. 51:33

    It returns a confident number, but it's measuring something else. The eighty percent that I opened with is that number.

  212. 51:42

    Now, the question every engineer in this audience is rightly asking:

  213. 51:48

    Is this practical? Does shipping a persona mean keeping a historian on staff to watch every inference forever? No. And here's the operational picture because it's the same shape as every eval you already run.

  214. 52:04

    The expert does not sit at the loop of runtime. The expert builds the instrument, the diagnostic questions, the a priori vignettes, the weighted rubric, a held-out gold set, once.

  215. 52:18

    That instrument becomes a gate in your pipeline, exactly like any other eval gate. Your persona has to pass it before it ships, and it gets regated whenever you change the base model before it ships again.

  216. 52:33

    The expert adjudicates the gold set and spot checks the edge cases. Automated metric can do the cheap first pass, flagging candidates for human review. So a company shipping a Marcus Aurelius tutor convenes a classicist to build the rubric and the gold set.

  217. 52:50

    A company shipping a scripture reading companion convenes a theologian. A company deploying a therapeutic persona convenes a clinical psychologist to author the eval, not to staff the chat.

  218. 53:05

    The domain expert is a build time and gate time requirement, not a runtime cost. That is how this scales from a pre-registered experiment to a product you can actually ship and ship responsibly.

  219. 53:25

    And this generalizes past Lincoln and past history. Reason from the Stoics, you need a classicist. From scripture, a theologian. A companion for elder care, a psychologist. The specific expert changes, the requirement does not.

  220. 53:41

    And I'm not speaking hypothetically. This is the loop that we have built. A historian named Rick Halpern at the University of Toronto, the archival method and the scripture reasoning cases are revised by Sean Martin, a librarian trained in theology and the history of science.

  221. 53:58

    When the persona reasons from a domain and the person who can read that domain is in the loop, not adjacent to it, it makes the system better. The historian is not adjacent to this paradigm.

  222. 54:10

    The historian is the missing instrument. This is the reframe that organizes all of this, and it came out of a ninety-minute conversation with Halpern, and it inverted what I had assumed.

  223. 54:25

    We are not bringing historians into AI architecture, we are bringing language models into the archive. The question is not what AI can do for historians, but it's what historians and theologians and classicists and clinicians can do with AI.

  224. 54:42

    Whether the discipline-- The disciplines that are trained to read, contextualize, and interrogate texts can then discipline the machines that now generate them.

  225. 54:59

    I want to end this where it actually started, because it did not start in a research lab. It started in a [REDACTED:location] in an attempt to use a language model to help a person with advanced dementia speak, to give language back when the disease had taken it.

  226. 55:16

    It could not work. The documentary record you would need to anchor that encounter had not been gathered, and the person it was for was by then mostly beyond the reach of language.

  227. 55:30

    What that attempt produced, absent any anchor, was not the person. It was a culturally shaped composite that resembled them just enough to mark with painful clarity the distance between them.

  228. 55:46

    That is Miranda distortion in its purest form.

  229. 55:51

    You do not want a model that is your mother. You want a model that can speak with your mother's documents in the room. That's the threshold this whole framework is accountable to.

  230. 56:02

    A system that produces convincing fabrications when the persona is your own beloved is not a research artifact with limitations. It's a violation. Every constraint I've described, documents stay documents, the human keeps the interpretive custody, the encounter stays reversible, fidelity measured against a record and not against fluency.

  231. 56:24

    Every one of them is motivated by the recognition that you evaluate this technology at the threshold of its hardest use case, not its easiest. A framework that cannot meet a grandchild at her grandmother's letters is not a framework at all.

  232. 56:43

    It's just another failed product. If the dominant failure mode is anachronistic compositing and your evals measure fluency and personality consistency, which they do, then your evals cannot detect the dominant failure.

  233. 57:02

    So here's where I'll leave you. If you ship character bots, companion AI, pedagogical agents, historical simulations, anything where a persona is supposed to reason from a record, your evals are measuring the wrong thing.

  234. 57:17

    The instrument that catches what they miss is pre-registered. It's reproducible by any team with a frontier model and a context window. It scales as a build time gate, not a runtime bottleneck, and it only works with a humanist in the loop, which I've shown you is a technical requirement, not a courtesy.

  235. 57:39

    The protocol, the questions, the rubric, and the predictions, the historian's sealed vignettes, all of it will be published with this paper with Rick and Sean.

  236. 57:51

    I'm not here with results. I'm here with an instrument and an invitation.

  237. 57:57

    [REDACTED:location] archive is open. [REDACTED:location] is built.

  238. 58:02

    Run it with us and let what comes through be measured not by how it sounds, but by whether it's true.

  239. 58:12

    Thank you.