AI Engineer World's Fair 2026
Build for the Memo, Not the Demo — Notes from 200 Investment Committees
Read the talk
Build for the Memo, Not the Demo
Financial AI earns trust when a skeptical reader can trace claims, reconcile numbers, challenge assumptions and identify the person accountable for the decision.
From a talk by Shawn Chan
Would you trust this paper with $100 million?
Before his company commits $100 million, Shawn Chan has to answer a deceptively simple question: does he believe the paper in front of him? For fifteen years, a confident person handed him that recommendation. Now, sometimes, a chatbot supplies it. The grammar is better, the formatting is cleaner, and follow-up questions never provoke defensiveness. None of those qualities establishes that the recommendation is right.
A financial AI product must survive scrutiny, not merely create a good first impression. A demonstration can impress a room for five minutes; an investment committee is a room whose job is to resist being impressed. The same evidence discipline that makes a product credible also helps its founders raise money. That distinction becomes urgent when a CEO sees a compelling demo, a major customer brings in compliance, or the next funding round is three months away.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Trust in the number
Chan’s perspective comes from cross-border mergers, IPOs and strategic investments spanning Hong Kong, Mainland China, the UK and the US. He reports attending about 200 investment committee meetings and reading hundreds of founder and banker decks—including 47-slide seed decks. The recurring lesson is that a number does not get a deal approved on its own. Trust in that number does.
Early in his career, Chan watched a polished, expensive banker present beautiful slides containing confidently wrong numbers. A senior person asked, “Where does this number come from?” The banker stalled. The presentation had no usable answer to the question that mattered most. Chan approaches AI from that side of the table: as the buyer and evaluator deciding whether to trust a system and the founder selling it, rather than as its builder.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Two different machines
A demo and a memo operate under different conditions:
| Demo | Investment memo | |
|---|---|---|
| Input | One clean document | Filings, transcripts, broker notes, spreadsheets, call notes |
| Evidence | A convenient example | Hundreds of pages, often in conflict |
| Output | A fluent answer | A defensible recommendation |
| Success | Impress the audience | Survive an argument |
The memo is the document a real committee reads before real money moves. Its inputs include rushed notes from last Tuesday’s call alongside formal filings. Disagreement is part of the working material, not an exceptional condition that can be removed from the demonstration.
The everyday comparison is a phone summarizing a long email versus that same phone standing before a bank to defend your mortgage application. A plausible summary may satisfy the first task. The second requires the system to be right and to prove why its answer deserves belief.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
When the demo becomes a financial event
In Chan’s February 2023 example, a large technology company promoted an AI assistant with an incorrect answer about a space telescope. Chan reports an approximately 8% one-day share-price decline and roughly $100 billion in lost market value. He connects the selloff to the unchecked sentence; contemporaneous reporting on Alphabet’s decline also identifies competitive and advertising concerns, so the sentence should not be treated as the sole established cause.
The missing discipline is familiar to a junior analyst: establish where the claim came from and whether anyone checked it before publication. Once money is watching, even a promotional sentence can face the memo’s standard of scrutiny.
Inside a product, that scrutiny becomes six questions: which source deserves belief, whether the numbers agree, whether contradictions remain visible, whether a statement is a fact or a guess, how quickly it can be substantiated, and whose name is attached to the decision.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Numbers must agree with the document and reality
Chan’s failed-memo example gives revenue growth as 18% on page one and 17.4% in a table on page eleven. The difference is 0.6 percentage points, but its significance is larger than its size. If nobody checked the easy arithmetic, a committee has reason to doubt whether anyone checked the harder assumptions. In his account, the memo failed because of what the discrepancy revealed about the process.
The larger-scale example is an American real-estate business that used an algorithm to buy houses. Chan describes around half a billion dollars written off, closure of the unit and a quarter of staff let go. Those rounded figures require care: the matching Zillow Offers episode involved forecasts as well as realized charges; Zillow’s subsequent shareholder letter reports $405 million in combined third- and fourth-quarter inventory write-downs, a category distinct from total business losses.
Chan characterizes the failure as insufficient supervision rather than a lack of model intelligence. His engineering requirement is continuing reconciliation with reality after launch. Internal agreement between two pages is one control; checking whether the system’s numerical expectations remain aligned with actual outcomes is another. Fluency and confidence provide neither.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Put the disagreement in front of a human
A CEO’s growth figure on an earnings call can disagree with the official filing. That gap is valuable diligence material: it creates a question worth investigating. A system optimized to produce smooth prose may instead select whichever version reads better, leaving the reader unaware that a conflict existed.
Chan recounts a discrepancy caught only because someone happened to have both documents open. “Luck is not a control.” The builder’s responsibility is to preserve the disagreement and put it before a human. Automatically choosing a preferred answer is not equivalent to resolving the underlying issue; it can simply erase the evidence that a review is needed.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Keep estimates from becoming facts
Consider the sentence: “The company will likely receive approval next quarter.” It describes an estimate, not a completed event. Committees need to evaluate such judgments separately from the facts supporting them. When a system blends both into seamless prose, readers lose the boundary between what is established and what they are being asked to believe. A $100 million decision can then become an assessment of the document’s overall tone.
Chan watched an expectation of approval soon become a statement that approval had been received across three memo drafts. Nobody needed to deliberately falsify the document: successive rewrites removed the uncertainty until the estimate looked like a fact. The approval did not arrive on schedule.
The control is a durable uncertainty label. A tag or color must remain recognizable when the content is copied into someone else’s slides three weeks later. For the approval example, the pending status has to remain attached to the claim through editing and reuse; better wording alone cannot turn an expected approval into a received one.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The click-through is the product
A correct claim still has limited practical value if nobody can find its source. Chan illustrates the danger with the New York legal brief containing six nonexistent court cases: convincing citations gave invented authorities the appearance of verifiability. The incident corresponds to Mata v. Avianca.
The lawyer also asked the chatbot whether the cases were real and received reassurance. Chan places that exchange before filing; the court’s opinion records conflicting accounts of its timing, with the later declaration placing it after the first show-cause order. The verification failure is the same: asking the generator to endorse its own output does not independently establish that the cited authority exists. The court imposed sanctions on the attorneys and their firm. Realistic names, page numbers and formatting had made the unsupported material easier to believe.
The verification path is part of the product. Chan proposes a 30-second test: when someone challenges a sentence, one click should reach the exact source paragraph. The alternative is opening seven browser tabs and scrolling while a room waits. He has been that person in a real meeting. A citation list that looks complete but leaves the reviewer searching has failed the practical test.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Someone has to sign
An airline chatbot told a grieving customer to buy a full-price ticket and claim a bereavement discount afterward. The error concerned retroactive eligibility, not whether bereavement fares existed. The customer took the airline to a tribunal. The corresponding Air Canada case shows how confidently delivered advice can become an organizational liability.
Chan recounts the airline’s rejected position as treating the chatbot as a separate legal entity responsible for its own actions. The tribunal awarded damages; responsibility did not disappear because the information came through software.
For a financial decision, the architectural consequence is a findable human signer. The system should support that person’s review and responsibility, not make it harder to determine who owns the result. Build around the accountable person.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Five requirements before a live deal
These failure stories become concrete acceptance requirements. Before allowing a vendor’s system near a live deal, Chan would demand:
- Receipts attached to claims. Each sentence links directly to its source paragraph and carries the source’s trust level. A citation tab at the end is insufficient.
- Visible facts and estimates. A reader can immediately distinguish established information from someone’s best judgment.
- Automatic numerical agreement. The system refuses to release a memo with inconsistent figures, rather than depending on a person catching them at two in the morning.
- Explicit contradictions. When sources disagree, the system raises the issue instead of silently selecting the friendlier answer.
- Logged human approval. The record identifies who reviewed the memo, what changed and when the reviewer signed. That record is the audit trail.
None of these requirements asks for a smarter model. Chan treats them as problems of evidence handling, controls and honesty. The competitive test is whether a tired, skeptical finance professional can trust the output at eleven at night without opening seven tabs—not whether the underlying model gained a few benchmark points.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Your pitch becomes somebody else’s memo
After the fundraising pitch ends, the deck becomes an internal investment memo. Someone on the other side of the table writes down the case for investing and checks the numbers the founder said aloud against the data room. The evidence discipline expected of a financial AI product therefore applies to the company selling it. Your presentation has to remain credible when you are no longer in the room to explain it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Further reading
- Zillow's Q4 2021 shareholder letterDocumentation
Company results reporting actual inventory write-downs against the earlier wind-down forecasts.
AP documents the market decline and the competitive concerns surrounding Google's Bard presentation.
The court's findings explain the invented authorities, verification failures and sanctions in Mata v. Avianca.
ABA commentary explains the incorrect bereavement-fare advice and the tribunal's negligent-misrepresentation decision.
Read the complete timestamped transcript
- 0:00
[upbeat music] Good afternoon. Thank you for being here.
- 0:15
Day four of our conference, a room with no windows, right just after lunch. You are the strongest people in this building. [laughing] Let me start with a confession. For fifteen years, my job has been one thing.
- 0:30
I sit in a room, a very smart, confident person hands me a piece of paper, and before my company spends a hundred million dollars, I have to decide, do I believe this paper?
- 0:44
This year, my job still exactly that. Except now, some of time, the very confident person handing me the paper is a chatbot. And honestly, the chatbot is often better written than humans.
- 1:03
Better grammar, nicer formatting. Never get defensive when you ask a follow-up question. Very, very sure of itself. The problem is, being sure of yourself and being right are two different skills.
- 1:19
Some of most confident people I have ever met in finance were also the most wrong. AI just learned that trick faster than rest of us. So here my whole wor- whole talk in two sentence.
- 1:36
First, almost every AI finance product is built to impress people for five minutes, and almost none are built to survive a room whose entire job is to not be impressed.
- 1:52
Second, and this is the part I promised the organizers I'd add, the exact same skills that fix your product is the skills that gets investors like me to write you a check.
- 2:09
Same muscle. I've proved both. And I'm telling this now, this week, because a lot of you are about to get pulled into exactly this. Your CEO saw a demo somewhere.
- 2:24
Your biggest cus- uh, customer suddenly has a compliance department, or you are three months from raising your next round. When any of those days arrive, I'd rather you to h- hear the hard part from someone who sit on the other side of the table, than discover them live in a room in front of the
- 2:49
people who bill by hour. A quick, a quick bit ab- about me. I promise this is the boring part, and I will keep it fast. Fifteen years of cross-border deals, Hong Kong, Mainland China, the UK, the US.
- 3:10
Mergers, IPOs, big strategic investment. On names you'd actually recognize, the kind of companies that go public and your LinkedIn feed won't shut up and about it for a week.
- 3:26
Along the way, I've sat in about two hundred investment committees, meetings. I've the [REDACTED:physical_attribute] to prove it. I checked. It's not genetics.
- 3:40
And here, the part that matters for second half of this talk, I've also read hundreds and hundreds of pitch decks, founders decks, bankers decks.
- 3:56
Forty-seven slide seed decks. I have seen fonts that should be illegal. [sighs]
- 4:05
Two hundred committee meetings taught me one thing.
- 4:10
No textbook says out loud, the number on the page is not what gets a deal approved. Trust is a number what gets a deal approved. Money doesn't f- flow intelligence.
- 4:24
Money fol- follows trust, and trust is fragile, especially when the thing that wrote the page has never once in its entire life said the words, "I'm not sure."
- 4:39
Let me give you one small taste of what those rooms feel like. Early in my career, a very polished, very expensive banker present a beautiful slides full of confident wrong numbers.
- 4:56
One senior person in room asked one qu- quiet question, "Where does this number come from?" The banker paused for what felt like an entire physical quarter. The pause taught me more about finance than three years of exams did.
- 5:14
So I'm not sure as a-- So I'm not here as a builder. I don't build th- these systems. I sit across the table from them and from the founders selling them, deciding whether to trust them.
- 5:32
Today, I'm going to give you both halves of that, what breaks trust in your product and what builds trust in your pitch.
- 5:46
Let's define two machines that most everyone keep confusing. Machine one is a demo. One clean document in, one fluent answer out. Its the whole job is to make a room go, "Oh," for five minutes.
- 6:03
Machine two is a dem-- is a memo. The real document is a real committee reads before real money moves. Hundreds of pages, filings, transcripts, broker notes, spreadsheets, someone's rushed notes from a call last Tuesday.
- 6:22
Half the sources disagree with each other. And the memo's job is not to make your go, "Oh." Its job's to survive an argument. Basically, a family dinner, except somebody's uncle brought a spreadsheet.
- 6:39
Here's the everyday version of, uh, the gap. Ask your phone to summarize a long email. That's a demo. It's not just sound pl-- uh, plausible for ten seconds. Now imagine that some same phone has to stand in front of a bank and defend out loud why you deserve a mortgage.
- 7:04
Suddenly, plausible isn't enough. Now it has to be right, and it has to prove it. That second situation is what a memo actually is.
- 7:19
Now, you might think the demo world and the memo world never touch. Let me tell you about the most ex-expensive typo in history. February twenty twenty-three, one of the biggest tech company on the earth launches its shiny new AI assistant with a promotional demo.
- 7:39
In that demo, the assistant answers a simple question about a space telescope, and it gets a wrong, uh, the wrong answer. One sentence, one wrong fact about a telescope in a marketing demo, the market noticed.
- 7:58
The company's stock dropped around eight percent in a day. That's roughly one hundred billion dollars of value gone because of unchecked sentence. One hundred billion dollars for one sentence.
- 8:13
Nobody in that company asked the one question. Every junior analyst on my team is trained to ask before anything leaves the building, "Wait,
- 8:28
where does this claim come from? Did everyone check it?" So here's the punchline. Even the demo fails the memo test. The moment
- 8:40
real money is watching, and the real money is always watching, every sentence become a memo sentence. There is no safe demo anymore.
- 8:51
Keep that story in your head because of the same trust-breaking moment happens in six smaller, quieter, very predictable ways inside of your product every day. Let's go through them fast.
- 9:10
Six ways trust quietly breaks. Which source do you believe? Do the numbers agree? Do you hide contradictions or show them? Is that a, a fact or a guess?
- 9:27
Can you prove it in thirty seconds? And whose name is actually on the decision? Six. Keep count with me. I will keep each one short, and I will bring receipts.
- 9:45
Model one, not every source deserves the same trust, but most AI system treat them like they do. Think of like this. A number from an audit filing is your accountant speaking under oath.
- 10:04
A number from an analyst note is a friend at a party confident, probably from someone's internal email is a thing you overheard in a, in a elevator. Most retrieval systems can't tell them-- this apart.
- 10:25
They grab whichever text is close to your question and hand it over like a gospel. True story, anonymized. I know-- I once watched a, um, very expensive AI tool confidently quote a number from a group chat, someone's rough guess,
- 10:46
te-text six months earlier. The model loved the confident phrasing. The real audit number was the three rows away in the actual filings. The AI just liked the group chat version better.
- 11:02
It sounded more enthusiastic. If your system can't tell an accountant under oath from a rumor in a group chat, it, it is not ready for real money.
- 11:22
Model two, the numbers have to agree. Each other,
- 11:28
everywhere, every time. Here's a memo that already died. Page one says revenue growth,
- 11:38
uh, eighteen percent, eighteen percent. Page eleven, in a little table nobody reads for fun, says seventeen point four. Nobody in the room cares about the missing zero point six.
- 11:55
They care about what it, what it means. If this person didn't check the easy mathematics, what did they-
- 12:04
Now check on the hard stuff. That memo didn't pass,
- 12:10
not because of the number, because of what number you implied. And if you want the industrial strength version of this failure, remember the giant American real estate company that let algorithm to buy houses at scale.
- 12:31
The algo- algorithm was extremely confident about house prices. The house disagreed.
- 12:41
The company ended up writing off around half a billion dollars, shut the whole unit down, and let a quarter of the staff go. The model wasn't stupid. The model was unsupervised.
- 12:57
Nobody built the boring machi- me- um, machinery that forces the numbers to keep agreeing with reality after lunch day.
- 13:09
Fluent and confident, remember, is not the same as right.
- 13:19
Model three, surprises people. A contradiction is not a bug. A contradiction is a gift. If the CEO says one growth number on the earning call and the official filing says the different, a different one, the gap is the single most interesting thing in the real story heist.
- 13:45
Real diligence lives for that gap. AI does the op- opposite. It is trained to sound smooth and helpful. So when it hits a conflict, it quietly picks whichever version reads nicer and moves on.
- 14:07
You never even learn there was a disagreement. I've sit through exactly this. CEO's number and the filing numbers
- 14:21
meaningfully different. Neither one, uh, nobody, nobody flag it. Everyone just used neither one.
- 14:31
We caught it because one person happened to have both documents open at once. Pure luck. Luck is not a control. Your job as a builder isn't resolve the argument.
- 14:47
It's to make sure that the argument happens in front of a human instead of quietly alone inside a box.
- 14:59
Model four, facts and guesses have to live in separate box.
- 15:08
And the fluent AI loves melting them into one smooth sentence. Example, the company will likely receive approval next quarter.
- 15:19
Reads like a fact, sounds like a fact. It is a guess. Somebody buys the estimate wearing fact-shaped clothing. A committee's entire job is to agree with the guesses while trusting the facts.
- 15:36
If your system melt them together, the committee can't find them sim- sims, and then all, all they can do is to, is, is approve or reject the general vibe of the document.
- 15:52
You should not spend a hundred million dollars on vibes. I watched approval except soon.
- 16:01
Turn across three drafts of a demo into approval received. Nobody lied. The guest just wore his fact costume a little longer.
- 16:14
Each rewrite until nobody remembered it stand-- it started as a guess. The approval did- didn't arri- arrive on schedule. That was an uncomfortable phone call. This fix almost embarrassingly cheap.
- 16:31
Label your guesses, a tag, a color, anything that survives being copy/paste into someone else slide three weeks later.
- 16:47
Model five, if nobody can find where a claim come from,
- 16:53
it doesn't matter how right it is. You've all heard about the New York's lawyer. Uh, he filed a l- legal brief written with a chatbot's help.
- 17:09
The brief cited six court cases, beautiful citations, prope- proper formatting, very convincing. One small issue, the cases didn't exist. The AI invested-- invented all six.
- 17:27
And here, my favorite detail, the part that should be taught in school. Before filing, the lawyer got suspicious, so he asked the chatbot, "Are these cases real?" And the chatbot said, "Yes."
- 17:44
That is like asking the guy who sold you the watch whether the watch is real. The judge fined him. The story went around the world. And my second favorite detail, the fake cases even had realistic, realistic surrounding names and page numbers.
- 18:07
The AI didn't just lie, it just, uh-The formatting. It cited itself beautifully. Wrong, but beautifully. The license is a 30-second test. When someone points at a sentence and says, "Show me where this come from,"
- 18:27
you either click once or land on exact source paragraph, or you open seven browser tabs and start s- swiping. I have personally been the guy with seven tabs in a real meeting while the room full of people watched my scroll.
- 18:49
10 out of 10 wouldn't re- would not recommend. If you remember only one sentence from this wh- whole talk, the click-through is the product. Everything else is, well, written packaging.
- 19:11
Model six, my favorite because it's the most human. Someone has to sign.
- 19:21
Here's the story you probably know. An airline website chatbot told a grieving customer he could book a full-price ticket now and claim a bereavement discount afterwards.
- 19:36
That policy didn't exist. The chatbot made it up, uh, pu- uh, politely, fluently, confidently. The customer took the airlines to a tribunal,
- 19:52
and the airline's defense, this is real, was, was that the chatbot is, quote, "A separate legal entity responsible for its own action." That is the corporate version of my dog ate my homework.
- 20:12
The tribunal didn't buy it. The airline paid, and every one of us in a boardroom quietly took a note that day. You can't outsource accountability to your own software.
- 20:27
At the bottom of every real decision, a human signs. If your architecture doesn't have a fundable human at the end of it, you ha- you have not built a product, you have, have built an excuse generator.
- 20:47
So build your AI around that accountable person, not instead of them.
- 20:56
Okay, the fix. Five things. Each one is a direct cure for a story you have just heard. This is almost word for word
- 21:10
what I'd demand from a vendor before letting their system near a live deal. One,
- 21:19
every claim comes with a receipt. Each sentence linked s- straight to its source paragraph with the source trust level attached, not a citation tab at the end.
- 21:34
Two, fact and the guesses stay visibility separate. At a glance at the page, I insistently see what's proven and what's somebody best estimate.
- 21:50
Three, numbers agree with each other automatically. The system refuse to ship a memo where the figures don't match. No human checking at 2:00 in the morning. Four, contradictions get surfaced, never smoothed over.
- 22:09
When sources disagree, the system raises its hand instead of picking the fr- front layer answer. Five, a real human approval gate and is logged. Who's, who reviewed?
- 22:26
What changed? When they signed? That log is audit trail. Notice what's not on the list. A smarter model.
- 22:39
Not one of these is a bigger brain problems. All five plumbing and honesty problems. The winners in the, in this category won't win on benchmark points.
- 22:55
They will win because a tired, skeptical finance pers- f- finance person can trust their output at 11:00 at night without opening seven tabs.
- 23:13
Now the part I promised, the money. Many of you are not just building AI products, you are raising for them,
- 23:24
or you will be. So let me tell you what actually happens after you leave the pitch meeting. Your deck becomes a memo. Literally, literally, literally someone like me sits down and writes an internal memo about you.
- 23:44
Every number you said out loud gets checked, get, get checked against your data room. What, which means everything I just told you about the documents apply to you personally for license.
- 24:00
Okay, my time's up. Thank you. [clapping] [outro jingle]