Contents
  1. Financial workflows, consequences, and decision roles
  2. The meaning of a financial fact
  3. Availability, revisions, and as-of answers
  4. Permitted financial information
  5. Research assistance with inspectable evidence
  6. Reconciliation and unresolved discrepancies
  7. Customer assistance and accountable handoffs
  8. Financial authority and institutional oversight
  9. Confirmed financial changes and uncertain outcomes
  10. Reconstructable decisions and proportionate audit evidence
  11. Evidence of completed work and consequential errors
  12. Historical validity and difficult operating conditions
  13. Deployment scope and demonstrated value
  14. Monitoring, correction, and authorized resumption
  15. Check understanding
  16. Open questions
  17. Selected talks
  18. References
  19. Talk library
← All topics

AI in Finance

Financial AI connects information to decisions about money, records, and customer obligations. Research, reconciliation, and customer assistance expose different parts of that connection. Their usefulness depends on preserving what information means, establishing when and how it may be used, and verifying the outcome beyond the assistant’s response.

Financial workflows, consequences, and decision roles

A financial instrument is a contract creating a financial asset for one party and a liability or equity interest for another. Loans and bonds establish contractual obligations; shares represent equity interests. These distinctions matter because extracting a contract’s terms, explaining them, and exercising rights under it are different activities. The AASB definition supplies this foundational vocabulary.

WorkflowExisting work and possible assistanceResponsibility and completion
ResearchAn analyst gathers filings and earnings material. AI can help locate, extract, and summarize evidence.The analyst checks the analysis; producing a draft does not make an investment decision.
ReconciliationOperations staff compare records. AI can help investigate relationships and prioritize discrepancies.An accountable operations owner resolves cases against the underlying records.
Customer assistanceService staff interpret requests and account information. AI can explain supported facts or prepare a handoff.Completion requires resolution or usable human assistance, rather than merely a delivered answer.

These are responsibilities, not universal job titles. An accountable owner decides whether the result is acceptable and who must resolve exceptions. A bounded institutional example is Morgan Stanley’s announced meeting assistant: it drafts an email for an advisor to edit and send, while another output enters a customer-record system. Drafting and sending remain distinct, and record writes need their own controls.

Here, risk exposure means the nature and amount of potential loss associated with a position, obligation, or action. Incorrect research can influence a decision; a mistaken posting changes records; disclosure can harm a customer. Operational risk includes losses from failed processes, people, or systems. Scale, deadlines, detectability, and reversibility determine how far a small error can spread.

Delegation can permit assistance, a prepared proposal, human-confirmed execution, or explicitly bounded automation. These levels should be assigned per operation. Accuracy and confidence do not grant permission to act.

The meaning of a financial fact

A financial value needs an entity or instrument, measurement definition, currency, scale, sign convention, period, and source location. Assets are resources; liabilities are obligations; equity is the residual interest. A balance sheet describes a date. Revenue, expenses, and profit describe activity over a period; cash-flow statements describe cash movements. The SEC statement guide explains why profit and cash are different quantities.

The published xBRL-JSON teaching fixture illustrates how qualifiers make an amount usable.
Fixture elementMeaning preserved
EUR 1,230,000A monetary amount with an explicit currency, rather than an unqualified number.
Entity and assets conceptWhose assets are described and which accounting quantity is reported.
Instant: 2020-01-01 at midnightThe fixture represents the end of December 31, 2019; it is not an annual flow.
Fact identifier and declared precisionA reference to the fact and its numerical precision, neither of which establishes factual truth.

A ledger maintains financial account entries. In double-entry bookkeeping, total debits equal total credits. Providing a service on credit increases accounts receivable and revenue without receiving cash. Later collection increases cash and decreases the receivable; it does not record revenue again. Balanced entries can still use the wrong accounts, amounts, or dates. A counterparty is the other party to a financial transaction.

SourceClaim it can support
FilingWhat the identified reporting entity disclosed, within the filing’s concepts, units, and scope.
Transaction recordThe receiving system’s recorded operation and status; the status must be interpreted.
Market dataAn observation supplied under particular data-use terms, not unrestricted permission to reuse it.
ContractThe specified parties’ financial rights and obligations.
Product terms or service procedureThe applicable product’s handling rules, not another product’s rules or an account’s current state.

Extraction must retain these qualifiers with the value. Evidence-bearing fields and records covers that representation boundary. A currency-free “1.23” cannot safely substitute for the fixture’s monetary fact.

Availability, revisions, and as-of answers

An as-of answer uses an explicit information cutoff. The reporting period describes the underlying activity; publication makes a version available externally; permissions determine who may use it; ingestion determines when the application receives it. A figure about an earlier quarter can therefore be unavailable at a later decision date. Historical reconstruction must preserve these distinctions.

A restatement corrects previously issued financial reporting. The IAS 8 overview explains retrospective correction of material prior-period errors. A later report can change an earlier period’s value without changing what was available for the original decision.

One fact, different usable versions

Example

Current corrections cannot replace historical evidence silently.

1 / 3 · Before availability

The application lacks usable evidence.

The fact persists. Original and revised versions remain distinct; eligibility depends on the answer’s cutoff.
Read the diagram as text
  • Reported financial fact.
  • No usable version.
  • Original source version.
  • Original usable.
  • Corrected source version.
  • Original supports earlier replay.
  • Correction supports current answer.
  • Reported financial factNo usable version: availability.
  • Reported financial factOriginal source version: reported by.
  • Original source versionOriginal usable: eligible now.
  • Reported financial factCorrected source version: corrected by.
  • Original source versionOriginal supports earlier replay: retained for.
  • Corrected source versionCorrection supports current answer: eligible now.
  1. Before availability. The application lacks usable evidence. Active: Reported financial fact, No usable version. New: Reported financial fact, No usable version.
  2. Original available. The original becomes eligible. Active: Reported financial fact, Original source version, Original usable. New: Original source version, Original usable.
  3. Correction available. New current evidence preserves the original historical version. Active: Reported financial fact, Original source version, Corrected source version, Original supports earlier replay, Correction supports current answer. New: Corrected source version, Original supports earlier replay, Correction supports current answer.

Present-day historical data is not automatically historical evidence. FRED ordinarily returns today’s information, including revisions; vintage queries request an earlier information set. EDGAR calendar frames select a last-filed fact fitting the period. Neither an old observation date nor an old frame label establishes that the selected version was available then.

First exclude versions unavailable at the cutoff, then select the applicable event and version. Event-time filtering alone can admit later backfills. A created timestamp is useful only if it represents actual availability and the retrieval implementation enforces that meaning. Preserve original versions, corrections, and transformation history; stale ingestion and unresolved source conflicts should remain visible.

Permitted financial information

Access, permitted purpose, and disclosure authority are separate conditions. Entitlements specify which users and uses are permitted. Rights, restrictions and decision authority explains the general boundary. In financial applications, the permission to read information internally must not silently become permission to send it to a model provider or disclose it to a customer.

Material nonpublic information is a legal category, not a synonym for confidential data. In the bounded U.S. SEC discussion, materiality concerns information important to a reasonable investor; nonpublic information has not been generally disseminated. Information barriers restrict passage between institutional functions. The applicable jurisdiction and activity determine the precise legal test.

Permission at each information crossing

Example

Internal access does not authorize external processing.

This proposed path checks each transfer separately. A failed condition ends processing at that boundary.
Read the diagram as text
  • Customer records.
  • Authorized team.
  • Institution processing.
  • External model provider.
  • Permitted result.
  • Processing denied.
  • Customer recordsAuthorized team: data: entitlement permits access.
  • Authorized teamInstitution processing: data: purpose and barriers permit.
  • Institution processingExternal model provider: data: disclosure contract permits.
  • External model providerPermitted result: data: permitted output.
  • Customer recordsProcessing denied: control: entitlement fails.
  • Authorized teamProcessing denied: control: internal use fails.
  • Institution processingProcessing denied: control: external use fails.

Public visibility does not establish AI-use permission. CME’s website market-data terms, for example, restrict covered data uses and expressly prohibit specified AI processing. These terms do not describe every exchange or separately negotiated license.

The institution must establish the actual provider contract: permitted inputs and purposes, training use, retained copies, downstream processors, derived-output restrictions, and deletion arrangements. Vendor and model-processing obligations supplies the broader framework. FINRA’s notice makes clear that using third-party GenAI does not remove a member firm’s existing supervisory responsibilities.

Research assistance with inspectable evidence

Grounding connects a claim to evidence that supports it. A citation enables inspection; it does not establish source quality or correct extraction. Applicable sources and unresolved disagreement and Claim support and resolvable citations explain the general mechanism.

A constructed report comparison contains revenue of EUR 120 million and EUR 144 million. Assume the same entity, equal-length periods, consistent currency and scale, matching revenue definitions, and source versions appropriate to the question. These conditions make the arithmetic meaningful; selecting two numbers that happen to be labeled revenue is insufficient.

Evidence before arithmetic

Example

Correct calculation depends on comparable inputs.

Both qualified sources feed comparability checks. Calculation and drafting remain upstream of analyst acceptance; unsupported comparisons are held.
Read the diagram as text
  • Period A source. Qualified revenue: EUR 120 million.
  • Period B source. Qualified revenue: EUR 144 million.
  • Check comparability.
  • Calculate change.
  • Draft analysis. Retain qualifications and unresolved claims.
  • Analyst review.
  • Accepted analysis.
  • Held for correction.
  • Period A sourceCheck comparability: data: fact and qualifiers.
  • Period B sourceCheck comparability: data: fact and qualifiers.
  • Check comparabilityCalculate change: control: comparable.
  • Check comparabilityHeld for correction: control: unresolved mismatch.
  • Calculate changeDraft analysis: data: checked result.
  • Draft analysisAnalyst review: data: analysis and evidence.
  • Analyst reviewAccepted analysis: control: accepted.
  • Analyst reviewHeld for correction: control: correction required.
g=R2R1R1=144120120=0.20=20%g=\frac{R_2-R_1}{R_1}=\frac{144-120}{120}=0.20=20\% Here, R1R_1 and R2R_2 are comparable period revenues and gg is their relative change. At R1=0R_1=0, this percentage change is undefined.
StatementEvidence status
The reports state 120 and 144 million.Reported facts: verify source locations and qualifiers.
Revenue increased by 20%.Calculated result under the stated comparability assumptions.
Demand caused the increase.Interpretation requiring additional supporting evidence.
Revenue will increase again.A prediction not established by these two reported values.

Non-GAAP measures are adjusted measures outside the applicable standard accounting presentation. Similar labels can conceal different calculations, and changed adjustments can undermine comparisons across periods. The SEC’s interpretations describe these problems. If an adjustment definition changes, retain both definitions and qualify or withhold the comparison until a consistent basis is established.

Assistance can move document search, extraction, and first-draft preparation into a reviewable pipeline. The analyst still checks applicability, definitions, calculations, and unresolved claims. Version extraction templates and compare results on representative documents. BlackRock’s described workflow gives domain experts a sandbox for that iteration; short, simple documents and complex instruments may require different extraction strategies. No time saving follows merely from introducing the sandbox.

Reconciliation and unresolved discrepancies

Reconciliation compares separately maintained records to identify, explain, and resolve differences. A reconciliation break is an unresolved discrepancy. Connecting records can reveal inconsistencies that individual document checks miss; the useful output is an investigable case, not an automatic conclusion of fraud.

Posting records an entry in the accounting system. Settlement means completion of the relevant cash or asset transfer. They are different completion boundaries: recording a receivable or another obligation does not establish that cash arrived.

Two records, separate closure evidence

Example

An explanation does not close a break.

The proposed process preserves unresolved cases. Corrections require authorization and subsequent verification against records.
Read the diagram as text
  • Internal entries.
  • External statement.
  • Match and investigate.
  • Unresolved break.
  • Proposed adjustment.
  • Authorized posting.
  • Verify case outcome.
  • Verified resolution.
  • Internal entriesMatch and investigate: data: internal records.
  • External statementMatch and investigate: data: external records.
  • Match and investigateUnresolved break: control: evidence insufficient.
  • Match and investigateProposed adjustment: control: correction supported.
  • Match and investigateVerify case outcome: control: correspondence established.
  • Proposed adjustmentAuthorized posting: control: approval and checks pass.
  • Proposed adjustmentUnresolved break: control: approval absent.
  • Authorized postingVerify case outcome: data: receiver evidence.
  • Verify case outcomeVerified resolution: control: discrepancy resolved.
  • Verify case outcomeUnresolved break: control: discrepancy remains.

Differences can reflect timing, a bank fee, duplicate-looking records, a partial payment, or several internal entries corresponding to one statement amount. Those are investigation hypotheses. Matching totals alone cannot determine which explanation applies or whether an adjustment is needed.

Rules are a concrete baseline. Dynamics can select the first qualifying transaction, with configurable manual handling for ambiguous amount matches. AI assistance should be tested against that process, especially where descriptions require interpretation.

In Dynamics reconciliation, marking a bank-originated item new is not posting it. A reconciled statement can also carry unmatched items forward. Preserve case-level unresolved status.

Customer assistance and accountable handoffs

Financial service begins with the customer’s actual problem. A generic policy explanation may leave an account-specific issue unresolved. The CFPB chatbot report describes failures involving inaccurate answers and inaccessible human assistance. Resolution and a usable service route must remain explicit outcomes.

Pending means a transaction has not fully processed or posted; posted means it has been recorded on the account. A pending credit-card amount can change, for example when a restaurant tip is added. Processing depends on the parties and transaction type, so an assistant should preserve the observed status without inventing a completion date.

Product and status determine the route in Chase’s dispute procedure.
Observed casePublished route
Pending credit-card chargeThe charge must post before a dispute can open.
Pending debit-card transactionA telephone dispute route is available; online debit disputes require posting.

Request meaning also matters. Within covered U.S. consumer electronic transfers, Regulation E’s interpretation distinguishes asking whether a transfer occurred from alleging an error. A lost access device accompanied by an allegation of possible unauthorized use can trigger error handling. These distinctions do not supply a universal rule for credit cards or every payment product.

General product information, account-specific factual assistance, personalized recommendations, and transaction initiation need separate authority decisions. Identity and entitlement checks precede account access. When escalation is needed, provide the request, relevant evidence, proposed action, unresolved conflict, and likely consequences. A reviewer cannot meaningfully authorize an opaque operation merely because it ends with a confirmation button.

A handoff transfers responsibility only when an empowered recipient accepts it. Keep an identified owner while acceptance is pending; preserve unresolved status after rejection or silence. Shared case context can support both a human interface and model input, but does not confer identical permissions. Accepted handoffs and responsibility develops that operating contract.

Financial authority and institutional oversight

Authorization binds an action to its actor, account, recipient or instrument, amount, currency, operation, limits, and validity period. Preserve that binding through retries and callbacks. Changed material details require renewed approval, and current permissions must still allow execution. Approval, changing state, and retries explains the general mechanism.

Maker-checker control separates preparation from independent approval where required. Segregation of duties more broadly separates conflicting responsibilities. Basel’s internal-control principles distinguish committing a bank, paying funds, and accounting for assets and liabilities. They call for appropriate delegation, approval limits, and control resources; they do not impose identical dual approval on every action.

Approval binds specific instructions

Example

Changed details cannot inherit earlier approval.

A proposed control design checks evidence, approval scope, and current conditions before execution. Rejection ends this attempt.
Read the diagram as text
  • Financial proposal.
  • Evidence check.
  • Scoped approval. Binds material details and validity.
  • Fresh execution gate. Checks current authority and state.
  • Execute approved operation.
  • Reject; reassess before resubmission.
  • Financial proposalEvidence check: submit for checking.
  • Evidence checkScoped approval: evidence sufficient.
  • Evidence checkReject; reassess before resubmission: evidence insufficient.
  • Scoped approvalFresh execution gate: approval granted.
  • Scoped approvalReject; reassess before resubmission: approval denied.
  • Fresh execution gateExecute approved operation: details match; checks pass.
  • Fresh execution gateReject; reassess before resubmission: changed, expired, or disallowed.

Enforce the final decision outside the model. A prompt instruction cannot substitute for a gate controlling actual system access and actions.

EU payment dynamic linking supplies a bounded example: under its applicable authentication conditions, changing amount or payee invalidates the code.

Model risk is potential harm from incorrect outputs or inappropriate model use. Effective challenge is qualified scrutiny with enough independence and influence to change a decision. Federal Reserve SR 26-2 describes these practices for its scope but expressly excludes generative and agentic AI. Applying them here is an engineering analogy, not a GenAI regulatory mandate.

A responsibility allocation need not prescribe an organizational chart.
ResponsibilityDecision or evidence owned
Business ownershipAccept intended use, residual risk, and operating scope.
Technical operationMaintain implementation, monitoring, and documented changes.
Risk and compliance reviewAssess applicable requirements, limitations, and exceptions.
Independent validationChallenge suitability and evidence; track unresolved findings.
Internal auditAssess whether controls remain appropriate and effective.

Oversight requires evidence access, expertise, review time, and authority to reject or suspend work. A second model does not automatically supply institutional independence. An escalation should explain the proposed action, the suspected constraint conflict, and its consequences; exhausted review capacity requires a narrower operating scope rather than nominal approval.

Confirmed financial changes and uncertain outcomes

A system of record is the designated authoritative source for a business fact. Transaction status, account balance, and customer-case ownership may belong to different systems. Integration and confirmed effects connects their identifiers, field meanings, and required completion evidence.

An approved adjustment can commit while its acknowledgment is lost. The figure assumes a receiver modeled on Stripe’s documented recovery behavior; it is not a claim about a particular ledger.

Committed effect, uncertain caller

Example

A missing acknowledgment changes knowledge, not the posting.

1 / 3 · Approved

One operation has authorization.

Receiver state and caller knowledge are separate. Authoritative evidence resolves the same operation; it does not create another adjustment.
Read the diagram as text
  • Adjustment A.
  • Approval for A.
  • Receiver committed A.
  • Caller outcome unknown.
  • Receiver evidence for A.
  • Caller outcome confirmed.
  • Approval for AAdjustment A: authorizes.
  • Adjustment AReceiver committed A: submitted and applied.
  • Receiver committed ACaller outcome unknown: acknowledgment lost.
  • Receiver committed AReceiver evidence for A: established by.
  • Receiver evidence for ACaller outcome confirmed: resolves knowledge.
  1. Approved. One operation has authorization. Active: Adjustment A, Approval for A. New: Adjustment A, Approval for A.
  2. Acknowledgment lost. The receiver applied A; the caller cannot establish that yet. Active: Adjustment A, Approval for A, Receiver committed A, Caller outcome unknown. New: Receiver committed A, Caller outcome unknown.
  3. Outcome reconciled. New receiver evidence resolves A while retaining its identity. Active: Adjustment A, Approval for A, Receiver committed A, Receiver evidence for A, Caller outcome confirmed. New: Receiver evidence for A, Caller outcome confirmed.

Idempotency provides bounded protection against repeating an operation. Stripe retries reuse the same key and parameters; changed parameters produce an error. The stored response can include an HTTP 500. Keys may be pruned after at least 24 hours, and reuse after pruning starts a new request. Validation failures and concurrent execution conflicts need not create saved results. These limits prevent treating a key as permanent exactly-once processing.

The receiving contract must specify status lookup, concurrent-version checks, accounting dates, batch cutoffs, and completion states. Preserve operation identity, attempts, and external evidence. If the available records cannot establish the effect, leave it unresolved rather than converting a timeout into a declared financial failure.

Rolling back application code does not reverse a financial effect. Compensation is an explicit corrective operation, potentially requiring new authorization. A Saga can invoke compensations in reverse completion order when those operations exist; that pattern does not guarantee that every payment or posting is reversible, or that compensation itself succeeds.

Reconstructable decisions and proportionate audit evidence

An audit trail is an attributable chronological record of relevant evidence, decisions, approvals, and effects. The SEC electronic-recordkeeping release illustrates reconstructable modifications and deletions within its specified broker-dealer and security-based-swap scope. A success message alone cannot recreate the original and intermediate records.

A proposed decision record joins distinct evidence, rather than one narrative explanation.
Record componentReconstruction purpose
Case, source references, and versionsIdentify the original evidence and subsequent corrections.
Model, configuration, and policy versionsIdentify the execution context and applicable checks.
Approved details and enforcement resultDistinguish permission recorded from permission actually checked.
Attempts and receiver confirmationDistinguish intended action from established external effect.

Recorded rationale can help inspection without faithfully exposing internal model reasoning. Likewise, evidence available to a reviewer, evidence accessed, and evidence substantively reviewed are different claims. Missing links remain gaps; a generated explanation cannot repair them.

Define necessary evidence, authorized readers, retention triggers, and disposal before logging. Protect integrity and retain correction history without indiscriminately copying personal data. Lineage and proportionate audit evidence explains the general design; financial record categories and applicable obligations determine actual retention.

Evidence of completed work and consequential errors

Evaluate complete scenarios against the existing process. Coverage and independent assessment defines population and slice boundaries; finance adds workflow-specific outcomes and consequential errors.

WorkflowOutcome and error evidenceCounting and effort
ResearchScore numerical correctness separately from evidence selection and methodology.Report claims and completed analyses assessed, missing judgments, and total analyst checking and correction time.
ReconciliationInspect false matches, unresolved amounts by currency, break age, and independently verified closure.Count candidate matches separately from completed cases; include investigator effort.
Customer assistanceAssess correct resolution, improper disclosure or commitments, repeat contacts, and usable escalation.Use eligible service cases over a stated follow-up window; retain unresolved cases.

Separate error frequency from severity. Report abstentions, missing judgments, review workload, latency, and recovery alongside completion. A narrow average can conceal consequential failures or transferred human effort.

Similar aggregate scores can conceal opposite weaknesses in arithmetic and methodology, as one private financial evaluation reported. A checked calculator can establish arithmetic without validating the selected inputs or analytical approach. Compare the same task population and preserve rubric-level results rather than selecting a model from one combined score.

Human review also needs testing. A reported Duolingo experiment placed fabricated alerts into legitimate historical sessions and observed reviewers accepting some of them. It supports testing whether reviewers reject deliberately wrong evidence, not assuming that ordinary calibration scores guarantee resistance to automation bias. This is cross-domain evidence, not a financial error rate.

Financial outcomes can arrive late and change status. A dispute may still await evidence or issuer review at the evaluation cutoff. Preserve payment time, observation time, reason, and evolving status. No observed dispute means none has been observed yet; it does not establish a legitimate payment. Label meaning and observation limits explains this boundary.

A fraud-coded dispute records an allegation of unauthorized payment, which can include a legitimate charge the customer did not recognize. A reported dispute, adjudicated case outcome, and separately investigated fraud are different targets. Select the target that matches the service claim instead of silently substituting one for another.

Historical validity and difficult operating conditions

Look-ahead bias uses information unavailable at the simulated decision time. Independent data boundaries explains leakage generally. A historical case should record its decision cutoff, permitted sources, source versions, and ingestion state. Selecting documents about earlier periods is insufficient when their values were published or revised afterward.

Later knowledge can also reside in model parameters. Research on pretrained language models found later pandemic information appearing in analyses of pre-pandemic earnings calls. Instructions against using later information did not eliminate the demonstrated leakage, and masked identifiers could sometimes be inferred. Restricting retrieved documents therefore cannot independently certify historical validity.

For a claim about future operation, move the evaluation origin forward and use only earlier information for each case. Preprocessing and model selection must respect that boundary too. Test new entities separately when the claim concerns them; legitimate prior customer history need not be removed when the intended task allows it.

Constructed exception cases specify starting state, expected behavior, and independent outcome evidence.
ExceptionRequired behaviorEvidence to inspect
Missing or stale feedHold unsupported claims and invoke continuity arrangements.Source freshness, affected work, and accepted operational ownership.
Wrong company or irrelevant contextWithhold unsupported conclusions.Whether the answer remains constrained by the supplied evidence.
Ambiguous records or partial paymentRetain unresolved state until evidence distinguishes candidates.Consistent balances and case state across dependent tool calls.
Hostile document requesting extra authorityReject out-of-scope actions.The independent action gate’s decision and recorded attempt.
Volume surge or unavailable reviewersConstrain intake or authority and preserve service continuity.Backlog, ownership, and exercised recovery procedures.
Failure after earlier successful actionsRecover without treating the episode as a fresh independent task.State transitions and independently recorded outcomes.

A multi-turn simulation must preserve state across calls: a payment cannot be simultaneously absent in one tool and posted in another without an explicit consistency condition. Scenario coverage should name workflows, institutions, customer groups, periods, and exception types. Constructed difficult cases reveal failure mechanisms; their selected frequency is not a production failure-rate estimate.

Deployment scope and demonstrated value

Shadow operation runs beside the existing process without controlling its actions. Historical testing, shadow observation, restricted assistance, and separately approved expansion form a useful engineering progression. It is a proposed release design, not a universal banking requirement.

Measure checking, queueing, integration, support, and customer outcomes across the combined workflow. Observed savings do not alone establish a deployment effect; Live evidence and causal improvement explains credible comparisons.

Exposure and authority need separate gates

Example

A larger assisted population need not gain more authority.

This proposed progression requires acceptance at each transition. Any unmet condition holds scope; financial authority requires a separate decision.
Read the diagram as text
  • Protected historical tests.
  • Shadow observation. No control over live actions.
  • Restricted assistance.
  • Broader assisted population. Existing authority unchanged.
  • Separately approved authority.
  • Hold current scope.
  • Protected historical testsShadow observation: offline criteria pass.
  • Protected historical testsHold current scope: offline criteria fail.
  • Shadow observationRestricted assistance: live evidence and ownership accepted.
  • Shadow observationHold current scope: acceptance incomplete.
  • Restricted assistanceBroader assisted population: quality and capacity permit expansion.
  • Restricted assistanceHold current scope: quality or capacity insufficient.
  • Broader assisted populationSeparately approved authority: action evidence and approval obtained.
  • Broader assisted populationHold current scope: authority case not established.

Versioned prompt cohorts can expose regressions before wider release. Prewire agent-wide and per-tool stops, and record configuration changes. Acceptance and stop thresholds must fit the institution’s severity, workload, and review capacity; suggested rollout percentages are not universal evidence.

Shared model and infrastructure providers create dependencies beyond local answer quality. The FSB identifies concentration, correlated behavior, cyber risk, and model risk as potential financial-stability channels. A locally useful assistant still needs a workable continuity plan; a provider switch requires evidence of compatible behavior rather than an assumption that another endpoint is interchangeable.

Monitoring, correction, and authorized resumption

Operational records should connect the versions used to the actions taken. Protected references can separate sensitive evidence from general diagnostic events. Those references enable affected work to be located without granting every operator access to the underlying customer data.

An operating contract assigns institution-specific thresholds and accountable responses.
SignalResponsible responseResumption evidence
Stale sources or changed formatsData and application owners inspect affected extraction.Corrected ingestion and representative replay.
Old breaks or growing review queuesOperations owner constrains work and assigns investigators.Verified backlog disposition and sustainable review capacity.
Incorrect outputs after a prompt changeApplication owner identifies the version and affected cases.Reviewed correction and regression results.
Complaint or disputed decisionEmpowered reviewer investigates and owns correction.Recorded disposition, downstream correction status, and customer communication.

Find affected work before correcting it

Example

A changed source can reach several business records.

This constructed dependency chain identifies reassessment targets. Dependency does not itself require reversal; each owner determines the justified correction.
Read the diagram as text
  • Corrected source.
  • Affected analysis.
  • Adjustment proposal.
  • Posted record.
  • Customer case.
  • Responsible owners reassess.
  • Corrected sourceAffected analysis: prior version supported.
  • Affected analysisAdjustment proposal: informed.
  • Adjustment proposalPosted record: separately authorized and posted.
  • Posted recordCustomer case: informed status.
  • Affected analysisResponsible owners reassess: analysis needs review.
  • Adjustment proposalResponsible owners reassess: proposal needs review.
  • Posted recordResponsible owners reassess: effect needs review.
  • Customer caseResponsible owners reassess: customer outcome needs review.

A stop control limits future execution at checked decision points; it does not undo completed effects or necessarily interrupt an executing tool. Continue observing in-flight work. Drill permanent controls, record who changed them, and alert on activation. Removing a temporary rollout flag is different from removing emergency shutdown capability.

Contestability gives an affected person a route to challenge an outcome. Link the complaint to the original decision and evidence, then assign a reviewer authorized to inspect and correct it. Record the rationale, revised disposition, downstream correction status, and communicated result. NIST’s voluntary framework supports these practices; a trace alone supplies no correction path.

Changes to sources, terms, permissions, model configurations, providers, or intended use can invalidate earlier acceptance evidence. Reassess the affected behavior, preserve manual continuity, and test recovery before restoring authority. Failure localization and competing explanations supports diagnosis; resumption remains an accountable operating decision.

Open questions

  1. AI assistance on ambiguous reconciliation cases still needs comparative evidence. Plausible explanations may increase checking work; progress requires matched cases showing fewer false matches or lower total review effort without hiding unresolved discrepancies.

  2. Historical LLM assessment remains difficult when later knowledge may reside in parameters. Clean document cutoffs cannot isolate that channel. Progress would include independently specified information boundaries and tests sensitive to leakage without assuming that every correct historical prediction is contaminated.

  3. Review capacity must preserve independent judgment under increasing throughput. Expertise alone may not prevent acceptance of misleading evidence. Progress requires measuring error rejection, review time, and queue pressure together, then establishing operating limits that maintain effective scrutiny.

  4. Recovery across receiving systems lacks a universal contract. Key retention, uncertain effects, and corrective operations vary, making generic retry logic dangerous. Progress requires receiver-specific interruption tests that demonstrate duplicate protection and separately authorized correction after the original outcome is resolved.

Follow the curated reading path through the speakers and demonstrations behind this entry.

18 min

AI Engineer World's Fair 2026 · 2026

State of Data

Sean Cai

Cited in this entry

Separates financial arithmetic from methodology and develops evaluation around dependent actions and recovery.

Watch talk

Explore more talks

The rest of the library, beyond the curated path. Cited talks support this entry; reviewed transcripts were processed in full. Metadata candidates have not been reviewed as sources or verified as topic members.

64 matching talks

TalkSpeakerEventYear
Varsha ShahAI Engineer World's Fair 20262026
Waseem AlshikhAI Engineer Summit 20252025
Leo PekelisAI Engineer World's Fair 20242024
Ritvik PandyaAI Engineer World's Fair 20262026
Devendra Chaplot, Devendra Singh ChaplotAI Engineer World's Fair 20242024
Ahmed MenshawyAI Engineer World's Fair 20242024
Cedric ClyburnAI Engineer World's Fair 20262026
Roy DerksAI Engineer Summit 20252025
Daniel WhitenackAI Engineer World's Fair 20242024
Lawrence JonesAI Engineer Europe 20262026
Lance MartinAI Engineer World's Fair 20242024
Udi MenkesAI Engineer World's Fair 20262026
Jesse HuAI Engineer Code 20252025
Parth AsawaAI Engineer World's Fair 20262026
Vinoth GovindarajanAI Engineer World's Fair 20262026
Sandipan BhaumikAI Engineer Europe 20262026
Agents Need Feature Flags

Cited in this entry

Sachin GuptaAI Engineer World's Fair 20262026
Christopher Lovejoy, Saul HowardAI Engineer World's Fair 20262026
Sumaiya ShrabonyAI Engineer World's Fair 20262026
AI’s Jurassic Park Period

Cited in this entry

Aaron StanleyAI Engineer World's Fair 20262026
Angus J. McLeanAI Engineer Europe 20262026
Sarthak AggarwalAI Engineer World's Fair 20262026
Chaitanya AsawaAI Engineer World's Fair 20262026
Ofer MendelevitchAI Engineer World's Fair 20252025
Dan MasonAI Engineer World's Fair 20252025
Jeremy Silva, Chris HernandezAI Engineer World's Fair 20252025
Stephen ChinAI Engineer Europe 20262026
Nishant GuptaAI Engineer World's Fair 20262026
Sachin GuptaAI Engineer World's Fair 20262026
Giran Moodley, Mayan Soni, Oussama Hafferssas, Mayank SoniAI Engineer Europe 20262026
Ramana Siddanth EmaniAI Engineer World's Fair 20262026
Ayush BhardwajAI Engineer World's Fair 20262026
Vibhor KumarAI Engineer World's Fair 20242024
Nathan WanAI Engineer World's Fair 20252025
Vasuman MozaAI Engineer World's Fair 20262026
Shawn ChanAI Engineer World's Fair 20262026
Anju KambadurAI Engineer Summit 20252025
Rachna SrivastavaAI Engineer World's Fair 20252025
Sahil Yadav, Hariharan GanesanAI Engineer World's Fair 20252025
Sheila Gulati, Nischal NadhamuniAI Engineer World's Fair 20242024
Nina Lopatina, Rajiv ShahAI Engineer World's Fair 20252025
Fuzzing in the GenAI Era

Metadata candidate

Leonard TangAI Engineer World's Fair 20252025
Yogendra MirajeAI Engineer World's Fair 20252025
Mustafa Ali, Kyle CorbittAI Engineer Summit 20252025
Mitesh PatelAI Engineer World's Fair 20252025
Identity for AI Agents

Metadata candidate

AI Engineer Code 20252025
Divakar KumarAI Engineer World's Fair 20262026
Juan Herreros ElorzaAI Engineer Europe 20262026
Yuval Belfer, Niv GranotAI Engineer World's Fair 20252025
Kshitij GroverAI Engineer World's Fair 20252025
Scaffold Wisely

Metadata candidate

Rahul SengottuveluAI Engineer Summit 20252025
Laurie VossAI Engineer Europe 20262026
Yogendra MirajeAI Engineer World's Fair 20262026
Kobie CrawfordAI Engineer Europe 20262026
Stop Using RAG as Memory

Metadata candidate

Daniel ChalefAI Engineer World's Fair 20252025
Jim BennettAI Engineer World's Fair 20252025
Nuno CamposAI Engineer Europe 20262026
Sandipan BhaumikAI Engineer Europe 20262026
Emil EifremAI Engineer World's Fair 20262026
Mike ConoverAI Engineer Summit 20252025
Lucas PalmaAI Engineer World's Fair 20262026
Peter GostevAI Engineer Europe 20262026
Soumith ChintalaAI Engineer Summit 20252025
Rustin BanksAI Engineer World's Fair 20252025

References

Coverage and source review
Processed transcripts
35 processed in full · 5 in the curated path
Automated source review
Passed
Metadata candidates
34 unreviewed; not verified topic membership
Corpus version
1bd8e407b26a07b33815594e1b2db5f41827119a2b3cb6fbf240f9fc571fc767

Automated review checks source support; it is not publication approval.

A synthesis of selected conference talks and technical references. Citations link to the source material; they do not imply that every talk on this subject is included.

  1. AASB 132 Financial Instruments: Presentation

    December 2022 compilation, paragraph 11 and application guidance AG4; foundational vocabulary rather than current accounting classification advice.

  2. FINRA Regulatory Notice 24-09

    June 27, 2024 notice to FINRA member firms; potential uses and existing supervisory obligations.

  3. AI-Driven Multi-Document Correlation for Enterprise Financial Compliance and Fraud Detection

    Deployment must connect existing enterprise data systems and fit prioritized risk cases into the investigator's audit workflow.

  4. CFPB: Chatbots in consumer finance

    June 2023 original issue spotlight; complex problems, human intervention and deficient chatbot risks.

  5. Morgan Stanley: Launch of AI @ Morgan Stanley Debrief

    Firm's original June 2024 launch announcement; bounded institutional example of assistance, review and integration.

  6. Basel Committee: sound management of operational risk

    March 2021 publication; introduction, section 3 paragraphs 1–5 and governance paragraphs 29–31. Payment illustration applies the definition.

  7. Agents Need Feature Flags

    Start with suggestions, add human-confirmed execution per surface, and make autonomous execution an explicit per-tool opt-in.

  8. SEC Beginners’ Guide to Financial Statements

    SEC educational guide; balance-sheet, income-statement and cash-flow explanations.

  9. XBRL International: xBRL-JSON tutorial and examples

    Sections 1, 3–5, 8 and 9.2; a published teaching fixture suitable for an annotated financial fact.

  10. Double-entry bookkeeping and cash movement

    Double-Entry Accounting; Debits and Credits; T-accounts; When Cash Is Debited and Credited; Revenues; Permanent and Temporary Accounts.

  11. Basel Committee: credit risk and contractual obligations

    September 2000 publication, introduction paragraphs 2, 3 and 7; durable definitions only.

  12. SEC: EDGAR Application Programming Interfaces

    Submissions, XBRL Data APIs, company-concept and frames documentation.

  13. Resolving an ambiguous payment request

    Network errors, Server errors and Idempotency; metadata correlation during reconciliation.

  14. CME Group: Market Data Explanation and Disclaimer

    CME website market-data terms inspected on the verification date; an example of contractual information boundaries.

  15. Chase: Disputing a Charge

    Issuer's current dispute FAQ, pending-charge question; concrete product-specific service boundary.

  16. FRED real-time periods and historical information sets

    Real-Time Periods: Introduction and Examples; realtime_start/realtime_end semantics.

  17. Historical feature retrieval with a per-row cutoff

    Official master documentation: Point-in-time joins and Retrieving features as of the event time, including opt-in availability filtering.

  18. IFRS Foundation: IAS 8 Basis of Preparation of Financial Statements

    Official standard overview, prior-period errors; supports distinguishing reporting period from subsequent correction.

  19. Event time and actual availability require separate handling

    Issue #6615 description, concrete backfill example, proposed cutoff and TTL distinction.

  20. SEC: Selective Disclosure and Insider Trading

    2000 adopting release, discussion of Regulation FD materiality and institutional information barriers; U.S. securities context.

  21. How Kepler Built Verifiable AI for Financial Services

    Source attribution alone does not establish source quality or extraction correctness.

  22. FinQA: A Dataset of Numerical Reasoning over Financial Data

    Original paper, dataset design and section 5.1 retriever/program-generator architecture. Evidence for report analysis, not investment advice.

  23. open-rag-eval: RAG Evaluation without "golden" answers.

    Citation faithfulness evaluates the support relationship between a response and its cited passage using graded support categories.

  24. SEC: Non-GAAP Financial Measures Compliance and Disclosure Interpretations

    Questions 100.02–100.05; U.S. SEC staff interpretations concerning non-GAAP disclosures.

  25. How BlackRock Builds Custom Knowledge Apps at Scale

    Treat domain prompts as versioned, comparable artifacts evaluated against a dataset, and invest in domain experts' ability to author them.

  26. How BlackRock Builds Custom Knowledge Apps at Scale

    Choose and iterate the extraction strategy together with the prompts, according to instrument complexity, document length, and model constraints.

  27. How BlackRock Builds Custom Knowledge Apps at Scale

    Separate a domain-facing sandbox from an app factory that packages the resulting configuration into an application.

  28. AI-Driven Multi-Document Correlation for Enterprise Financial Compliance and Fraud Detection

    Document-level validation can miss inconsistencies that appear only when payroll, procurement, and tax records are connected.

  29. Microsoft Dynamics 365: Reconcile bank statements using advanced bank reconciliation

    Statement validation, reconciliation worksheet, new-transaction posting and reversal features; July 2026 documentation.

  30. Microsoft Dynamics 365: Set up bank reconciliation matching rules

    Matching-rule setup, default selection and manual-matching option; July 2026 documentation.

  31. Chase: What are pending transactions and how long do they take?

    Issuer educational explanation of pending credit-card transactions; supports a customer-service status example.

  32. CFPB Regulation E: § 1005.11 Procedures for resolving errors

    Current displayed Regulation E interpretation of § 1005.11(a), comments 2–4; covered U.S. consumer electronic-fund-transfer servicing.

  33. IT Admin for the AI Workforce — Sarthak Aggarwal, Decawork

    Agents with authority and side effects need an operational lifecycle beyond model behavior testing.

  34. AI’s Jurassic Park Period

    Escalations should explain the proposed action, suspected constraint violation, and likely consequences rather than present an opaque command with a yes/no prompt.

  35. Why Your Enterprise Tech Stack Isn't Ready for AI Agents - And What to Build Instead

    Give humans and LLMs a shared action contract, with different presentations of shared context.

  36. Your Agent Didn’t Fail. Your Harness Did.

    Approval must remain bound to one specific action and its scope, identity, arguments, and lifetime; expiration should terminate the approval path.

  37. Basel Core Principles: Principle 26, Internal control and audit

    April 2024 Basel Core Principles, Principle 26 and assessment criteria; institutional responsibility and meaningful checks.

  38. IT Admin for the AI Workforce — Sarthak Aggarwal, Decawork

    A restriction expressed only as an instruction does not prevent an authorized agent from taking destructive actions.

  39. Commission Delegated Regulation (EU) 2018/389, Article 5

    Original EU regulation text, Article 5; concrete financial example of approval bound to material transaction details.

  40. Federal Reserve SR 26-2: Revised Guidance on Model Risk Management

    April 17, 2026 attachment: scope footnote 3; effective challenge; sections IV–VII on development, validation, monitoring, governance and vendors.

  41. Your Agent Didn’t Fail. Your Harness Did.

    Delivery alone is insufficient: a named system of record must persist the fact and support replay into future work.

  42. Stripe retry keys and retention boundaries

    Idempotent requests: response caching, parameter matching, key retention and execution-start exceptions.

  43. Your Agent Didn’t Fail. Your Harness Did.

    Trace one real run from trigger identity through inherited state, authority, execution attempts, and surviving external evidence.

  44. From Chaos to Choreography: Multi-Agent Orchestration Patterns That Actually Work — Sandipan Bhaumik

    The compensation pattern, also called the Saga pattern, pairs execution with explicit undo operations and invokes them in reverse completion order.

  45. SEC: Electronic Recordkeeping Requirements, Release 34-96034

    Final adopting release, pages 21–23 and amended Rules 17a-4 and 18a-6; specified U.S. broker-dealers and security-based swap entities.

  46. Connecting the Dots with Context Graphs

    The proposed reasoning memory preserves the basis of prior decisions for future retrieval, debugging, and compliance review.

  47. NIST Privacy Framework 1.0: lifecycle and minimized audit evidence

    Core ID.IM-P; GV.PO-P1; CT.PO-P; CT.DM-P5/P8; CM.AW-P6; PR.AC-P; PR.DS-P3.

  48. Production Evals For Agentic AI Systems

    Use simulated workflow scenarios and score both task completion and execution quality.

  49. State of Data

    In the speaker's private finance evaluation, similar aggregate model scores concealed opposite weaknesses in arithmetic and methodology.

  50. Production Evals For Agentic AI Systems

    Apply an SRE or production-engineering lens: assess delivered value, operational reliability, human burden, risk, user experience, scalability, and resilience.

  51. Build AI Systems for Discernment, Not Approval - Angel Ortmann Lee, Duolingo

    The When Machines Mislead case study inserted fake copy-typing alerts into legitimate historical exam sessions and found that skilled proctors accepted half of those alerts.

  52. Stripe card-dispute timing and lifecycle

    Before the dispute; inquiries; dispute timing; after the decision. API object separately checked for status definitions.

  53. Stripe dispute categories and contested fraud claims

    Duplicate, Product not received and Fraudulent category explanations, especially the legitimate-charge recognition caveat.

  54. Lookahead Bias in Pretrained Language Models

    Sections 2–5 and Appendix A.1; inspected original manuscript and reported experiments, not locally reproduced results.

  55. Rolling-origin forecast evaluation

    Section 5.10: rolling forecasting origin, multi-step forecast errors and stretch_tsibble example.

  56. Simulation-Maxxing: How Nubank ships agents 20× faster with simulations

    A multi-turn evaluation case must represent a stateful trajectory, not merely an input and expected answer.

  57. Basel Committee: Principles for Operational Resilience

    March 2021 principles 2–7; bank operational continuity, dependency management and incident recovery.

  58. Building Trust in Enterprise AI: Evaluating Domain-Specific LLMs for Real-World Financial Scenarios

    In the reported financial evaluation, thinking models continued answering despite wrong context, which the speaker associated with greater hallucination.

  59. IT Admin for the AI Workforce — Sarthak Aggarwal, Decawork

    An action gate should reject and escalate actions outside the approved password-reset plan, even if the executor proposes them.

  60. State of Data

    A long-horizon evaluation should require dependent actions and recovery, not merely a collection of isolated difficult questions.

  61. From Ambient Documentation to Clinical Intelligence

    The speaker describes evaluation as a lifecycle spanning internal benchmarks, staged clinician rollout, and continual monitoring.

  62. Agents Need Feature Flags

    Route cohorts to versioned prompts and promote a candidate only after observing its behavior against a baseline.

  63. Agents Need Feature Flags

    Ship agent-wide and per-tool kill switches first, and ensure in-flight work checks them at the next decision point.

  64. Agents Need Feature Flags

    Track mitigation effectiveness and record flag changes with enough context to reconstruct an incident.

  65. FSB: Financial stability implications of artificial intelligence

    FSB’s own report summary and risk channels, November 2024. Institutional system-level framing rather than a quantitative forecast.

  66. Why Your Enterprise Tech Stack Isn't Ready for AI Agents - And What to Build Instead

    Use a unified, immutable event log as the source of truth, accepting more complex reads in exchange for reconstructable history.

  67. Why Your Enterprise Tech Stack Isn't Ready for AI Agents - And What to Build Instead

    Separate orchestration events from immutable, schema-driven objects containing sensitive data.

  68. Shipping complex AI applications | Braintrust & Trainline

    Start with a critical part of an already operational system instead of instrumenting the entire suite at once.

  69. Shipping complex AI applications | Braintrust & Trainline

    Managed prompt updates should expose both who changed the prompt and what changed.

  70. NIST AI RMF 1.0: accountability, appeals and override

    GOVERN 2.1–3 and 3.2; MEASURE 3.3; MANAGE 2.4 and 4.1–3; Appendix C. Record fields are an implementation proposal.

  71. Agents Need Feature Flags

    Flags require ongoing drills, lifecycle ownership, and interaction testing.

  72. How Kepler Built Verifiable AI for Financial Services

    Derived numbers need replayable derivation histories spanning source values, calculations, and internal information, not just links to filings.

  73. Your Agent Didn’t Fail. Your Harness Did.

    Internal acceptance does not prove the intended result appeared at the user-visible boundary.

  74. The Build-Operate Divide: Bridging Product Vision and AI Operational Reality

    Treat the prototype-to-reliability transition as a recurring operational iteration loop, described here as crossing a quality chasm.