Contents
  1. Strategic problems and investment mandates
  2. Portfolio constraints and investment sequencing
  3. Lifecycle economics and credible returns
  4. Uncertainty and staged commitments
  5. Accepted responsibilities and decision authority
  6. Capability coverage and team placement
  7. Build, buy, and partner choices
  8. Supplier obligations, change, and exit
  9. Quality expectations and residual-risk decisions
  10. Evidence gates and continuing authorization
  11. Operating budgets and realized results
  12. Field evidence and revised investment assumptions
  13. Workforce commitments and organizational change
  14. Check understanding
  15. Open questions
  16. Selected talks
  17. References
  18. Talk library
← All topics

AI Engineering Leadership

AI engineering leadership connects technical possibilities to organizational commitments. An investment needs a worthwhile problem, credible benefits, capable people, bounded authority, and resources for continuing operation. Accountability means deciding what evidence justifies further spending or broader use—and changing or stopping work when those conditions no longer hold.

Strategic problems and investment mandates

A value hypothesis proposes how an intervention will improve an outcome. A business case compares that proposition with alternatives, costs, risks, and the existing baseline. Strategic fit means the improvement advances an organizational objective important enough to displace other work.

A hypothetical support organization wants faster customer resolution. Its proposed assistant drafts responses for staff review. Better documentation is a competing intervention. The investment mandate below defines the proposed commitment without assuming that response generation is the actual constraint.

Conditions between assistance and value

Example

Task improvement reaches customers only through additional operating conditions.

The support hypothesis connects assistance to resolution through lower total work and useful capacity. Each arrow names a necessary condition, not an observed effect.
Read the diagram as text
  • Response assistance.
  • Less drafting work.
  • Less total work.
  • Usable capacity.
  • Improved resolution.
  • Response assistanceLess drafting work: Acceptable drafts replace work.
  • Less drafting workLess total work: Checking does not absorb savings.
  • Less total workUsable capacity: Released time can be reassigned.
  • Usable capacityImproved resolution: Demand and downstream capacity exist.
Mandate elementSupport-assistant example
Intended benefitImprove resolution without degrading advice.
EvidenceCompare total work and outcomes with existing practice and documentation improvements.
Accountable ownerSupport director accepts outcome responsibility.
ResourcesReserve implementation, reviewer, and analyst capacity.
BoundariesDrafting only; exclude account changes and unsupported advice.
Next decisionReview evidence before committing to wider use.

Task correctness, workflow performance, and business value establish different claims. A correct draft can still require expensive checking or leave resolution unchanged. Value agreement and delivery boundaries develops this distinction. Define success before selecting a model.

Declining work can preserve strategic focus. Factory describes declining customer consulting work that would generate revenue without sufficiently improving its shared product. The general decision is whether the requested work advances the intended business, not merely whether engineers can deliver it.

Portfolio constraints and investment sequencing

A portfolio groups initiatives managed together because they compete for resources and share consequences. Portfolio management aligns their selection and sequencing with strategy and delivery capacity. Individually attractive projects can collectively exceed available expertise.

Opportunity cost is the benefit forgone by using a resource elsewhere. Assigning already-paid specialists to a pilot may leave payroll unchanged while displacing valuable work. Available budget therefore does not establish available delivery capacity.

Shared constraints across initiatives

Example

Separate business cases can promise the same scarce capacity.

Arrows identify required capacity or prerequisites, not execution order. All three initiatives need domain reviewers; the assistant also needs reviewed knowledge and access authorization.
Read the diagram as text
  • Assistant pilot.
  • Knowledge cleanup.
  • Required access review.
  • Domain-review capacity. Finite availability.
  • Reviewed knowledge.
  • Access authorization.
  • Assistant pilotDomain-review capacity: Requires capacity.
  • Knowledge cleanupDomain-review capacity: Requires capacity.
  • Required access reviewDomain-review capacity: Requires capacity.
  • Assistant pilotReviewed knowledge: Depends on.
  • Assistant pilotAccess authorization: Depends on.
Investment categoryReason to fundSequencing implication
Operating improvementAn identifiable service problem.Establish the outcome baseline first.
Uncertain opportunityA potentially valuable new capability.Fund learning before broad commitment.
Required obligationNecessary authorization or data stewardship.Reserve capacity before discretionary expansion.
Shared capabilityRecurring needs across teams.Fund common work without forcing every local requirement into it.

An unavailable reviewer or prohibited data use is a hard constraint; a preferred launch date may be negotiable. Sequence knowledge cleanup and access review before expanding the assistant when they are prerequisites. Attribute their shared contribution once rather than giving every dependent project the entire benefit.

Concentration risk arises when several initiatives depend on the same supplier or information source. They can be affected together by one change. Multiple endpoints do not necessarily diversify that exposure; provider features and supported behavior require inspection.

Lifecycle economics and credible returns

Total cost of ownership covers acquiring, integrating, operating, changing, and retiring a capability over a declared period. Include evaluation, expert review, training, support, failure handling, and exit—not just supplier charges. Complete-task economics explains why subsequent attempts and required review belong in the workload cost.

ROI=100×BCC,C>0\mathrm{ROI}=100\times\frac{B-C}{C},\qquad C>0 Here, return on investment (ROI) is a percentage; BB is attributable monetary benefit and CC is fully loaded program cost for the declared appraisal boundary.

Assume one year of attributable benefits is forecast at $120,000 and program costs at $100,000. Forecast ROI is 20%; the benefit-cost ratio is 1.2. These are assumptions, not realized results. The ROI convention distinguishes the two measures.

Claimed improvementWhat must be established
Less drafting timeChecking and downstream work have not absorbed the reduction.
Usable capacityReleased time can perform identified valuable work; spending may remain unchanged.
Avoided expenditure or cash savingsA credible future expense is avoided, or actual spending falls.
More salesCount attributable contribution after associated costs, not all sales revenue.

Observed improvement may also reflect staffing, demand, or process changes. Isolate the intervention's contribution before monetizing it, and avoid counting the same improvement twice. Benefits realization connects the claim to allocation decisions and continuing verification.

Keep unpriced benefits visible alongside ROI. A favorable ratio does not settle whether remaining risk is tolerable. Economic comparability also requires an explicit benefit period and consistent treatment of direct and indirect costs.

Higher model charges can coexist with lower overall expense. Dan Bjornn reports that a rebuilt application cost more per message but required less maintenance. Without a normalized cost breakdown, that account illustrates a cost category to investigate rather than a transferable savings estimate.

Uncertainty and staged commitments

Staged funding commits enough resources to resolve uncertainty before a larger decision. Technical feasibility, appropriate adoption, economic value, and operating readiness are different uncertainties: a prototype can work while none of the others is established. Experimentation should produce evidence about the investment hypothesis, not just completed features.

The pilot agreement specifies learning objectives, reserved staff, a spending ceiling, a decision date, and stop conditions before results arrive.

Evidence determines the next commitment

Example

Further funding is conditional, not automatic.

These branches govern investment. Expansion still requires separate permission for user exposure. Revised experimentation ends this decision; it does not imply unlimited retries.
Read the diagram as text
  • Bounded experiment.
  • Assess investment evidence.
  • Fund expansion.
  • Fund revised experiment.
  • Defer commitment.
  • Stop investment.
  • Bounded experimentAssess investment evidence: Produces evidence.
  • Assess investment evidenceFund expansion: Case supported; capacity available.
  • Assess investment evidenceFund revised experiment: Case uncertain; useful test remains.
  • Assess investment evidenceDefer commitment: Case viable; prerequisite unavailable.
  • Assess investment evidenceStop investment: Case fails; no justified next test.
Assumption changedEffect on the one-year case
Less eligible demandFewer opportunities to earn the forecast benefit; repeat the appraisal at that volume.
More expert reviewAdditional labor can consume the margin despite unchanged model charges.
No useful redeploymentTime reduction alone cannot support the claimed capacity benefit.
Later useful operationA shorter benefit period can leave too little contribution to cover costs.

Sensitivity analysis varies consequential assumptions. Real-options reasoning values preserving choices to expand, defer, change, or stop; preserving those choices also costs money and time. The Green Book explains these appraisal tools.

Sunk expenditure is already incurred and cannot be recovered by stopping. It does not justify further funding; compare future consequences and alternative uses of resources.

Accepted responsibilities and decision authority

Decision rights specify authority over particular decisions. Responsibility vocabulary distinguishes doing work, answering for outcomes, providing advice, and receiving updates. A delivery assignment does not automatically grant launch or risk-acceptance authority.

A risk owner manages and monitors a specified risk, securing acceptance from the appropriate authority rather than necessarily accepting it personally.

Decision or dutyProposed assignmentAuthority boundary
FundingSponsor with finance inputCommit within written delegation; escalate excess.
Quality assessmentDomain quality leadJudge evidence; identify unacceptable behavior.
Risk managementNamed risk ownerManage exposure; seek authorized acceptance.
DeploymentDesignated service authorityDecide within approved scope and escalation limits.
Incident interventionAssigned operating engineerObserve production; intervene or roll back when necessary.
Benefits verificationBenefits owner with financeVerify whether intended improvement materializes.

These assignments require accepted duties, time, evidence access, authority limits, an available substitute, and an escalation recipient. Combining roles can be practical; conflicting incentives warrant separate challenge. Escalation is necessary when the decision exceeds the current owner's authority.

If delivery wants to launch but the quality lead identifies unresolved consequential errors, the disagreement goes to the designated authority. Delivery cannot resolve it merely by declaring its implementation complete.

Capability coverage and team placement

An operating model arranges work, people, authority, information, and technology to deliver outcomes. AI and the organizational operating model establishes that context; leadership chooses where those capabilities and decisions sit.

Capability coverage extends beyond model development: domain judgment, product discovery, software and data integration, evaluation, operations, security, privacy, economic measurement, and organizational change all require capacity. Record who supplies each capability, their availability, access to evidence, and uncovered duties. A list of job titles cannot establish that the work is covered.

ArrangementUseful characteristicTradeoff to investigate
CentralizedA common team concentrates expertise and delivery.Competing business demands share its prioritization and capacity.
EmbeddedBusiness-unit teams keep development near domain decisions.Repeated needs may receive separate implementations.
FederatedLocal applications use centrally maintained capabilities.Shared and local owners must coordinate boundaries and changes.

Train existing staff when recurring work benefits from their domain knowledge. Hire for durable gaps; temporarily embed specialists or obtain external help when urgency exceeds available capability. Require knowledge transfer and retain internal ability to judge acceptance. These are conditional sourcing choices, not evidence for one optimal staffing mix.

Review and support need explicit capacity. Domain experts must have time to examine consequential cases and diagnose failures. The AI quality-lead proposal combines customer understanding with systematic evaluation work; it does not require that the same person implement production code.

A shared platform serves recurring needs through supported interfaces; it need not implement every local capability. Shared services and workload needs explains that boundary. Common ownership should follow useful reuse rather than a mandate to centralize all work.

Build, buy, and partner choices

Build versus buy concerns implementation responsibility: build takes primary responsibility; buy adopts a supplied capability; partner explicitly shares delivery or specialist work. Apply the choice to components, as the Sourcing Playbook recommends, rather than treating the entire system as indivisible.

Compare arrangements for the same support workload, quality requirements, operating duties, and one-year horizon. Include retained work when comparing complete costs.

ArrangementRetained organizational workDecision tradeoff
Build the workflowIntegration, quality, operation, maintenance, and support.Greater implementation control requires continuing engineering capability.
Buy a workflow productFit assessment, local integration, acceptance, and supplier management.Faster access to existing capability can constrain adaptation and exit.
Partner on implementationDomain decisions, acceptance, and agreed continuing duties.Specialist delivery helps only if knowledge and maintenance obligations remain manageable.

Buying model access supplies a component, not the entire support workflow. Operating an open model also adds infrastructure, lifecycle, observability, and control work. Neither route removes the need to decide which capabilities differentiate the organization and which can be supplied economically.

Forward deployed engineering places implementation close to customer operations. It can bridge a customer's implementation gap. Shared primitives reduce repeated construction, while genuinely customer-specific behavior remains local; neither arrangement eliminates maintenance.

Every arrangement retains an internal owner capable of judging usefulness and acceptance. Supplier expertise supplements that capability; it cannot replace the organization's own domain judgment.

Supplier obligations, change, and exit

Vendor lock-in is the practical burden of changing suppliers or architectures. Custom data formats, training interfaces, and maintenance demands can make switching unaffordable even when the organization owns its data. Bjornn's deployment account illustrates this burden without establishing a universal result for fine-tuning.

Similar APIs reduce some integration work but do not establish equivalent behavior. Provider changes can require renewed evaluation and prompt adjustment. Hosting providers may also differ in tool calling, structured outputs, or caching, so an alternative endpoint is not demonstrated continuity.

Exit requires receiving capability

Example

Cancellation cannot substitute for accepted continuity.

Knowledge and reserved capacity jointly enable readiness exercises. Failed readiness holds the transition. Accepted operation permits transition; verified replacement service supports retirement under the contingency plan.
Read the diagram as text
  • Transfer operating knowledge.
  • Reserve receiving capacity.
  • Exercise continuity. Requires both prerequisites.
  • Accept operating responsibility.
  • Hold transition.
  • Authorize transition.
  • Retire former service.
  • Transfer operating knowledgeExercise continuity: Supplies knowledge.
  • Reserve receiving capacityExercise continuity: Supplies staff.
  • Exercise continuityAccept operating responsibility: Acceptance criteria met.
  • Exercise continuityHold transition: Acceptance criteria unmet.
  • Accept operating responsibilityAuthorize transition: Enables decision.
  • Authorize transitionRetire former service: Replacement operation verified.

Record support coverage, escalation, change notification, custom-work ownership, knowledge transfer, and exit duties. Name the transition authority, budget, staff, and acceptance standard. Contractual commitments require fulfillment evidence.

Access to execution records helps the organization investigate and change its system. Factory emphasizes retaining that access alongside control over information flows. Operating ownership and service retirement places these continuing duties at explicit service boundaries.

Assign an owner to review vendor data obligations, including permitted handling and end-of-service responsibilities. Purchasing authority alone does not settle appropriate information use.

Quality expectations and residual-risk decisions

Risk appetite is the exposure an organization is willing to accept in pursuit of its objectives.

Residual risk is exposure remaining after controls. Its tolerability depends on consequences and frequency, and on those controls actually working.

Quality expectations define acceptable behavior for the intended use. Passing a fixed set of examples leaves other inputs unexamined; nearby wording can expose different behavior. Evaluation therefore supplies bounded evidence, not an automatic permission to proceed.

Risk-record elementSupport-assistant application
Affected people and harmIncorrect advice may mislead customers; identify consequential cases.
Mitigation and capacityPut capable reviewers at consequential decisions, with time to intervene.
Remaining exposureRecord failures still possible and obtain acceptance within delegated authority.

Data governance assigns authority, handling rules, and evidence obligations. Purpose and accountable data use develops those responsibilities. A disputed use of customer records goes to the accountable data owner and appropriate specialists; the investment sponsor must resource the resulting mitigation rather than treating access as an engineering preference.

Evidence gates and continuing authorization

An evidence gate is a decision point with declared requirements and an authorized decision maker. Evaluations systematically assess intended-use criteria; release and revision decisions explains their design. Leaders must connect that evidence to the particular permission being requested.

PermissionRequired judgmentDecision maker
SpendFund the forecast commitment or revise it.Delegated budget authority.
Expose usersAuthorize intended use and remaining risk.Designated deployment authority.
Expand action authorityEvidence supports the additional exposure.Authority responsible for expanded scope.
Accept operationReceiving staff can fulfill continuing duties.Receiving operations owner.

Permission changes; system identity persists

Example

Historical approval does not authorize changed conditions.

1 / 4 · Authorized

A governs current operation.

Assume policy requires suspension after an unreviewed supplier change. Old approval remains recorded; reassessment and new authority precede resumption.
Read the diagram as text
  • Support assistant.
  • Authorization A. Supplier A only.
  • Authorized with A.
  • Supplier changed. Unreviewed supplier B.
  • Suspended.
  • Reassessment evidence.
  • Authorization B. New bounded permission.
  • Authorized with B.
  • Authorization ASupport assistant: Records prior permission.
  • Authorization AAuthorized with A: Authorizes.
  • Support assistantAuthorized with A: Current state.
  • Supplier changedSuspended: Requires pause.
  • Support assistantSuspended: Current state.
  • Reassessment evidenceAuthorization B: Supports decision.
  • Authorization BAuthorized with B: Authorizes.
  • Support assistantAuthorized with B: Current state.
  1. Authorized. A governs current operation. Active: Support assistant, Authorization A, Authorized with A. New: Support assistant, Authorization A, Authorized with A.
  2. Suspended. Supplier change triggers the required pause. Active: Support assistant, Authorization A, Supplier changed, Suspended. New: Supplier changed, Suspended.
  3. Reassessed. New evidence arrives; suspension remains. Active: Support assistant, Authorization A, Supplier changed, Suspended, Reassessment evidence. New: Reassessment evidence.
  4. Reauthorized. Authorized judgment permits bounded resumption. Active: Support assistant, Authorization A, Supplier changed, Reassessment evidence, Authorization B, Authorized with B. New: Authorization B, Authorized with B.

Exposure can increase gradually: shadow operation cannot affect outcomes; advisory use recommends to people; bounded autonomy permits specified actions within limits. Evidence must support each expansion. Human agreement supplies feedback, not an infallible correctness label.

Record the decision's scope, conditions, owner, unresolved issues, and reassessment triggers. An architecture decision record preserves a technical choice and its rationale; enforcement remains separate. Discoverable records let people and agents recover the reason for a rule rather than reconstructing it from memory.

Intercom's automated-review account describes bounded eligibility, an option to request human review, and retained engineer responsibility for production observation and rollback. This illustrates changing an approval practice while preserving operating duties. Its reported pilot does not establish that automated approval is safer for every workload.

Operating budgets and realized results

An operating budget funds continuing service work rather than only initial delivery. Reserve resources for evaluation, expert review, support, changes, failure handling, and eventual retirement. Full ownership cost includes management and labor alongside technology charges.

A forecast states expected spending and value under documented assumptions. FinOps forecasting connects engineering, product, finance, and leadership. The budget owner must manage the commitment or seek additional funding; connect this responsibility to those controlling demand and realizing benefits.

Cost allocation assigns shared expense to responsible groups. Allocation policies can use central funding, fixed shares, or usage-related proportions. Make the basis visible so project accounts neither hide common costs nor pretend that accounting shares describe avoidable expenditure.

Review itemCompareRequired follow-through
Monetary benefitForecast contribution against attributable results.Benefits owner revises the case.
Review workloadExpected checking against actual work and quality.Delivery and operations adjust capacity.
Supplier spendingForecast charges against demand and actual expense.Budget owner changes scope or funding.
Shared supportAllocation basis against continuing obligations.Shared-service owner updates funding agreements.

Variance analysis compares actual results with expectations. State the eligible population and observation period; separate failures, successes, and unresolved outcomes. Missing evidence remains unavailable. Selectively observed outcomes cannot automatically represent all users. Task metrics and usage accounting covers implementation.

Recent cases have had less time to produce downstream outcomes. Compare consistent follow-up horizons rather than treating an absent complaint as success. Each material departure from expectations needs an action, an owner, and a date for reviewing the result.

Stopping the assistant might remove cancellable supplier charges while leaving shared-team salaries unchanged. The assigned salary expense is then reallocated, not saved. Whether released staff time has another valuable use is a separate decision.

Field evidence and revised investment assumptions

Field evidence preserves the operating problem, customer variation, outcomes, and maintenance burden. It helps distinguish local requirements from reusable capabilities. Recurring work can justify shared investment; a single customer's unusual requirement need not become everyone's platform obligation.

An after-action review compares expectations, observations, explanations, and resulting changes. Double-loop learning goes beyond repairing execution: it revises the assumptions and decision rules governing the work. Faster drafting with unchanged resolution can challenge the investment premise rather than indicate a need for still faster drafting.

Possible explanationEvidence that distinguishes itOrganizational response
Failed value hypothesisTarget behavior improves; intended business outcome does not.Reconsider the intervention and success assumptions.
Delivery failureImplementation violates an established requirement.Repair execution and its feedback mechanisms.
Adoption barrierUseful capability is not incorporated effectively into work.Investigate fit, practice, and role expectations.
Measurement gapEligible outcomes remain unobserved or selectively recorded.Improve evidence before assigning a cause.
External changeSupplier behavior or supported features changed.Reassess compatibility and operating assumptions.

Preserve the original decision and why it changed, including unused systems and stopped investments. Decision records support later review; storing them outside an active conversation also makes the rationale recoverable when people or agents lose context.

In the support example, deferring expansion can release reviewers for knowledge cleanup. That changes the portfolio without proving cash savings. The next allocation follows the value of future work, not the effort already spent.

Workforce commitments and organizational change

Adoption means sustained, appropriate use in everyday work. When access fails to produce benefit, investigate workflow fit, role expectations, practical training, review capacity, and decision authority. More accounts or mandatory usage cannot distinguish these explanations.

WorkBeforeDuring transitionIntended continuing arrangement
DraftingStaff compose responses.Staff practice assistance on their own work.Use assistance where it improves the task.
LearningExisting duties consume available time.Protect role-specific practice time.Retain capability as responsibilities change.
Expert contributionSpecialists work within established boundaries.Build collaboration and prototyping skills.Preserve depth while supporting broader delivery.

Leadership must fund the transition. Intercom describes pairing explicit role expectations with dedicated enablement staff, immersion days, and shared learning. Automattic's reported experiment protected time by pausing roadmap work in stages. These accounts describe concrete commitments; their bundled interventions do not isolate AI's effect.

Involve affected workers in changed duties, training, and career expectations, including displaced work. OECD surveys associated consultation with more favorable reported outcomes, but the findings were observational. Participation can reveal workload and implementation problems; it does not itself establish productivity or cash savings.

Make reporting errors safe and useful rather than penalizing every rejection of AI output. Reward useful outcomes and learning, using activity counts diagnostically. Incentives and cross-team consequences explains why usage targets can reward behavior that fails to improve the intended service.

Repeated observations must distinguish temporary learning effort from persistent additional work. For the support assistant, unchanged total effort after sustained practice calls for revisiting fit, checking duties, or the intervention itself—not indefinitely labeling the added workload a transition cost.

The revised mandate can defer expansion, retain necessary checking, assign released time to knowledge cleanup, and fund the remaining support duty. Continue only when the resulting benefit and obligation are credible; otherwise change or stop the investment.

Open questions

  1. The conditions favoring centralized, embedded, or federated teams remain unsettled. Local context and shared expertise create competing demands. Progress would compare similar workloads while recording handoff delay, expert availability, operating cost, and quality—not infer organizational superiority from one successful deployment.

  2. Temporary transition work remains difficult to separate from permanent added burden. Tool capability, staffing, and employee experience change together. Repeated observations of duties, total workload, outcomes, and employment expectations would help establish whether gains persist after intensive support ends.

  3. Practical supplier exit remains harder to establish than interface portability. Behavioral differences, missing expertise, and continuity duties can dominate switching effort. Progress requires a completed transition with accepted coverage, observed replacement behavior, and full costs, rather than an export demonstration alone.

Follow the curated reading path through the speakers and demonstrations behind this entry.

Explore more talks

The rest of the library, beyond the curated path. Cited talks support this entry; reviewed transcripts were processed in full. Metadata candidates have not been reviewed as sources or verified as topic members.

10 matching talks

TalkSpeakerEventYear
Sandipan BhaumikAI Engineer Europe 20262026
Sonny Merla, Mauro Luchetti, Mattia RedaelliAI Engineer Europe 20262026
Rossella Blatt Vital, Deepsha MenghaniAI Engineer World's Fair 20252025
Addy OsmaniAI Engineer World's Fair 20262026
Alex AtallahAI Engineer World's Fair 20252025
Amir HaghighatAI Engineer World's Fair 20252025
Fuzzing in the GenAI Era

Cited in this entry

Leonard TangAI Engineer World's Fair 20252025
Michal CichraAI Engineer Europe 20262026
Martin Harrysson, Natasha ManiarAI Engineer Code 20252025
Sanja GrbicAI Engineer World's Fair 20262026

References

Coverage and source review
Processed transcripts
16 processed in full · 6 in the curated path
Automated source review
Passed
Metadata candidates
0 unreviewed; not verified topic membership
Corpus version
1bd8e407b26a07b33815594e1b2db5f41827119a2b3cb6fbf240f9fc571fc767

Automated review checks source support; it is not publication approval.

A synthesis of selected conference talks and technical references. Citations link to the source material; they do not imply that every talk on this subject is included.

  1. HM Treasury: The Green Book (2026)

    Appraisal framework, sunk costs, uncertainty, sensitivity analysis and evaluation sections. Supports the investment mandate and staged-commitment reasoning.

  2. Guidelines for Managing Projects: How to organise, plan and control projects

    Scope definition, stakeholder analysis, change control, and benefits realization sections; supports a lightweight customer delivery agreement.

  3. The Production AI Playbook: Deploying Agents at Enterprise Scale

    Define business success and build a representative evaluation dataset before comparing models; reuse that dataset to assess provider upgrades.

  4. How Forward Deployed Engineering is done at Factory

    Factory positions deployed engineers as a customer-to-product feedback channel rather than a team that executes consulting projects for customers.

  5. APM: What is portfolio management?

    Portfolio definition, shaping, planning and risk sections. Supports explaining why individually attractive initiatives may collectively exceed available capacity.

  6. ACCA: Relevant costs

    Relevant-cost principles and labor example. Supports continue/stop economics and distinguishing cost allocation from avoidable spending.

  7. Most Enterprise Agentic Projects Are Doomed — Here’s Why

    The speakers recommend that finance 'think like a VC': fund a portfolio of AI experiments rather than require every project to promise a fixed solution and predictable payback upfront.

  8. UK Government: Data Ownership Model

    Principles, role responsibilities and Appendix B. Supplies concrete vocabulary and an organizational allocation example.

  9. CNCF Platforms White Paper

    Brief distinction between field implementation and shared platform investment; contextual link to /topics/ai-platform-engineering.

  10. HM Treasury and Government Finance Function: The Government Efficiency Framework

    Sections 4.1–4.10 and 6.1–6.4; supports benefits realization, net value and cross-functional cost accounting.

  11. fun stories from building OpenRouter and where all this is going

    Providers hosting models can differ in supported features, prices, and performance, requiring provider-level compatibility handling.

  12. FinOps terminology: ownership, depreciation and utilization

    Capitalization; Depreciation; Fixed Cost; Cost Allocation; Total Cost of Ownership; Activity Based Costing; Shared cost.

  13. ROI Institute: Introduction to the ROI Methodology

    Steps 6, 9 and 10. Supports the chapter's explicit ROI convention and distinction between observed improvement and attributable benefit.

  14. ROI Institute: ROI Basics

    Original authors' October 2007 explanatory article, page 2. Supports monetary contribution and declared benefit-horizon vocabulary.

  15. HM Treasury: The Orange Book, October 2004

    Sections 3.4 and 4.4–4.6. Used solely for durable first-use vocabulary and the distinction between ownership and execution.

  16. Your Fine-Tuned Model Is Tech Debt: A 50x ROI House of Cards

    The speaker reports that the rebuilt system cost more per message but less overall because maintenance effort fell.

  17. Most Enterprise Agentic Projects Are Doomed — Here’s Why

    Use hypothesis-driven delivery organized around increasing statistical confidence through short build, evaluate, and iterate loops.

  18. Department for Work and Pensions: Accounting officer system statement, 2025

    Sections 4.3 and 4.7–4.15. Concrete institutional example of delegated authority, risk appetite and escalation.

  19. The Build-Operate Divide: Bridging Product Vision and AI Operational Reality

    An AI quality lead combines customer and domain understanding with systematic quality diagnosis, regardless of whether they write production code.

  20. Intercom: AI is approving our pull requests

    April 2026 first-person implementation account. Concrete example of changing operating practices while retaining accountability and escalation.

  21. Operating Model Canvas

    Original framework authors’ explanation; supports a plain-language operating-model definition and a before-and-after organizational map.

  22. Government Digital Service: Artificial Intelligence Playbook for the UK Government

    Building the team and Creating the AI support structure. Supports capability coverage and concrete leadership resource commitments.

  23. AWS: Generative AI operating models in enterprise organizations with Amazon Bedrock

    Centralized, decentralized and federated operating-model sections and concluding ownership discussion. LOB means line of business: an organizational business unit.

  24. Cabinet Office: The Sourcing Playbook

    Chapters 3, 4 and 13. Supports capability-level sourcing, bounded experiments and resourced exit obligations.

  25. The Rise of Open Models in the Enterprise

    An open model, inference engine, and GPUs do not by themselves constitute production inference.

  26. Forward Deployed Engineering 101

    Selling an implemented business outcome can address the adoption burden that remains when customers buy a platform but lack the skills to build with it.

  27. Forward Deployed Engineering 101

    The speaker defines FDE as an enterprise-scale design partnership built on shared platform primitives, rather than independent software projects for every customer.

  28. Forward Deployed Engineering 101

    Keep genuinely customer-specific behavior local, and move generalizable capabilities into the platform over time.

  29. Your Fine-Tuned Model Is Tech Debt: A 50x ROI House of Cards

    The speaker's 'calcification tax' describes maintenance complexity that made both model switching and architectural change too costly.

  30. The Rise of Open Models in the Enterprise

    The speaker reports that switching frontier providers is feasible, but still requires renewed evaluations and prompt tuning.

  31. How Forward Deployed Engineering is done at Factory

    The speaker recommends model independence, access to execution data and traces, and centralized information-flow governance.

  32. Fuzzing in the GenAI Era

    Static examples can miss brittleness: small changes to an input can cause large changes in application behavior.

  33. The Build-Operate Divide: Bridging Product Vision and AI Operational Reality

    Prioritize human-in-the-loop review at consequential decision points in high-risk, high-trust workflows.

  34. FinOps Foundation: Forecasting

    Definition and forecasting-strategy sections. Supports named budget ownership and revising expectations when systems or plans change.

  35. NIST AI Risk Management Framework: Core

    GOVERN 2 and 6; MANAGE 1–4. Supports accountability, continuing authorization and distinction between assessment evidence and a decision to proceed.

  36. Most Enterprise Agentic Projects Are Doomed — Here’s Why

    Use progressive autonomy through an exposure ladder: shadow mode, advisory mode, controlled autonomy, then wider autonomy gated by outcome evidence.

  37. APM Body of Knowledge, Seventh Edition: Transition into Use

    Sections 2.3.1–2.3.3, printed pages 88–92. Supports readiness, receiving-team acceptance, support and transferred-work accounting.

  38. BDD, ADR, PRD, WTF: Capturing Decisions for Humans and AI Alike — Michal Cichra, Safe Intelligence

    Use Architecture Decision Records (ADRs) to preserve both the reason for a rule and its enforcement, then make enforcement failures point back to those records.

  39. FinOps Foundation: Allocation

    Shared-cost strategy discussion. Supports transparent funding of common platforms and support obligations.

  40. Moving away from Agile: What's Next?

    Use the talk's proposed MECE measurement framework to connect inputs, operational outputs, developer experience, quality, and economic outcomes.

  41. National Academies: inference with missing outcomes

    Chapter 4: missing-data mechanisms; complete-case analysis; weighting; sensitivity to assumptions.

  42. National Academies: informative censoring and sensitivity analysis

    Chapter 5, TIME-TO-EVENT DATA; assumptions about informative censoring and sensitivity parameters.

  43. Chris Argyris: Double-Loop Learning, Teaching, and Research

    Opening definition in the reproduced 2002 Academy of Management Learning & Education article, pages 206–218. Supports first-use vocabulary for revising investment assumptions and decision rules.

  44. BDD, ADR, PRD, WTF: Capturing Decisions for Humans and AI Alike — Michal Cichra, Safe Intelligence

    Keep the work-feedback-repair loop stable, but use task-specific skills to select relevant documents, checks, and review artifacts.

  45. Moving away from Agile: What's Next?

    Pair rollout with explicit role expectations and hands-on coaching using employees' own work, especially during the first few sprints.

  46. BDD, ADR, PRD, WTF: Capturing Decisions for Humans and AI Alike — Michal Cichra, Safe Intelligence

    Retrieving decision documents can consume substantial context, but the speaker reports that rediscoverable rules make repeated context compaction tolerable.

  47. 500 people vibe-coded for 30 days. I was one of them.

    Radical Speed Month combined protected experimentation time, small-team autonomy, and a requirement to ship something real.

  48. 500 people vibe-coded for 30 days. I was one of them.

    Tool access was complemented by role-specific training, hands-on practice, and documented development and security processes.

  49. From Hype to Habit: How We’re Building an AI-First SaaS Company—While Still Shipping the Roadmap

    Invest in T-shaped talent that retains deep expertise while adding prototyping, cross-team collaboration, and end-to-end building skills.

  50. How Building with AI Can Double the Throughput of Your Engineering Team

    Pair explicit role expectations with full-time enablement staff and opportunities for employees to learn together.

  51. OECD: The impact of AI on the workplace

    Chapters 4, 5 and 7, especially pages 77–83. Supports worker participation and explicit discussion of changed duties and employment expectations.

  52. KCS v6 Practices Guide: Summary

    Lessons Learned section; supports an organizational alternative to rewarding article production or tool activity alone.

  53. How Forward Deployed Engineering is done at Factory

    Define an outcome or ROI story at the beginning that connects changes in engineering behavior to core business goals.

  54. Brynjolfsson, Li and Raymond: Generative AI at Work

    Sections 2.2–4.1 and introductory findings in the opened revision. Concrete evidence connecting knowledge access, decision rights, onboarding and measured service work.