AI Engineer World's Fair 2026
Every AI Company Is Accidentally Building a Bank — Dor Sasson, Stigg
Read the talk
Every AI Company Is Accidentally Building a Bank
Dor Sasson explains why AI spending needs a decision before inference, a reservation while work runs, and settlement afterward—and why shared credits, autonomous agents and enterprise budgets make this an infrastructure problem.
From a talk by Dor Sasson
At a glance
Ideas worth remembering
Check entitlement synchronously before inference; reconcile and settle actual usage asynchronously afterward.
Concurrent consumers need coordinated reservations, because several requests can otherwise authorize spending against the same balance.
Preserve credit sources and consumption ownership: a single total cannot express drawdown priority or team, user and agent spending controls.
Recheck and reserve as agents take on more work, because the full cost may be unknown when the task begins.
Consumption behavior can change effective spending even when the per-unit price stays fixed, making pricing and runtime infrastructure closely connected.
A subscription can hide an uncontrolled compute bill
In April, a series of AI product changes looked like pricing emergencies: Anthropic restricted third-party agents such as OpenClaw from using subscription plans, OpenAI changed its Pro offering, and GitHub restricted free Copilot access. Dor Sasson, co-founder and CEO of Stigg, uses these episodes to introduce his claim that every AI company is accidentally building a bank. Selling access to software now means deciding who can spend a shared resource, how much they can spend, and when to stop them.
The OpenClaw example makes the mismatch concrete. Under Claude Max, a subscription payment covered agent-driven consumption that could impose much greater costs on Anthropic. In Sasson’s account, the system could not effectively distinguish consumption appropriate to a subscription from consumption that belonged on the API, and the response was to stop access. These incident descriptions and the diagnosis of the companies’ internal controls are Sasson’s account; they do not establish the precise costs or implementation failures inside those companies.
The same exposure can appear inside a customer organization. Sasson cites Uber exhausting its annual AI consumption budget within weeks and a Replit example in which three users could consume an organization’s entire credit pool. The important relationship is between a large shared allowance and a small number of consumers capable of spending it quickly. A contract approved for a year does not, by itself, make the resource last a year.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Move permission checks ahead of inference
The architectural failure is a timing problem. A user or agent receives access, burns tokens, and generates usage records. Only afterward does the system reconcile the spending against what that consumer was entitled to use. By then, the compute has already been consumed. An invoice can describe the loss, but it cannot prevent it.
That ordering creates problems on both sides of the sale. Customers face sticker shock or unexpectedly exhausted budgets. Vendors face overages, compute spikes and potentially lower margins when consumption exceeds what they planned to supply. Usage logging remains necessary, but a record of completed work cannot serve as permission for work that has already happened.
The proposed ordering checks access synchronously on the request’s hot path, before inference starts. Reconciliation and settlement happen asynchronously afterward. The ATM analogy captures why: permission to withdraw cash must be checked before the machine dispenses it. AI inference has the same irreversible-spending problem, even though final usage accounting can finish later.
Where does the decision have to move to prevent spending? The comparison below separates the permission decision from the later accounting step. Its useful distinction is temporal: settlement can follow inference, while authorization has to precede it.
Compute is consumed before entitlement is checked.
Checking completed usage discovers an overage. Checking the request first can prevent unauthorized consumption.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
One request needs several policy decisions
Sasson points to an architecture published by OpenAI’s financial engineering team as an example of this request-time approach. He describes a decision waterfall that considers the client, user or agent and evaluates several conditions before granting access. A balance alone cannot answer whether a particular request is allowed.
The conditions address different parts of the decision:
- Feature access: Is this product or capability included for the consumer?
- Rate limits: Does the request fit the limits of the consumer’s plan?
- Trial and promotional status: Do special program rules apply to this consumption?
Together, these conditions determine whether the user or agent can draw down resources now.
Two properties make the waterfall useful: it runs as a single synchronous evaluation at the request, and its rule priority is deterministic. The system knows which rules take precedence before it encounters the request. Sasson treats this as a financial decision system that operates ahead of billing: an invoice records a commercial outcome, while this evaluation controls whether the outcome may occur.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A balance check alone cannot stop double spending
Consider Sasson’s account with ten dollars and two parties entitled to withdraw from it. If both inspect the account at the same time, both can see the same ten dollars. Each request looks affordable in isolation. The failure appears when both act on that observation: one balance has supported two spending decisions.
A hold changes what the next consumer can spend. Once one withdrawal reserves the funds, those funds are committed to that work rather than freely available to another withdrawal. Settlement follows when the actual consumption is known. Applied to AI, the reservation must affect subsequent access decisions before other requests can spend the same pool. With ten thousand agents reaching for a shared balance, checking early helps only if concurrent checks and reservations cannot all authorize the same money.
Sasson groups hold-and-settle, double-entry accounting, idempotency and auditability among the banking properties AI products need. They address related concerns: reserving spend, recording its movement, avoiding duplicate effects from repeated requests, and retaining a history that can be inspected. Double-entry records alone do not specify how competing reservations are coordinated; the talk identifies the required properties without developing a transaction or locking implementation.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Credits have sources, and spending has owners
Even a correctly coordinated balance can discard information the product needs. Sasson’s example is a pool of 1,700 credits assembled from different sources. Some credits might have been granted earlier; others might be promotional. The total answers how much exists, but it cannot answer which credits should be consumed first. That requires preserving the sources and defining a drawdown order.
This is why storing credits as a single integer becomes limiting. Once sources have different treatment, consumption is also an allocation decision. An oldest-first policy and a promotional-first policy can produce different remaining pools after the same amount of work. The talk presents those as choices to resolve, rather than prescribing one universal ordering.
The ownership model also grows beyond a flat organization-to-user relationship. An enterprise may approve a seven-digit annual pre-commitment, then distribute its consumption across teams, users and agents. Each level can require allocations, budgets and spend caps. The person approving the contract needs to see both how much remains and who is consuming it over the contract’s life.
The two dimensions belong together: credit provenance determines which funds a request consumes, while organizational attribution determines whose activity consumes them. A top-level balance cannot supply either form of fine-grained control on its own.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Four patterns for selling AI at scale
The architecture comes together in four recurring patterns. They cover the timing of spending, the location of accounting data, the uncertainty of agent work, and the customer’s ability to understand consumption.
- Reserve before inference; settle actuals afterward. Permission and reservation precede costly work. Accounting then reconciles the reservation with actual usage. This connects request handling to the concurrency problem: simultaneous workloads must not each claim the same available credits.
- Keep metering and ledgers inside the VPC when needed. Sending usage events to an external service raises latency, cost and data-sovereignty concerns. Sasson expects more companies to place these systems within their own virtual private cloud. The placement choice responds to those concerns; it does not remove the need for metering or settlement.
- Recheck balances as agents continue working. An agent’s eventual actions and total cost may be unknown at the start. As it takes on further work or spawns more activity, the system needs additional checks and reservations, followed by asynchronous reconciliation.
- Make consumption visible across the business. CFOs and CIOs need views across models, features and products, alongside attribution to teams, users and agents. Sasson expects this visibility to become a prerequisite for enterprise purchases.
Agentic checks add an important limit to the initial ATM analogy. A cash withdrawal has a known requested amount; an agent task may keep discovering more work. Admission at the beginning cannot authorize an unknown amount of future spending. The check-reserve-settle pattern therefore has to recur as the task progresses.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Consumption rate can change the price customers experience
Near the end, Sasson describes an April change in which the per-unit cost stayed the same while token consumption increased. His spoken account names OpenAI, while the displayed slide identifies Anthropic and a Claude tokenizer change. That provider attribution is inconsistent across the presentation; the general mechanism is the supported point: the unit price alone does not describe what customers spend when the same input consumes more tokens.
The mechanism is straightforward: at the same cost per token, consuming more tokens increases total cost. Executing such a change therefore involves the infrastructure that governs consumption, rather than only editing a price list. Rate of use, credit drawdown and the customer’s budget are connected parts of the product’s economics.
Sasson closes by comparing the opportunity to the early-2010s “Stripe moment,” when reusable payment infrastructure made checkout easier for developers shipping internet products. His proposed equivalent for AI would make reservations, settlement, ledgers and spending controls easy to adopt. Coming from Stigg’s co-founder, this is also a product thesis: teams should be able to keep shipping AI without independently reconstructing these financial mechanisms.
The practical warning is to recognize the bank-like responsibilities while the system is still being built. Shared pools, concurrent consumers and uncertain agent costs become harder to correct after revenue and usage scale. You may not have set out to build a bank. Once your product grants access to spendable resources, its spending rules are part of its architecture.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:12
Yeah, we're good. Okay. So, hi, everyone. Uh, given there's-- I hear there's a match right now, uh, as we speak, so given that you are here means a lot to me. Uh, so at least I'm, I'm winning the World Cup right now. So that's, that's a, that's a big thing. Um, thank you for joining me. I flew all the way over. I'm the co-founder and CEO of Stigg. I flew all the way over to make a, kind of a funky statement on an AI conference, which is literally every AI company right now is accidentally building a bank. Um, and when I
- 0:42
mean bank, I'm gonna try to kinda unpack this thing, uh, not as like a weird marketing metaphor, but effectively what ha- what is really happening right now with the AI economy and why a lot of the constructs and the things that we are seeing, using, consuming, and paying for in AI are actually, uh, really behaving like, uh, constructs that we know from the banking financial systems. Um, so cool. Thank you for, for, for having me. Um, and let's, uh, get it going.
- 1:13
By the way, I'm the first talk today, so the clicker is not working. It's gonna be interesting, uh, here now as I move through this. So, um, let me take you through a little bit of background. Uh, as I was working through this, uh, talk, uh, back in April, really every single, uh, code-generating, uh, product platform out there just broke under this new type of economy and under this new type of c- uh, consumer behavior. In, in just a, you know,
- 1:43
matter of, like, five weeks, everything felt like pricing emergency one after the other. And the case I wanna make here is that these pricing emergencies go way beyond just emergencies from a financial commercial side of the house. These really are infrastructure c- uh, emergencies that are emerging because of how those systems were built to begin with. So just to give you some perspective, we had in, in April, uh, we had Anthropic, uh, basically eliminating access for OpenClaw, uh, and other third-party
- 2:13
agents to basically use the subscriptions plans. Then immediately we had OpenAI basically changing how the pro tier is being priced, bumping the price overnight, five, five X, which is quite intense. And then, then we had G- GitHub freezing and then completely eliminating access, trial access, free access to their-- Whoa, we have some echo here. Um, to their, uh, free, uh, plans for, for, for Copilot. And effectively, like all these companies all at once in
- 2:43
a single month experienced what it means to actually sell on scale, uh, uh, AI, and what actually happens when you don't have the fundamentals in place. Um, so let me take one, one step farther. So let's zoom in to even just one of the use cases, right? So, um, if you think about Anthropic and you think about the OpenClaw use case, right? What happened was effectively Anthropic was subsidizing every consumption of every user using OpenClaw
- 3:13
within the CLO- Claude Max subscription. So effectively, what the customers were paying daily was, you know, dollars and cents. Um, but on the, uh, on the cost side for Anthropic, it was like hundred and fiftys to even seven hundred and fifty subsidized, which is effectively not an economic that you can scale, even for the company at, at the size, at the scale of Anthropic, right? So what it seemed to be a business problem really was an infrastructure problem because Anthropic didn't have a,
- 3:43
a- an effective way to delineate between users who are using their APIs and users who are using their subscription, um, uh, programs. And so they had to immediately seize and stop access, right? They had to, to, to, to react. Um, cool. So...
- 4:02
And this is not unique to Anthropic, guys. Like, uh, we're seeing this, like, I don't know if you all heard-- read the news most recently with a company, uh, undisclosed company burning through like half a billion of their AI, uh, cloud credits, um, just merely having like a few employees burning through the entire organization contract. Uh, we had, uh, Uber saying they're basically torched down their entire, uh, AI usage consumption budget for the year in just a few
- 4:31
weeks into the year. And then we had Rep- Replit giving an example of what could happen if, uh, a single org has three users that are basically burning down the entire credit pool for Replit for the entire org. So really what happening here is, uh, this is not just an, an, an Anthropic problem. And I think what's, what's mutual and kinda like switching gears from the examples to the actual what's really happening here, um, the problem basically with, with all these three
- 5:00
examples is that the checks for what you're entitled, what you're allowed to do, whether you're an agent or a user with your AI product, happened, uh, after, uh, the invoice and not before. So basically, with each and every one of those examples, you were allowed to consume, you were allowed to use the product, you were allowed to burn tokens, and checks happen only after the invoice. So basically, the spend and the after effect gets reconciled only after it's too late, only after the invoice
- 5:30
is, is sent, and this is effectively bad for business. It's sticker shock. It's bad for practice of how we use the systems. These systems are now mission critical, and we're basically none of these companies had in place the financial infrastructure that allows, exactly like banks, guys, to check before we draw down, before we consume, right? So I think what's, what we're seeing is effectively- Uh, really is a change in how we architect, and how we ship, and how
- 6:00
we build software. AI is not just changing what is value, it's not just changing how we interact, who is the users, it also effectively changing what exactly is getting paid and how we actually architect for systems that can, uh, handle runtime, handle the complexity of usage happening before the, the, the invoice, before the checks, and how we actually can architect for that. So I think what's interesting, and what you see in front of you, um, is
- 6:30
basically on the left side, what we're seeing is how most sha- teams ship this architecture today, and you can see that basically the check and the settlements are happening after, a- after the inference, after the usage is being logged. And that's typically too late. We see situations like spikes, overspends, overages, like experiences that are bad for business, bad for users, and quite frankly, in many cases, also bad for the vendor. Because if you didn't allocate enough resources and compute to
- 7:00
actually allow for this to happen, you al- you also have a margin problem potentially, right? On the right side, we're basically saying, or I'm, I'm trying to make the case for you all that with AI, with the current infrastructure and how AI is being built and sold, we actually need to enforce and, and do this synchronously before we actually update the balances, and reconcile and settle in, in async way after, after effect. So hot path actually get checked
- 7:30
synchronously, and everything that comes after, uh, is basically reconciled after. I think the best example, guys, is if... Imagine you go to the ATM, right? And you draw cash. When you draw the cash, it do- the check if you are allowed to draw it doesn't happen after you get the cash. It happens before you draw the cash. And so that's exactly the same thing. Our industry effectively misses this layer, this financial infrastructure that allows for this type of behavior to happen, uh, na- naturally, right?
- 8:00
Um, and so obviously this sounds like a great idea, Dor, right? Like you're, you're pitching this infrastructure, sounds like you have something to do with this. Um, but guys, this is not just me. Uh, in February, OpenAI, they have a quite a, a great team that their team call financial engineering, and these guys basically, uh, uh, uh, published how they think about the architecture and what they had to build for this to work at the scale of OpenAI. And what they had to say reads very
- 8:30
similarly to how, uh, transactional systems of banks work. Basically, you have a decision waterfall that needs to consider in runtime what is the client, the user, the agent is allowed to do, and it needs to take into account many different conditions and policies of what is really allowed, right? So are you, are allowed or you have access to a certain feature or product, um, what are the certain rate limitations under your plan, um, whether you are in a
- 8:59
trial or in a certain promotional program. Like you need all these decisions to be calculated immediately at runtime, and then ultimately decide, "Can I give this user access? Can they draw down? Can they consume or not?" And all those decisions really can happen after the invoice. They can, because you need to actually be able to, uh, calculate and prioritize this. So I think we have three things to notice from here, right? The first thing is this is a single synchronous evaluation. It happens
- 9:29
at the request. It doesn't happen after effect. I think the second thing is the priority is deterministic. So we know in advance what are all the different rules that apply for certain things to be consumed, right? And I think lastly, and here, guys, this is my opinionated view, you can disagree, I don't think this is billing. I don't think this is necessarily, uh, this entire architecture is just related to the idea of invoicing and billing. This is a financial system that actually make decisions ahead
- 9:59
of any invoice, ahead of any billing element, and I think we're gonna see more and more of this becoming a standard, uh, for other products and other AI systems as well. Um, really just nerding just a little bit on... I'm sorry, I just skipped one s- slide. So yeah, just, just nerding just a little bit about, um, what other traits or s- or elements we see and are used to in banks, and we take them for granted. But in AI, they're not that granted. In
- 10:29
software, they're not that granted. So just one example, right? Like if two separate parties have access to the same banking account, and that banking account has like ten dollars, and they're entitled to draw down those ten dollars, what happens and when they both of them try to draw down the same time, right? So banking systems has the idea of holding, of settling. You can really... Like concurrency is effectively solved with banks. But with AI and agents, it's not that case. You could have,
- 11:00
whatever, like ten thousand agents trying to reach out to the same dollar pool, and they're gonna try to do that a- at once. And superficially, it might look like they should be able to draw down because there is balance in the pool. But what happens when they all draw down the sa- from the same pool simultaneously? Concurrency become an issue. So how do you solve for that? Um, this entire idea in financial systems actually called, uh, like it's a classic double, double spend problem. Uh, it comes from the accounting world. It's a double entry bookkeeping.
- 11:30
And as I'm sure you all are building shipping today with AI, as the revenue scales, as your company scale, the idea of double entry accounting, double entry billing, the idea that you can solve concurrency, these things would matter, uh, as you scale, and they're not so easily solved down the line. So think about things like, um, you know, can you basically introduce traits like hold and settle? Think about idempotency and auditability of those, um,
- 12:00
uh, uh, request and et cetera. Um-
- 12:06
Another thing that I think is very, very, uh, known in, in, you know, how banking systems works, but not s- not necessarily enough established in AI, is the idea that, uh, all the different pools have different sources, right? So in bank, you have your debit, uh, account, you have your cash account, you have your savings, right? And if you draw down, the bank actually knows how to calculate where are you drawing down from and what does it mean. Here, y- you know, say I have a pool of, you know, one, you know, one
- 12:36
thousand seven hundred credits, uh, but they're not the same. They're not alike. Each one of them is sourced from a different, uh, source. So how do you decide which, which one of them do you draw down against? Do you start with the ones that were, you know, given granted earlier? Do you start with the ones that were promotional? So how do you really calculate what's getting drawn down to the users? I think that's, that's a real problem. Most companies today try to solve it like a single integer, uh, databases, and that's not really going to work as you think
- 13:06
forward. Um, cool. Um, so really tying the knot, the, the knot here, I think, um, if you think about how we interact with software, it used to be flat. So it used to be just users and organizations, and there wasn't a lot of meat in between them, right? The, the, the, the hierarchy of how you consume software was quite flat, and it was quite straightforward in terms of who see the value. But today, I think effectively credit pools are no longer flat. Uh,
- 13:36
you have a lot of different hierarchies in between who consumes different credits, and also they demand different governance, different logics, allocations, budgets, spend caps. And effectively what happens is that budgets are set at the top where, for instance, if you're selling to the enterprise, you're maybe selling like a pre-commit annual contract. Uh, C-suite or somebody who approved that s- seven digits contract, they expect to see how that credits are getting drawn down along the contract life. But they don't just
- 14:06
wanna see that they were used. They wanna know who actually used them, which team, which users, w- uh, which agents, uh, to what extent, and they wanna have fine grain control over that. That's literally the bare bones of how more and more companies thinking about selling AI workloads and how they, they effectively work in production. I'm gonna run through some of this. Um, I think this idea of like, what does it mean, like high cardinality graphs and dimensions, being able to
- 14:36
slice and dice against model types, users, like all this type of, uh, dimensionality of, uh, of, uh, consumption is, is probably interesting too to talk about. But what I actually wanna expand on is ultimately those four patterns. So as you build AI and start selling AI at scale, you're effectively like four patterns that keep repeating themself, and they're very, very much alike and, and very similar to the idea of, you know, banking systems. I think the first one we talked about earlier is basically being able to reserve before
- 15:06
inference and settle actuals after. Um, this is like going to s- to be like a pattern that again and again we're seeing more and more often, and it has ramifications on how you do accounting, how you do, uh, basically handle all those, uh, different concurrencies workloads, and what have you. I think the second thing is we're seeing more and more AI companies unexcited about the idea of sending all this usage data to, to somewhere else over the cloud. Uh, for latency reasons, for cost reasons, for, uh, data sovereignty
- 15:36
reasons, uh, there's a lot of good reasons for why these companies actually rather to keep the events data inside their VPC, and I think we're gonna see more and more AI companies seriously considering deploying this, this architecture, this ledger, this metering within their VPC rather the, the contrary of sending it over the internet. Um, I think for agentic checks, so think about how agents spawn, uh, the idea of checking balances. When agent does something, you don't know in advance what they're ultimately going to do,
- 16:06
how that a- how that task is going to basically cost. So if you don't know the cost, you need to be able to check those balances as the agent continues to get work, and you need to be able to, uh, reserve and account for them in advance and then asynchronically after effect reconcile. Um,
- 16:25
quickly, uh, moving forward. So ultimately, if you have those systems in place, what you really gain out of it is visibility. So it used to be given that for usage-based products you need visibility, but I think what happens today, you know, it goes beyond just plain visibility. Um, it, it means that effectively, like the ability to see for a CFO, for a CIO, consumption of AI workloads across different models, across different features, acro- across different products is not just, you know, a
- 16:55
nice to have, it becoming a table stakes. It becomes something that you can't do business without, and we're gonna see more and more companies expect to have this off the bat before they actually do business with you. Um, and so really just giving some examples to how some of the best frontier AI companies are approaching some of this. I think you're, you're seeing the patterns already there. So I think a lot of these companies already appreciate that you're no longer just selling software in the c- you know, commodity
- 17:25
way that we used to think about s- subscriptions and usage. It's becoming more and more of a financial systems, and financial systems require some of these complex ideas that weren't, were not there before AI. Um, and I think just to give like one small example that I talked about earlier, but I think might have gone a little bit, uh, hidden, um, y- One of the changes in April that OpenAI did is lier- really, they didn't change the pricing. All they did was change the pace
- 17:55
in which tokens get burned. So effectively, we use the same model, we pay the same cost, uh, per unit, but we burn more tokens. So that effectively was a pricing change. How do you execute such a change? Um, because it's not just financial change, it's an infrastructure change. You basically change the rate of how the software gets consumed. So there's a lot of complexity, not in terms of just how we change pricing, but also how the infrastructure underneath allows for more complex ideas into how value is getting, uh, uh, introduced to the users. And
- 18:25
really just, you know, before wrapping up, you know, there-- our team likes to say, and this is something we started to say more and more often than before, it really feels like, you know, there was in the, in the early 2010s, there was this Stripe moment where every company wanted to ship product into the internet, and you needed a checkout, and you needed to be secure, and it... And developers need to, to, to, to make it easy for them to quickly get them up and running. I feel like where we are today is those ideas of, like, building a banks, it's okay if you feel
- 18:55
like, I don't know, you like wearing suit and tie, maybe you wanna build a bank. But if you are building a bank, you probably want to be aware to all of these traits and elements that has to be there, so you don't meet them down the line where, where it gets really, really more complex to fix in a hindsight. And so feels like the industry really misses its Stripe moment, uh, where all these constructs just exist, and they're so easy to use. And, you know, there's a lot of excitement from our team when it comes to the potential of having such systems ready
- 19:25
and help AI companies keep shipping. Um, if any of that was interesting, really come by our booth. We are in a quite of a weird location, right at the start, but still there. Uh, I know there's a game right now. Come visit. We're demoing the product. Feel free to ask any questions. And thank you for coming today. It was awesome to, to have you, have you with me. And yeah, uh, love to nerd some more, uh, by our booth.