AI Engineer World's Fair 2026
Small Models, Big Results: Training a Finance Agent for Under $500 — Charles Dickens, Snorkel AI
Read the talk
Small Models, Big Results: Training a Finance Agent for Under $500
Charles Dickens explains how verified financial data and reinforcement learning helped a 4-billion-parameter agent outperform its 235-billion-parameter sibling—and why disciplined tool use mattered more than harder training examples or elaborate rewards.
From a talk by Charles Dickens
At a glance
Ideas worth remembering
On 290 held-out financial questions, the trained 4B model reached almost 60% accuracy versus 51% for its 235B sibling. The advantage is specific to this simulated financial task.
Schema hallucination, context flooding, and repeated failed strategies make tool discipline a concrete training target.
The reported training run cost about $420 for compute and $40 for judging. A binary final-answer reward outperformed the more elaborate reward designs tested.
Single-table training produced the greatest lift among the tested data mixtures and transferred to questions involving two to five tables, while general tool-calling performance held up.
Enterprise evaluations should match the intended environment complexity, autonomy horizon, and output complexity.
The financial task behind the model comparison
A financial agent can know plenty about finance and still fail at the job. It must work with the tables and APIs actually available, finish a workflow without compounding mistakes, and leave behavior that someone can audit. Charles Dickens, a research scientist at Snorkel AI, frames those requirements as the practical meaning of the question, “How well does the model or the agent reason?”
Snorkel and UC Berkeley’s Sky Computing Lab tested whether specialization could meet those requirements more effectively than parameter scale. Their trained 4-billion-parameter Qwen model reached almost 60% on financial questions, compared with 51% for its 235-billion-parameter variant. This is a result in a simulated financial environment, rather than evidence that the smaller model is generally more capable.
The project sits within Snorkel’s open research and benchmarking work, alongside its work on frontier-model data and enterprise deployments. Its starting judgment is deliberately practical: a specialist that knows the forms, tools, and rules may be the right choice for a narrow workflow. As Dickens puts it, you would not call Terence Tao for a tax audit. The experiment asks whether training can give a small model that kind of operational competence.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Turn annual filings into questions with checkable answers
Snorkel’s FinQA data begins with SEC 10-K annual reports pulled from EDGAR. These filings supply the financial material analysts use to identify risks and inform investment decisions. A model converts that material into approximately 6,900 SQL tables; each table then supplies one generated question-answer pair.
Generation follows a question taxonomy developed with financial experts who work with these documents. That taxonomy gives the model context about the kinds of questions to produce. The generated examples also carry metadata, including lineage, table names, and columns. Those details make it possible to check whether a question refers to the data that was actually extracted.
Verification has three distinct jobs:
- Programmatic consistency checks: Use the metadata to catch invented table or column references and check the example’s connection to its source data.
- Independent agent reviews: Add automated review beyond the original generation step.
- Expert manual review: Have humans assess whether the questions and answers are realistic and verified.
The resulting tasks require planning, tool calling, and reasoning, but finish with a single verifiable answer. That choice matters later: training can reward a final outcome without having to assign a separate score to every intermediate action. The reported splits contain 4,000 training examples, 500 validation examples, and 290 held-out benchmark examples. Companies in the benchmark do not overlap with the other splits, so evaluation tests work on different companies rather than another question about a company already represented in training.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A plausible query can still derail the workflow
The benchmark exposed recurring failures even in frontier models. These failures concern how an agent uses its environment:
- Schema hallucination: Assume a table or column exists, potentially carrying a familiar schema from pretraining into an unfamiliar database.
- Context flooding: Issue poorly planned queries that return more information than the agent can use within its context limits.
- Poor recovery: Repeat a failed strategy instead of using the error message to change the next attempt. Dickens reports this behavior in the 235-billion-parameter model.
Consider the talk’s concrete example: SELECT *. The query asks for all columns, so a tool response can fill the agent’s context with material that was never selected for relevance to the question. The causal problem starts before the final calculation: a broad retrieval produces an oversized observation, and that observation consumes the space needed to continue the task. Better query planning would control what comes back. The training results later show an overall improvement in answering questions; they do not isolate a measured before-and-after change for this particular query behavior.
Error recovery creates another opportunity to learn from the environment. An error message can tell the agent that its chosen operation will not work; repeating the same strategy wastes that information. Dickens also reports seeing these kinds of failures in insurance underwriting. That recurrence motivates training for reliable tool use rather than treating every failure as a shortage of financial knowledge.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Reward the answer produced by the whole tool-use loop
The team used rLLM, developed by its Sky Computing Lab collaborators, to train the agent. Dickens describes a CLI-first framework that connects to existing agent frameworks through a decorator pattern with very little code change. It supports multiple reinforcement-learning algorithms and training backends; this experiment uses GRPO.
Inside the training environment, the agent runs a ReAct loop: it reasons about what to do, calls a tool to interact with the generated tables, receives the result, and continues toward an answer. The environment contains roughly 7,000 tables. A language-model judge compares the final answer with a reference and supplies a binary correctness reward. The reward therefore evaluates the result of a trajectory, rather than separately paying the agent for accessing a table or completing a query.
Where does feedback enter this system? The diagram separates two kinds of feedback: tool responses guide the next action within an attempt, while the reference-based correctness reward guides training across attempts. That distinction explains how a simple final score can train a workflow containing several decisions.
The base model is Qwen3 4B. Roughly 1,000 concurrent environments generated training trajectories, with training run on eight H100s. The reported cost was about $420 for compute and $40 for the judge API, keeping the run below $500. Those figures describe the training compute and judging costs; they do not establish an all-in cost for building and reviewing the financial dataset.
On the 290 expert-curated held-out examples, the trained model’s accuracy more than doubled relative to its base model. Its almost-60% result exceeded the 235-billion-parameter model’s 51%. The observable change is substantial, but the remaining errors matter too: the successful specialist still answers only about six in ten questions correctly in this evaluation.
Runs a ReAct loop to choose actions and produce an answer.
The agent interacts with tables through tools. Final-answer correctness supplies the binary reward used by GRPO.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Single-table training helped with multi-table questions
The first benchmark asks one question about one table. A harder extension, FinQA reasoning data, introduces dependencies across two to five tables and requires sequential decisions and planning. Its reported splits contain roughly 1,000 training examples, 120 validation examples, and 80 benchmark examples.
Training on the simpler single-table data improved performance on this multi-table benchmark without explicit multi-table training. The useful interpretation is compositional: a model that becomes more reliable at using a table can apply those skills repeatedly when a question requires several tables. This transfer result makes the original tool-discipline diagnosis more plausible, because the benefit survives an increase in planning demands.
Specialization also raised a different concern: would finance training damage general tool-calling ability? On BFCL, a benchmark for general tool calling, overall accuracy slightly improved, with minor gains in multi-turn and memory categories. For this model and evaluation, reinforcement-learning fine-tuning preserved broader tool-use competence rather than sacrificing it for the financial task.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Harder examples and richer rewards did not win
The data ablation compared three ways to train:
- Single-table FinQA only: Concentrate on the simpler tasks.
- A single- and multi-table mixture: Include both kinds of questions during training.
- A curriculum: Start with single-table questions, then move to multi-table questions.
The simpler dataset produced the most lift. Dickens interprets this as evidence that reasoning depth was not the main bottleneck in this experiment. The model first needed to become reliable at the fundamentals of tool use; once those improved, it could compose them into longer solutions. Adding harder examples did not beat addressing the failures already visible in the easier tasks.
Reward design tested a separate idea. Perhaps an agent would learn faster if it received credit for intermediate behavior, such as table access or query completeness. The most elaborate variant used an expert-developed rubric with multiple weighted components. That creates a more detailed definition of a good trajectory, but also requires decisions about which behaviors deserve credit and how much each should count.
Again, the simplest version won: a single binary correctness reward produced the greatest lift on this data. The result supports starting with the task’s verifiable outcome before investing in a complicated reward rubric. It does not make intermediate rewards useless everywhere; here, the additional reward engineering did not outperform pass or fail.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Evaluate the environment, the workflow length, and the output
The enterprise blueprint starts with data that represents the target tasks and outcomes, then uses reinforcement learning to improve a small model within that setting. Dickens connects this approach to better deployment economics and reports related results in healthcare, law, and insurance. The concrete finance experiment supplies the detailed example: verified questions, an interactive table environment, a checkable final answer, and a training objective that improves completion.
Reliable evaluation must then grow beyond a short question with one final answer. Dickens proposes three axes:
- Environment complexity: Test how realistic, complex, and dynamic the agent’s working environment is. A tidy task does not exercise all the difficulties of an enterprise stack.
- Autonomy horizon: Evaluate the lengths of workflows people actually intend to delegate, including whether the agent makes safe, useful decisions on its own and improves its judgment as a copilot over time.
- Output complexity: Cover the range of outputs produced in day-to-day work. Dickens sees rigorous evaluation beyond the prevailing focus on text artifacts as an underexplored opportunity.
Those axes expose the next challenge for the specialist. Answering a verifiable financial question is a useful training target, but dependable enterprise work also requires handling changing environments, longer sequences of decisions, and richer deliverables. The benchmark should expand along the dimensions that the intended deployment actually demands.
Snorkel’s Open Benchmarks grants funded this project as part of a multimillion-dollar commitment to open benchmarks. The program’s described path runs from submission and selection through development, publication, launch, and promotion. The talk closes by pointing readers toward the released training scripts, synthetic data, and model artifacts—a practical invitation to build on the experiment and to create evaluations that test the next set of real workflows.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Explores a complementary way to improve agent behavior: use execution traces and error messages to edit prompts and programs, rather than updating model weights from a scalar reward.
Read the complete timestamped transcript
- 0:12
Everyone, thanks for joining. I'm Charles, and I'm a research scientist at Snorkel AI. And today I'm gonna be talking about how we got a four billion parameter model to outsmart its two hundred and thirty-five billion parameter variant. And this is a specific result in a simulated financial environment, but I think there are general takeaways that I wanna make sure are taken home. And the first is for enterprise AI workloads,
- 0:42
reliability and specialization, uh, can actually outweigh raw parameter scale. And this is possible with high-quality data that actually represents the domain and task distribution that you're targeting. And, uh, this can result in smaller models that can achieve parity with frontier models at just a fraction of the cost. So before I get to it, I'm gonna take this opportunity to introduce ourselves a little bit. This is our team. We're a frontier
- 1:12
AI data lab, and our team, we have founders from labs at University of Washington, Wisconsin, Stanford, and an internal team of really talented FDEs, researchers. And we're in a fortunate position to play, and by that I mean build and deploy datasets and environments, uh, for frontier AI. And that's at the intersection of academia and industry research, but also grounded in enterprise deployments. Um, and this is all to say that we have a team that's
- 1:42
built to focus on building the best data possible to define and advance AI. And the way we do this is threefold. First is open research and benchmarks, so we're constantly contributing to open source benchmarking, and the project I'll be talking about today actually follows under this pillar. Um, but I also wanna acknowledge frontier lab data and enter-- environments as well as enterprise, uh, deployments are the things we work on.
- 2:12
And, um, under open research and benchmarks, we support research through grants, and I'll talk about opportunities here at the end of this talk. So, um, get ready for that. And then co-authoring with academic, uh, or industry partners, um, and leading independent research and publishing open artifacts. So what I'm talking about today is, uh, a partnership with UC Berkeley Sky Computing Lab, and we built, like I said, a four billion
- 2:42
parameter Qwen model to outperform its two hundred and thirty-five billion model variant on financial tasks. We saw the four billion parameter model, spoiler alert, uh, achieve almost a sixty percent pass at run, while the two thirty-five billion model, uh, achieved fifty-one percent. So exciting results, and this is all, uh, in collaboration with-- I've gotta give credit where credit's due, Manan Rungta, Shujun Tan from the Sky Computing Lab, as well as Bhavishya and Chris Glaze and myself, uh, from Snorkel. And you can
- 3:12
find all the assets, the, uh, training scripts, synthetic data, uh, on our GitHub. So check those out. And this is all open source and, um, available to you. So I'll cover today what, uh, we see financial institutions need, um, from AI, and this is from our enterprise deployments, uh, learnings, as well as the FinQA open source benchmark, and this is expert-validated financial question answering data that we trained our, uh-- and evaluated our model on.
- 3:42
And finally, uh, the training process for the RLLM FinQA 4B model. So let's get started. This is, uh, the one and only slide with the Charles Dickens reference, I promise. But this is a tale of two models, I'd say. First, the large generalist versus small specialist. And the point is here, when financial institutions ask us, "How well does the model or the agent reason?" What they typically mean is, can it operate in my stack of complex schemas, legacy
- 4:12
APIs, and tech debt? And can we trust it in realistic, uh, and long workflows without compounding errors? And can we reliably audit the agent? And that's important in financial and legal domains, for example. Which brings us to the question is, uh, how can we evaluate and improve agents to make them reliable in these, uh, real-world environments? So, um, the reflex may be to scale. We've all seen scaling law results. Uh, but today I'm gonna argue that specialized
- 4:42
workflows don't actually need generalists. We don't need polymaths. You wouldn't call Terence Tao for a tax audit, um, if you're lucky enough to have him in your contacts. You'd call the specialist that knows the forms, the tools, and the rules. So that's why we built FinQA. This is expert-validated financial question answering data, and I'll walk through the multi-stage pipeline of how we built this benchmark and training data. So first is schema and data extraction. This is coming from 10-K
- 5:12
reports. Uh, these are annual reports required by the SEC for all public companies, and they're used, uh, by analysts for identifying risks and informing investment decisions. Uh, we pull these from the EDGAR system and use the Qwen three thirty billion model to produce tables, SQL tables, approximately six thousand nine hundred of those. And we use each table to generate a single question-answer pair. And
- 5:42
this is, uh, produced along with a question taxonomy that was developed with financial experts that work with these documents, and these are used as context, uh, for the model to generate a question-answer pair along with metadata that we'll see, uh, what that's used for later. And finally, the third step is verification, and this is a three-layer verification process. First is programmatic consistency checks. That's where we use, um, the metadata, uh, which includes
- 6:11
lineage and table names and columns to make sure we're not hallucinating those. The second is automated reviews from independent agents. And finally, expert manual review, so humans actually going through these and making sure the questions and answers are realistic and verified.
- 6:29
And the outcome is Snorkel's FinQA data, and this is our first pass at this. Uh, realistic queries requiring planning, tool use, uh, tool calling, and reasoning. And, um, answers are single, verifiable, final answers. And from this, um, we get from our, uh, original set of roughly seven thousand tables, four thousand in train, five hundred in val, and two hundred and ninety in the
- 6:59
held-out benchmark. And the data is split so that no company in the benchmark overlaps, um, in any split.
- 7:08
And on this data, we identified, um, even with frontier models, uh, reoccurring discipline gaps, uh, failure modes. And the first is schema hallucination. So a model will actually assume that tables exist, column names exist, and, um, this can be an artifact of pre-training data, something it's seen, um, in its past. The next is context flooding, so actually flooding its own context, um, with poorly planned qu-queries. For example, just
- 7:39
calling select star, um, and this could overwhelm its context limits successful. And finally, poor recovery, um, from errors. So Qwen three two thirty-five would actually repeat the same failed strategy, um, instead of considering the error messages and adapting. And these failure modes are not actually just seen in the financial environment that I'm talking about today, but also what we saw in insurance underwriting. And this is, uh, work that I'm calling out that was
- 8:09
presented at the CAIS conference, um, a few weeks ago. So now I'll get into actually training the RLLM FinQA four billion model. Uh, for this, we used the RLLM, uh, framework, and this was developed by the-- our Sky Computing Lab collaborators. And if you haven't used this yet, um, this is a really user-friendly tool, and I'll-- A few core features, um, I wanna
- 8:38
call out is it really works with any agent framework, requires near zero code changes, just using, uh, a decorator pattern, CLI-first workflow, battle-tested results. So, um, we've seen RLLM actually improve not just on this finance, um, case study, but also in math reasoning, for example. Um, multiple RL algorithms are built in, GRPO, Reinforce RL OO, for example, and you can customize those, and then supports multiple, uh, training backends.
- 9:10
All right, now for our training environment. Um, this is the, uh, training environment we use for our model. The agent runs in a ReAct loop and has access to tools, uh, to interact with the tables we generated. The environment includes the roughly seven thousand tables we created, and finally, reward is binary correctness, uh, determined by a language model as judge, in this case, GPT-5 nano using reference-based evaluation.
- 9:41
Training details are listed here. This is, uh, I guess would be interesting to folks. We have the Qwen three four billion model is our base model, GRPO and binary reward, uh, for optimization, eh, the RLLM framework. Um, we ran roughly a thousand concurrent environments to generate trajectories, and this all was a cost under five hundred dollars to get the results that we saw, um, broken down between
- 10:11
compute, roughly four hundred and twenty, and forty dollars for a judge, uh, the judge API. And this was on eight H100s. The first result, um, is the central question: Can a four billion model, uh, armed with the specialized tools and training actually compete with its larger variant? And the answer is yes. So what we saw was actually the base model's accuracy, um, more than doubled, and it
- 10:41
outperforms its larger variant, the two thirty-five parameter model, and this is, uh, despite being a fraction of the size. And we evaluate all models shown here, um, on the expert-curated held-out two hundred and ninety samples.
- 10:57
So a natural concern is whether this actually transfers to harder problems. So this is just single question-answer pairs on a single table. Um, and to investigate this, we, uh, looked at our FinQA reasoning data. So this is a development on the FinQA, um, data I presented earlier. And the difference here is that now we're having, uh, dependencies on multiple tables, two to five tables, and requires sequential, um, decision-making and planning.
- 11:26
And we evaluated on this dataset. You can see the breakdown, um, as well. We have roughly a thousand in train, a hundred and twenty in validation, eighty in benchmarked. Um, and what we see is the tool use discipline actually, uh, generalized directly from training on the simple data. And no explicit training on multi-table examples was needed actually to see lift in, uh, this multi-table, um, variant of the dataset. And
- 11:57
we'll investigate why with ablation studies. Um, but the next concern we had was whether this, um, result would generalize to, um, more general use cases. So for this, we used the BFCL benchmark, and, um, this measures general tool-calling capability. And overall accuracy, uh, was, um, actually slightly improved. And then, um, we had also minor gains on multi-turn and memory. So specialization didn't actually
- 12:26
erode, uh, the model's broader tool use competence when we did RL fine-tuning for this model.
- 12:35
So to isolate what drove this performance, we ran an ablation study on data mixture, so different ways we can, um, combine the training data we had. The first is just training on FinQA data, and the second single and multi-table, and finally a curriculum where we start with single and then move to multi. And surprisingly, just training on the simpler set, uh, resulted in the most lift. Um, and this partially explains why, um, the generalization result we saw earlier, the bottleneck was never reasoning depth, it was just tool use, meaning the
- 13:05
model where it was actually failing and improving the reliability of that. And once you mastered the fundamentals, it could compose those skills.
- 13:14
Um,
- 13:18
yeah. So, oops. Let me go back. Um, the next natural question is whether, um, more informative reward signals, um, for example, rewarding intermediate steps, uh, table access, query completeness, things like that could actually accelerate learning. Uh, so this is different from curriculum, for example. And the most sophisticated variant we came up with was a, a rubric, and this uses fine grain scoring with, uh, multiple weighted components and this is actually coming-- came up with, with experts.
- 13:48
Um, and despite this investment in reward engineering, we actually again saw, um, simpler was better. A single binary reward signal, um, was what got us the most lift on this data, which led us to a blueprint, um, that we've been developing for enterprise agents. Um, and generally, I would say this study showed us that small models, uh, can successfully be trained, uh, with RL, um, when you have quality data to, uh, represent the
- 14:18
target, um, outcomes and tasks. And this can result in better deployment economics, uh, for specialized tasks and domains. And, um, we've seen this and repeated this result in other areas, including healthcare, law, and insurance, um, which I pointed to earlier. So check that out. Um, and our point of view as far as building, um, axes for evaluating agents is this, is environment
- 14:48
complexity is, um, one of the most important axes. How complex, how realistic and dynamic are the environments, uh, the agent's operating in? And the next is autonomy horizon. Um, are agents being tested and evaluated at the lengths and the horizon lengths and the way people actually wanna use agents? Um, are they making safe and good decisions by themselves? And can they actually improve, um, their judgment as a co-pilot over time? And finally, output complexity.
- 15:18
We wanna capture the wide range of outputs in our day-to-day work. Um, and I would say that this is currently under explored. Um, new benchmarks are coming up in this direction, but, uh, still I see a big, uh, primary focus on, um, text artifacts and, um, creating more rigorous evaluations in, um, these settings is, uh, a big opportunity. Uh, progress is being made here. So on that note, um, I'm gonna
- 15:48
advertise our Open Benchmarks grants, which funded this project, and this is a multimillion-dollar commitment, um, to producing open benchmarks like this. Um, we've had a lot of success, Agent, uh, Last Exam, um, Judgment Bench, um, others are coming out and this is, um, really accelerating our ability to measure the frontier and shape the frontier. Um, so to apply, uh, check out our website. Um,
- 16:18
you'll go through a submission, uh, selection review process, develop, and then finally publication, launch, um, and promotion. And also wanna call out that we're hiring in professional research, um, you name it, AI engineers and, um, yeah, check that out. You can, uh, check out our website. And finally, thank you so much for attending. Um, you can use this QR
- 16:48
code to see the blog as well as pointers to all the artifacts that released, uh, on GitHub, Hugging Face. Um, yeah. Thank you all.