Why We Deleted Our MCP Server and Rebuilt It — Abhi Arya, Reducto

Read the talk

Why We Deleted Our MCP Server and Rebuilt It

Reducto’s first MCP server could build document workflows confidently, but nobody could explain them. The rebuild moved product knowledge into tools, exposed uncertainty and used real sessions to improve future work.

From a talk by Abhi Arya

At a glance

Ideas worth remembering

  • A tool for every endpoint gives an agent access, but can leave workflow meaning and assembly entirely to the model. Encode known product patterns in tools that users can understand.

  • Pair structured operations with live-state snapshots so the agent can reference the actual pipeline, including specific errors, instead of guessing its current configuration.

  • Handle output uncertainty with extraction feedback and task uncertainty with clarifying questions. Evaluations should penalize unsupported assumptions rather than reward confident completion.

  • Logged overrides, low confidence and questions can feed longer-term guidance. The proposed learning loop improves the context around future agent work.

  • Design for the person accountable downstream: they need to understand what the agent built and discover mistakes quickly enough to act.

Document automation makes agent mistakes somebody’s problem

A clean document is the easy case. Production brings scanned documents, rotated pages and complex layouts that a model must interpret before it can extract useful data. Abhi Arya, who works on product at Reducto, introduces the company as the layer between those documents and the tools or agents consuming them. He reports that Reducto processed over three billion documents in the preceding three years, across customers including Harvey and Scale AI. Building that document layer also meant building agents around it—and discovering that access to the system did not make those agents reliable collaborators.

Selected presentation frame from Why We Deleted Our MCP Server and Rebuilt It — Abhi Arya, Reducto at 90 secondsOpen full source frame
A presentation slide shows a document-processing example beside the speaker.

The question of what counts as agent-first software arrived through a Mean Girls-inspired laptop sticker: “Get in, loser, we're building agent first software.” Asked to explain the phrase, Arya found that a good API, harness or prompt mostly described ways to improve a model’s performance. The architectural question was harder: who pays when the agent is wrong?

That question produces three designs, distinguished by what the agent can do and what the human can verify:

  • Auto mode. Give the agent tools and a prompt, then optimize for autonomy. A confident mistake can reach customers before the person responsible understands what happened.
  • User-first. Give the human all the configuration controls, then add an agent with little ability to act. The interface remains laborious, while the agent becomes a limited chatbot.
  • Agent experience in service of the user. Give the agent useful, scoped capabilities and enough context to act, while keeping its work understandable to the human who must answer for the result.
0:120:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

One tool per endpoint made a good demo and an opaque workflow

Reducto’s document pipelines combine operations such as parsing, classification, splitting and extraction for work including invoicing and contract management. Users, especially the sales team, spent too much time configuring those operations. An MCP server offered a way to move that assembly work into an agent: the user could describe the desired use case, and the agent could build the pipeline.

The first implementation used Claude Code to create an MCP tool for every API endpoint that powered the app, all in one large file. A carefully chosen prompt produced a pleasing Slack demo. But its author already understood both the product and the intended workflow. That knowledge supplied the clarity the interface itself lacked.

Team-wide use exposed the difference. Someone would supply incomplete notes from a customer call or a Notion page, and the agent would construct an entire workflow. It would confidently announce success, yet neither the user nor the agent could explain what the workflow actually did. The failure was an inability to inspect the meaning of the assembled system, even though the agent could call its APIs. There was no useful place for the human to question an assumption or validate the work.

Selected presentation frame from Why We Deleted Our MCP Server and Rebuilt It — Abhi Arya, Reducto at 363 secondsOpen full source frame
A slide presents three cards labeled Auto mode, User-first, and Agent Experience in service of the user.
3:424:12
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

3:42 · section reference included

Put the known workflow inside the tool

Arya deleted the implementation in one commit. Its replacement had fewer tools, with more product knowledge inside each one. The concrete example is classification followed by extraction: classify a document, then route it into an extraction node that pulls out downstream information. Previously, a generic step-creation tool had to be called repeatedly, leaving the model to assemble the relationship correctly. The rebuild made that relationship a single classify-to-extract tool.

Selected presentation frame from Why We Deleted Our MCP Server and Rebuilt It — Abhi Arya, Reducto at 423 secondsOpen full source frame
The slide contrasts repeated generic API step calls with tools that carry expertise: snapshot, classify, extract, and guidance.

This changes where the decision happens. The agent still helps construct the pipeline, but engineers encode the connections they already know should hold. Tool responses supply guidance about what comes next, and some tools chain operations automatically. The agent has less freedom to invent an arrangement, while the resulting arrangement is easier for a user to follow.

A separate snapshot tool addresses a different source of guessing: what currently exists? One call returns the live pipeline state, including the whole flow and specific error codes. That lets the agent reference actual state instead of inventing variable names or reconstructing the pipeline from its conversational context. The two changes work together: structured tools encode valid patterns, and snapshots provide the facts needed to use them.

Where did the responsibility for connecting classification to extraction move? The comparison below separates model assembly from the product-aware tool, with live state supplying context to the rebuilt agent. The important change is that a known relationship no longer has to be rediscovered during every conversation. This deliberately narrows possible pipeline shapes so the correct path is also the simplest path available.

Compare the ideasFrom repeated assembly to a product-aware operation

Choose how repeated generic step calls connect.

The rebuilt tool carries the classify-to-extract relationship; the snapshot supplies the pipeline’s actual state.

6:126:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:12 · section reference included

A valid pipeline can still produce a quietly wrong answer

The rebuilt tools made construction understandable. They did not resolve every mistake in the outcome. A pipeline can run successfully with a bad extraction schema or an incomplete description of the intended use case. Arya’s example is a run over a hundred documents that misses a scan on page thirty. There need not be a crash or an explicit error to alert the user: successful execution can conceal unsuccessful document work.

Two kinds of uncertainty need different responses:

  • Uncertainty in the output. Extraction returns confidence scores and bounding boxes, giving the system information about the result and its location in the document. Tools feed that information back to the model to help improve the extraction schema over time.
  • Uncertainty in the task. An ambiguous field or unfamiliar document type calls for a question to the user. Choosing between two plausible readings silently can produce a coherent pipeline for the wrong interpretation.

The evaluation criteria changed alongside the interface. Agents failed evaluations for building on shaky assumptions, and harness guidance directed them to use ask-user-question tools available in Claude Code and other harnesses. Asking became rewarded behavior. The MCP guidance also required explanations of what the agent created and why, so the user could review the work rather than accept a success announcement.

A clarifying question closes the gap between the words in a request and the outcome the user intended. It supplies missing context before the assumption spreads into the schema or workflow, while giving the person responsible a chance to exercise judgment. This adds an interaction to the process, but that interaction can prevent the agent from confidently completing the wrong task.

8:429:12
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:42 · section reference included

Keep the agent’s reach within what the human can inspect

Opaque construction, silent wrong answers and agents that pass tests by hiding behavior share a common shape: the agent has more autonomy and access than the human can verify. Pipeline construction, schema generation and context gathering are attractive places to use agent intelligence because the human is nearby and can catch a mistake while it is still comparatively cheap to change.

The intended division of work gives the agent room to build and explain, while relying on Reducto for document parsing and extraction. Arya reports higher user satisfaction as the work became easier to see, and describes the document layer as providing accurate results for trusted parts of the pipeline. These are qualitative product claims here, without accuracy measurements or stated guarantee conditions; inspectable construction therefore remains a reason to review results, rather than a promise that every extraction is correct.

The harness is what makes this division usable. Tools, current state, validators, questions and explanations surround the model with the information and constraints needed to deliver work downstream. Model intelligence supplies the ability to reason through construction; the surrounding system makes that reasoning useful to someone who must own the result.

11:4112:11
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:41 · section reference included

Turn overrides and questions into guidance for the next pipeline

Once the work was guided and inspectable, Reducto instrumented its use. MCP sessions logged prompts sent by the model, and team members exported Claude Code sessions. Those records exposed where humans overrode the agent, where confidence ran low and where the agent asked questions. Each signal identifies something useful about the interaction: an override reveals a decision the user rejected, while a question reveals context the agent needed before continuing.

That information flowed into the MCP’s longer-term guidance. Arya describes a self-improving loop per organization or user, where building more pipelines improved subsequent pipeline construction without the team rewriting guidance by hand. The recording does not explain the update algorithm or how proposed guidance changes were validated. The supported mechanism is a feedback loop through surrounding guidance, rather than a described model-weight training procedure.

How does one session influence the next? The cycle below shows the route through session records and persistent guidance. Human intervention becomes information the next attempt can use, so an interruption does more than repair the current workflow: it can help the system avoid repeating the same uncertainty.

How it fits togetherPipeline use feeds future pipeline guidance

The agent works within guided, inspectable capabilities.

Session behavior supplies feedback to organization- or user-specific guidance used in later construction.

13:1113:41
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:11 · section reference included

Adoption meant people could own what the agent built

The sales team’s behavior supplied a practical signal that the redesign mattered. Arya had exposed the agent as Reducto tools without explicitly telling the team to use it for pipeline construction. People incorporated it into their preparation for customer calls: gather customer context from meeting notes or Slack, use the MCP to ask questions about the customer, then build a pipeline for that customer.

Compare that with the first team-wide rollout. In both cases, customer information initiated workflow construction. Initially, incomplete notes led to a confident result nobody could explain. After the redesign, people chose the tools themselves and trusted the result enough to show it to high-value customers. The observable change was ownership: a person could present the work as something they had built and understood. Arya reports that it paid off in sales calls, without giving a numerical sales result.

The ending revises the meaning of agent-first. Better models, faster inference and more connectors do not settle the design question. The agent operates alongside people who must answer for its work downstream. When it is confidently wrong, who discovers the mistake, and how quickly? Structured tools make construction understandable; uncertainty signals and questions make questionable decisions visible; instrumentation carries those lessons into future attempts. Those are the conditions under which people can start trusting an agent with real work.

14:105:12
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

14:10 · section reference included

Resources

Read the complete timestamped transcript
  1. 0:12

    So I'm Abhi, and I work on product here at Reducto. And before I dive in right away, I wanted to give you guys some more context on what Reducto is. Uh, we're an agentic document platform, which basically means that we help teams turn messy, unstructured documents into data that your tools and your AI agents can understand. I'll jump into the agents themselves in a second, but I'd like to talk about where we're at first really quick. In the last three years, we've

  2. 0:42

    processed over three billion documents across customers like Harvey, Scale AI, and top five global tech companies and hedge funds. Running at that scale means that we've seen basically every way a document can break in production, and we've iterated our product and our models to handle basically anything that you can throw at them. And getting AI to work on a clean document with a frontier model is not super hard anymore. We basically--

  3. 1:12

    like, you can throw something in the claw and you might get like sixty percent of the way there, but the hard part is production. Scanned documents, rotated pages, et cetera. The model starts hallucinating because a large majority of enterprise data is unstructured. And so Reducto is a layer that sits between your model and that data. We handle all the complex layouts and post-processing, and that way we end up working with and building a lot of agents ourselves while designing these pipelines. So now that

  4. 1:42

    you know a little more about what we do, I'm gonna dive into our main topic with something seemingly unrelated. Um, I visited my girlfriend this past weekend, and she gave me a sticker from the movie Mean Girls, and basically it was an agentified version of it. So it says, "Get in, loser, we're building agent first software." And in true software engineer fashion, I had to put a sticker like that on my laptop. And she got this from work, but she asked me, like, "What is agent first software?" 'cause she doesn't work in AI. And as, as I

  5. 2:12

    was, like, doing so, as I was thinking about it, I realized I have a lot of different answers. Building a really good harness, building a really good API, writing or iterating on prompts, all of these don't really, um, land on the agent being put first. They're more so just optimizing the model. Agent first software is inherently an architecture problem. And so I think that structure, that architecture really comes down to one question, and that is

  6. 2:42

    who pays for the agent being wrong? And to that, I think there are like three buckets that you can put this into. So you have auto mode to start, where you optimize for autonomy, you write a prompt, and you let the agent do whatever it wants. And when it messes up, which it will, you won't have any idea because it's really confident about what it's doing, and you're the one that's paying for it being wrong to your customers, to your users, et cetera. Then you have user first. This is where you optimize the human interface, give the user all the

  7. 3:12

    knobs, and then you bolt on an agent that can't really do much, which is basically just a chatbot experience. You've both capped the user's understanding of the product because now they can't prompt it, and the agent's capability within the product as well. And lastly, the bucket that I think wins here is agent experience in service of the user. This is where the agent's capabilities are structured to give the model, um, rich and scope capabilities where it can do things that the user really wants,

  8. 3:42

    aiding and enabling them, and even where the human can still verify the work that it's doing. This framing is what turns agent experience into less of a question of the model and more of a question of context, capability, and overall self-improvement. And here's how I discovered these buckets and came to that last point. At Reducto, we've been building tools that basically automate end-to-end document work through pipelines. Think that we first parse a document, then we

  9. 4:12

    classify it, or maybe split it into sections and then extract from it. Um, and that's in use cases like invoicing, contract management, and all these things that require human input. One thing that I noticed early on was that any of the users, especially our sales team at Reducto, was spending all their time in configuration instead of outcomes. And that's a larger problem across the board because legacy interfaces were designed for kind of building that pipeline. Meanwhile, in this kind of agentic work world,

  10. 4:42

    we're increasingly focused on the final deliverable and how can we get the user to that point as fast as possible. So the clearer move to get here for us was an MCP server. Let the agent do the heavy lifting and only burden the user with thinking about their use case and writing a prompt. To get some validation for the direction, I threw together an MVP with Claude Code, and every API endpoint that drove our app was given an MCP tool in one massive file. I gave it a really pointed

  11. 5:12

    prompt. I put it into Slack, and honestly, it was pretty nice. Um, to some extent, it should have worked in that setting. I knew exactly what I was doing. I was building a really curated demo, and that's why the standalone demo told me nothing at all. The real test happened when I gave the MCP to everyone on the team, and immediately it broke. And here's how. Someone would feed it a half-baked description from a customer call or maybe a Notion page,

  12. 5:42

    and the agent would just start going. It would build the entire workflow end to end. But very confidently, it would tell the user that everything it did was right, but neither the user nor the agent could explain what the workflow actually did. And that is the rude awakening of auto mode in production. The agent never really said that it's not sure. It handled with every tool that it had, and there was nowhere for a human to really validate or think about what was going on. We built something that could communicate with our

  13. 6:12

    system, but we hadn't built something that could actually work with it. And you'd think that the fix here is more detail in the system prompt or more model guidance, but realistically, it's architecture. The whole file that we use to talk to the MCP, I deleted the entire thing in one commit, and what replaced it wasn't more tools. It was mostly fewer tools, but they carried the structure of actual work. So for example, instead of one create step

  14. 6:42

    of this workflow to, uh, call over and over and just praying that the agent assembled it right in the pipeline, I built tools that understood how our internal dashboard or our product works. For example, a common use case with Reducto is to route our classify endpoint that classifies documents into an extract node. And what that does is basically extract information downstream. And instead of basically letting the model guess this or figure it out through the prompt, I turn it into

  15. 7:12

    a single classify to extract tool. And after each step, the agent gets contextual guidance on what actually comes next with certain tools even chaining automatically like this one, rather than the agent having to guess or think about what is right versus being told what's right. Now, the agent doesn't get to freelance the parts that we know are true, and since agents are stateless, we even added a snapshot tool. So this is one call that basically hands back

  16. 7:42

    the live state of our pipeline from the context of the entire flow to the errors that were surfaced with, like, really specific error codes, and that way they can reference what's actually there instead of guessing variable names, hallucinating, et cetera. By moving context out of the prompt and into structure, the agent stopped guessing and started referencing. And because the tools only let pipelines take the shape that we as the engineers on the product wanted, the human could follow

  17. 8:12

    exactly what happened, and we could explain the agent's outputs to downstream users of this MCP tool. So to avoid auto mode, you don't hand your agent an API, you hand it tools that carry the overall expertise, such as the right patterns, the state, and the validators. So the correct path also becomes the simplest path that the MCP can take. So we can build good pipelines now. We're done, right? Not really. Thinking back to auto mode,

  18. 8:42

    the agent was confidently wrong and no one could see it. Um, the-- what we just talked about fixed that for building. The tools made it structurally legible and the human could follow what the agent did. But now the problem showed up at runtime. When the outcome itself that the agent is trying to build goes wrong, who can tell? For example, if we had a really bad schema or maybe the user didn't really specify the actual use case, how can we tell that the agent did something wrong? And with agents it goes wrong quietly.

  19. 9:12

    So instead of a segfault or some actual error, um, the pipeline runs across a hundred documents and maybe it missed a scan on page thirty or maybe the user didn't specify what types of documents were going into the pipeline. And basically both of these things are the same failure. The a-agent itself is very confidently wrong and the uncertainty that the agent has is completely invisible. So everything that we did to fix this basically came down

  20. 9:42

    to one, uh, change which was making uncertainty really legible. In our system or with Reducto, uh, failures carry information. So with extract we output confidence scores, bounding boxes, all of this information about what went wrong and we were able to build tools that feed that information back into the model which actually helped optimize the extract schema over time and get more accurate document outputs. But that's just the outcome uncertainty. The agent

  21. 10:12

    also has to surface its own uncertainty. When it hits something ambiguous, um, two ways to read a field or a document type that it hasn't seen, it shouldn't guess and it shouldn't just pause exactly where it is and rather it should ask, "I can read this two ways," and tell the user, "What did you mean?" And so basically in our evals we started failing the agent for building any shaky assumptions and we gave our harness, uh, prompt guidance to use the ask user question tools et

  22. 10:42

    cetera that Claude Code and other like frontier harnesses offer instead of guessing. So we rewarded asking instead of guessing. And judgment doesn't live in the agent alone. This entire pipeline is still overseen by the user and the confidence and the reasoning of the agent was then outputted as basically prompt guidance in the MCP. So now every time you make something new it explains to the user what it did and why so the user can then review it, inherently acting in

  23. 11:11

    service of the user. The agent basically satisfies what you wrote and not the, um, outcome that you wanted. So what happens if you make the uncertainty basically a bit more legible and you keep the final judgment where it can actually be accountable in the hands of the human, then that one clarifying question does the work of all three, giving the model context, capability and also adding a human in the loop to actually build that judgment and

  24. 11:41

    verify things as well. So going back to kind of all of these failure modes overall, we had initially an auto mode that nothing-- nobody could in-inspect. We had wrong answers that basically arrived in silence and then an agent that passed the test by reward hacking or hiding what it did. And underneath they're basically one shape where the agent had more autonomy and more access than the human could verify. And that

  25. 12:11

    intelligence lives in the construction overall. So building the pipeline, generating the schema and pulling in the right context is where the agent is more, um, most powerful and where a mistake is the cheapest because the human is right there. By building for that interface we were able to also have user satisfaction go up with the MCP because they could really see what was going on. And by using Reducto for the document layer, we can guarantee accurate parses and extractions

  26. 12:41

    for parts of the pipeline that we can trust agents with. Where the agent operates, it earns trust by being verifiable, correct, and building trust in service of the user. And the whole conversation we had about legibility overall is what inherently pays off. Agents are just intelligence in a box, and the harness that you surround them with is what provides them with the ability to deliver these outcomes downstream. And so once we reach this

  27. 13:11

    point, we focus on instrumentation, where every MCP session, the prompts that were sent from the model were logged, and we even asked, like, our sales team and other people on the team to basically export their Claude Code sessions to understand how the model was behaving, where humans overrode the agent, where confidence of the model ran low, and where it asked questions. And because we added all this instrumentation, um, it flowed back into the long-term guidance that the MCP had overall.

  28. 13:41

    So the system got better at building pipelines by building more pipelines. Without us rewriting by hand, this then became a self-improving loop per org or per user that was building these pipelines. The foundation became overall the same decision, where we kept the model very guided and very inspectable, and the information around it became everything else that actually guided the model. And so one thing that kind of told me that we were doing something right with these improvements

  29. 14:10

    is I never, like, explicitly told the sales team to use the agent to build pipelines, but rather I just exposed it in our cloud enterprise as Reducto tools. And the sales team and various other people on the team started basically wiring it into their workflow. So before they had a customer call, they would use CircleBack or Slack and basically learn about the customer and then use the MCP to basically ask them questions about the customer and then build a pipeline for that customer downstream.

  30. 14:41

    So when an agent's experience was built really in service of the users that were using it, we weren't pushing people directly toward it, and rather they reached for it and actually started using it as a company-wide resource. And that's where they also started to trust the output enough to put this in front of high-value customers because the demo didn't just work, but the human actually said, "Hey, I built this. Come check it out." And it actually paid off in these sales calls. And I hope the Slack reactions can

  31. 15:11

    also help gauge that a little bit. So overall, a few months ago, I would've told you that agent first is-- means that you have the best model, the best inference, speed, or intelligence of the model. And with models getting better from closed source to open source, from expensive to cheap, I think it now means almost the opposite of what it sounds. To some extent, the agent doesn't go first. And to build for agent experience means you have to consider the agent being

  32. 15:41

    in the loop with the people who have to answer for the work downstream. So it's not just what does the agent connect to. Everyone's agent connects to everything nowadays, and Claude can make a new connector for you immediately. But regard-- beyond that, it's when your agent is really confidently wrong, who finds out and how fast do they find out? If you build for that, then you can build for something that people will actually trust with real work. And yeah,

  33. 16:11

    that's about it. Thank you guys. And also we're at-- Reducto's at booth P8, and we are hiring for lots of positions, so please come check it out at reductoai.careers. Thank you.