We Mapped 115 Microservices for Our Coding Agents — Kamalakannan Nandagopal, Postman
Read the talk
We Mapped 115 Microservices for Our Coding Agents
Postman’s API context graph connects endpoints, callers, implementations and production traces so coding agents can reason across repositories. Kamalakannan Nandagopal explains where that context improves discovery and redesign, why stale dependencies break it, and how cheaper analysis could move architecture checks into everyday development.
From a talk by Kamalakannan Nandagopal
At a glance
Ideas worth remembering
API context connects callers, endpoints, implementation code and data dependencies across repositories, giving agents a way to reason about distributed workflows.
Consumer-aware redesign requires both request shapes and their purpose. The graph's links into implementation context helped recover both in the free-form JSON example.
Freshness is part of correctness: new consumers caused an impact-assessment failure, and removed APIs or dependencies must also disappear from the graph.
Lower token costs could make architectural analysis frequent enough to run during PR review and local development, where problems can be caught earlier.
An autonomous agent still needs to understand the other services
Postman’s cloud features run on more than 115 microservices and thousands of REST API endpoints. For Kamalakannan Nandagopal, a staff engineer at Postman, that makes a familiar coding-agent success—implementing a change inside one repository—only part of the engineering problem. The change may depend on code, callers and deployed behavior scattered across hundreds of repositories.
The team’s agents progressed from line and function completion to repository-wide refactors, then to autonomous workflows that make changes, verify them, submit pull requests and review pull requests. As the scope of action grew, the scope of required knowledge grew with it. Production services can run in different environments and versions, while front ends, CLIs and API clients all use the same backend infrastructure. A locally plausible implementation can therefore miss a dependency elsewhere in the system.
Several familiar context tools address pieces of this problem:
- Documentation: A useful starting point, but writing and maintaining it costs engineering time. Nandagopal’s concern is that delegating all maintenance to an LLM still leaves a need for human curation.
- Skills and MCPs: They help agents use knowledge and tools, but their usefulness depends on the context they receive.
- Agent memory: It can preserve knowledge that previously lived in an engineer’s head. In the team’s experience, however, that memory remained tied to individual users rather than becoming a shared engineering resource.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
APIs connect the knowledge that repositories separate
The useful precedent came from Postman’s go-to-market team. Its business agents needed customer information from multiple sources. After several iterations, a shared RAG and GraphRAG pipeline turned that information into structured context for purpose-built agents. Engineering borrowed the idea of a centralized context layer, then chose APIs as the organizing structure.
That choice follows the shape of a distributed system. A microservice owns responsibilities associated with a domain or model; its APIs expose those responsibilities. Calls between APIs connect the services into workflows and user journeys. Mapping those calls gives an agent a route through the system that a single repository cannot supply.
The graph began with an inventory of production microservices and their exposed endpoints. Each endpoint was connected to the code that implements it, down to the implementation line. The team also mapped service-to-service and endpoint-to-endpoint calls, traced data through databases and caches, and connected front ends, CLIs and other API clients to the backend. LLMs indexed this information into the shared context graph.
What makes an architectural relationship useful enough for an agent to act on? The diagram separates the relationships being mapped from the evidence supporting them. Every graph data point had to rest on what Nandagopal calls a “hard truth”: an implementation line or a production telemetry trace. The LLM organizes the information; code and observed execution ground it.
The resulting map revealed a tightly interconnected central cluster of core services, typically classified as tier zero or tier one, and one or two heavily bloated services toward the edges. These patterns made the architecture visible at a scale larger than any individual service.
Applications that interact with backend infrastructure.
The graph connects application entry points to backend behavior while retaining code or telemetry evidence for its data points.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Measure real engineering tasks, starting with API discovery
To evaluate the graph, the team returned to real pull requests and important project discussions. For each PR, it identified the developer’s intent and the eventual outcome. For design discussions, it used the decision actually taken as the retrospective target. These became evaluations of whether an agent could recover useful engineering answers.
The comparison used three setups: Postman’s AI agent connected to the graph, Claude Code using GitHub code search, and Claude Code with the same graph exposed as a skill. Each evaluation received a score from zero to five, alongside token consumption. On the presented plots, high scores toward the left meant good answers with fewer tokens; high scores toward the right meant good answers at greater token cost. These are reported internal task results: the recording does not specify the evaluation count or detailed scoring rubric, so the comparisons should not be read as a general agent ranking.
API discovery asks a deceptively small question: which endpoint should return a particular user’s profile name? In a system with hundreds of APIs, several may appear to do the same thing. The graph-connected agents scored well on this search task, while pure code search performed poorly. The difference between Postman’s agent and Claude Code with graph context was small enough that the team did not investigate it further. Here, access to the shared map mattered more than that difference between agents.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The temporary JSON field becomes a redesign problem
The redesign example starts with a feature that is ready to ship. Its API has a request schema, an implementation and a tested product integration. Then a dependent team asks to pass one additional property for future use. With little time to work through the schema, the developer adds a free-form JSON context object and a TODO to revisit it two weeks later.
Two weeks later, the field intended to carry one attribute now carries fifteen. It may have callers in one, five or ten services—the example describes possible growth, rather than a fixed consumer count. The observable change is that an escape hatch has become a shared input format. Adding a schema now requires discovering what each caller sends and why it sends it; otherwise, the redesign risks rejecting data an existing consumer depends on.
Three different information problems sit inside that task:
- Consumer coverage: Finding every caller can be impractical across a real distributed system.
- Request shape: A service dependency map identifies connections, but does not necessarily reveal the fields each caller supplies. Production logs may lack that detail, including for privacy reasons.
- Intent: Even a complete list of request shapes does not explain why a caller chose a particular representation.
The graph gave Postman’s agent a way to work through those problems in order. It identified consumers, followed their connections into implementation code, and used the surrounding code to reason about the purpose of each input. In the presented evaluation, the agent recovered the distinct request shapes used by the consumers. That supplies the information needed to replace the free-form field with a considered schema; the result described here is the analysis, rather than a demonstrated production migration.
How does the agent get from an endpoint to a schema decision? The diagram shows why a list of dependencies alone is insufficient: each consumer must lead to both its request shape and the code context explaining its use. Pure code search performed poorly in this evaluation. Claude Code with the graph skill improved somewhat, but tended to stop after enough turns and wait for the user to decide how to proceed. Supplying context and carrying the analysis through to completion remained separate concerns.
The example grows from one intended attribute to fifteen.
Consumer discovery supplies coverage; implementation context supplies the request shapes and reasons needed for redesign.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Impact assessment works only while the dependencies stay current
Impact assessment extends the same reasoning to a proposed change. Before replacing a V2 API with a V3 API, an engineer needs to understand how V2 is used. Before changing an endpoint in a PR, the engineer needs to know which other services will be affected. In the presented successful impact tasks, Postman’s agent produced strong answers while using roughly half to one-third as many tokens as code search.
Then one evaluation reversed the pattern. Asked to identify services affected by an endpoint change, Postman’s agent slowed down and consumed many tokens while failing the task. Investigation found that new consumers had appeared between writing the evaluation and running it. The graph no longer represented the current dependency set.
This failure changes the maintenance requirement. Building the graph once is insufficient: additions and removals of APIs and dependencies both need to reach it. An omitted caller leaves an impact analysis incomplete; a removed dependency that remains in the map can send the analysis toward a relationship that no longer exists. Nandagopal connects those gaps directly to developer trust and, eventually, whether developers keep using the tool.
The team’s retrospective of PRs in its top ten repositories found that almost 75% affected APIs, either by changing an API’s external behavior or by integrating a new API. That is a finding about those repositories, but it explains why dependency-aware analysis mattered to this team: changes with possible cascading effects were a routine part of development.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The shared map makes an ecosystem-wide review possible
Once the graph existed, the team could ask a larger question: perform a complete engineering analysis and architecture review of the ecosystem. Postman’s agent produced a 21-page report with findings and action items for the CEO, CTO and staff engineers. Its value was the ability to inspect relationships across the system, beyond the scope of a single-repository review.
The reported findings covered several distinct kinds of risk:
- Dependency cycles: Interdependent services that can complicate an incident.
- Long-running migrations: Risky migrations that may signal accumulated technical debt.
- Single points of failure: Places where a failure could create a larger system problem.
- Telemetry blind spots: Behavior present in implementation code that did not appear in telemetry.
The next proposed layer adds business meaning and design rationale. Knowing which API calls another answers a structural question. Knowing why a data model exists, why a domain is organized a certain way, or why one API should be chosen over another helps an agent make subsequent design decisions. The team was exploring how to connect that semantic knowledge to the coding agent’s existing context.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Use cheaper analysis more often, and earlier
The closing ambition is to use efficiency to increase the frequency of engineering checks. If an analysis takes fewer resources, it becomes more practical to run it at the PR level or even in the engineer’s development environment. Earlier feedback could catch dependency problems while a change is still being shaped, giving developers more confidence in the code their agents generate. This is the intended direction, rather than a demonstrated rollout of all those checks.
At the time of the talk, Nandagopal described the capabilities as available through Postman and early access. For a practical starting point, the supplied Context Graph documentation explains connecting and indexing data sources, then querying the graph. It also distinguishes exploration from answers: the visualization helps users explore, while queries return grounded results. That fits the talk’s central requirement—an architectural map becomes useful to a coding agent when it can lead the agent to the relationships and evidence needed for a decision.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
A practical guide to connecting and indexing sources and querying architectural dependencies. It explains the graph's data model, background indexing and query interfaces, and cautions that indexing results may be incomplete until jobs finish.
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Explores a complementary way to improve coding agents: learning repository skills from execution traces and carrying that knowledge into later attempts and other models.
Read the complete timestamped transcript
- 0:12
Today I'm going to talk about how you can take your coding agents beyond code generation and how API context might be the answer for that. I'm Kamal. I'm a staff engineer at Postman. And Postman, as you might know, is the end-to-end API platform. But beyond the desktop app that allows you to test your REST APIs and build your APIs, uh, Postman also has, uh, a lot of cloud features. And powering all of those cloud features is an extensive microservices architecture. We have more than one
- 0:42
hundred and fifteen microservices and thousands of, uh, REST API endpoints. Uh, so I'm going to be using our real, uh, engineering use cases and workflows on, on top of our coding agent experiences and what we learn from it and how we have been improving the experience and the effectiveness of our coding agents with API context graph. So our journey with the coding agents s-started similar to how many of you might have gone through the same journey as well, starting with line and function-level auto-completion. This is the Copilot, uh,
- 1:12
days eras. Uh, moving on to full code base, uh, refactors and one-shot implementations and one-shot fixes. Uh, this is your Composer eras and Claude Code eras. Now moving on to fully autonomous coding agent workflows. Uh, we have agents that, uh, end-to-end start the workflow and then make the changes, verify the changes, send pull requests, and even verify pull requests as well. What we see as the coding agents excelling really well at are greenfield projects or fixes and improvements within
- 1:42
the context of a single project. But production systems tend to be slightly more complex. They tend to be distributed systems split across multiple microservices. They're often deployed across multiple environments, even multiple different versions in each of these environments. We have several applications, front-end, CLI, API, all of them talking to our back-end infrastructure. And all of this is spread across hundreds of repos where the code is distributed across all of them. When we think about improving coding agent performance, the first area that we start
- 2:12
looking into is documentation. Documentation is a great place to start. Uh, but as we all know, developers, we hate writing and maintaining documentation. It's an added overhead, it costs time, and it becomes tricky to go and maintain over time. But if you offload documentation purely to LLMs to author and maintain, research points out that purely LLM maintained and written documentation tends to degrade exponentially over time. Some sort of human curation is definitely needed to make it more effective. Skills and MCPs are
- 2:42
a really good alternative to look at, but we all know skills and MCPs are only as good as the context that they're provided to. Agent memory is actually a really good solution for, uh, this scenario. Um, in terms of real-world, uh, experiences, most of this knowledge is probably already tribal knowledge in people's heads, and now it's getting translated into agent memory within the coding agent or the harness. The biggest limitation factor that we see with this is agent memory is still very limited to an individual user, and no standard patterns have now emerged yet in terms of sharing agent memories
- 3:12
across your broader teams or across your entire engineering organization. While we were trying to solve this problem for the coding agents on the engineering side, our go-to-market teams were also building business agents for go-to-market agents for their end-to-end executive use cases. So they've been building, uh, custom agents that are purpose-built with customer data that is brought in from multiple data sources, and they've done several iterations to streamline all of this unstructured data coming in from all of these data sources into a structured data on top of which these
- 3:42
agents operate. What they have ended up with is a really optimized like RAG plus GraphRAG pipeline that serves as a centralized context layer that makes all of these individual, uh, business agents work really well and solve the problem for all of them. So we started looking at how can we learn from this, and what can we do to make the coding agent performance really better. And when we looked at this, what we realized is in a distributed, uh, systems and microservices architecture, APIs are that context layer for you. Complex engineering architecture is often broken down into
- 4:11
domains and models, and each microservice is responsible for owning one single responsibility as per a model or a domain, and APIs are what defines what that responsibility is. And API calls between systems is usually how you identify workflows and user journeys that span across multiple systems. So agents are only as good as the context that they're given. So we set out on a journey to build the best context layer for the agents. So we went out to start building an API context graph for what
- 4:41
Postman's engineering architecture looks like. The journey started by identifying and cataloging every single microservice that was on production. And then we identified every endpoint that is exposed by each of these REST API services. And then we captured how each and every one of them is implemented right down to the exact line of code where it was implemented. And we also mapped how each and every service talks to each other and how each endpoint connects to each ev-- each and every other endpoint. We also tracked how the data goes all the way
- 5:11
down across the stack, including databases and caches and all the way down, uh, in the entire stack. We also identified how applications like front-end, CLIs, APIs, how they interact with the whole back-end architecture and back-end in-- uh, uh, infrastructure. We get-- give all of this data and we index all of this data using LLM to produce a centralized and effective context graph layer. But one key principle that we had in mind always was every single data point that landed on the context graph always had to be grounded
- 5:41
down in a hard truth. And that hard truth is either a line of code that is implemented in your code base, or it could be a, a, a, an exact trace that is coming from a production telemetry. With this in place, we set out on a journey to populate our context graph, and this is snap-snapshot of what our, um, engineering architecture kind of looks like. So it is, uh, some of the very common patterns that you would expect start to emerge. You start to see a central cluster of core services, of very tightly intertwined, uh, in-intertwined
- 6:11
APIs and services that are tightly dependent on each other. You find one or two heavily bloated services that are on the outer edges of these systems. The ones at the center is typically what you would classify as tier zero or tier one services.
- 6:24
So now that we have the context graph, when we start looking at how does-- how do we measure what does it add in terms of improvements to the coding agents? So we took real examples from the Postman, uh, Git organiz-- uh, from, from our organization's Git repos. We took real PRs that were raised to the GitHub repos. And for each and every one of them, we identified what the developer was trying to do and what the actual outcome was, uh, at, at the end of the PR. We also took, like, important projects and what were the key design decisions and takeaways,
- 6:54
uh, that were being discussed in those design discussions. We mapped each and every one of them into an eval, and we mapped what the actual decision that was taken in retrospect as the ground truth that we wanted the evals to confirm into. To, to run the validations, we were able to get the Postman's AI agent connect to the API context graph, and we were able to validate all of these against the Postman AI agent as well. But to keep the comparison fair, we also validated this against a, a generic coding agent, in this case, Claude Code, and we tried to verify
- 7:24
these use cases simply by, like, a straightforward GitHub code search. We also took the same context graph, and we made it available as a skill and also compared that against the generic, uh, coding agent, in this case, uh, Claude Code. So you're gonna see the results of this in a graph plotted like this. Uh, so we scored every eval on zero to five scale, and we also measured the exact amount of tokens it took for each of these runs. So what you will see, if something shows up on the top right, it means that the agent did a good job. It was
- 7:54
able to score high on the evaluation, but it spent a lot of tokens doing so. If it's on the top left, it did a great job not only scoring high, but it also did it much more efficiently. If something goes beyond the dashed line, uh, it just did not meet the bar of what a good, um, score should be on that evaluation. So then we bucketed like com-- some of the common use cases and scenarios that we see in day-to-day development and where do they fall under. One key, uh, uh, scenario that we always notice is API discovery. Any developer who is
- 8:24
making a change either to the front end or back end, integrating with other systems, what is the right API to call? And, and in real-world systems, you probably have hundreds of different of-- hundreds of different APIs, and m-most likely you have multiple APIs that probably do the same thing. So in this case, uh, a simplified example that I've taken is what is the right API to, to call to get the profile name for a particular user. So in this evaluation, this is a pure search use case, and the results are what you would expect. Like, with the context graph, the agents score really well as compared to, like,
- 8:54
pure code search, where the agents don't do really well. And there's a slight deviation between, like, Postman and Claude, but the deviation is not that big enough, uh, for us to be, uh, diving deeper into. The second, second use case is where it gets really, uh, interesting. So this is what we call API design and redesign scenarios. So in this scenario, let's take a developer who's building a new feature, and to build that feature out, they've designed a new API, and they've designed a schema for the request input. They have defined the, uh, implementation of it. They have tested it out from their side, and then they go
- 9:24
integrate it with the product, and they are happy with the product, and they are ready to ship. At the last minute, one of the dependencies comes to them and then say, "Hey, would it be nice if I just had this one additional property that I would like to pass to this API that I would like to use at a later point in time?" So the developer is like, "Maybe we don't have enough time to evaluate all the different ways in which it can be done. What is the right schema to do this?" So they just add this one additional context, which is a free-form JSON object, and they add a to-do in the code saying, "Revisit this and add a schema two weeks
- 9:54
later." The feature ships to production. Everybody's happy. Two weeks later, the developer comes back and finds out that that one particular property which is supposed to send one attribute now has fifteen attributes, and it's being passed by one service, or it might be five services, it might be ten services. Now, the developer has no idea what exactly is the shape in which the request is being passed down, uh, to the API and what are all the different ways in which every single consumer has been passing data down to. So it can be, uh... So this is what I call,
- 10:24
like, analysis of a shape of input for each and every API and why they are used like that. So this can be hard for a bunch of reasons. First, tracking down every single consumer of your API can be impractical in real-world scenarios. Even if you had service dependencies and service maps, finding out the exact shape of input used by each and every consumer can get really tricky. Engineers often lack the detail or depth of information in production logs, and for privacy reasons, it may not even be available at a lot of times. Even if the shapes are
- 10:53
identified, finding out exactly why the consumer decided to pass it down like this is a very, very tricky problem. But in this scenario, when we took this evaluation and then ran it with the Postman agent, what was surprising was the Postman agent almost got it really well. And the thing that helped in this case is, first of all, the context graph having an awareness of every single consumer of an API, but not just that, but having the drill down or the ground truth of where exactly each of these APIs were implemented in code, and the context surrounding that code
- 11:23
helped the agent understand why the change was done. And very quickly, it was able to figure out exact unique shapes for every consumer that was calling this API, and you could be able to make the changes, uh, quite reliably and confidently. And you can see that, um, the pure code search does not do really well, but with Claude and the skill, um, it does slightly better, but it tends to give up, like, it-- after it's done, like, enough number of turns and waiting for the user to kind of confirm what to do next. So the other very common, uh, use cases that we start
- 11:53
to see is impact assessment. How do you figure out what is the impact of one particular change? So if I change this API, what are the po-potential use cases? Or maybe there is an existing API, there's a V2 API, and I'm interested in introducing a V3 API. How do I understand how the V2 is API-- V2 API is used first so that I can design a better V3 API? So this is one scenario where we see the context graph really shine well, and you can see that, like, Postman Agent not only does really well in terms of the quality of the output, but also it's
- 12:23
able to do it with almost, like, half or even one-third when compared to your code search. But this is not smooth sailing all the time. Sometimes we see failures, and in this case, one of the use cases that we threw at it was find out the impact of a change in one particular PR, where it in-introduce a change to an endpoint, and figure out all the, all the downstream services that are affected by that endpoint. And we started seeing a failure, and you can see the graph, like, almost inverted. Uh, we see Postman not only slowing down, but also, like, consuming a lot of tokens while it did that.
- 12:53
And when we investigated in, uh, into it, and then tried to understand why it was failing, what we realized was that, uh, there were new consumers between the time where the eval was written and the time it, the evaluation was run. So the context graph is not only important to populate, but it's equally important that you keep it up to date. And any delays in syncing your context graph with the ground truth, uh, creates gap in the result. And if there's a gap in the result, it immediately leads to a loss of trust by the developer, and that translates into lack of usage.
- 13:24
Not only detecting new APIs and new dependencies, detecting removals of APIs and removal of dependencies is equally important, and that can also lead to confusions and problems. So you might think that these use cases may or may not be common to you, but when we went back and retrospected on all the PRs that were raised in the top ten repos, we noticed that almost seventy-five percent of the PRs had some effect on the APIs. So either these changes are introducing a change to the external surface area of an API, or they are intro-introducing an integration of a new
- 13:54
API. And both of them have, can have, like, really, uh, cascading impact downstream.
- 14:01
So these use cases were, um, clear for us to do, like, a before and after comparison with the context graph and without the context graph. But as soon as we had this information, we were able to ask questions that we pre-- we were previously not able to ask. And these questions asked allowed us to be able to take the Postman Agent and say, "Hey, do a complete engineering analysis and an end-to-end architecture review of our entire s- uh, ecosystem." So this is where the light bulb moment happened for us. Uh, when we asked the Postman Agent to do this, it produced a twenty-one-page engineering report
- 14:31
going through every single weakness and, uh, uh, points of vulnerability. It found out interesting insights for us, including cycles of dependencies, which can be really problematic in terms of an incident. Uh, risky long-running, uh, migrations, which are signals of technical debt, and then single points of failure, which can be really problematic or issues that are waiting to happen. Uh, blind spots in telemetry, where stuff that is in the implementation does not reflect in telemetry. And then it came out with clear action items for, hey, this is what your CEO, this is what your CTO, and this is what your staff
- 15:01
engineers, uh, might need to do. On top of all of these things, when we look at what we can do for further improvements, uh, one of the areas that stands out is bringing in the semantic layer of what does your business do and connecting it into the context layer of what your coding agents take as input, uh, for their workflows. Another layer of improvement that we are looking into is how can we improve and add additional reasoning of why does a data model exist? Why does a domain map like this? Why does this API exist versus other APIs? Why,
- 15:31
why should you use one API over the other API? And all of these helps, uh, in subsequent coding decisions, and all of that translates into really effective, uh, output when you're using them in your day-to-day development workflows. So takeaways and learnings. Uh, we've realized that API context graph serves as a context layer for coding agents in a real-world distributed systems. Um, gives a high-level view that helps you answer questions about your engineering architecture that you were not able to answer in the context of a single repo or a single microservice. It shortens the time to find the right way to
- 16:01
do a particular problem. And all of this means that do the, doing this either with the right amount of quality or being much more efficient in being able to do this. Um, so efficiency translates either into your cost savings, but the way that we see it is if we can do one particular task a lot more efficiently, it means that we can do it more often. And this allows us to do the same evaluations at a lot more frequency, including bringing it down at the PR level or even bringing it down, right down to the development, uh, machine or in the
- 16:31
development environment where the, uh, where the engineer is working on. So the sooner we can move the decisions and catch problems earlier in the life cycle, we believe that the agents who are generating code can translate into peace of mind for the developer and help them ship, uh, changes to production with a lot more confidence going forward. So that's everything from my side. Uh, all of these capabilities are available within the Postman as a platform. These are also available, uh, to use, uh, as early access. Uh, p- I'll be around, uh, outside, uh, right after the talk if you have
- 17:01
any questions or follow-ups. We also have a booth, S thirty-two, right across the hall from here. That's everything from my side. Thank you so much.