AI Engineer Code 2025
Dashboards Are Dead — Sarah Simionescu, Composio
Read the talk
Dashboards Are Dead
Sarah Simionescu explains why moving work into an agent requires more than an MCP connection: tools must be discoverable, their dependencies must be clear, and data must move across apps without filling the model’s context.
From a talk by Sarah Simionescu
At a glance
Ideas worth remembering
MCP provides communication with a service; tool discovery, prerequisite guidance, and coordination across apps still require interface design.
Returning a plan alongside relevant tools helps an agent resolve dependencies, such as finding a Slack channel ID before querying its messages.
The cross-app analytics example saves a PostHog cohort outside model context, inspects Metabase’s structure, and generates SQL that uses the saved IDs.
Preparing an application for agents means making its operations usable toward a goal, beyond making its website discoverable or adding an AI button.
Six months of Datadog, without learning the dashboard
Sarah Simionescu, whose team is responsible for the Composio dashboard, opens with an awkward discovery. She had used Datadog every day for six months. Yet when a coworker asked to see an alert and she opened its dashboard, she froze: she did not know where anything was. Daily use of the service had stopped requiring familiarity with its visual interface.
That experience gives “Dashboards Are Dead” its specific meaning. An agent can become the place where a person asks questions and gets work done, while the applications supplying the data recede from view. The provocation comes with a joke at her own expense: this is great news for the audience, perhaps less comfortable news for someone building a dashboard.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Five interfaces stand between a bug report and a fix
The postmortem begins in 2022, the jokingly named Dark Ages. A Slack message reports an error when searching for prod. Resolving it means reading the complaint in Slack, querying Datadog, checking the session in PostHog, changing code in VS Code, and opening a pull request in GitHub. The task is one investigation; its execution spreads across five tools and five interfaces.
Each interface also asks the user to translate intent into its own language. Datadog has query syntax, Jira has JQL, and Slack has search modifiers. Even tools that offer SQL can disagree about the dialect. Redesigns add another learning cost, but the deeper burden is expressing the same investigation through several different systems. In this framing, dashboards and query languages are translation devices between a person’s question and the data that could answer it.
In the talk’s 2023 chapter, the “sparkle button” moves some of that translation into an LLM. A user asks a question, and the model writes a query inside the application. Simionescu’s joke about questions requiring more than two database joins captures her skepticism about the reliability of these early features; it is not a demonstrated technical ceiling. The user still visits the dashboard, and the generated query is only sometimes correct.
The next change is where the interaction happens. With Anthropic’s November 2024 announcement of MCP, an agent such as Claude could connect to a service, generate a query, execute it, and return the answer directly. The question no longer had to begin inside that service’s dashboard. But a communication protocol leaves the service responsible for deciding what tools and guidance the agent receives.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
MCP opens the doors; the agent still needs a map
Connecting a dozen MCP servers exposes a second interface problem. The agent can reach the applications, but reaching them does not tell it how to complete a task. Simionescu identifies three recurring failures in that setup:
- Repeated discovery. In the fresh-conversation setup she describes, yesterday’s mistake formatting a Slack link does not teach today’s agent how to do it. Skills can carry instructions forward, but those instructions consume context.
- Too many tool definitions. Connecting enough servers can load thousands of definitions into the context window. Composio’s GitHub toolkit alone has over 200 tools. The model may select the wrong operation or struggle to determine which prerequisite call must happen first.
- Disconnected applications. Each server describes its own app. A task spanning multiple apps still needs someone to connect the steps and move the relevant information between them.
The tradeoff is between supplying useful guidance and overwhelming the agent with everything it might need. More tools and more instructions can increase the search burden before any work begins. The talk’s memorable image is an agent standing in thousands of separate rooms: MCP provides entry, while the missing map contains tool selection, prerequisites, and relationships across applications.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Search returns tools and the order in which to use them
The bug-report example returns in a 2026 mock-up. The user pastes a Slack message link into Claude and asks it to use Sentry and Datadog, find the root cause, and create a draft pull request. The mock-up depicts a fix in less than five minutes; that timing illustrates the proposed experience rather than establishing an independently measured debugging result.
Claude first calls Composio Search with three goals: fetch Slack messages, search Sentry issues, and search Datadog logs. Search returns the relevant tools together with a plan for using them. The Slack prerequisite makes the difference concrete: the agent must find the channel ID before querying messages. A tool name alone would leave that dependency for the model to discover.
Once the agent retrieves the complaint, it can investigate Datadog and Sentry in parallel while scanning the codebase. Those steps supply the evidence needed to identify a root cause and propose a code change. The observable change in the workflow is from a human carrying context through several windows to an agent carrying out a coordinated investigation and producing a draft PR. Simionescu says she did not build a custom workflow or write a skill for this task.
The talk then presents early, unreleased comparisons against apps’ native MCP servers listed in the Claude marketplace, using the same tasks and model. Simionescu reports a clear difference, but the supplied explanation gives neither numerical results nor an evaluation protocol, so it supports no quantified advantage. The proposed reason is the interface: Composio translates changing, sparsely documented APIs into operations that agents can discover and use together, building on MCP, CLI, and native tools.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Carry a cohort across apps without carrying every row into context
The analytics example starts with a straightforward question: which verticals do users select during onboarding? Claude uses Composio Search, writes a PostHog query, executes it, and returns the distribution. The follow-up is more interesting: among users who selected e-commerce, which toolkits do they use? Here, “toolkits” means the apps used through Composio. PostHog identifies the cohort, while Metabase supplies the data needed to examine its toolkit usage.
Claude queries PostHog for the e-commerce users’ IDs and saves the results without loading the full list into its context. It then queries Metabase for the database schema and samples data to understand the structure. This separates two needs: the model needs enough information to decide how to query, while the operation needs the actual IDs that define the cohort.
Composio Remote Workbench then dynamically generates SQL containing a regex string that searches for those user IDs, again without placing the entire list in the model’s context. What moves between the systems, and what does the model need to see? The diagram traces the saved cohort separately from the schema and samples used to plan the query.
The result is a view of the most popular toolkits for the e-commerce persona, assembled from both sources in less than a few minutes as described in the demonstration. The important mechanism is the separation of planning from bulk data handling: a saved cohort can participate in a later query without every ID becoming model input. The recording does not explain the regex construction or matching semantics, so it does not establish that this method works for arbitrary identifier formats.
Select IDs of users who chose e-commerce during onboarding.
The user IDs remain outside the model’s full context. Schema and sample queries inform the SQL that uses those IDs.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The application has a new kind of user
The ending turns from individual workflows to product design. Making a landing page friendly to AI discovery does not prepare the application itself for agent use. Composio began by turning popular applications into tools for developers’ agents; Simionescu now describes requests from startups whose clients want to use their services through their own agents.
That user arrives with a goal and a set of tools. In the talk’s phrasing, it has no eyes and will not click the sparkle button. Its useful interface consists of operations it can find, understand, and combine to finish the job. The two demonstrations give that product requirement substance: expose prerequisites for a debugging investigation, and support cross-app data handling without forcing the model to read every intermediate result.
Simionescu’s forecast is that the next era favors products that are easy for agents to use. Her opening moment changes meaning under that forecast. Being lost in Datadog initially felt like falling behind; after months of accomplishing the work elsewhere, it looked more like a preview of how application use could change.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Beating RL With Reflection: GEPA and Optimize Anything
Develops the learning question raised here: how execution traces can improve prompts and repository skills so later agent attempts repeat less discovery.
Read the complete timestamped transcript
- 0:12
Hello everyone. My name is Sarah, and I have a confession to make. So I've been using Datadog every day for the last six months, and the other day, my coworker comes up to me and he's like, "Hey, Sarah, can I look at this alert?" I'm like, "Okay." So I open up Datadog, and I'm logged out. It's kinda awkward, so I have to log in, and he's, like, looking over my shoulder, and I'm, like, overly conscious about everything that I'm doing on my computer. And I open up the Data- the Datadog dashboard,
- 0:43
and I just froze. I had no idea where anything was, and this is actually not a knock on Datadog. The problem was I had genuinely never opened the dashboard. I'd been using it every day without ever looking at it. And so today, I wanna make an argument that sounds kind of insane, and then show you that it's obvious. The dashboard is dead, and I think that's great news for everyone in this room, except probably for me, because I am--
- 1:13
My team is responsible for the Composio dashboard.
- 1:18
Alas, how did we get here? I believe a postmortem is in order. Shall we? The year is 2022. It is the Dark Ages. Someone pings me on Slack. They say, "Hey, I get this error when I try and search for prod." I would've instantly spawned five different windows, read the context in Slack, write a query in Datadog, check PostHog for this session, uh, fix the bug in VS Code, open the PR in GitHub. That is five tools and five UIs I have to learn and
- 1:48
relearn every time they ship a new redesign. But frankly, redesigns are the least of my problems because every single tool has its own query language. So Datadog has its own query syntax, Jira has JQL, Search- Slack has its search modifiers. And it gets worse because even when these tools claim to speak the same language, let's say SQL, they don't even agree on SQL.
- 2:15
Every dashboard you've ever used, every weird, obscure query language ever written, was a translation device between you and your data because the machines on the other end cannot understand what you actually wanted. You never wanted a dashboard or its cursed query language. You wanted the answer. And in 2023, everything changed.
- 2:37
The sparkle button.
- 2:40
The sparkle button was born, and with a click of this magical button, an LLM would write a sometimes correct query to get you the answer you need. That is, if your question does not require more than two database joins.
- 2:55
And soon enough, these sparkling buttons were everywhere, so dashboards became AI native. Problem solved, right? Well, not quite. In November of 2024, Anthropic announced the MCP protocol with a promise to create an open standard for connecting AI systems with its data sources, and that it did. Instead of using a sparkle button, your own agent, Claude, whom you've already been using every day, could generate the query for you, execute it on your behalf, and give you the answer directly. And you might be thinking,
- 3:26
problem solved, right? This is the end of the story. Well, far from it because MCP is a protocol. It's a channel for communication between agents and your service, and it's up to you, the service, to choose how to communicate with the agent, and it turns out that makes all the difference. If you've ever actually wired up a dozen MCP servers, you will know the reality is a mess for three reasons. Agents don't learn. Every conversation starts from zero. It has no memory of how it
- 3:56
fumbled formatting links properly in Slack yesterday, and so it's just gonna fumble again today. And we patch this with skills, but skills is just a Band-Aid because agents become dumber with the more context you provide, and loading more tools, more skills, takes more context. And if you connect enough servers, you're dumping thousands of tool definitions straight into the context window. I mean, our GitHub, like, toolkit alone has over 200 tools. The model just drowns. It grabs the wrong tool or struggles to resolve dependencies, and it can't figure out
- 4:26
which one it needs to call first. And third, every app is isolated. Each MCP server knows about itself and nothing else. So the moment a task spans two apps, and like they always do, it is your job to piece that together. And so MCP gave agents a door into every app, but it left them standing in thousands of separate rooms with no map and no memory of ever being there.
- 4:54
I'm gonna show you a mock-up demo here of how I would solve this bug report, uh, from the very beginning, how I would solve it in 2026. So imagine in this mock-up I have my Claude connected to the Composio MCP. What I would do is I would literally just copy and paste the link to the message from Slack and say, "Please use Sentry, Datadog, find the root cause, create a draft PR, make no mistakes." And the first thing it does is it calls Composio's search to state the task it wants to accomplish. So
- 5:24
in this case, it stated three. It wants to fetch Slack messages, it wants to search Sentry issues, it wants to search Datadog logs. And for each of these three, Composio returns not only the correct tools it needs, but actually a plan of how to use them. So for example, for Slack, you actually have to find the Slack channel ID before you can start querying for messages.
- 5:46
Now that it's armed with all the context it needs, it has a plan. It starts pulling the message from Slack to see what the user complaint was. Then it begins pulling from data sources in parallel, Datadog, Sentry, and it begins scanning the code base. Once it identifies the root cause, PR is up with a fix in less than five minutes. I didn't have to build a workflow or write a skill to teach it how to do this. This is literally just Claude using Composio's MCP. And it has become so much easier to do things this way
- 6:16
that opening up Claude has just become like muscle memory for me to do literally anything. And the dashboards I used to open every day are becoming increasingly unfamiliar to me.
- 6:27
And the impact of designing for agents is measurable. So these are some early unreleased, uh, results comparing Composio against each app's own native MCP that's listed in the Claude, uh, marketplace. So it's the same tasks, it's the same model, and we see a clear difference.
- 6:44
But why? Why is this experience so much better than native MCPs? It's because at Composio, we are developing a brand-new interface specifically designed for agents. We translate messy, sparsely documented, ever-changing APIs from all the apps that you live in every day into something that agents really love to use. So we build on top of MCP, CLI, and native tools to deliver a cohesive, unified, and excellent experience for your agents that allows them
- 7:14
to perform complex operations across these apps seamlessly. I'm gonna show you a more classic example of, like, what I usually use it for every day. So let's say I want to dig into some user data. I want to see the distribution of what vertical my users select during onboarding. I can simply just ask Claude. Claude will again run Composio Search. It will write a query in PostHog, execute it, nice and simple. There's my results. But what if I wanted to take this up a notch? So I want to
- 7:44
look at that e-commerce section right there, and I want to see what toolkits they like to use. Toolkits is another word for apps we like to use at Composio. So let's see if Claude can figure out how it can pull this data from Metabase, uh, given the user IDs in, in PostHog. So once again, it uses Composio Search. It queries the user IDs of those who selected e-commerce from PostHog, and it actually saves it without loading the full results into its context.
- 8:16
Then it makes some queries, uh, in Metabase to get the database schema, sample some data, get a feel for how things are structured. And once it's confident, it actually uses a very cool tool called Composio Remote Workbench to dynamically generate an SQL query with a regex string to search for all those user IDs, again, without ever loading the entire thing into its context window. And there we go.
- 8:42
So we can see the most popular toolkits amongst this user persona in less than a few minutes, uh, and that's data pulled from both sources. If you want to play with the live version, come visit us at the Composio booth. Uh, I'd love to have you play with this live.
- 9:00
And so the dashboard has died, and I'm thinking more and more about what this means for the future. Many startups have begun making their landing pages and their websites AI-friendly for GEO purposes, but so few have really prepared their applications to be used by agents. Composio began as a tool to help developers. We were taking these popular apps that users like and turned them into these tools that agents could use. But the humans are tired of dashboards, and their
- 9:29
agents a-are tired of poorly designed MCP servers. And now we're getting requests from startups saying their clients are begging for a way to use their services through their agents.
- 9:41
So this is a lesson for everyone building anything right now. You are now serving a new species of user. They don't have eyes. They are not going to click your sparkle button. It shows up with a goal and a set of tools, and it judges you on exactly one thing, whether it can get the job done. For a decade, we've built pro-- uh, we've built products to be easy for humans to use, and the next era belongs to those that are easy for agents to use. So
- 10:11
picture me again. I'm frozen in front of that Datadog dashboard, and I was thinking to myself in that moment that I'm falling behind. Uh, but now I think of that moment as a preview of what's to come. Thank you.