← All AI Engineer talks

AI Engineer Summit 2023

Climbing the Ladder of Abstraction

Read the talk

Climbing the Ladder of Abstraction

From spreadsheet formulas to books and lodging searches, AI can help us move between useful views of the same information—and keep acting at every level.

From a talk by Amelia Wattenberger

What changed when accounting became a spreadsheet?

What did accounting look like in the early 1900s? Handwritten letters and numbers, annotations in margins, calculations performed by hand. Crossed-out entries and ink blots make the effort visible: maintaining the document was itself a substantial part of the work.

VisiCalc brought the spreadsheet to personal computers in 1979; Lotus 1-2-3 followed in 1983. The breakthrough was not automatic arithmetic. Calculators and computers already supplied that. It was a structured interface that connected calculations through formulas. Change a cell or add a row, and the dependent results update live.

Green-on-black VisiCalc payroll table with employee rows, tax columns and a populated totals row, shown in a YouTube player.
VisiCalc’s payroll spreadsheet with calculated totals.

That structure changed where an accountant could spend attention. Instead of repeatedly calculating values and maintaining rows and columns, they could concentrate on what the numbers meant. The interface made many small automations work together as a tool for a larger task.

0:160:22
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:16 · section reference included

Augmentation is built from smaller automations

The corresponding question for AI is how to combine new capabilities with interfaces people use today—and interfaces they might want tomorrow. Two goals often frame that discussion:

  • Automation: Do a tedious task for the user, such as copying data into a table or calculating values.
  • Augmentation: Give the user a new ability, or improve an existing one, especially in creative or nuanced work such as analyzing data.

Concerns about jobs being automated are valid, Amelia Wattenberger acknowledges. But treating automation and augmentation as opposing goals misses their relationship: augmentation can be composed of smaller automations. Aggregating data into a table and generating visualizations can both serve a larger activity that remains directed by a person: answering the question that motivated collecting the data.

The spreadsheet already works this way. Its cell calculations are automated, while Excel remains a tool people use to understand financial data. Automatically assembling the table or writing its formulas would extend that pattern, removing more of the preparation involved in analysis.

Slide with the statement “Each cell is automated, the overall task is augmented” above a multicolored spreadsheet.
Each cell is automated; the overall task is augmented.

This is the reason to look beyond chatbots as the default AI interface. General tools such as calculators and chatbots are useful, but surrounding their capabilities with structure can make them much more powerful for particular tasks. The desired arrangement leaves the user directing the work while the model handles frustrating subtasks inside the interface.

1:381:46
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

1:38 · section reference included

Zoom changes which information matters

The ladder of abstraction describes how the same object can be represented at different levels of detail. Digital maps make this familiar. At a close view of the Monterey Bay Aquarium, buildings, names, icons, and routes help someone navigate within the aquarium. That representation is less useful for getting to the aquarium in the first place.

Zooming out does more than make buildings smaller. The map hides some information and foregrounds other information as the task changes.

ViewInformation emphasizedSupported task
NeighborhoodCity streets and restaurantsFind and reach a local destination
RegionHighways and terrainTravel to and from Monterey
Broad geographyState and country shapesUnderstand the wider geography

Keeping every building label and street visible at every scale would make the broad views incomprehensible. Screen space and attention are limited, and most fine detail is irrelevant to the task at that level. Could AI bring this behavior to other kinds of information?

4:455:00
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:45 · section reference included

Zooming out on Peter Pan

What would it mean to zoom out on a book? Reading usually exposes every word, but remembering a book or explaining it to a friend involves topics and plots rather than exact wording. Language models can summarize and transform text, making these different representations available inside a reading interface.

Wattenberger’s demo begins with the first five chapters of Peter Pan. Initially, she scrolls through the original first chapter. Zooming out replaces each paragraph with a one-sentence summary. A minimap on the right makes the reduction in text volume visible. At the next level, a summary stands in for a group of roughly ten paragraphs, reducing the text further.

PenPal interface showing Peter Pan, short text blocks under Chapter 1 and Chapter 2 headings, and a narrow minimap on the right.
Peter Pan condensed into summaries under chapter headings.

At the highest level, each chapter becomes one sentence, so all five chapters fit on a single page. The unit of attention has moved from words, to paragraphs, to groups of paragraphs, to chapters.

That reading demonstration suggests a writing workflow: inspect the chapter-level representation, adjust pacing or plot structure there, and then zoom back in to examine how the full text changes. This editing behavior is a proposal, rather than a demonstrated round trip. Its attraction is concrete: the writer could work on the structure without holding every part of the book in memory while editing word by word.

6:537:00
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:53 · section reference included

From a text summary to a story arc

A book can also be represented by its emotional trajectory. A story arc maps mood across the narrative rather than summarizing its wording. Kurt Vonnegut’s familiar Man in a Hole shape follows a character into trouble, out again, and into better circumstances. Wattenberger offers The Hobbit, The Wizard of Oz, and Alice in Wonderland as examples of that pattern.

The proposed interface would assign sentiment values to sections of a book and plot them on a graph. A writer could then adjust part of that curve and inspect the resulting changes to the prose. Here the higher-level representation is no longer shorter text: it is a different editing surface for a property of the story. This is a direction for future writing tools, not a capability established by the reading demo.

8:318:44
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:31 · section reference included

Get information, reason about it, act

Combining smaller automations with movement between abstraction levels suggests a broader product direction. At the time of the talk, Wattenberger was on Adept’s design team, working in a company focused on training AI to read screens and use software through actions people perform. The goal was to make knowledge work—work done on a computer—easier.

Interviews about people’s daily work revealed a recurring sequence:

  1. Obtain information.
  2. Transform it or reason about it.
  3. Act on it.

Adept’s design exploration asks what it would mean to zoom out on information throughout that sequence. The lodging workflow that follows is a hypothetical interface sketch, including its booking and messaging actions.

9:319:43
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:31 · section reference included

A lodging page organized around your decision

Suppose you are attending a conference in San Francisco. You find Airbnb listings near the venue and open one. Its detail page contains information intended for everyone, but your decision depends on particular questions: How close is it to the venue? Is there a coffee maker? Is the Wi-Fi good? Answering them requires digging through the page.

Zooming out slightly would remove branding and generic detail that do not help with this decision. The new view would retain the listing name, rating, a short summary, and total price. Then it would add information specific to the trip: walking minutes to Hotel Nikko, where the conference is being held. Instead of repeatedly consulting a map, you could compare that value directly. If the walk is too long, the view could also identify the nearest BART station and the walk to it as an alternative.

Wi-Fi matters because the night before the conference may still involve working on the talk. AI could extract relevant reviews and summarize whether their evidence is positive or negative. That gives the traveler a quick judgment from reported experiences, rather than a measured connection speed. The interface could also select a useful bedroom, living-room, or kitchen photo from the larger gallery, reducing another search through presentation material.

The abstraction must preserve the ability to act. From this personalized view, the traveler would still be able to reserve the listing or message its host without returning to Airbnb. Removing detail should simplify the decision while leaving the person in control of what happens next.

10:3310:44
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:33 · section reference included

From individual listings to a comparison table

The next question is whether a hotel would be a better deal. A shared representation could present hotel and Airbnb information in the same form, regardless of the source website. Putting those views side by side would make their relevant differences easier to compare.

But fifty individual views would still be cumbersome. Zooming out again would turn the hypothetical set of fifty listings into a spreadsheet, with the deciding factors laid out as columns. The traveler could scan the distribution of total prices, compare walking times, and inspect counts of positive Wi-Fi reviews. Booking would remain available at this level: a clear winner could be reserved directly from the table without returning to Airbnb or Hotels.com.

12:5013:05
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

12:50 · section reference included

Select a cluster, then ask a new question

Sometimes neither the cheapest nor the closest listing is an obvious winner. A further abstraction would represent each listing as a circle on a scatterplot, making several criteria visible together.

EncodingMeaning
Horizontal positionPrice increases to the right
Vertical positionListings closest to the hotel sit near the bottom
ColorWi-Fi review evidence

This view would reveal a cluster of inexpensive, nearby listings with favorable Wi-Fi reviews. The traveler could identify that promising group without comparing every table cell individually.

Then another constraint appears: the flight arrives at 9 AM. The traveler would circle the promising cluster and message every selected property about early check-in. The proposed decision rule is to book the first property that replies that an 11 AM check-in is possible. Selection on the plot therefore becomes an input to an action across several listings; the visualization is more than a passive summary.

13:5514:07
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:55 · section reference included

Make it easy to move between levels

Each task benefits from a different representation. Today, much of that transformation happens in the user’s head: inspect one listing, remember its relevant details, inspect another, and try to maintain a comparison across fifty sources. That manual abstraction consumes memory and mental energy.

Wattenberger credits Bret Victor’s Up and Down the Ladder of Abstraction, whose verified publication is an interactive essay. The direction of travel matters: higher abstraction is not inherently better. The opportunity is to use AI to generate different levels, connect them, and make movement between them easy. A detailed listing, a table, and a scatterplot can remain useful parts of the same workflow.

Three diagonally arranged views show a hotel listing, a hotel comparison table and a colored scatterplot, each beside a vertical plus-and-minus control; a Send message popup appears over the plot.
Hotel information presented as a listing, comparison table and scatterplot.

Adept’s work here was an ongoing exploration. Its design ambition returns to the spreadsheet: automate the smaller, tedious parts so a person can do the larger task more effectively. For AI interfaces, that includes generating useful representations of information and preserving the ability to act through them. The gain comes from making the right level available for the task at hand.

14:5615:01
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

14:56 · section reference included

Resources

From the talk

Read the complete timestamped transcript
  1. 0:00

    [upbeat music] I wanna start with a question.

  2. 0:16

    Does anybody remember what accounting looked like in the early 1900s?

  3. 0:22

    Yeah, me neither. But from what I gather, it was super frustrating- I'm having some trouble with the connection. Please try again in a moment ... and tedious, and it involved a lot of, like, writing letters and numbers, uh, annotating in margins, performing calculations by hand.

  4. 0:37

    Um, you can probably look at these pages and sense how frustrating it is by looking at how many things are crossed out and all the ink blots on the page.

  5. 0:45

    So thankfully, this isn't how the job's done these days. So in 1979, VisiCalc totally changed the game, and this was the first spreadsheet for personal computers.

  6. 0:56

    It became an essential tool for accountants, at least until Lotus 1-2-3 was launched, uh, four years later. And the innovation here wasn't performing the calculations automatically. We already had calculators and computers to do that for us.

  7. 1:11

    But instead, the innovation was having the structured interface that stacked those automatic calculations together into formulas, so that when you change the value of a cell or you add a row to your spreadsheet, uh, all of the spreadsheet numbers would be updated live.

  8. 1:27

    So instead of spending all day doing calculations or manually updating the rows and columns, accountants could now spend more time, uh, worrying about the actual numbers.

  9. 1:38

    Okay. Fortunately or unfortunately, this isn't a spreadsheet conference, so let's get back to talking about AI.

  10. 1:46

    So one of the things I'm most interested in is what are the best ways to combine our new AI superpowers with the interfaces that we use today, or more importantly, the interfaces that we wanna use tomorrow.

  11. 2:00

    So often when people talk about building interfaces with AI, they refer to these two distinct goals, whether it's automation or augmentation. In essence, automation takes rote tasks and does them for the user, which is really great for anything that's super tedious or boring, like copy and pasting data into a table or doing calculations by hand.

  12. 2:21

    And in contrast, augmentation gives the user a new ability or it improves their existing abilities, which is awesome for things that are creative or nuanced, things we don't really trust models with yet, uh, like analyzing data.

  13. 2:35

    And I think this contrast often ignores how related these two concepts really are. Um, automation has become a bit of a buzzword, um, or a trigger word where people are worried about their s- jobs being automated.

  14. 2:52

    And I think this is a very valid concern, and I kind of wanna reframe this dichotomy. So, uh, instead, I think augmentation is composed of smaller automations. If our end goal is to automate tasks or jobs, we'll still need to automate parts of them.

  15. 3:10

    So for example, if the end goal is analyzing data, automating the smaller tasks like aggregating the data into a table or generating visualizations from that table are gonna help focus on your end goal, which is answering the question that motivated the data collection in the first place.

  16. 3:26

    So if we go back to our spreadsheet example, we can think of each cell, the calculations that, uh, that create them as having been automated away. And no one really thinks of spreadsheets as taking people's jobs.

  17. 3:41

    Instead, uh, Excel, what I'm showing here, which is kinda like the current king of spreadsheets, uh, is an essential tool for people who interact with things like financial data.

  18. 3:52

    If we automate these parts behind the scenes, uh, that's the first step towards achieving the goal of augmenting, uh, working with data. So in the future, we can easily imagine having this table aggregated automatically or writing the formulas for us, and having all this work done helps augment us in our greater goal of analyzing and understanding the

  19. 4:13

    data. This is one of the reasons why you might hear me say some things like chatbots aren't necessarily the future. Um, I think that these flexible general tools like calculators and chatbots are wonderful, but then adding the structured interface around them makes them so much more powerful for a ton of different use cases.

  20. 4:34

    What we want is something where the technology behind chatbots is embedded into the interfaces where we're still driving, but the model's automating away the smaller tasks that we find so frustrating.

  21. 4:45

    So what might these interfaces look like? Before answering that question, I wanna introduce one more concept: the ladder of abstraction. So the basic idea here is that the exact same object can be represented at many different levels of detail.

  22. 5:00

    So I think maps are a good example of this.

  23. 5:04

    Um, we take this interface for granted, but Google Maps and other digital maps are incredibly compelling interfaces. They're so well designed, um, and they help represent different tasks involving navigation and localization at different scales.

  24. 5:18

    So here we are at the most zoomed in scale, and we can see all of the different structures within the Monterey Bay Aquarium. We can see individual buildings, the names, the icons for them, maybe routes between the buildings, and this is great for navigat-navigating around the aquarium, but maybe not so great for getting to the aquarium.

  25. 5:38

    As we zoom out, all of these buildings get smaller because they're further away, but that's not the only thing that happens. So at these more zoomed out levels, uh, Google Maps actually starts hiding information.

  26. 5:50

    So I can't see the buildings inside of the aquarium anymore or their icons or names, but instead I can see city streets and, uh, different restaurants, and this will support a different te- set of tasks, like finding a restaurant or a destination and getting to that place

  27. 6:08

    Zooming out even further, we lose those city streets and stores, and instead, we look at highways and terrain. And again, we have a different task here. This level supports longer range travel, getting to and from Monterey.

  28. 6:21

    And then if we go all the way out, we're mostly looking at the shape of states or countries. So if we tried to keep all of that information at higher zoom levels, it would be completely incomprehensible.

  29. 6:34

    There's really only so much information we can fit in our brains and so many pixels on a screen, and most of that detail isn't relevant for the task we're trying to do anyway.

  30. 6:45

    So you could wonder, can we use AI to bring these kinds of principles to other types of interfaces?

  31. 6:53

    For example, what would happen if I zoomed out on a book? What would that even look like?

  32. 7:00

    Typically, when we, we read a book, we're looking at every single word, but that's not the only level we think about. When remembering books we've read in the past or summarizing a book for a friend, we're more concerned with overall topics and plots than specific wording.

  33. 7:15

    And now that we have access to language models, which are amazing at summarizing and transforming text, how can we use them to change the way we read and write?

  34. 7:26

    So here's a quick demo I put together of the first five chapters of Peter Pan,

  35. 7:31

    and [laughs] there's no tricks here. I'm just scrolling through the first chapter. So if we take this and we use an LLM to zoom out, we can see each paragraph change to a one-sentence summary.

  36. 7:44

    And we have a, a mini map to the right, and you can kind of see, uh, how much fewer words there are in the page and how much more quickly I could read this.

  37. 7:54

    If we zoom out another level, we can see summaries of, say, 10 paragraphs at once. And again, you can see in the mini map, we have way less text to read.

  38. 8:04

    And then finally, at that highest zoom level, we've reduced each chapter in one sentence, and here we can fit five chapters on one page.

  39. 8:15

    So if I were writing Peter Pan and I wanted to do something like tweak the pacing or modify the plot structure, viewing the text at this highest zoom level, editing it, and then zooming back in to see how that changed the raw text would be a much nicer workflow than keeping all the parts in your head as

  40. 8:31

    you change it word by word. So another way to think about a book at a high level is with a story arc, um, and this describes the mood mapped over an entire story.

  41. 8:44

    Uh, you might be familiar with Kurt Vonnegut's graphical representation of the most common story arcs. For example, we have Man in a Hole, uh, where the main character gets in trouble, gets out of it, and ends up better for the experience, um, which you'll see in stories like The Hobbit or The Wizard of Oz or Alice in

  42. 9:01

    Wonderland. What if we could, uh, take the semantic value of all the sections in a book and plot that on a, on a graph? And then if we wanted to edit the story, we could go ahead and tweak parts of that graph and see how the raw texts change.

  43. 9:22

    I mainly highlight this because I'm super excited to see how we use AI to innovate on writing tools within the next few years.

  44. 9:31

    But first, let's combine the two concepts. Um, so the first concept is augmentation as stacked automations, and the second concept is traversing the ladder of abstraction for different tasks.

  45. 9:43

    How might this look in a more general product? So I'm on the design team of a startup here in SF named Adept, and at Adept, we're focused on training AI to use software, read screens, and take actions the way humans do.

  46. 9:58

    And our end goal is to make knowledge work easier, so any work on a computer.

  47. 10:03

    So after speaking with a lot of people about what they do day to day at their jobs, we found that much of knowledge work involves getting information, transforming or reasoning about it, and then acting on that information.

  48. 10:15

    So given this really common workflow, one of the things we've been thinking about is, what might it mean to zoom out on any piece of information? So we have some sketches where we're exploring what that might feel like or what it might enable us to do.

  49. 10:28

    I thought it'd be really fun to share one of those with y'all today.

  50. 10:33

    All right, so completely hypothetical situation. Let's say I was going to an awesome conference in San Francisco. What I would do first is I would go to Airbnb. I'd find listings near the venue.

  51. 10:44

    I'd click into the detail page of one of the listings, and there's all this generic information that should work for everybody. But I have specific criteria that'll help me decide whether or not it's the right thing to book.

  52. 10:56

    So I'm gonna be l- digging through this page, looking for things like, how close is it to the venue? Is there a coffee maker? Does it have good Wi-Fi?

  53. 11:05

    That kind of thing. This kind of decision would be much easier if I could zoom out just a little, get rid of all the branding and standard information that isn't really important to me right now, and focus on my deciding factors.

  54. 11:19

    So to start, I can see the name of the listing, maybe the rating, a quick summary, and the total price. And this is all pretty generic so far, but I know this conference is at the esteemed Hotel Nikko.

  55. 11:31

    And I'm typically gonna be looking at a map to find places near that venue. But if I could just, uh, extract the, the walking minutes to the hotel and put that right on the page, that'd be really helpful.

  56. 11:43

    And maybe if that's a little bit far, I can figure out what is the closest BART station, uh, to the listing and then add the walk to BART there as well as a backup way to get to the hotel.

  57. 11:55

    Another thing that's really important to me is the Wi-Fi speed. Um, I know I'm going to be working on my talk the night before, true story [laughs], so I'm gonna need really fast internet.

  58. 12:06

    So I can use AI to pull out the relevant reviews and summarize them as positive or negative to really quickly judge whether the Wi-Fi is gonna work or not.

  59. 12:15

    Additionally, usually Airbnb has like fifty vanity photos for any given listing, and I really just want one photo of the bedroom or living room or kitchen. Um, so if I could just pull those out and put them on the page, that would help me a lot.

  60. 12:31

    And then most importantly, at this higher zoom level, preserving the ability to act on this information. So directly from this page, I can go ahead and reserve this listing or send a message to the host without going back to Airbnb.

  61. 12:46

    That would be really helpful and keep me in control.

  62. 12:50

    And I never really know whether staying at an Airbnb or a hotel is gonna be a better deal, so typically, I'll also look at hotel listings. And it's pretty great to be able to see that same elevated view no matter which site I'm looking at.

  63. 13:05

    Additionally, if I'm gonna compare the hotel with the Airbnb listing, have these, having these similar views side by side is gonna give me a really easy comparison between the two of them.

  64. 13:17

    But what if I wanted to look at fifty listings? Comparing fifty of these individual views would still be a lot of work.

  65. 13:26

    Zooming out a level, I can look at a spreadsheet for all fifty listings with my deciding factors all laid out for easy comparison. So I can quickly eyeball the distribution for total price, uh, get a sense of how quick the walks are for each of the listings, how many positive Wi-Fi reviews there are.

  66. 13:44

    Importantly, I can still take action on this level. So if I see a listing that's a clear winner, I can go ahead and book it right here instead of going back to Airbnb or Hotels.com.

  67. 13:55

    But sometimes the decision isn't so clear-cut or it's more multifaceted than, uh, having the cheapest or the closest listing.

  68. 14:07

    So if I zoom out another level, each listing has been abstracted into a circle on a scatter plot, and these are colored by the Wi-Fi reviews. You can see the, the cheapest listings on the left of this plot with the most expensive ones on the right, and the closest ones to the hotel near the bottom.

  69. 14:25

    And I can pretty quickly see that there's this cluster of listings, uh, that are the cheapest and the closest, and they also have good Wi-Fi. But I just realized my flight gets in at nine AM.

  70. 14:38

    But thankfully, I can still initiate actions from this view. So I can circle these, send a message to all the listings when this-- within this cluster, asking about their policy on early check-ins, and whichever one responds first that I can check in at eleven AM, I'm gonna go ahead and book.

  71. 14:56

    All right, so as we saw, there's so many tasks that are best suited by a specific zoom level.

  72. 15:01

    And what we're currently doing is we're manually abstracting that information in our heads. So in this example, digging through fifty different Airbnb or hotel listings, uh, we're keeping all of the previous ones in our heads to try to find the best one, and this takes a lot of mental energy.

  73. 15:19

    I know I titled my talk Climbing the Ladder of Abstraction. [laughs] That was partially to not rip off Bret Victor, who has a talk titled Up and Down the Ladder of Abstraction.

  74. 15:30

    It's a great talk. But I'm not trying to argue that higher levels are better. Instead, what I'm trying to argue is that we can use AI to generate these different levels, glue them together, and make it easy to move between them.

  75. 15:44

    And I think this could completely change the way that we work with information.

  76. 15:50

    So this is one of the many great explorations we're doing at Adept to make all computer work easier. We're gonna have a lot more to share in the near future.

  77. 15:57

    Stay tuned. Um, and then to sum up, there's three things that I would love for you to take away from this talk. The first is augmenting tasks are gonna look a lot like automating smaller, tedious parts.

  78. 16:09

    No one's thinking of spreadsheets as taking people's jobs, and digital spreadsheets is exactly the kind of innovation that I wanna see in the next few years. Secondly, we often think about information at different levels of abstraction, and let's make this easier by using AI to generate and act on these different levels.

  79. 16:25

    And then lastly, this is the kind of thinking we're doing at Adept. Uh, feel free to follow us, um, or, uh, follow along, check in, uh, and we're at Adept.ai.

  80. 16:38

    All right, thanks for listening. [audience clapping] [upbeat music]