AI Engineer World's Fair 2024
Breaking AI’s 1 Gigahertz Barrier
Read the talk
Breaking AI’s 1 Gigahertz Barrier
Faster inference can change more than response time: it can make research, multimodal interfaces and continuously running agents practical parts of everyday computing.
From a talk by Sunny Madra
When speed changes the computer
What changes when a computer crosses a speed threshold? Sunny Madra opens with Intel’s 1999 gigahertz milestone, roughly twenty-five years before this talk. The milestone was a technology demonstration, not a commercial processor launch. The slide’s Intel news release gives the analogy a concrete starting point: a memorable number that invites a larger question about what faster hardware makes possible.
Madra connects that milestone to the later shift toward multicore processors. In his historical comparison, microprocessor speeds improved by three orders of magnitude in about two decades. He then asks what a comparable trajectory would mean for LLMs, invoking Jensen Huang’s view that AI innovation is moving beyond the Moore’s-law curve. The useful question is architectural: when does a faster component change how the whole system is designed?
For a nearer-term example, Madra reports that Groq increased the speed of Llama 3 8B by over 50% between April and June 2024. Groq’s April launch account provides contemporary context, but the talk supplies neither absolute throughput endpoints nor workload conditions for this improvement. Faster inference matters because it expands the amount of work an application can perform before the user has to wait.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A pizza trip instead of a hundred tabs
Input processing is one half of that opportunity. Madra sketches an input rate of 10,000 tokens per second and mentions roughly one-third of a second of processing. He does not specify the input length, so those figures do not establish an end-to-end task time. His point is that a model can ingest and analyze information much faster than a person can read through it.
The concrete example is Globe.Engineer. Give it a task such as planning a trip to New York to try the best pizza, and it breaks the request into several research streams. Madra describes it working against the live Internet: finding flights, taxi options, hotels and food, then assembling an itinerary. The displayed result gathers those branches into one interface, with a navigation tree, travel text and flight cards.
Madra estimates that the Globe.Engineer trip-planning example completes in perhaps less than five seconds. Doing the same research himself might involve tens or even hundreds of browser tabs, each representing a separate line of inquiry. The interface consolidates that work because the model can both process incoming material quickly and generate the resulting plan quickly. Faster output alone would leave the research bottleneck intact; the example depends on both sides of token processing.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Industrializing digital work
Once an application can absorb those research streams, the LLM begins to look less like an isolated feature and more like a computing core—or even an operating system. That would change how people program computers and what they expect computers to analyze. Madra frames this as an expansion of human capabilities, setting aside the question of AGI to focus on a different computing paradigm.
His analogy is the move from craft production to industrial production. He groups carmaking, farming and clothing together: a few cars made in a day versus hundreds or thousands, food produced for a village versus national distribution, and a sweater made in a day or a week versus clothing manufactured at scale. These are broad illustrations of increasing productive capacity, rather than a precise chronology of the Industrial Revolution.
For computing, Madra borrows a framework from Paul Maritz, whose career included Microsoft, VMware and Pivotal, where the two met. Earlier eras changed the representation, connectivity or location of work:
| Computing era | What changed | Familiar expression |
|---|---|---|
| Digitization | Paper processes became digital | Files, folders, inbox, outbox |
| Internet | Digital processes became connected | Networked information and services |
| Cloud and mobile | Work moved to new infrastructure and devices | Cloud scale and phone access |
| AI | Digital production can become automated at scale | Generating and analyzing artifacts |
The paper metaphors survive in operating systems because the first transformation largely reproduced existing office processes. Connectivity followed, then what Madra describes as roughly fifteen years of cloud and mobile form-factor changes. AI introduces the possibility of scaling the production of digital work itself.
Presentation artwork makes the distinction tangible. Madra’s illustrative comparison contrasts a designer making one or two artifacts a day, eighteen to twenty-four months earlier, with Midjourney producing a thousand in a minute. Those are his scale illustrations, not established production measurements. The shift he is pointing to is from commissioning individual artifacts to generating many candidates on demand.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
An LLM at the core
How fast would inference need to become before it could occupy that central role? Madra proposes a hypothetical decision time of 0.1 milliseconds, equivalent to 10,000 sequential decisions per second:
This is a thought experiment about complex decisions, not a demonstrated inference result. At that scale, repeated model calls could become part of ordinary software execution rather than conspicuous pauses in an interaction.
The resulting change would reach how software is built, how it runs and how it scales. Present-day latency makes it easy to assume that an LLM must remain outside the computer’s inner loop. Imagining CPU-like speed growth removes that assumption. Madra credits Andrej Karpathy for a diagram placing an LLM at the center of connections to video, audio, browsers, other LLMs, code interpreters and file systems. The model becomes a coordinator across capabilities that currently have separate interfaces.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Personal interfaces need more than fast speech
The first visible change would be a move from responses arriving around reading speed to decisions that feel instantaneous. Returning to the Globe example, Madra contrasts a few seconds of automated research with an afternoon, an evening or several evenings of personal effort. The benefit is not just reading the same answer sooner; it is completing a larger task within the time previously spent waiting for a small part of it.
Personalization adds another requirement. OpenAI’s early memory features provide Madra’s starting point: a system can remember the names of a user’s pets, children or spouse. He cites Bill Gurley and Brad Gerstner’s podcast discussions of personalization as a major frontier. For those details to support a seamless experience, the system must retrieve and use them quickly enough that personal context does not become another source of friction.
The interface then moves beyond keyboards, pointing devices and mobile touch. At the time of the talk, GPT-4o voice demonstrations showed Madra the possibilities of mixed interaction, although he did not regard the full experience as achieved. His shorthand is XRX: any type of input, reasoning, and any type of output. Input and output need not use the same modality.
Consider booking a haircut. A user can ask aloud which appointments are available, but a spoken list is difficult to hold in memory:
| Part of the interaction | Useful modality | Content |
|---|---|---|
| Request | Voice | Ask for available haircut appointments |
| Response | Text | 9 AM, 11 AM, 3:30, 5:30 |
Displaying the options lets the user compare them without remembering a sequence of spoken times. No appointment has been booked at this point; the system has only presented availability. A natural interface chooses the output that helps the user act, rather than simply mirroring the input.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Spend the speed on more reasoning
Complex task scheduling requires more than a fluid interface. Madra expected increased work on advanced assistants from model providers in the second half of 2024. He connects that prospect to evaluation: LLMs were commonly assessed through a single attempt, and he sees performance constraints as part of the reason. A system that has time for only one response has fewer opportunities to reconsider its work.
Faster inference creates room for repeated attempts and collaboration. Madra refers to unnamed papers in which repeated reasoning or multiple agents let models with fewer parameters compete with larger models. The talk does not identify the papers, models or evaluation conditions, so it does not establish a particular benchmark comparison. The mechanism is nevertheless clear: additional inference can be spent on solving the problem instead of merely returning the first answer faster.
He treats Apple AI’s on-device and off-device interaction as an early example of systems cooperating across compute locations, while anticipating more sophisticated collaboration. That leads into predictive analytics: instead of waiting for a person to ask a question or trigger an action, an agent could keep working in the background.
That continuous operation depends on an economic assumption as much as a latency assumption. Madra imagines compute cycles becoming nearly free, explicitly describing a future condition rather than the situation at the talk. The same constraint applies to context: a large context window does not make processing its contents free. Lower costs could make both persistent agents and richer context practical, allowing a system to notice changes before a user asks about them.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Custom stories and continuous analysis
Creative tools offer a personal version of this expanded capacity. Madra enjoys asking an LLM to write a Seinfeld episode around modern events. What interests him is the model’s assignment of situations to characters: it can identify who would experience the awkward or comic part of a scenario. Extending that customization from written episodes into multimedia is the next possibility he raises.
Business analysis provides a more operational example. Before Groq acquired Definitive Intelligence, Madra’s team worked beyond text-to-SQL on Pioneer, an automated data-science agent developed with his colleague Rick. Pioneer was intended to work on a problem almost indefinitely. Its starting point was a business KPI and the data used to understand it.
Ordinarily, people compare incoming data with KPIs, investigate changes, and turn the analysis into spreadsheets or PowerPoint presentations for management. Pioneer’s proposed loop makes the investigation continuous:
- Define the KPI the business wants to understand.
- Examine new data as it arrives.
- Ask additional questions prompted by the findings.
- Investigate those questions and continue the analysis.
The distinguishing feature is the ongoing inquiry, rather than a single translation from a natural-language question into SQL.
Madra recalls applying Pioneer to a dataset of workers and their performance reviews. It surfaced a relationship involving age, the type of review received and subsequent productivity decline. He qualifies the recollection and invites Rick to correct it; the talk gives no ages, effect sizes or causal evidence. The anecdote illustrates the kind of unexpected correlation a continuing investigation might surface, not a validated rule for employment decisions.
The same idea extends to dynamic optimization. Drawing on experience at Ford after its acquisition of Autonomic, Madra points to vehicle production and shipping: sophisticated supply-chain software can still leave substantial inefficiencies. Faster, continuing analysis might help those systems adapt. He presents this as an opportunity for former Ford colleagues, not a deployed optimization result.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Spare compute for work that can wait
Not every useful agent needs an immediate response. In discussing edge and decentralized AI, Madra names hyperspace.ai as a project that makes unused GPU capacity available to others, drawing comparisons with SETI@home and Render. The historical domain should not be treated as current documentation for that network: today it presents an AI tools service, and its connection to the described project is unresolved.
The workload distinction matters more than the project name. Tasks without strict real-time requirements can use distributed spare capacity even when interactive tasks need predictable latency. Madra connects this possibility to improving throughput and latency in existing systems, then raises power consumption as another reason to consider distributing the work. He suggests a potentially useful arrangement, without quantifying energy savings.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A convincing voice still needs authentication
Speed also matters on the defensive side. Madra describes a colleague receiving a phishing call from someone who sounded formal and possessed extensive personal information. AI, he argues, can help scammers construct deeper, more convincing narratives than the scripts associated with older call-center scams. That creates a need for protective systems operating on the recipient’s side of the interaction.
The colleague’s decisive test was to request a formal message through the HSBC app. The caller could not provide one. That shifted the interaction from judging the plausibility of a story to asking for confirmation through a trusted channel. As voice cloning improves and more personal information becomes available online, Madra expects protective assistance to need very low latency: it must help while the suspicious conversation is still happening.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Tutoring around the learner’s interests
Education brings speed, cost and personalization together. Madra describes cheaper, more widely available tokens as a priority at Groq, then invokes Sal Khan’s discussion of personalized tutoring and the two-sigma result. The relevant historical finding in Bloom’s two-sigma paper is an average achievement difference of about two standard deviations between human tutoring with formative testing and corrective feedback and conventional classroom instruction. The cited studies involved probability and cartography in grades 4, 5 and 8 over eleven instructional periods across three weeks. That is narrower than a guaranteed improvement for every student, and it is not a measured result for AI tutors.
The application opportunity is to make individual attention affordable and adapt explanations to a learner’s interests. Madra describes a homeschooling service that frames arithmetic around unicorns and ponies for a child who likes them. Addition, subtraction and multiplication retain their mathematical content, but the surrounding story becomes personally engaging. His playful example combines three ponies and two unicorns in a multiplication prompt. The useful design principle is to customize the context in which a concept is taught, so the learner has a reason to stay with it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The cost of making software work together
The final application is less visible than a voice assistant or personalized tutor: enterprise integration. Madra argues that most enterprise software deployment and maintenance spending is tied to interconnectivity, interoperability and compatibility. He offers no spending breakdown, but identifies a familiar source of work: getting systems to cooperate and keeping them working together.
Fast, inexpensive AI could reduce that burden. The closing prospect is economic as well as technical: if more interpretation and coordination can happen cheaply enough, organizations may spend less effort maintaining the connections between their tools. Crossing the speed barrier would matter here because it changes which work software can afford to do continuously.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Further reading
An April 2024 account of Llama 3 availability on Groq, with contemporary throughput and latency figures.
The 1984 paper describing human tutoring results and the search for group instruction methods with comparable effectiveness.
A layered approach to combining model responses, with benchmark results, latency tradeoffs and links to an implementation.
Read the complete timestamped transcript
- 0:00
[on-hold music] What we really wanted to pay homage to today is, um, actually, you know, just twenty-five years ago, we crossed the one gigahertz speed barrier, uh, in microprocessors.
- 0:24
What's really crazy is, um, when, when we started thinking about this talk, I actually thought it happened a lot before nineteen ninety-nine. Uh, and I just kinda remember my own, uh, arc of, um, getting involved with computers.
- 0:36
But really, it was nineteen ninety-nine. I had to kinda double and triple-check it. This is the exact press release when, uh, Intel broke the one gigahertz speed barrier. And obviously, that was interesting, you know, for a couple of perspectives.
- 0:49
One, it was this, you know, really big number and moment, but two, it was really after this that, um, you know, Intel started to change about how they think about processors would be used, and they went for, I guess, you know, multicores and things like that.
- 1:02
And, and it's really something that we need to think about w- in terms of what's gonna happen with LLMs. And, and really, if you go back to the, the rate of increase, it only took, uh, you know, about two decades to get three u- orders of magnitude speed improvement in, in microprocessors.
- 1:19
And so if we s- take a step now and look at where we are with LLMs and we think about anywhere close to the speed of innovation, and in fact, you know, what we hear a lot of people talk about, um, you know, including Jensen, is that we're beyond the sort of curve of Moore’s law.
- 1:33
So we're actually innovating even faster than that in, in LLMs today. Um, you know, just to look at what we've been able to do at Groq, just in a short amount of time, um, s- you know, this is between April and June of this year, you know, we were able to increase the speed of Llama Three, uh,
- 1:51
eight B by over fifty percent. And so, uh, the improvements that are happening in this area are really, really quick and, and super exciting, and we're really kinda keen to kind of dive into what could happen here.
- 2:05
Um, and so let, let's think about, like, the state-of-the-art, right? And so, um, you know, there's models today that, you know, we can process and others can process that's a huge inputs.
- 2:15
Say, on the equivalent of, you know, ten thousand input tokens per second, which gets you down to, say, a third of a second across, you know, processing all of those.
- 2:23
And when you do that, you actually end up with these capabilities, um, from a, you know, speed perspective that far exceed human capabilities for both integrating and analyzing information, and it's happening, um, you know, really, really fast.
- 2:37
The example I like to talk about here, um, and I don't know if you've used this, but I highly recommend it. It's this, um, you know, really cool service called Globe.Engineer.
- 2:48
And what it does is you give it a task, uh-- or, you know, so it says I'm here. And I think the example I use here, "Help me plan a trip to New York to try, you know, the best pizza," or something like that.
- 2:58
And what it will do is it, and, you know, I, I couldn't even capture the whole screen here, but it'll basically figure out all the different elements that have to happen, and it's doing this live online.
- 3:07
It's connected to the Internet. So everything from the flights to the taxi options to the hotel options and then the food options and then itinerary and how I can do it.
- 3:17
And it, you know, it does it all in, you know, maybe less than five seconds. And if you think about what's really happening there, and I like to, you know, think about when I tr- uh, plan for trips myself, I end up basically opening, you know, tens to sometimes even hundreds of tabs, and those tabs each have,
- 3:34
like, a, like, a research stream happening for me. And now all of that is solved in, like, you know, a simple interface, you know, really enabled by these LLMs being able to, one, input, uh, process tokens, input tokens faster, and then ultimately output tokens faster.
- 3:49
And it's really giving us, um, a huge edge up in how we operate as humans. Um, and you know, where does this all go? Like, if we start thinking about, um, you know, human superintelligence, uh, and optimizing and accelerating models, it really takes us to, like, interesting paradigms here.
- 4:07
And, you know, we'll talk about this more in a second, but, like, the, you know, the, the high-level way to think about it is, what if an LLM, you know, really becomes either, like, an operating system or, like, the core of, you know, how we think about compute today, and we com- we think about it completely differently
- 4:23
than any of the approaches that we've had before. Um, you know, the way we program these things, the way our expectations are over how they analyze things. And so we're really...
- 4:32
You know, that's interesting in terms of where this is going in terms of superintelligence and staying away from AGI, but more about changing the paradigm from where we are today.
- 4:43
And, you know, the thing that crosses my mind here is what happened in the Industrial Revolution. You know, if we think about three industries, let's think about making food, making cars, and making clothes.
- 4:55
All of those before the Industrial Revolution were bespoke, right? So you'd have, you know, people that would make one or two cars a day. You'd have people work on farms that could, you know, maybe farm for less than a city, even a small village, or someone that was making sweaters could, you know, make them, you know, one,
- 5:11
one a day or maybe even one a week. And when we had the Industrial Revolution show up, we basically had this ability to make hundreds or thousands of cars a day, food, farming at a scale that could be national, uh, clothing that could be made at national scale.
- 5:24
And we're really... You know, we haven't had that in technology. Um, the arc of technology has been, um... And this isn't my own framework. It comes from Paul Maritz, uh, you know, who was a, a longtime Microsoft guy and then, uh, VMware and then Pivotal, where, where he and I met.
- 5:41
Um, you know, he said, "The first era of computing was just taking paper processes and making them digital." And he goes, "That's evident in the way if you think about how the operating system is structured: files, folders, inbox, outbox.
- 5:54
Those are all paper processes that got turned into, you know, digital processes. The next era for us was basically making those things connected," right? That's the Internet era. And what we've been through now, you know, maybe in the last fifteen years, is form factor changes, right?
- 6:09
Either pushing things into the cloud for scale or mobile so you can do it on your phone. But finally, with AI, we're, we're starting to get to a place where we have the industrialization in the same way we saw for those, you know, manufacturing and physical industries.
- 6:22
We see that for technology. So, you know- 18 or maybe 24 months ago, if you needed to have a, um, a Photoshop made of some kind of artifact that you were going to put in a presentation, you'd go to your designer, and maybe the designer would make one or two a day for you.
- 6:39
Now you can go to Midjourney and get a thousand made in the next minute if you want to. So we're going through that same kind of industrialization for tech- technology.
- 6:47
And if we just dive in deeper here into, you know, where we go as we can get into like ten thousand complex decisions per second just by getting this down to, you know, point one milliseconds.
- 6:57
And then if we, if we really, really kind of start increasing that, it does become viable to think about the core of our computing becoming an LLM. And I think this is a real challenge for a lot of people because we, you know, obviously we have existing paradigms that we're really, really locked into.
- 7:13
But this paradigm shift is fundamentally different in terms of how software will be built, how software will run, and how software will scale. And we don't think about it too much today because we think about the speed associated with, um, you know, running LLMs and their capabilities.
- 7:30
But if we can imagine the same, uh, growth that we saw in CPUs happen in this era, um, we can imagine that the, the core of these devices change to become, you know, something...
- 7:42
And this is, again, hat tip to Karpathy. This is a diagram that he drew. But we can imagine an LLM being a core at, you know, whether what happens in video and audio.
- 7:50
We're starting to see that today, what happens in our browsers, how we interact with other LLMs, how we interact with, you know, code interpreters and even our file systems and how we interact with those type of things.
- 8:02
And so what is the art of possible if we start doing this? And so, uh, I'll just kind of rattle off some things here that, you know, crossed our minds as we were putting this presentation together.
- 8:12
Um, you know, we really don't spend a lot of time thinking about it, but many responses today, um, in LLMs are, are sort of near real time. They're at sort of reading speed.
- 8:23
But if we go to like instantaneous responses and decision-making, this becomes a lot faster. Again, this is really evident when you think about something like that globe example I showed.
- 8:32
What you're really able to do there is take a task that would probably take you either an afternoon or evening or a number of evenings, and it's done in just a few seconds for you.
- 8:41
Um, and then there's personalized experiences. You know, today we don't really have a lot of personalized experiences happening. We're starting to see elements of it. You know, I think, uh, OpenAI has started to launch a number of features that allow it to understand, you know, specifics of your world.
- 8:56
It could be your pets' names or kids' names or spouses' names. But really, I think, you know, where this goes to, and a lot of people push on this.
- 9:03
I know, you know, two of my friends, uh, you know, Bill Gurley and Brad Gerstner, they talk about this a lot on their pod, where they really view personalization as the next major frontier.
- 9:13
And personalization and speed are going to go hand in hand if we're going to make that work kind of seamlessly for folks. I think next is kind of a, a universal natural language processing.
- 9:24
And so if we think about our interface today to software, it's, you know, it-- we, you know, we started with sort of point and click and keyboards. Uh, we've gone to touch with our, you know, mobile devices.
- 9:37
But really, you know, you start to see the power of this and, you know, I think everyone's been super excited for the release of GPT-4o, uh, the voice agents.
- 9:45
We, we-- I don't think we've fully got there yet, but I think we've showed the art of the possible there with what they were able to do with voice and then that kind of mixed interaction.
- 9:54
I would say like, you know, we refer to it as sort of like XRX, where it's like any type of input reasoning and any type of output. Um, you know, the example I like to tell people there, if you're trying to order something, you may want to interact with an agent in voice, but you may want to
- 10:08
see the responses in text. And so think about if you're trying to book your haircut and you want to say, "Well, tell me what times are available." And then, you know, it tells you, "Well, there's nine AM and eleven AM and three thirty and five thirty."
- 10:20
That's hard to remember if it's just coming back to you in voice. So you want to basically have these interactions that are multimodal and kind of touches on my second point there.
- 10:28
And I think we're going to start to see a lot more of those, uh, interface changes as well.
- 10:34
Um, you know, advanced virtual assistants, this is like complex t-task scheduling. I think a lot of what we'll see in the back half of just this year is, uh, you know, agents start to become much more, uh, complex and a lot of focus from LLM providers as well, I think on making, uh, you know, complex tasks something
- 10:52
that are solved. It's, it's interesting today because we measure the efficacy of a LLM through generally single shot. And I think we do that because of, you know, going back to that where we, you know, the start of the conversation, which is the performance barrier.
- 11:07
But naturally, if you even take any existing LLM today and multi-shot it, its scores get a lot better. And there was a couple papers that came out recently that showed if you just had multiple agents working together on a problem, they can far ex-- of, of a less, uh, ca-- you know, less parameter model, they can compete
- 11:24
with higher parameter models just by doing sort of multi-shot reasoning or working together. And so I think we'll see a lot more of that as the speed improves, and I think there's, there's an incredible, incredible optionality there.
- 11:35
You know, we saw the first, um, I think first cut of collaborative AI agents with Apple AI, you know, where you see something maybe running on device, interacting with something off device.
- 11:46
It's-- I think it's a very early implementation, and I think these things will get much more sophisticated and better. Um, an area, you know, we've spent a lot of time within our career is like analytics and predictive analytics.
- 11:57
I think today everything is, uh, you know, pretty much action-oriented and derived off a human action. So I think if we get to a place where the speed goes up, it can be a lot more predictive.
- 12:07
You know, what does that really mean? It's just an agent that's always running in the background because the compute cycles are next to free. We don't s-see that today, but I think we get there as we get, you know, higher up the curve.
- 12:17
Um, you know, context-aware as well. And today we, again, we are generally limited to how much context we can provide, and we're having to, even with, with models with bigger context windows, we still have to, you know, be conscious of, you know, how much compute cycles we're going to use.
- 12:31
But I think if that becomes next to free, becomes quite powerful for us.
- 12:36
Um, you know, creative tools and, and customizable content. Uh, I'll focus on the second one here. This is, this is an area where I think many of us would, would like to see things go.
- 12:48
You know, the example I always like to-- You know, one of my favorite shows was Seinfeld, and obviously, you know, it's not on anymore. But one of the things I like to do, uh, you know, when I, when I'm bored is go into, you know, LLM of choice and have it write a Seinfeld episode, but made up
- 13:03
of, like, modern-day things that are happening. And if you ever try that, it's super fun because it does an incredible job of, you know, identifying which character in those scenarios that you give it would, would have, uh, you know, sort of the funny or odd thing happen to them.
- 13:18
And so, the idea of, you know, taking that beyond sort of writing and taking that to multimedia forms is, is gonna be really, you know, really, really powerful going forward.
- 13:27
Um, you know, complex decision-making. You know, before our company was acquired by Groq, uh, you know, we were building a company called Definitive Intelligence, so we spent a lot of ar-- a lot of time in this space, um, not only, uh, doing sort of, say, natural language to re-- uh, you know, analysis of, of SQL, um, right,
- 13:46
text-to-SQL, as a lot of people would call it. But, um, you know, Rick, who's sitting here with us, like, you know, he was working on this really cool product for us called Pioneer, which was a automated data science agent where it's really meant to run almost endlessly on a problem.
- 14:01
And, uh, you know, you sort of define a KPI. And if you think about how a business runs, a business has a bunch of KPIs, and then a business has a bunch of data that's coming in, and then usually humans are taking that data and analyzing it to KPIs and creating PowerPoints and spreadsheets and telling either senior
- 14:18
management or the world how well they're doing. Well, there's no reason that just shouldn't happen automatically, right? And where there's an agent just constantly, you know, looking at the new data that's coming in, asking additional questions, diving into it.
- 14:29
And I think we had a lot of interesting things emerge. You know, we had let Pioneer loose on a dataset of human workers and their performance reviews. And one of the things that we saw was it was able to correlate really interesting, uh, things that we couldn't think about in terms of, you know, depending on your age
- 14:51
and depending on your performance review, it really affected your, um, I guess your output, your productivity. And so it was able to kind of discover that if you're of a certain age and you got a certain type of [chuckles] performance review, your productivity would fall off.
- 15:06
And maybe Rick can correct me if I'm wrong later, but it was something along those lines, uh, which I-- It was always an interesting example for us. And then obviously, uh, a lot of, you know, um, really interesting things around dynamic optimization.
- 15:18
Um, you know, this, this an area m- we're familiar from before. Um, you know, when a bunch of us were at Ford, um, after the acquisition of Autonomic, we really saw, you know, for the supply chain, if you think about how, you know, cars are produced and how they're shipped, um, you know, there's, you know, pretty sophisticated
- 15:36
software that does this, but it's still not efficient, right? And I think, um, you know, the art of the possible with sort of what we were talking about earlier could be very, very interesting for some of our old, old colleagues at Ford.
- 15:49
Um, I'll touch on a couple more things and then leave a couple minutes for questions if there's any. But edge AI and decentralized AI, this is pretty cool. Um, you know, there's a, a really cool project called, you know, hyperspace.ai.
- 16:02
What they're doing is, um, they actually have a lot of, uh, you know, taking, you know, sort of like SETI@home or even Render and where they're basically allowing people to take their unused GPU compute and make it available in the cloud, uh, or I guess, yeah.
- 16:19
And, um, and why that's interesting is there's certain use cases that necessarily don't require something to be real-time, and so I think we'll see a lot more of that.
- 16:27
Now, this intersects really well with us getting more throughput and getting lower latency out of existing systems, so I think we'll see a lot more of that as well, especially 'cause the amount of power consumption that's required.
- 16:39
If you distributed that, it could be really interesting. Um, and a couple more here is, uh, enhanced security and privacy. This is a big area. You know, I was, I was talking, uh, to one of our colleagues last night, and he was subject to a really, really scary type of, um, I guess maybe phishing call where, um,
- 17:00
you know, someone had called in, um, sounded very formal, uh, and had a lo- access to a lot of his information. Now, you, you know, we've all seen, um, there's, you know, these kind of, uh, uh, people that run scam call centers and people that go and attack them.
- 17:14
But tho- these folks armed with AI are much more sophisticated because they can create stories and narratives that are much deeper than sort of the call center worker of past.
- 17:24
And, uh, now I think in order to protect against these systems, you'll almost need to have something on your side, um, so that you can, you know, you can think about it.
- 17:33
You know, with our colleague, he was just so confused because the narrative was so good. The only way he could really figure out that this person was a scammer other than hanging up on them was saying, "Hey, well, send me some kind of formal message through the HSBC app, uh, and then, then I'll know it's you."
- 17:47
And, and, you know, the person wasn't able to do that. And so I do think, um, you know, as voice cloning, as more of our information is online, we have to be really careful, and we'll, we'll need these protective systems that we can use, um, and we need them to run incredibly fast.
- 18:03
And so, um, and I think this is the last set of them here is, uh, you know, education is, is something that's really important to us, you know, broadly at Groq.
- 18:11
We, we think about this, and we think about, you know, making tokens available cheaper and more broadly, um, and being able to personalize. You know, Sal Khan has a very good TED Talk from a couple years ago where he really highlights, um, you know, it's the two sigma talk, and he says you can take any student at
- 18:26
any level, the highest levels or even someone performing lower, and if you give them a personalized tutor, they can improve their test scores two standard deviations. And so imagine doing that, you know, obviously with AIs that are, um, you know, can be, one, very cheap to use and that can be personalized to their learning experience.
- 18:43
Um, you know, I was speaking to someone recently who was building, uh, an AI service for homeschooling, and what was, what was powerful about that particular service is, let's say you have a, a young child, and they're really into unicorns or ponies, and you wanna teach them about, you know, math and so, you know, math, subtraction, addition,
- 19:02
multiplication. It's a lot easier if you frame it in the context of those things. Hey, you know, you have three ponies times two unicorns, and what, what do you get from it? [chuckles]
- 19:10
And so I never thought about that before, but for learning and customizing that for the interest of the person is quite powerful, so we'll see more of that. Um, and then you just, interoperability and compatibility, right?
- 19:21
I think this is an area, if you've ever been in enterprise software, the majority of money spent in deploying and maintaining enterprise software is really related to, you know, interconnectivity and interoperability and compatibility.
- 19:36
And so, um, you know, having really fast and cheap, um, you know, AI technologies will help us [chuckles] really reduce a huge burden that exists on the enterprise today. So, um, that's it.
- 19:48
Hopefully, you guys enjoyed that. [audience applauding] [upbeat music]