AI Engineer World's Fair 2026
Physical AI's Next Bottleneck Is Finding the Right Video — Rafael Levi, Bright Data
Read the talk
Physical AI’s Next Bottleneck Is Finding the Right Video
Rafael Levi of Bright Data walks through the gap between abundant web video and useful robot training data, then demonstrates a workflow that searches for actions before collecting clips.
From a talk by Rafael Levi
At a glance
Ideas worth remembering
Collecting physical-action data has several constraints: instructed recordings can change behavior, simulation raises questions about physical fidelity, and teleoperation requires operator time.
Frame-to-frame changes can supply motion information, but relevant human video still needs processing before it becomes useful robot training material.
Bright Data’s proposed workflow searches an action index before collection, returning trimmed clips so customers can avoid downloading entire videos to find a useful moment.
Action search can find incidental events that titles miss, from dishwashing and brand appearances to driving maneuvers and game-level completions.
A robot needs examples of the world changing
A system that can see, understand and act needs examples of actions and their consequences. Rafael Levi, who works on data collection at Bright Data, opens with that requirement: improving the AI still leaves the question of what to feed it. His shorthand is memorable: “AI without data is just a box.” The presentation focuses on finding video that can help train robots and world models.
The opening history traces a progression from robots learning through images and real-world actions in 2022, to AI controlling multiple robots in 2023, to wider access through open source in 2024. As more labs gain access to models, collecting suitable experiences becomes a practical constraint. Levi also distinguishes the excitement around humanoids from their autonomy: many of the robots people encounter are still remotely controlled.
The scale comparison explains why he turns to video discovery. Levi puts language-model training material at trillions of words, image training material at billions of labeled images, and existing robot-action video at only about a million videos. These are his broad estimates rather than a like-for-like dataset census, but the distinction matters: a large supply of text or pictures does not automatically supply sequences showing how objects move when someone acts on them.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Why recording more actions has costs of its own
Opening a door seems like an easy action to collect. Pay someone to record it, then add the video to a training set. The complication is that the instruction changes the circumstances: demonstrating how to open a door for a camera may produce a different movement from walking into the house during an ordinary day. Camera awareness can change behavior too. Levi calls this instructed data and considers it biased; the recording does not establish how much that difference affects robot performance.
Other collection methods trade different constraints:
- Simulation: Virtual environments are cheap to use, but Levi questions whether their physics is faithful enough for the robot behaviors he wants to train. The attraction is inexpensive experience; the concern is whether that experience carries over to the physical world.
- Teleoperation: A person can control a real robot and record its actions, but collecting more hours requires more operator time. His example of recording eight hours a day makes the scaling problem concrete: a large dataset still depends on many people spending many days operating machines.
- Existing datasets: Reusing a prepared dataset avoids fresh collection, but the available supply is limited in the account he gives.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Web video supplies motion, but still needs a path to robot control
The alternative is to look for actions people have already recorded. Web video contains people handling objects, opening doors and encountering the consequences of motion and gravity. That makes it a potential source of physical experience without commissioning every demonstration. It also creates a discovery problem: useful interactions sit among enormous amounts of unrelated footage.
Levi cites a Meta training result as motivation: about a million hours of real-world video followed by 62 hours of robot data, after which the system could control a real robot. The result is presented without the model identity, task or evaluation conditions, so the useful lesson is narrower than a universal recipe. A large video stage can supply experience that a smaller robot-specific stage then connects to control; the cited example still includes real robot data.
How can a video of a person pouring water help a robot? The explanation starts with change between frames. Compare two images, measure the movement between them, then repeat across the sequence. The changing positions provide information about motion; Levi describes extracting angles and distances from that progression. A still image captures a state, while the sequence exposes how the state changes during an action.
That is a high-level account of extracting motion, not a complete method for turning human movement into robot commands. Levi explicitly acknowledges that downloaded video does not get a training system “100% there”: further processing is necessary. Keep that distinction in mind through the product demonstration. Finding a relevant clip supplies material for the training pipeline; it does not finish the pipeline.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Move action selection ahead of the download
Collecting video first and deciding what is useful afterward can be expensive. Levi reports that NVIDIA discards about 96% of video collected for Cosmos training, and that Stable Video Diffusion discards 74%. He uses a million-hour download as an illustration of the first percentage: only 4% would remain useful. The rates are cited claims rather than independently established measurements here, but they identify the cost at issue. Footage rejected after collection has already consumed bandwidth, storage and processing.
Bright Data’s proposed order is “search first collect second.” Describe the action you need, search an existing video index for that action, then receive the relevant snippets. The intended unit of collection becomes the action clip rather than the entire source video. Queries can describe a person washing dishes or folding a T-shirt, even when those actions are not the subject of the video’s title.
Where does selection happen, and what work remains afterward? The flow below follows the dishwashing query developed in the next demonstration. Searching and trimming happen before the customer collects the training material. Motion and distance processing still follow retrieval. The proposed savings come from collecting a smaller, more relevant slice of video; the index is the infrastructure that makes that earlier selection possible.
Specify the action and visible hand interactions.
The video index moves relevance selection before collection. A trimmed result still needs downstream processing for robot training.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The dishwashing query describes a sequence, not a topic
The demonstration makes the query concrete: a person washing dishes by hand in a sink, removing grease from a plate with a sponge and soap, rinsing, and placing dishes in a drying rack, with close-up hand interactions. Each added detail narrows what useful footage should contain. A video broadly about kitchens is insufficient; the desired material shows the manipulation and the steps around it.
The search then returns clips of people washing dishes. The observable change is from a written description of the desired activity to candidate video snippets containing it. Those snippets may come from videos about cleaning a kitchen or another broader subject. That is the point of indexing actions: an incidental moment can become a search result even when the source video was named for something else.
Levi says the demonstration was recorded with about 100 million indexed videos and that the index had since reached 1.1 billion. The playback is sped up, so it does not establish search latency, and the examples do not quantify retrieval accuracy. A second query asks for a human folding different types of clothes and returns folding footage. Together, the examples illustrate the intended interface: describe a visible action in detail, then collect clips that can feed subsequent training work.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The API returns clips and information about the match
The same workflow is available through an API, so repeated searches need not depend on a person using the demonstration interface. Levi describes returning a video snippet and its URL, along with a link to the original video. The result also carries information about where and how the query matched:
- Timestamp: Locates the relevant moment in the source video.
- Match score: Describes how close the result is to the query.
- Frame count: Reports how many frames concern the requested content.
Brand discovery uses the same gap between a video’s title and its visible contents. Imagine the example Levi gives: a person applies makeup, and a particular brand is visible on the table, while the recording’s title concerns something else. Searching the brand name in titles could miss that appearance. Searching indexed video content is meant to find the moment when the product is present or used. Here the useful result is a brand appearance rather than a robot demonstration.
For training, the intended benefit is less irrelevant footage entering the collection and preparation process. Levi connects noisy data to model problems and describes the product as removing that noise. The demonstrated capability supports a more limited conclusion: action retrieval can select material before bulk collection. Match scores describe query relevance; the talk does not establish that retrieval eliminates all unsuitable training data or downstream model errors.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Search for the event hidden inside the video
Driving footage extends the idea from household manipulation to road events. Dashcam videos can contain someone running a red light, stopping at one, turning left or turning right. These are distinct behaviors within a larger recording, and action search could collect examples of each. Levi proposes them as useful material for self-driving systems and other world models; the recording does not demonstrate a driving model trained on the returned clips.
The ending returns to physics and everyday movement: sitting, walking and moving through the world. Levi’s frustration is with commissioning huge numbers of recordings while relevant actions already exist online. His practical objection is discovery. A title search asks what the uploader called the video; a training-data search needs to ask what happened inside it.
The proposed product brings discovery across public video providers, including YouTube and Vimeo, into one place. Its usefulness can extend beyond training and brand appearances. The final example is a gamer looking for footage of someone beating a particular level, even when no video is titled for that specific event. The same mechanism applies: find the relevant action within a longer recording, collect that portion, and process it for the task at hand.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:01
[music]
- 0:13
Hi everybody.
- 0:15
Welcome.
- 0:16
For those that came to actually see me,
- 0:19
thank you. For those that are just
- 0:21
hanging out here, also thank you. I'm
- 0:23
going to try to entertain you guys,
- 0:25
teach you something new.
- 0:28
Um again my name is Rafael. I work at
- 0:31
Bright Data. I've been with Brighta for
- 0:33
over eight years. And in Bright Data we
- 0:35
are collecting data and uh we're trying
- 0:37
to innovate. So today I want to talk
- 0:40
about video discovery for agentic world
- 0:42
models training. And uh we're going to
- 0:45
explain what that is in a little bit.
- 0:48
So
- 0:49
what does it take to take to build a
- 0:52
system that can see understand and act
- 0:56
right? I mean, we figured out what an AI
- 1:00
is. We're constantly improving it. Like,
- 1:02
you know, that's kind of handled. But
- 1:04
now, the hard part is what data do you
- 1:07
provide to it? Because everybody knows
- 1:10
AI without data is just a box, right?
- 1:15
So, let me take you a little bit back in
- 1:17
history, right? In 2022,
- 1:20
Google taught a robot by images, real
- 1:25
world actions, right? Then going
- 1:28
forward, there was a leap uh in 2023. Uh
- 1:32
we had uh an AI actually control
- 1:35
multiple robots.
- 1:37
In 2024, the open source came out. Now,
- 1:40
this where the game changed. Once we had
- 1:42
an open source, all the labs had access
- 1:45
to AI and they started already putting
- 1:48
AI into robots, right? And so now these
- 1:52
models are already driving humanoid
- 1:55
robots. I'm you seeing some of course
- 1:58
most here are remote controlled, but I
- 2:01
mean in San Francisco you got weo
- 2:03
driving without any drivers. How cool is
- 2:06
that? First time I got in the car, I
- 2:09
thought I was going to die, but it was
- 2:10
actually very nice. Right now, how do
- 2:14
you think it learned how to drive? By
- 2:17
actually teaching itself. By by learning
- 2:20
on dashboard cameras, right? Every car
- 2:23
has a dashboard. So, if you provide that
- 2:25
to an AI, it can figure out how to
- 2:28
actually operate a machine.
- 2:31
So, as I said, the AI is no longer the
- 2:33
hard part. The data is. And this is what
- 2:36
I want to talk about. Right? So for chat
- 2:42
LLMs, there's trillions of words and
- 2:44
texts, right? So we have so much text
- 2:47
data to train the LLMs. For image
- 2:50
generation, we got billions of images,
- 2:53
labeled images. So that's also a huge
- 2:55
database.
- 2:57
But for robotics, there's only about a
- 2:59
million videos. It's a very small
- 3:02
limited data sets of robots doing
- 3:05
things, right?
- 3:09
And so the idea here is that of course
- 3:14
you a lot of companies out there what
- 3:16
they do is they pay people to record
- 3:19
actions, right? So for example, hey
- 3:22
record me how you open a door or record
- 3:25
to me how you're sitting on the chair.
- 3:28
But then the problem becomes is that
- 3:30
when somebody is told to do something,
- 3:33
they do not do it naturally. So I call
- 3:35
it instructed, right? If you're told to
- 3:38
record how you open the door and you're
- 3:40
doing it for somebody, it's not going to
- 3:41
be the same as if you just walk into the
- 3:43
house. The movement is totally
- 3:45
different. So is that data good for
- 3:48
robotic training? In my assumption, no.
- 3:50
It's biased data and it doesn't deliver
- 3:52
the same results. So obviously the robot
- 3:55
is not going to be as effective as if
- 3:57
it's intuitive. Right? So I just threw
- 4:00
in some examples on the slide, right? So
- 4:03
uh
- 4:06
the built by hand data is not the same
- 4:08
as the data that is kind of on the site,
- 4:11
right? So and also the fact is that
- 4:14
people behave very differently when they
- 4:16
on the camera. Some of you know some
- 4:17
people are camera shy. So you know as
- 4:20
soon as you put a camera on them you
- 4:22
they change but in the natural world
- 4:25
it's a totally different thing. So this
- 4:27
is what we are trying to solve in bright
- 4:29
data right
- 4:32
um
- 4:35
again so why is the usual data sources
- 4:37
are not enough right simulations
- 4:39
uh virtual it's it's cheap but it's um
- 4:42
you know everybody's playing video games
- 4:44
the physics in video games are almost
- 4:46
there but they're not good enough where
- 4:48
to train a robot on it right a hand
- 4:51
control the person controlling a robot
- 4:53
it's doable but how many hours a day can
- 4:56
you record the robot, how many people
- 4:58
you need to actually get huge scale
- 5:00
data, right? I mean, you can record
- 5:02
maybe eight hours a day every day. How
- 5:05
many hours are you going to get in a
- 5:06
year? It's not really scalable.
- 5:09
And of course, there's already pre-built
- 5:11
data sets, but as I said before, they're
- 5:13
small. There's only about a million
- 5:14
videos at this point. So, what is the
- 5:16
alternative? The alternative is the web.
- 5:20
Let's just think about YouTube. How many
- 5:22
videos are there on YouTube? I don't
- 5:25
know, five billion videos. How many
- 5:28
actions are in those videos that a robot
- 5:32
can replicate?
- 5:34
I mean, let's think about it. How many
- 5:36
do you think videos are there of a
- 5:37
person opening the door firsterson view?
- 5:42
Millions of hours. You'll be surprised.
- 5:46
So,
- 5:50
every day the web video shows gravity
- 5:52
motions, right? So um cause and effect
- 5:56
obviously accidents big the biggest
- 5:59
cause and effect right how handle how
- 6:01
people handle objects and billions of
- 6:03
hours great training material for robots
- 6:07
but what is the problem the problem is
- 6:09
that there's a lot of noise right oh and
- 6:13
for example right so one of the AI
- 6:15
models that meta trained they gave it
- 6:17
about a million hours of real world
- 6:20
videos
- 6:21
and then all it took is 62 hours of real
- 6:24
robotics data to actually control to
- 6:27
control a real robot. So if you think
- 6:30
about it,
- 6:31
the data was collected publicly, right?
- 6:33
So it trained the the the robot on just
- 6:37
random things and then in 62 robotic
- 6:40
hours, it was already moving. So there
- 6:42
no simulations needed, autonomous
- 6:45
robotics delivered fast.
- 6:52
But the videos don't show how the robots
- 6:54
move, right? So you might think, well,
- 6:56
listen, how useful is it a person
- 6:58
pouring water into a cup for a robot. So
- 7:02
there is methods actually out there how
- 7:04
an AI can distinguish and actually learn
- 7:07
from the actions that are in a video. If
- 7:09
you take two frames frame by frame and
- 7:12
the AI measures the movement this
- 7:14
difference, right? So you have an image
- 7:16
A and image two and there's a small
- 7:18
movement that is changing between the
- 7:20
two images and AI can actually measure
- 7:23
that change. So if you keep doing that
- 7:26
through the whole video and AI can
- 7:28
actually figure out angles, distance and
- 7:31
everything that it needs to train and
- 7:33
robot right. So we don't necessarily
- 7:36
need the sensor's data but of course the
- 7:39
video itself that you download from
- 7:41
video from YouTube or the video that uh
- 7:44
you extract the data from doesn't
- 7:46
necessarily get you 100% there. You do
- 7:48
need to have some processing on top of
- 7:50
that but
- 7:52
it's available. It's out there and all
- 7:55
you need to do is process it.
- 7:58
So a few more examples right? So for
- 8:00
example, Nvidia when they train in the
- 8:03
Cosmos robot, they're throwing out about
- 8:05
96% of the video. So what does that
- 8:08
mean? They download a million hours of
- 8:11
video and the only part is 4% is
- 8:14
actually useful. Everything else gets
- 8:16
thrown out. That's wasted compute.
- 8:18
That's wasted bandwidth. That's wasted
- 8:20
storage. That's just a lot of wasted
- 8:22
money. Okay. Stable video diffusion is
- 8:25
throwing out 74% of the videos that
- 8:28
they're downloading.
- 8:30
So in bright data we are trying to solve
- 8:33
that problem and we're trying to take it
- 8:36
a step forward. So what we propose is
- 8:39
search first collect second. What we do
- 8:42
is we are actually indexing videos and
- 8:45
we are allowing you to search for
- 8:48
specific actions. Since we're talking
- 8:51
about a person opening a door, let's
- 8:52
keep using that uh example. What you can
- 8:56
do on our platform is actually input
- 8:59
detailed information, right? A query,
- 9:02
person washing dishes, person folding a
- 9:04
t-shirt and so on. And we will provide
- 9:08
the snippets from the videos of these
- 9:11
actions.
- 9:14
So what does that mean?
- 9:16
That means that
- 9:19
there's less waste of data collection.
- 9:22
there's less waste of storage and
- 9:24
bandwidth, right? So, a little more on
- 9:26
how it works. Obviously, you define what
- 9:28
you're looking for.
- 9:32
We search through billions and billions
- 9:34
of pre-indexed videos, not by keywords,
- 9:38
but by actions, and we provide you with
- 9:41
ready to use clips.
- 9:44
And obviously, they're already trimmed,
- 9:46
prepared for your training. So all you
- 9:47
got to do is process them a little bit
- 9:49
for maybe motion distance sensors and so
- 9:52
on and it's ready to be ingested into a
- 9:55
robot.
- 9:57
I have a small example of a video of how
- 10:00
it actually works on our platform. So
- 10:02
let me just hit play here if I can find
- 10:05
it.
- 10:07
So here we see person washing dishes by
- 10:09
hand in sink removing grease from the
- 10:11
plate using sponge and soap including
- 10:14
rinse and placing dishes into drying
- 10:16
rack. close-up hand interactions.
- 10:19
So, right now what is happening is it's
- 10:21
searching through
- 10:23
billions of videos. Now, on this video,
- 10:25
it's a bit old. We only got about 100
- 10:27
million index there, but now we're up to
- 10:29
1.1 billion videos. And it will
- 10:31
literally like this is a little sped up,
- 10:33
so I didn't want to waste your guys
- 10:34
time, but it will provide to you clips
- 10:38
from videos of people washing dishes.
- 10:41
And you can see right here all the
- 10:43
videos in
- 10:46
That would be fine. Now, some of these
- 10:48
videos might not have to do anything
- 10:49
with washing dishes. They might have
- 10:51
cleaning kitchen. They might have to do
- 10:53
something else. And here's another
- 10:55
example. For example, human folding
- 10:57
different types of clothes,
- 10:59
right? So, a little more description.
- 11:01
The more description you give, the more
- 11:03
you actually get back.
- 11:06
So, again, a little speed up. Let's see.
- 11:14
I want you to understand it could be
- 11:15
anything. A person putting on makeup.
- 11:17
Let's say you have a brand and you want
- 11:18
to find videos where your brand is being
- 11:20
found. All that can be also provided. So
- 11:24
you see people folding clothes. So now
- 11:27
you can train a robot on how to fold
- 11:29
clothes.
- 11:34
Of course everything is available via
- 11:36
API. We don't expect people to actually
- 11:37
do anything manually, right? So you can
- 11:40
trigger it by API. you get a returned
- 11:43
snippet of the video, the URL. We give
- 11:46
you the link to the original video if
- 11:48
you want to watch it. But the main thing
- 11:50
here is that
- 11:52
we can give you the snippets of the
- 11:54
video to train your robotics. Now, I'm
- 11:58
not sure if any of you training any
- 11:59
robots, so I'm going to give you another
- 12:02
example of how useful it could be. Let's
- 12:05
say you have a brand and you want to
- 12:07
know where your brand is being demoed on
- 12:11
videos.
- 12:13
The video titles doesn't matter because
- 12:15
it doesn't actually consume. It's not
- 12:17
about your product. It's just a person
- 12:20
doing some podcast recording something.
- 12:23
But you see a woman putting on makeup
- 12:25
and she you see on the table that this
- 12:27
is your brand of a makeup. So now all of
- 12:30
a sudden you can find all videos for
- 12:32
your brand where it's being used
- 12:34
whatever it is right by searching name
- 12:37
of the video you won't find it by by
- 12:40
video indexing you can actually find
- 12:42
specific things that you are interested
- 12:44
in apple falling from the tree and so
- 12:48
on. So what comes back? We provide to
- 12:51
you this time stamp. We provide to you
- 12:53
the score matching, how close it is to
- 12:54
your query, and of course the frame
- 12:57
count. How many frames are about what
- 13:00
you're looking for?
- 13:04
So what does this change?
- 13:06
This changes a lot. Okay, less waste.
- 13:09
You don't need to download million of
- 13:11
hours of videos. you can actually get
- 13:13
exactly the the specific actions that
- 13:17
you are interested in.
- 13:21
Again, it's very crucial for robotic
- 13:23
training because noise creates problems,
- 13:27
hallucinations, right? This removes the
- 13:30
noise.
- 13:34
What is it useful for? Not only robots,
- 13:36
but obviously self-driving for example,
- 13:38
Wimo, right? trained on dash cam video
- 13:42
cameras and there's millions, hundreds
- 13:45
of millions of hours on YouTube of dash
- 13:47
cams. How many of you like to watch
- 13:49
accidents? Dash cam c accidents. I'm
- 13:52
sure everybody watched them. Crazy
- 13:54
driving in Russia, right? So, this is
- 13:57
what we're talking about, right? The
- 13:59
videos that are out there, okay? A
- 14:01
person running a red light, person
- 14:03
stopping on red light, making a left
- 14:05
turn, making a right turn, and so on.
- 14:07
All of that can be a very useful data
- 14:10
for training and any other world models.
- 14:13
Physics, right? Robots needs to
- 14:15
understand physics. What is the gravity?
- 14:18
How to sit? How to walk? How to move?
- 14:21
Again, of course, you can pre-record it.
- 14:24
Um, a lot of times right now I speak to
- 14:27
some of the people and they have
- 14:28
millions of people recording videos for
- 14:30
them. Hey, can you guys record me how
- 14:32
you open doors? That's crazy. Why do you
- 14:35
need a million people doing it when
- 14:37
everything is available online? Online
- 14:40
is one of the hugest databases out
- 14:42
there. It's just people don't really
- 14:44
like it's it's not that easy to use it,
- 14:46
right? I mean, right now the only search
- 14:48
available to us is by the keyword of the
- 14:51
name of the video.
- 14:53
And of course, one source, the whole
- 14:56
public video web in one place. I mean,
- 15:00
we got YouTube, we got Vimeo, we got so
- 15:02
many different uh video providers out
- 15:05
there, billions and billions of hours
- 15:09
of training data just there waiting for
- 15:12
you to grab it, collect it and process
- 15:14
it. So basically this is a new product
- 15:16
that we're doing a bright data right
- 15:18
indexing videos so that you guys can
- 15:21
find specific actions anything you can
- 15:24
think about you know again it could be
- 15:26
useful for brands uh as well it's not
- 15:29
only for training data anything you can
- 15:32
think of could be useful
- 15:35
maybe you are a gamer and you want to
- 15:38
know how do you beat this level but
- 15:41
there's no exactly video like that so
- 15:43
you can search okay people beating the
- 15:45
level on the video game, right?
- 15:47
Anything.
- 15:49
The world is yours.
- 15:52
So,
- 15:53
I'm going to wrap it up. I don't know if
- 15:55
you have any questions, but uh if you
- 15:57
want to connect with me, talk about it
- 15:58
more, feel free to to add me on
- 16:00
LinkedIn. And um yeah, I'm going to
- 16:04
thank you guys for your attention. And
- 16:07
uh I'm at a bright data booth if you
- 16:09
want to talk more about it. And thanks
- 16:12
for coming.