Training Taste — Thais Castello Branco, Taste Labs

Read the talk

Training Taste: Measuring Slop and Giving Agents Better Judgment

Thais Castello Branco explains how Taste Labs breaks design into measurable properties, detects repetitive patterns, and combines deliberate variation with brand guidance and verification.

From a talk by Thais Castello Branco

At a glance

Ideas worth remembering

  • Design evaluation becomes more tractable when broad judgments are decomposed into specific properties. Context can increase expert agreement, while small classifiers can detect combinations of patterns associated with slop.

  • Castello Branco reports that website homogenization preceded AI and that AI increased repetition across contexts. Her classifier advantage is also a reported result; the talk supplies no numerical metrics or detailed comparison protocol.

  • Better generation requires both variation and fit. The proposed creativity system departs from selected category rules, while structured brand guidance preserves the choices that should remain consistent.

  • Inference-time interaction and verification complete the system: clarify user intent, guide the agent with structured context, and check whether the output follows it. The immediate goal is a higher minimum quality standard.

Breaking design into problems that can be solved

Selected presentation frame from Training Taste — Thais Castello Branco, Taste Labs at 38 seconds
Breaking design into problems that can be solved

After a brief opening, Thais Castello Branco introduces Taste Labs and its mission to reduce AI slop. Design is the company’s first focus within a broader ambition to improve model capabilities in subjective domains such as design and writing. The engineering challenge is to break a large, fuzzy domain into smaller problems whose failures can be identified and addressed.

One part of that work happens with frontier labs: evaluate models, locate weaknesses, and construct post-training data or reinforcement learning environments to address them. The choice of method depends on the property being improved. Palette selection, contrast, and alignment can become almost deterministic when the task and context are specific enough. Castello Branco qualifies this objectivity as an answer that most experts would agree with. Aesthetics produces more expert disagreement, so she argues for leaning more heavily on data there.

The second part operates at the application layer without changing the underlying model. An application using an off-the-shelf model still has to counter a tendency toward average styles, understand user and brand preferences, and keep outputs consistent with those preferences. Context, judgment, and verification therefore remain application responsibilities even when the model itself improves.

0:150:16
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:01 · section reference included

Easy generation does not make judgment easy

Castello Branco next asks what makes something great. She contrasts tasks with an objective answer against writing and design, where correctness alone cannot define excellence. A poem, a coffee shop, or a website can feel special through a combination of distinctiveness, attention to detail, craft, and authenticity. Her goal for generation is consequently more demanding than reproducing an average example: it includes producing work that deliberately departs from familiar patterns.

Slop offers a more tractable starting point because she believes people can agree more readily about repetition and a sense of lifelessness than about greatness. She also recognizes the benefit of generation: someone without design or engineering expertise can create a PowerPoint, a website, or a web app with a button. As generation becomes extremely cheap, however, the ability to produce an artifact becomes separated from the ability to judge it.

Design judgment develops through prolonged exposure, pattern recognition, a point of view, the courage to depart from convention, and restraint about what to leave out. Expecting every user to acquire that expertise is unrealistic: people do not have the time or skills to develop taste in every domain. Taste Labs instead aims to make it easier for ordinary users to create better work and understand their own preferences.

2:282:32
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:25 · section reference included

Repetition, lack of fit, and low intent

Selected presentation frame from Training Taste — Thais Castello Branco, Taste Labs at 371 seconds
Repetition, lack of fit, and low intent

Slop predates AI, but Castello Branco describes AI as an accelerator because it makes thoughtless generation so easy. She identifies three recurring characteristics, beginning with repetition: seeing the same design choices over and over. The second, lack of fit, concerns whether an artifact feels appropriate for a particular context, moment, and person.

Her concrete example connects those first two characteristics. One user requests a pet shop website and another requests a finance firm website, yet the designs converge. The problem is that the system repeats a pattern across contexts that should influence its choices. Careful design would not necessarily produce the same result for both businesses.

The third characteristic is low intent. Rapid prompting and attempts to obtain a finished result in one shot contribute to it, but she also assigns responsibility to the system. An application can help users understand and express what they want, interpreting intent and adding context before creation. This makes intent clarification part of the product’s work rather than simply a prerequisite users must satisfy on their own.

4:494:51
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:49 · section reference included

Measuring how website design converges

Selected presentation frame from Training Taste — Thais Castello Branco, Taste Labs at 424 seconds
Measuring how website design converges

The research begins with a measurement question: can slop be detected quantitatively? Taste Labs analyzed over 2 million websites from roughly the preceding 10 years to examine changes in design over time. The team also generated a synthetic set of websites to compare human-made sites with AI-generated ones. This supplies both a historical view of design trends and a comparison with generated output.

Her reported finding is that the internet was already becoming more homogeneous before AI, with increasingly similar palettes and layouts. She suggests faster circulation of design trends as a possible explanation, rather than establishing it as a measured cause. With AI, she reports more repetition and similar patterns appearing even across completely different categories. The distinction is between convergence that already existed and generation that repeats patterns with less regard for context.

6:146:16
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:12 · section reference included

Using small classifiers to detect combinations of patterns

Selected presentation frame from Training Taste — Thais Castello Branco, Taste Labs at 501 seconds
Using small classifiers to detect combinations of patterns

Taste Labs turns the website data into structured characteristics through pattern mining. Castello Branco names colors, typography, layout, and audience as dimensions to extract. The team then trains probes: small classifiers, each responsible for spotting one characteristic. This decomposes a broad judgment about a website into narrower detection tasks.

The predictive signal comes especially from combinations of detections and the frequency with which multiple characteristics occur together. The team identifies probes associated with sites judged likely to be AI slop, then uses their combined occurrence to predict slop. A detected characteristic is therefore evidence within a pattern, rather than necessarily a verdict on its own.

Castello Branco reports strong predictive performance and says this approach outperformed most LLM-as-judge methods in the comparison. Those methods ask a language model to distinguish great human-quality work from AI-generated slop. The claim remains qualitative: she supplies no numerical performance metrics, named baselines, or detailed evaluation protocol, so the size and scope of the advantage cannot be established here.

Detection is only a first step. As production becomes cheaper, she argues that judgment becomes more valuable: discerning what is appropriate, decomposing a problem, and understanding how to solve it. She explicitly prefers the word judgment here. Human judgment remains valuable, while tools can address specific pieces of the larger problem.

7:377:39
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

7:37 · section reference included

Creating variation deliberately at inference time

Castello Branco argues that inference-time behavior is at least as important as improving the model. Inference is when the system interacts with its user and exchanges information about context and intent. Better underlying capabilities cannot fully replace that conversation; leaving it unresolved allows slop to persist.

To address repetition, she describes a system provisionally nicknamed the creativity API. It would act as an inspiration machine for an agent, helping it produce designs outside the familiar distribution instead of converging on the same average. The proposed intervention is intentional variation, with the aim of improving quality rather than merely making outputs different.

She distinguishes this from simply increasing model temperature and hoping randomness yields something useful. A startup pitch deck illustrates the intended reasoning: understand what makes a good deck in that category, then choose a few expectations to depart from while preserving others. The tradeoff is novelty with fit. Breaking selected rules can create a distinctive result; ignoring the category altogether can produce something that feels inappropriate. She describes this design principle without specifying the system’s internal implementation.

9:319:35
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:27 · section reference included

Turning a brand into guidance and evaluation criteria

For fit, existing brands offer a concentrated source of design judgment. Castello Branco describes great brands as the product of dozens of designers investing craft, thought, and care. That work has already defined what is appropriate for a particular company, yet agents often fail to use it well. Maintaining the brand gives the system a specific quality target. She also suggests using the small classifiers as a gate to prevent an agent from shipping slop.

The brand API is the first product she says Taste Labs is releasing publicly, with beta testing already underway among design partners. It takes a brand URL and extracts specific, structured components that an agent can follow. The structure serves two purposes: it guides generation, and it provides criteria against which the result can be judged.

Verification remains essential after supplying that guidance. The application needs to determine whether the agent is staying on track, how well it adheres to the brand, and where it fails. This closes the loop between instructions and observed output. Castello Branco reports that this flow is helping improve quality, but does not provide a component schema or a numerical adherence score.

11:2111:23
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:21 · section reference included

Retrieving cohesive systems and preserving an existing brand

Selected presentation frame from Training Taste — Thais Castello Branco, Taste Labs at 833 seconds
Retrieving cohesive systems and preserving an existing brand

For users without an existing brand, Taste Labs is creating an index of prepared brand systems. A user asking for something dreamy could retrieve a system already designed to be cohesive, rather than generating every choice at that moment. The proposed benefit is to reuse considered design work and reduce the risk of an incoherent or repetitive result.

She then presents a slide-deck example using the branding of the General Intelligence Company of New York. Her comparison places the original brand beside a default generated result and a result produced using the extraction process. She describes the extracted-brand result as much closer to the original, including its details. This is an illustrative claim about improved fidelity; the spoken explanation does not identify the precise visual differences or quantify the improvement.

12:5312:56
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

12:53 · section reference included

Starting with a higher minimum standard

Selected presentation frame from Training Taste — Thais Castello Branco, Taste Labs at 866 seconds
Starting with a higher minimum standard

Castello Branco closes by affirming the value of human taste and exceptional craft. Her immediate aim is to raise the basic standard of generated work before debating whether models can reach the pinnacle of human taste. Decomposing design problems and measuring their properties are practical steps toward that higher quality floor. The talk ends with thanks, applause, and music.

14:0414:05
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

14:04 · section reference included

Read the complete timestamped transcript
  1. 0:01

    [music]

  2. 0:12

    >> Test.

  3. 0:13

    Okay, amazing.

  4. 0:15

    It's great to meet everyone. I'm Taís.

  5. 0:16

    I'm the founder of Taste Labs. Uh for

  6. 0:19

    those of you who don't know us, we came

  7. 0:20

    out of stealth a few weeks ago and our

  8. 0:22

    whole mission is basically how do we end

  9. 0:24

    AI slop? I that's my personal enemy. Um

  10. 0:27

    and so we really believe that to solve

  11. 0:30

    this problem of slop, we have to like

  12. 0:32

    decode subjective domains. Uh there's

  13. 0:34

    been so much effort being put into

  14. 0:36

    getting models and agents amazing at

  15. 0:38

    things like coding and math. Uh and it's

  16. 0:40

    time that we put all that same effort

  17. 0:41

    into making them great at things like

  18. 0:43

    design uh and writing. And so design is

  19. 0:45

    this first pillar that we're starting

  20. 0:46

    with and it's been it's been incredibly

  21. 0:48

    exciting. Um

  22. 0:50

    We work primarily in two ways. So we

  23. 0:52

    work a lot with the frontier labs on how

  24. 0:54

    do we evaluate their models, understand

  25. 0:56

    where they're breaking, understand what

  26. 0:58

    could be better about them, and then

  27. 0:59

    construct the right either post-training

  28. 1:01

    data or RL environments to basically fix

  29. 1:03

    that problem. And part of this is like

  30. 1:05

    how do you take something as fuzzy and

  31. 1:06

    large as design and break it down to a

  32. 1:09

    level that you can identify what is best

  33. 1:11

    solved through each method. What are

  34. 1:13

    elements of design that are almost like

  35. 1:15

    once you kind of boil down the problem

  36. 1:16

    becomes so specific that they almost

  37. 1:18

    become deterministic. So for example, uh

  38. 1:20

    if you're trying to train a model to be

  39. 1:21

    good at selecting color palettes or have

  40. 1:23

    contrast or alignment, those are things

  41. 1:26

    that if you define the problem and the

  42. 1:27

    context in a specific enough way, uh you

  43. 1:29

    you can get to an answer that's like

  44. 1:31

    pretty objective or that at least most

  45. 1:32

    experts would agree to. But maybe other

  46. 1:34

    things like uh aesthetics, you naturally

  47. 1:37

    will see this expert disagreement. And

  48. 1:38

    so then you want to lean on to things

  49. 1:40

    that are closer to to data. So anyway,

  50. 1:42

    we spend a lot of time thinking about

  51. 1:43

    all those problems. Uh but on the other

  52. 1:44

    side is also

  53. 1:46

    without even touching the model layer,

  54. 1:47

    right? How do we actually help agents

  55. 1:49

    and app layer companies produce better

  56. 1:51

    things? And there's a lot that goes into

  57. 1:53

    that, right? You have these different

  58. 1:54

    sets of problems at the application

  59. 1:55

    layer because you're using an

  60. 1:56

    off-the-shelf model that tends to

  61. 1:58

    collapse in terms uh uh of style tends

  62. 2:00

    to collapse to the mean. So, how do we

  63. 2:02

    force that creativity back to the

  64. 2:03

    system? How do we avoid these patterns

  65. 2:05

    of slop, which we'll talk about a lot

  66. 2:07

    today? Uh how do you understand like

  67. 2:09

    user preferences or brand preferences

  68. 2:11

    preference so that you can uh maintain

  69. 2:13

    endurance to that style? Uh so, there's

  70. 2:15

    lots of things that are actually need to

  71. 2:17

    be solved as context or judgment or

  72. 2:19

    verification at the app layer, which is

  73. 2:22

    why we kind of work across both.

  74. 2:25

    Maybe I'll start with more of a a

  75. 2:27

    philosophical question of like how how

  76. 2:28

    do you define something that is great?

  77. 2:30

    Like how do you define greatness? And

  78. 2:32

    for something like math, it's easier,

  79. 2:34

    right? Because there's kind of one

  80. 2:36

    objective answer, and uh great is the

  81. 2:38

    same as correct. But then for something

  82. 2:40

    like writing or design,

  83. 2:43

    it's much harder, right? Like how do you

  84. 2:44

    define what's like a great tweet or

  85. 2:45

    what's a great art piece or what's a

  86. 2:47

    great website? Um I don't know what's

  87. 2:50

    the last time that you interacted with a

  88. 2:52

    poem or walked into a coffee shop and

  89. 2:53

    for some reason it kind of like hit

  90. 2:55

    different and it felt

  91. 2:56

    very special. Uh but probably it's a

  92. 2:58

    combination of things that it it felt

  93. 3:00

    very unique. It felt almost a little

  94. 3:02

    different. It kind of called your

  95. 3:03

    attention. Uh it felt like there it was

  96. 3:04

    made with a lot of care and attention to

  97. 3:07

    to detail and craft, and it almost had

  98. 3:08

    the sense of of like authenticity. Um

  99. 3:11

    and I think that's a lot of what AI is

  100. 3:12

    missing today. It's like how do we take

  101. 3:14

    uh things that are not necessarily

  102. 3:16

    average, right? How do we produce things

  103. 3:17

    that are purposely like out of

  104. 3:18

    distribution? Um and slop is kind of the

  105. 3:21

    opposite of that, right? I think it is

  106. 3:23

    hard to define what is great sometimes,

  107. 3:24

    but I think it's pretty pretty easy to

  108. 3:26

    define what is slop in the sense that

  109. 3:27

    most people would agree. I think the

  110. 3:29

    sense of like repetition of kind of

  111. 3:31

    soullessness is something that all of us

  112. 3:32

    feel right now when using AI, and I

  113. 3:34

    think it's quite magical, by the way,

  114. 3:35

    that AI has gotten to a point that any

  115. 3:38

    human on the planet that is not even a

  116. 3:39

    designer, that is not an engineer, can

  117. 3:41

    click a button and suddenly make an

  118. 3:43

    entire PowerPoint or make a website or

  119. 3:45

    make a web app. That's pretty cool. But

  120. 3:47

    it comes with consequences, right? Uh it

  121. 3:49

    comes with consequences of suddenly now

  122. 3:51

    the cost of generation is basically

  123. 3:53

    going to zero. Uh and But the average

  124. 3:55

    person hasn't necessarily honed their

  125. 3:57

    taste. Like I does think about the

  126. 3:59

    amount of effort and work that a

  127. 4:01

    designer puts in throughout their life

  128. 4:03

    to like build up their taste, right?

  129. 4:04

    Like there's all this process of like

  130. 4:06

    getting exposed to many things and

  131. 4:08

    learning to like spot patterns and

  132. 4:10

    learning to develop a point of view and

  133. 4:11

    like kind of do things in a in a

  134. 4:13

    courageous way that maybe are a little

  135. 4:14

    bit against the norm. Learning what not

  136. 4:16

    to do and how to like have restraint and

  137. 4:18

    that's very hard. Like the average

  138. 4:20

    person doesn't necessarily have the the

  139. 4:22

    time nor the skills to go and develop

  140. 4:24

    taste in everything, let's say in

  141. 4:25

    design. And so

  142. 4:27

    um

  143. 4:28

    I think it would be a bad case scenario

  144. 4:29

    for us to just like be like, "Okay, the

  145. 4:30

    way to fix slop is for everyone to have

  146. 4:32

    taste." cuz I don't think that's

  147. 4:33

    necessarily realistic. Um I think how do

  148. 4:35

    we how can we understand this better so

  149. 4:37

    that we can make even for the average

  150. 4:39

    person the ability to create something

  151. 4:41

    great and to understand maybe their own

  152. 4:42

    taste um

  153. 4:44

    easy more more more easy. So that's

  154. 4:47

    that's a lot of what we're we're

  155. 4:48

    focusing on. Um

  156. 4:49

    So yeah, I mean this phenomenon of slop,

  157. 4:51

    by the way, is not new. Uh if you were

  158. 4:53

    in the internet uh as social media

  159. 4:55

    emerged, you probably saw a lot of slop

  160. 4:57

    before that. But I do think that AI has

  161. 4:59

    been this kind of like accelerating

  162. 5:00

    force, right? Of like being able to

  163. 5:01

    create things very easily uh with a

  164. 5:03

    click of a button and that like

  165. 5:04

    thoughtlessness around it. And there's

  166. 5:06

    kind of these three characteristics that

  167. 5:07

    I I would say repeat in slop. Uh so A,

  168. 5:10

    repetition. So you start seeing the same

  169. 5:13

    thing many many many times. Um the

  170. 5:15

    second is lack of fit, which I actually

  171. 5:16

    think is is very related. So

  172. 5:19

    fit is kind of this ability for

  173. 5:20

    something to feel correct for a specific

  174. 5:22

    context, right? For a specific moment in

  175. 5:24

    time, for a specific person. Uh but

  176. 5:26

    suddenly if you have repetition and

  177. 5:27

    let's say one person asks for a website

  178. 5:30

    for their pet shop and the other one

  179. 5:31

    asks for a website for their finance

  180. 5:33

    firm and somehow those designs converge

  181. 5:36

    and look the same.

  182. 5:37

    That's quite odd, right? Like if that

  183. 5:39

    was in if you were actually crafting

  184. 5:40

    that with care, that wouldn't you

  185. 5:42

    wouldn't converge necessarily on those

  186. 5:43

    things. And so this lack of fit and lack

  187. 5:46

    of understanding of context is actually

  188. 5:47

    huge problem that like leads to slop. Um

  189. 5:50

    and the third is maybe low intent, which

  190. 5:51

    is probably a mix of Yeah, you're going

  191. 5:53

    to have a bunch of people prompting

  192. 5:54

    really quickly and maybe just wanting to

  193. 5:55

    one shot something. But I think there's

  194. 5:57

    actually this like intent interpretation

  195. 5:59

    piece that's missing in the systems that

  196. 6:00

    we're building. Like how can you help

  197. 6:02

    your user, right? Like how can you help

  198. 6:03

    them better understand the intent that

  199. 6:04

    they have um so that you can add more

  200. 6:07

    color and add more context on onto what

  201. 6:09

    you're trying to create.

  202. 6:12

    Okay. And I I'm a big believer by the

  203. 6:14

    way that you in order to fix something

  204. 6:16

    first have to measure it and you first

  205. 6:18

    have to understand it. I think that's

  206. 6:20

    exactly why we're so focused on like how

  207. 6:21

    do we

  208. 6:22

    uh

  209. 6:23

    turn these domains into something a bit

  210. 6:25

    more verifiable so that we can attach a

  211. 6:27

    measure to it. So uh

  212. 6:29

    you'll you'll go on a little bit of a

  213. 6:30

    research journey with me here now, but

  214. 6:31

    we basically wanted to figure out can we

  215. 6:33

    measure slop? Like can we actually

  216. 6:35

    measure this quantitatively and spot

  217. 6:36

    this and what does that like look like?

  218. 6:41

    So we analyzed over 2 million websites

  219. 6:43

    from the past like 10 years kind of like

  220. 6:45

    way back machine style to try to

  221. 6:46

    understand all the trends across like

  222. 6:48

    design, how is the internet changing, uh

  223. 6:50

    how are

  224. 6:51

    how is like design changing over time?

  225. 6:53

    And two things were interesting. And we

  226. 6:55

    also met by the way then kind of

  227. 6:57

    synthetically generated a set of uh

  228. 7:00

    design websites so we could kind of like

  229. 7:01

    compare like how does human-made sites

  230. 7:04

    compare to AI-generated ones? And there

  231. 7:07

    were a few things that were interesting.

  232. 7:08

    So one was that you already kind of saw

  233. 7:11

    a a bit of like a collapse

  234. 7:13

    uh on the internet before even AI. So

  235. 7:15

    you saw kind of the internet becoming

  236. 7:17

    more homogeneous, using more similar

  237. 7:19

    color palettes, using more similar

  238. 7:20

    layouts, uh which is probably a function

  239. 7:22

    of more uh

  240. 7:24

    I would say this like kind of trends

  241. 7:25

    spreading more and more and more

  242. 7:26

    quickly, let's say. Uh but with AI I

  243. 7:28

    think you saw this repetition happening

  244. 7:30

    a lot more and being almost more like um

  245. 7:32

    identified kind of regardless of

  246. 7:33

    context. So even in completely different

  247. 7:35

    buckets you saw patterns that were very

  248. 7:36

    similar.

  249. 7:37

    So we we built this I I call this

  250. 7:39

    probes, but basically we uh we did two

  251. 7:42

    things. So we did this like pattern

  252. 7:43

    mining on all this data to understand

  253. 7:45

    like what are features that we can

  254. 7:46

    extract from all these sites. What are

  255. 7:47

    all these characteristics that we can

  256. 7:49

    make more objective, right? Colors,

  257. 7:51

    typography, layout, audience. Like, how

  258. 7:53

    can we like distill this down into

  259. 7:54

    things that become almost like uh

  260. 7:56

    structured? And then how do we uh train

  261. 7:58

    up these like probes? So, think of these

  262. 8:00

    as like baby classifiers. Like, how do

  263. 8:02

    we uh train the ability to spot this one

  264. 8:04

    characteristic?

  265. 8:05

    And for all these slop sites, we

  266. 8:07

    identify we started identifying like

  267. 8:09

    what are the probes that basically mean

  268. 8:11

    the site is very likely to be AI slop.

  269. 8:14

    Um and especially when you start

  270. 8:16

    combining them use and you see the

  271. 8:17

    frequency of multiple of these happening

  272. 8:19

    at once, it became very likely that you

  273. 8:21

    could actually like measure

  274. 8:23

    uh and predict slop. And we saw a a

  275. 8:25

    super high basically ability to do that

  276. 8:27

    prediction, which was really cool to

  277. 8:29

    see. This performed better, by the way,

  278. 8:30

    than like most LLM as a judge methods of

  279. 8:32

    like asking an LLM to like judge if that

  280. 8:35

    uh is like great human quality versus

  281. 8:36

    like AI-generated slop. Uh so, that was

  282. 8:39

    pretty cool to see. And I think kind of

  283. 8:40

    shows this pattern that we see in AI

  284. 8:43

    really being uh an actually quantitative

  285. 8:46

    thing that we can see in slop, uh which

  286. 8:48

    I find really cool. But, obviously, we

  287. 8:49

    don't want to stop there, right? We

  288. 8:50

    don't want to just measure slop. We want

  289. 8:52

    to also solve it. And so,

  290. 8:54

    um there's a few I I think I mentioned

  291. 8:56

    this before, but like the as the cost of

  292. 8:58

    production basically goes to zero, I

  293. 9:00

    think the thing that becomes

  294. 9:02

    expensive and matters more than ever is

  295. 9:04

    judgment. Um I don't even want to use

  296. 9:06

    the word taste here.

  297. 9:07

    Is judgment. I think it's this ability

  298. 9:09

    to discern what's right. It's this

  299. 9:10

    ability to break down a problem so that

  300. 9:12

    you can actually understand it and

  301. 9:13

    create solutions for it. And so, yes,

  302. 9:16

    there's the side of judgment that is

  303. 9:17

    human judgment that I actually think is

  304. 9:18

    more valuable than ever. But, there's

  305. 9:20

    also the side of like how do we build

  306. 9:21

    the right tools and systems to like fix

  307. 9:23

    pieces of this problem, right?

  308. 9:27

    So, yeah, how do we how do we fight

  309. 9:28

    slop, my my enemy?

  310. 9:30

    Um

  311. 9:31

    And, by the way, I think there's there's

  312. 9:33

    a lot of conversation going around how

  313. 9:35

    do you fight slop at the model layer?

  314. 9:37

    Like, how do we make models better? How

  315. 9:39

    do we make models have a higher bar?

  316. 9:40

    Which don't get me wrong, it has to be

  317. 9:42

    solved and we're working very hard to

  318. 9:43

    solve that, too. But I actually think

  319. 9:45

    this problem of inference time is

  320. 9:46

    equally, if not even more important.

  321. 9:49

    Because that's actually when you

  322. 9:50

    interact with the end user. And this

  323. 9:52

    kind of back and forth of how do you

  324. 9:54

    understand this context and intent

  325. 9:55

    happens at the moment of inference time.

  326. 9:57

    So, I don't think that we can ignore and

  327. 9:58

    just make models better and not solve

  328. 10:00

    this, otherwise slop will keep existing.

  329. 10:02

    Um so, maybe breaking down a few of

  330. 10:04

    those pieces and kind of

  331. 10:06

    um

  332. 10:07

    a few of the ways that we've thought

  333. 10:08

    about solving this or a few solutions

  334. 10:09

    that we built to solve this. But I

  335. 10:10

    think, for example, for something like

  336. 10:11

    repetition, one of the things that we're

  337. 10:13

    working on is I I've nicknamed it, I

  338. 10:15

    don't know if that's going to be the

  339. 10:15

    official name, but like the creativity

  340. 10:17

    API. How can we create a system that

  341. 10:19

    almost becomes an inspiration machine

  342. 10:21

    for your agent? So that it can produce

  343. 10:23

    something that's actually out of

  344. 10:24

    distribution instead of something that

  345. 10:25

    is

  346. 10:26

    in that same average and kind of mean

  347. 10:28

    that we're seeing happen with like the

  348. 10:29

    slop sites. Um so, this is one of the

  349. 10:31

    ways that practically, if we can

  350. 10:33

    intentionally produce something that's

  351. 10:34

    out of distribution, you can improve

  352. 10:36

    this like overall uh quality. And

  353. 10:39

    by the way, I I don't think that this

  354. 10:41

    can be something just like randomness.

  355. 10:43

    So, it's not just about like turning up

  356. 10:44

    the temperature of the model and and

  357. 10:45

    kind of

  358. 10:46

    fingers crossed hoping for the best. I

  359. 10:47

    think it's much more like how do we

  360. 10:48

    understand um even like what are rules

  361. 10:52

    or expectations in specific domains?

  362. 10:53

    Like let's say that you asked for a

  363. 10:55

    slide deck for for the pitch of your

  364. 10:57

    startup. Like

  365. 10:58

    what is a what does a good pitch deck

  366. 11:00

    look like? And then how do you almost

  367. 11:01

    like intentionally break rules uh to

  368. 11:04

    create things that are more creative,

  369. 11:05

    right? Because usually creativity isn't

  370. 11:06

    like randomness, isn't doing something

  371. 11:08

    that completely feels off for that

  372. 11:10

    situation. It's like you intentionally

  373. 11:12

    maybe diverge on a couple of things

  374. 11:14

    while maintaining kind of

  375. 11:16

    um

  376. 11:16

    adherence to to expectations of that

  377. 11:19

    category, let's say for for others. So,

  378. 11:21

    that's one of the things we're working

  379. 11:21

    on. The second one of this problem of

  380. 11:23

    fit, I think um it's interesting, but

  381. 11:26

    brands, as probably a lot of you who are

  382. 11:28

    designers know, takes so much effort to

  383. 11:30

    create great brands. Like great brands

  384. 11:32

    are the work of

  385. 11:34

    dozens of designers putting in a lot of

  386. 11:36

    like craft and thought and care.

  387. 11:39

    And so we've almost like already

  388. 11:40

    pre-done the work of defining what is

  389. 11:41

    great for that specific company and then

  390. 11:43

    we're not using it well. So this like

  391. 11:45

    brand endurance actually think is a huge

  392. 11:46

    problem and one of the things that can

  393. 11:48

    very

  394. 11:49

    more easily let's say like raise that

  395. 11:51

    bar of quality. So I'll I'll touch on an

  396. 11:52

    example on this one specifically. And

  397. 11:54

    then same with like intent and judgment

  398. 11:56

    I think the baby classifiers was a good

  399. 11:57

    example. Um

  400. 11:59

    like how it how we can actually like use

  401. 12:01

    this to be even become a gate for slop

  402. 12:03

    and and not let your agent ship slop.

  403. 12:06

    But so the brand API is the first

  404. 12:08

    product that we're releasing to to the

  405. 12:09

    public. This is already in in beta

  406. 12:11

    testing with a bunch of our design

  407. 12:13

    partners. And essentially what it does

  408. 12:14

    is it can take let's say a brand URL and

  409. 12:18

    extract this into like very specific

  410. 12:20

    components that are good for an agent to

  411. 12:21

    follow. So basically how do we turn

  412. 12:23

    something as fuzzy as a brand into

  413. 12:25

    something so structured that it becomes

  414. 12:26

    easy to

  415. 12:28

    for your agent to follow that but also

  416. 12:29

    for you to judge against it, right?

  417. 12:30

    Because I think the piece that we can't

  418. 12:32

    forget here is this judgment and

  419. 12:33

    verification. So yes, this goes and

  420. 12:36

    helps your agent to produce something

  421. 12:37

    better.

  422. 12:38

    But how can we also add a way for you to

  423. 12:41

    judge okay, is the agent actually

  424. 12:42

    staying on track? Is it actually

  425. 12:43

    performing well to adhere to this brand

  426. 12:45

    or how is it failing or where is it

  427. 12:46

    failing? So this is the first flow I

  428. 12:49

    would say that we we're seeing that is

  429. 12:51

    really helping to improve quality.

  430. 12:53

    And what's cool is of course we're

  431. 12:54

    talking here about an example of a brand

  432. 12:56

    that already exists. But let's say you

  433. 12:58

    have an agent or you have an app and the

  434. 13:01

    person that is using your app actually

  435. 13:02

    doesn't have a brand. Let's say they're

  436. 13:03

    an average consumer. Can we actually One

  437. 13:05

    of the things that we're creating is

  438. 13:06

    basically like a repository, like an

  439. 13:08

    index of brands

  440. 13:10

    of pre almost like pre-created brand

  441. 13:12

    system so that if they want something

  442. 13:13

    that feels dreamy, why not retrieve a

  443. 13:15

    dreamy brand system that already has

  444. 13:17

    been thought out to be cohesive instead

  445. 13:19

    of doing like a generative approach the

  446. 13:21

    moment of that might end up not so great

  447. 13:23

    or might end up again in those pillars

  448. 13:25

    of slop.

  449. 13:26

    And I want to show you a real example of

  450. 13:28

    this in action. So, um

  451. 13:29

    there's this company that I think is

  452. 13:30

    awesome called the General Intelligence

  453. 13:31

    Company of New York. They have a sick

  454. 13:32

    website, you guys should check it out.

  455. 13:34

    Um but basically, if you ask Cloud

  456. 13:35

    Design to create a slide deck uh in

  457. 13:38

    their branding,

  458. 13:40

    the the middle one is basically what it

  459. 13:41

    comes up with. So, the one on the left

  460. 13:42

    is is the original brand. Uh this is the

  461. 13:45

    kind of the the default. And if you kind

  462. 13:47

    of use this extraction actually in the

  463. 13:49

    process, it creates something that's way

  464. 13:50

    more high fidelity with the original. Um

  465. 13:53

    and that even like in the details, I

  466. 13:55

    would say like feels right. So, this is

  467. 13:56

    just to show an example of it in in

  468. 13:58

    action. Um

  469. 14:00

    but yeah, I think we

  470. 14:04

    I think all of us would agree that like

  471. 14:05

    human human taste and kind of the peak

  472. 14:07

    of human craft is always going to be

  473. 14:10

    like deeply valuable. And that right

  474. 14:13

    now, I think the challenge is we are

  475. 14:15

    almost even not earning the right to

  476. 14:17

    debate this like how can we have

  477. 14:19

    uh models like reach this like pinnacle

  478. 14:21

    of taste. I don't think it's about that

  479. 14:22

    at all. It's like how do we first just

  480. 14:24

    like

  481. 14:24

    raise the bar. Like the bar is

  482. 14:26

    currently, I would say, on the ground.

  483. 14:27

    And so, I think all of this work that

  484. 14:29

    we're putting into like how do we

  485. 14:31

    decompose a problem and how do we

  486. 14:32

    measure it is exactly so that we can at

  487. 14:34

    least like improve this bar um of

  488. 14:36

    quality. And I think we have to start

  489. 14:37

    with that.

  490. 14:41

    That's it. Uh thank you very much for

  491. 14:43

    for the time. Uh this is this is

  492. 14:45

    awesome.

  493. 14:46

    >> [applause]

  494. 15:03

    [music]