← All AI Engineer talks

AI Engineer World's Fair 2026

Ending AI Slop

Read the talk

Ending AI Slop

Better creative output starts with decomposing quality: verify the constraints you can specify, preserve the preferences you cannot average, and collect expert judgment with enough context to train on.

From a talk by Thais Castello Branco

Before you start: Basic familiarity with model training, reinforcement learning and preference data will help; no design background is required.

Which problems belong in the model?

How do you improve an AI’s design, writing, personality or emotional intelligence when a correct answer is not enough to define success? These are the subjective domains at the center of Thais Castello Branco’s work as founder of Taste Labs. The starting point is to decompose what quality means before deciding how to train for it.

That decision spans two layers. At the foundation-model layer, evaluation and benchmarking can expose failures that become candidates for reinforcement learning environments or post-training data. At the application layer, a disappointing result may instead reflect missing context or misunderstood user intent. Castello Branco anticipates a world in which many more people create without being domain experts; helping them express what they want matters alongside improving the model’s capabilities. The training problem is the focus here, but it does not absorb every problem an application must solve.

0:170:27
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:17 · section reference included

Quality needs a target

Writing, design, sales and marketing frequently admit several valid answers. The difficulty is not just choosing among them: it is defining what would make one excellent. Without a clear target, both evaluation and training become harder.

Dark slide reading “Most of the world is subjective. Yet few are studying these domains.” with the speaker inset at lower left.
Most of the world is subjective, yet few are studying these domains.

Code offers a useful comparison. It can be decomposed, executed and checked. Those are properties of the domain that make training easier, rather than capabilities supplied by a model alone. Design and writing do not arrive with an equivalent general-purpose test for goodness. Castello Branco frames this as capability follows measurability: making some part of quality measurable creates a route toward improving it. A second problem remains even after that work—an average or familiar answer may not be the best creative answer.

1:582:11
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

1:58 · section reference included

Great for whom, where and when?

To make a fuzzy request such as “great design” more verifiable, first specify the person, taste and situation it must serve. The same slide could be excellent for a startup and inappropriate for a finance firm. Its quality depends on its use, not only on its visual properties.

Taste also changes over time. A design that feels right today may differ from what worked five years ago or what will work five years from now. Compared with a stable correctness test in code or mathematics, this adds another moving condition to evaluation. Audience, intended use and time therefore belong in the definition of the task, before an evaluator is asked to judge the result.

Slide titled “Transform from fuzzy into more verifiable.” above a green and gray graphic with connected white points and blurred numbers.
Transform from fuzzy into more verifiable.
3:253:41
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

3:25 · section reference included

Turn brand adherence into a task

Suppose a coding agent has produced an internal dashboard or landing page. How do you decide whether it is good? A company already has a useful source of constraints: the work its designers put into defining its brand. Color combinations, typography, spacing and texture embody decisions about which elements work together and why. Asking for something on brand is a more bounded problem than asking for something great.

Castello Branco uses Reducto as the example. Decompose the brand into colors, typography, motion, animation and textures, and the broad judgment starts to expose specific properties an evaluator can inspect. A holistic LLM-as-a-judge prompt still has to infer what the brand means and decide whether the output expresses it. An explicit specification makes the target less implicit.

The next step is task design. Castello Branco flags reward hacking and hallucination patterns as reasons not to assume an LLM judge is always the right evaluator. To shape the problem into an RL environment, the task needs a usable ground truth. The difficult work is establishing that ground truth from the fuzzy goal, not merely attaching a score to an output.

For the proposed Reducto task, the sequence is:

  1. Establish the specification. Use the decomposed brand components as the ground truth.
  2. Request a new page. Ask the agent to adhere to the brand while producing something different from the existing page.
  3. Evaluate the constraints. Grade the result against the specification, rather than treating similarity to the original as the whole objective.

This leaves room for new combinations of familiar components. A page can be on brand without reproducing the reference layout. The specification should constrain the design without making copying the route to success. This is an illustrative environment design, not a reported training result.

4:274:37
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:27 · section reference included

Why familiar output can still be poor output

Design properties sit along a spectrum. Alignment and aspects of typography are closer to objective assessment. Style, fit and creativity are harder to measure. Decomposing a design problem therefore creates a routing problem: which method best addresses each component? It does not make every component equally verifiable.

The creative end of that spectrum exposes the second difficulty. Castello Branco contrasts answering two plus two with writing or design. For the arithmetic question, the shared correct answer is desirable. For a creative task, the most familiar continuation may be precisely what makes the output feel repetitive. Her account of collapse toward the mean is a way to describe that repetition, rather than a claim that a model literally computes an arithmetic average of designs.

Creative excellence can involve deliberately breaking a rule or departing from an established pattern. The departure must be intentional: merely becoming unusual does not establish quality. Once the question becomes whether an unconventional choice works, human preference and judgment may provide a better training signal than an environment with a readily verifiable answer.

Castello Branco reports that Taste Labs works with over a thousand experts across different media and styles. The purpose of that breadth is to preserve variation when decomposing the domain, rather than let one dominant style stand in for the entire distribution of good work. Expert diversity is one part of the proposed response to repetitive output, not a complete solution by itself.

6:567:09
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:56 · section reference included

Pull what you can toward verification

The more a task can be verified—especially programmatically—the more suitable it becomes for an RL environment. Establishing ground truth and designing the task can move a previously fuzzy requirement in that direction. As context, changing taste or individual preference becomes decisive, the task retains a stronger need for human judgment.

Brand adherence illustrates the movement. Taken as a whole, it can require an expert’s judgment. Extracting the components that matter makes part of that judgment observable and measurable. The routing decision follows the decomposition:

ComponentEvaluation direction
Explicit, observable constraintsVerification
Codifiable parts of brand adherenceExtract constraints, then verify
Context, taste and preferenceHuman judgment

The useful unit of choice is the component, not the entire domain. A single design task can contain both checks and judgments.

Framework diagram listing design criteria, with connecting lines to Verify and Judge boxes and an Extract box on the route from brand adherence toward Verify.
A framework routes design criteria toward verification or judgment.
9:179:26
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:17 · section reference included

Keep the person attached to the preference

For the more subjective components, Taste Labs’ position is that human taste remains ahead of LLM judging. The training challenge is to capture that judgment as useful data. Collecting preferences from many people is insufficient if the collection process loses who they are, what they like, why they like it and when that preference applies.

Two people can prefer different styles without either label being wrong. Nor does their disagreement establish that a design halfway between the two would satisfy either person. If those preferences are pooled without context, legitimate differences can look like inconsistent supervision.

Castello Branco proposes attaching something like a preference vector to preference data. The aim is to make matching possible: a judgment can be understood relative to a person’s preferences rather than treated as a universal vote. Training could then preserve multiple preferences intentionally, instead of treating the differences as noise. This is a proposal for representing preference, not a specified vector format or a demonstrated training implementation.

10:2310:32
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:23 · section reference included

Make expert feedback specific and locatable

Once human judgment becomes the data source, another question appears: what makes the data useful to a model? Some aspects of subjective quality remain difficult to define, but the collection process contains concrete controls. Decompose the domain into meaningful categories, select experts deliberately across those categories, inspect their qualifications and apply strict selection criteria. Coverage across media and styles should be a collection decision, not an accident of who happens to annotate.

Then route each problem to the appropriate process and inspect the resulting annotations for signal. Castello Branco emphasizes specificity: an expert’s reasoning should identify precisely what is good or bad and explain the observation. A broad verdict leaves the model with much less information about which feature drove the judgment.

For a landing page, the next improvement is to connect the observation to the relevant code component. A paragraph describing the visual result can leave the model to infer which code produced the criticized feature. Linking the commentary to the exact component makes that relationship explicit. A TypeScript annotation record can express the connection directly:

typescript

type DesignAnnotation = {
  assetId: string;
  component: {
    file: string;
    exportName: string;
    selector: string;
  };
  criterion: string;
  observation: string;
};

const annotation: DesignAnnotation = {
  assetId: "landing-page-review",
  component: {
    file: "src/components/Hero.tsx",
    exportName: "Hero",
    selector: "[data-component='hero-title']",
  },
  criterion: "alignment",
  observation:
    "The heading begins to the left of the body copy and primary button.",
};

Here, the example record preserves both the specific visual observation and where to look in the implementation. It records a critique; it does not apply a design change. The collection workflow and its QA checks should work together to produce this kind of clear, grounded signal.

11:4011:50
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:40 · section reference included

Quality does not require universal agreement

Human QA has two different jobs. In a collection of slide-design preferences, reviewers can check whether annotations follow the specifications and contain useful reasoning. Asking a second designer to agree with the first designer’s preference is a different test. The meaning of disagreement depends on what they disagree about.

Disagreement aboutQA response
Alignment or another relatively objective propertyInvestigate a possible error
Style or aesthetic preferencePreserve a potentially valid difference

Alignment disagreement may expose a flaw in the observation or annotation. Aesthetic disagreement can reveal the preference distinctions the dataset ought to capture. Screen for sound fundamentals without requiring a single shared taste. Consensus is useful when applied to the right question.

The final test is how the data interacts with models. Taste Labs conducts research on that interaction, but Castello Branco acknowledges that work with frontier labs does not always provide a feedback loop precise enough to identify the downstream effect of a particular data contribution. That limitation makes controllable properties—expert selection, decomposition, specificity and annotation quality—especially valuable to measure within the data workflow.

Castello Branco closes by advocating quality over quantity. High-quality data in subjective domains is expensive, difficult and dependent on deep expertise. Her recommendation is to invest in carefully decomposed tasks and strong expert judgment rather than accumulate noisy, poorly specified examples. She argues that this produces better results, without supplying a quantified comparison. The expense buys a more useful training signal: judgments whose domain, reasoning and intended meaning have been preserved.

13:5814:11
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:58 · section reference included

Resources

From the talk

  • The public website used as the talk's brand-adherence example; its current design may differ from the presentation.

Read the complete timestamped transcript
  1. 0:00

    [on hold music] Hello, everyone. It's great to meet you all.

  2. 0:17

    I'm Thais. I'm the founder of Taste Labs. For those of you who don't know us, we came out of stealth a few weeks ago, uh, and our whole mission is basically how do we end AI slop?

  3. 0:27

    And we believe that to really solve this problem, we have to first decompose and understand subjective domains, right? I think as probably all of you know, AI's gotten quite good at things like coding and math, uh, but it's still super behind on things like design, creative writing, personality, emotional intelligence.

  4. 0:44

    And to understand these domains, I think we have to take a little bit of a different approach than we do with, um, with objective ones. So our idea is, like, how do we become this data and infrastructure layer to really, uh, understand these problems and to become the solution for them across the stack?

  5. 0:59

    So from the foundation model layer all the way to how do we build solutions for agents as well. So we work primarily in two ways. We work with the top frontier labs on how do we evaluate, benchmark their models, understand where they're breaking, uh, understand how we can fix them, and how do we determine also which problem

  6. 1:16

    is better fixed through each method. So what things should be turned into RL environments? Which things should be turned into post-training data problems? Um, but then we also go and work with a lot of agent and application layer companies on what are things that we actually don't believe should be solved at the foundation model layer and that

  7. 1:32

    might be better solved through methods like context, uh, or understanding user intent, right? We're basically betting on a world where suddenly you're gonna have billions of people creating, uh, that are not necessarily experts.

  8. 1:43

    So this understanding of user intent and contest-- and context is equally as important as how do we get these models to improve. For today, I'm gonna focus on the model training part.

  9. 1:51

    For those that are here tomorrow, I'll also be giving a chat on the design track where I'll cover more on what we're doing on the agent side of the house.

  10. 1:58

    But there we go. Okay. So most of the world is subjective, as I was mentioning. Uh, a lot of the world is subjective, right? If we talk about these domains of writing, even workflows within companies, right, of sales, marketing.

  11. 2:11

    Uh, a lot of the times there's this multitude of answers. There's not one clear right answer, and it's very hard to define what great even means. I think at the end of the day, like, those are why these domains are so difficult.

  12. 2:22

    Um, and we oftentimes, I would say, forget to mention, like, why. Uh, we treat, for example, the fact that code is verifiable and measurable as something that is a property about models, and models are great at, at coding, um, because we've made them great at coding.

  13. 2:42

    But realistically, it's actually a fact about code. Code is something that decomposes, it verifies, it executes, and so it makes it a lot easier for us to be able to train on these domains.

  14. 2:51

    For something like design or writing, like, how do you decompose it? How do you verify it? [laughs] How do you judge if it's actually good? So that's why they become so difficult.

  15. 2:59

    Um, so there's two characteristics that I wanna touch on today on why fundamentally subjective domains are harder. One is that capability follows measurability. So if we can solve the measurability problem, or at least part of it, then we can solve a big portion of these domains.

  16. 3:14

    Uh, the second, which I'll touch on later, is basically this, like, collapse to the mean and why the mean is not necessarily optimal in subjective domains.

  17. 3:25

    Okay. So to really start solving this problem, we have to turn something that feels fuzzy, like if I ask you, "What is great design?" into something that is more verifiable.

  18. 3:41

    So there's a few questions here, right? 'Cause if I ask you this, of what is great design, um, you could ask yourself, okay, um, do you mean great for which type of person, for which type of taste, for which situation?

  19. 3:52

    The same slide could be amazing, for example, if you are a startup and completely inappropriate if you are a finance firm. So it's contextual, first of all. Second of all, it has this property that it changes over time, which is different from other domains.

  20. 4:06

    What is considered good today is different than five years ago and different than five years from now. In code, that's not necessarily true, or in math, right? That's something that is way more consistent over time.

  21. 4:15

    So our ability to, again, decompose it and understand how is this good for a specific audience, how is this good today, how is this good in context, um, is some of the things that we've been thinking about in terms of how to, how to break this down.

  22. 4:27

    But I wanna give you a very specific example because, of course, this can mean many things. So let's talk about brand. Um, if you're at a company and you've used coding agents, you've probably shipped an internal dashboard.

  23. 4:37

    You've probably shipped an internal, like, landing page. And oftentimes you might wonder, okay, how do I determine if this is slop, if this is actually good? And you have kind of this secret weapon at your disposal, which is really all the work that probably designers at your companies, for example, put into defining your brand.

  24. 4:53

    A brand to define takes a lot of effort, takes a lot of care. You're defining all these components about it, when it's good, why you're choosing certain combinations of colors, of typography, of spacing, of texture.

  25. 5:03

    Um, but if I just ask you to be like, "Okay, create something great," that's very hard. But suddenly if I'm like, "Okay, make something that is on brand," that is a much easier problem to define, and a brand is something that can become decomposable.

  26. 5:14

    Uh, so for example, if we t-- I'm using the Reducto brand as an example here because I, I, I like their website. Um, let's say that we decompose this brand into the colors, the typography, the motion, the animation, the textures.

  27. 5:26

    Suddenly you have these very codified things that you can verify against. Verifying in general if something's on brand, and you can try this, uh, by prompting an LLM-as-a-judge to do it, is quite hard.

  28. 5:36

    But once you start picking apart the exact elements that represent what great is, then it suddenly becomes the shape of something that is codifiable and verifiable. So if you want to turn this into a shape of an RL environment, for example, right, how would you train a model for a capability like brand adherence?

  29. 5:53

    Uh-

  30. 5:55

    LLMs, LLM-as-a-judge might not necessarily always be the best method. We know that there's a lot of reward hacking. We know that there's, uh, interesting hallucination patterns there too. And so we oftentimes try to create methods of basically how do we turn a task that feels fuzzy into one where there's a clear ground truth so that it can

  31. 6:12

    become the shape of an environment? So in this case, the task design itself is really kind of the hardest part of the problem of how do you turn something that appears very fuzzy into something that actually can be RL'd.

  32. 6:23

    Um, and so in this case, the decomposition that I mentioned becomes the ground truth. So let's say that you start by tasking an agent to create a new page that is gonna adhere to the Reducto brand, but be completely net new and different.

  33. 6:34

    Um, you would want that output to not only be graded versus the original, but to be graded con-- on this ground truth, right? Because it could come up with completely new ways of using these components that are still valid but are different from the original.

  34. 6:47

    So you don't necessarily want to just see if it's replicating the original. So this is one example of, like, how to turn this into a problem of environment shaped,

  35. 6:56

    um, so that we can make it more verifiable. But in a way, all of these things are, I would say, like spectrums, right? You have, uh, in a problem like design, you have these elements of things like vision, uh, alignment, typography that are closer to objective.

  36. 7:09

    Once you start moving up that scale onto things like style, fit, creativity, how do you judge and measure something like creativity, right? That's much harder. And so you kind of need to think of this as like a routing problem of how do you understand this, like, vast fuzzy problem, break it down into smaller components, and what is

  37. 7:27

    the best solution for each of these components?

  38. 7:30

    So why does something feel like slop, for example, when we're talking about something like creativity, right? I think this is the second reason why, um, subjective domains are so much harder to solve because if we're talking about the properties of models, they're basically predicting what's the most likely outum-- outcome to show up next, and they assume that

  39. 7:46

    that outcome is the ideal outcome. And for something like math and coding, that is true, right? You want the answer that your model gives you to be the average answer if you're asking what two plus two is, uh, which also happens to be the right answer and the optimal answer.

  40. 8:00

    But for something like writing or design, you don't necessarily want the average answer, right? The average, meaning the most likely, does not necessarily coincide with, like, the optimal. Uh, a lot of, like, w- the-- what I describe it is a lot of greatness in creativity happens actually at the ends of the distribution.

  41. 8:16

    It's not the most likely outcome. It's when you actually actively break from rules and actively break from patterns that you can create things that are subjective and, and great.

  42. 8:24

    Um, and so the reason why this feels like slop and that we have this feeling that we're surrounded by, by slop is exactly because of this collapse to the mean and this repetition.

  43. 8:34

    And so we have to find ways of, okay, how do we break these patterns? How do we break from the mean? Uh, but in a way that's also intentional.

  44. 8:40

    So then you kind of shift the problem onto things that are not so easily maybe verifiable, uh, but that are more questions of human preference and judgment, and that might be better solved by data, for example, than by environments.

  45. 8:51

    And so again, this kind of like mode collapse is, is really the thing that we're trying to solve and how do we force that distribution back. Uh, oftentimes, by the way, we, we work, for example, with a community of designers, um, over like a thousand experts that are experts in different types of medium, different styles, and we

  46. 9:08

    purposely want to force that distribution when bre-- we're breaking down the problem exactly so we don't end up in this mode collapse. But there's of course a lot of other pieces of that puzzle.

  47. 9:17

    Uh, but I think this is an interesting framework is like the closer you are to something that it becomes verifiable, especially programmatically, the better for something like RL, right?

  48. 9:26

    And I think the challenge is how do we turn things that feel fuzzy into things that become more verifiable by establishing this ground truth and designing tasks in a way that allow for that.

  49. 9:35

    But then the more that it does shift to things that are contextual or that depend on that time element that we talked about or this, like, distinction in preference, the more this shifts towards something that requires human judgment.

  50. 9:48

    So, uh, I, I, I like this analogy of basically kind of pulling things toward verification. So brand adherence on its own would be something that's very hard and that is probably better judged by a human than, for example, by an LLM as a judge or something deterministic.

  51. 10:02

    But by codifying it, by understanding which pieces matter and how do I turn that into something that is observable and measurable, we can kind of pull it into this realm of verification.

  52. 10:12

    Uh, so th- this routing logic is, I would say, if you have one takeaway, uh, take this away, is like how do we break down this problem into something that you understand what's actually the best method to solve it?

  53. 10:23

    And so when we are talking about these things that are more subjective, uh, we do require human judgment. I think this is something that we, um,

  54. 10:32

    we believe human judgment is still at a much higher level than any LLM as a judge. And this, like, human taste is really how do we encapsulate this in a way that, uh, can be turned into high quality data so that we can train these models better?

  55. 10:46

    And I think one of the tricky things here is oftentimes in the past you had kind of this, uh, collection of preference data that would collapse again to the mean because you would collect it from a bunch of different people without necessarily understanding who they are, what they like, why they like it or when they like it.

  56. 11:01

    And if you don't break up that problem accordingly, you then end up again with preferences that kind of don't agree with each other. 'Cause that naturally happens in the world, right?

  57. 11:08

    I bet that some of you might like one style better than another. And that doesn't mean either of those things are wrong or it doesn't mean that the best answer is the average of what two people might like.

  58. 11:17

    It means that we need to fundamentally understand that the world is multi-preference and how do we do that matching accordingly. So this almost like understanding of how do we create like a preference vector, let's say for someone and attach that to even something like preference data can help us to train in a way that allows for that

  59. 11:33

    more pluralism of preferences intentionally instead of that data being turned into something that's noisy.

  60. 11:40

    When we talk about data quality as well for these domains, I think it becomes very, uh... I don't know if any of you have bought data, for example, for these domains or, or have tried to curate data yourselves.

  61. 11:50

    Um, but there's-- it's very hard to define, okay, like now I'm roaming to this side of data. What is actually good? What is actually gonna be helpful to my model?

  62. 11:58

    Um, and there's obviously a few things that are harder to define, but a few that I think become more controllable. And how do we actually understand patterns of quality so that we can measure them?

  63. 12:07

    So one of them, I would say, is that problem decomposition. How do we force, like, true distributions of what you see in the world? How do you force true, like, expert selection across these, um, different buckets?

  64. 12:16

    And these are things that you can totally control, right? You can see who those experts are. You can do a selection pattern that is very strict. Um, and you can decompose that problem, and this is something that is completely in your control and that totally helps with the, um, results, let's say, of the experiment being, being good.

  65. 12:33

    Um, the second is I would say that, like, flow of, like, how do you determine then the right problem routed to the right solution? Um, we also, I would say, do a lot of what we call

  66. 12:45

    essentially, like, QA on this data, and I think there's two ways to do that, is understanding what are properties about that data point that correlate with being, it being rich and high signal?

  67. 12:54

    So for example, specificity. When you're trying to ask an expert to define, is this good, is this bad, or put reasoning behind it or create a whole observation system around, uh, how they would judge an asset, the specificity of their language and of how precise they're being able to be with how they're doing that description is what

  68. 13:15

    will determine that data quality. Or for example, if you can tie their commentary with actually which piece in the code does this relate to? Like, let's say you have an expert that's judging a landing page.

  69. 13:26

    Uh, they might be able to just, like, write a paragraph describing this. But we know that models have a tricky time kind of actually connecting the piece of the code to the visual.

  70. 13:35

    And so if you can find, for example, a method to tie that exact code component to the commentary of the expert, suddenly you have data that is way less noisy and way more clear.

  71. 13:44

    So these are some things that we can do both in terms of the flow of how you connect the data, uh, of how you collect the data, but also in these, like, QA checks of how do you determine characteristics about it that will, uh, correlate highly, let's say, with, with valuable data.

  72. 13:58

    And same with human QA. I think-- By the way, human QA is tricky here because there's two sides to it. There is a side of, for example, let's say you're collecting, uh, a bunch of preference data about slide design.

  73. 14:11

    Um, you could have human QA be like, "Okay, is this, again, high-quality data? Is it following the specs?" Which most people would agree on. But suddenly, if you ask for expert consensus and you try to have another expert see if they agree with the initial designer's, like, votes, you might start seeing some disagreement there.

  74. 14:27

    And you-- I think the key is understanding, is that disagreement something that's actually a flaw in the data? Meaning, are they, for example, disagreeing on something that they should be agreeing on, such as alignment?

  75. 14:37

    Like, alignment's something that's pretty objective, so it would be kind of odd to see experts disagreeing on that front. But suddenly, if they're disagreeing on things like sty-style or, um, aesthetics, that is not necessarily bad data.

  76. 14:51

    That's actually good data. It shows you that there is a distinction for what people like. And so, uh, this almost, like, analysis of how do you run human QA in a way that both screens for kind of the fundamentals, but then when you use consensus, I think in an intentional way, is another big, big piece of this.

  77. 15:06

    Um, and then obviously kind of seeing actually how this data is, is interacting with models. Uh, we run a lot of research on our side. Obviously, when you're interacting with labs as we do, um, it-- there's a lot more involved than we oftentimes don't get that feedback loop of exactly what affected this cause, which is why we're

  78. 15:22

    so adamant on focusing on these things we can control in terms of data quality, 'cause these are things that, uh, we can completely measure on, on our side. Um, so I think one other, uh, piece of advice or message of the day is, I think especially when it comes to subjective domains,

  79. 15:41

    I would advocate for a quality over quantity approach. I think creating high-quality data is expensive. It's difficult. It takes a lot of understanding and depth around a specific domain.

  80. 15:51

    And having that be incredibly high quality done by people that also are incredibly high taste or whatever you wanna call it in that domain, uh, yield far better results than getting a bunch of noisy data or a bunch of messy data, uh, that was not necessarily intentional or didn't have all those things we talked about of, like,

  81. 16:06

    the problem breakdown. So that is my, my message of, of the day. Thank you. [audience applauding] [upbeat music]