← All AI Engineer talks

AI Engineer World's Fair 2024

Personality-Driven Development: Exploring the Frontier of Agents with Attitude

About this talk

Perpetual's Benjamin Stein explains how giving AI agents recognizable personalities, voices, forms, and workplace roles makes their capabilities easier to understand while raising customer expectations and introducing gender-related design biases. He argues that specialized virtual teammates also offer a practical engineering decomposition: narrower inputs and responsibilities reduce opportunities for LLM and tool-calling errors. The talk concludes with a vision of customizable AI employees whose behavior can be molded to individual business needs instead of fixed SaaS product assumptions.

Chapters

  1. 0:00Personality-driven development and anthropomorphized agents
  2. 1:31Voice, perceived gender, and AI product personas
  3. 6:03Customer expectations and familiar AI employee roles
  4. 10:38Agent specialization, tool calling, and generated-persona bias
  5. 17:11Customizable virtual employees and the future of business software

Talk transcript

  1. 0:00

    [upbeat music] All right, everybody. Uh, my name is Ben.

  2. 0:15

    Um, I'm, I'm going to talk about, uh, anthropomorphized agents. Um, I'm calling it personality-driven development, which is a kind of a cute name. And what we do at Perpetual is we build AI agents.

  3. 0:27

    We call them virtual teammates or AI employees. And we have really leaned into the idea of giving them forms and form factors. So level set on, on terminology, right?

  4. 0:38

    Anthropomorphization, which I had to practice saying a whole lot. I like A18N, which I don't know if it will catch on. That is when you give human traits to non-human entities, right?

  5. 0:47

    So Yogi Bear or Lightning, uh, McQueen, or Nemo, right? These are all anthropomorphized creatures. Um, fun fact, zoomorphization is when you do the same thing for animals. Not really relevant to this talk, but I thought it was kinda cool.

  6. 1:02

    Uh, Aslan is a good example of this, right? Aslan's a lion from Narnia, so it was a non-human entity that... Or a, whatever, a non-animal entity that got animal characteristics.

  7. 1:12

    You can tell your boss on Monday this is what you learned at the conference. So these are not new ideas, right? We've seen this in software for, for decades, right?

  8. 1:20

    Clippy, uh, you know, Her. I actually put this on the slide before, uh, you know, the recent, um, uh, OpenAI fiasco. But anthropomorphizing, giving technology a human form, very, very common, right?

  9. 1:31

    And now we're seeing this obviously a lot more with, with agents, but again, not new ideas. Um, right, so MailChimp, you know, twenty years, they've had a monkey as the form of a human that, like, sends email for you, right?

  10. 1:45

    So again, these are not new. But I would ask, right, so we have Siri, which is clearly not anthropomorphized, right? There's, there's no persona here. However, I'm gonna make you raise your hands.

  11. 1:56

    How many think Siri, or you could do Alexa too, is female?

  12. 2:02

    Anyone think it's male? Got one. Anyone think it's non-binary or think this is just a stupid question? Right. It's very weird, right? And this is sort of a, a lot of, like, my learnings and my observations is that we're all people, and we really like to ascribe, uh, uh, human characteristics to things, right?

  13. 2:21

    So well, suddenly this AI became female because it had a voice, and that's strange, right? So, you know, GPT-4o, is this female?

  14. 2:31

    Male? What about when you're talking to it? Suddenly it's like, oh, well now it has a voice, and so suddenly it has a prescribed gender. And these are just weird concepts, right?

  15. 2:42

    And they're, they're normal. There-- nothing profound here. But it's interesting when you're thinking about product development and software development, what your customers or what an audience is going to perceive on the other side.

  16. 2:54

    The foundational models, interestingly, have really moved away from any of these concepts. And, uh, you can guess some of the reasons. Some of them we'll, we'll talk about. But these are, like, the most modern, uh, representations of the foundational models, at least the ones that have consumer experiences, right?

  17. 3:10

    You have Copilot, you have OpenAI, you have, uh, Meta AI, Gemini. Like, this is how they're being represented, right? They clearly all share a, a single designer for some reason.

  18. 3:22

    Um, actually, I wanna go back to the female thing for a second. The, uh, we'll keep it here. When we got a Google Home a couple years ago, right?

  19. 3:28

    The little screen you put in your, your kitchen that shows your photos, and I was super excited, and I was, you know, showing my wife. And I said, "Hey, look, we can set a timer, and it can play music," and all of these things.

  20. 3:38

    And she looked at me with just this, like, look in her eye. She's like, "You cannot have a [REDACTED:gender] in the kitchen that you boss around." She's like-- And she was dead serious.

  21. 3:47

    She was like, "That is unacceptable. You can't just tell a [REDACTED:gender] what to do." And I was like, "It's not a [REDACTED:gender], right?" And she was like, "I don't care."

  22. 3:55

    And she was very, very serious. And I was like, "Wow, this is remarkable how, like, deeply ingrained this are." So I changed the voice to, like, an Australian [REDACTED:gender], and, like, now we're cool, and, like, everyone's happy. [laughing]

  23. 4:04

    But, like, a true story, right? So, uh, I think it's also helpful to think about the con- contrapositive, the, the opposite. Like, what is, like, an AI that's, like, doesn't have a form, right?

  24. 4:15

    So I thought this was just a great example, right? The, the AI that's inside Google Photos is just mind-blowing, right? Just, like, search for a dog, all the pictures of my beagle, like, lovely.

  25. 4:25

    Um, but there's no form here. You don't think about Google Photos as having, like, personality, right? It just, it just is. And the algorithm is, like, under the covers.

  26. 4:33

    Um, oh, another fun fact as long as I'm talking my family. So I showed this to, to my [REDACTED:age], Zeke, and he was like, "Oh my god, Google Photos has ChatGPT inside." [laughing]

  27. 4:44

    So if you wonder how, uh, you know, Gemini's branding is going with the youth, it's not. [laughing]

  28. 4:50

    Not good at all. [chuckles] So at Perpetual, we've really leaned into this idea of giving agents forms and personalities, and we really took it a couple steps, and we tell our customers, "You can give your agents their own forms."

  29. 5:03

    And so here's tech lead, right? Kind of a cyborg android type of a, you know, persona. And, you know, a member of the team writes code, does code reviews, things like that.

  30. 5:15

    Um, but really leaned in to say, "Listen, they can have a personality. They can have a form factor. They can have preferences." Um, and then things get real weird because our recruiter is, like, an artichoke, and it's like, oh, it's kinda like technology that's zoomorphized into a...

  31. 5:31

    Or whatever you do with a vegetable that now has human characteristics. And it's all very bizarre. But on the other hand, it like, oh, it actually kinda makes sense.

  32. 5:37

    You're like-- I mean, it makes no sense, but it's also like, I understand this. Like, it's a recruiter who has this form, and we have teams of agents, and these are hamsters, and they run the business.

  33. 5:47

    And we have the general manager and the graphic designer, and they, they represent, um, their own roles, right? On the one hand, it's very amusing. So the question would be, like, why are we doing this?

  34. 5:58

    And, well, I'll get to why we're doing it in a second. Let me talk about the expectations 'cause this was-

  35. 6:03

    This was surprising. So customer expectations, as soon as you put a form onto something, like get real, and they get real very quickly. So assumption number one is that you can chat with it.

  36. 6:15

    And there's nothing about workflow. I mean, when we're at the end of the day, like our agents, all we're do-- we're talking about is workflow, if we're being real, right?

  37. 6:21

    It's just like, it's just smart workflow. That's what we're all doing. There's no reason that you should be able to chat with, but all of a sudden, oh, it has a face, I must be able to chat with it.

  38. 6:29

    Oh, do I talk with it in, in Slack or in Teams? Like, wait, why do you think that should even be a thing? But like one hundred percent it is.

  39. 6:37

    Personality. Everyone assumes it's gonna have a personality, right? And generally, the baseline is like, you're a helpful assistant, right? So assume, oh, you're gonna be friendly, a helpful assistant, and we very, very quickly adopted that mental model.

  40. 6:50

    By, by we, I mean customers like have adopted this mental model of just like, "Oh, helpful assistant, that's great." Um, you know, we let customers make, you know, make them snarky, make them funny, but like at the end of the day, there's an assumption, and again, als-also weird.

  41. 7:02

    Like, why does software have personality? And suddenly with a face, it just needs it.

  42. 7:07

    Uh, users have no patience with these things going wrong, right? We all know this stuff goes off the rails, it gets wrong, but like all of software fails all the time, right?

  43. 7:14

    But you never hear people being like, "Fuck you, Google Sheets," right? But like, they curse those hamsters, you better believe it, right? It's just like, it's weird. You suddenly have this thing that you can get mad at, and like you can...

  44. 7:26

    A-a-and this is all just psychology. It's user psychology that is really, really innate. Um, and thinking about this from the perspective of product development, of, you know, software development, it's like, well, is there any reason that we actually do it?

  45. 7:40

    Like, it seems like there's... I mean, I already named some, some downsides, actually listed a lot more at the end. But, so why do we, why do we even bother?

  46. 7:49

    First off, these are really easy concepts to understand, right? When I tell someone, "Oh, you have an AI software engineer." Okay, like you instantly know what we're talking about.

  47. 7:59

    "Oh, it's an AI recruiter. Oh, okay, well, it'll probably like read resumes, and it'll probably coordinate interviews." And without any additional words, and we found that to just be an incredibly powerful, um, uh, what's the word?

  48. 8:12

    Like a jargon. Not even jargon. It's just like the terse way to describe the things that we're all doing without getting into really complicated discussions about React frameworks, right?

  49. 8:21

    It's just great. Branding, right? This is, I would say it's, I don't know if branding is quite the right word, but having a handle, like something to describe what it is, is also really, really powerful, right?

  50. 8:34

    I think most people think of like... I should just talk about AI for a second. Like call it like the algorithm, right? People th-- like know the word algorithm, and everyone thinks an algorithm is just like why my news feed is not in order, right?

  51. 8:44

    Like, oh, it's the algorithm, right? It's just this concept that is just very hard to, to grasp, very hard to grok. But giving something a name or, and a name that's something that we c- are familiar with, really, really powerful because now you can talk about it.

  52. 8:58

    And I was really struck by, I don't know, for the Android folks, uh, Google Now used to be-- All the phones are gonna buzz, or I guess all the phones that haven't been updated in three years are gonna buzz.

  53. 9:07

    Uh, Google Now was essentially what Google Assistant became, and it did all the same stuff. It set your timers and it, your reminders, and you would, you would talk to Google Now.

  54. 9:15

    But like, what was it? It was like, it was just like this weird conceptual thing. But all of a sudden it became an As- Google Assistant, and you're like, oh, again, it's a thing that like works on my behalf, and it, like you have this like almost corporeal understanding of like what it is.

  55. 9:30

    And I've actually been curious if Gemini makes this better or worse. Um, but again, Google Now and Assist- they're the same thing, but just that one nuanced difference made it really easy to understand.

  56. 9:41

    Price anchoring. So this is an interesting one. I don't know that this is like well tested in the, the field yet. All of this is so new. But if what we're talking about again is just like workflow and all of our agents are just doing like workflow, like what is the m-mindset for what workflow should cost, right?

  57. 9:56

    We look at comps and it's like, I don't know, fifty bucks a month. I mean, it's sort of like there's, again, making up the numbers, but if you think about it as a percentage, or if you're price anchored on a junior employee, wow, it changes the conversation, right?

  58. 10:09

    So you talk to an executive and it's like, yeah, one-twentieth of the cost, one-one hundredth of the cost, right? Suddenly it changes that nature of that pricing conversation. Again, I don't know that this is fully battle tested or if it will withstand tests of time, right?

  59. 10:22

    But like right now, it's fantastic. Is that a ferret? What is that? [laughing]

  60. 10:28

    Otter.

  61. 10:29

    An otter? Hmm. [laughing] Um, that's an awesome picture. Uh,

  62. 10:38

    so the other, uh, uh, reason that this is a really helpful construct is it is a way to decompose problems, right? So thinking about if, you know, wearing an engineering hat for a moment, right?

  63. 10:49

    What is engineering? It is abstractions about the real world, it is getting the right levels of abstractions, and it's about decomposing problems. Like that's just all we do all day long, right?

  64. 10:58

    And in a sense, this is like an arbitrary way to break down a problem. On the other hand, it's a really useful way to break down a problem, which is what we've learned, right?

  65. 11:06

    So specialized agents have... Heard a bunch of talks today talk about specialized agents versus generalized and how the specialized ones perform better, right? It's like, oh, we have a finite set of tools.

  66. 11:16

    We have a small number of inputs. Like there's just less chance for, uh, LLMs who are trying to interpret or do tool calling to get things wrong. Um, and so it just happens to be a really convenient way to break down problems and also to scale because we can keep subdividing, um, like, uh, agent problems into more

  67. 11:33

    and more specializations. And so it's just very natural. Again, arbitrary, but like even for me, just like, oh, this is very helpful. I understand that my AI engineer writes code, and I understand that my AI copywriter, you know, writes compelling copy, and like great.

  68. 11:46

    I can get my head around that very, very easily.

  69. 11:50

    And lastly, it's just fun, right? Which is, you know, that's sort of a, a company branding question, but like we like making our video game avatars, right? You spend more time making your like eyebrows correct in like the Nintendo Switch than you do playing the game.

  70. 12:02

    We roll characters in D&D. Like, so like why not have fun, right? So that's sort of, uh, my personal, um, perspective on this. Let's talk about the downsides. And there's not gonna be any cute pictures because, you know, this is, like, the sad part.

  71. 12:15

    So [laughs] okay. So one of the big glaring ones is, like, we are just inviting inclusivity and, uh, stereotype challenges, right? We, we're just asking for it. Um, one thing that was fascinating is, you know, all those avatars were, were generated, right?

  72. 12:30

    So a customer can pick their form, and we, we ought to generate it. One hundred percent of the software engineers are generated with, like, neckties, like, looking like men.

  73. 12:39

    That's just, like, what comes right out of DALL-E. And, like, I'm not gonna have any, like, perspectives, but, like, it's just like, it's all of them do. So we are just inviting this onto ourselves, and is that worth it, right?

  74. 12:49

    Is that worth it in a, in a work context, um, to really just invite those questions?

  75. 12:55

    Um, expectations of performance. Uh, I don't know what it is because we assume, like, software's gonna always just, just work. Um, but there's just this expectation that, like, these agents are gonna perform really, really well, right?

  76. 13:07

    There's just this bar. It's like, "Oh, well, my..." You know, even though our, you know, our junior employees don't perform well, there's an expectation that these things are going to perform at a very high bar.

  77. 13:15

    It's just what we've seen. It's like, yes, of course, it's gonna get it right a hundred percent of the time. Um, the features I alluded to before, it's like, why are we spending our time building, like, chat interfaces and all of these, these things that, like, it, it, it's, it's just, it's strange.

  78. 13:28

    Like, you almost have to if you're building this type of personified agent. Um, but it's just 'cause it's ex-expected.

  79. 13:36

    Um, certainly as a, a startup, we haven't had to deal with this, but I'm going to guess that when we want to walk this back and rebrand, like, holy shit, right?

  80. 13:44

    Like, there's just, like, walking back when all of our customers have these perso- is gonna be very, very difficult, right? This is not just, like, changing some colors, right?

  81. 13:51

    This is a major, uh, you know, stake in the ground that we, we would be planting.

  82. 13:57

    Uh, and lastly, it's, it could be a distraction, right? Like, all of these fun stuff that we're talking about around personalities and, like, um, chatting and, um, preferences, it's a distraction from the actual business value, which is, like,

  83. 14:11

    document review and data extraction, or the thing that actually someone would pay for. Um, it can be distracting. Um, but at the end of the day, honestly, the biggest downside right now is that it is just a very stark reminder that you're replacing jobs.

  84. 14:24

    And, like, the, uh, outside the scope of the talk, whether or not that is a good thing or a bad thing or an inevita- inevitable thing, right? However, it is a reality that as soon as you, you bring this up with a, a, a prospect, like, first thing on their mind.

  85. 14:38

    And I'll, I'll just be real. I'll tell a story. So very recently, I was, you know, pitching sort of, you know, C-suite pitch, right, to a CEO. It was like, "Oh, this is, uh, you can't scale your business.

  86. 14:48

    You, you don't have enough people. You could never hire, uh, enough people to scale the business to meet your aspirations. Have we got a solution for you? It's also shit work, like, no one wants to do it.

  87. 14:56

    Like, it's great." And he was just loving it. He's like, "Yeah, this is, like, exactly what I need, and we-- at the cost, it'll be great." And so set up our first, uh, design session, and we get into the meeting, and he has, you know, one of his, you know, like, an IC, uh, on the team.

  88. 15:12

    He's like, "Oh, yeah, well, I don't do any actual work. I'm, like, an executive," right? "I brought the person who knows what they're doing. This is Ben, and he has software that has virtual employees.

  89. 15:21

    Can you please tell him what your job is?" Uh, like, my jaw dropped. I was like, "Oh, sh..." Like, I was not at all prepared for that, right? Just like, "Oh, I can't talk to you with a straight..."

  90. 15:30

    Like, whether or not I'm going to actually make you more productive and you will do better in your... Like, just to lead that, with that in the conversation was just, like, really, really hard.

  91. 15:39

    So I have to, like, very, very quickly, like, on the fly, try to, like, walk back that concept of like, "Oh, no, no, no, like, this is, this, this is actually gonna help you."

  92. 15:44

    And, like, whether, whether we don't know what the future's going to hold, but, like, we are definitely, like, uh, leaning heavily into this, um, and it's potentially a huge hurdle.

  93. 15:55

    Um, got a couple more minutes. I wanna talk about, you know, so this was the, was the title, Personality-Driven Design. And there's this, uh, I don't actually even have, like, the words for this.

  94. 16:03

    I'd be curious afterwards if, if folks do. The software we're building and our ability to, you know, create these, uh, roles, these virtual employees that have job descriptions and forms and personal preferences and, like, there's no, like, checkboxes in, inside, like, the configuration.

  95. 16:20

    It's all very prompt-driven, right? But it's a way to inject nuance and business logic into these, um, into these agents with zero configuration, which means that every single instance is, like, one hundred percent, excuse me, bespoke for each customer, which is a really wild concept.

  96. 16:38

    When you think about, like, "Oh, can I reproduce this bug? Does it work on my machine?" It's like, well, of course not. Like, of course, this doesn't work. It's, it, it, it, every single instance is bespoke, right?

  97. 16:47

    So here's just, like, maybe this is not a real one. This is an example, right? A giraffe, right, who has a personality. It's charismatic. It's entertaining. You review resumes.

  98. 16:55

    You read cover letters. You know, you do that, like, operational work, and here's where it gets super interesting is the preferences, right? Oh, as a hiring manager, I want to tell my recruiter my priorities, and I like people who went to Ivy League schools, and I don't like job hopper.

  99. 17:11

    I, I sound like such an old [REDACTED:gender]. And I don't like people who hop jobs, and, like, I like cover letters. Like, okay, well, I can train, you know, lowercase t, train my virtual employee, uh, the, the, the way I want to work and the way I want them to work.

  100. 17:25

    And so this is just an amazing concept that, like, I, again, I really don't have my head around, like, what it means for, like, the future of software when an entire thing can be molded to meet the business's need, not just, like, oh, what the PM of the SaaS platform, like, happened to, like, think was a good

  101. 17:39

    checkbox, right? And so this is sort of, like, the future that I'm really, really excited to be working on, um, period, full stop. Awesome. Thank you, everybody. [upbeat music]