← All AI Engineer talks

AI Engineer World's Fair 2024

AI Music Generation: From Prompt to Production

Phlo Young54:34

Read the talk

AI Music Generation: From Prompt to Production

Follow Phlo Young from voice conversion and text-generated songs to personalized prompts, lyric preparation, and hands-on control of separated vocals and instruments.

From a talk by Phlo Young

Start with the listener’s taste

Have you used a music-generation model—and what would you want it to make? Phlo Young opens with those questions. About half the room has tried music generation. Boomy draws little recognition; Young remembers it as an earlier, less complex tool. Suno and Udio, the workshop’s main tools, are more familiar. Asked for a favorite artist, one listener names Miles Davis, and the room supplies the genre: jazz. That exchange establishes the practical problem: translate someone’s musical taste into something they can hear.

Young approaches the task as a self-taught songwriter, not a classically trained musician. He wrote his first song at thirteen and welcomes corrections. QR-linked notes, live questions, and a short poll let participants steer the workshop toward their interests.

His experimental approach predates generative music. After learning Python for automation during the pandemic, he also began testing music distribution: release the same recording under different artist names, vary the cover art, marketing, and audience targeting, and observe how it performs. Young reports that one instrumental pen name accumulated about five million streams over the preceding couple of months. His own background spans R&B and hip-hop, with EDM and country-influenced experiments; generation adds another way to explore those musical directions. Distribution and artist income motivate him, although the workshop does not develop a release or monetization procedure.

The poll favors prompting and tweaking. A more advanced option—download generated music, separate stems, then mix and master—also attracts interest. Young points to Suno and Udio’s Discord communities, including an unnamed contributor whose unusual banjo and bagpipe prompts produce striking results. The useful premise is that effective musical descriptions may be less intuitive than simply naming an instrument.

1:361:53
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

1:36 · section reference included

Separate three meanings of AI music

The input tells you which part of music-making the model is doing. Young found “AI music” frustratingly ambiguous because it could describe a generated song, an expanded recording, or a familiar voice applied to someone else’s performance.

ApproachInputTransformation
Text to musicDescription, optionally lyricsGenerate a sample or song
Audio to musicExisting soundTransform or develop it into music
Voice conversionRecorded voiceChange its vocal identity

Young uses style transfer for the third category. In the examples that first attracted his attention, consumers often called voice conversion “AI music” without distinguishing it from generating the composition itself.

10:0010:18
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:00 · section reference included

A human performance, a different vocal identity

The first video makes the distinction audible. Its creator, whom Young tentatively identifies as Roberto Nickson, begins with a Kanye-style voice and a converted excerpt of Day ’N Nite. He then finds a Kanye-style beat on YouTube, writes eight bars, and records them himself. The workstation view shows the human performance entering an ordinary audio-production workflow.

Audio workstation with waveform tracks and an inset of a man performing into a microphone.
Recording vocals alongside an audio-track timeline.

Next comes the conversion: the same recorded performance plays back with a Kanye-like vocal identity. The beat, writing, and reference performance were already supplied by a person. Voice conversion changes the voice without replacing the entire songwriting process. Seeing this sequence pushed Young into Discord servers and subreddits to learn how to reproduce it.

Ghostwriter’s viral Drake-and-The-Weeknd-style song illustrates the same distinction. In Young’s account, Ghostwriter wrote and produced the song and converted his own vocals; the workshop plays an excerpt rather than demonstrating its creation. Young recalls finding it below 1,000 views and estimates nine million streams in its first 24 hours. Those are his recollections, not verified platform totals. His attribution of the removal to the RIAA also differs from contemporaneous reporting, which identifies UMG in the YouTube copyright notice. The mechanism remains clear even where the viral history is uncertain: an authored recording became associated with recognizable voices through conversion.

11:5812:06
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:58 · section reference included

Voice conversion in the artist’s hands

Listening again to the earlier Kanye example, Young now hears conspicuous artifacts. What initially astonished him sounds less natural after subsequent improvements in voice models. He credits experimentation within communities that included teenage model trainers; his description of the tools becoming “10 times better” expresses that listening impression, not a measured quality improvement.

The Randy Travis example changes the purpose of the technique. Travis had lost his voice, and his team used conversion to make a new recording available in his vocal identity. Young plays an excerpt and describes unusually positive listener reactions, including a comment from an older fan moved to tears. Against the hostility he often encounters around AI music, this artist-directed use offers a different relationship between performer and technology: the recognizable voice helps the artist continue making music.

16:4216:54
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

16:42 · section reference included

When the prompt supplies the starting material

Text-to-music moves the starting point upstream. Young turns from converted recordings to full generated songs, beginning with BBL Drizzy in the context of the Drake–Kendrick dispute. The distinction matters more than the meme’s backstory: this example starts with words rather than a recorded singer. The taxonomy slide separates that route from transforming existing audio or converting a voice.

Slide titled “What is AI Music?” listing text-based generation, transformation of existing sounds, and audio filtering or voice conversion.
Three approaches: text to music, audio to music, and style transfer.

Young describes comedian King Willonius entering a prompt into Udio and, as he recalls less confidently, supplying lyrics. In his account, there was no studio vocal performance to replace: text went in and a song came out. He characterizes it as a one-shot generation and emphasizes its subsequent reuse in other songs. That is the workshop’s clearest contrast with Ghostwriter’s human-authored reference performance. Copyright questions arise here but remain unresolved.

A second example comes from ElevenLabs: a text-generated song about programming and GPUs. Its lyrics describe writing code and hoping to make it work, making the result a self-referential example of an AI system singing about AI development. Young reports that it circulated widely on Twitter and groups it with the meme-song examples.

20:1020:20
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

20:10 · section reference included

Beatboxing and desk percussion become arrangements

Audio-to-music keeps a performed sound at the beginning of the process. A Stable Audio example starts with beatboxing and transforms that input into music. Young remembers it as the first publicly available audio-input feature of this kind he encountered, while distinguishing a musical output from a necessarily complete song. A second clip, from the Suno team, starts with fingers drumming on a desk and produces rock music with guitar and other instruments. The attraction for a musician is immediate: a rhythm can be performed directly instead of translated entirely into descriptive text.

An audience member hears a Tom Petty resemblance in the preceding generation. That observation arrives alongside Young’s discussion of complaints alleging that Suno and Udio trained on copyrighted music. This is a historical workshop, held at the AI Engineer World’s Fair on June 26, 2024; the interface behavior, account options, and service integrations that follow belong to that demonstration. Young has no affiliation with the featured services. He points technically curious participants toward music-model papers in his notes, but follows the poll’s preference for practical work rather than explaining model architectures.

24:5325:04
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

24:53 · section reference included

Give everyone a way to generate

To let participants work alongside him, Young distributes business-card codes. Entering a code through the green Gateway button reveals credentials for workshop accounts. The demonstrated login choices include Google, Discord, Twitter, and Apple for Udio, and Discord, Google, and Microsoft for Suno. He has paid for accounts to expand the room’s generation allowance, although his “Premier” and “unlimited” terminology should not be read as Udio’s documented plan names or limits: its original subscription announcement describes finite-credit Standard and Pro plans.

Apple’s two-factor authentication may obstruct some logins, but participants confirm that they can get in. The participation loop is straightforward: generate a song, copy its share link, and submit it through the workshop’s Google Form. Young proposes randomly selecting two participants to keep one account each. He has purchased a month of access rather than annual subscriptions.

29:0329:26
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

29:03 · section reference included

Build Patrick’s birthday song

Patrick has a birthday that week, so the first live task becomes a personalized birthday song. Young initially includes a Randy Travis reference and anticipates artist-name filtering. An audience member suggests changing the spelling and adding the associated musical style; another supplies a more useful personal detail: Patrick is a patent attorney. Patrick names The Band as a favorite, and the room settles on rockabilly as a prompt description.

Young starts generation and carries the same prompt across services. The artist reference becomes a concrete failure case: a rejection names Randy Travis, after which he tries the altered spelling Tandy Ravis. This does not establish a reliable way around moderation. The useful prompt ingredients are already independent of that attempt: the occasion, the person, his occupation, and a musical direction.

32:1232:30
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

32:12 · section reference included

A songwriting partner and an in-house producer

Young thinks of Udio as a songwriting partner. He describes its default workflow as generating roughly 30 seconds of music at a time, with an experimental model available in the paid account shown. That short result becomes an object to revise rather than an answer to accept wholesale.

ActionPurpose
ExtendContinue a promising idea
RemixExplore a different version
DiscardAbandon an unhelpful result
InpaintModify a selected section
Upload audioStart from an existing recording

For someone with original music, Young recommends bringing the recording, tempo, key, and lyrics into the process. He has tried this with his own songs and found the reinterpretations surprisingly compelling.

Suno, by contrast, feels to him like an in-house music producer: its output often has a polished, Top 40 character. His suggestion that this might come from a hidden system prompt or the training data is speculation, not knowledge of its internals. While onboarding friends and family, he noticed a more actionable issue in custom mode: lyric lines with poorly controlled syllable counts could fall badly out of alignment with the beat. He reports that his more recent generations handled timing and song structure better. The comparison is about his working experience—iterative development on one side, production polish on the other—not a controlled quality ranking.

35:2935:43
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

35:29 · section reference included

Generate an event song and listen to the birthday results

The next prompt is a song about the AI Engineer World’s Fair. Organizer Benjamin Dunphy’s taste suggests old-school hip-hop and boom bap. The demonstrated procedure is short:

  1. Sign in to Udio and open Create.
  2. Enter the song description and musical direction.
  3. Choose generated lyrics, supply custom lyrics, or request an instrumental.
  4. Choose an available model and start generation.

The free-account default shown is Udio 32; the longer experimental option belongs to the paid experience in the demonstration. A compact prompt expressing the spoken request is:

A song about the AI Engineer World's Fair.
Old-school hip-hop, boom bap.

Young then uses the same prompt in the other service. Before starting, he refreshes the page because he has encountered a state problem: repeated generation could unexpectedly extend a previous song instead of starting a new one.

Returning to Patrick’s songs, Young first encounters another moderation error associated with the Randy Travis reference. Other results do play: the vocals name Patrick, celebrate his birthday, and refer to rock and roll. Patrick calls one result the “Theme song to my sitcom.” On Suno, clicking the song name exposes its lyrics. A later alternate playback seems to repeat the same song, and Young leaves open whether the service generated it twice. The session produces enjoyable personalized music, but neither the birthday playback nor the event-song experiment becomes a systematic comparison of the services.

39:1639:24
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

39:16 · section reference included

Prepare and edit lyrics before generating audio

With time running short, Young turns to the workshop notes, where he also plans to publish the recording and transcript. A GPT-4 prompt template takes a Song Description and asks for structured lyrics. The displayed notes include style tags, an intro, verses, a chorus, an instrumental break, and an outro. Young flags a GitHub rendering problem around brackets, but the workflow is still clear: have ChatGPT draft the lyrics, edit them yourself, then give the music model the revised text. Separating lyric preparation from audio generation creates an explicit revision step before the musical performance is generated.

GitHub notes showing a GPT prompt with song description, style tags, intro, verses, chorus, instrumental break, and outro, followed by music prompt examples.
Workshop notes provide a structured lyric prompt.

For Patrick’s example, that preparation step can be expressed as a reusable instruction:

Song Description:
A rockabilly birthday song for Patrick, a patent attorney
who loves rock and roll.

Draft lyrics with an intro, verses, chorus,
instrumental break, and outro.
Keep the lines singable and the chorus easy to repeat.
Return the lyrics so I can edit them before music generation.

The song description supplies the subject and direction; the separate lyric draft gives the user a chance to change phrasing before spending another generation on audio. Young also points to official Suno and Udio tips and returns briefly to the community “Prompt Master,” without completing that deeper walkthrough.

44:0244:12
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

44:02 · section reference included

Generic voices, emotional delivery, and descriptive labels

Philip asks whether conversion can produce a generic rapper’s voice instead of a recognizable performer such as Eminem. Young proposes two approaches:

  • Blend voice sources: combine material from several speakers rather than targeting one identity.
  • Use a synthetic source: obtain a TTS voice and use that output as training material.

Young suggests that some unspecified tools can begin with six seconds of voice, illustrating a blend with equal short samples from different people. He names no model, training conditions, or quality criterion for that minimum. Neither blending nor synthetic input by itself establishes permission to use the source material or clears rights in the result.

A responsible-use question gets a non-lawyer disclaimer rather than substantive ethical guidance. Another question asks about preserving emotional tone, citing the changing delivery in Kendrick Lamar’s Euphoria. Young recommends RVC as his preferred voice-conversion tool at the time. He emphasizes that cloud execution is an alternative to running it locally, where hardware requirements may be an obstacle. He also confirms that participants can have the slides.

After an audience member reports that the shared Google document requires an access request, Young acknowledges the permissions mistake. His final concrete prompting technique is to search Rate Your Music by artist, genre, or desired sound, collect the descriptive labels attached to relevant music, and use those labels in Suno or Udio prompts. He tentatively extends the suggestion to Stable Audio. The practical move is to replace a vague taste reference with musical vocabulary. His further claim that the models were trained on Rate Your Music labels comes from unnamed reverse engineering and is not established training-data provenance.

Young acknowledges the distinction between ethics and law but defers that discussion for lack of time. He reports twelve song submissions and proposes a later drawing, possibly streamed, followed by changing the shared credentials before handing accounts to the winners. That handoff remains a plan. The apparent closing then plays the World’s Fair song, whose lyrics include copper wires, coders, blueprints, and robots. Young reacts positively and wonders what Udio produced, without completing the comparison.

45:5346:10
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

45:53 · section reference included

Separate the song into parts you can control

A late question brings the workshop back to production: how do you pull out stems and edit them? Young returns to Patrick’s birthday song and demonstrates the historical Suno-to-DAW route. Starting from the song’s link, he changes the domain from suno.com to sunodaw.com, retaining the song path. He explicitly corrects his initial wording: add daw to the domain, not to the end of the whole URL. After signing in, the service imports the song for separation. Young estimates 30–90 seconds for import and separation, sometimes longer for more complicated songs.

A digital audio workstation, or DAW, gives the generated song a different form of editability. The imported mix becomes source-separated audio for vocals and instrumental parts, with individual level controls. These are estimates extracted from the mixed recording, not recovered original recording tracks or guaranteed isolation of every instrument. Young first describes keeping music he likes while reducing unwanted vocals, then plays vocals alone and identifies the bass. The task has changed from asking for another whole song to listening to, and adjusting, particular components of this one.

WavTool timeline with stacked audio tracks, a blue selected region, a vertical playhead, and two waveform channels below.
WavTool displays multiple audio tracks and an expanded waveform.

Young identifies the service as WavTool and tentatively describes its accounts as free at the time. For a local alternative, he recommends Ultimate Vocal Remover, or UVR5, an open-source application with a more involved setup and hardware considerations. Its input does not have to be AI-generated: a conventional mixed song can also be separated. The demonstration stops short of a complete mix or master, but it reaches an important production boundary. Patrick’s birthday song is no longer only an output to accept or reject; its vocal and instrumental parts can now be auditioned and balanced separately.

50:3750:47
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

50:37 · section reference included

Resources

From the talk

Read the complete timestamped transcript
  1. 0:00

    [upbeat music] Good morning, everyone.

  2. 0:15

    Good morning. [laughs] Um, I'm excited to be here. I, I really have, uh, one main goal in mind, and that's for this talk to be-- or this workshop to be interactive.

  3. 0:30

    I want it to be fun, and I want you [REDACTED:gender] to walk away with something tangible. Um, so I, I... We, we can actually switch, switch to the slides.

  4. 0:38

    Uh, slideshow. It should be live now. If you [REDACTED:gender] wanna go ahead and go to this link that I have up on the screen or use your phone to go to the QR code, I have some notes for you [REDACTED:gender] at that URL.

  5. 0:52

    Um, but, but tho-those are my stated goals for this workshop. I wanna really probably start by, uh, asking you [REDACTED:gender] some questions so I can get some interactions going.

  6. 1:03

    But I'm glad you [REDACTED:gender] came. Thank you for, for, for coming. I really do appreciate it. Um,

  7. 1:08

    the music stuff is something I have been into for a long time, like lifelong musician, artist. Uh, but the AI music stuff is something I just started paying attention to, um,

  8. 1:19

    with, with regard to the tools I'm gonna show you today maybe about a month or so ago. So basically, what this workshop is gonna be about is I'm gonna start with the insights that I've gained by, like, doing all the research over the last thirty days, and then we'll get straight into the tools so you [REDACTED:gender] can,

  9. 1:32

    again, uh, walk away with something tangible. Um,

  10. 1:36

    let's start with... Okay, by, by-- I know it's early and, and, uh, I'm still groggy myself, [laughs] so, uh, you know, you don't have to necessarily go crazy here. But by show of hands, uh, how many of you have ever used a music generation model, period?

  11. 1:53

    Okay, we got... Okay, nice. About half, half of you [REDACTED:gender]. That's dope. Um,

  12. 2:00

    b-by show of hands, how many of you [REDACTED:gender] used-- I had someone in the, in the crowd already say that they used Boomy. Have you [REDACTED:gender] ever used Boomy before?

  13. 2:07

    No, no, no Boomy? Okay, that'[REDACTED:url]- You're an OG if you used Boomy 'cause it was from like two years ago. [laughs] But it was a music generation model that, um, was actually pretty cool, but, but didn't get as complex as the, the music generation models we have today.

  14. 2:20

    How many of you [REDACTED:gender] have, have-- H-how many of you [REDACTED:gender] have ever used, uh, Suno or Udio, by show of hands? Okay, nice. Okay, so you [REDACTED:gender] will be familiar with, with what I'm gonna do today.

  15. 2:30

    Uh, hopefully, I can show you [REDACTED:gender] some, like, advanced prompting techniques at the least. But for the most part, that'[REDACTED:url]- tho-those are the tools that we're gonna, uh, be focused on.

  16. 2:38

    Um, I was doing this earlier, asking some of the AV crew what you [REDACTED:gender]' favorite artists are. My [REDACTED:gender] here, right, right here in the middle, if you could name, uh, your favorite artist of all time, what, what would it be?

  17. 2:49

    Who would it be? Me? Yes. Miles Davis. Miles Davis, okay. And correct me if I'm wrong 'cause that's a little bit out of my league. Like [REDACTED:url] Blues, R&B, jazz?

  18. 3:02

    Jazz. Okay, okay. I, I think I have something special for you today then. So,

  19. 3:07

    uh, let's, let's start. I can get into,

  20. 3:14

    uh, a little bit of what we're gonna do as well.

  21. 3:18

    And I, I actually would like to say-- 'cause I've already talked about my workshop goal. What I would actually like to say is that, um, I am not a classically trained musician.

  22. 3:26

    I've-- I, I wrote my song, my first song when I was [REDACTED:age]. But I welcome any discourse you [REDACTED:gender] have in terms of, like, uh, accuracy, right? So if you [REDACTED:gender] feel like, uh, you're actually wrong, like, I-- that's cool.

  23. 3:38

    I, I welcome that. [laughs] And, uh, I actually have, uh, something for you on that link that I showed earlier where you can send me questions or comments li-- like, live while we're doing this, um, to either correct me or, or send me the question or whatever.

  24. 3:50

    So that, that was one of the main things I wanted to say is, like, I'm-- I, I try and be as accurate as possible, but if I say something wrong, you know, feel free to call me out.

  25. 3:57

    I, I'm, I'm comfortable with that. I'm cool with that. Um, and actually, let me make sure I have my stuff pulled up.

  26. 4:06

    If you [REDACTED:gender] can go to this link, there's a link on there that says Poll. If you [REDACTED:gender] c-can go, uh, fill out that poll. I think there's, like, two or three questions on there.

  27. 4:15

    Uh, if you [REDACTED:gender] could go to that link now and start submitting your answers, I'd really, really appreciate that 'cause it, it-- I wa-- again, I want this to be really interactive, but I also, uh, wanna tailor it to specifically what you [REDACTED:gender] wanna know about, uh, AI music generation.

  28. 4:32

    So if you [REDACTED:gender] could make your way over to that link and [REDACTED:url]- and, uh, submit, would really, really apprecia-- uh, appreciate it.

  29. 4:40

    Now that I've done some housekeeping, I can kinda get into, um, my journey here. I would say I'm way more of a... My-- First of all, my name's Phlo, as you [REDACTED:gender] may or may not know.

  30. 4:51

    Um, I'm way more of a songwriter than, like, a software engineer. Uh, but I have a lot of respect for what you [REDACTED:gender] do. I, I, I know enough Python to be dangerous, but beyond that, [laughs] I'm not really that good at, uh, software design or, or engineering.

  31. 5:04

    Um, but I, I, over the pandemic, uh, took a lot of time out to learn how to write scripts. And a-- Um, I'm really, really big into automation, so I use Python for automation for a lot of my stuff.

  32. 5:15

    Da-daily personal life stuff, uh, and, and, uh, stuff for work as well. Um, the other thing I would say is

  33. 5:25

    I'm really, really big into experimentation. Um, Phlo is actually my artist name, and if you go on Google or YouTube or Spotify, Apple Music, you can actually look me up and find my music.

  34. 5:34

    Uh, but the truth is that, like, I have a very eclectic taste, and I, like, have, uh, experimented in a, in a lot of different genres. So I, um...

  35. 5:45

    Couple years ago, I decided that I wanted to do some split testing, which is pretty odd for an artist. Uh, but I, I did wanna do some split testing.

  36. 5:52

    And what, what I really found out was that the same song under two different artists can behave wildly diff-- or can perform wildly differently. So I literally would take the same recording, like I'm a self-taught engineer, would record at my home or other studios, and upload it under two different artist names.

  37. 6:09

    And depending on the cover art, the marketing, uh, who you try and target when you're, when you're posting online in terms of the daily content, I would see that songs would perform differently.

  38. 6:18

    Um, so what I basically did was start, uh, creating these quote-unquote pen names and putting out music that way. And I actually have a pen name, uh, that has way more streams than my actual artist name.

  39. 6:28

    So if you look up me as an artist [laughs], I don't have that many streams, but I have a couple pen names that, uh, have been, like, viral. I have one that's got, like, 5 million streams over the last couple months.

  40. 6:37

    Uh, and it's just instrumental music. So I, I'm an artist that does, like, R&B and hip hop, but I've experimented with, like, EDM. I actually have a song with some, like, country influence, believe it or not.

  41. 6:45

    Uh, so I've just b- experimented a lot, um, and, and, and that's kinda how we got here today. I was-- I like to say I was experimenting with, uh, multiple genres way before AI music got here, but I, I love AI music for, for that reason.

  42. 6:58

    Um, in this workshop we'll cover, uh, AI music generation and demystifying it. I kinda like distilling better, um, but you know, ChatGPT gave me this word, so I put it in the slides. [laughs]

  43. 7:09

    Um, and we're gonna get hands-on with Suno and Udio, uh, and talk about turning ideas into songs. And, and then, uh, actually I, I have a question about this later on, but, um, I really wanna talk about opportunities for using AI music.

  44. 7:22

    I'm super passionate about, like, distribution, um, and, and for, like, regular everyday artists to be able to make, um, an income from, from their music. So if any of you [REDACTED:gender] are actually artists, like, I would love to get into, uh, those type of, uh, discussions.

  45. 7:36

    Um, let's see. Uh, have you [REDACTED:gender]... Did anybody actually fill out the thing that I was asking about? Let's see.

  46. 7:46

    Oh, I'm disconnected. Break, uh, leg. 'Cause I really can start throwing in little tidbits here and there about, um,

  47. 7:56

    what you [REDACTED:gender] wanna learn about. I put what, like, one of the options that I put for the poll is, like, an advanced option where we can, like, actually pull down some of the music that Suno is generating, uh, get stems for that, that music, and start to, like, mix and master.

  48. 8:09

    Um, so I don't know if you n- if, if any of you [REDACTED:gender] would actually be interested in that, but I put that as- That'd be great. Yeah? Yeah.

  49. 8:14

    Okay. Awesome. Yeah. Th- th- that was the hardest thing about, uh, preparing for this, is like, I just don't know exactly who's gonna be in the crowd, and what, like, what level they're at.

  50. 8:22

    So I tried to, like, uh, come up with a couple different tracks just for this, this, uh, presentation alone, but, um, yeah, I just needed you [REDACTED:gender]' feedback. So, oh, nice.

  51. 8:30

    We did get a bunch of submissions. And it looks like, um, prompt techniques. Okay. Tweaking. A lot of prompt techniques. Okay. Cool. I have this [REDACTED:gender] that I call, like, the prompt master.

  52. 8:42

    H- how many of you g- by show of hands, how many of you [REDACTED:gender] are in the Suno or Udio Discord server? Nobody. Oh, yeah. Oh, you're in there?

  53. 8:48

    Okay. Nice. Yeah. Even for you, I think I'll have a treat. 'Cause like, uh, there's a couple people in there that have posted, uh, since the beginning of, uh, Suno and Udio, and then [REDACTED:url] there's this one [REDACTED:gender] who's, like, been able to do some really, really cool, like, banjo music and bagpipe music that's insane.

  54. 9:02

    But hi- the prompts are very unintuitive in, in my opinion. So you kinda have to see it to, you know, understand what it's doing.

  55. 9:10

    Um, so we'll definitely get into the... I'll, I'll make sure that I include a lot of prompting stuff. Thank you for everyone that just came in. I appreciate you [REDACTED:gender] joining.

  56. 9:18

    Um, right. So this actually gets into-

  57. 9:25

    You might notice that I sound-- You might notice that I sound-

  58. 9:27

    Okay. It did actually work, but I'm, I still might play it from my computer. So

  59. 9:35

    this is where we start with the insights that I've learned by studying AI music and how it got to this point over the last, uh, couple months. And then I'll get into...

  60. 9:45

    It, it, it's, uh, kinda painful, but I have to play some videos for you, for you [REDACTED:gender] to catch up to where I'm at in terms of, like, understanding and how this whole thing developed.

  61. 9:52

    Um, so we'll have some videos to watch, uh, but maybe five or six minutes of, of, of videos. Um,

  62. 10:00

    when people say AI music, I think that can mean a lot of different things. And when I first started doing my research, uh, it was really frustrating that someone, uh, e- everyone, everyone meant something slightly different.

  63. 10:18

    So what I've tried to do is compartmentalize the definitions of AI music, and tidy up, like, like, I guess really tidy up the definitions, 'cause there was a lot of, like, blurred lines, uh, when I started doing the research.

  64. 10:29

    Um, and, and the way that I, I, uh, compartmentalized this was to say that we have, uh, text to music, uh, where you're generating short, short samples, um, and full songs using text, text prompts.

  65. 10:42

    You have, um, audio to music, and this is where you're taking an existing sound and transforming it into, into, uh, actual music. And you have, like, style transfer, uh, which is basically voice conversion.

  66. 10:52

    Um, and voice conversion is actually the first thing that I came across, uh, when I found AI music, uh, last year, actually. So it's been, like, a long time coming, but I didn't really start doing the research until recently.

  67. 11:04

    But I think what most people call AI music is really actually just voice conversion. What, what... It's hard to say that now, but maybe when I started doing research it was, but, but I would say that, um, the, the general consumer, w- what they call AI music is, is, is mostly just style transfer, where someone, like, took

  68. 11:22

    their voice and converted it to another voice. And that actually gets into the first video, and I'll go ahead and play it now.

  69. 11:30

    You [REDACTED:gender] may have seen this video. It's actually pretty cool. Text to music.

  70. 11:38

    Uh.

  71. 11:45

    You might notice that I sound like Kanye West. No, you-

  72. 11:48

    Could we, uh, turn it up just a little bit in the house? 'Cause this one's a little bit low. He gets a little louder later, but you'll, you'll catch it

  73. 11:58

    You might notice that I sound like Kanye West. No, Yeezy didn't record a voiceover for me for this video. I didn't learn how to do impressions. This is AI.

  74. 12:06

    So let me come back to my original voice for a second because this is crazy. Today on Metaverse, I posted AI Kanye covering popular songs. Here's an example of him singing Day 'N Nite.

  75. 12:15

    Day 'n Nite. Whoa. I toss and turn, I keep stress in my mind.

  76. 12:19

    So that's clearly crazy, and I started thinking, you know, "What are the implications of this for the music industry?" Now all you have to do is record reference vocals and replace it with a trained model of any musician you like, which is exactly what I did.

  77. 12:30

    I found this Kanye-style beat on YouTube. I wrote eight bars, and I'm gonna record them now, and then I'm gonna have AI Kanye replace me. I got a fantasy that's beautiful, that's dark and twisted.

  78. 12:41

    But I attacked [REDACTED:location] whole religion all because of my ignorance. What was I thinkin'? That was some bitch shit. I lost Adidas, but I'm still Yeezy. Back in [REDACTED:location], [REDACTED:gender], I'm a genius.

  79. 12:50

    [REDACTED:gender] in the hood just like I'm Eazy. Kanye Weezy, [REDACTED:location] of Chicago, life ain't easy. All praise be to Lord Jesus. Ghanda, please rest easy. All right, let me cut it there.

  80. 13:00

    Let me cut it there. So let's hear those vocals I just recorded now with Kanye over them. I got a fantasy that's beautiful, that's dark and twisted. But I attacked the whole religion all because of my ignorance.

  81. 13:12

    What was I thinkin'? That was some bitch shit. I lost Adidas, but I'm still Yeezy. Back in [REDACTED:location], [REDACTED:gender], I'm a genius. [REDACTED:gender] in [REDACTED:location] hood just like I'm Eazy.

  82. 13:20

    Kanye Weezy, [REDACTED:location] of Chicago, life ain't easy. All praise be to Lord Jesus. Ghanda, please rest easy. All right, let me cut it there. Let me cut it there.

  83. 13:29

    Yo, that new- So you [REDACTED:gender] kinda get the, uh, the gist of what he was doing there. But I think, um, a lot of times when we hear AI music, AI music, uh, that's what it is.

  84. 13:38

    It's like someone taking, uh, the, the regular songwriting process where you go get production or do the production yourself, uh, you write the lyrics and then record your own vocals.

  85. 13:48

    And then what people are calling AI music is really just style transfer or voice conversion, where they're taking that voice and turning it, it into someone else's. So this is actually the first-- Shout out to, this [REDACTED:gender]'s name's, uh, Roberto Nixon, if I'm not mistaken.

  86. 14:00

    And I saw this video, and it kind of... I think it just definitely changed the trajectory of, of what I was doing at the time. I dropped everything and was like, "I have to figure out how the heck he just did this Kanye West, uh, voice thing."

  87. 14:10

    Uh, so I, I pretty much, uh, followed, uh, the steps in his video. I dove into some Discord servers and some subreddits, and that's how I really figured out how to, uh, do this for myself.

  88. 14:20

    Um, so, so I think, uh, that's, that's one compartment that I've kind of, like, created for AI music and put it, um, uh, to the side. The other one-- Or actually, I, I have a, a couple more examples of, of this as well.

  89. 14:32

    Let me... I think of, the, the next one is-- How many of you [REDACTED:gender] were familiar, by show of hands, how many of you [REDACTED:gender] heard that Drake AI song that went viral last year?

  90. 14:44

    Okay, we got a couple people. Nice. Um, so I'll, I'll play it. I don't wanna play the whole thing, but, uh, again, this is another example of, uh, voice conversion.

  91. 14:52

    And a lot of people were saying, "Oh my God, he pressed a button and generated a whole Drake song," and that's not really the truth. Uh, this kid is, his name is Ghostwriter, the, the person behind the song, and he did an interview where he talked about the process of, uh, creating the song, and it was all

  92. 15:04

    him except for the actual vocals. He, like, changed his voice to sound like Drake and, and The Weeknd. I'll play a little bit of it. [singing]

  93. 15:19

    I came in with my ex like Celine on the flex, ayy. Bumping Justin Bieber, the fever ain't left, ayy. She know what she need, yeah, I need her, she blessed, ayy.

  94. 15:30

    Giving you my best, ayy, yeah. You got my heart on my sleeve with a knife in my back. What's with that? Ayy. [REDACTED:age], I love him, that my [REDACTED:gender].

  95. 15:39

    That's my slatt, ayy. Metro made the beat, so you know that it's gon' slap, ayy, yeah. It's gon' slap, ayy. Tell 'em run it back. Talking to a diva, yeah, she on my nerve.

  96. 15:51

    That's actually my favorite part, but I'ma cut it off. [laughs]

  97. 15:54

    Um, so the Drake song was interesting to me 'cause I kind of saw it unravel. Like I, I think I found the video or heard the song when it was, like, under 1,000 views, uh, p- being posted in some of those sub- subreddits that, that I was talking about.

  98. 16:07

    Uh, but it went super viral. I think he got like 9 million streams in the first 24 hours. He was on his way, this [REDACTED:gender], uh, Ghostwriter, was on his way to charting, of course, before the RIAA got ahold of it [laughs] and was like, "We're shutting this down."

  99. 16:19

    So he got DMCA'd, and the song got wiped from everywhere. Uh, you can still find it on the internet, but it, it pretty much got wiped from everywhere. And, uh, I just thought, I said, "[REDACTED:gender], this is really gonna change everything."

  100. 16:29

    But again, everybody was calling it and saying it was AI music without, uh, being specific about what ki- kind of, uh, AI music it was. Um, so I have a couple more videos to show you [REDACTED:gender], but I just wanted to go over the Drake one 'cause I thought that was pretty cool.

  101. 16:42

    This next one is actually my favorite example of voice conversion that I've seen to date. Let, let me backtrack a little bit. These two songs that I just played, or videos that I just played, are from a rough- roughly a year ago.

  102. 16:54

    I think March of last year or April of last year. Since then, the voice conversion tools have gotten 10 times better, and, like, 'cause it-- When I listened to the Kanye West one back then, it was amaz- it literally blew my mind, no, no lie.

  103. 17:06

    And today I kind of listen to it and I'm like, "Wow, it sounds... It's really artifacty. It's, it's bad. It doesn't exactly sound like a natural human Kanye West."

  104. 17:13

    But today they have stuff that sounds almost exactly, you know, like Kanye West. I think-- And I say the kids because when I joined the Discord server, it was literally a bunch of, like, [REDACTED:age], [REDACTED:age], [REDACTED:age] doing this, [laughs] uh, these, these conversions.

  105. 17:25

    Um, but, uh, the kids have gotten a lot better at, like, t- uh, training, training, uh, these, um, models, these voice models.

  106. 17:35

    Uh, so the next one is actually my favorite, uh, example of voice conversion. Uh, how many of-- By show of hands, how many of you [REDACTED:gender] in here are familiar with Randy Travis?

  107. 17:44

    Yeah, I knew I was gonna get y'all back there. Yes, sir. Yes, sir. Okay. So Randy Travis, um, and forgive me 'cause I'm not exactly sure how this happened, but he lost his voice some time ago.

  108. 17:55

    And, uh, Randy Travis and his team kind of saw this AI thing unraveling or un- unfolding and decided to sit down, write a song

  109. 18:04

    Uh, produce a song, and then use this voice conversion technique for, uh, Randy to sing the song without him actually singing the song. So this is one of my favorite examples.

  110. 18:13

    I think also in the last year or so, the,

  111. 18:17

    the not-- maybe temperament is not the best word, but

  112. 18:23

    the general public opinion about AI music has shifted a little bit. It hasn't gotten too much better. It's really taboo, honestly. Like, I was a little bit, uh, nervous about doing this conference 'cause, like, people just do not like AI music at all, don't want it in their lives.

  113. 18:35

    But since then, this, this R-Randy Travis song came out a couple of months ago, or a couple weeks ago, I think it was.

  114. 18:42

    And when I looked at the YouTube comments of this video, the feedback was overwhelmingly positive. People were, like, in the comments saying, "I'm c-- I'm a [REDACTED:age] [REDACTED:gender], I'm bawling tears right now listening to this Randy Travis song."

  115. 18:55

    Um, like, they were like, "I-- We don't care how we get Randy, as long as we get him back. We'll, we'll take AI or not." You know what I mean?

  116. 19:00

    So i-it just was a, a really cool example of, I think, how voice conversion can be used in the, in the right way, you know, by, by the original artist.

  117. 19:07

    So I'll go ahead and play some of this as well. [singing]

  118. 19:54

    Hello, hello. Okay. Sorry, I apologize to the Randy Travis fans in the back. I'm, I'ma cut him off a little bit 'cause I, I do wanna keep it going.

  119. 20:00

    But, um, even though I'm not super into country music, I, I, I thought it was really touching that he was able to, like, get with his team and, and use, uh, AI to, uh, to create, um, new original music.

  120. 20:10

    I thought that was really cool. Uh, so that, that pretty much, uh, takes care of the style transfer stuff. I've showed you [REDACTED:gender] some examples. The next thing I wanna get into is, uh, text to music.

  121. 20:20

    Um, and I'll go over full songs as opposed to, uh, um, uh, the, the short samples that I was talking about. If you, uh, go-- A-again, just wanna remind you [REDACTED:gender], if you go to the link, aietalk.com/music, and send me the question on there, I can see it pop up on my screen.

  122. 20:37

    Um, and so, so I can, like, answer it in a, in a, in motion. I also wanna repeat it out loud for the transcription and all that good stuff.

  123. 20:43

    Um, okay, so text to music. Text to music. Text to music. I'm gonna show you [REDACTED:gender] a couple examples. The w-- I, I broke text t-to music down even a l- a little bit further, uh, by

  124. 20:57

    having... I have a couple folders here, and I think the way I did text to music was-- Yeah, m-meme songs and samples. Uh, okay, by show of hands, how many of you [REDACTED:gender] have heard about this Drake and, uh, Kendrick Lamar beef going on?

  125. 21:10

    Okay, dang. Okay, y'all are, y'all are up on game. I appreciate y'all. [laughs] Okay. Um, so

  126. 21:19

    for anyone who hasn't been paying attention, just know that one [REDACTED:gender] said very disparaging words about another [REDACTED:gender], and then the internet ran with it. So I don't have to catch you up on eighteen million songs that these [REDACTED:gender] have recorded now. [laughs]

  127. 21:33

    I couldn't even keep up with myself, to be honest with you. Um, but, but yeah, so, so w-- basically, the internet ran with this whole concept of BBL Drizzy, and Kendrick Lamar said a lyric in one of these diss songs that they had back and forth over the last couple weeks where he called Drake BBL Drizzy.

  128. 21:47

    If-- I, I don't wanna explain BBL to y'all. I ain't gonna lie to you. Um, so ho-hopefully you, you understand what a BBL is, and they call him BBL Drizzy.

  129. 21:53

    And it's funny because, like, uh, it can be taken a, a couple different ways, but I'll just play the music so I'm not rambling too much. [singing]

  130. 22:13

    My grandma would probably cuss me out for, uh, turning that off early, but I'ma cut it there and, um, basically say that, that this song was gener-- Th-- To me, this is

  131. 22:38

    w-really what gets to what AI music is. Uh, because, uh, there's a comedian named King Willdonius who logged into Udio, one of the tools I'm gonna tell you about, typed in a prompt, I think typed in some lyrics, if I'm not mistaken, and generated this song.

  132. 22:51

    He didn't get in a studio. He didn't record his vocals. There was no voice conversion. He just typed in some words and got this BBL Drizzy song. Um, and I, I think, uh, this is, like, a monumental moment in music as well because it was one of the first times I saw something like this go viral.

  133. 23:08

    And I think we'll get int- we-- I mean, we can, if you [REDACTED:gender] are interested, we can get into, like, some of the legality of AI-generated music and the copyrights around it.

  134. 23:16

    Um, but, but I think this, this moment where he, he created this BBL Drizzy song and it went viral, it's been used in a bunch of different songs now.

  135. 23:25

    It's, like, it's, like, monumental. But, but, but this is text to music. There was no prior recording. He just generated, uh, this song. One, one-shot, uh, prompt to a song.

  136. 23:36

    I think I might have one more good example for-- Y'all, I'll, I'll play this for a little bit.

  137. 23:45

    ElevenLabs is another, uh, AI startup that has now entered the text-to-music space, and this is a song that they generated. Uh, again, no prior recordings. This is just text.

  138. 23:58

    Uh, music came out the other end, and it went viral on Twitter. [singing]

  139. 24:14

    One, programming young. Dreaming that one day we'd make it work. Lines of code we'd write all night. Hoping that one day we'd get it right.

  140. 24:33

    Can we teach the-

  141. 24:36

    Oh, he was about to get into it. I, I cut off a little too early. But, but basically this is a sort of meta song because it's an AI model singing about GPUs. [laughs]

  142. 24:43

    Um, and I thought that that was pretty cool. Uh, but, but yeah, another text-to-music song that I would kind of put into the, uh, meme, the meme song category.

  143. 24:53

    Uh, with that being said, I think I had... Oh, yeah, so, uh, audio to music. These are actually the most interesting samples-- examples, and also, uh, the shortest if I'm not mistaken.

  144. 25:04

    Um, so I'll get into some of these. These are really cool. I, I think the, the people who are like, uh, musicians or artists who have, uh, a lot of talent already, uh, like m- musical talent already are, are into, um...

  145. 25:22

    Oh, let me play this. Into this style of, uh,

  146. 25:26

    music generation because it takes an existing sound and transforms it into-

  147. 25:28

    You might notice that it sounds like Ky-

  148. 25:30

    A new Yeezy track goes hard-

  149. 25:31

    Into another song. Uh, an, an existing sound and, and turns it into music, sorry. [beatboxing]

  150. 25:58

    So th- uh, this is, uh, using Stable Audio. It's another tool. It's not Suno or Udio, but it is called Stable Audio. And I think this was the first tool I saw that had this feature publicly available where you could put audio in and get a, a song out or get music out.

  151. 26:12

    Maybe not a, a full song. Um, but I, I, I really like this because it-- I think it unlocks a lot of different ways to create music. Um, so I th- I think that'[REDACTED:url]- I might have one more example to show you [REDACTED:gender].

  152. 26:26

    This one was cool. This is also from the, the Suno team as opposed to Stable Audio. [drum beating] This one's pretty good. [rock music]

  153. 26:51

    Nice. Okay. Um, so I thought that was pretty cool because literally all he recorded was his fingers drumming like that on the desk, and then out came this full song with the guitar and all these other instruments.

  154. 27:02

    Um, so I think that'[REDACTED:url]- You [REDACTED:gender] kind of get the, the gist of like what, uh, audio, uh, audio to audio, um, model generation or music generation can be like.

  155. 27:11

    Um, and I wanna say with that, we can get directly into actually generating some songs ourselves. Okay, we could talk about how it works, but, uh, it-- You [REDACTED:gender] didn't really seem so much interested in like the technical side [chuckles] when I was looking at the poll, so I could briefly touch on it.

  156. 27:29

    But I think, um, h- how-- By show of hands, how many of you [REDACTED:gender] saw that Suno and Udio have been, uh, are getting sued this week? You [REDACTED:gender] saw that?

  157. 27:39

    Already?

  158. 27:39

    Yeah. The RIAA just, uh, filed a complaint against both services. So, um... And, and what they're alleging is that these, uh, AI models are trained on massive music datasets that are copyrighted, and we are-- Well, not we.

  159. 27:53

    Uh, don't, don't put me in no indictments. I do not wanna be [laughs]

  160. 27:57

    But-- 'Cause I, I don't-- That, that's another thing I, I didn't mention at the beginning, but I have no affiliation to Suno, Udio, Stable Audio, Boomy, no one. Like I, I, I just like this stuff and learned it and, uh, figured it out and wanted to do a presentation on it.

  161. 28:09

    Um, but, but what they are alleging is that Suno and Udio are infringing on their copyrights by training on this, this music.

  162. 28:20

    That, that last song that you played before, that just really sounded like Tom Petty.

  163. 28:23

    Yeah. Yeah. [laughs] Yeah, we had a, a comment from the audience saying that the last, uh, generation sounded like Tom Petty. Um, if you are more interested in how these models actually work, the architecture, transformers versus diffusion, uh, I have a link on the page that I mentioned earlier.

  164. 28:39

    If you click on the Notes button and then scroll all the way to the bottom, I have compiled a list of all of the, uh, music model papers. So if you [REDACTED:gender] are interested, uh, there's some-- you, you can find it there.

  165. 28:52

    Um, okay. So what I wanna do now as we get into, uh, Udio and Suno, there is a lot less people than I thought there was gonna be, so this'll work well, hopefully.

  166. 29:03

    But there's a business card on the table in front of you, and that business card has a secret code on it. This is a little PVP, uh, survival of the fittest.

  167. 29:15

    Like may, may, may the best [REDACTED:gender] win. But if you take that secret code, go to the link that I showed earlier, let me put it on the screen again,

  168. 29:26

    and click on the Gateway button, there should be a green Gateway button that looks like this. If you put in that code, what will pop up is a username and a password.

  169. 29:39

    If you use that username and password to sign into any of these services...

  170. 29:44

    Let me try and make it a little bigger than that.

  171. 29:49

    If you wanna get into Udio, you can use Google, Discord, Twitter, Apple. I've signed up for all the services. If you wanna get into Suno, you can do Discord, Google, or Hotmail, Microsoft, whatever they're calling it nowadays.

  172. 30:03

    And then you can sign into-- So what I did for this presentation so you [REDACTED:gender] could generate music, uh, freely, because I, I, I know i- in the, the show notes, uh, for this workshop, I put that you [REDACTED:gender] should have your own account, but I also thought it would be cool if we could do like unlimited

  173. 30:15

    generations as opposed to what they give you on the free, uh, account. So what I did was I signed up and paid for a Premier account for Udio and Suno and gave you [REDACTED:gender] the login just now.

  174. 30:25

    So if you go log in, you can generate music as we're talking about, uh, the next couple steps.

  175. 30:33

    Just know that, uh, I, like, I think 2FA is on Apple, if I'm not mistaken, so there might be some difficulty getting in, but just try and get into one of these services, then log into Udio or Suno.

  176. 30:45

    And, and the w- the, uh, URL for that, I think I have it here, is U-D-I-O.com once you've signed into one of the other, uh, authorization services, or [REDACTED:url]U-N-O.com will get you into, uh, one of these, one of these sites.

  177. 30:59

    It worked.

  178. 31:00

    Worked? Nice. Awesome. We got somebody in. Oh, another person. Oh, awesome, awesome, awesome. Hey, PVP, some of you [REDACTED:gender] might not make it. I'm not gonna lie to you. [laughs]

  179. 31:10

    But yeah, I just wanted to do that so you [REDACTED:gender] had, um... And actually what I'm gonna do as well, like when you [REDACTED:gender] are in your seats generating songs, there's another link on this, uh, uh, aietalk.com/music page where it says Submissions.

  180. 31:22

    What I wanna do, uh, we might be tight on time, but what I wanna do is have everyone who generates songs here in person to submit to this link.

  181. 31:32

    Uh, it's basically just a Google Form. So you generate some music,

  182. 31:36

    take the link, sh- the share link to that music on either service, uh, click on this link here, Submissions, and submit your song. And I think what I'm gonna do at the end is just take like a random number generator, and, uh, two people will be selected to get access, like, keep access to these two accounts that

  183. 31:52

    I set up. So one of you [REDACTED:gender] will walk away with like a Suno account, one of you [REDACTED:gender] will walk away with a Udio account.

  184. 31:59

    And like I said, I, I paid for the Premier. It's, I didn't pay for the whole year. After they got sued, I was like, "Yeah, I might not wanna invest." [laughs] [laughs]

  185. 32:07

    So I just went ahead and did 30 days.

  186. 32:12

    Okay, so we can get into some of the techniques. W- while, while we're on this topic, uh, and you [REDACTED:gender] are hopefully generating in your seats as well, does anyone here have a birthday this week?

  187. 32:23

    Nice, my [REDACTED:gender]. Randy Travis. Um, what's your name?

  188. 32:30

    Patrick.

  189. 32:31

    Patrick. Okay, I think what we're gonna do is generate a, a birthday song for Patrick,

  190. 32:35

    and I'm gonna do that from my account. [laughs]

  191. 32:40

    Wait, what's... Uh, one more time.

  192. 32:42

    I'll be immortalized.

  193. 32:43

    Yes, absolutely. [laughs] Patrick, my [REDACTED:gender]. Uh, Randy Travis. Oh, they're gonna block that. Fan and, uh, music.

  194. 32:59

    I'm not even gonna try. [laughs] There you go.

  195. 33:07

    Yeah, they just... I don't know, U- Udio may not be as bad about this, but, uh, if you put, uh, artist names, and obviously now because they're being sued, but if you put artist names, they like, they, they have a content filter, you know what I mean?

  196. 33:19

    They don't-

  197. 33:23

    Yeah, but like

  198. 33:26

    Absolutely. I'm, I'ma talk about that as well, like ways to kinda get around the filters. Um, but yeah, that definitely works. Uh, we had a comment from the audience that basically said if you, like, change the spelling up a little bit and then put like the style of music that the artist is related to, you can, uh,

  199. 33:44

    you can skate.

  200. 33:46

    He's an attorney.

  201. 33:47

    An attorney?

  202. 33:49

    He's a patent attorney, so put that in there.

  203. 33:50

    Oh, shit. Oh, you're the one suing, uh, Suno and Udio. [laughs] [laughs]

  204. 34:02

    Uh, birthday song for Patrick. Let's put in some...

  205. 34:08

    Well, yeah, Patrick, what do you like? What kind of music you like?

  206. 34:14

    It's a weird question. Uh, it would be the band.

  207. 34:18

    Which is what genre?

  208. 34:21

    Rockabilly.

  209. 34:22

    Rock? One more time.

  210. 34:25

    Rockabilly.

  211. 34:26

    Rockabilly. I don't... I'm, I'm gonna be honest with you, I'm lost.

  212. 34:31

    Did you say the band? The band. The band. Yeah. Yeah. That's kind of like... I guess you would call it rockabilly.

  213. 34:37

    Let's go ahead and generate a song for my [REDACTED:gender] Patrick.

  214. 34:41

    Cool.

  215. 34:41

    And then I'm gonna take the same prompt.

  216. 34:45

    I'm kinda getting into some of the, um, prompting techniques here, but

  217. 34:50

    you [REDACTED:gender] are able to see this screen, I hope, and see what I'm doing.

  218. 34:56

    And then we'll play a song. Uh, couldn't generate that. Song description, name Randy Travis.

  219. 35:03

    What about Tandy Ravis? Ravis.

  220. 35:19

    Who is Ravis? Listen, [REDACTED:gender], just give me my song.

  221. 35:29

    All right. All right, so I'm not monologuing the whole time. I'm gonna go ahead and wrap it up. But I, I... The way I think of these two different services is that, uh, Udio is more like a songwriting partner.

  222. 35:43

    Um, they have an experimental model that you [REDACTED:gender] now ha- have access to if you log into that Premier account that I paid for. Uh, but, but the default Udio experience is that you generate 30 sec- 30 seconds of music at a time.

  223. 35:54

    Um, and you can either extend that idea, remix that idea, or there's a third option I didn't write here, and you can scrap it, 'cause a lot of times, you know, you get weird stuff, uh, when you, when you're not great at prompting.

  224. 36:05

    Um, but you can also use inpainting, uh, to, to modify [REDACTED:url] uh, specific sections of a song, like if you wanna change or tweak one or, or two little things.

  225. 36:16

    And then now they have this feature o- on, uh, within Udio where you can do the audio to audio, and it, it allows you to upload music. Um, so I, I- If you are an artist and you write songs or have written a song before, or spoken word or anything,

  226. 36:32

    I would urge you to try to take one of your creations, especially if you've recorded it, it gets crazy. If you've recorded some music before, you should take the tempo, uh, the key that that song is in, and your lyrics and put it into Udio and see what comes out the other end.

  227. 36:49

    It's blown my mind a couple of times. I've taken a couple of my own songs and putting them in, and it, and it gets crazy what, what comes out the other side.

  228. 36:55

    Like, Udio does me better than I do me sometimes,

  229. 36:59

    honestly. So, so it, it's a really cool experience if you have, like, e- pre-existing music, you can, uh, really play with Udio a lot. And then the way that I think about Suno, um, is your in-house music producer.

  230. 37:12

    Uh, I like to describe it in a way that I, I feel like

  231. 37:17

    somewhere in the system prompt of Suno's model, they have, like, Top 40 music in there, 'cause it always comes out really polished and clean, and maybe it's the way that the, the data that they put into the mu- into the model in the first place.

  232. 37:27

    But I, but, uh, my experience with Suno has that been, been that, uh, it just comes out really, really clean, um, Top 40-ish. One thing it's gotten really, really good at lately, this is like a, a, a side note, is that when I first ...

  233. 37:41

    I, I personally have onboarded, like, I think a dozen or so people over the last couple weeks because I wanted to, um, understand what is the actual use case, uh, for these services.

  234. 37:51

    So I just started onboarding friends and family, some of my music friends, some people that are not into music at all. And what I found when, uh, a couple weeks ago, was that, like, when I would generate songs, uh, if you're not super specific and counting the syllables in each line of the song, when you're getting into

  235. 38:07

    the, the custom mode of Suno, it'll ... The timing of the beat and the artist's, quote unquote, the vocal's, uh, lyrics, gets super off. Really, really, really off. Um, but they've gotten really good about this lately.

  236. 38:20

    Like, uh, the last couple songs I've generated from Suno aren't off. Um, they've-- The song structure's a lot better, 'cause that's another tip we'll get into. The, the song structure really, um,

  237. 38:31

    has improved a lot, um, um, in Suno over the last couple weeks. But, but it's more or less the same, just the outputs are a little bit different from, uh, Udio.

  238. 38:40

    And now we get into the tutorial. Um, so have you [REDACTED:gender], anyone submitted any songs yet?

  239. 38:48

    Let me see. Let me recheck. I was just, uh, generating some music for my [REDACTED:gender] Patrick. We're gonna play that in a second. But that was, that was the basic tutorial.

  240. 38:57

    You kind of ... It's about as simple as it gets. Like, it really, really lowers the barrier to entry for creating music, and I think we, we can get into some, some other stuff a- about, like, more conversations about that.

  241. 39:07

    But I think it ... For now I'll just say that it, it lowers the, the barrier and makes it easier to, uh, create stuff for [REDACTED:url] you know, specific situations.

  242. 39:16

    Okay, nice, we have a couple submissions. We do have a couple submissions. But the gist of it is you go to Udio.com, you log in,

  243. 39:24

    you go to this, uh ... Oh, they used to have ... Oh, I guess here at the top it al- it always says Create. Then you can type in,

  244. 39:33

    oh, what's another one? The AI Engineer World's Fair.

  245. 39:41

    Let me start over. A song about the AI Engineer World's Fair.

  246. 39:49

    And I know, uh, the organizer Benjamin Duffey, Dunphy really likes old school hip hop.

  247. 39:58

    Boom bap. So I'll generate a song. So you come in, you type in your prompt, you hit Create. Uh, there's a couple of different options down here. Again, if you have a free account you won't th- see this experimental longer song.

  248. 40:10

    Um, uh, you'll, you'll be defaulted to the Udio 32 model. But you can come in and write your own lyrics on top of this prompt that you just put in.

  249. 40:20

    Or, or you can select to only have an instrumental generated as opposed to, uh, a full song with lyrics and everything. Um, so a- at, at its core, that's how it works, for pretty much all these services as well.

  250. 40:30

    There's not too much beyond that that you need to know to start generating music. I would argue that there's a lot more that you need to know if you wanna get good at it, but that's, that's pretty much, uh,

  251. 40:38

    how we, um, generate music on both services.

  252. 40:42

    And what I like to do, just to get a nice even comparison, is use the same exact prompt. I'll refresh the page, 'cause, like, for some reason they're really buggy, uh, sometimes, where if you generate song after song without refreshing the page, you get, like, extensions of songs as opposed to a brand-new song.

  253. 40:59

    But I will go ahead and create based on this prompt.

  254. 41:06

    And then I'ma come back because we've now gone over the basic tutorial. I'ma come back and look at some questions, 'cause I know some questions came in. Nice.

  255. 41:15

    And in the meantime, while I'm getting my thoughts together, I'm going to, uh, play my song, or the song that, that was created for Patrick, wishing him a happy birth- Oh, moderation error.

  256. 41:26

    They didn't like your, um, or my, my Randy Travis.

  257. 41:32

    Let's go ahead and play this one for a little bit. [upbeat music]

  258. 41:40

    It's your day, Patrick. We're singing loud for you. Oh yes, oh yes, it's true. It's your time, Patrick. Celebrate the whole day through. Oh yeah, oh yeah, you know.

  259. 41:55

    Ooh. Yeah. Yeah.

  260. 42:13

    Happy birthday. Ooh yeah. Oh, we know, Patrick, you love that rock and roll. Oh yeah, yeah. Oh yeah. Hey

  261. 42:32

    Patrick, what'd you think?

  262. 42:33

    It's great. Thank you. [laughs] Theme song to my sitcom. [laughs]

  263. 42:40

    Patrick said it's the, the theme song for his sitcom. I'm gonna play these other two quick. I just wanna hear what they sound like. [singing]

  264. 42:52

    Uh, real quick, if you click on the actual song name, you can see the lyrics on the Suno side.

  265. 43:01

    Birthday candles flame so bright. For a connoisseur of music's light.

  266. 43:12

    Rocking through the day and night, night. Patrick's world is pure delight.

  267. 43:25

    I'm gonna try the other version, see if, uh... Was that more your speed, Patrick?

  268. 43:29

    Yeah, they're both great.

  269. 43:29

    Oh, awesome. Okay. All right. Oh, we got five minutes left on the presentation.

  270. 43:42

    Birthday candles fla-

  271. 43:44

    Okay. So either they generated the same song twice or...

  272. 43:51

    Yeah, maybe. Yeah. Let's see what the AI Engineer World's Fair song. Uh, and actually, l- let's address these questions 'cause we only have a couple minutes left.

  273. 44:02

    And then I'll get into the notes real quick 'cause, uh, I did wanna show you [REDACTED:gender] some of the prompt stuff. The truth is that, like, I have everything in, um, under the Notes section, like pretty much everything I talked about today.

  274. 44:12

    Also, when the recording becomes available, you can come back to the same page, get, uh, access to the recording, and then I'll hopefully have the transcript up in, like, an hour, um, because I recorded it myself.

  275. 44:23

    But if you come here to this page, aiietalk.com/music, and click on the Notes section, um, and you wanna not listen to me talk about the prompts, you can find the prompt stuff yourself, uh, not quite all the way at the bottom, but here, uh, close to the bottom.

  276. 44:39

    And, uh, this is a GPT prompt, meaning you literally can go to ChatGPT, use this prompt. It, it got a little bit weird. Like, for some reason when I use these closing brackets in GitHub, um, it, like, basically disappears the whole...

  277. 44:52

    And, like, you won't even see lyrics anymore. But basically, what you can do is... I, I think even this would work. Like, GPT-4 is now smart enough to do this.

  278. 44:58

    You can take and you can put in this, uh, template, and then actually fill in the top here where it says Song Description, and GPT will actually produce some coherent lyrics.

  279. 45:06

    And it allows for better music generation when you, uh, come in with the lyrics, as opposed to giving it a general topic and hoping it comes up with good lyrics itself.

  280. 45:14

    You can kind of tweak the lyrics yourself before you start to generate. So that's the GPT prompt. And then for the music model prompts, uh, these are pretty much just examples of what we, it, uh, just did with, uh, Patrick and, um, the AI Engineer World's Fair song.

  281. 45:27

    It's just, like, a general description of a song, um, that you wanna hear, that you wanna generate, and then it pops out the other end.

  282. 45:35

    Uh, super pressed for time. Um, there are some official Udio and Suno tips, um, that are very useful as well. And then here, I wa- I kinda wanna get into the Prompt Master thing, actually.

  283. 45:53

    Philip asked, "Is there a way to transfer-- to do voice transfer of a more generic voice rather than a specific v- voice?" Referring to, uh, the voice conversion stuff that we talked about earlier.

  284. 46:04

    And I think the end of that question got cut off

  285. 46:08

    'cause you said, "Like..."

  286. 46:10

    Yeah, like, like for copyright reasons. I don't wanna be Eminem. I wanna do, like, generic [REDACTED:origin] rapper voice.

  287. 46:16

    Are you talking about doing voice conversion training, like training models yourself? Like going to one of the services or downloading the tools? You literally can just take and blend multiple voices.

  288. 46:25

    So the way that it works is, like, you need, depending on what service or model you're us- or what, um, um, training tool set you're using, you need, like, minimum six seconds of someone's voice to start training.

  289. 46:37

    Um, and, and you get, like, a kind of coherent model after that. But what you can essentially do is just six seconds of your voice, six seconds of his voice, six seconds of Emi- Eminem, and it kind of, like, blends it together.

  290. 46:46

    Or you can just go to a TTS service and get, like, an actual, like, robot voice, like someone-- a, a voice that doesn't belong to a real person, and then use that to train on.

  291. 46:54

    So yeah, for copyright purposes, yeah, you can, you can do that.

  292. 46:59

    Dang, they gotta wrap it up. Okay, um,

  293. 47:02

    what recommendations do you have to help advocate for the responsible and ethical use of...

  294. 47:09

    Yeah, sorry about that. I don't know why the questions are getting cut off.

  295. 47:15

    Oh, ethical use of AI music. Um, [sighs] that's a good one. I'ma be honest, I'ma plead the Fifth 'cause I'm not a lawyer.

  296. 47:27

    I'ma, I'ma be completely honest. [laughs] I'ma be completely honest.

  297. 47:30

    Maybe the other thing to do is law.

  298. 47:32

    Oh, yeah, yeah. Good point. Um, what tool generates the best cloned voice that can match emotional tone, like Kendrick Lamar's Euphoria three switches? What is the best cloned voice?

  299. 47:43

    I would definitely look into RVC. If you can-- You can use RVC in the cloud. You don't have to... Like, when you Google RVC, there comes up a bunch of tutorials on how to run it yourself locally.

  300. 47:51

    You do not have to do that. You can run it in the cloud 'cause y- you need, um, some decent hardware to do, like computer hardware to do it.

  301. 47:58

    But, but definitely look into RVC. It's the best, uh, right now. Can we have access to the slides? Yeah. And my time is up shortly. Oh, oh, my-- Did my thing go to sleep?

  302. 48:08

    It did go to sleep.

  303. 48:10

    I think I heard a question about how to log into your accounts.

  304. 48:15

    Oh, to log into Udio, Udio or Suno?

  305. 48:17

    No, to get the Google Docs.

  306. 48:20

    Really? Uh-

  307. 48:23

    You said request access through Google.

  308. 48:25

    Oh, okay, okay. Sorry. That's my blunder. Um, that was pretty much it. Oh, r- Rate Your, Rate Your Music is something you definitely should check out. Like, what someone basically reverse engineered is that- These models were trained on labels from Rate Your Music, so you can almost reverse engineer exactly what sound you want by going to Rate

  309. 48:46

    Your Music, typing in an artist name, typing in a genre, um, or r- really just typing in what you wanna hear, and then using whatever labels they put on that music.

  310. 48:56

    You put it back into the model itself, Suno, Suno or Udio, Stable Audio sometimes, I think, um, and you can get, like, very, very close to a specific, uh, sound.

  311. 49:05

    I don't wanna say specific artist, but, like, yeah, you, you can get pretty close. Um,

  312. 49:10

    yes. And I, I was gonna get into, like,

  313. 49:14

    the law thing. I don't-- I-- You, you want me to differentiate between ethics and law, and I understand. But I was gonna get into it a little bit, but we, we kinda ran out of time, so I apologize for that.

  314. 49:21

    If you [REDACTED:gender] would like, I'd love to talk about this stuff all day. If you, if you wanna approach me after the workshop, and we chop it up all day about this stuff, I, I'd love it.

  315. 49:29

    Um, yeah. And, and then what I'll do for giving away the accounts, maybe I'll do, like, a live stream or something like that and do, like, a-- grab everybody's email who submitted, 'cause I think we had-- Nice, we had 12 submissions.

  316. 49:42

    Let me refresh. I think we had 12 submissions, and I'll pick somebody and just, like, send you [REDACTED:gender]-- I'll change the username and password so everybody can't log into your stuff. [laughs]

  317. 49:50

    And then I'll give you, uh, I'll give you the accounts. But yeah, I, I appreciate you [REDACTED:gender] for coming. Uh, hope you [REDACTED:gender] got something out of this and walked away with some music, generated some stuff.

  318. 49:58

    And I think I'll play a song to go out. [singing]

  319. 50:04

    Walking through the fair. Copper wires in the air. Coders in the screens. Dreams turning into scenes. Blueprints in my hand. Robots take a stand.

  320. 50:19

    Lines of code they shift. Tech minds give the gift. AI dreams they see.

  321. 50:29

    This fair-

  322. 50:31

    Dang, that's dope. I wonder what Udio generated.

  323. 50:37

    Oh, oh, first-- 'Cause I just saw the question that someone said, uh, stems, how to, like, pull the stems out and, like, edit them. There's a tool, again, not, not, um, like I have no affil- affiliation with them.

  324. 50:47

    But if you just go to-- Actually, the, the easy way to do this is let's say we just take this, uh, happy birthday song by Patrick. We go to suno.com.

  325. 50:56

    All you have to do is, uh, [REDACTED:url].com or [REDACTED:url] [REDACTED:url] [REDACTED:url]. Like, literally go to the link where you just generated the song. Type in D-A-W, digital audio workstation, at the end of the, uh, URL.

  326. 51:09

    Not the end of the URL, but the end of the, the domain name, and hit Enter. And I think it takes anywhere from 30 seconds to... Let's go ahead and sign in real quick.

  327. 51:22

    To 90 seconds. I've had to wait longer for some of the songs 'cause it's more complicated to do. Uh, [REDACTED:url] sorry, [REDACTED:gender], I, I forgot to mention this. But yeah, like, if you go and do that, type in D-A-W at the end of the, uh, Suno song, it'll pull that [REDACTED:url] that exact song that you just generated

  328. 51:37

    into this digital audio workstation and give you stem by stem, and then let you control the volume of each, each specific stem. And by stem, I just mean each, uh, specific instrument.

  329. 51:45

    They do all the different instruments and then the vocal itself. So this is a really, really cool tool that I think is, like, [REDACTED:url] a startup that, uh, hasn't been around for that long.

  330. 51:54

    Um, but it's, it's a really, really cool tool for, like,

  331. 51:59

    uh, being able to, to, to get more control out of your generations.

  332. 52:03

    So I just wanted to- D-A-W? D- D-A-W, so digital audio workstation.

  333. 52:10

    And you can, like- [singing]

  334. 52:11

    Birthday candles flame so bright.

  335. 52:15

    Like, let's say I don't like the vocals, but I really like the music. [singing]

  336. 52:18

    Connoisseur of mu- Rock and- Patrick's world is pure delight.

  337. 52:40

    In the code and in the show.

  338. 52:45

    And this is just vocals. [singing]

  339. 52:46

    Has a patent when he's on go.

  340. 52:52

    It's the bass. [singing]

  341. 52:54

    Rock tunes flowing in his veins. Patrick dances through the rain.

  342. 53:07

    Happy birthday, Patrick dear. Let the music keep you near.

  343. 53:22

    Rock and roll all through.

  344. 53:26

    So yeah, that, that's, uh, pretty much how you use Wave. It's called Wave Tool, that, uh, that service. So that's how you can get the stems. Th- there's also another tool called UVR5, but it gets really involved.

  345. 53:36

    You do have to have, like, dec- decent hardware. But you can also use, like, uh, instead of using, like, an online service where-- I think all the accounts on Wave Tool right now are free.

  346. 53:45

    But you can use UVR5 locally forever. It's like an open source software project that, uh, allows you to pull stems just like this, literally. You can get-- A- a- and that, that's not just, uh, AI-generated music.

  347. 53:56

    Any song you can put into UVR5 and get, like, the stems from it. It's, it's actually pretty cool. I know- [audience applauding] Yeah, appreciate it.

  348. 54:04

    I know you-- They're-- They want you [REDACTED:gender] to go over to the, uh, what was it called again? The general session. Yeah, the general session. So I'm, I'ma, I'ma shut the hell up now.

  349. 54:13

    I appreciate you [REDACTED:gender]. Thank you. Thank you for coming. [outro music]