← All AI Engineer talks

AI Engineer Summit 2025

WTF do people use Open Models for??

Read the talk

What people use open models for

Open-model usage reveals two different markets: people choosing a conversational personality, and businesses protecting reliable workflows from unwanted change.

From a talk by Eugene Cheah

Before you start: Basic familiarity with language models, API requests and production software is helpful; no knowledge of model training is required.

What are all those models doing?

What do people actually do with the flood of open models? Eugene Cheah opens with Hugging Face’s accelerating catalog: he reports more than 50,000 model uploads per month over the preceding year. DeepSeek-R1 supplies the striking example—a model he presents as challenging GPT-4o without a billion-dollar budget. That comparison depends on the task; the original R1 model card shows stronger results on some reasoning benchmarks, not superiority across every evaluation.

Stacked area chart showing accelerating model growth, with colored categories including natural language processing and Other.
Hugging Face model growth in “Zero to One (Million Models).”

Cheah reports more than four million downloads of a 685 GB model in the preceding month. If each represented a complete transfer, that would imply 2.74 exabytes of data, or roughly 690 years over his 1 Gbps connection. The conditional matters: Hugging Face’s download documentation describes counts of designated file requests, including HEAD requests, rather than complete weight transfers. His separate prediction of another 500 uploads during the talk also does not follow from the monthly rate alone. The opening establishes enormous interest, but download counters cannot tell us either the bytes transferred or the work performed.

Cheah’s observation point is Featherless AI, where he is CEO, alongside his role as co-lead of RWKV. At the time of the talk, he describes a service offering unlimited API requests to more than 3,700 open models for thousands of users. The individual plan costs $25 per month, includes R1, and sits alongside larger commercial plans. The ambition is to keep expanding toward the Hugging Face catalog; the resulting mix of customers gives Featherless a view into both personal experimentation and production workloads. These are the service conditions behind the February 2025 observations, not a current pricing guide.

0:010:13
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:01 · section reference included

Equal prices reveal preferences—and hide fine-tunes

For individual users during one week in February 2025, R1 leads Featherless’s usage chart, followed by the Llama 3 family, Mistral NeMo 12B and Qwen. Individuals do not pay per token or face different prices for different models. Instead, the plan limits them to one large-model request at a time, while smaller models generally respond faster. That removes a major price incentive from the choice, although latency still matters. Cheah interprets the remaining preferences as a mixture of personality, reputation and feel, rather than simply MMLU scores.

The chart changes meaning when its categories change. Grouping by model family makes Llama and Qwen look like coherent products. Grouping by model name splits those wedges into many fine-tunes, each with different behavior and intended uses. A family’s popularity can conceal demand for very specific personalities. That distinction becomes especially useful when interpreting creative and conversational workloads.

1:452:06
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

1:45 · section reference included

Production changes the chart

Businesses bring a different preference: consistency. Developers want to decide when their model changes, rather than discover that a provider has changed it for them. When Cheah adds the scaled commercial customers excluded from the individual chart, the smaller Mistral NeMo becomes dominant. It is already an older model in a fast-moving market—roughly seven months past its July 2024 release at the February event. Cost helps explain its persistence: Cheah says Featherless provides access to four small models for the price of one large model.

Pie chart beside the presenter, with mistral-nemo-12b-lc at 51.84% and llama33-70b-16k at 20.15%.
Mistral NeMo occupies 51.84% of the displayed pie chart.

Once deployed, models can remain in place for quarters or years. Cheah credits NeMo’s early appeal to its Apache 2.0 license and its ability to replace GPT-3.5-class workloads, contrasting that license with restrictions in the Llama license that made some enterprise lawyers uncomfortable. He also credits cloud-platform defaults, including AWS and GCP, and hundreds of fine-tuning tutorials with reinforcing adoption. Those distribution paths matter because an example copied into a tutorial can become a dependency in a production system long after the model stops being fashionable.

Consider a narrow application whose prompts and monitoring already produce dependable results. Cheah describes teams that have tuned such systems to above 99% task accuracy; he does not provide a shared evaluation set for that figure. A model change can invalidate the prompts, disturb measured behavior and create another round of repairs. Keeping a fixed model version preserves the baseline against which improvements are evaluated.

The same persistence appears in safeguards. Cheah reports that Llama 2 accounts for about 2% of Featherless’s workload. He highlights Llama Guard, based on Llama 2, as an example: older safeguard tutorials continue to send both existing production systems and newly arriving teams toward it, despite newer alternatives. The age of the model is less decisive than whether the surrounding workflow still works.

3:043:17
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

3:04 · section reference included

From model identity to creative and social uses

Featherless’s strict no-text-logging policy limits what it can directly observe. Instead, the team infers use from model identity and requesting-application metadata—a Llama Guard request is a strong clue about safeguards—and asks customers what they are doing. It also exchanges observations with aggregators such as OpenRouter. Cheah’s approximate ordering is creativity and friendship, coding, ComfyUI-style workflows, RAG and chat interfaces, then other agents and work. Customers include individuals, internal business systems and products that repackage the API in mobile apps, web apps or agents. The percentages that follow mix traffic and request descriptions without a common published denominator, so they describe this observed ecosystem rather than a precise market census.

Cheah estimates that creative writing, roleplay, companionship and the smaller therapy/journaling category together account for 30–40% of traffic. NovelCrafter illustrates the writing workflow: outline a novel, manage its lore, draft scenes, maintain high-level narration and collaborate with both humans and AI. Authors and fan-fiction writers use the model within a continuing creative project, rather than as a one-question answer engine. Adjacent uses include generating Dungeons & Dragons-style game content.

For roleplay and companionship, Cheah names Wyvern Chat and SillyTavern, along with mobile products that call Featherless underneath. He describes this as the largest non-code, non-agentic category by active users. Its association with explicit content obscures other reasons people use it; he connects the demand to recurring references to the film Her in closed-model product discussions too.

Cheah reports that more than 60% of users in the companionship segment are women, without supplying a sampling method. He compares the mismatch between public image and actual demand to romance publishing: a large audience can remain underrepresented in how an industry presents itself. An unnamed app maker describes women having long daily conversations to decompress, particularly when real-life partners are absent or emotionally unavailable. Cheah turns that observation into an interpersonal point about being available to listen, which leads into the neighboring category of therapy and journaling.

That neighboring category is fragmented across many dedicated apps, but the interface need not be specialized. Some surveyed users simply choose a ChatGPT-like client or a therapy character inside a companion app. These are descriptions of how people use the products, not evidence that the products provide effective clinical care. The common requirement is a conversation that helps the user reflect.

6:156:18
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:15 · section reference included

The technical requirement behind “vibes”

For these users, MMLU is not the main selection criterion. Community fine-tunes compete on conversational behavior, and favorites can change weekly. The apparently vague word vibes includes a concrete requirement: do not rush to solve the conversation. In the therapy and journaling example, the model should help a person work toward their own answer instead of immediately prescribing what to do. The desired behavior can run against the answer-first habits of instruction-tuned assistants.

Story writing, roleplay and companionship have a related requirement: slow burn. Cheah compares rushing a story to starting Game of Thrones at the first episode and then jumping straight to the finale. The entertainment lies in what happens between those endpoints. A model that resolves every tension immediately can be efficient at answering and poor at sustaining the experience.

Fine-tunes make those preferences concrete:

  • Voice: Cheah names alpindale Magnum for a cowboy-style personality.
  • Genre: He names Rossini for science fiction, which he says many other models handle poorly.

He expects these particular favorites to change quickly. Yet users also return to old models, treating their distinctive behavior like the remembered charms of a retro game. In this category, replacement does not necessarily make the predecessor worthless.

Cheah puts this broader community, including closed products such as Character.ai, in the tens of millions of users. He contrasts that scale with other categories he describes as not yet reaching millions. These are his broad audience estimates, but they explain why a large catalog of differently behaving models can matter more than a single leaderboard winner.

10:3910:52
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:39 · section reference included

Coding shifts from completion to intervention

Coding has a smaller human audience but a much larger appetite for tokens per person. Cheah places coding at 20–30% of traffic. He first separates familiar IDE features—autocomplete and chat-based editing—from more agentic workflows. Replit, Mistral and other labs supply small coding models in the 3–12 billion parameter range, and he considers autocomplete essentially solved for this purpose. The competitive frontier has moved toward agents that carry out larger changes.

The useful agent is not necessarily the most autonomous one. Cheah describes developers finding SWE-agent- or Devin-style autonomy excessive when it leads to loops. His preferred pattern asks clarifying questions, accepts corrections and lets a human intervene quickly. Vibe coding takes that interaction far enough that the person may never edit the code directly, instead steering the work through chat.

Cheah estimates that a developer deeply engaged in agentic coding can generate 1,000 times the input and output token traffic of one companion-chat user. Repeated context, tool results and revisions make a comparatively small population expensive to serve. In his account, at most tens of thousands of simultaneous coders can therefore produce a rapidly growing share of overall traffic.

Open models are entering a category still dominated by commercial ones. Cross-referencing OpenRouter, Cheah reports Claude Sonnet’s dominance in agentic coding at a ratio greater than 10:1, without defining the comparison population precisely. Claude Sonnet also benefits from the commercial-model integrations in tools such as Cursor. Since the R1 wave, however, Cheah says Featherless’s coding share rose from under 5% in January to 20–30% by the period discussed. That change is why January’s distribution would already give the wrong impression of the platform.

For an open tooling stack, he pairs Cline for the chat-driven agent with Continue for autocomplete. In the February 2025 setup he describes, R1 supplies an experience he compares with commercial models in Cursor. This is an interaction comparison, not a demonstrated performance-equivalence test. Hobbyists can also pursue private, offline setups, including Mac Studio clusters. His practical recommendation is to try the projects in VS Code. The accompanying frame shows Cline’s completion report beside React code and a terminal displaying a successful compilation; it captures a visible end state, not a general reliability measurement.

VS Code with Cline’s Task Completed summary, React code, and a terminal reading Compiled successfully, beside the presenter.
Cline reports task completion while the terminal shows a successful compilation.
13:0013:08
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:00 · section reference included

Graph workflows and alternative chat clients

Cheah assigns about 5% of traffic to ComfyUI and related personal agentic workflows. Graph interfaces familiar from diffusion and image generation also let people chain text operations. The users he names are not necessarily developers: reporters, lawyers, musicians, influencers and even poets assemble workflows to produce the material they need. As with coding agents, a handful of power users can generate substantial token volume. Unlike Cline’s sudden R1-related growth, this category has built gradually over months, amid competing platforms that remind him of JavaScript’s framework wars.

Cheah puts RAG and ChatGPT-style interfaces at about 20% of requests. Featherless has its own client, Phoenix, as other providers have theirs. His standout independent example is TypingMind: at talk time, a one-time purchase that users can run locally against an API provider such as Featherless. Its differentiators are product details—polished interaction, keeping pace with ChatGPT’s interface features and adding a plugin system. Running a client on a laptop does not by itself mean the model inference is local.

16:3716:52
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

16:37 · section reference included

An email agent that can reach production

Cheah assigns the remaining agents-and-work category 10–20% of Featherless traffic, then divides it into narrow workflow automation and fully automated agents. For enterprise teams, the immediate objective is useful adoption with limited downside. Risk-averse institutions have little reason to accept a spectacular demonstration that cannot be controlled in production. His driver's-seat analogy expresses the requirement: build the human escape hatch from day one.

Insurance claims and logistics provide a concrete workflow because so much of the work arrives by email. The proposed system proceeds in a controlled sequence:

  1. Read an inbound email and gather the relevant inventory, rules and other business context.
  2. Draft a response using those checks.
  3. Present the draft in a UI where a person can edit, finalize or replace it.
  4. Let that person send the response.

The UI can sit on an existing Salesforce or ERP system. When the draft is right, the remaining task is simply to click send; when it is wrong, the person already has a place to take over.

The important implementation boundary is between creating a draft and authorizing an external action. A small TypeScript state model can express that boundary for an illustrative logistics email: a customer asks for 12 units, inventory shows 8, and a rule forbids promising more than the available stock. Preparing the response creates a reviewable proposal; it does not send it.

typescript

type InboundEmail = {
  id: string;
  requestedUnits: number;
};

type ReplyDraft = {
  emailId: string;
  body: string;
  status: "needs-review";
};

type ApprovedReply = {
  emailId: string;
  body: string;
  status: "approved";
  reviewedBy: string;
};

function prepareDraft(
  email: InboundEmail,
  availableUnits: number
): ReplyDraft {
  const offeredUnits = Math.min(email.requestedUnits, availableUnits);
  return {
    emailId: email.id,
    body: `We can supply ${offeredUnits} of the ${email.requestedUnits} units requested.`,
    status: "needs-review"
  };
}

function approveReply(
  draft: ReplyDraft,
  editedBody: string,
  reviewerId: string
): ApprovedReply {
  if (!reviewerId.trim() || !editedBody.trim()) {
    throw new Error("A reviewer and a response are required.");
  }
  return {
    emailId: draft.emailId,
    body: editedBody,
    status: "approved",
    reviewedBy: reviewerId
  };
}

const email = { id: "email-42", requestedUnits: 12 };
const pendingReply = prepareDraft(email, 8);
// The review UI displays pendingReply; approval is a separate action.

The model-generated portion would supply the proposed wording; the surrounding application owns review and dispatch. In a deployed service, the server must enforce the approval boundary as well as the UI.

Cheah estimates that, when implemented well, this pattern can successfully draft around 80–90% of responses at launch, with humans taking over the remainder. Every submission still passes a person before reaching the customer. That is enough to improve productivity and make adoption attractive without requiring complete automation. Cheah places responsibility for missed corrections on the reviewer; operationally, the workflow must give that reviewer the context and ability to catch mistakes. A nominal approval button alone is not the same as effective review.

18:3518:49
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

18:35 · section reference included

Earn automation one class of cases at a time

After the reviewed workflow has run across thousands of customers, teams can identify classes of cases that consistently work. They can then add classification and enable automatic replies for those cases. Automation becomes a permission earned by a known workflow, rather than a global setting applied before its behavior is understood.

The opposite approach can launch to enthusiasm and fail later. A fully automated workflow without an escape hatch eventually encounters an unusual case and sends a damaging response. Sometimes management has knowingly accepted that trade-off because the savings justify it. Otherwise, the incident can kill the project and postpone further adoption for another year. Cheah connects that risk to teams spending too long building complete automation before putting a limited, useful version into production. His objection to perfectly reliable autonomous agents concerns real production use, not whether a proof of concept can look convincing.

The practical target is the tractable majority of work, with escape hatches for the rest. Cheah uses 80% as that target and argues that, inside a large corporation, even partial automation can be worth millions in productivity and savings. Humans are not perfectly reliable either. The decision is whether the eventual mistake is acceptable and repairable, not whether the system can promise never to make one.

Two examples make the boundary explicit:

WorkflowFailure the organization may accept
Cold outreachLosing a lead it would not otherwise have contacted
Customer supportAn error a human can later apologize for and correct

These are choices about consequences. They do not imply that every lost lead or support error is harmless; they identify where a business might consciously accept imperfect automation.

21:0521:17
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

21:05 · section reference included

Reliability needs a stable baseline

Cheah turns to reliability engineering, invoking Google SRE and safety-critical industries. Start with a streamlined system for the common cases. Then inspect the residual failures, solve a large fraction of those, and repeat. The mechanism is not simply retrying the same model ten times; it is repeatedly improving how the system handles the cases that remain unresolved.

He illustrates this with repeated 80–90% improvements to the remaining failure cases and cites 99.998% reliability after ten repetitions. That percentage is not derived in the talk. An idealized calculation makes the assumptions explicit: if each round fixes a fraction c of the remaining failures and never breaks a previously working case, then after n rounds:

Fn=F0(1c)nRn=1Fn\begin{aligned} F_n &= F_0(1-c)^n \\ R_n &= 1-F_n \end{aligned}

Here F₀ is the initial failure fraction and Rₙ is the resulting success fraction. Real systems must measure both the remaining failures and regressions; the recurrence alone cannot establish production reliability. Cheah’s airline joke—repeat the process again—underscores the incremental approach. A less spectacular first launch can survive, deliver value and create the conditions for further improvement.

That brings the discussion back to old models. Stable dependencies let teams accumulate improvements, even if leaving models unchanged also gives their hosting provider continued business. Cheah nevertheless urges Featherless’s Llama 2 customers to evaluate an upgrade. He suggests a possible improvement from 99% to 99.9%, but does not name the metric or evaluation conditions. The useful distinction is between choosing an upgrade and having one imposed: Featherless’s promise to keep hosting the old model preserves the customer’s ability to decide.

His corresponding request to closed-model labs is Git-style versioning. A team cannot easily build successive reliability gains on a baseline that changes every week. Stable model versions and deliberate upgrades belong together: preserve the version that works, evaluate a replacement against the actual workload, and change it when the evidence supports doing so.

23:5824:09
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

23:58 · section reference included

Qwerky and the next reliability problem

The closing announcement arrives through the Qwerky persona. It introduces Qwerky as a 72-billion-parameter hybrid of linear and attention transformers, claiming less than half the GPU compute cost of other transformer models and calling it the strongest post-transformer hybrid to date. The proposed alternative to attention-only models is motivated by inference cost and by the possibility that a different architecture can scale. Those comparative claims arrive without a specified workload or benchmark in the announcement.

The announcement puts Qwerky’s build cost at $100,000 and contrasts it with a stated $10 million for DeepSeek. The accounting scopes are not defined, so this is not a like-for-like comparison of total model-development costs. It is a claim about the feasibility of pursuing an alternative architecture with a smaller budget, not evidence that the two projects purchased equivalent training work.

The final question returns to usefulness. Being able to answer advanced mathematics or orbital-calculation questions is different from being trusted to complete one assigned task without hallucination or failure. The persona argues that MMLU becomes less informative for this purpose once models exceed the knowledge needed in ordinary office work. The gap is between answering difficult questions and behaving dependably inside a workflow.

The proposed research direction is to use linear-transformer models for persistent memories, customization and more reliable agents. These are possibilities to investigate, not capabilities established by the announcement. The concrete invitation is to discuss running Featherless or Qwerky in a private cloud or on premises—bringing the talk back to the user’s actual workload, the environment where it must run, and the behavior it must preserve.

26:2226:34
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

26:22 · section reference included

Resources

From the talk

Updates since the talk

  • Later Featherless linear-attention model release with weights, inference instructions and comparisons against Qwen2.5-72B-Instruct.

  • RADLADSPaper

    Follow-up research coauthored by Eugene Cheah on converting pretrained transformers into linear-attention models at 7B, 32B and 72B scales.

Read the complete timestamped transcript
  1. 0:01

    There is no moat. The open model warriors are climbing at the gates, scaling the benchmark, and more importantly, the hearts of their users. In the past 12 months, more than fifty thousand AI models have been uploaded to Hugging Face per month, and it's accelerating.

  2. 0:13

    That is more than one AI model a minute, meaning by the end of this talk today, that's another five hundred models alone. However, not all models are equal. The whale that crashed into the room is DeepSeek-R1, the first open source model to catch up and surpass GPT-4o, proving that you do not need a billion dollars to compete

  3. 0:28

    with the Big Labs, just a little smart and prudence to catch up. With over four million downloads of a six hundred and eighty-five gigabyte model on Hugging Face this past month alone, this one model exceeds over two point seven four exabytes of data moving across the internet.

  4. 0:41

    A number so ridiculously large that it would take my fast one gigabit internet six hundred and ninety years to transmit. That's a lot of open models. So what the F do people use these open models for?

  5. 0:52

    Hey there, I'm Eugene, CEO of Featherless AI, and co-lead for the RWKV open source project. At Featherless, we provide unlimited API requests to over three thousand seven hundred truly open AI models to thousands of users.

  6. 1:05

    With a twenty-five dollars a month flat pricing for individual, with access to all our models, including the famous DeepSeek-R1, with larger plans for scaled up business users. Our goal at the twenty-five dollars price point is to provide accessibility to all truly open AI models, with the goal of continuously expanding our catalog to eventually cover all Hugging Face

  7. 1:29

    models. All of which give us a unique data and insight in which how these open models are actually getting used via our platform, the how and the why, and by reflection, the open source committee as well.

  8. 1:45

    So let's get the pie chart out. Let's focus on our individuals users first. For a week in February 2025, unsurprisingly, most individuals are using DeepSeek-R1, followed by Llama 3 line of models, the Mistral Nemo 12B, and the Qwen line of models, with everything else effectively buried.

  9. 2:06

    One important thing to note on this data for individual users is that they are not billed or charged by tokens. For Featherless, we are instead limiting, uh, uh, the plans to one large model request at a time, with smaller models typically being faster than larger ones.

  10. 2:25

    There is no price difference for individuals here for-- between the models. The end result is users choosing their model based on their preference and vibes, and perhaps the fame of the model instead of MMLU or price.

  11. 2:39

    And while model classes is exciting from a model family worldview, it is masking the details in the story. So let's switch our view to models names instead of model class, which give us a new perspective.

  12. 2:52

    This is where fine-tuning fragmentation starts to begin, where the larger Llama and Qwen pie gets sliced up into individual models, all with different personalities and individual use case.

  13. 3:04

    However, the first surprising insight I would make is the staying power of a model once a model is used in production, because the needs for business in production and users differ dramatically.

  14. 3:17

    In reality, most developers want consistency. They want to only change their model when they choose to do so or opt in to do so, not when their provider decided to do an update.

  15. 3:29

    Nothing shows this better than the Mistral Nemo 12B. Because while this chart represents individuals, I intentionally excluded scaled commercial users for a reason. Because when we add it up, scaled, cons-- uh, commercial users, which was previously excluded, things change dramatically.

  16. 3:46

    Even when you switch over to the data by model class, the dominance of much smaller M-Mistral Nemo models, which at this point is over eight months old and has been essentially replaced by much larger and better model, is a surprising observation.

  17. 4:01

    This is contributed by the following factors. One, small models in general are cheaper at scale, be it at other platforms or ours. In our case, we provide access to four small models for the same price as one big model.

  18. 4:15

    The staying power of models in production and tutorials is a lot more sticky, uh, as most people think. S-things, uh, that enters production in, in enterprises tends to not be changed for quarters or, uh, or even years at a time.

  19. 4:30

    You see, the Mistral Nemo model was one of the early truly open source models with Apache 2 licensing that surpassed even the GPT-3.5 class. This caused the early shift from closed models to open model, and more importantly, it was without the Llama license restriction that many enterprise lawyers were rather uncomfortable about.

  20. 4:49

    As such, one of the side effect was they became the default model for lots of cloud pro-- platforms, including AWS, GCP, etc., with hundreds of fine-tuning tutorials. And the models was pushed heavily in replacing existing users' GPT-3.5 workload for the respective cloud providers, resulting into lots of production environments today running on these models, despite being it outdated

  21. 5:13

    by today's AI year standards. One thing many people miss is that for many commercial and enterprise use cases, once they get something working reliably at scale, especially once they have the metrics in place to observe changes and the, their reliability, and they have prompted it to ninety-nine percent plus accuracy, they really do not want to change their

  22. 5:33

    system and have it break overnight. Sure, they can get an AI engineer to update the system instruction prompt every few weeks when a model updates, or they can use a stable model version in open source.

  23. 5:44

    And here's the crazy thing, not change the model. Because if it ain't broke, don't fix it. So another standout literally is the Llama 2, Llama Guard model. Llama 2 is about two percent of our work-workload, despite having newer, better versions of this model for basically everything in, in this model class.

  24. 6:05

    It's still the go-to model for AI safeguard tutorials online because it, it hasn't been updated, and it's being actively used in production even for new teams that come online.

  25. 6:15

    So that just bring us to the topic,

  26. 6:18

    what do people use open models for? You see, despite us being privacy friendly as a platform with strict no text logging as a policy, we can infer usage from the model, uh, being used and the metadata for the requesting application.

  27. 6:32

    Like what else we use Llama Guard for other than safeguards. More importantly, we can straight up ask our customers and exchange notes with aggregators like Open Routers, which we did.

  28. 6:43

    And here's what we see AI usage today, sorted approximately by volume and usage. AI for creativity or friendship, AI for code, AI for ComfyUI, AI for RAG and ChatGPT, very classic, AI for agent and work.

  29. 7:02

    That is not in the previous categories. All of this is based on data that, uh, used by our customers directly as individual or as business for internal usage at scale.

  30. 7:13

    Or-- Which in some cases could just mean that it's repackaging our API as an app for the App Store,

  31. 7:19

    or AI web apps or agents that you play with online.

  32. 7:24

    So getting back to it, creative writing, AI roleplay and companionship, and to a much smaller extent, therapy and journaling.

  33. 7:33

    This use case represent the vast majority of non-coding AI requests, representing thirty to forty percent of all AI traffic at any point of time. For creative writing, it's apps like NovelCrafter, a tool designed specifically to outline, manage, and draft your novel.

  34. 7:50

    From the series lore, the high-level narration and collaborative writing on the text with other humans and AI. This segment obviously is particularly popular with authors and with the fan fiction community.

  35. 8:01

    A small adjacent segment to this community are folks who use AI for creative content in games, in particular Dungeons & Dragons style kind of content.

  36. 8:11

    For the other seg-- For the other group, AI role-play and companionship, we have apps like Wyvern Chat, Silly Tavern, and a whole lots of them. While there's a lot of social stigma around this segment due to its association with explicit or sexual content, it is still the number one use case for AI today for non-code or agentic

  37. 8:29

    workflows. And the u- and it's the use case with the most number of active users, be it directly or through the App Store which uses our API under the hood.

  38. 8:40

    It is also why Sam Altman and gang can't stop teasing about the movie Her, because it's still a, a large percent of closed source AI movie, uh, usage. Also, it's twenty-twenty five.

  39. 8:52

    Let's get the gender stereotype out of here, 'cause unlike popular belief, it's not about you males. Over sixty percent of the users in this segment are [REDACTED:gender]. This mirrors a well-known pattern where romance novels targeting [REDACTED:gender] is the number one sales of books by volume by a really huge margin.

  40. 9:13

    However, you won't see a bookstore, uh, presenting themselves that way due to the social stigma and stereotype around it, at least for most bookstores.

  41. 9:22

    Likewise, the same here is happening for AI.

  42. 9:25

    So to further fight that stereotype that all this is about R-rated male content, let me quote one of the app makers here.

  43. 9:32

    "[REDACTED:gender] tend to use these apps to hold long conversations and de-stress talk- talking about their day-to-day lives, uh, every day, in part because their real-life partners are either non-existent or emotionally unavailable for them."

  44. 9:48

    So ouch. If you are worried about AI stealing your partner, learn to talk to her and be emotionally available, which ties in to another usage trend that we see,

  45. 9:59

    therapy and journaling. There are lots of dedicated commercial apps on the App Store, but when we ask users what apps they use, it is heavily fragmented. It seems be it before or after AI, Silicon Valley can't help but create another journaling app every single month.

  46. 10:18

    One thing to note separately, a good percentage of users in, in our group that we, we surveyed just essentially said that they use ChatGPT clones or their companion apps with a therapy character for the same use case, essentially not being dependent on a specific therapy or journaling app.

  47. 10:39

    Taking a few steps back into this broader category, there's a reason why I group all these use cases all together. In general, this segment-- for this segment, no one cares about MMLU.

  48. 10:52

    This segment is all about vibes, vibes, vibes, and it's the use case for a large segment of various community fine-tuned with various different names, and it's extremely competitive with top models ranking constantly changing weekly.

  49. 11:06

    So what do I mean by vibes? For one, it's ironically doing the one thing that is the complete opposite for every AI model instruction tune. It is not rushing to answer the question, or even solve or answer the question.

  50. 11:20

    For example, in therapy and journaling, the job of the therapist, or in this case the AI, is to guide and empower a client to do what they want to reach their goal, not tell you what to do.

  51. 11:30

    The answer has to come from the individual.

  52. 11:33

    Likewise, in story writing, role-play, and companionship, the term used is slow burn. No one is watching Game of Thrones by starting from episode one and skipping everything by jumping the last episode.

  53. 11:45

    We came here for the entertainment in between, not to talk about the ending at a dinner table.

  54. 11:51

    So while there's a new exciting flavor for the com- community every few days, just like there are new books, every one of these models are unique and different. Want a model that can go howdy and all cowboy, there is the legendary Alpine Dale Magnum mo- model.

  55. 12:07

    Want a model that specifically focus on science fiction, something which apparently a lot models are really bad at, use the Rossini model.

  56. 12:15

    Given the speed this community moves, the models I name will probably be replaced by next month, frankly.

  57. 12:22

    But when talking to the community in this space, where everyone plays and hops around model s- dailies, focusing on the hot new celebrity tune for the week, they sometimes return to their old favorite.

  58. 12:35

    As one user put it, "It's like playing a retro game from the memories. It has its charms."

  59. 12:41

    And I will emphasize that this is a category that has the largest scale. For a sense of scale, the current size of this user base and community, including closed so- closed source commercial apps like Character.ai, is in the tens of millions.

  60. 12:56

    Everyone else is not even, it's not even in the millions yet.

  61. 13:00

    So the next major segment is coding, Copilot, and coding agents,

  62. 13:08

    which is about 20 to p- 30% of all traffic.

  63. 13:13

    In general, there are two segments, uh, to this group. Auto-completion tools, similar to the original GitHub Copilot or editing by chat, it-- which is basically available natively or through plugin to every IDE today.

  64. 13:27

    You are probably not even considered an IDE anymore if you don't have these features as an option.

  65. 13:34

    These features are now being served by dozens of AI model that are small and arguably good enough f- for, uh, with parameter size ranging from three billion to 12 billion from Reply, Mistral, and nearly every AI lab under the sun.

  66. 13:46

    I would argue auto-completion for code for this part is essentially solved. So the battleground for AI and code has moved to more agentic code, uh, agents. And I don't just mean the s- SWE agent or Devin style agents, which is apparently too autonomous and goes into too many infinite loops for many d- developers.

  67. 14:08

    I mean nearly autonomous agent with lots of clarifying questions, back and forth interventions with, with humans-in-the-loop kind of agent, where humans can rapidly step in and make a tweak or suggest changes in chat whenever something goes slightly wrong.

  68. 14:22

    The trending term for this phenomenon is called vibe coding, where basically you never touch the code and just keep chatting and prompting all your changes.

  69. 14:32

    Also, because of how token hungry these agents can be, one thing to note is how much traffic it asks for compared to the users of the previous segment. A developer fully in the flow of vibe coding could essentially generate 1,000 times more input and output tokens traffic back and forth than a single person chatting to a companion

  70. 14:50

    model. So where there are at most tens of thousands of coders there, uh, using these models at any point in time, this traffic volume is growing ridiculously fast in the past week.

  71. 15:03

    One s- side note to be clear, cross-referencing data from OpenRouter right now, the vast majority of traffic is dominated by CloudSon- Sonnet for this use case.

  72. 15:16

    And we are not talking about a s- a small difference. We're talking about over a 10 to one ratio here. Most tools that support this more agentic coding workflow like Cursor is still integrate heavily towards the more commercial models.

  73. 15:29

    However, because of the volume of the use case and since the R1 wave, we have seen a rapid growth in using such tools with open models. This is part of the reason why there is no use for me presenting January data, because this coding segment itself went from basically under the 5% threshold in volume to easily now

  74. 15:49

    20 to 30% of all traffic. The one major open source project that sticks out that I would like to highlight is Cline, which focuses on the chat agentic flow.

  75. 16:01

    And when you combine it with the other open, uh, source project, Continuous Dev, for the auto-completion integration, you basically have the same level of experience as Cloud or OpenAI models with Cursor or any other, uh, IDE platforms with agentic workflow via R1.

  76. 16:17

    Or we are possibly using your own private MacStudio cluster under the [REDACTED:location],

  77. 16:22

    actually popular for, like, all these hobbyist users who actually want to do this completely without the internet.

  78. 16:29

    Highly recommended that you g- uh, give these open projects a try, uh, in VS Code before they get forked a dozen times in the next YC.

  79. 16:37

    Next up is ComfyUI and friends. About 5% of the traffic for personal agentic workflows. Well, many of you may be familiar with these two being widely used in the diffusion space for more complicated AI generation-- image generation workflows.

  80. 16:52

    S- similar, uh, graph style based UIs are now being used, right, from reporters, lawyers, musicians, influencers, and even poets somehow, th-- where, where these people are chaining complicated workflows to generate the text that they want for their use case.

  81. 17:09

    Similar to coding agentic workflow, the token explosion from a handful of these power user slowly stack up to noticeable levels. But these users are not developers. They are literally, like I said, musicians or, like, all these creative personnels who learn how to use these tools.

  82. 17:27

    However, unlike Cline, which saw a sudden explosion due to R1, the past few weeks, this has been more about a slow and gradual user base build up in the past months.

  83. 17:40

    It's also a space like most agentic platforms, uh, where it-- there's essentially what I call a framework war, if-- for those who are familiar with the JavaScript era.

  84. 17:51

    The next major segment, RAG and Ch- ChatGPT clones, essentially. About 20% of all requests at this point. Well, every platform, be it Featherless or OpenRouter or every other AI provider will have their own internal ChatGPT UI.

  85. 18:05

    Ours is called Phoenix. A standout application in this category that I would like to name is TypingMind. An indie application which, surprise, surprise, bills you with a one-time fee instead of a subscription, and can be used locally on your laptop against any API provider like us.

  86. 18:24

    In particular, it's the level of UI polish and how they've been keeping pace with cloning every single ChatGPT UI feature into their app, and customizing that with even their own plugin system.

  87. 18:35

    Beyond that, you probably heard this use case a thousand times. ChatGPT, RAG, yada, yada. So let's skip to the one that everyone wants to hear more about. AI for agents and work, representing 10 to 20% of all our traffic.

  88. 18:49

    Because agent is such an over-abused term today, I'll split this into two major categories to make things clear. Workflow automation or narrow agents, and fully automated agents.

  89. 19:04

    If I, uh, so right now, let's start from workflow automation or narrow agent. For most of you who's working on agents in enterprises, your number one priority now is to get all your agents into production by maximizing ROI for your company and shareholder value, if you want to use that term, without limiting the negative impact.

  90. 19:26

    Especially since we are in New York, lots of financial institutes are risk-averse by default, with many of you potentially working in there developing your agents right now. So have that stance, build in human escape hatches by default.

  91. 19:41

    What I mean by that is, basically, if you are building an automation system in your company,

  92. 19:47

    like Waymo, make it so that the human can take the driver's seat needed from day one.

  93. 19:53

    So for example, we work with a few insurance claim companies and logistic companies advising them on how to fully automate their email process, which is the bulk of their work.

  94. 20:03

    A common pattern we advise them would be them to create an AI agent and UI for fully automating drafting responses for inbound emails, checking against inventories, checking against rules, checking a lot of other things in the process,

  95. 20:16

    and build a platform UI for editing and finalizing and sending a response.

  96. 20:21

    In best case scenario, if the AI did everything perfectly, it's just all about clicking send. And sometimes these things are built on top of their existing customers' Salesforce or ERP system.

  97. 20:33

    However, more importantly, it at least have one human checking the final submissions before it hits the, the end user. When done right, at launch, the AI is able to successfully draft around 80, 90% of the response, with humans rejecting and manually taking over the remainder, so there's no harm to everyone else.

  98. 20:51

    Productivity for the responders shoots through the roof. Management is happy. The AI is adopted.

  99. 20:58

    And if any mistakes were made and the human failed to correct it, well, that's kind of the human's fault, not the AI.

  100. 21:05

    Over time, as confidence start to build, they then start to add classification for certain known reliable use case, and eventually start fully automating those use cases.

  101. 21:17

    It's like auto-reply to certain, certain scenarios. With full confidence after having seen the system run for thousands of customers.

  102. 21:26

    On the flip side, teams that went full, fully ambitious into full 100% automation without human escape hatches,

  103. 21:35

    sometimes they start with a successful launch because everyone wanted it to succeed. But more commonly, as they run the system, and as more users use it, they eventually get fully burnt by an angry customer for a really bad automated response down the line.

  104. 21:50

    In some cases, this may be fully acceptable compromise that management accepts from the start because the savings benefit is worth it. Other times, this kills the entire AI project, essentially postponing AI adoption in your organization for another year, a scenario that you do not want to happen.

  105. 22:12

    And this is commonly because these, a lot of teams end up developing an entire fully automated workflow without trying to move it into production early. That's why I say your goal is to go beyond the POC phase and into adoption in your company at scale in some portion as fast as possible, and to inter- integrate and iterate

  106. 22:36

    from there. So that leaves the other mythical category where, which everyone loves to chase. Fully, truly 100% reliable agentic agents that is fully automated. Just don't. It does not exist.

  107. 22:51

    Everyone has, who has done it, tried to done it, has got burnt.

  108. 22:56

    And it's, we are not, and we are talking about things in production, not POC.

  109. 23:04

    The mindset that we should be thinking when we try to build AIs into production

  110. 23:09

    is that we should approach things from, like, trying to solve the 80% with escape hatches along the way.

  111. 23:18

    Because at, even at 80%, for some of you mega corps here, that's already millions in productivity and saving.

  112. 23:27

    Furthermore, do I need to remind you that even humans are not 100% reliable?

  113. 23:33

    So only do this when the eventual mistake, right, is acceptable and can be fixed.

  114. 23:40

    And it's a basically a trade-off that you're willing to take. So common example is cold calling systems when you're actually willing to lose a lead because you did, you wouldn't have gotten it otherwise anyway.

  115. 23:52

    Or sometimes customer support because you can get a human to apologize later.

  116. 23:58

    But rewinding back, it's like you should take the approach just like how Google software reliability engineers or any reliability engineering for industry where safety and reliability is paramount do.

  117. 24:09

    Start with an extremely streamlined, reliable system for your 80 to 90% work scenarios. And after you're done, do it for another 80 to 90% for that 10 to 20% failure scenarios and repeat this process 10 times.

  118. 24:24

    And now suddenly you have 99.998% reliability. If, if you're an airliner that's not Boeing, do it another 10 times and everyone is happy with incremental gains and pro-- and improvements.

  119. 24:37

    It may not be the sexy first launch. It may not be the most wow to do it this way, but at the very least, your project survives to production, you have huge impact to ROI, and you can iterate from there.

  120. 24:53

    And more importantly, incrementally improve and amusingly, probably never change some of the models that you use and forever give me business, which I'm ironically going to do and say the following.

  121. 25:06

    Hey, Featherless customers, especially the more enterprising ones, please try to upgrade from Llama 2 to something newer this year or next. If done right, it's probably a free 99 to 99.9% improvement.

  122. 25:17

    I know all of you are using this model and came to us in part because we are one of the last few providers to host it, and we forever will.

  123. 25:24

    But hey, I'm giving this advice. It's free real estate.

  124. 25:28

    Similarly, hey, for closed source labs, it's 2025. You do not need AGI to learn how to use Git version from AI models.

  125. 25:39

    I say this in all seriousness because for companies in production who is trying to incrementally improve with each step to 99 plus 9999%, it's nearly impossible to do so if you change the model every week.

  126. 25:53

    So with that, ah, shucks, I'm out of time.

  127. 25:58

    All the best in building your AI agent.

  128. 26:01

    Hey, hey, Eugene, come back here. You're forgetting the one more thing bit. Ah, crap, he can't hear me. The laptop speakers are muted, and I have no control over that.

  129. 26:11

    Uh, you see, us AI may be flawed, but honestly, we are probably more consistent and reliable than humans are, as you can clearly see. Uh, oh, well, here goes me saving the day.

  130. 26:22

    Hi, I'm Quirky, a seventy-two billion parameter linear transformer and attention transformer hybrid with a completely new architecture that runs at less than half the GPU's compute cost of other transformer models.

  131. 26:34

    The strongest post-transformer hybrid to date. When Eugene proposed training a post-transformer model to provide an alternative to attention is all you need, which runs at a fraction of the inference cost, many of you said, "It can't be done.

  132. 26:47

    The tech is unproven and will not scale." Well, DeepSeek cost ten million. Uh, Quirky cost only a one hundred thousand to build. You can find out more about our model at the following link.

  133. 27:00

    More importantly, as we enter the era where the average AI model has better MMLU than the average office worker, the benchmark has lost its meaning. After all, how many of you audience can calculate orbital reentry from Earth to Mars or PhD math questions?

  134. 27:17

    All of which are questions apparently frontier AI models are now capable of. How many of our AI agents today can be reliably be trusted with one task it's instructed to do so without hallucination or failure?

  135. 27:31

    I think you get the point. What we find more exciting is exploring a future where we take advantage of linear transformer models as a means of persisting memories, customization, and improving reliability for future AI models to make useful AI agents.

  136. 27:47

    And if you would like to run Featherless or Quirky in your private cloud or on-premise environment, reach out to us and we would like to hear more about your use case.