AI Engineer Summit 2025
WTF do people use Open Models for??
About this talk
Featherless.ai CEO Eugene Cheah examines how people and businesses actually use open-weight models, contrasting individual demand for DeepSeek-R1, Llama, and Qwen with enterprise reliance on stable, permissively licensed Mistral NeMo deployments. He surveys companionship, creative writing, coding, retrieval, and agent applications, emphasizing model stability, token consumption, production reliability, and future memory-oriented linear-transformer architectures. A second unidentified speaking cluster appears near the closing segment.
Chapters
- 0:01Open-model growth, DeepSeek-R1, and Featherless.ai
- 1:45Individual preferences and enterprise model stability
- 6:43Companionship, creativity, and fragmented AI applications
- 14:32Coding-agent token demand and production reliability
- 26:01Closing: benchmark limitations and memory-oriented models
Talk transcript
- 0:01
There is no moat. The open model warriors are climbing at the gates, scaling the benchmark, and more importantly, the hearts of their users. In the past 12 months, more than fifty thousand AI models have been uploaded to Hugging Face per month, and it's accelerating.
- 0:13
That is more than one AI model a minute, meaning by the end of this talk today, that's another five hundred models alone. However, not all models are equal. The whale that crashed into the room is DeepSeek-R1, the first open source model to catch up and surpass GPT-4o, proving that you do not need a billion dollars to compete
- 0:28
with the Big Labs, just a little smart and prudence to catch up. With over four million downloads of a six hundred and eighty-five gigabyte model on Hugging Face this past month alone, this one model exceeds over two point seven four exabytes of data moving across the internet.
- 0:41
A number so ridiculously large that it would take my fast one gigabit internet six hundred and ninety years to transmit. That's a lot of open models. So what the F do people use these open models for?
- 0:52
Hey there, I'm Eugene, CEO of Featherless AI, and co-lead for the RWKV open source project. At Featherless, we provide unlimited API requests to over three thousand seven hundred truly open AI models to thousands of users.
- 1:05
With a twenty-five dollars a month flat pricing for individual, with access to all our models, including the famous DeepSeek-R1, with larger plans for scaled up business users. Our goal at the twenty-five dollars price point is to provide accessibility to all truly open AI models, with the goal of continuously expanding our catalog to eventually cover all Hugging Face
- 1:29
models. All of which give us a unique data and insight in which how these open models are actually getting used via our platform, the how and the why, and by reflection, the open source committee as well.
- 1:45
So let's get the pie chart out. Let's focus on our individuals users first. For a week in February 2025, unsurprisingly, most individuals are using DeepSeek-R1, followed by Llama 3 line of models, the Mistral Nemo 12B, and the Qwen line of models, with everything else effectively buried.
- 2:06
One important thing to note on this data for individual users is that they are not billed or charged by tokens. For Featherless, we are instead limiting, uh, uh, the plans to one large model request at a time, with smaller models typically being faster than larger ones.
- 2:25
There is no price difference for individuals here for-- between the models. The end result is users choosing their model based on their preference and vibes, and perhaps the fame of the model instead of MMLU or price.
- 2:39
And while model classes is exciting from a model family worldview, it is masking the details in the story. So let's switch our view to models names instead of model class, which give us a new perspective.
- 2:52
This is where fine-tuning fragmentation starts to begin, where the larger Llama and Qwen pie gets sliced up into individual models, all with different personalities and individual use case.
- 3:04
However, the first surprising insight I would make is the staying power of a model once a model is used in production, because the needs for business in production and users differ dramatically.
- 3:17
In reality, most developers want consistency. They want to only change their model when they choose to do so or opt in to do so, not when their provider decided to do an update.
- 3:29
Nothing shows this better than the Mistral Nemo 12B. Because while this chart represents individuals, I intentionally excluded scaled commercial users for a reason. Because when we add it up, scaled, cons-- uh, commercial users, which was previously excluded, things change dramatically.
- 3:46
Even when you switch over to the data by model class, the dominance of much smaller M-Mistral Nemo models, which at this point is over eight months old and has been essentially replaced by much larger and better model, is a surprising observation.
- 4:01
This is contributed by the following factors. One, small models in general are cheaper at scale, be it at other platforms or ours. In our case, we provide access to four small models for the same price as one big model.
- 4:15
The staying power of models in production and tutorials is a lot more sticky, uh, as most people think. S-things, uh, that enters production in, in enterprises tends to not be changed for quarters or, uh, or even years at a time.
- 4:30
You see, the Mistral Nemo model was one of the early truly open source models with Apache 2 licensing that surpassed even the GPT-3.5 class. This caused the early shift from closed models to open model, and more importantly, it was without the Llama license restriction that many enterprise lawyers were rather uncomfortable about.
- 4:49
As such, one of the side effect was they became the default model for lots of cloud pro-- platforms, including AWS, GCP, etc., with hundreds of fine-tuning tutorials. And the models was pushed heavily in replacing existing users' GPT-3.5 workload for the respective cloud providers, resulting into lots of production environments today running on these models, despite being it outdated
- 5:13
by today's AI year standards. One thing many people miss is that for many commercial and enterprise use cases, once they get something working reliably at scale, especially once they have the metrics in place to observe changes and the, their reliability, and they have prompted it to ninety-nine percent plus accuracy, they really do not want to change their
- 5:33
system and have it break overnight. Sure, they can get an AI engineer to update the system instruction prompt every few weeks when a model updates, or they can use a stable model version in open source.
- 5:44
And here's the crazy thing, not change the model. Because if it ain't broke, don't fix it. So another standout literally is the Llama 2, Llama Guard model. Llama 2 is about two percent of our work-workload, despite having newer, better versions of this model for basically everything in, in this model class.
- 6:05
It's still the go-to model for AI safeguard tutorials online because it, it hasn't been updated, and it's being actively used in production even for new teams that come online.
- 6:15
So that just bring us to the topic,
- 6:18
what do people use open models for? You see, despite us being privacy friendly as a platform with strict no text logging as a policy, we can infer usage from the model, uh, being used and the metadata for the requesting application.
- 6:32
Like what else we use Llama Guard for other than safeguards. More importantly, we can straight up ask our customers and exchange notes with aggregators like Open Routers, which we did.
- 6:43
And here's what we see AI usage today, sorted approximately by volume and usage. AI for creativity or friendship, AI for code, AI for ComfyUI, AI for RAG and ChatGPT, very classic, AI for agent and work.
- 7:02
That is not in the previous categories. All of this is based on data that, uh, used by our customers directly as individual or as business for internal usage at scale.
- 7:13
Or-- Which in some cases could just mean that it's repackaging our API as an app for the App Store,
- 7:19
or AI web apps or agents that you play with online.
- 7:24
So getting back to it, creative writing, AI roleplay and companionship, and to a much smaller extent, therapy and journaling.
- 7:33
This use case represent the vast majority of non-coding AI requests, representing thirty to forty percent of all AI traffic at any point of time. For creative writing, it's apps like NovelCrafter, a tool designed specifically to outline, manage, and draft your novel.
- 7:50
From the series lore, the high-level narration and collaborative writing on the text with other humans and AI. This segment obviously is particularly popular with authors and with the fan fiction community.
- 8:01
A small adjacent segment to this community are folks who use AI for creative content in games, in particular Dungeons & Dragons style kind of content.
- 8:11
For the other seg-- For the other group, AI role-play and companionship, we have apps like Wyvern Chat, Silly Tavern, and a whole lots of them. While there's a lot of social stigma around this segment due to its association with explicit or sexual content, it is still the number one use case for AI today for non-code or agentic
- 8:29
workflows. And the u- and it's the use case with the most number of active users, be it directly or through the App Store which uses our API under the hood.
- 8:40
It is also why Sam Altman and gang can't stop teasing about the movie Her, because it's still a, a large percent of closed source AI movie, uh, usage. Also, it's twenty-twenty five.
- 8:52
Let's get the gender stereotype out of here, 'cause unlike popular belief, it's not about you males. Over sixty percent of the users in this segment are [REDACTED:gender]. This mirrors a well-known pattern where romance novels targeting [REDACTED:gender] is the number one sales of books by volume by a really huge margin.
- 9:13
However, you won't see a bookstore, uh, presenting themselves that way due to the social stigma and stereotype around it, at least for most bookstores.
- 9:22
Likewise, the same here is happening for AI.
- 9:25
So to further fight that stereotype that all this is about R-rated male content, let me quote one of the app makers here.
- 9:32
"[REDACTED:gender] tend to use these apps to hold long conversations and de-stress talk- talking about their day-to-day lives, uh, every day, in part because their real-life partners are either non-existent or emotionally unavailable for them."
- 9:48
So ouch. If you are worried about AI stealing your partner, learn to talk to her and be emotionally available, which ties in to another usage trend that we see,
- 9:59
therapy and journaling. There are lots of dedicated commercial apps on the App Store, but when we ask users what apps they use, it is heavily fragmented. It seems be it before or after AI, Silicon Valley can't help but create another journaling app every single month.
- 10:18
One thing to note separately, a good percentage of users in, in our group that we, we surveyed just essentially said that they use ChatGPT clones or their companion apps with a therapy character for the same use case, essentially not being dependent on a specific therapy or journaling app.
- 10:39
Taking a few steps back into this broader category, there's a reason why I group all these use cases all together. In general, this segment-- for this segment, no one cares about MMLU.
- 10:52
This segment is all about vibes, vibes, vibes, and it's the use case for a large segment of various community fine-tuned with various different names, and it's extremely competitive with top models ranking constantly changing weekly.
- 11:06
So what do I mean by vibes? For one, it's ironically doing the one thing that is the complete opposite for every AI model instruction tune. It is not rushing to answer the question, or even solve or answer the question.
- 11:20
For example, in therapy and journaling, the job of the therapist, or in this case the AI, is to guide and empower a client to do what they want to reach their goal, not tell you what to do.
- 11:30
The answer has to come from the individual.
- 11:33
Likewise, in story writing, role-play, and companionship, the term used is slow burn. No one is watching Game of Thrones by starting from episode one and skipping everything by jumping the last episode.
- 11:45
We came here for the entertainment in between, not to talk about the ending at a dinner table.
- 11:51
So while there's a new exciting flavor for the com- community every few days, just like there are new books, every one of these models are unique and different. Want a model that can go howdy and all cowboy, there is the legendary Alpine Dale Magnum mo- model.
- 12:07
Want a model that specifically focus on science fiction, something which apparently a lot models are really bad at, use the Rossini model.
- 12:15
Given the speed this community moves, the models I name will probably be replaced by next month, frankly.
- 12:22
But when talking to the community in this space, where everyone plays and hops around model s- dailies, focusing on the hot new celebrity tune for the week, they sometimes return to their old favorite.
- 12:35
As one user put it, "It's like playing a retro game from the memories. It has its charms."
- 12:41
And I will emphasize that this is a category that has the largest scale. For a sense of scale, the current size of this user base and community, including closed so- closed source commercial apps like Character.ai, is in the tens of millions.
- 12:56
Everyone else is not even, it's not even in the millions yet.
- 13:00
So the next major segment is coding, Copilot, and coding agents,
- 13:08
which is about 20 to p- 30% of all traffic.
- 13:13
In general, there are two segments, uh, to this group. Auto-completion tools, similar to the original GitHub Copilot or editing by chat, it-- which is basically available natively or through plugin to every IDE today.
- 13:27
You are probably not even considered an IDE anymore if you don't have these features as an option.
- 13:34
These features are now being served by dozens of AI model that are small and arguably good enough f- for, uh, with parameter size ranging from three billion to 12 billion from Reply, Mistral, and nearly every AI lab under the sun.
- 13:46
I would argue auto-completion for code for this part is essentially solved. So the battleground for AI and code has moved to more agentic code, uh, agents. And I don't just mean the s- SWE agent or Devin style agents, which is apparently too autonomous and goes into too many infinite loops for many d- developers.
- 14:08
I mean nearly autonomous agent with lots of clarifying questions, back and forth interventions with, with humans-in-the-loop kind of agent, where humans can rapidly step in and make a tweak or suggest changes in chat whenever something goes slightly wrong.
- 14:22
The trending term for this phenomenon is called vibe coding, where basically you never touch the code and just keep chatting and prompting all your changes.
- 14:32
Also, because of how token hungry these agents can be, one thing to note is how much traffic it asks for compared to the users of the previous segment. A developer fully in the flow of vibe coding could essentially generate 1,000 times more input and output tokens traffic back and forth than a single person chatting to a companion
- 14:50
model. So where there are at most tens of thousands of coders there, uh, using these models at any point in time, this traffic volume is growing ridiculously fast in the past week.
- 15:03
One s- side note to be clear, cross-referencing data from OpenRouter right now, the vast majority of traffic is dominated by CloudSon- Sonnet for this use case.
- 15:16
And we are not talking about a s- a small difference. We're talking about over a 10 to one ratio here. Most tools that support this more agentic coding workflow like Cursor is still integrate heavily towards the more commercial models.
- 15:29
However, because of the volume of the use case and since the R1 wave, we have seen a rapid growth in using such tools with open models. This is part of the reason why there is no use for me presenting January data, because this coding segment itself went from basically under the 5% threshold in volume to easily now
- 15:49
20 to 30% of all traffic. The one major open source project that sticks out that I would like to highlight is Cline, which focuses on the chat agentic flow.
- 16:01
And when you combine it with the other open, uh, source project, Continuous Dev, for the auto-completion integration, you basically have the same level of experience as Cloud or OpenAI models with Cursor or any other, uh, IDE platforms with agentic workflow via R1.
- 16:17
Or we are possibly using your own private MacStudio cluster under the [REDACTED:location],
- 16:22
actually popular for, like, all these hobbyist users who actually want to do this completely without the internet.
- 16:29
Highly recommended that you g- uh, give these open projects a try, uh, in VS Code before they get forked a dozen times in the next YC.
- 16:37
Next up is ComfyUI and friends. About 5% of the traffic for personal agentic workflows. Well, many of you may be familiar with these two being widely used in the diffusion space for more complicated AI generation-- image generation workflows.
- 16:52
S- similar, uh, graph style based UIs are now being used, right, from reporters, lawyers, musicians, influencers, and even poets somehow, th-- where, where these people are chaining complicated workflows to generate the text that they want for their use case.
- 17:09
Similar to coding agentic workflow, the token explosion from a handful of these power user slowly stack up to noticeable levels. But these users are not developers. They are literally, like I said, musicians or, like, all these creative personnels who learn how to use these tools.
- 17:27
However, unlike Cline, which saw a sudden explosion due to R1, the past few weeks, this has been more about a slow and gradual user base build up in the past months.
- 17:40
It's also a space like most agentic platforms, uh, where it-- there's essentially what I call a framework war, if-- for those who are familiar with the JavaScript era.
- 17:51
The next major segment, RAG and Ch- ChatGPT clones, essentially. About 20% of all requests at this point. Well, every platform, be it Featherless or OpenRouter or every other AI provider will have their own internal ChatGPT UI.
- 18:05
Ours is called Phoenix. A standout application in this category that I would like to name is TypingMind. An indie application which, surprise, surprise, bills you with a one-time fee instead of a subscription, and can be used locally on your laptop against any API provider like us.
- 18:24
In particular, it's the level of UI polish and how they've been keeping pace with cloning every single ChatGPT UI feature into their app, and customizing that with even their own plugin system.
- 18:35
Beyond that, you probably heard this use case a thousand times. ChatGPT, RAG, yada, yada. So let's skip to the one that everyone wants to hear more about. AI for agents and work, representing 10 to 20% of all our traffic.
- 18:49
Because agent is such an over-abused term today, I'll split this into two major categories to make things clear. Workflow automation or narrow agents, and fully automated agents.
- 19:04
If I, uh, so right now, let's start from workflow automation or narrow agent. For most of you who's working on agents in enterprises, your number one priority now is to get all your agents into production by maximizing ROI for your company and shareholder value, if you want to use that term, without limiting the negative impact.
- 19:26
Especially since we are in New York, lots of financial institutes are risk-averse by default, with many of you potentially working in there developing your agents right now. So have that stance, build in human escape hatches by default.
- 19:41
What I mean by that is, basically, if you are building an automation system in your company,
- 19:47
like Waymo, make it so that the human can take the driver's seat needed from day one.
- 19:53
So for example, we work with a few insurance claim companies and logistic companies advising them on how to fully automate their email process, which is the bulk of their work.
- 20:03
A common pattern we advise them would be them to create an AI agent and UI for fully automating drafting responses for inbound emails, checking against inventories, checking against rules, checking a lot of other things in the process,
- 20:16
and build a platform UI for editing and finalizing and sending a response.
- 20:21
In best case scenario, if the AI did everything perfectly, it's just all about clicking send. And sometimes these things are built on top of their existing customers' Salesforce or ERP system.
- 20:33
However, more importantly, it at least have one human checking the final submissions before it hits the, the end user. When done right, at launch, the AI is able to successfully draft around 80, 90% of the response, with humans rejecting and manually taking over the remainder, so there's no harm to everyone else.
- 20:51
Productivity for the responders shoots through the roof. Management is happy. The AI is adopted.
- 20:58
And if any mistakes were made and the human failed to correct it, well, that's kind of the human's fault, not the AI.
- 21:05
Over time, as confidence start to build, they then start to add classification for certain known reliable use case, and eventually start fully automating those use cases.
- 21:17
It's like auto-reply to certain, certain scenarios. With full confidence after having seen the system run for thousands of customers.
- 21:26
On the flip side, teams that went full, fully ambitious into full 100% automation without human escape hatches,
- 21:35
sometimes they start with a successful launch because everyone wanted it to succeed. But more commonly, as they run the system, and as more users use it, they eventually get fully burnt by an angry customer for a really bad automated response down the line.
- 21:50
In some cases, this may be fully acceptable compromise that management accepts from the start because the savings benefit is worth it. Other times, this kills the entire AI project, essentially postponing AI adoption in your organization for another year, a scenario that you do not want to happen.
- 22:12
And this is commonly because these, a lot of teams end up developing an entire fully automated workflow without trying to move it into production early. That's why I say your goal is to go beyond the POC phase and into adoption in your company at scale in some portion as fast as possible, and to inter- integrate and iterate
- 22:36
from there. So that leaves the other mythical category where, which everyone loves to chase. Fully, truly 100% reliable agentic agents that is fully automated. Just don't. It does not exist.
- 22:51
Everyone has, who has done it, tried to done it, has got burnt.
- 22:56
And it's, we are not, and we are talking about things in production, not POC.
- 23:04
The mindset that we should be thinking when we try to build AIs into production
- 23:09
is that we should approach things from, like, trying to solve the 80% with escape hatches along the way.
- 23:18
Because at, even at 80%, for some of you mega corps here, that's already millions in productivity and saving.
- 23:27
Furthermore, do I need to remind you that even humans are not 100% reliable?
- 23:33
So only do this when the eventual mistake, right, is acceptable and can be fixed.
- 23:40
And it's a basically a trade-off that you're willing to take. So common example is cold calling systems when you're actually willing to lose a lead because you did, you wouldn't have gotten it otherwise anyway.
- 23:52
Or sometimes customer support because you can get a human to apologize later.
- 23:58
But rewinding back, it's like you should take the approach just like how Google software reliability engineers or any reliability engineering for industry where safety and reliability is paramount do.
- 24:09
Start with an extremely streamlined, reliable system for your 80 to 90% work scenarios. And after you're done, do it for another 80 to 90% for that 10 to 20% failure scenarios and repeat this process 10 times.
- 24:24
And now suddenly you have 99.998% reliability. If, if you're an airliner that's not Boeing, do it another 10 times and everyone is happy with incremental gains and pro-- and improvements.
- 24:37
It may not be the sexy first launch. It may not be the most wow to do it this way, but at the very least, your project survives to production, you have huge impact to ROI, and you can iterate from there.
- 24:53
And more importantly, incrementally improve and amusingly, probably never change some of the models that you use and forever give me business, which I'm ironically going to do and say the following.
- 25:06
Hey, Featherless customers, especially the more enterprising ones, please try to upgrade from Llama 2 to something newer this year or next. If done right, it's probably a free 99 to 99.9% improvement.
- 25:17
I know all of you are using this model and came to us in part because we are one of the last few providers to host it, and we forever will.
- 25:24
But hey, I'm giving this advice. It's free real estate.
- 25:28
Similarly, hey, for closed source labs, it's 2025. You do not need AGI to learn how to use Git version from AI models.
- 25:39
I say this in all seriousness because for companies in production who is trying to incrementally improve with each step to 99 plus 9999%, it's nearly impossible to do so if you change the model every week.
- 25:53
So with that, ah, shucks, I'm out of time.
- 25:58
All the best in building your AI agent.
- 26:01
Hey, hey, Eugene, come back here. You're forgetting the one more thing bit. Ah, crap, he can't hear me. The laptop speakers are muted, and I have no control over that.
- 26:11
Uh, you see, us AI may be flawed, but honestly, we are probably more consistent and reliable than humans are, as you can clearly see. Uh, oh, well, here goes me saving the day.
- 26:22
Hi, I'm Quirky, a seventy-two billion parameter linear transformer and attention transformer hybrid with a completely new architecture that runs at less than half the GPU's compute cost of other transformer models.
- 26:34
The strongest post-transformer hybrid to date. When Eugene proposed training a post-transformer model to provide an alternative to attention is all you need, which runs at a fraction of the inference cost, many of you said, "It can't be done.
- 26:47
The tech is unproven and will not scale." Well, DeepSeek cost ten million. Uh, Quirky cost only a one hundred thousand to build. You can find out more about our model at the following link.
- 27:00
More importantly, as we enter the era where the average AI model has better MMLU than the average office worker, the benchmark has lost its meaning. After all, how many of you audience can calculate orbital reentry from Earth to Mars or PhD math questions?
- 27:17
All of which are questions apparently frontier AI models are now capable of. How many of our AI agents today can be reliably be trusted with one task it's instructed to do so without hallucination or failure?
- 27:31
I think you get the point. What we find more exciting is exploring a future where we take advantage of linear transformer models as a means of persisting memories, customization, and improving reliability for future AI models to make useful AI agents.
- 27:47
And if you would like to run Featherless or Quirky in your private cloud or on-premise environment, reach out to us and we would like to hear more about your use case.