AI Engineer World's Fair 2025
Memory Masterclass: Make Your AI Agents Remember What They Do! — Mark Bain, AIUS
About this talk
Mark Bain leads a workshop on durable AI-agent memory, arguing for causal relationships, knowledge graphs, and GraphRAG. Guest presenters Vasilije Markovic of Cognee, Alex Gilmore of Neo4j, and Daniel Chalef of Zep/Graphiti demonstrate graph-based agent workflows, memory-server integration with Claude Desktop, and temporal graphs. Bain compares MCP integrations across Neo4j, Graphiti, Cognee, and Mem0 before introducing a GraphRAG chat arena and discussing agentic firewalls and audience questions.
Chapters
- 0:15Workshop introduction, speakers, and AI memory agenda
- 2:35Memory foundations, GraphRAG, and causal relationships
- 17:16Cognee demonstration: agents, GitHub repositories, and graph data
- 24:53Neo4j memory server and Claude Desktop
- 27:35Graphiti, Zep, temporal graphs, and MCP comparisons
- 40:19Agentic firewalls and the GraphRAG chat arena
- 46:19Audience questions and workshop closing
Talk transcript
- 0:00
[upbeat music] Woo!
- 0:15
I'm super excited to be here with you. Um, this is my first time speaking at AI Engineer and, um, we have an amazing, um, group of speakers, guest speakers.
- 0:28
Vasilije Markovic from Cognee. Vasilije, um... Ooh, there is Vasilije. Daniel Chalef from Graphiti and Zep AI, and Alex Gilmore from Neo4j. Um, the,
- 0:44
the plan looks like this. I will do a very quick power talk and about, about the topic that I'm super passionate, um, the AI memory. Next, we'll have four live demos, uh, and we'll move on to some new solution that we are proposing, a GraphRAG chat arena, uh, that I will be able to demonstrate
- 1:08
and I would like you to follow along once it's being, um, demonstrated. And at the very end, uh, we'll have a very short Q&A session. Um,
- 1:22
there is, um, a Slack channel that I would like you to join. Um, so please scan the QR code right now before we begin, and let's make sure that everyone has access to the, um, to these materials.
- 1:37
There is, um, a walkthrough short on the channel that will go through
- 1:45
closer to the end of our workshop, but I would like you to start setting it up if you, if you may, on your laptops if you want to follow along.
- 2:01
All right, it's, uh, workshop-graphraggchat. You can also find it on Slack and you can, uh, join the channel. So a little bit about myself. Uh, so hi everyone again.
- 2:16
I'm Mark Bain, and I'm very passionate about the memory, what is memory, the deep physics and applications of memory across different technologies. Um, you can find me at [REDACTED:username], uh, on social media or on my website.
- 2:35
And let me tell you a little bit of a story about myself. Uh, so when I was, um, [REDACTED:age], I was very good at maths and I did math Olympiads with many brilliant minds, including, uh, Wojciech Zaremba, the co-founder of OpenAI.
- 2:54
And thanks to that deep understanding of maths and physics, I did have many great opportunities to be exposed to the problem of AI memory. So first of all, I would like to recall, um, two conversations that I had with Wojciech and Ilya in 2014 in September.
- 3:20
When I came here to study at Stanford, um, at one party we met with Ilya and Wojciech, who back then worked at Google, and they were kind of trying to pitch me that there will be a huge revolution in AI.
- 3:37
And I kind of, like, followed that. I was a little bit unimpressed back then. Right now I probably, um,
- 3:45
kind of take it as a very big excitement when I look back to the times and I was really wishing good luck to, to the guys who were doing deep learning because back then I, I didn't really see this prospect of, uh, GPUs giving that huge edge, uh, in compute.
- 4:04
Uh, however, uh, during that conversation, it was like 20 minutes, at the very end, I asked Ilya, "All right, so there is going to be a big AI revolution, but how will these AI systems communicate with each other?"
- 4:23
And the answer was very perplexing and kind of sets the stage to what's happening right now. Uh, Ilya simply answered, "I don't know. I think they will invent their own language."
- 4:37
So that was 11 years ago. Fast-forward to now, um, the last two years I've spent doing very deep research on physics of AI and kind of like delve into all of these most modern AI architectures, incude- including attention, diffusion models, VAEs, and many other ones.
- 4:58
And I realized that there is something critical, something missing,
- 5:06
and this power talk is about this missing thing.
- 5:13
So over the last two years, I kind of followed on on my last years of doing a lot of research in physics, computer science, information science, and I came to this conclusion that memory, AI memory, in fact, is any data in any format, and this is important, including code,
- 5:39
algorithms and hardware, and any causal changes that affect them.
- 5:47
That was something very mind-blowing to, to reach that conclusion, and that conclusion sets the tone to this whole track, the GraphRAG track.
- 6:01
In fact, I was also perplexed by how biological systems use memory and how different cosmological structures or quantum structures, they in fact have a memory. They kind of remember.
- 6:17
And let's get back to maths and to physics and geometry. When I was doing science olympiads, I was really focused on two, three things: geometry, trigonometry, and algebra.
- 6:33
And I realized in the last year that more or less the volume of
- 6:43
loss in physics perfectly matches the volume of loss in mathematics. And also the constants in mathematics, if you really think deeply through geometry, they match the constants both in mathematics and in physics, and if you really think even deeper, they kind of like transcend over the, all the other disciplines.
- 7:08
So that made me think a lot, and I found out
- 7:13
that the principles that govern LLMs are the exact same principles that govern neuroscience,
- 7:23
and they are the exact same principles that govern mathematics. I studied,
- 7:29
I studied papers of Perelman. I don't know if you've heard who is Perelman. Perelman is this mathematician who refused, um, to take a one million dollar award for proving the,
- 7:46
um, for, for proving the, one of the mo- mo- most important conjectures, mm, about symmetries of three spheres. Um,
- 8:01
and once I realized that this deep math of
- 8:10
spheres and circles is very much linked with how attention and diffusion models work, basically the formulas that Perelman reached are linking
- 8:26
entropy with curvature. And curvature, basically, if you think of curvature, it's attention, it's gravity. So in a sense, there are multiple disciplines where the same things are appearing multiple times, [tongue click]
- 8:42
and I will be publishing a series of papers with some amazing supervisors who are co-authors of two of these,
- 8:56
uh, method, methods, methodologies, um, the transformers and VAEs.
- 9:03
And I came to this realization that this equation governs everything, governs maths, governs physics, governs our AI memory, governs neuroscience, biology, physics, chemistry, and so on and so forth.
- 9:20
So, [tongue click] I came to this equation that memory times compute would like to be a squared imaginary unit circle.
- 9:35
If that existed ever, we would have perfect symmetries, and we would kind of not exist because for us to exist, these asymmetries needs to show up. And in a sense, every single LLM, through weights and biases, the weights are giving the structure, the compute that comes and transforms the data in sort of the raw format,
- 10:00
the compute turns it into weights. The weights are basically, if you take these billions of parameters, the weights are the sort of like matrix structure of how this data looks like, uh, when, when you really find relationships in the raw data.
- 10:17
All right. And then there are these biases, these tiny shifts that are kind of like trying to like in a robust way adapt this model so that it doesn't break apart, but still is, still is very well reflecting the reality.
- 10:35
So something is missing. So when we take weights and biases and we apply scaling clause and we keep adding more data, more compute, we kind of get a better and better and better understanding of the reality.
- 10:47
In a sense, if we had infinite data, we wouldn't have any biases. And this understanding is, again, the principle of this RAG, of GraphRAG.
- 11:02
The disappearance of biases is what we are looking for when we are scaling our models.
- 11:09
So in a sense, the amount of memory and compute should be exactly the same. It's just slightly expressed in a different way. But if there are some, uh, there are in- any imbalances,
- 11:23
then something important happens. [tongue click] And I came to another conclusion that our universe is basically a network database.
- 11:34
It has a graph structure, and it's a temporal structure, so it keeps on moving, following some certain principles and rules.
- 11:44
And these principles and rules are not necessarily fuzzy. They have to be fuzzy
- 11:54
because otherwise everything would be completely predictable. But if it would be completely predictable It means that me myself would know everything about every single of you, about myself from the past and myself from the future.
- 12:10
So in a sense, it's impossible, and that's why we have this sort of like heat diffusion entropy models. They allow us to exist, but something is preserved.
- 12:27
Relationships. Any single asymmetry that happens at the quantum level,
- 12:36
any single tiny asymmetry that happens preserves causal links.
- 12:44
And these causal links are the exact thing that I would like you to have as a takeaway from this workshop.
- 12:55
The difference between simple RAG, hybrid RAG, any types of RAG, and graph RAG, is that we are having the ability
- 13:06
to keep these causal links in our memory systems. Basically, the relationships are what preserves causality. That's why
- 13:19
we can solve hallucinations. That's why we can optimize
- 13:28
hypothesis generation and testing. So we will be able to do amazing research in biosciences, chemical sciences, just because of understanding that this causality is preserved within the relationships.
- 13:46
And these relationships, when there are these asymmetries that are needed, they kind of create this curvature, I would say. So we under... We intuitively feel every single of you is choosing some specific workshops and talks that you guys go to.
- 14:06
Right now, all of you are attending to the talk and workshop that we are giving. It means that it matters to you, and it means that potentially you see value, and this value, this information, is transcended through space and time.
- 14:28
It's very subjective to you or any other object,
- 14:32
and I think we really need to understand this. So LLMs are basically these weights and biases or correlations. They give us this opportunity to be fuzzy.
- 14:47
You know? A- actually, one thing that I learned from Wojciech 10, 8, 11 years ago,
- 14:55
was that hallucinations are the exact necessary thing to be able to solve a problem where you have too little memory or too little compute, compute for the combinatorial space of the problem you are solving.
- 15:08
So you're basically imagining. We are taking some hypothesis basing... based on your history, and we are kind of trying to project it into the future. But you have too little memory, too little compute to do that, so it can be as good as the amount of memory and compute you have.
- 15:24
So it means that the missing part is something that you kind of can curve thanks to all of these causal relationships and this fuzziness. And... Oops.
- 15:38
Reasoning is reading of these asymmetries and the causal links.
- 15:47
Hence, I really believe that agentic systems are
- 15:55
sort of the next big thing right now because they are following the network database principle.
- 16:05
But to be causal, to recover this causality from our fuzziness, we need graph databases. We need causal relationships. And that's the major thing
- 16:19
in this emerging trend of GraphRAG that we are here to talk about. And I would like to, at this moment, invite on stage our three amazing guest speakers, and I would like to start with Vasilije.
- 16:38
Vasilije, please come over to, to the stage. Next will be Alex and Daniel, and I will present something myself. All right. So, uh, Vasilije will show us how to lurch a search and optimize memory based on certain use case at hand.
- 16:59
All right. [audience applauding]
- 17:03
Test, test. Um, so let's just make sure this works.
- 17:16
Okay. So, um, nice to meet you all. Uh, and I'm Vasilije. I'm originally from Montenegro, a small country in the [REDACTED:location]. Um, beautiful, so if you wanna go there, my cousins Igor and Milos are gonna welcome you.
- 17:34
Everyone knows everyone So, uh, you know, if in case you're just curious about memory, I'm building a memory tool on top of the graph and, uh, vector databases. My background's in business, big data engineering, and clinical psychology, so a lot what Mark, uh, talked about kind of connects to that.
- 17:53
Um, I'm gonna show you a small demo here. Uh, the demo is to do a Mexican standoff between two developers where we are analyzing their GitHub repositories and these, uh, data from the GitHub repositories is in the graph, and, um, this Mexican standoff means that we will, um, let the crew of agents go analyze, look at their
- 18:12
data, and try to compare them against each other and give us a result that should represent how, uh, who should we hire, let's say, uh, ideally out of th- these two people.
- 18:21
So, uh, what we are seeing here currently is how Cognify works in the background. So Cognify is working by, uh, adding some data, turning that into a semantic graph, and then we can search it with wide variety of options.
- 18:33
We plugged in CrewAI on top of it, so we can pretty much do this on the fly. So, um, here in the background, I have a client running. This client is connected to the, to the system.
- 18:43
So, um, it's now currently, uh, searching the datasets and, uh, starting to build a graph. So let's, uh, see. It takes a couple of seconds, but, uh, in the background, uh, we are effectively ingesting the GitHub, uh, data from the GitHub API, building the semantic structure, and then, uh, letting the agents actually search it and, and make
- 19:03
decisions on top of it. So, uh, as every time with live demos, uh, things might go wrong, so I have a video version in case this does. Let's see.
- 19:15
And I'll switch to the vi-- Oh, here we go. So, um, the semantic graph started generating, and as you can see, we have activity log where the graph is being continuously updated on the fly.
- 19:27
Data's being stored in memory, and then, uh, data's being enriched, and the agents are going and making decisions on top. So what you can see here on the side is effectively the agentic logic that is reading, writing, analyzing, and using all of this, uh, let's say pre-configured, uh, set of weights and benchmarks to, to analyze any, uh,
- 19:48
person here. So Cogni is a framework that's modular. You can build these tasks. You can ingest from any type of a data source, thirty plus data sources supported now.
- 19:55
You can build any type of a custom graph. You can build graph from relational databases, semi-structured data, and we also have these memory association layers inspired by the cognitive science approach.
- 20:05
And then effectively, um, as we kind of build and, and enrich this graph on the fly, we see that, uh, you know, it's getting bigger, it's getting, uh, more popular, and then we are storing the data back into the graph.
- 20:16
So this is the, uh, stateful temporal aspect of it. We kind of build the graph in a way that we can add the data back, that we can analyze these reports, that we can search them, and that we can let other agents access them on the fly.
- 20:29
The idea for us was let's have a place where agents can write and continuously add the data in. So, um, I'll have a look at the graph now so we can inspect it a bit.
- 20:39
So if we click on, on any node, we can, uh, see that, uh, the details about the commits, about the information from the, from the developers, the PRs, whatever they did in the past and, and which repos they contributed to.
- 20:52
And then at the end, as the graph is pretty much, uh, filled, we would see the final report kind of starting to come in. So let's see how far we got with this.
- 21:01
So it's taking-- It's preparing now the final output for the hiring decision task, so let's have a look at that when it gets loaded.
- 21:12
We just finished this this morning. I hoped to have a hosted version for you all today, but didn't work. CrewAI is, uh, causing some troubles. So, uh, let's, uh, we have to resolve this one.
- 21:23
So let's see. Yes. So I will just
- 21:34
show you the video with the end so we don't wait for it.
- 21:39
So here you can see that towards the end, um, we can see the graph and
- 21:49
we can see the final decision, which is a green node. And in the green node, we can see that we decided to hire Laszlo, our, uh, developer who has a PhD in graphs, so it's not really difficult to make that call.
- 22:03
And we see why, and we see the, the numbers and the benchmarks. So thank you. This has been very fast three-minute demo, so hope you enjoyed. And if you have some questions, I'm here afterwards.
- 22:12
We ha-- We are open source, so happy to see new users and if you're interested, try it. Thanks.
- 22:17
Woo-hoo. [audience applauding] Thank you. Thank you, Vasilije. Um, next up is Alex. So Vasilije, uh, showed us something I call semantic memory. So basically you take your raw data, you, uh, load it and Cognify it, as they like to say.
- 22:36
Come on, come on up, Alex. And that's the base. That's something we already are doing. And next up is Alex, who will show us Neo4j MCP, uh, server.
- 22:52
The stage is yours.
- 22:59
Test, test. Test, test. Five, four, three, two, one. We good? Okay.
- 23:09
All right. Okay. So hi, everyone. My name's Alex. Um, I'm an AI architect at Neo4j. Um, I'm gonna demo the memory MCP server that we have available. Um, so there is this walkthrough document that I have.
- 23:26
Um, we'll make this available in the Slack or by some means so that you can do this on your own. Um, but it's pretty simple to set up. Um, and what we're gonna showcase today is really, like, the foundational functionality that we would like to see in a agentic memory sort of application.
- 23:40
Um, primarily, we're gonna take a look at semantic memory in this, um, MCP server, but we are currently developing it, and we're gonna add additional memory types as well, um, which we'll discuss, uh, probably later on in the presentation.
- 23:52
Um, so in order to do this, we will need a Neo4j database. Neo4j is a graph-native database that we'll be using to store our knowledge graph that we're creating.
- 24:02
Um, they have a Aura option, which is, um, hosted in the cloud, or we can just do this locally with the Neo4j Desktop app. Um, additionally, we're gonna do this via Claude Desktop, and so we just need to download that, and then we can just add this config to the, um, MCP configuration file in Claude.
- 24:22
And this will just connect to the Neo4j instance that you create. Um, and what's happening here is we're going to-- uh, Claude will pull down, um, the memory server from PyPi, and it will host it in the backend for us, and then it'll be able to use the tools that are accessible via the MCP server.
- 24:39
And the final thing that we're gonna do before we can actually have the conversation is we're just gonna use this brief system prompt, and what this does is just ensure that we are properly recalling and then logging memories after each interaction that we have.
- 24:53
Uh, so with that, um, we can take a look at a conversation that I had, um, in Claude Desktop using this memory server. Um, and so this is a conversation about starting an agentic AI memory company.
- 25:06
Um, and so we can see, um, all these tool calls here. And so initially, we have nothing in our memory store, which is as expected. But as we kind of progress through this conversation, we can see that at each interaction, it tries to recall memories that are related to the user prompt.
- 25:24
And then at the end of this interaction, it will create new entities in our knowledge graph, um, and relationships. And so in this case, an entity is going to have a name, a type, and then a list of observations, and these are just facts that we know about this entity, and this is what is going to be
- 25:41
updated, um, as we learn more. In terms of the relationships, these are just identifying how these re-- uh, how these entities relate to one another. And this is really the core piece of why using a, uh, graph database as k- sort of the context layer here is so important because we can then-- we can identify how these,
- 26:00
um, entities are actually related to each other. It provides a very rich context. And so as this goes on, we can see that we have quite a few interactions.
- 26:10
We are adding observations, creating more entities. And at the very end here, we can see we have quite a lengthy conversation, and we can say, you know, "Let's review what we have so far."
- 26:21
And so we can read the entire knowledge graph back as context, and Claude can then summarize that for us. And so we have all of the entities we've found, all the relationships that we've identified, and all the facts that we know about these entities based on our conversation.
- 26:35
And so this provides a nice review of what was discussed about this company and our ideas about how to create it. Now, we can also go into Neo4j Browser, um, and this is available both in Aura and local, and we can actually visualize this knowledge graph.
- 26:49
And we can see that we discussed Neo4j, we discussed MCP and LangGraph, and if we click on one of these nodes, we can see that there is a list of observations that we have.
- 26:59
And this is all the information that we've tracked throughout that conversation. And so it's important to know that, like, even though this knowledge graph was created with a single conversation, we can also take this and use it in additional conversations.
- 27:10
We can use this knowledge graph with other ID-- um, uh, clients, such as Cursor IDE or Windsurf. And so this is really a powerful way to, um, create a, like, memory layer for all of your applications.
- 27:25
Um, and so with that, um, I'll pass it on. [chuckles] Thank you.
- 27:29
All right. Give a round of applause to Alex. [clapping]
- 27:35
Thank you, Alex. The next up is Daniel. Um, I-I will just share personal, um, beliefs about MCPs. Um, I was testing MCPs of Neo4j, Graphiti, Cogni, Mem0 just before the workshop.
- 27:51
And I'm a strong believer that this is our future. We'll have to work on that. And in a second, I will be showing a mini GraphRAG chat arena. And next up, something very, very important that Daniel does is temporal graphs.
- 28:06
Daniel, uh, is co-founder of Graphiti and Zep. They have ten thousand stars on GitHub and growing very fast. The stage is yours, Daniel. Please show us what you do.
- 28:16
Thank you. So five, four, three, two, one. Did that work? Seems to have, right? So, um,
- 28:29
I'm here today to tell you that there's w- no one size fits all memory, um, and why you need to model your memory after your business domain. So if you saw me a little bit earlier and I was talking about Graphiti, Zep's open source temporal graph framework,
- 28:50
um, you might have seen me just speak to how you can build custom entities and edges in the Graphiti graph for your particular business domain. So business objects from your business domain.
- 29:05
What I'm gonna demo today is actually how Zep implements that and how e-easy it is to use from Python, TypeScript, or Go. And what we've done here is we've solved a fundamental problem plaguing memory, and we're enabling developers to build out memory that is far more cogent and capable for many
- 29:30
different use cases. So I'm gonna just show you a quick example of
- 29:36
where things go really wrong. So many of you might have used ChatGPT before. It generates facts about you in memory, and you might have noticed that it really struggles with relevance.
- 29:48
Sometimes it just pulls out all sorts of arbitrary facts about you. And unfortunately, when you store arbitrary facts and retrieve them as memory, you get inaccurate responses or hallucinations.
- 30:02
And the same problem happens when you're building your own agents. So here we go. We have an ex-example media assistant, and it should remember things about jazz music, NPR podcasts, The Daily, et cetera, all the things that I like to listen to.
- 30:18
But unfortunately, because I'm in conversation with the agent or it's picking up my voice when I'm, you know, it's a voice agent, um, it's learning all sorts of irrelevant things, like I wake up at seven AM, my dog's name is Melody, et cetera.
- 30:32
And the point here is that irrelevant facts pollute memory. They're not specific to the media player business domain. And so the technical reality here is as well that many frameworks take this really simplistic approach, approach to generating facts.
- 30:52
If you're using a framework that has memory capabilities, agent framework, it's generating facts and throwing it into a vector database. And unfortunately, the facts dumped into the vector database or Redis mean that when you're recalling that memory, it's difficult to differentiate what should be returned.
- 31:08
We're gonna return what is semantically similar. And here we have, um, a bunch of facts that are semantically similar to my request for my favorite tunes. Um, we have some good things, and unfortunately, Melody is there as well because Melody is a dog named Melody, and that might be something to do with tunes.
- 31:30
Um, and so bl-bunch of irrelevant stuff. So basically, semantic similarity is not business relevance,
- 31:43
and this is not un-unexpected. I was speaking a little bit earlier about how vectors and are just basically projections into an embedding space. There's no causal or relational, uh, relations between them.
- 31:59
And so we need a solution. We need domain-aware memory, not better semantic search.
- 32:06
So with that, I am going to unfortunately be showing you a video because the Wi-Fi has been absolutely terrible. Um. [laughs]
- 32:18
And let me bring up the video. Okay. So
- 32:27
I built a little application here, and it is a finance coach. And I've told it I wanna buy a house.
- 32:37
And it's asking me, well, how much do I earn a year? It's asking me about what student loan debt I might have. And we'll see that on the right-hand side, what is stored in Zep's memory are some very explicit co-business objects.
- 32:59
We have financial goals, debts, income sources, et cetera. These are defined by the developer, and they're defined in a way which is really simple to understand. We can use Pydantic or Zod or Go Structs, and we can apply business rules.
- 33:20
So let's go take a look at some of the code here. We have a TypeScript financial goal schema using Zep's underlying SDK. We can define these entity types. We can give a description to the entity type.
- 33:34
Uh, we can even define fields, the business rules for those fields, so the values that they take on.
- 33:41
And then we can bu-build tools for our agent to retrieve a financial snapshot which runs multiple Zep searches at the same time concurrently and filters by specific node types.
- 33:58
And when we start our Zep application, what we're gonna do is we're gonna register these particular goals, uh, sorry, objects, with, uh, Zep, so it knows to build this ontology in the graph.
- 34:12
So let's do a quick little addition here.
- 34:18
I'm gonna say that I have five thousand dollar a month rent.
- 34:23
I think it's rent. And in a few seconds, we see that Zep's already parsed that new message and has captured that five thousand dollars. And we can go look at the chart, the graph.
- 34:35
This is the, the Zep front end. And we can see the knowledge graph for this user has got a debt account entity. It's got fields on it, um, that we've defined as a developer.
- 34:50
And so again, we can really get really tight about what we retrieve from Zep by filtering. Okay. So we're at time. So just very quickly, we wrote a paper about how this, all of this works.
- 35:01
You can get to it, uh, by that link below and appreciate your time today.
- 35:08
You can look me up afterwards. [clapping]
- 35:11
Great paper, by the way. All right, so whilst I'm getting ready, um, I would appreciate if you confirm with me, uh, whether you have access to Slack. Uh, is the Slack working for you, the Slack channel?
- 35:24
All right. I think we are slowly running out of time, so I'd appreciate if you have any questions to any of the speakers, please, uh, write these questions on Slack, and we will be outside of this room, and we are happy to answer more of these questions just after the workshop.
- 35:41
I right now move on with, um, a use case that I developed and to this GraphRAG, uh, chat arena. Mm.
- 35:51
To be specific, before delving into agentic memory, into knowledge graphs
- 36:03
I led a private cybersecurity lab and worked for defense clients, very big clients with very serious problems on the security side. And I used to...
- 36:18
In one project, I had to navigate between something like 27, 29 different terminals and shells. And it requires knowing lots of languages.
- 36:33
Like, if you think of, like, different Linux distros, every firewall and networking devices usually has its own shell, proprietary often. There is PowerShell. So you need to know, like, lots of languages to communicate with these machines to, to work with such clients.
- 36:47
And I realized that LLMs are not only amazing to translate these languages, but they are also very good to kind of create a new type of shell, a human language shell.
- 36:59
There are such shells. But such shells, they would really be excellent if they have episodic memory, the sort of temporal memory of what was happening in this shell historically.
- 37:14
And if we have access to this temporal history, the events, we kind of know what the users were doing, what their behaviors are. We kind of can control every single code execution function that's running, including the ones of agents.
- 37:30
So I spotted with some investors and advisors of mine, I spotted a niche,
- 37:37
something we call agentic firewall, and I wanted to do a super quick demo of how it would work.
- 37:44
So basically you would, um, run commands and type PWD, and in a sense we... I suppose lots of us had computer science classes or, or we worked in shell, and we have to remember all of these commands.
- 38:01
Like, um, show me running Docker containers. Like, it's Docker PS, right? But if you go for more advanced commands-
- 38:12
We cannot see your screen
- 38:13
Uh, I think it's for a reason. Yeah, I think it's for a reason. Um, oh, it's... One sec.
- 38:22
Sorry about that.
- 38:32
Yeah, there.
- 38:32
All right. It's there. Okay. Thank you. In general, I would need to know right now some command that can extract me, for instance, the name of the container that's running and its status.
- 38:48
Show me just, um, image and status. I can make mistakes, like human language, fuzzy mistakes. Um,
- 39:00
show if Apache is running. All right. Show the command
- 39:10
we did three commands ago. So basically, if you plug in the agentic...
- 39:21
If you plug in the agentic memory to things like that, I thi- I think it got it wrong, but you, you get me, right? So if I get through, like, different shells and terminals, um, and I have this textual context that what was done and the context of the certain machine of what is happening here,
- 39:41
mm, and it kind of spans across all the user... All the machines, all the users, and all the sessions in PTYs, TTYs. I think that we can really have a very good context also for security.
- 39:57
So that space, um, the temporal logs, the episodic logs, is something that I see will boom and emerge. So I believe that all of our agents that will be executing codes in terminals will be executing it through a l- maybe not all, but the ones that are running, uh, on the enterprise gate.
- 40:19
They will be going through agentic firewalls. I'm, I'm close to sure about that. So that's my use case. Um, and now let's move on to GraphRAG chat arena. So you have on Slack, uh, a link to this doc, and this doc is allowing you to set up a repo that we've created for this workshop, and we'll be
- 40:43
promoting it afterwards. So about a year ago, I met with, uh, Jerry Liu from LlamaIndex, and we were chatting quite a while about, like, how to evolve this conversational memory, and he gave me two pieces of advice.
- 40:57
One of them, think about data abstractions, the other, think about evals. Data abstractions, I kind of quickly solved within, like, two months. Evals, I realized that there won't be any evals in form of a benchmark.
- 41:10
This... All of these hot potatoes and all of that, it's fun. I know that there are great papers written by our guest, guest speakers and other folks about hot potatoes.
- 41:19
But it's not the thing. You can't do a benchmark for a thing that doesn't exist. Basically, the agentic GraphRAG memory will be this type of memory that evolves, so you don't know what will evolve.
- 41:33
So if you don't know what will evolve, you will need a simulation arena, and that will be the only right
- 41:40
eval. So one year fast-forward, and we've created a prototype of such agentic memory arena. Think about it like Web Arena, but for memory. And let me quickly show you that.
- 41:54
You can go to this repository. I did a fork of that. There is Mem0, there is Graphiti, there is Cognee. Um, and there will be two approaches. One approach will be, um-
- 42:08
Sort of the repo, the, the library itself and the other is through MCPs, because we don't really know what will work out better. So whether repos or the MCPs will work out better.
- 42:17
So we'll need to test these different approaches. But we need to create this arena for that. So we basically clon- cloned that repo, and we use ADK for that.
- 42:28
So we get this nice chat where you can talk to these agents, and you can switch between agents. So I want to talk with Neo, and there is a Neo4j agent running behind the scenes.
- 42:43
There is a Cypher graph agent running behind the scenes, and they can, kind of for now, switch between these agents. Maybe I'll increase the font size a little bit.
- 42:53
So the Neo agent's basically answering the questions about this amazing technology, the graphs, specifically Neo4j.
- 43:02
And I can switch to Cypher, and then an agent that is excellent at running Cypher queries talks with me. And I'm writing, "Add to graph that I'm Mark, and I'm passionate about memory architectures."
- 43:16
And basically what it does is it runs these layers that are created by Cogni, by Mem0, by Graphiti, and all the other vendors of semantic and temporal memory solutions,
- 43:31
or specifically created by an MCP server that Alex was demonstrating, the Neo4j MCP server. So I'm really looking forward to how this, uh, technology evolves. But what I really, but what I quickly wanted to show you is that it already works.
- 43:50
It has this signs of being this agentic memory arena. So I can ask my graph through questions, and the agent goes to the connection. This is just one... You know what's amazing?
- 44:04
It's just one Neo4j graph. It's just one Neo4j graph on the backend, and all of these technologies that can be tested, how the graphs are being created and retrieved.
- 44:16
It's, it's like, when I think of that, it's like the most brilliant idea that we can do with agentic memory simulations. So I get answers from the graph. Here is the graph.
- 44:29
I can basically rerun, uh, the commands to see what's happening on this graph. And let me just move on.
- 44:38
And next thing is I would like to add to the graph that Vasilije will show how to integrate Cogni and da, da, da. So I add new information, and the Cypher writes it to the graph.
- 44:51
And then I want to do something else. It's, it's super early stage still, but then I transfer to Graphiti, and I can repeat the exact same process. So I can right now, using Graphiti, search what I just added, and I can switch between these different memory solutions.
- 45:08
So that's why I'm so excited about that. And we do not have time to, like, practice it together, do the workshop, but I'm sure we'll write some articles, so please follow us.
- 45:19
And I would appreciate, if you have any questions, pass them on to Slack. I, I will ask Andreas whether we have time for a short Q&A or we need to move it to, to, like, breakout or outside of the room.
- 45:34
We can take, like, five minutes.
- 45:35
Five minutes. All right. So, um, that's all for, for now, for today. I, I really, uh, would like, um, Vasilije, Daniel, and Alex to come back to stage so we can ask any of us.
- 45:48
Please, uh, direct the questions to, to any of us, and we'll try to, uh, answer them. Yeah, let's go.
- 45:55
Hi. I'm Lucas. Um-
- 45:57
Hi, Lucas
- 45:57
... I wanna ask a, a fundamental question. How do you decide what is a bad memory over time? Uh, because you, you could, like, as a developer and as a person, we evolve the, the line of thought, right?
- 46:11
So one thing that you thought was good, like, three years, 10 years ago may not be good right today. Uh, so how do you decide?
- 46:19
Sure. A very good question. So, um, I, I, I'll answer in... Maybe you guys can help. I will answer in a very scientific way. So basically, the one that causes a lot of noise.
- 46:30
The noisy one doesn't make a lot of sense. So you decrease noise by redundancy and by relationships. So the less relationships and the more noisiness, the... So, so in a sense, and not con- not, not well-connected node has a potential of not being correct, but there are other ways to validate that.
- 46:53
And would you like to, uh, follow?
- 46:56
Yeah, sure. Uh, a practical way, um, we, we let you model the data with PyDantic so you can kind of load the data you need and add weights, uh, to the edges and nodes.
- 47:06
So you can do something like temporal weighting. You can add your custom, let's say, logic, and then effectively would know how your data is kinda evolving in, in time and, and how it's becoming less or more relevant and what is the set of, uh, algorithms you would need to apply.
- 47:19
So this is the idea. Not solve it for you, but let help you solve it with tooling. Um, but yeah, there is-- Depends on the use case, I would say.
- 47:26
Would-
- 47:27
Yeah. No, I have nothing to add. I think that's a great explanation.
- 47:29
I, I think I'd-- what I would add is that there is missing causal, causal links. Missing causal links is what is most probably a good indicator of fuzziness. Yeah, next question.
- 47:42
All right. Um, can you hear me? How would you bet- embed in, um, security or privacy into the network or the application layer if there's a corporate, they have top secret data, or I have personal data that is a graph, I wanna share that, but not all of it?
- 47:59
Oh, that's a, that's a really good one. I, I think I'll answer that, um, very briefly. So basically, you do have to have that context. You do have to have that, these decisions, intentions of colonels, of majors, and anyone, like, in the enterpr- like CSOS and anyone's in, in the enterprise stack.
- 48:17
And in a sense, it also gets kind of like fuzzy and complex, so I expect this to be a very big challenge. That's, that's why I want to work in that.
- 48:25
But I'm sure that applying ontologies, the right ontologies first of all to this enterprise cybersecurity stack, really kind of provides these guard la- guardrails for navigating this challenging problem and, and decreasing these fuzziness and errors.
- 48:41
Thank you.
- 48:42
Yeah. I would also just add, like, all these applications are built on Neo4j, and so in Neo4j you can, like, do role-based access controls, and so you can prevent users from accessing data that they're not allowed to see.
- 48:55
So it's something that you can configure with that.
- 48:58
And one more thing.
- 48:58
Hi. This question is for Mark. Yeah.
- 49:01
Uh, yeah. Go on, go on, go on.
- 49:03
You were about to say something. Please go ahead first.
- 49:05
Yeah, just one thing. Like, we also noticed that if you isolate per graph per user or kinda keep it, like, very physically separate, for us it works, really works well.
- 49:13
People react to that really well. So that's one way.
- 49:15
Yes, independent graphs, personal graphs.
- 49:19
Yeah, Mark, in your earlier presentation you mentioned in this equation that related gravity, entropy, and something, and also memory and compute-
- 49:27
Yes
- 49:27
... to I square.
- 49:28
Yes.
- 49:28
Could you show those two again and explain them again?
- 49:30
Of course, yeah, uh, if, if we have time. Other than that, um, it's probably for a series of papers to properly explain that. So that's one, memory times compute equals I square.
- 49:41
The other one is that if you take all the attention diffusion and VAs which are doing the smoothing, it preserves the sort of asymmetries.
- 49:49
So very briefly speaking, let's set up the vocabulary. So first of all, curvature equals attention equals gravity. This is the very simple, most important principle here. I, I will need to when writing these papers, we are really tightly trying to define these three.
- 50:05
Next, diffusion, heat, entropy. It's the exact same thing. We just need to align definitions, and if it's not exact same thing, if there are other definitions, we need to show what's really different.
- 50:16
And now, if you think about attention, it kind of shows the sort of like pathways towards certain asymmetries. If you take a sphere, if you start bending that sphere and make it like, you know, like, you, you kind of try to extend it, two things happen.
- 50:33
Entropy increases and curvature increases, in a sense. And, and Perelman, what he did, he proved that you can, like, bend these spheres in any way, 3D spheres. 4D and 5D and higher le- like level spheres were already solved, so he solved for 3D sphere.
- 50:49
And these equations are proving that basically there won't be any other architectures for LLMs. It will be just attention diffusion models and VAs. Like s- maybe not just VAs, but like kind of like something that smooths, uh, leaves room for biases.
- 51:05
All right. Thank you all. Uh, I really appreciate you coming. I hope it was helpful. Thank you, the guest speakers, and we'll answer the questions, uh, outside of the room.
- 51:15
Appreciate that. [clapping] [outro music]