agent-memory
Every LLM app I've built runs into the same wall eventually: the model has no memory beyond whatever fits in the current context window, and once a conversation gets long enough, or a user comes back tomorrow expecting the app to remember them, you need somewhere to actually store what happened.
Every LLM app I've built runs into the same wall eventually: the model has no memory beyond whatever fits in the current context window, and once a conversation gets long enough, or a user comes back tomorrow expecting the app to remember them, you need somewhere to actually store what happened. Not a vector database with an ops team behind it. Just a place to put "the user said X" and get it back later, in whatever shape the conversation needs it in: a sliding window of recent turns, or a searchable pile of things said further back.
agent-memory is that place. Three interchangeable storage backends behind one interface: a plain in-memory array for quick scripts and tests where persistence doesn't matter, a JSON file for small projects that just need to survive a restart, and SQLite via better-sqlite3 for anything that needs to hold real conversation history without falling over. Swap between them by changing one constructor call. Memory.sqlite(path) or Memory.file(path), same API either way.
What the API actually looks like
The high-level Memory class wraps whichever backend you picked and gives you remember() to store something, getContext() to pull recent history back out in OpenAI's {role, content}[] message format (ready to drop straight into a chat completion call) and search() to look further back than the sliding window covers. That last part matters more than it sounds: most memory needs aren't "give me the last ten messages," they're "did this user ever mention their preferred deployment region," which requires actually searching, not just windowing.
The SQLite backend is written to be optional in practice even though it's declared as a dependency. getDb() lazy-requires better-sqlite3 only when you actually instantiate the SQLite backend, and throws a friendly, specific error with an install hint if the native module isn't there. That way a project that only ever uses the in-memory or file backend never pays the cost of a native dependency it doesn't need. Small detail, but it's the difference between "this package quietly requires a native compile step for everyone" and "this package requires a native compile step only for the people who chose the backend that needs it."
The gap between the README and what's actually built
I have to be straightforward about this one, because it's a real gap and not a small one. The README describes a five-type memory taxonomy (fact, preference, context, episodic, semantic) like the library has some kind of structured classification system for what it stores. It doesn't. What actually exists is role, timestamp, and search filtering on flat records. The taxonomy language is aspirational; I was thinking about how memory should eventually be organized while writing the docs, and wrote that vision down before building it, the same pattern that shows up in next-on-windows's README too. There's also a Redis backend mentioned that was never implemented: three backends exist, not four.
The search itself is honest in the code, if you read the code: it's literal substring matching, not embeddings, not vector search. There's a comment in the source that says so plainly: for real semantic search, bring your own embedder. I didn't fake a vector search implementation to make the package look more sophisticated than it is. I just described, in the README, the more sophisticated version I was imagining building, and never fully closed the gap between that vision and the substring-matching reality that shipped. The package.json keywords even list "vector" and "rag," which overstates it further. That's on me, and it's worth saying plainly rather than letting the keywords do quiet false advertising.
Why I still think the honest version is useful
Substring search sounds unglamorous next to "semantic memory," but for a huge fraction of what I actually need (did the user already tell me their name, did they already say which framework they're using, has this exact question come up before in this session) literal matching is completely sufficient and costs nothing to run. Real vector search needs an embedding model, a similarity index, and meaningfully more infrastructure than a solo project's chat memory usually justifies. I'd rather ship the honest, cheap version that solves the actual problem in front of me than half-build a vector search layer that's more impressive to describe and worse to depend on while it's unfinished.
Where it's actually used
This is the backbone underneath any conversational feature I build that needs to feel like it remembers you across more than one exchange. Nyxera's early design work leans on the same underlying idea, even though Nyxera's own memory layer is far more elaborate and purpose-built than this general package. agent-memory is the version I reach for when a project needs "just enough memory to not feel amnesiac," without needing the full weight of a dedicated memory architecture built specifically for it. Not every conversation needs Nyxera's containment-first, multi-tier recall system. Most of them just need to remember what you said five minutes ago, reliably, without a database migration.
The sliding window, and the decision to keep it dumb
getContext() returns recent history in the exact shape a chat completion call expects, windowed to however many turns you configure. I kept the windowing logic deliberately simple: a fixed count of most-recent messages, no attempt to summarize older turns or compress them down to save context space. Plenty of more sophisticated memory systems do exactly that: summarize what falls out of the window instead of just dropping it, so a long conversation doesn't lose earlier context entirely, just loses the raw detail of it. I didn't build that here, on purpose, because summarization is itself another model call, another point of failure, and another place where the summary can subtly misrepresent what was actually said. A fixed window that drops old turns cleanly is at least honest about what it's doing: nothing before the window exists in getContext()'s output, full stop, no illusion of continuity where there isn't any. If a project genuinely needs long-conversation summarization, that's a deliberate addition on top of this, not something I wanted baked silently into the default behavior.
Why three backends and not just SQLite everywhere
I could have made SQLite the only option and called it done: it's the most capable of the three by a wide margin. I kept the in-memory and file backends anyway because they solve genuinely different problems. In-memory is for tests and quick scripts where you don't want a database file left behind after every run. File-based JSON is for small tools where "readable by opening the file in a text editor" is worth more than query performance. I've debugged more than one early prototype by just opening the JSON file directly and reading it, which you can't do with a SQLite database without another tool in the loop. Three backends, three real reasons to reach for each one, not just three options for the sake of having options.
Full stack developer. Founder of Yashveer Labs. The docs describe the memory system I want to build eventually. The code is the honest, working version of what I've built so far.
Start a conversation about this.
Whether it's agent-memory itself or the next system worth building, the lab is reachable.