A portable pattern for giving LLM chatbots memory — done cheaply and reliably, without a vector database.
TL;DR — Keep two things apart:
- Memory = what the bot knows about this user (a growing, per-user record in a normal SQL DB).
- Knowledge = what is generally true (a small, curated, read-only set of facts looked up by exact key).
The LLM is not where either lives. It's the glue that combines them at answer time. Expensive generation runs only when the user's state actually changed (snapshot + hash → cache). That last part is why it's called fast-memory.
Most "give my bot memory" tutorials reach for a vector database + RAG on day one. For a large, unstructured, fast-changing corpus that's right. For a bounded, structured domain (a product catalog, a set of reference values, a user's own history) it's overkill: slower, pricier, and it hallucinates more than an exact lookup does.
This repo is the distilled, domain-neutral version of a memory system that runs a real health assistant in production. All domain specifics have been stripped and replaced with a toy "personal recommendation assistant" so you can read the shape and drop in your own domain.
| Path | Read it for |
|---|---|
AGENTS.md |
Start here if you're an AI agent. Deploy checklist + invariants. |
ARCHITECTURE.md |
The two layers, the request flow, why no vector-RAG. |
spec/ |
Precise spec of each layer + the synthesis step. |
schema/001_init.sql |
Portable SQL DDL for the memory tables. |
reference/ |
Minimal, readable Python reference implementation. |
examples/walkthrough.md |
End-to-end trace on the toy domain. |
- Provision Postgres, apply
schema/001_init.sql. - Copy
.env.example→.env, fill in your LLM key + DB URL. - Replace the toy knowledge in
reference/knowledge/catalog.jsonwith your curated facts (keyed by a stable id). - Adapt the fields in
reference/models.pyandreference/snapshot.pyto what your bot needs to remember. - Run the walkthrough to confirm the cache/hit path works.
MIT — see LICENSE.