A production-shaped GraphRAG system: documents are ingested with parent-child (small-to-big) chunking, an LLM extracts a typed knowledge graph with provenance back to individual chunks, and a FastAPI service fuses vector search with N-hop graph traversal to stream grounded, citable answers into an interactive Next.js chat UI.
docker compose up --build # UI :3000 · API :8000 · Neo4j :7474 · Qdrant :6333
- Architecture
- How retrieval works
- Graph schema
- Setup
- API reference
- Sample queries
- Tests
- Design trade-offs: retrieval accuracy vs. latency
- Project layout
flowchart TB
subgraph client["Frontend · Next.js + Tailwind"]
UI["Chat UI<br/>streamed tokens · citation cards"]
GI["Graph Inspector<br/>Cytoscape"]
end
subgraph api["Backend · FastAPI"]
ING["POST /ingest<br/>async job"]
CHAT["POST /chat<br/>SSE stream"]
SUB["GET /graph/subgraph"]
end
subgraph pipe["Ingestion pipeline"]
LOAD["Loaders<br/>pdf · docx · md · html · csv"]
CHUNK["Hierarchical chunker<br/>parent ~1000tok / child ~200tok"]
EMB["FastEmbed<br/>bge-small-en-v1.5"]
EXT["Extraction<br/>Claude structured output"]
end
subgraph ret["Retrieval"]
AG["LangGraph agent<br/>plan → retrieve → reflect"]
FUSE["Fusion<br/>vector → parent → N-hop graph"]
end
QD[("Qdrant<br/>child vectors")]
NEO[("Neo4j<br/>hierarchy + knowledge graph")]
LLM["Claude Opus 5"]
UI --> CHAT
GI --> SUB
ING --> LOAD --> CHUNK
CHUNK --> EMB --> QD
CHUNK --> NEO
CHUNK --> EXT --> LLM
EXT --> NEO
CHAT --> AG --> FUSE
FUSE --> QD
FUSE --> NEO
AG --> LLM
SUB --> NEO
CHAT -.->|"status · sources · graph · token · done"| UI
flowchart LR
D["Document"] --> S["Split on headings<br/>(breadcrumb path)"]
S --> P["Parent blocks<br/>~1000 tokens<br/>contiguous, no overlap"]
P --> C["Child chunks<br/>~200 tokens<br/>sliding, 40-token overlap"]
C -->|embed| QD[("Qdrant")]
P -->|text + hierarchy| NEO[("Neo4j")]
C -->|hierarchy| NEO
P -->|"one LLM call per block"| E["Entities + typed edges<br/>+ verbatim evidence span"]
E -->|"evidence offsets ∩ child spans"| PROV["Child-level provenance"]
PROV --> NEO
Extraction runs per parent block, not per child: a ~1000-token block carries
enough context to resolve "it acquired them" into real entity names, which a
200-token window cannot. Child-level provenance is then recovered exactly — the
model returns a verbatim evidence span for every edge, that span is located in
the parent's character range, and it is intersected with the child windows'
recorded offsets. Every relationship therefore points at both the parent
block and the specific child chunks that assert it.
sequenceDiagram
participant U as Browser
participant A as FastAPI
participant G as LangGraph agent
participant Q as Qdrant
participant N as Neo4j
participant C as Claude
U->>A: POST /chat
A-->>U: event: status
A->>G: run graph
G->>C: plan sub-questions
loop per hop (bounded)
G->>Q: search child vectors (top-k)
Q-->>G: child hits + scores
G->>N: expand children → parent blocks
G->>N: seed entities → N-hop traversal
N-->>G: subgraph (nodes + typed edges)
G->>C: sufficient? follow-up query?
end
G-->>A: merged evidence bundle
A-->>U: event: sources (citations)
A-->>U: event: graph (subgraph)
A->>C: stream grounded answer
C-->>A: token deltas
A-->>U: event: token ×N
A-->>U: event: done
The fusion is sequential, not two rankers merged. Vector hits decide which region of the graph is relevant, so the traversal is seeded by (a) entities mentioned in the child chunks that actually matched and (b) entities the question names outright, resolved through a Neo4j full-text index. That keeps the subgraph tight and on-topic instead of drifting toward whatever node happens to be well-connected.
Small-to-big in one line: search the ~200-token child for precision, answer from its ~1000-token parent for context.
(:Document {id, title, source, ingested_at})
-[:HAS_PARENT]-> (:Parent {id, doc_id, text, ordinal, section_path, token_count})
-[:HAS_CHILD]-> (:Child {id, doc_id, parent_id, text, ordinal,
start_char, end_char})
(:Entity {id, name, type, description, doc_ids})
-[:APPEARS_IN]-> (:Parent) // provenance: context block
-[:MENTIONED_IN]-> (:Child) // provenance: exact vector chunk
(:Entity)-[:ACQUIRED|FOUNDED_BY|PARTNERED_WITH|… {
type, description, evidence, confidence,
doc_id, parent_id, child_ids
}]->(:Entity)Entity-to-entity edges use the model's own SCREAMING_SNAKE_CASE relation as
the actual Neo4j relationship type (via apoc.merge.relationship), so
MATCH (a)-[:ACQUIRED]->(b) works natively. The type is repeated in a type
property so every query in the codebase behaves identically on the no-APOC
fallback path.
Entities resolve by slug (Northwind Analytics → northwind_analytics), so the
same company mentioned across blocks and documents merges into one node while
provenance lists accumulate. Constraints, range indexes and a full-text index on
Entity(name, description) are created idempotently at startup.
Structural edges (HAS_PARENT, HAS_CHILD, APPEARS_IN, MENTIONED_IN) never
connect two :Entity nodes, so MATCH (s:Entity)-[r]-(m:Entity) traverses only
knowledge edges without needing a type filter.
cp .env.example .envThen pick one LLM provider and set its key:
| Provider | .env |
Notes |
|---|---|---|
| Anthropic (default) | LLM_PROVIDER=anthropicANTHROPIC_API_KEY=sk-ant-… |
Claude Opus 5. Best extraction quality. |
| Google Gemini | LLM_PROVIDER=geminiGEMINI_API_KEY=AIza… |
Free key. Defaults to gemini-3.5-flash-lite — see the free-tier note in trade-offs. |
Switching is an environment variable, not a code change — everything goes
through the three calls in core/llm.py. /api/v1/health reports which
provider is live and whether its credentials are present.
No embedding API key is needed either way — embeddings run locally on CPU via FastEmbed and the model is baked into the backend image at build time, so on Gemini's free tier the whole system costs nothing to run.
docker compose up --build| Service | URL | Notes |
|---|---|---|
| Chat UI | http://localhost:3000 | |
| API docs | http://localhost:8000/docs | OpenAPI / Swagger |
| Health | http://localhost:8000/api/v1/health | per-dependency status |
| Neo4j Browser | http://localhost:7474 | neo4j / graphrag_password |
| Qdrant | http://localhost:6333/dashboard |
First build takes a few minutes (ONNX runtime + embedding model). Neo4j is gated behind a healthcheck, so the backend starts only once Bolt answers.
# Ingest the sample corpus
curl -sS -X POST http://localhost:8000/api/v1/ingest \
-F "files=@data/samples/northwind_acquisition.md"
# -> {"job_id":"…","state":"queued","poll_url":"/api/v1/ingest/jobs/…"}
# Watch progress
curl -sS http://localhost:8000/api/v1/ingest/jobs/<job_id>
# Ask (streamed)
curl -N -X POST http://localhost:8000/api/v1/chat \
-H 'Content-Type: application/json' \
-d '{"message":"Who founded the company that Northwind Analytics acquired?"}'Or just open http://localhost:3000, drop a file into the Documents tab, and ask in the chat.
# stores only
docker compose up neo4j qdrant
# backend
cd backend
python -m venv .venv && source .venv/bin/activate
pip install -r requirements-dev.txt
export ANTHROPIC_API_KEY=sk-ant-... \
NEO4J_URI=bolt://localhost:7687 \
QDRANT_URL=http://localhost:6333
uvicorn app.main:app --reload
# frontend
cd frontend && npm install && npm run dev| Method | Path | Description |
|---|---|---|
POST |
/api/v1/ingest |
Async ingestion. Accepts JSON {text, title, source} or multipart files / text. Returns 202 with a job id. |
GET |
/api/v1/ingest/jobs/{job_id} |
Job state, stage, progress and counts. |
POST |
/api/v1/chat |
SSE stream. Events: status, sources, graph, token, done, error. |
POST |
/api/v1/chat/sync |
Same pipeline, single JSON response — for scripts and eval harnesses. |
GET |
/api/v1/graph/subgraph |
Node-edge JSON. Seed with ?entity=… (repeatable), ?q=… (text), or omit both for a corpus overview. Params: hops, limit, doc_id. |
GET |
/api/v1/graph/stats |
Document / parent / child / entity / relationship counts. |
GET |
/api/v1/documents |
Indexed documents. |
DELETE |
/api/v1/documents/{doc_id} |
Removes the document from both stores and drops orphaned entities. |
GET |
/api/v1/health |
Per-dependency health; degraded when a store is unreachable. |
| Event | Payload | When |
|---|---|---|
status |
{step, label, detail} |
Once per agent node, live |
sources |
{citations[], sub_questions[], hops_used} |
After retrieval settles |
graph |
{nodes[], edges[]} |
After retrieval settles |
token |
{text} |
Per model delta |
done |
{answer, grounded, citation_index, latency_ms, retrieval} |
Terminal |
error |
{message} |
Terminal on failure |
sources and graph are always emitted before the first token, so the UI
can render citation cards and the subgraph while the answer is still streaming.
Against data/samples/northwind_acquisition.md:
| Question | What it exercises |
|---|---|
| "Who founded the company that Northwind Analytics acquired?" | Multi-hop: ACQUIRED → FOUNDED_BY. The answer (Priya Raman) never appears in the same passage as the acquisition. |
| "How is Crestline Cloud connected to Northwind Analytics?" | Graph traversal over two distinct paths — direct SUPPLIES, and via Halcyon Data PARTNERED_WITH. |
| "What did Halcyon Data acquire, and what happened to the product?" | Entity resolution across sections plus passage detail. |
| "Why was integration expected to take nine months?" | Pure passage retrieval — the graph adds nothing, and the answer should cite [P*] only. |
| "What was Northwind's FY2023 revenue and margin?" | Numeric detail from the parent block; shows why small-to-big matters (the child chunk alone loses the surrounding context). |
| "Who is the CEO of Tessellate Metrics?" | Negative case — not in the corpus. The system should say so rather than invent an answer. |
Toggle Multi-hop agent off in the UI to compare single-pass latency against multi-hop recall on question 1.
cd backend
pip install -r requirements-dev.txt
python -m pytest ../tests -v124 tests, fully offline — no Neo4j, Qdrant or API key required. Coverage:
| File | Focus |
|---|---|
test_chunking.py |
Parent/child linkage, character-offset integrity, size budgets, overlap, deterministic ids, hard-splitting of oversized segments |
test_extraction.py |
Type normalization to safe Neo4j tokens, evidence→child provenance mapping, entity merging, per-block failure isolation |
test_retrieval.py |
Parent ranking, small-to-big expansion, graph seeding, multi-pass merging, context assembly, triple deduplication and citation indexing |
test_graph_store.py |
Cypher construction — brace-literal integrity, APOC and fallback write paths, bounded per-hop traversal, node caps |
test_agent.py |
Planning, hop budgets, follow-up loops, repeat-query suppression, graceful degradation when the planner or reflector fails |
test_api.py |
Ingestion job lifecycle, the full SSE contract and event ordering, graph endpoints, validation and error paths |
test_loaders.py |
Format detection, encoding fallbacks, malformed input |
test_providers.py |
Provider selection, Gemini request shaping and schema wiring, retry classification, rate-limiter pacing |
Run them in the container instead, matching the shipped environment:
docker compose build backend
docker run --rm -v "$PWD/tests:/tests:ro" -v "$PWD/backend/app:/app/app:ro" -w /app \
graphrag-backend:latest \
sh -c "pip install -q pytest pytest-asyncio httpx; python -m pytest /tests -q"Chunk sizes (200 / 1000 tokens). Small children keep embeddings dense and on-topic, which is what drives precision; large parents give the model room to reason. The cost is index size — roughly 5× more vectors than parent-level chunking, and 40-token overlap adds ~20% on top. Qdrant absorbs this easily at this scale; at tens of millions of chunks you would drop the overlap first.
Parents never span a section boundary. 1000 tokens is a ceiling, not a
target: a 300-token section produces a 300-token parent rather than being
padded with unrelated text from the next section. Coherent context beats full
context blocks, and it keeps the section_path breadcrumb on every citation
truthful. The corollary is that heading-only sections would otherwise become
useless parent blocks containing nothing but a title, so any section under half
a child chunk is merged into the section that follows it.
Extraction per parent, not per child. One LLM call per ~1000 tokens instead
of per ~200 tokens is ~5× fewer calls and better coreference resolution. The
cost is that ingestion is the slow path: a 20-page document is ~20 calls, run at
concurrency 4. This is exactly why /ingest is asynchronous with a polled job.
Local embeddings (FastEmbed / bge-small, 384-dim). No second API key, no
per-chunk network hop, and ingestion stays fast on CPU. A larger hosted
embedding model would buy a few points of recall; for chunk-level retrieval
feeding a strong reader model, that was not worth the added key, cost and
latency. EMBEDDING_MODEL / EMBEDDING_DIM swap it.
BFS traversal instead of variable-length Cypher. N hops = N bounded queries
rather than one -[*1..N]- match. Slightly more round trips (each is ~1-3 ms on
a warm graph), in exchange for exact hop numbering in the UI, a hard cap on
fan-out per round, and no string-built Cypher. A MATCH (a)-[*1..3]-(b) on a
hub entity can explode combinatorially; this cannot.
Gemini's free tier is capped per day, per model — that is the real
constraint. Measured on a live free key, gemini-3.6-flash allows 20
requests/day; since extraction is one call per parent block, that is four
documents' worth of ingestion before the model is unusable. The *-flash-lite
models have far more headroom, so they are the default here: a weaker extractor
that runs beats a stronger one that 429s. Check your own limits at
AI Studio — Google no longer publishes
a static table. Budgeting per question: MAX_AGENT_HOPS=3 costs up to three
calls (plan + reflect + generate); MAX_AGENT_HOPS=1 costs one.
Pacing still guards the per-minute limit. Free keys are metered per
minute and answer with 429 instead of queueing, so the provider puts every
request through a token bucket sized by GEMINI_RPM (default 10) with
exponential backoff behind it as a safety net. The practical consequence is
that ingestion is the slow part: extraction is one call per parent block, so
at 10 rpm a 20-block document takes ~2 minutes to index. Querying is unaffected
— that is 1–3 calls. On a paid key set GEMINI_RPM=0 to disable pacing.
Choosing gemini-3.6-flash over a Pro model is deliberate here: extraction is a
high-volume, well-specified task where a fast model at high throughput beats a
slow one that exhausts the quota.
The multi-hop agent is opt-in per request. Planning and reflection are two
extra model calls. With max_agent_hops = 1 the agent skips straight to a
single retrieval pass — the right default for lookups. Relational questions pay
the planner call to earn substantially better recall. The UI exposes this as a
toggle so the trade-off is visible rather than buried.
Rough shape of the latency budget on a warm system:
| Stage | Typical | Notes |
|---|---|---|
| Query embedding | ~5 ms | local, in-process |
| Qdrant search (k=12) | ~10 ms | HNSW |
| Parent fetch | ~5 ms | indexed lookup |
| Graph traversal (2 hops) | ~15 ms | 2 bounded queries |
| Retrieval subtotal | ~35 ms | reported per-request in timings_ms |
| Planner call | ~1-2 s | skipped when max_agent_hops = 1 |
| Reflection call | ~1-2 s | skipped on the final hop |
| Answer generation | ~2-6 s | streamed; first token much earlier |
Retrieval is not the bottleneck — model calls are. That is why the agent is bounded, why status events stream immediately, and why the answer streams token by token instead of arriving whole.
Generation streams outside the LangGraph agent. The graph owns evidence gathering only; the API streams the final answer from the merged bundle. Keeping generation out of a graph node means the first token reaches the client as soon as evidence settles, rather than after a node boundary.
Prompt caching on the system prompt. Extraction reuses one large system
prompt across every block of every document, marked cache_control: ephemeral,
so repeated blocks hit cache rather than re-paying for the prefix.
- Qdrant client and server versions are coupled. The client refuses a
server more than one minor version away, so
qdrant-clientinrequirements.txtand theqdrant/qdranttag indocker-compose.ymlmust be bumped together. Note that Qdrant storage is not backward compatible across such a bump — an existingqdrant_datavolume must be recreated, which for this system means re-ingesting (Neo4j keeps the text, so nothing is lost but time). - Job state is in-process. A job is bound to the worker running it, so
horizontal scaling of the API needs Redis or a real queue behind
services/jobs.py. The interface is already narrow enough to swap. - Entity resolution is slug-based.
Northwind AnalyticsandNorthwindresolve to different nodes. Embedding-based entity linking or an LLM canonicalization pass would fix this and is the highest-value next step. - No reranker. A cross-encoder over the top-k children before parent expansion would improve precision at ~50-100 ms cost.
- Scanned PDFs are rejected, not OCR'd — the loader fails loudly rather than indexing an empty document.
- Ingestion is not transactional across stores. Vectors are written before the graph, so a crash mid-extraction leaves a searchable but graph-less document rather than orphaned nodes. Re-ingesting the same file is safe: ids are content-derived, so every write is an upsert.
.
├── docker-compose.yml # UI + API + Neo4j + Qdrant, one command
├── .env.example
├── backend/
│ ├── Dockerfile
│ ├── requirements.txt
│ └── app/
│ ├── main.py # app factory, lifespan, CORS
│ ├── config.py # env-driven settings
│ ├── schemas.py # Pydantic contracts
│ ├── api/ # ingest · chat (SSE) · graph · health
│ ├── core/
│ │ ├── chunking.py # parent-child hierarchical splitter
│ │ ├── extraction.py # entities, typed edges, provenance mapping
│ │ ├── retrieval.py # vector → parent → N-hop fusion
│ │ ├── agent.py # LangGraph plan → retrieve → reflect
│ │ ├── graph_store.py # Neo4j: schema, writes, BFS traversal
│ │ ├── vector_store.py # Qdrant: child vectors
│ │ ├── embeddings.py # local ONNX embeddings
│ │ ├── loaders.py # pdf · docx · md · html · csv · json
│ │ ├── llm.py # provider-agnostic façade
│ │ ├── providers/ # anthropic · gemini, one narrow interface
│ │ └── prompts.py
│ └── services/ # ingestion pipeline · job registry
├── frontend/
│ ├── Dockerfile
│ ├── app/ # Next.js App Router
│ ├── components/ # ChatPanel · MessageBubble · CitationCard
│ │ # GraphInspector (Cytoscape) · IngestPanel
│ └── lib/ # typed API client + SSE framing
├── tests/ # 88 offline tests
└── data/samples/ # synthetic corpus for the sample queries
Claude Opus 5 (Anthropic SDK) or Gemini 3.6 Flash (google-genai) — structured outputs either way · LangGraph · FastAPI · Pydantic v2 · Neo4j 5 + APOC · Qdrant · FastEmbed (bge-small-en-v1.5) · Next.js 14 · Tailwind · Cytoscape