Skip to content
MuhammedNihaleeyPublic

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

GraphRAG with Hierarchical Indexing

A production-shaped GraphRAG system: documents are ingested with parent-child (small-to-big) chunking, an LLM extracts a typed knowledge graph with provenance back to individual chunks, and a FastAPI service fuses vector search with N-hop graph traversal to stream grounded, citable answers into an interactive Next.js chat UI.

docker compose up --build     # UI :3000 · API :8000 · Neo4j :7474 · Qdrant :6333

Contents


Architecture

flowchart TB
    subgraph client["Frontend · Next.js + Tailwind"]
        UI["Chat UI<br/>streamed tokens · citation cards"]
        GI["Graph Inspector<br/>Cytoscape"]
    end

    subgraph api["Backend · FastAPI"]
        ING["POST /ingest<br/>async job"]
        CHAT["POST /chat<br/>SSE stream"]
        SUB["GET /graph/subgraph"]
    end

    subgraph pipe["Ingestion pipeline"]
        LOAD["Loaders<br/>pdf · docx · md · html · csv"]
        CHUNK["Hierarchical chunker<br/>parent ~1000tok / child ~200tok"]
        EMB["FastEmbed<br/>bge-small-en-v1.5"]
        EXT["Extraction<br/>Claude structured output"]
    end

    subgraph ret["Retrieval"]
        AG["LangGraph agent<br/>plan → retrieve → reflect"]
        FUSE["Fusion<br/>vector → parent → N-hop graph"]
    end

    QD[("Qdrant<br/>child vectors")]
    NEO[("Neo4j<br/>hierarchy + knowledge graph")]
    LLM["Claude Opus 5"]

    UI --> CHAT
    GI --> SUB
    ING --> LOAD --> CHUNK
    CHUNK --> EMB --> QD
    CHUNK --> NEO
    CHUNK --> EXT --> LLM
    EXT --> NEO
    CHAT --> AG --> FUSE
    FUSE --> QD
    FUSE --> NEO
    AG --> LLM
    SUB --> NEO
    CHAT -.->|"status · sources · graph · token · done"| UI
Loading

Ingestion

flowchart LR
    D["Document"] --> S["Split on headings<br/>(breadcrumb path)"]
    S --> P["Parent blocks<br/>~1000 tokens<br/>contiguous, no overlap"]
    P --> C["Child chunks<br/>~200 tokens<br/>sliding, 40-token overlap"]
    C -->|embed| QD[("Qdrant")]
    P -->|text + hierarchy| NEO[("Neo4j")]
    C -->|hierarchy| NEO
    P -->|"one LLM call per block"| E["Entities + typed edges<br/>+ verbatim evidence span"]
    E -->|"evidence offsets ∩ child spans"| PROV["Child-level provenance"]
    PROV --> NEO
Loading

Extraction runs per parent block, not per child: a ~1000-token block carries enough context to resolve "it acquired them" into real entity names, which a 200-token window cannot. Child-level provenance is then recovered exactly — the model returns a verbatim evidence span for every edge, that span is located in the parent's character range, and it is intersected with the child windows' recorded offsets. Every relationship therefore points at both the parent block and the specific child chunks that assert it.


How retrieval works

sequenceDiagram
    participant U as Browser
    participant A as FastAPI
    participant G as LangGraph agent
    participant Q as Qdrant
    participant N as Neo4j
    participant C as Claude

    U->>A: POST /chat
    A-->>U: event: status
    A->>G: run graph
    G->>C: plan sub-questions
    loop per hop (bounded)
        G->>Q: search child vectors (top-k)
        Q-->>G: child hits + scores
        G->>N: expand children → parent blocks
        G->>N: seed entities → N-hop traversal
        N-->>G: subgraph (nodes + typed edges)
        G->>C: sufficient? follow-up query?
    end
    G-->>A: merged evidence bundle
    A-->>U: event: sources (citations)
    A-->>U: event: graph (subgraph)
    A->>C: stream grounded answer
    C-->>A: token deltas
    A-->>U: event: token ×N
    A-->>U: event: done
Loading

The fusion is sequential, not two rankers merged. Vector hits decide which region of the graph is relevant, so the traversal is seeded by (a) entities mentioned in the child chunks that actually matched and (b) entities the question names outright, resolved through a Neo4j full-text index. That keeps the subgraph tight and on-topic instead of drifting toward whatever node happens to be well-connected.

Small-to-big in one line: search the ~200-token child for precision, answer from its ~1000-token parent for context.


Graph schema

(:Document {id, title, source, ingested_at})
  -[:HAS_PARENT]-> (:Parent {id, doc_id, text, ordinal, section_path, token_count})
  -[:HAS_CHILD]->  (:Child  {id, doc_id, parent_id, text, ordinal,
                             start_char, end_char})

(:Entity {id, name, type, description, doc_ids})
  -[:APPEARS_IN]->   (:Parent)      // provenance: context block
  -[:MENTIONED_IN]-> (:Child)       // provenance: exact vector chunk

(:Entity)-[:ACQUIRED|FOUNDED_BY|PARTNERED_WITH|… {
    type, description, evidence, confidence,
    doc_id, parent_id, child_ids
}]->(:Entity)

Entity-to-entity edges use the model's own SCREAMING_SNAKE_CASE relation as the actual Neo4j relationship type (via apoc.merge.relationship), so MATCH (a)-[:ACQUIRED]->(b) works natively. The type is repeated in a type property so every query in the codebase behaves identically on the no-APOC fallback path.

Entities resolve by slug (Northwind Analytics → northwind_analytics), so the same company mentioned across blocks and documents merges into one node while provenance lists accumulate. Constraints, range indexes and a full-text index on Entity(name, description) are created idempotently at startup.

Structural edges (HAS_PARENT, HAS_CHILD, APPEARS_IN, MENTIONED_IN) never connect two :Entity nodes, so MATCH (s:Entity)-[r]-(m:Entity) traverses only knowledge edges without needing a type filter.


Setup

1. Configure

cp .env.example .env

Then pick one LLM provider and set its key:

Provider .env Notes
Anthropic (default) LLM_PROVIDER=anthropic
ANTHROPIC_API_KEY=sk-ant-…
Claude Opus 5. Best extraction quality.
Google Gemini LLM_PROVIDER=gemini
GEMINI_API_KEY=AIza…
Free key. Defaults to gemini-3.5-flash-lite — see the free-tier note in trade-offs.

Switching is an environment variable, not a code change — everything goes through the three calls in core/llm.py. /api/v1/health reports which provider is live and whether its credentials are present.

No embedding API key is needed either way — embeddings run locally on CPU via FastEmbed and the model is baked into the backend image at build time, so on Gemini's free tier the whole system costs nothing to run.

2. Run everything

docker compose up --build
Service URL Notes
Chat UI http://localhost:3000
API docs http://localhost:8000/docs OpenAPI / Swagger
Health http://localhost:8000/api/v1/health per-dependency status
Neo4j Browser http://localhost:7474 neo4j / graphrag_password
Qdrant http://localhost:6333/dashboard

First build takes a few minutes (ONNX runtime + embedding model). Neo4j is gated behind a healthcheck, so the backend starts only once Bolt answers.

3. Ingest and ask

# Ingest the sample corpus
curl -sS -X POST http://localhost:8000/api/v1/ingest \
  -F "files=@data/samples/northwind_acquisition.md"
# -> {"job_id":"…","state":"queued","poll_url":"/api/v1/ingest/jobs/…"}

# Watch progress
curl -sS http://localhost:8000/api/v1/ingest/jobs/<job_id>

# Ask (streamed)
curl -N -X POST http://localhost:8000/api/v1/chat \
  -H 'Content-Type: application/json' \
  -d '{"message":"Who founded the company that Northwind Analytics acquired?"}'

Or just open http://localhost:3000, drop a file into the Documents tab, and ask in the chat.

Local development (without Docker)

# stores only
docker compose up neo4j qdrant

# backend
cd backend
python -m venv .venv && source .venv/bin/activate
pip install -r requirements-dev.txt
export ANTHROPIC_API_KEY=sk-ant-... \
       NEO4J_URI=bolt://localhost:7687 \
       QDRANT_URL=http://localhost:6333
uvicorn app.main:app --reload

# frontend
cd frontend && npm install && npm run dev

API reference

Method Path Description
POST /api/v1/ingest Async ingestion. Accepts JSON {text, title, source} or multipart files / text. Returns 202 with a job id.
GET /api/v1/ingest/jobs/{job_id} Job state, stage, progress and counts.
POST /api/v1/chat SSE stream. Events: status, sources, graph, token, done, error.
POST /api/v1/chat/sync Same pipeline, single JSON response — for scripts and eval harnesses.
GET /api/v1/graph/subgraph Node-edge JSON. Seed with ?entity=… (repeatable), ?q=… (text), or omit both for a corpus overview. Params: hops, limit, doc_id.
GET /api/v1/graph/stats Document / parent / child / entity / relationship counts.
GET /api/v1/documents Indexed documents.
DELETE /api/v1/documents/{doc_id} Removes the document from both stores and drops orphaned entities.
GET /api/v1/health Per-dependency health; degraded when a store is unreachable.

SSE event contract

Event Payload When
status {step, label, detail} Once per agent node, live
sources {citations[], sub_questions[], hops_used} After retrieval settles
graph {nodes[], edges[]} After retrieval settles
token {text} Per model delta
done {answer, grounded, citation_index, latency_ms, retrieval} Terminal
error {message} Terminal on failure

sources and graph are always emitted before the first token, so the UI can render citation cards and the subgraph while the answer is still streaming.


Sample queries

Against data/samples/northwind_acquisition.md:

Question What it exercises
"Who founded the company that Northwind Analytics acquired?" Multi-hop: ACQUIRED → FOUNDED_BY. The answer (Priya Raman) never appears in the same passage as the acquisition.
"How is Crestline Cloud connected to Northwind Analytics?" Graph traversal over two distinct paths — direct SUPPLIES, and via Halcyon Data PARTNERED_WITH.
"What did Halcyon Data acquire, and what happened to the product?" Entity resolution across sections plus passage detail.
"Why was integration expected to take nine months?" Pure passage retrieval — the graph adds nothing, and the answer should cite [P*] only.
"What was Northwind's FY2023 revenue and margin?" Numeric detail from the parent block; shows why small-to-big matters (the child chunk alone loses the surrounding context).
"Who is the CEO of Tessellate Metrics?" Negative case — not in the corpus. The system should say so rather than invent an answer.

Toggle Multi-hop agent off in the UI to compare single-pass latency against multi-hop recall on question 1.


Tests

cd backend
pip install -r requirements-dev.txt
python -m pytest ../tests -v

124 tests, fully offline — no Neo4j, Qdrant or API key required. Coverage:

File Focus
test_chunking.py Parent/child linkage, character-offset integrity, size budgets, overlap, deterministic ids, hard-splitting of oversized segments
test_extraction.py Type normalization to safe Neo4j tokens, evidence→child provenance mapping, entity merging, per-block failure isolation
test_retrieval.py Parent ranking, small-to-big expansion, graph seeding, multi-pass merging, context assembly, triple deduplication and citation indexing
test_graph_store.py Cypher construction — brace-literal integrity, APOC and fallback write paths, bounded per-hop traversal, node caps
test_agent.py Planning, hop budgets, follow-up loops, repeat-query suppression, graceful degradation when the planner or reflector fails
test_api.py Ingestion job lifecycle, the full SSE contract and event ordering, graph endpoints, validation and error paths
test_loaders.py Format detection, encoding fallbacks, malformed input
test_providers.py Provider selection, Gemini request shaping and schema wiring, retry classification, rate-limiter pacing

Run them in the container instead, matching the shipped environment:

docker compose build backend
docker run --rm -v "$PWD/tests:/tests:ro" -v "$PWD/backend/app:/app/app:ro" -w /app \
  graphrag-backend:latest \
  sh -c "pip install -q pytest pytest-asyncio httpx; python -m pytest /tests -q"

Design trade-offs: retrieval accuracy vs. latency

Chunk sizes (200 / 1000 tokens). Small children keep embeddings dense and on-topic, which is what drives precision; large parents give the model room to reason. The cost is index size — roughly 5× more vectors than parent-level chunking, and 40-token overlap adds ~20% on top. Qdrant absorbs this easily at this scale; at tens of millions of chunks you would drop the overlap first.

Parents never span a section boundary. 1000 tokens is a ceiling, not a target: a 300-token section produces a 300-token parent rather than being padded with unrelated text from the next section. Coherent context beats full context blocks, and it keeps the section_path breadcrumb on every citation truthful. The corollary is that heading-only sections would otherwise become useless parent blocks containing nothing but a title, so any section under half a child chunk is merged into the section that follows it.

Extraction per parent, not per child. One LLM call per ~1000 tokens instead of per ~200 tokens is ~5× fewer calls and better coreference resolution. The cost is that ingestion is the slow path: a 20-page document is ~20 calls, run at concurrency 4. This is exactly why /ingest is asynchronous with a polled job.

Local embeddings (FastEmbed / bge-small, 384-dim). No second API key, no per-chunk network hop, and ingestion stays fast on CPU. A larger hosted embedding model would buy a few points of recall; for chunk-level retrieval feeding a strong reader model, that was not worth the added key, cost and latency. EMBEDDING_MODEL / EMBEDDING_DIM swap it.

BFS traversal instead of variable-length Cypher. N hops = N bounded queries rather than one -[*1..N]- match. Slightly more round trips (each is ~1-3 ms on a warm graph), in exchange for exact hop numbering in the UI, a hard cap on fan-out per round, and no string-built Cypher. A MATCH (a)-[*1..3]-(b) on a hub entity can explode combinatorially; this cannot.

Gemini's free tier is capped per day, per model — that is the real constraint. Measured on a live free key, gemini-3.6-flash allows 20 requests/day; since extraction is one call per parent block, that is four documents' worth of ingestion before the model is unusable. The *-flash-lite models have far more headroom, so they are the default here: a weaker extractor that runs beats a stronger one that 429s. Check your own limits at AI Studio — Google no longer publishes a static table. Budgeting per question: MAX_AGENT_HOPS=3 costs up to three calls (plan + reflect + generate); MAX_AGENT_HOPS=1 costs one.

Pacing still guards the per-minute limit. Free keys are metered per minute and answer with 429 instead of queueing, so the provider puts every request through a token bucket sized by GEMINI_RPM (default 10) with exponential backoff behind it as a safety net. The practical consequence is that ingestion is the slow part: extraction is one call per parent block, so at 10 rpm a 20-block document takes ~2 minutes to index. Querying is unaffected — that is 1–3 calls. On a paid key set GEMINI_RPM=0 to disable pacing. Choosing gemini-3.6-flash over a Pro model is deliberate here: extraction is a high-volume, well-specified task where a fast model at high throughput beats a slow one that exhausts the quota.

The multi-hop agent is opt-in per request. Planning and reflection are two extra model calls. With max_agent_hops = 1 the agent skips straight to a single retrieval pass — the right default for lookups. Relational questions pay the planner call to earn substantially better recall. The UI exposes this as a toggle so the trade-off is visible rather than buried.

Rough shape of the latency budget on a warm system:

Stage Typical Notes
Query embedding ~5 ms local, in-process
Qdrant search (k=12) ~10 ms HNSW
Parent fetch ~5 ms indexed lookup
Graph traversal (2 hops) ~15 ms 2 bounded queries
Retrieval subtotal ~35 ms reported per-request in timings_ms
Planner call ~1-2 s skipped when max_agent_hops = 1
Reflection call ~1-2 s skipped on the final hop
Answer generation ~2-6 s streamed; first token much earlier

Retrieval is not the bottleneck — model calls are. That is why the agent is bounded, why status events stream immediately, and why the answer streams token by token instead of arriving whole.

Generation streams outside the LangGraph agent. The graph owns evidence gathering only; the API streams the final answer from the merged bundle. Keeping generation out of a graph node means the first token reaches the client as soon as evidence settles, rather than after a node boundary.

Prompt caching on the system prompt. Extraction reuses one large system prompt across every block of every document, marked cache_control: ephemeral, so repeated blocks hit cache rather than re-paying for the prefix.

Known limitations

  • Qdrant client and server versions are coupled. The client refuses a server more than one minor version away, so qdrant-client in requirements.txt and the qdrant/qdrant tag in docker-compose.yml must be bumped together. Note that Qdrant storage is not backward compatible across such a bump — an existing qdrant_data volume must be recreated, which for this system means re-ingesting (Neo4j keeps the text, so nothing is lost but time).
  • Job state is in-process. A job is bound to the worker running it, so horizontal scaling of the API needs Redis or a real queue behind services/jobs.py. The interface is already narrow enough to swap.
  • Entity resolution is slug-based. Northwind Analytics and Northwind resolve to different nodes. Embedding-based entity linking or an LLM canonicalization pass would fix this and is the highest-value next step.
  • No reranker. A cross-encoder over the top-k children before parent expansion would improve precision at ~50-100 ms cost.
  • Scanned PDFs are rejected, not OCR'd — the loader fails loudly rather than indexing an empty document.
  • Ingestion is not transactional across stores. Vectors are written before the graph, so a crash mid-extraction leaves a searchable but graph-less document rather than orphaned nodes. Re-ingesting the same file is safe: ids are content-derived, so every write is an upsert.

Project layout

.
├── docker-compose.yml          # UI + API + Neo4j + Qdrant, one command
├── .env.example
├── backend/
│   ├── Dockerfile
│   ├── requirements.txt
│   └── app/
│       ├── main.py             # app factory, lifespan, CORS
│       ├── config.py           # env-driven settings
│       ├── schemas.py          # Pydantic contracts
│       ├── api/                # ingest · chat (SSE) · graph · health
│       ├── core/
│       │   ├── chunking.py     # parent-child hierarchical splitter
│       │   ├── extraction.py   # entities, typed edges, provenance mapping
│       │   ├── retrieval.py    # vector → parent → N-hop fusion
│       │   ├── agent.py        # LangGraph plan → retrieve → reflect
│       │   ├── graph_store.py  # Neo4j: schema, writes, BFS traversal
│       │   ├── vector_store.py # Qdrant: child vectors
│       │   ├── embeddings.py   # local ONNX embeddings
│       │   ├── loaders.py      # pdf · docx · md · html · csv · json
│       │   ├── llm.py          # provider-agnostic façade
│       │   ├── providers/     # anthropic · gemini, one narrow interface
│       │   └── prompts.py
│       └── services/           # ingestion pipeline · job registry
├── frontend/
│   ├── Dockerfile
│   ├── app/                    # Next.js App Router
│   ├── components/             # ChatPanel · MessageBubble · CitationCard
│   │                           # GraphInspector (Cytoscape) · IngestPanel
│   └── lib/                    # typed API client + SSE framing
├── tests/                      # 88 offline tests
└── data/samples/               # synthetic corpus for the sample queries

Stack

Claude Opus 5 (Anthropic SDK) or Gemini 3.6 Flash (google-genai) — structured outputs either way · LangGraph · FastAPI · Pydantic v2 · Neo4j 5 + APOC · Qdrant · FastEmbed (bge-small-en-v1.5) · Next.js 14 · Tailwind · Cytoscape

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages