A local Retrieval-Augmented Generation (RAG) API built with Python and FastAPI.
This project implements a complete RAG pipeline that retrieves relevant information from a custom knowledge base and uses a local LLM to generate grounded answers β running entirely on the local machine with zero cloud/API costs.
Documents
β
Text Chunking
β
Vector Embeddings
β
ChromaDB
β
Semantic Retrieval
β
Relevant Context
β
Prompt Augmentation
β
Qwen LLM
β
AI-Generated Answer