Generative AI · Document intelligence
Self-Correcting Multi-Document RAG
A retrieval-augmented generation system that doesn’t trust its first answer: it filters what it retrieves, grades what it writes, and tries again when the grade is low.
- Type
- Python application: CLI and Gradio web UI
- Stack
- LangChain · OpenAI · FAISS · Sentence Transformers
- Inputs
- PDF, TXT and Markdown
- License
- MIT
Overview
You give the system several documents. It can summarize them, or build a knowledge base you can question from the command line, an interactive session or a Gradio web interface. The demo runs on Hugging Face Spaces, which puts it to sleep when idle, so the first request may take a moment.
Pipeline
- 1 · RetrieveVector similarity search over FAISS (top 5 chunks)
- 2 · GuardrailAn agent drops irrelevant chunks; checks run in parallel
- 3 · GenerateA context-aware answer from the filtered chunks
- 4 · EvaluateFactual consistency scored 1–5; below 3, back to step 3
Regeneration is capped at three attempts, and the best-scoring answer is returned. The threshold, attempt count, models and chunking are all set in YAML: by default, GPT-4o generates answers and GPT-4o-mini summarizes. Summaries use a map-reduce strategy, so long documents are summarized in chunks and then combined.
Engineering
- Modules for loaders, processing (summarizer and vector store), agents and pipeline, behind a CLI with
summarize,create-kb,query,interactiveandmetricscommands. - Input validation with path-traversal protection and suspicious-pattern detection.
- Structured, rotating logs, plus metrics for per-stage latency, score distribution, token usage and filter rejection rates, saved to
logs/metrics.json. - Caching for the vector store and embeddings.
Limitations
It needs an OpenAI API key and an internet connection, and it accepts PDF, TXT and Markdown only. A self-assigned consistency score is a useful filter, not a guarantee. The evaluator is itself a language model, so a confident wrong answer can still pass.