Generative AI · Document intelligence

Self-Correcting Multi-Document RAG

A retrieval-augmented generation system that doesn’t trust its first answer: it filters what it retrieves, grades what it writes, and tries again when the grade is low.

Type
Python application: CLI and Gradio web UI
Stack
LangChain · OpenAI · FAISS · Sentence Transformers
Inputs
PDF, TXT and Markdown
License
MIT
retrieve guardrail generate evaluate score < 3 · retry ≤ 3 best answer FAISS · top-5 chunks

Overview

You give the system several documents. It can summarize them, or build a knowledge base you can question from the command line, an interactive session or a Gradio web interface. The demo runs on Hugging Face Spaces, which puts it to sleep when idle, so the first request may take a moment.

Pipeline

  1. 1 · RetrieveVector similarity search over FAISS (top 5 chunks)
  2. 2 · GuardrailAn agent drops irrelevant chunks; checks run in parallel
  3. 3 · GenerateA context-aware answer from the filtered chunks
  4. 4 · EvaluateFactual consistency scored 1–5; below 3, back to step 3

Regeneration is capped at three attempts, and the best-scoring answer is returned. The threshold, attempt count, models and chunking are all set in YAML: by default, GPT-4o generates answers and GPT-4o-mini summarizes. Summaries use a map-reduce strategy, so long documents are summarized in chunks and then combined.

Engineering

  • Modules for loaders, processing (summarizer and vector store), agents and pipeline, behind a CLI with summarize, create-kb, query, interactive and metrics commands.
  • Input validation with path-traversal protection and suspicious-pattern detection.
  • Structured, rotating logs, plus metrics for per-stage latency, score distribution, token usage and filter rejection rates, saved to logs/metrics.json.
  • Caching for the vector store and embeddings.

Limitations

It needs an OpenAI API key and an internet connection, and it accepts PDF, TXT and Markdown only. A self-assigned consistency score is a useful filter, not a guarantee. The evaluator is itself a language model, so a confident wrong answer can still pass.

Contact

Have something interesting in mind?

I’m always interested in thoughtful collaborations, interesting research problems, and opportunities to build useful things.