Teach Your AI to Think Twice: Building a Self-Evaluating RAG Bot
A feedback-driven RAG assistant that evaluates and revises its own answers before responding.

Representational image
Imagine asking an AI assistant a question and having it double-check its own answer, identify flaws, and improve itself—all before you see the response.
That’s what I set out to build. The result? A feedback-driven RAG assistant that reads documents, thinks before it speaks, and refines its responses until a second AI (Google Gemini Flash) is satisfied.
This isn’t just a chatbot; it’s an experiment in self-improving AI assistants leveraging RAG principles. This is your personal research assistant that:
- Runs on your machine
- Talks like whoever you want (Sherlock Holmes, an SRE mentor, a Roman senator…). I prefer Sherlock and constantly change the conversational rules for fun.
- Learns from your PDFs—technically, only to generate responses.
- And above all, knows when it may be wrong—and fixes it.
Let’s build it. Full code at repo.

Output representation. You can change the BOT name and config as needed.
🔧 What We’re Building
At its core, this project is a Retrieval-Augmented Generation (RAG) bot:
- 🔍 Reads your
.pdfand.txtfiles. I plan to expand the supported formats in the future. - 🧠 Embeds them using HuggingFace + FAISS
- 🤖 Answers your questions using either OpenAI or a local LLM via LM Studio
- 🧪 Validates each answer with Gemini Flash
- ♻️ If Gemini isn’t happy, the answer is revised—up to 5 times. Configure as you see fit.
And the twist? You can define the assistant’s entire personality in one place and the whole system will follow it, including the evaluator.
🛠️ Tech Stack
+-----------------+--------------------------------+
| Tool | Purpose |
+-----------------+--------------------------------+
| Python | Language of choice |
| Gradio | Frontend UI |
| LangChain | Prompt chaining + LLM handling |
| HuggingFace BGE | Embeddings for vector search |
| FAISS | Vector DB |
| LM Studio | Local LLM hosting |
| Gemini Flash | Evaluator + Feedback Writer |
| tiktoken | Token-aware truncation |
+-----------------+--------------------------------+
📐 Architecture

Flow diagram (Left), HLD (Right)
Data flow within the RAG pipeline, illustrating how user queries are processed and refined through embedding search, LLM generation, and Gemini evaluation.
Gemini reads the assistant’s personality (THOTH_PERSONA). It checks for hallucinations, tone, clarity, and factual grounding. If it rejects an answer, it rewrites it and sends back a better response.
Everything happens silently. The user only sees the final output.
⚙️ Core Functional Units
| Component | Function |
| -------------------------------- | -------------------------------------------------------------------------- |
| `answer_query()` | Orchestrates end-to-end flow with retries, logging, and evaluator guard |
| `evaluate_with_gemini_flash()` | Evaluates output against persona, dynamically sourced from `THOTH_PERSONA` |
| `get_feedback_and_rewrite()` | Rewrites using Gemini if response violates persona or hallucination rules |
| `perform_ingestion_check()` | Auto-refreshes FAISS index by hashing PDF/txt documents |
| `truncate_context()` | Trims context for LLM input respecting `MAX_TOKENS_FOR_CONTEXT` |
| `query_lm_studio()` | Sends context + question to LM Studio endpoint |
| `setup_llm_chain()` | Uses LangChain with OpenAI API and persona injected as system prompt |
| `start_periodic_faiss_monitor()` | Runs periodic background thread to monitor doc changes |
⚙️ Setting It Up
✅ Prerequisites
- Python 3.10+
- Use UV to install dependencies: uv install .
- .env file with: OPENAI_API_KEY, GOOGLE_API_KEY, CHECK_INTERVAL_MINUTES=5 (default)
- Optional: LM Studio running on port 1234
📂 Folder Structure in the Repository
Some files are part of .gitignore file and hence providing the file structure for reference.

faiss_index is recreated as needed; fill the docs folder with PDF or text files.
.env file is structured as below. Keep your keys private and never commit to repository.
OPENAI_API_KEY=sk-proj-a*<your_key>
GOOGLE_API_KEY=AI*<your_key>
🏁 Run the Bot
python main.py
📂 Add Files To
/docs/*.pdf or *.txt

We can see the timer job is picking up the new files and updating the database automatically.
The script will auto-detect changes, refresh the FAISS index, and go live with new content—no restart needed. If you are using UV as I did, the setup is automated. Else, ensure you have Git installed for cloning the repository. Familiarity with virtual environments is recommended to manage dependencies.
🧠 Persona Control
The magic of this setup is the THOTH_PERSONA string. Define your assistant in natural language, like:
THOTH_PERSONA = """
You are Sherlock Holmes -- analytical, precise, unemotional. You do not guess. You deduce.
Respond strictly using provided context. If unsure, state so with calm clarity.
"""
I’ve added a few more personalities to the repository. I also plan to improve how personalities are handled and add a few more features.
Even Gemini’s evaluation uses this personality to check tone, format, and logic.
Want to switch to a historian or SRE mentor? Just change the prompt. Here’s how you incorporate it into your prompt:
prompt = f"{THOTH_PERSONA}\n\n{question}" # Incorporating persona into prompt
🔁 The Self-Correction Loop
Every response goes through:
- LLM Generation
- Gemini Evaluation
- If unsatisfactory → Rewrite
- Try again (up to 5 attempts)

Typical evaluation by Gemini of the response provided by LLM
Only the final, validated response is shown to the user. All evaluation logs stay in the console. Gemini provides a textual critique of the LLM’s response, highlighting areas for improvement such as factual inaccuracies, tone inconsistencies, or lack of clarity. This feedback is used to guide the rewrite process, often involving rephrasing sentences, adding context, or correcting errors. These logs provide valuable insights into common errors and areas where the system can be further refined.
💡 Use Cases
- Chat with policy or SRE documents (this was my primary use case, although the project later moved in a different direction)
- Turn your research papers into a conversational assistant
- Teach kids with an assistant that adapts tone
- Build compliant bots that stay within legal/technical bounds. For example, a legal bot could be configured to avoid generating responses that contain specific disclaimers or confidential information.
🔒 Why This Matters
In an age of hallucinations and brittle bots, this project proves that:
✅ Feedback matters—AI should check its own work
✅ You can run it locally with no cloud lock-in. The evaluator is optional and controllable. At the time of publication, I used Gemini 2.0 Flash because it was free.
✅ Personality is programmable—not hardcoded
✅ It’s fun—and surprisingly powerful. By incorporating self-evaluation and feedback loops, this approach aims to reduce hallucination rates and improve the reliability of AI assistants in production environments.
🚀 Try It Yourself
This is a project anyone can build and customize in a weekend. Whether you’re an SRE, a researcher, or just AI-curious—this one’s for you.
📎 Resources
- Code at GitHub link
- Read about UV here
- Read about LM Studio here
- LM Studio running model Nous Hermes 2—Solar 10.7B—read details here
- Gemini 2.0 flash details here
- Digital art generated using ComfyUI. Read about it here
- Embedder model details here
Originally published on Medium on July 26, 2025.