Files
povtornyy-ekzamen-faq-bot-c…/README.md
T

119 lines
3.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# FAQ Bot ChromaDB + MCP-style Tool
This project implements a simple FAQ bot that answers questions about a machine learning course.
The bot uses:
- **ChromaDB** to store and retrieve FAQ documents.
- **Ollama** embeddings (`nomic-embed-text`) for vectorization.
- **LangChain** to build an agent that routes queries to the appropriate tool.
- **MCP-style HTTP tool** (`fetch_course_meta`) that returns course metadata from a local JSON file.
## Project Structure
```
.
├── chroma_faq/ # Persisted Chroma vector store
├── data/
│ ├── faq1.md
│ ├── faq2.md
│ ├── faq3.md
│ └── course_meta.json
├── src/
│ ├── __init__.py
│ ├── agent.py
│ ├── cli.py
│ ├── main.py
│ └── tools.py
├── requirements.txt
└── README.md
```
## Setup
1. **Install Ollama**
Download and install Ollama from https://ollama.ai/.
Pull the required models:
```bash
ollama pull nomic-embed-text
ollama pull llama3
```
2. **Create a virtual environment** (optional but recommended):
```bash
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
```
3. **Install Python dependencies**:
```bash
pip install -r requirements.txt
```
## Running the Bot
### Preset Questions
Run the script without arguments to execute three preset questions (two for the FAQ tool, one for the metadata tool):
```bash
python -m src.main
```
You should see output similar to:
```
Running preset questions:
Q1: What is the deadline for Assignment 1?
A1: The deadline for Assignment 1 is August 31, 2026. source: chroma
Q2: How many lectures are there in the course?
A2: There are 12 lectures in the course. source: chroma
Q3: What is the course schedule for next week?
A3: The course schedule for next week is:
- 2026-09-01: Lecture 1 Introduction to ML (Room 101)
- 2026-09-08: Lecture 2 Data Preprocessing (Room 102)
- 2026-09-15: Lecture 3 Linear Regression (Room 103)
source: mcp_meta
```
### Interactive Mode
Start an interactive session:
```bash
python -m src.main --interactive
```
You can type any question, and the bot will answer using the appropriate tool. Type `exit` or `Ctrl+C` to quit.
## How It Works
1. **Data Loading**
`src/tools.py` contains `load_faq_to_chroma()` which reads all `.md` files in `data/`, chunks them, embeds them with `nomic-embed-text`, and persists the vector store in `chroma_faq/`.
2. **Tools**
- `search_course_docs(query, k)` searches the Chroma vector store for relevant FAQ snippets.
- `fetch_course_meta(query)` reads `data/course_meta.json` and returns schedule or instructor information based on the query.
3. **Agent**
`src/agent.py` builds a LangChain agent that:
- Uses a system prompt to decide which tool to call.
- Adds a `source:` tag to the final answer indicating whether the answer came from the FAQ (`chroma`) or the metadata tool (`mcp_meta`).
4. **CLI**
`src/cli.py` provides a simple commandline interface to run preset questions or an interactive session.
## Extending the Bot
- **Add more FAQ documents** Place additional `.md` files in `data/` and rerun the script to rebuild the vector store.
- **Add more metadata** Update `data/course_meta.json` or modify `fetch_course_meta` to call a real HTTP endpoint.
- **Change the LLM** Replace `Ollama` with another LLM provider in `src/agent.py`.
## License
This project is provided as-is for educational purposes. Feel free to adapt and extend it for your own use cases.