Files
povtornyy-ekzamen-faq-bot-c…/README.md
T

118 lines
3.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# FAQ Bot QDrant Vector Store
This project implements a simple FAQ bot that uses **QDrant** as the vector store instead of ChromaDB.
The bot can ingest a text file containing FAQ content, embed the text using OpenAI embeddings, store the embeddings in QDrant, and answer user questions by retrieving the most relevant passages.
## Features
- **Vector Store** QDrant (via `qdrant-client`)
- **Embeddings** OpenAI `text-embedding-ada-002`
- **CLI** Ingest data, query the bot, delete the collection
- **API** `get_response(question: str, top_k: int = 5)` for integration with tools like MCP-tool
## Prerequisites
- Python 3.9+
- QDrant server running locally or accessible via network
- OpenAI API key
## Setup
1. **Clone the repository**
```bash
git clone https://git.brojs.ru/kuzakhmetovartur/povtornyy-ekzamen-faq-bot-chromadb-odin.git
cd povtornyy-ekzamen-faq-bot-chromadb-odin
```
2. **Create a virtual environment (optional but recommended)**
```bash
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
```
3. **Install dependencies**
```bash
pip install -r requirements.txt
```
4. **Set environment variables**
Create a `.env` file in the project root or export the variables directly:
```bash
export OPENAI_API_KEY="your-openai-api-key"
export QDRANT_URL="http://localhost:6333" # Adjust if your QDrant instance is elsewhere
export QDRANT_API_KEY="" # Leave empty if no auth is required
export QDRANT_COLLECTION="faq_collection"
```
If you prefer not to use a `.env` file, you can set the variables in your shell session.
## Usage
### 1. Ingest Data
Prepare a plain text file (`faq.txt`) containing your FAQ content. Then run:
```bash
python src/index.py ingest faq.txt
```
The script will:
- Split the text into chunks (max 500 characters per chunk)
- Generate embeddings for each chunk
- Store the embeddings in QDrant under the collection name defined by `QDRANT_COLLECTION`
### 2. Query the Bot
```bash
python src/index.py query "What is the return policy?"
```
You can adjust the number of results returned with `--top_k`:
```bash
python src.index.py query "What is the return policy?" --top_k 3
```
### 3. Delete the Collection
> **Warning:** This will permanently delete all data in the collection.
```bash
python src/index.py delete
```
### 4. Integration via API
If you want to use the bot programmatically (e.g., from MCP-tool), import the `get_response` function:
```python
from src.index import get_response
answer = get_response("How do I reset my password?")
print(answer)
```
## Troubleshooting
- **QDrant Connection Errors**
Ensure the QDrant server is running and reachable at the URL specified by `QDRANT_URL`. Check firewall settings if accessing remotely.
- **OpenAI API Errors**
Verify that `OPENAI_API_KEY` is correct and has sufficient quota. Check the OpenAI dashboard for usage limits.
- **Large Documents**
The ingestion script splits documents into 500character chunks. Adjust `max_chunk_size` in `split_text_into_chunks` if you need larger or smaller chunks.
## License
This project is provided under the MIT License. Feel free to modify and extend it for your own use cases.
## Contact
For questions or support, contact Artur Kuzakhmetov at `artur@example.com`.