Update README.md

This commit is contained in:
RK-A committed 2026-06-02 07:11:07 +00:00
1 parent ccf74a120f
commit 6f89ea202a
1 file changed
+38 -42
+38 -42
View File
@@ -1,66 +1,62 @@
# RAG Agent with ChromaDB and Tavily
# RAG Agent with ChromaDB and Web Search
## Overview
This repository implements a **RAG (Retrieval‑Augmented Generation) agent** that can answer questions by searching a local knowledge base stored in **ChromaDB** or by fetching up‑to‑date information from the web using **Tavily**. The agent automatically chooses the most appropriate source based on the query and returns the answer together with the source identifier.
This repository implements a simple RAG (Retrieval‑Augmented Generation) agent that can answer questions using a local knowledge base stored in **ChromaDB** and also perform real‑time web search via **Tavily**. The agent automatically decides which source to use and reports the chosen source in the answer.
## Features
- **Local Knowledge Base** – Vector store backed by ChromaDB with embeddings from Ollama (`nomic-embed-text`).
- **Web Search** – Uses Tavily API for real‑time web queries.
- **Automatic Routing** – The agent decides whether to use the local KB or the web search.
- **CLI** – Simple command‑line interface for interactive queries.
- **Persistence** – The ChromaDB store is persisted between runs.
- **Local knowledge base** – Text files (.txt, .md) are loaded, chunked, and stored in a persistent ChromaDB collection.
- **Semantic search** – Uses Ollama embeddings (`nomic-embed-text`).
- **Web search** – Powered by Tavily.
- **Automatic source selection** – The agent chooses between local and web search based on the query.
- **CLI** – Simple chat loop with `exit` to quit.
## Setup
1. **Install Ollama** and pull the required models:
```bash
ollama pull llama3
ollama pull nomic-embed-text
```
```bash
# 1. Create a virtual environment (optional but recommended)
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
2. **Install Python dependencies**:
```bash
pip install -r requirements.txt
```
# 2. Install dependencies
pip install -r requirements.txt
3. **Set up the Tavily API key**. Create a `.env` file in the project root with:
```env
TAVILY_API_KEY=your_api_key_here
```
# 3. Pull required Ollama models
ollama pull llama3
ollama pull nomic-embed-text
4. **Add documents** you want to index into the `documents/` folder. The script will automatically load `.txt` and `.md` files.
# 4. Set your Tavily API key
export TAVILY_API_KEY=YOUR_KEY # Windows: set TAVILY_API_KEY=YOUR_KEY
```
## Usage
```bash
python main.py
```
1. **Load documents** – Place your `.txt` or `.md` files in the `documents/` folder.
2. **Run the agent**
```bash
python agent.py
```
3. **Chat** – Type your question. Type `exit` to quit.
You will be prompted for a query. Type `exit` to quit.
## Example
Example:
```
Query: Какие последние новости про AI-агентов?
Answer:
[Web Search]
1. AI Agents are ...
https://example.com
...
Source: tavily
[Web Search] ...
Источник: tavily
Query: Что в наших конспектах про LangGraph?
[Local KB] ...
Источник: chromadb
```
## Project Structure
- `vectorstore.py` – Helper functions for creating and populating the ChromaDB vector store.
- `agent.py` – Defines the tools and initializes the LangChain agent.
- `main.py` – CLI entry point.
- `vectorstore.py` – Functions to create and load the ChromaDB vector store.
- `rag_tools.py` – Two LangChain tools: `search_local_kb` and `web_search`.
- `agent.py` – Main script that sets up the agent and runs the chat loop.
- `requirements.txt` – Python dependencies.
- `README.md` – Documentation.
- `README.md` – This file.
## Notes
## License
- The agent uses the **Zero‑Shot React** strategy. It may call both tools if the query is ambiguous. You can tweak the prompt or the routing logic if needed.
- The ChromaDB store is persisted in `./chroma_db`. Delete this folder to re‑index.
- Ensure the Ollama server is running locally when executing the agent.
MIT License.