From 8c20d6b82957fa11dd83e0d3de6413611b69eb7e Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?=D0=9A=D0=B8=D1=80=D0=B8=D0=BB=D0=BB=20=D0=A0=D0=BE=D0=BC?= =?UTF-8?q?=D0=B0=D0=BD=D0=BE=D0=B2?= Date: Tue, 2 Jun 2026 07:17:32 +0000 Subject: [PATCH] Update README.md --- README.md | 93 ++++++++++++++++++++++++++++++++++--------------------- 1 file changed, 57 insertions(+), 36 deletions(-) diff --git a/README.md b/README.md index 9373fb5..2c1445f 100644 --- a/README.md +++ b/README.md @@ -1,61 +1,82 @@ # RAG Agent with ChromaDB and Web Search -This repository implements a simple RAG (Retrieval‑Augmented Generation) agent that can answer questions using a local knowledge base stored in **ChromaDB** and also perform real‑time web search via **Tavily**. The agent automatically decides which source to use and reports the chosen source in the answer. +This repository implements a simple RAG (Retrieval-Augmented Generation) agent that can: -## Features +1. Search a local knowledge base stored in **ChromaDB** using semantic embeddings from **Ollama**. +2. Perform real‑time web search via **Tavily**. +3. Decide automatically which source to use and indicate the source in the final answer. -- **Local knowledge base** – Text files (.txt, .md) are loaded, chunked, and stored in a persistent ChromaDB collection. -- **Semantic search** – Uses Ollama embeddings (`nomic-embed-text`). -- **Web search** – Powered by Tavily. -- **Automatic source selection** – The agent chooses between local and web search based on the query. -- **CLI** – Simple chat loop with `exit` to quit. +## Prerequisites -## Setup +- Python 3.10+ (recommended via `pyenv` or `conda`). +- Ollama installed locally and the following models pulled: + ```bash + ollama pull llama3 + ollama pull nomic-embed-text + ``` +- A Tavily API key. Create a `.env` file in the project root with: + ```text + TAVILY_API_KEY=YOUR_KEY_HERE + ``` + +## Installation ```bash -# 1. Create a virtual environment (optional but recommended) +# Optional: create a virtual environment python -m venv venv source venv/bin/activate # Windows: venv\Scripts\activate -# 2. Install dependencies +# Install dependencies pip install -r requirements.txt - -# 3. Pull required Ollama models -ollama pull llama3 -ollama pull nomic-embed-text - -# 4. Set your Tavily API key -export TAVILY_API_KEY=YOUR_KEY # Windows: set TAVILY_API_KEY=YOUR_KEY ``` -## Usage +## Preparing the Knowledge Base -1. **Load documents** – Place your `.txt` or `.md` files in the `documents/` folder. -2. **Run the agent** - ```bash - python agent.py - ``` -3. **Chat** – Type your question. Type `exit` to quit. +Place any `.txt` or `.md` files you want the agent to know about in the `documents/` folder. +Run the following command once to load them into ChromaDB: -## Example +```bash +python -c "from vectorstore import create_vectorstore, load_documents; store=create_vectorstore(); load_documents('./documents', store)" +``` + +The vector store is persisted in the `chroma_db/` directory, so the data will be available for subsequent runs. + +## Running the Agent + +```bash +python main.py +``` + +You will see a simple chat loop. Type your questions and the agent will answer. ``` -Query: Какие последние новости про AI-агентов? -[Web Search] ... -Источник: tavily +Welcome to the RAG agent. Type 'exit' to quit. -Query: Что в наших конспектах про LangGraph? -[Local KB] ... -Источник: chromadb +User: What is LangGraph? + +Assistant: LangGraph is a framework for building ... +Source: chromadb ``` +If the information is not present locally, the agent will automatically perform a web search and label the answer with `Source: tavily`. + ## Project Structure -- `vectorstore.py` – Functions to create and load the ChromaDB vector store. -- `rag_tools.py` – Two LangChain tools: `search_local_kb` and `web_search`. -- `agent.py` – Main script that sets up the agent and runs the chat loop. -- `requirements.txt` – Python dependencies. -- `README.md` – This file. +``` +├── agent.py # Core agent logic +├── rag_tools.py # Tool implementations +├── vectorstore.py # ChromaDB utilities +├── main.py # Entry point +├── requirements.txt +├── .gitignore +└── README.md +``` + +## Extending + +- Add more tools by creating new functions decorated with `@tool`. +- Replace the LLM or embeddings with other Ollama models. +- Switch to a different vector store (e.g., Qdrant) by updating `vectorstore.py`. ## License