add README.md

This commit is contained in:
2026-05-28 09:27:58 +00:00
parent 975b2c9222
commit ae37539085
+41 -35
View File
@@ -1,46 +1,52 @@
# Агент с RAG‑памятью
# RAG Agent with ChromaDB
## Что делает проект
Реализован AI‑агент, использующий локальное векторное хранилище ChromaDB и Ollama для эмбеддингов. Агент умеет добавлять документы (разбивая их на чанки) и выполнять семантический поиск.
## Overview
This repository implements a simple RAG (RetrievalAugmented Generation) agent that uses **ChromaDB** as the vector store and **Ollama** for embeddings and LLM inference. The agent can:
## Структура файлов
| Файл | Назначение |
|------|------------|
| `main.py` | Точка входа, демонстрация работы агента. |
| `requirements.txt` | Пакеты‑зависимости. |
| `README.md` | Это описание. |
| `chunker.py` | Разбиение текста на чанки (500/100). |
| `vector_store.py` | Работа с ChromaDB: добавление и поиск. |
| `tools.py` | Декорированные инструменты RAG‑агента. |
| `agent.py` | Создание агента через `create_agent`. |
| `cli.py` | CLI для интерактивного теста. |
| `init_documents.py` | Скрипт загрузки всех файлов из каталога в базу. |
1. Add documents to the knowledge base.
2. Search the knowledge base semantically.
3. Interact via a lightweight CLI.
## Установка
## File Structure
- `requirements.txt` Python dependencies.
- `chunker.py` Text chunking utilities (RecursiveCharacterTextSplitter).
- `vector_store.py` Wrapper around ChromaDB collection.
- `tools.py` LangChain tools for search and add operations.
- `agent.py` Agent creation with LangChain `create_agent`.
- `cli.py` Simple commandline interface.
- `init_documents.py` Helper to load all `.txt` files from a directory into the vector store.
## Installation
```bash
# Pull Ollama models
ollama pull llama3
ollama pull nomic-embed-text
# Install Python packages
pip install -r requirements.txt
# Ollama: pull nomic-embed-text (для эмбеддингов)
```
## Примеры использования
## Usage
1. **Load documents** (optional):
```bash
python cli.py add path/to/file.txt # Добавить документ
python cli.py search "Python" # Поиск по базе
python init_documents.py
```
## Архитектура
1. **LLM** `ChatOpenAI` через BroJS, но для эмбеддингов используется Ollama.
2. **Embeddings** `OllamaEmbeddings(model="nomic-embed-text")`.
3. **Vector store** – ChromaDB коллекция «rag_collection» в папке chromadb.
4. **Chunking** `RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=100)`.
5. **Tools** `search_knowledge_base`, `add_to_knowledge_base` с декоратором `@tool`.
6. **Agent** `create_agent` с системным промптом и доступом к инструментам.
7. **CLI** – простая команда‑интерфейс для добавления/поиска.
---
### Тесты
Для проверки можно запустить:
2. **Run the CLI**:
```bash
python cli.py add sample.txt
python cli.py search "Python"
python cli.py
```
Type any question or `/quit` to exit.
## Architecture
- **Chunking**: `chunker.split_text()` splits large texts into 500char chunks with 100char overlap using LangChains `RecursiveCharacterTextSplitter`.
- **Vector Store**: `vector_store.ChromaVectorStore` handles adding documents and similarity search. Embeddings are generated by `langchain_ollama.OllamaEmbeddings` (`nomic-embed-text`).
- **Tools**: Two tools decorated with `@tool`: `search_knowledge_base` and `add_to_knowledge_base`. They interact with the vector store.
- **Agent**: Created via LangChains `create_agent`, configured to use the two tools and a simple system prompt. The LLM is an Ollama `llama3` instance.
## Testing
Run the CLI and try:
```
/quit
Hello, what can you do?
```
The agent should respond using the knowledge base or add new documents if prompted.