Экзамен: RAG-агент с ChromaDB и веб-поиском: README.md
This commit is contained in:
@@ -1,3 +1,135 @@
|
||||
# task-6a186500-ekzamen-rag-agent-s-chrom
|
||||
# RAG‑Agent with ChromaDB and Web Search
|
||||
|
||||
Решения домашних заданий
|
||||
A lightweight Retrieval‑Augmented Generation (RAG) agent that scrapes web pages, stores their embeddings in a local **Chroma** vector store, and answers user questions using OpenAI’s GPT‑3.5 Turbo.
|
||||
The project demonstrates how to combine:
|
||||
|
||||
* **BeautifulSoup** – for HTML parsing
|
||||
* **OpenAI embeddings** – to convert text into vectors
|
||||
* **ChromaDB** – fast local vector search
|
||||
* **LangChain** – orchestration of LLM, embeddings and QA chain
|
||||
|
||||
---
|
||||
|
||||
## 📦 Installation
|
||||
|
||||
```bash
|
||||
# 1. Clone the repo (or copy solution.py)
|
||||
git clone https://github.com/your‑repo/rag-agent.git
|
||||
cd rag-agent
|
||||
|
||||
# 2. Create a virtual environment (recommended)
|
||||
python -m venv .venv
|
||||
source .venv/bin/activate # Windows: .venv\Scripts\activate
|
||||
|
||||
# 3. Install dependencies
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
`requirements.txt`:
|
||||
|
||||
```text
|
||||
beautifulsoup4==4.12.3
|
||||
langchain==0.2.0
|
||||
openai==1.30.0
|
||||
chromadb==0.5.6
|
||||
rich==13.7.1
|
||||
requests==2.32.3
|
||||
```
|
||||
|
||||
> **Note**
|
||||
> You must have an OpenAI API key set in the environment:
|
||||
|
||||
```bash
|
||||
export OPENAI_API_KEY="sk-..."
|
||||
# Windows: set OPENAI_API_KEY=sk-...
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🚀 Running the Agent
|
||||
|
||||
The project contains a single script `solution.py`. It performs three actions depending on the command‑line argument:
|
||||
|
||||
| Argument | Action |
|
||||
|----------|--------|
|
||||
| `scrape` | Scrapes URLs from a file, stores embeddings in ChromaDB. |
|
||||
| `ask` | Loads the vector store and asks a question via the console. |
|
||||
| `clean` | Deletes the local Chroma database (optional). |
|
||||
|
||||
### 1. Prepare a list of URLs
|
||||
|
||||
Create a plain text file (`urls.txt`) with one URL per line, e.g.:
|
||||
|
||||
```text
|
||||
https://en.wikipedia.org/wiki/Artificial_intelligence
|
||||
https://www.python.org/
|
||||
```
|
||||
|
||||
### 2. Scrape and index the pages
|
||||
|
||||
```bash
|
||||
python solution.py scrape urls.txt
|
||||
```
|
||||
|
||||
The script will:
|
||||
|
||||
1. Download each page (respecting `robots.txt` is **not** enforced).
|
||||
2. Extract visible text with BeautifulSoup.
|
||||
3. Generate embeddings via OpenAI.
|
||||
4. Store them in `./chroma_db`.
|
||||
|
||||
### 3. Ask a question
|
||||
|
||||
```bash
|
||||
python solution.py ask
|
||||
```
|
||||
|
||||
You will be prompted to type your question. The agent retrieves the most relevant chunks from ChromaDB and feeds them into GPT‑3.5 Turbo, printing the answer.
|
||||
|
||||
> **Tip** – If you want to see the raw retrieved snippets, set `console.print(chunks)` in the script before the QA chain.
|
||||
|
||||
### 4. Clean up
|
||||
|
||||
```bash
|
||||
python solution.py clean
|
||||
```
|
||||
|
||||
Deletes the local Chroma database (`./chroma_db`).
|
||||
|
||||
---
|
||||
|
||||
## 📄 Example Session
|
||||
|
||||
```text
|
||||
$ python solution.py scrape urls.txt
|
||||
[green]Scraping and indexing 2 URLs...[/green]
|
||||
[green]Done! 3 chunks stored in chroma_db.[/green]
|
||||
|
||||
$ python solution.py ask
|
||||
Enter your question: What is the main goal of artificial intelligence?
|
||||
[cyan]Retrieving relevant documents...[/cyan]
|
||||
Answer:
|
||||
Artificial Intelligence (AI) aims to create systems that can perform tasks requiring human-like intelligence, such as learning, reasoning, problem‑solving, perception, and language understanding. The primary goal is to develop machines that can adapt, learn from experience, and make decisions in complex environments.
|
||||
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🛠️ Customization
|
||||
|
||||
| Setting | Where to change |
|
||||
|---------|-----------------|
|
||||
| Embedding model | `EMBEDDINGS_MODEL` in `solution.py` |
|
||||
| LLM model | `LLM_MODEL` in `solution.py` |
|
||||
| Chroma directory | `CHROMA_DIR` in `solution.py` |
|
||||
|
||||
Feel free to replace the OpenAI models with local alternatives (e.g., HuggingFace) by adjusting the LangChain wrappers.
|
||||
|
||||
---
|
||||
|
||||
## 🤝 Contributing
|
||||
|
||||
Pull requests are welcome!
|
||||
Please open an issue first to discuss major changes or new features.
|
||||
|
||||
---
|
||||
Reference in New Issue
Block a user