54 lines
2.4 KiB
Markdown
54 lines
2.4 KiB
Markdown
**What was implemented**
|
||
The script `src/index.py` now uses **ChromaDB** as the persistent vector store instead of Qdrant.
|
||
It loads documents from a folder, splits them into chunks, embeds them with OpenAI embeddings, and stores the vectors in a Chroma collection.
|
||
A Retrieval‑QA chain is built with LangChain’s `RetrievalQA` and OpenAI’s GPT model, and a lightweight web‑search tool (`DuckDuckGoSearchRun`) is kept for quick queries.
|
||
|
||
**Why the main parts satisfy the assignment**
|
||
* The vector database is explicitly ChromaDB – the `initialize_vectorstore()` function creates a `chromadb.PersistentClient` and wraps it with LangChain’s `Chroma` wrapper.
|
||
* All required stack components are present: `chromadb`, `langchain`, `openai`, and `python-dotenv`.
|
||
* The agent can ingest, query, and perform web search, matching the functional requirements of the exam task.
|
||
|
||
**Key code excerpts**
|
||
|
||
`src/index.py` – imports and vector store initialization
|
||
```python
|
||
import chromadb
|
||
from langchain.embeddings import OpenAIEmbeddings
|
||
from langchain.vectorstores import Chroma
|
||
...
|
||
def initialize_vectorstore() -> Chroma:
|
||
client = chromadb.PersistentClient(path=CHROMA_DB_PATH)
|
||
client.get_or_create_collection(name=COLLECTION_NAME)
|
||
vectorstore = Chroma(
|
||
client=client,
|
||
collection_name=COLLECTION_NAME,
|
||
embedding_function=OpenAIEmbeddings(model=EMBEDDING_MODEL),
|
||
)
|
||
return vectorstore
|
||
```
|
||
|
||
`src/index.py` – ingesting documents into Chroma
|
||
```python
|
||
def ingest_documents(folder_path: str, vectorstore: Chroma) -> None:
|
||
raw_texts = load_documents_from_folder(folder_path)
|
||
chunks = split_text(raw_texts)
|
||
vectorstore.add_texts(chunks)
|
||
print(f"Ingested {len(chunks)} chunks into collection '{COLLECTION_NAME}'.")
|
||
```
|
||
|
||
`src/index.py` – web‑search helper
|
||
```python
|
||
def perform_web_search(query: str) -> List[Dict[str, str]]:
|
||
search_tool = DuckDuckGoSearchRun()
|
||
results = search_tool.run(query)
|
||
if isinstance(results, list):
|
||
return results
|
||
return [{"title": "Search Result", "url": "", "body": results}]
|
||
```
|
||
|
||
**Limitations**
|
||
* No unit tests are included.
|
||
* Error handling is minimal (e.g., missing environment variables or empty folders).
|
||
* The script is single‑threaded and may not scale for very large corpora without further optimization.
|
||
|
||
Overall, the implementation now adheres to the required stack and fulfills the RAG agent functionality described in the assignment. |