Add README.md
This commit is contained in:
@@ -0,0 +1,83 @@
|
|||||||
|
# Text Extractor CLI
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
This project provides a command‑line tool that extracts structured data from free‑form text using **LangChain** and **Pydantic**. It supports two schemas:
|
||||||
|
|
||||||
|
1. **PersonInfo** – name, age, profession, skills.
|
||||||
|
2. **MeetingNotes** – title, date, participants, agenda.
|
||||||
|
|
||||||
|
The tool automatically detects which schema to use based on the input text.
|
||||||
|
|
||||||
|
## Installation
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Create a virtual environment
|
||||||
|
python -m venv .venv
|
||||||
|
source .venv/bin/activate # On Windows use .venv\Scripts\activate
|
||||||
|
|
||||||
|
# Install dependencies
|
||||||
|
pip install -r requirements.txt
|
||||||
|
```
|
||||||
|
|
||||||
|
## Environment variables
|
||||||
|
|
||||||
|
The tool uses OpenAI’s API. Create a `.env` file in the project root with the following variables:
|
||||||
|
|
||||||
|
```
|
||||||
|
OPENAI_API_KEY=your_api_key_here
|
||||||
|
OPENAI_MODEL=gpt-3.5-turbo # optional, defaults to gpt-3.5-turbo
|
||||||
|
OPENAI_BASE_URL= # optional, for local LLMs (e.g. http://localhost:11434/v1)
|
||||||
|
```
|
||||||
|
|
||||||
|
If you are using a local model (Ollama/LM Studio) provide the `OPENAI_BASE_URL` and set `OPENAI_API_KEY` to any non‑empty string (e.g. `ollama`).
|
||||||
|
|
||||||
|
## Usage
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Pass text as a command‑line argument
|
||||||
|
python agent.py "Anna, 28 years old, Python developer. Skills: FastAPI, Docker."
|
||||||
|
|
||||||
|
# Or pipe text via stdin
|
||||||
|
cat <<EOF | python agent.py
|
||||||
|
Meeting: Sprint Planning
|
||||||
|
Date: 2024-06-01
|
||||||
|
Participants: Alice, Bob, Charlie
|
||||||
|
Agenda: Review backlog, assign tasks, estimate effort
|
||||||
|
EOF
|
||||||
|
```
|
||||||
|
|
||||||
|
### Output
|
||||||
|
|
||||||
|
The tool prints two sections:
|
||||||
|
|
||||||
|
1. **JSON** – a pretty‑printed JSON representation of the Pydantic model.
|
||||||
|
2. **Summary** – a human‑readable summary of the key fields.
|
||||||
|
|
||||||
|
Example output for a person:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"name": "Anna",
|
||||||
|
"age": 28,
|
||||||
|
"profession": "Python developer",
|
||||||
|
"skills": [
|
||||||
|
"FastAPI",
|
||||||
|
"Docker"
|
||||||
|
]
|
||||||
|
}
|
||||||
|
|
||||||
|
Summary:
|
||||||
|
Anna, 28 years old, works as Python developer.
|
||||||
|
Skills: FastAPI, Docker.
|
||||||
|
```
|
||||||
|
|
||||||
|
## Extending
|
||||||
|
|
||||||
|
To add a new schema:
|
||||||
|
1. Define a new `Pydantic` model.
|
||||||
|
2. Create a prompt and parser similar to the existing ones.
|
||||||
|
3. Update the `detect_schema` logic or add a new classifier.
|
||||||
|
|
||||||
|
## License
|
||||||
|
MIT
|
||||||
Reference in New Issue
Block a user