3b760b5103975ecdd148ad62f854693e33c468dc
Text Extractor CLI
Overview
This project provides a command‑line tool that extracts structured data from free‑form text using LangChain and Pydantic. It supports two schemas:
- PersonInfo – name, age, profession, skills.
- MeetingNotes – title, date, participants, agenda.
The tool automatically detects which schema to use based on the input text.
Installation
# Create a virtual environment
python -m venv .venv
source .venv/bin/activate # On Windows use .venv\Scripts\activate
# Install dependencies
pip install -r requirements.txt
Environment variables
The tool uses OpenAI’s API. Create a .env file in the project root with the following variables:
OPENAI_API_KEY=your_api_key_here
OPENAI_MODEL=gpt-3.5-turbo # optional, defaults to gpt-3.5-turbo
OPENAI_BASE_URL= # optional, for local LLMs (e.g. http://localhost:11434/v1)
If you are using a local model (Ollama/LM Studio) provide the OPENAI_BASE_URL and set OPENAI_API_KEY to any non‑empty string (e.g. ollama).
Usage
# Pass text as a command‑line argument
python agent.py "Anna, 28 years old, Python developer. Skills: FastAPI, Docker."
# Or pipe text via stdin
cat <<EOF | python agent.py
Meeting: Sprint Planning
Date: 2024-06-01
Participants: Alice, Bob, Charlie
Agenda: Review backlog, assign tasks, estimate effort
EOF
Output
The tool prints two sections:
- JSON – a pretty‑printed JSON representation of the Pydantic model.
- Summary – a human‑readable summary of the key fields.
Example output for a person:
{
"name": "Anna",
"age": 28,
"profession": "Python developer",
"skills": [
"FastAPI",
"Docker"
]
}
Summary:
Anna, 28 years old, works as Python developer.
Skills: FastAPI, Docker.
Extending
To add a new schema:
- Define a new
Pydanticmodel. - Create a prompt and parser similar to the existing ones.
- Update the
detect_schemalogic or add a new classifier.
License
MIT
Description
Languages
Python
100%