A fully local document processing pipeline that ingests PDF/text files, classifies them, extracts structured fields, and supports semantic search — with no paid or hosted AI APIs.
| Capability | Description |
|---|---|
| Ingestion | Reads all .pdf and .txt files from a folder |
| Classification | Invoice, Resume, Utility Bill, Other, Unclassifiable |
| Extraction | Type-specific structured fields via regex patterns |
| Semantic Search | SentenceTransformers + FAISS local vector index |
| Interfaces | CLI (main.py) and optional Web UI (server.py) |
cd ai_document_intelligence1
python3 -m venv .venv
source .venv/bin/activate # Linux/macOS
# .venv\Scripts\activate # Windowspip install -r requirements.txtFirst run only: SentenceTransformers downloads the all-MiniLM-L6-v2 model (~90 MB) from Hugging Face. After that, everything runs offline.
Ingest all PDFs/text files, classify, extract fields, and build the search index:
python main.py process --folder ./datasetThis produces:
output.json— classifications and extracted fieldssearch_index.pkl(+.faiss) — semantic search index
python main.py search --query "payments due in January"
python main.py search --query "software engineer with 5 years experience" --top 3python server.pyOpen http://localhost:5000 to:
- Scan a folder with live progress
- Filter results by document class
- Classify a single uploaded PDF
- Export results as JSON
output.json maps each filename to its class and extracted fields:
{
"invoice_1.pdf": {
"class": "Invoice",
"invoice_number": "1001",
"date": "2025-06-16",
"company": "Pioneer Ltd",
"total_amount": 2073.0
},
"resume_1.pdf": {
"class": "Resume",
"name": "John Doe",
"email": "john.doe@example.com",
"phone": "+1-555-799-6125",
"experience_years": 5
}
}dataset/*.pdf
│
▼
┌─────────────────┐
│ Text Extraction │ pypdf (+ pdfminer fallback)
└────────┬────────┘
│
▼
┌─────────────────┐
│ Classification │ Keyword/regex scoring heuristics
└────────┬────────┘
│
▼
┌─────────────────┐
│ Field Extraction │ Regex patterns per document type
└────────┬────────┘
│
▼
┌─────────────────┐
│ Semantic Index │ SentenceTransformers → FAISS
└─────────────────┘
Documents are scored against keyword lists for each category. Filename hints (e.g. invoice_1.pdf) provide additional signal. A confidence margin threshold prevents ambiguous classifications. Empty or unreadable files become Unclassifiable.
| Class | Fields |
|---|---|
| Invoice | invoice_number, date, company, total_amount |
| Resume | name, email, phone, experience_years |
| Utility Bill | account_number, date, usage_kwh, amount_due |
| Other / Unclassifiable | class only |
- Embeddings:
all-MiniLM-L6-v2via SentenceTransformers (384-dim vectors) - Index: FAISS
IndexFlatL2for exact nearest-neighbor search - Fallback: TF-IDF + cosine similarity (scikit-learn) if FAISS/ST unavailable
| Library | Purpose |
|---|---|
| pypdf | Primary PDF text extraction |
| pdfminer.six | Fallback PDF extraction |
| SentenceTransformers | Local embedding model |
| FAISS | Vector similarity search |
| scikit-learn | TF-IDF fallback search |
| Flask | Optional web UI/API |
| NumPy | Vector operations |
ai_document_intelligence1/
├── main.py # CLI entry point
├── server.py # Optional Flask web UI
├── requirements.txt
├── output.json # Generated results
├── search_index.pkl # Generated search index
├── src/
│ ├── processor.py # Ingestion, classification, extraction
│ └── search_engine.py # Semantic search engine
├── ui/
│ └── index.html # Browser UI
└── dataset/ # Sample PDFs (20 documents)
After the initial model download, the system runs entirely offline. No OpenAI, Claude, Gemini, or other hosted AI services are used.