vLLM Processing for Unstructured Historical Documents
-
Updated
Jun 22, 2025 - Python
vLLM Processing for Unstructured Historical Documents
Modular OCR pipeline for live-camera documents/screens. Auto-aligns perspective, blocks obstructions, and extracts text using interchangeable deep-learning models.
Optical Character Recognition, OCR pipeline, Arabic OCR, Deep Learning OCR, Computer Vision text extraction, Text recognition system, AI document processing, Multilingual OCR, Transformer OCR, OCR benchmarking, Bounding box detection, Ground truth evaluation.
🧠 AI-powered pipeline for cleaning scanned documents. Removes noise, enhances text, auto-tunes model weights, and returns OCR-optimized PDFs via CLI or cloud API.
Serverless OCR & PDF Text Extraction microservice for Personal AI Factory v1. Built with TypeScript and Vercel Serverless Functions, using pdf-parse, and node-fetch for high-performance parsing of machine-readable PDFs. Supports extracting clean text from textual PDFs and exposes a clean HTTP API returning structured JSON output for downstream n8n.
Arabic document OCR parser and structured data extractor with ZATCA QR decoding in pure PHP
Serverless OCR & PDF Text Extraction microservice for Personal AI Factory v1. Built with TypeScript and Vercel Serverless Functions, using pdf-parse, and node-fetch for high-performance parsing of machine-readable PDFs. Supports extracting clean text from textual PDFs and exposes a clean HTTP API returning structured JSON output for downstream n8n.
Webtoon translation workflow built with OCR and LLMs for people who need practical localization tools.
AI-powered document extraction system using OCR and intelligent field extraction.
High-performance RAG API with AI, multi-format docs, Gemini integration, security, CLI.
Offline-first local OCR pipeline powered by Ollama with two-tier caching and hierarchical Notion synchronization. Zero cloud API costs and zero confidential data leakage.
Composable OCR pipelines on a visual canvas — self-hosted, typed, and runnable headless via API. Built on Docling, Paddle & Surya.
To associate your repository with the ocr-pipeline topic, visit your repo's landing page and select "manage topics."