Skip to content
baselhusamPublic

About

Composable OCR pipelines on a visual canvas — self-hosted, typed, and runnable headless via API. Built on Docling, Paddle & Surya.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

OCRFlow

Composable OCR pipelines, fully under your control.

Drag open document-AI models onto a canvas, wire them together with typed connections,
and run the exact same graph headless through the API — on your own hardware.

License: MIT CI Python 3.11+ Next.js 16 Self-host first Status: in development

Quick start · How it works · Models · API · Docs · Contributing

An OCRFlow canvas: Load PDF → Detect Layout → Recognize Text → Structured Extract, producing invoice JSON

Why OCRFlow?

Most document-AI tooling is either a closed SaaS you send your documents to, or a pile of one-off scripts glued around a single library. OCRFlow sits in between:

  • Mix the best model for each step. Detect layout with Docling, recognize text with Surya, pull tables with Paddle, and extract JSON with a local LLM — in one graph.
  • Catch broken pipelines before they run. Every wire is typed (PageArtifact, Region[], TextLine[], JSON, …). Nodes only connect when the output actually fits the next input.
  • Prototype visually, ship headless. The graph you build on the canvas is the graph the API, batch jobs, and analytics execute. No export step, no rewrite.
  • Keep documents in your building. The default catalog is open-weight and runs on-prem — NVIDIA, AMD, Apple Silicon, or CPU — including fully air-gapped.

🚀 Quick start

You need Docker (Compose v2), make, and Git. GPU drivers are optional.

git clone https://github.com/baselhusam/OCRFlow.git
cd OCRFlow
cp backend/docker/.env.example backend/docker/.env

make up            # frontend :3000 + API :8000 (no OCR engines yet)
make db-migrate    # apply database migrations

Open http://localhost:3000, create an account (the email matching ADMIN_EMAIL becomes admin), and open a project canvas.

Then start the OCR engines you want. GPUs are detected automatically:

make detect        # what OS / GPU / compose overlay will be used
make ocr-up        # every engine — or one at a time:
make ocr-docling   # :8102
make ocr-surya     # :8101
make ocr-paddle    # :8103
make ocr-liquid    # :8104

Offline engines appear greyed out in the canvas palette and come online as soon as their service reports healthy. make help lists every target.

Note

Apple Silicon: Surya, Docling, and Liquid run on the host (PyTorch MPS) rather than in Docker, and Paddle runs in Docker on CPU. make ocr-up handles this for you. See GPU & accelerators.

Host-based development (no Docker for the app)

Requires Python 3.11+, Node.js 20+, PostgreSQL 14+, and Redis 6+.

# Backend
cd backend
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt            # + requirements-docling.txt / requirements-surya.txt
cp .env.example .env                       # DATABASE_URL, REDIS_URL, secrets, …
alembic upgrade head
uvicorn app.main:app --reload              # API on :8000

# Background worker (second terminal)
celery -A app.celery_app:celery_app worker --loglevel=info

# Frontend (third terminal)
cd frontend
npm install
cp .env.local.example .env.local
npm run dev                                # app on :3000

The same steps are wrapped as make install, make be-api, make be-worker, and make fe-dev. Full details in Installation.


🧭 How it works

                    ┌──────────────── control plane ────────────────┐
  Next.js canvas ──►│  FastAPI gateway · Celery worker              │
  REST / API keys ─►│  PostgreSQL · Redis                           │
                    └──────────────────────┬────────────────────────┘
                                           │ typed JSON over HTTP
                    ┌──────────────── data plane (opt-in) ──────────┐
                    │  Surya :8101 · Docling :8102 · Paddle :8103   │
                    │  Liquid :8104 · Ollama :11434 · hosted LLMs   │
                    └───────────────────────────────────────────────┘

make up starts only the control plane. Each OCR engine is its own microservice that loads weights lazily on first use, so you ship and run only what you need. The gateway never imports PyTorch; it forwards typed requests to whichever provider owns the model.

You work with three objects:

Object What it's for
Project A free-form canvas for experimenting. Drop in a PDF or image loader, run one node or the whole graph, preview regions and text.
Pipeline A reusable graph with a defined input and output, promoted from a project canvas. Pipelines can be nested as nodes in other canvases.
Job A pipeline run over many documents at once, with a per-document trace.

A few design rules keep graphs composable:

  • Atomic tasks. One node = one task = one endpoint. Docling's layout model and Docling's OCR are separate nodes, so you can swap either one independently.
  • Adapter, don't fork. Runners wrap upstream inference instead of reimplementing it; weights and licenses stay with the provider.
  • Collections are first-class. Run a page-level node across every page of a document, or fan a list out through an Items node.

More in How OCRFlow works.


🧩 Models

The catalog is task-level. Everything below runs today; more providers (Tesseract, RapidOCR, docTR, Florence-2, …) are listed in the registry as planned.

Provider Port What it gives you License
Loaders — PDF → pages, image → page, pick a page MIT
Docling 8102 Layout (Heron), OCR, TableFormer, figure classification & captioning, formulas, Granite-Docling VLM, full convert MIT code · Apache-2.0 weights
Surya 8101 Layout, text detection & recognition, reading order, tables, LaTeX OCR GPL-3.0 code · OpenRAIL-M weights ⚠️
Paddle 8103 DocLayout, PP-OCR text recognition, PP-Structure Apache-2.0
Liquid AI 8104 LFM2.5-VL-1.6B: prompted vision and schema-validated extraction LFM Open License v1.0
Ollama 11434 Local Qwen3 0.6B / Qwen3.5 0.8B for text and vision prompts and structured JSON Apache-2.0
Connected LLM / VLM — OpenAI, Anthropic, or any compatible endpoint for prompts and structured extraction Provider terms

Structured-extraction nodes take a schema you build field by field (or start from an Invoice / Receipt preset), and every output is validated before it moves downstream.

Warning

Surya's code is GPL-3.0. Check license compatibility before shipping it inside a proprietary product.

Full list with IDs, inputs, and outputs: Model catalog · backend/docs/MODEL_CATALOG.md.


🔌 Use it from code

Anything built on the canvas runs headless. Admins and Developer users can create API keys in Account & settings → API keys, then send documents to any pipeline:

import requests

BASE = "http://localhost:8000/api/v1"
headers = {"X-API-Key": "ocrflow_..."}  # read this from an env var in real code

pipeline_id = requests.get(f"{BASE}/developer/pipelines", headers=headers).json()["items"][0]["id"]

with open("invoice.pdf", "rb") as f:
    queued = requests.post(
        f"{BASE}/developer/pipelines/{pipeline_id}/documents",
        headers=headers,
        files=[("files", ("invoice.pdf", f, "application/pdf"))],
        data={"output_format": "json"},
    )
run_id = queued.json()["runs"][0]["id"]

result = requests.get(f"{BASE}/developer/pipelines/{pipeline_id}/runs/{run_id}", headers=headers).json()

Uploads accept up to 50 files per request. Interactive OpenAPI docs are served at http://localhost:8000/docs, and the full route list is in the API reference.


📚 Documentation

The full docs ship inside the app at http://localhost:3000/documentation (press ⌘K to search). The sources live in frontend/src/content/docs/:

Get started Quick start · Installation · GPU & accelerators · Air-gapped deploy
Concepts How it works · Projects · Pipelines · Jobs · Input & output
Guides Canvas · Connecting nodes · Connect models · Running jobs · Analytics · Admin
Reference Make commands · API · Environment variables · Keyboard shortcuts

🛠️ Development

make test          # backend pytest (non-GPU) + frontend vitest
make fe-lint       # eslint
make logs          # follow all container logs
make ocr-ps        # which OCR engines are up
make down          # stop the stack (make ocr-down for engines)
Repository layout
OCRFlow/
├── backend/                 FastAPI gateway, model runners, Celery worker
│   ├── app/                   API, model registry, runners, schemas, db
│   ├── alembic/               database migrations
│   ├── docker/                Dockerfiles, compose overlays, env example
│   ├── docs/                  model catalog, containerized serving, air-gapped notes
│   ├── tests/                 pytest suite
│   └── requirements*.txt      core / dev / docling / surya / paddle / liquid
├── frontend/                Next.js 16 app (React 19, React Flow, Tailwind v4, shadcn/ui)
│   └── src/content/docs/      in-app documentation
├── branding/                logos, brand guidelines, design system
├── scripts/                 accelerator detection, host OCR runner
├── docker-compose.yml       full-stack compose entrypoint
└── Makefile                 every dev and ops command (make help)
Layer Stack
Frontend Next.js 16, React 19, TypeScript, React Flow, Tailwind CSS v4, shadcn/ui, Framer Motion, Recharts
Backend FastAPI, SQLAlchemy 2 (async), Alembic, Pydantic, Celery
Data PostgreSQL, Redis
Inference Docling, Surya, PaddleOCR, Transformers, Ollama — each in its own service

🗺️ Roadmap

  • Typed canvas with project → pipeline → job promotion
  • Docling, Surya, Paddle, Liquid, Ollama, and hosted LLM/VLM providers
  • Auto-detected NVIDIA / AMD / Apple GPU serving and air-gapped deploys
  • Developer API keys, per-key usage, workspace and admin analytics
  • Collections: apply-to-all-pages, Items fan-out, structured schema builder
  • Server-side map and fan-in for collections
  • More standalone engines (Tesseract, RapidOCR, docTR, Florence-2, …)
  • Export nodes and pipeline presets

🤝 Contributing

Contributions are welcome. Read CONTRIBUTING.md for setup and conventions, and open an issue before large changes. New models follow the per-task checklist in backend/docs/MODEL_CATALOG.md: each task gets its own endpoint, schemas, validation, and tests.

Please follow the Code of Conduct, and report vulnerabilities privately as described in SECURITY.md.


📄 License

OCRFlow is released under the MIT License. Model weights and upstream libraries keep their own licenses — see the model table.


Built by Basel Husam · Brand and design system in branding/

About

Composable OCR pipelines on a visual canvas — self-hosted, typed, and runnable headless via API. Built on Docling, Paddle & Surya.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages