Metadata-Version: 2.5
Name: titanocr
Version: 0.1.2
Summary: General-purpose OCR + document-structure engine: any PDF/image to text + structure with pixel-level provenance
Project-URL: Homepage, https://github.com/Phantom-IN/titanocr
Project-URL: Repository, https://github.com/Phantom-IN/titanocr
Project-URL: Issues, https://github.com/Phantom-IN/titanocr/issues
Author: TitanOCR
License-Expression: MIT
License-File: LICENSE
Keywords: document-ai,layout-analysis,ocr,pdf,table-extraction
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Image Recognition
Classifier: Topic :: Text Processing
Requires-Python: >=3.13
Requires-Dist: fastapi>=0.141
Requires-Dist: httpx>=0.28
Requires-Dist: numpy>=2.5
Requires-Dist: omegaconf==2.4.0.dev15
Requires-Dist: onnxruntime>=1.28
Requires-Dist: opencv-python-headless>=5
Requires-Dist: pdf-inspector>=1.14
Requires-Dist: pillow-heif>=1.5
Requires-Dist: pillow>=12
Requires-Dist: pydantic-settings>=2.7
Requires-Dist: pydantic>=2.13
Requires-Dist: pypdfium2>=5.12
Requires-Dist: python-multipart>=0.0.32
Requires-Dist: pyyaml>=6.0
Requires-Dist: rapid-layout==1.2.1
Requires-Dist: rapid-table==3.0.2
Requires-Dist: rapidocr>=3.9
Requires-Dist: uvicorn>=0.52
Description-Content-Type: text/markdown

# TitanOCR

A standalone, general-purpose OCR and document-structure engine. Give it any PDF or image —
born-digital, scanned, or phone-photographed; printed, hybrid, or handwritten — and it returns
faithful text and structure with pixel-level provenance: every value carries its page, bounding
box, origin and confidence. Deterministic, auditable, and **CPU-only** on the primary path.

> **Scope:** a product in its own right, not a component of any particular domain pipeline.
> Consumers take the Document IR (or its Markdown projection) and do their own downstream
> processing.

---

## The split this project is built on

Sending document images to a vision LLM collapses two different problems into one opaque call:

| Problem | Right tool | Cost profile | Verifiable? |
|---|---|---|---|
| **Transcription** — what characters are on this page, where | OCR (detection + recognition + layout + tables) | ~₹0.00X/page, CPU | yes — bboxes + per-line confidence |
| **Understanding** — what the text means for a given consumer | the caller's own logic/LLM, on **text** | caller's choice | caller's problem — now grounded in auditable text |

A vision-LLM-only approach pays frontier rates on every page, with no bounding box, no calibrated
confidence, and no way to tell a correct answer from a fluent hallucination. This engine makes
transcription deterministic, nearly free, and verifiable — and keeps a VLM only as a budgeted
**escalation** for what CPU OCR cannot yet read confidently (today: handwriting). The roadmap
(phase 17) is to read handwriting on CPU too, with a deterministic model, shrinking escalation to
a safety net.

## Design invariants

These are load-bearing. Changing one is an ADR, not a refactor.

1. **OCR always runs.** Every raster page, no pre-classification gate. It is cheap enough that
   running it unconditionally is free insurance — and its confidence output *is* the triage signal.
   This is why the handwritten/printed split never needs to be known up front.
2. **The VLM is an escalation, never the default path.** It fires on a measured signal, on the
   narrowest possible unit (field > region > page > document), and always with OCR output supplied
   as grounding context.
3. **Every emitted value is traceable to pixels.** Page, bbox, source stage, confidence. If you
   can't point at where a value came from, it doesn't ship.
4. **No accuracy claim without the golden set.** Route mix, cost/page and per-stratum accuracy are
   measured artifacts, not vibes.
5. **A fallback that fires often is not a fallback.** Escalation rate has an explicit budget and
   alerts when breached.
6. **Every stage is a swappable implementation behind an interface**, wired by config, pure over the
   IR, with no module-level state. "Modular" is enforced by the dependency graph and a shared
   contract test suite, not by intent. → [08](docs/08-modularity-and-interfaces.md)
7. **Nothing blocks the event loop.** CPU work runs in a process pool; the async front door stays
   responsive under saturation so health probes never fail mid-document. → [07](docs/07-concurrency-and-scaling.md)
8. **Abstention over guessing.** Below-threshold content escalates or is flagged — never emitted as
   if certain.

## Layout

```
new/
├── README.md                  ← you are here
├── implementation_plan.md     build-level plan (titanocr + API + demo UI)
├── plans/                     numbered phase plans + master tracker (execution source of truth)
├── docs/
│   ├── 01-architecture.md     pipeline, components, service topology, stack + licences
│   ├── 02-document-ir.md      ⭐ the Document IR schema — the central contract
│   ├── 03-triage-and-escalation.md   confidence signals, routing policy, budget
│   ├── 04-handwriting-roadmap.md  the end-state: deterministic CPU handwriting (the roadmap)
│   ├── 05-eval-harness.md     golden set, metrics, release gates
│   ├── 06-deployment-and-operations.md  container, probes, sizing, operational behaviour
│   ├── 07-concurrency-and-scaling.md  execution models, thread/memory arithmetic, backpressure
│   ├── 08-modularity-and-interfaces.md  stage protocols, registry, module boundaries
│   ├── 09-success-criteria.md ⭐ what "good" means, how it's proven, kill criteria
│   └── adr/                   architecture decision records (0001-…, append-only)
├── src/titanocr/            ONE Python project (3.13) — ir/ (the contract), stages/, engines/,
│                              render_md/, service code, demo UI static/
├── pipelines/ · fixtures/ · models/ · scripts/ · tests/
├── packaging/pypi-placeholder/   the `titanocr` PyPI name reservation (phase 18)
└── eval/                    ⭐ the accuracy-proof layer (phase 16)
    ├── gates.yaml             release gates — `target: null` means "not measured yet"
    ├── golden/                golden sets: manifest + per-document truth (smoke · hard)
    ├── harness/               run · score · diff · report · thresholds
    └── scripts/               confidence-histogram batch tooling (threshold bootstrapping)
```

**v1 is a single project** ([ADR 0003](docs/adr/0003-v1-build-simplifications.md)): the IR, stage
interfaces and engines live as modules inside `titanocr`, with dependency direction
(`ir → stages → engines → service`) enforced by import-linter. Extracting `ir/` into a standalone
package is the recorded revert path if the contract ever gains a second consumer.

## Status

| Area | State |
|---|---|
| Design docs | 01–09 drafted and maintained as law (adr/ records decisions) |
| Document IR | **signed off — `ir_version` 1.8** (phase 02 signed off 1.0, 2026-08-13; minor bumps [ADR 0004](docs/adr/0004-table-blocks-carry-lines.md): table blocks may carry their recognised lines, [ADR 0005](docs/adr/0005-table-failed-flag.md): `table_failed` degrade flag, [ADR 0006](docs/adr/0006-quality-ink-ratio.md): `quality.ink_ratio`; [ADR 0007](docs/adr/0007-escalation-output-is-line-shaped.md) fixes what escalation may claim; later minors add the block types `figure`/`signature` without lines [ADR 0010](docs/adr/0010-visual-blocks-carry-no-lines.md), `rule` [ADR 0011](docs/adr/0011-rule-blocks.md), `connector` [ADR 0012](docs/adr/0012-connector-blocks.md) + its label anchors [ADR 0016](docs/adr/0016-arrows-anchor-to-labels.md), and `formula` [ADR 0018](docs/adr/0018-formula-blocks.md)) |
| `titanocr` service | **released `v0.1.0` (2026-08-13)** — pipeline stages 0–5 feature-complete and containerised: digital fast path (exact, ₹0), spawn pool with thread pinning + warm-up-gated readiness, PP-OCRv5 det/rec chosen accuracy-first by fixture bake-off (scan CER 0.15%, photo 0.30%), layout + reading order, geometric table reconstruction with ML fallback (cells exact on every fixture), triage `confidence-v1` + reject path + signals JSONL, VLM escalation (cropped, OCR-grounded, additive merge, caps + breaker; live-provider smoke blocked on an endpoint/key), demo UI at `/demo`, 976 MB `linux/amd64` image starting to ready offline. Full per-phase detail: [plans/](plans/README.md) |
| Accuracy | **still no accuracy claim.** The [phase-16](plans/phase-16-eval-harness.md) harness is built and in CI (`make eval` — stratified CER, cell accuracy at `(row,col)`, detection recall, route mix, ₹/page, abstention honesty, determinism, 14 release gates, clickable HTML report), and it *enforces* the honesty rules rather than stating them: a model-produced label is refused, a synthetic set reports `unclaimable`, an unset target can never report `pass`. **The golden set — real, human-labelled documents — does not exist yet**, so the verdict is `incomplete` by construction and phases 17/18/21 stay gated. First run's smoke baseline and the real defect it found: [note](docs/notes/eval-smoke-baseline-2026-08-26.md) |
| Packaging | the PyPI name **`titanocr`** is reserved (0.0.1 placeholder, no OCR code — `make pypi-placeholder`). Release automation is in place ([Releasing to PyPI](#releasing-to-pypi)). Shipping the engine itself as `pip install titanocr` is [phase 18](plans/phase-18-pypi-distribution.md), gated on the accuracy proof ([ADR 0009](docs/adr/0009-titanocr-rebrand-and-distribution.md)) |
| Structure fidelity | **phases 19–20 done (2026-08-26)** — the engine no longer asserts structure a page does not have (a stacked formula is not a grid; an equation number stays in its own y-band), the born-digital route types its regions with the layout stage in a pool worker ([ADR 0017](docs/adr/0017-layout-on-the-digital-fast-path.md), ≈0.26 s/page), and mathematics is a block type excluded from structure-guessing **by type** ([ADR 0018](docs/adr/0018-formula-blocks.md)). Renderer 1.12 |
| Roadmap | phase 16 (accuracy proof) → phase 17 (**handwritten OCR on CPU, deterministic**), phase 18 (**`pip install titanocr`**) and phase 21 (**formula → LaTeX**), all gated on it — see [plans/](plans/README.md) |

## Conventions

- Non-obvious or reversible-at-cost decisions get an ADR in `docs/adr/`. Numbered, append-only,
  superseded rather than edited.
- The Document IR is versioned. Breaking changes bump the major and require a migration note.
- Anything measured (thresholds, costs, route mix) cites the run that produced it. No folk numbers.

## Releasing to PyPI

Releases are published by [`.github/workflows/release.yml`](.github/workflows/release.yml) when a
`v*` tag is pushed. It uses PyPI **trusted publishing** (OIDC), so no API token is stored in the
repository or its secrets.

> **Gated:** publishing the engine is [phase 18](plans/phase-18-pypi-distribution.md), and phase 18
> waits on phase 16's accuracy proof (hard rule 14). The pipeline is ready; pushing a real version
> tag is the release decision.

**One-time setup**

1. On [pypi.org](https://pypi.org/manage/project/titanocr/settings/publishing/): *titanocr →
   Settings → Publishing → Add a new publisher → GitHub* with owner `Phantom-IN`, repository
   `titanocr`, workflow `release.yml`, environment `pypi`.
2. Optional, for dry runs: the same on [test.pypi.org](https://test.pypi.org) with environment
   `testpypi` (the name must be registered there too).
3. On GitHub: *Settings → Environments* → create `pypi` and `testpypi`. Adding yourself as a
   required reviewer on `pypi` makes every release wait for a manual approval before upload.

**Cutting a release**

```bash
uv version 0.2.0                      # bump the version in pyproject.toml
git commit -am "release 0.2.0"
git tag v0.2.0
git push origin main v0.2.0           # the tag push starts the release
```

The workflow then runs, in order — and nothing is uploaded unless every step passes:

1. the full CI gate (`ci.yml`: lint, typecheck, tests, real-engine contract suites, eval, container);
2. a check that the tag matches the `pyproject.toml` version;
3. `uv build` (sdist + wheel) and `twine check --strict`;
4. a clean-venv install and import of the wheel on Linux and macOS;
5. upload to PyPI (after approval, if the `pypi` environment requires it).

**Dry run:** *Actions → release → Run workflow* runs the same steps and publishes to **TestPyPI
only**, never PyPI.

**Rules**

- PyPI never accepts the same version twice — even after a deletion. A failed or wrong release is
  fixed by publishing the next version, not by re-tagging.
- Keep versions above the published ones (the placeholder is `0.0.1`; the engine is at `0.1.0`).
- Models are never in the distribution; they are fetched and checksummed (`make models`).

## Read order

New to the project: this file → [01-architecture](docs/01-architecture.md) →
[02-document-ir](docs/02-document-ir.md). Those three are the whole idea.

Before writing service code, also read [07](docs/07-concurrency-and-scaling.md) and
[08](docs/08-modularity-and-interfaces.md) — they constrain how the code is structured, and
retrofitting either is expensive.

Reviewing the project: this file → [09-success-criteria](docs/09-success-criteria.md). That doc
carries the targets, how each claim gets proven, and the conditions under which we stop.
