Handwriting-aware PDF/image OCR service with pluggable HuggingFace model backends. Extracts text and images from PDFs into structured JSON so downstream AI agents can process the content (e.g., building an Obsidian vault).
- Text layer check — PyMuPDF tries to extract embedded text from each page. If a page has enough text (≥10 chars by default), it's used directly — no model needed.
- OCR fallback — Pages with little or no text (scanned documents, handwritten notes) are rendered to images and sent through an OCR model.
- Image extraction — Embedded images (diagrams, figures, photos) are extracted and saved alongside the text.
# Install (includes both GOT-OCR2 and Qwen2.5-VL backends)
uv sync
# With the marker backend (structured Markdown/JSON for typed/scanned PDFs)
uv sync --extra marker
# Dev tools (pytest, ruff)
uv sync --extra dev# Parse a PDF (uses text layer where possible, OCR for the rest)
docparse parse ./lecture.pdf
# Force OCR on all pages
docparse parse ./handwritten_notes.pdf --force-ocr
# Choose model and output format
docparse parse ./notes.pdf --model got-ocr2 --format markdown --output ./parsed/
# List available models
docparse models# Start the server
uvicorn document_parser.server:app --reload
# Parse a file
curl -X POST http://localhost:8000/parse \
-F "file=@./lecture.pdf" \
-G -d "model=got-ocr2"
# Check available models
curl http://localhost:8000/modelsfrom document_parser import DocumentParser
parser = DocumentParser(model="got-ocr2")
result = parser.parse("./lecture.pdf", output_dir="./output")
for page in result.pages:
print(f"Page {page.page} ({page.source}): {page.text[:100]}...")
for img in page.images:
print(f" Image: {img['id']} ({img['width']}x{img['height']})")Beyond single files, docparse can process folders of class PDFs and feed them into one
long-lived, concept-first Obsidian vault. The pipeline has two halves:
- Deterministic processing (code) — loop a class folder, route each PDF to the right engine,
and write per-document JSON + a
batch_index.jsonstatus queue. - Vault integration (agent) — the
vault-buildskill drains that queue into a single vault where concepts are notes, subjects are folders, and courses are provenance.
The input contract is one folder per class (the folder name is the course, each PDF filename is the lecture title):
# Scaffold a batch.toml (suggests an engine per PDF + normalized titles to edit)
docparse batch init "~/Probability"
# Process the class folder → _parsed/<stem>.json + _parsed/batch_index.json
docparse batch run "~/Probability" # engine per manifest
docparse batch run "~/MachineLearning" -e marker # or override for the whole folder
# Inspect status
docparse batch status "~/Probability"Engine routing is manifest-driven: batch init suggests an engine from a text-layer probe
(typeset → marker, scanned/handwritten → qwen-vl-3b), which you confirm/override in batch.toml.
The vault is a dedicated standalone folder (default ~/CMU-Vault/), kept outside the repo and
the class folders. Its concept index is maintained deterministically:
# Scan the vault → .vault-index.json (concepts, aliases, topics, topic-dependency graph)
docparse vault index --vault ~/CMU-Vault # path remembered in ~/.docparse.tomlThen invoke the vault-build skill (in Claude Code) to integrate the processed documents:
it dedups each concept to one canonical note, merges new lectures into existing notes, and links
applied → foundational while keeping cross-topic links acyclic. Classes can be integrated in any
order — forward references become Obsidian dangling links that resolve on re-index. See
.claude/skills/shared/vault-conventions.md for the full vault model.
A local, click-driven UI over the whole flow — upload, route, process with live progress, review
parsed output, then a Vaultify button that runs the Claude Code vault-build skill in an
embedded terminal. Everything runs on your machine (FastAPI + your MPS); the finished vault is a
local folder you open in Obsidian.
# one-time: build the frontend
cd web && npm install && npm run build && cd ..
# launch (serves UI + API at http://localhost:8765, opens the browser)
uv run docparse webFor frontend development, run the Vite dev server (cd web && npm run dev) alongside
uv run docparse web — the dev server proxies /api (and WebSockets) to the backend.
Adding a class — two ways:
- New class — upload PDFs — makes an empty class and you upload PDFs into the workspace.
- Use a folder I already have — a native Choose folder… picker (macOS) registers an
existing folder of PDFs. They're read in place and never moved; all generated artifacts
(
batch.toml,_parsed/, the index) are written into the workspace instead, leaving your source folder pristine.
Classes live under a workspace (~/Documents/docparse-classes/ by default); the workspace and
vault paths are remembered in ~/.docparse.toml.
Screens: Library (classes), Process (per-PDF engine/title + run with per-page progress),
Review (source page vs parsed text/images side by side), Vaultify (embedded Claude Code
terminal running vault-build). The backend lives in src/document_parser/webapp.py; the UI in
web/.
{
"filename": "lecture.pdf",
"pages": [
{
"page": 1,
"text": "Convex Optimization\n\nConvex Sets...",
"source": "text_layer",
"images": []
},
{
"page": 5,
"text": "A set C is convex if...",
"source": "text_layer",
"images": [
{ "id": "p5_img0", "width": 640, "height": 480, "path": "output/images/p5_img0.png" }
]
}
],
"metadata": {
"total_pages": 73,
"ocr_pages": 0,
"text_layer_pages": 73,
"images_extracted": 15,
"model": "qwen-vl-3b",
"force_ocr": false,
"elapsed_ms": 320.5
}
}source is "text_layer" for pages where embedded text was used, or the model name (e.g., "qwen-vl-3b") for pages that went through OCR.
| Model | Kind | Memory | Best for |
|---|---|---|---|
qwen-vl-3b (default) |
image OCR | ~6GB | Messy handwriting, complex layouts; sized for 12–16GB Apple Silicon |
qwen-vl-7b |
image OCR | ~14GB | Best handwriting accuracy; needs 16GB+ RAM/VRAM |
got-ocr2 |
image OCR | ~4GB | General OCR, typed + handwritten text, tables, formulas |
marker |
document pipeline | ~3–5GB | Typed/scanned PDFs: structured Markdown/JSON, tables, equations (install with --extra marker) |
The OCR engines (qwen-vl-*, got-ocr2) are image → text: the parser renders each
page and routes it through the model per page. marker is a document pipeline — it
owns the whole document (its own layout analysis, reading order, table/equation
recognition, and image extraction), so the per-page text-layer routing, --force-ocr,
and --text-threshold don't apply to it. Marker pages carry extra structured fields
(blocks, and markdown when available) in the JSON output.
License note: marker's code is GPL-3.0 and its model weights are non-commercial (cc-by-nc-sa / OpenRail-M, with a waiver for small orgs). Fine for personal/research use; review the upstream license before any commercial use. The other backends are unaffected.
--use-llm (CLI) / use_llm=true (API) enables marker's optional LLM augmentation
(better tables/equations/forms). It needs ANTHROPIC_API_KEY and adds latency + cost.
Create a new file in src/document_parser/models/ and use the @register_engine decorator:
from document_parser.engine import OCREngine, register_engine
@register_engine("my-model")
class MyEngine(OCREngine):
def load(self):
# Load model weights
...
def run(self, image):
# Run OCR, return text string
...Then import it in src/document_parser/models/__init__.py so the decorator runs on startup.
src/document_parser/
├── __init__.py # Public API
├── engine.py # OCREngine / DocumentEngine ABCs, ModelRegistry, @register_engine
├── extractor.py # PyMuPDF: text extraction, page rendering, image extraction
├── parser.py # Smart routing orchestrator (per-page OCR + document backends)
├── batch.py # Folder batch processing: manifest, engine suggestion, batch_index.json
├── vault.py # Concept index (.vault-index.json) + topic-dependency graph
├── models/
│ ├── got_ocr2.py # GOT-OCR2 backend
│ ├── qwen_vl.py # Qwen2.5-VL backends (default)
│ └── marker.py # marker document backend (optional, --extra marker)
├── server.py # FastAPI server (single-file /parse API)
├── webapp.py # FastAPI app for the local web UI (classes, process, review, terminal)
├── jobs.py # Background processing jobs + progress events for the web app
└── cli.py # Typer CLI (parse, batch, vault, web, models)
web/ # Vite + React + xterm.js frontend (built to web/dist, served by webapp.py)
Vault building is driven by Claude Code skills in .claude/skills/: vault-build (batch
orchestrator) over vault-from-marker / vault-from-ocr / vault-from-handwriting, all sharing
shared/vault-conventions.md.