← Explore
TOPIC

#pdf-extraction

Open source repositories tagged with #pdf-extraction, ranked by health score.

Zipstack
Zipstack/unstract
Python
89
health

LLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows

7.2k
xberg-io
xberg-io/xberg
Rust
89
health

Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.

9.2k