Awesome AI Agent Stack
Document Processing & Data Ingestion
Getting real-world files into a form agents can use.
Document parsing & conversion
- microsoft/markitdown - Convert Office files, PDFs and more to Markdown.
- docling-project/docling - IBM's document parser with layout-aware PDF, tables and OCR; best open option for RAG ingestion.
- Unstructured-IO/unstructured - Partition and chunk any document type for LLM pipelines.
- opendatalab/MinerU - High-quality PDF-to-Markdown extraction, strong on scientific papers.
- datalab-to/marker - Fast PDF to Markdown with GPU acceleration.
- firecrawl/anydoc - Rust converter for Word, PowerPoint, Excel, EPUB, CSV and PDF to clean Markdown.
- allenai/olmocr - Toolkit for linearizing PDFs for LLM datasets/training.
- run-llama/liteparse - Fast open source document parser from LlamaIndex.
- oomol-lab/pdf-craft - Converts PDF files into other formats such as Markdown and EPUB.
- opendataloader-project/opendataloader-pdf - PDF parser that outputs AI-ready structured data.
- breezedeus/Pix2Text - Open-source tool recognizing layouts, tables, formulas, and text as Markdown.
- deepdoctection/deepdoctection - Document AI toolkit for layout analysis and information extraction.
- apache/tika - Toolkit detecting and extracting metadata and text from 1000+ file types.
- mwilliamson/python-mammoth - Convert Word documents (.docx files) to HTML.
- jgm/pandoc - Universal markup converter.
- firecrawl/pdf-inspector - Rust library for PDF classification and fast text extraction.
- KOUISAmine/file-converter-tools - Collection of file conversion utilities.
- gotenberg/gotenberg - Docker API for converting documents to PDF.
- Kozea/WeasyPrint - Convert HTML and CSS to PDF.
- paperless-ngx/paperless-ngx - Self-hosted document management with OCR.
- mayan-edms/Mayan-EDMS - Free open-source document management system with OCR and workflow support.
- xberg-io/xberg - Rust-core document intelligence extracting text, tables, and metadata.
OCR
PDF libraries
Structured extraction
Chunking & data pipelines
Media ingestion
Web content extraction
Web archiving