PDF to Markdown for RAG: 2026 Comparison of Top Tools
Most RAG pipelines start the same way: convert the PDF to Markdown, chunk it, index it, retrieve, answer. Converting PDF to Markdown for RAG is a mostly solved problem in 2026, with several excellent open-source converters.
For AI agents, though, Markdown is only the first step. Once a 300-page manual becomes a pile of chunks, each chunk forgets which section it came from, which table backs up the sentence it contains, and which page a reviewer should open to check the answer. Agents need four things that plain Markdown doesn’t carry:
- Structure: the section hierarchy, so an agent can browse an outline instead of guessing.
- Tables and images: kept as evidence and linked to the text that refers to them.
- Citations: a pointer back to the document, section and page.
- Evidence: the original page, especially for scans, drawings and slides where text extraction loses meaning.
This guide compares four tools on those criteria: Marker v2, MarkItDown, MinerU (including 4.0), and Knowhere by Ontos AI, an open-source document memory layer for AI agents. (Not to be confused with the vector search engine inside Milvus, which shares the name.) Disclosure: this article is published by the Knowhere team. We’ve tried to be fair, and we cover cases where another tool is the better fit.
Quick answer: which should you pick?
| If you need… | Pick | Why |
|---|---|---|
| The fastest, lightest way to turn files into Markdown for an LLM | MarkItDown | One pip install, MIT, many formats, token-efficient output |
| High-fidelity Markdown/JSON from hard PDFs (tables, math, multi-column), CPU or GPU | Marker v2 | Published olmOCR-bench results, fast mode for CPU, JSON with bounding boxes |
| A fully local parser with strong OCR/VLM and an agent-readable local library | MinerU 4.0 | Many formats, local CPU/GPU runtimes, stable page/block citation locators |
| A persistent, multi-document memory that agents can navigate and cite over MCP | Knowhere | Parsing + hierarchy + linked assets + agent-driven retrieval + page-level citations in one stack |
These tools aren’t mutually exclusive. A common pattern is to use a converter for raw extraction and a memory layer on top. Knowhere’s own V1 PDF path actually calls MinerU’s API for initial extraction.
Marker v2 (Datalab)
Marker converts PDF, image, PPTX, DOCX, XLSX, HTML and EPUB files to Markdown, JSON, HTML, or chunks. Version 2.0.0 (July 2026) was a rewrite focused on speed and CPU support. It offers a balanced mode (default on GPU) and a fast mode built for CPU, plus an optional --use_llm hybrid mode that merges tables across pages and handles inline math and forms.
What it does well
– Accuracy you can check yourself. Datalab reports 76.0% overall on olmOCR-bench (balanced mode), a third-party benchmark of 1,403 PDFs, and ships instructions for reproducing it. They also report roughly 5× more pages per second than MinerU’s pipeline backend. Both figures are Datalab’s own measurements.
– Rich output. JSON is a block tree with page bounding boxes, and the chunks format flattens top-level blocks for easy RAG chunking.
– Licensing for most teams. Code is Apache-2.0. Model weights use a modified OpenRAIL-M license that is free under $5M in funding or revenue.
Where it stops
Marker is a converter. It doesn’t keep a persistent index, a cross-document view, or retrieval tools for agents. You bring your own chunk store and search. The hosted Datalab API is priced at $4 per 1,000 pages (fast/balanced) or $10 per 1,000 pages (accurate).
MarkItDown (Microsoft)
MarkItDown is a lightweight Python utility that converts PDF, PowerPoint, Word, Excel, images, audio, HTML, CSV/JSON/XML, ZIP, EPUB and even YouTube URLs into Markdown. It is MIT-licensed and one of the most-starred projects in this space.
What it does well
– Simplicity. pip install 'markitdown[all]', then markitdown file.pdf -o file.md.
– Token efficiency. Markdown is compact, and MarkItDown keeps headings, lists, tables and links.
– Agent access to conversion. A markitdown-mcp package exposes conversion over MCP, and an OCR plugin uses an LLM vision client for text in embedded images.
Where it stops
The output is Markdown only. The README itself says it may not be the best option for high-fidelity conversion. Complex layouts and tables in difficult PDFs can lose fidelity, and there is no retrieval layer, citation model, or page evidence. It’s a great first step for your pipeline, but it’s not built to be agent memory.
MinerU and MinerU 4.0 (OpenDataLab)
MinerU is a mature open-source multimodal parsing engine focused on VLM and OCR. It supports PDFs, images, Office formats, RTF, ODF, EPUB, OFD, HTML and CSV, and outputs Markdown, HTML, LaTeX, DOCX, EPUB and structured content lists. It runs locally on CPU (ONNX, llama.cpp) or GPU (vLLM, LMDeploy), in Docker, or through an official hosted API.
MinerU 4.0 goes beyond Markdown. It adds “a local document library and agent reading: discover files, cache results, search content, continue by page or block, and preserve stable citation locators” in the form doc:/tier:/page:/block:, plus an installable agent skill. That’s a real step toward agent-ready documents.
What it does well
– Strong parsing on hard documents, many input formats, and a fully local deployment.
– Agents can read long documents incrementally and cite precise page/block locations.
Where it stops
MinerU describes itself as “not a RAG framework, vector database, or chat-with-doc application.” Its library is centered on reading parsed documents. It doesn’t publish a multi-document memory with a section hierarchy, cross-document relationships, and hosted MCP access. Its license is Apache-2.0 with extra terms: attribution in online services, and a commercial license above 100M MAU or $20M monthly revenue.
Knowhere by Ontos AI
Knowhere describes itself as “a document parsing and retrieval system that turns complex, dirty files into persistent, navigable memory for AI agents.” The team’s framing: “We’re not developing the next MinerU — we’re building document memory infrastructure that agents can effectively consume.”
What it does well: parsing, hierarchy, a document graph, publishing into searchable namespaces, and agent tools in one open-source product (Apache-2.0; SDKs, CLI and MCP are MIT), with page images kept as evidence and page-level citations. A hosted API costs $0.015 per page, with bring-your-own-model support on v2 endpoints.
Where it stops: it’s young (open-sourced May 2026), its benchmarks are internal, and self-hosting is heavier than a converter. Details below.
Other tools worth knowing
- Unstructured: ETL-style partitioning, chunking, embedding and 20+ destination connectors, with enterprise compliance. Hosted pricing is $0.015/page after 10,000 free pages.
- LlamaParse / LlamaIndex: a closed cloud parser feeding the LlamaIndex RAG/agent framework (1,000 credits = $1.25).
- Reducto: a closed enterprise platform for parsing, schema extraction with citations, and compliance (Parse at $10 per 1,000 pages).
- Docling: an MIT-licensed converter (LF AI & Data) with a rich document model, hierarchical/hybrid chunkers and an MCP server.
Feature comparison table
| Marker v2 | MarkItDown | MinerU 4.0 | Knowhere | |
|---|---|---|---|---|
| Output | Markdown, JSON (block tree + bboxes), HTML, chunks | Markdown | Markdown, HTML, LaTeX, DOCX, EPUB, structured content lists | chunks.json (text/table/image/page chunks), optional content.md, images, HTML tables; published into namespaces |
| Structure / hierarchy | Per-page block tree | Headings, lists, tables in Markdown | Page/block structure with locators | Section tree (section_path), page ranges, summaries, entities, cross-document graph |
| Tables & images | Tables as HTML blocks; images extracted | Markdown tables; image text via LLM OCR plugin | Tables, formulas and images parsed | Separate table/image chunks linked to their host section; read-only SQL over table grids |
| Citations / provenance | Page bounding boxes in JSON | None built in | Stable doc/tier/page/block locators |
Document → section path → page numbers → rendered page image / asset URL |
| Agent access / MCP | None (converter + API server) | markitdown-mcp (conversion) |
CLI, local library, agent skill | Hosted MCP (Cursor, Claude Code, Codex, VS Code…), self-hosted /mcp, SDKs, CLI |
| Self-host | Yes (CPU/GPU/MPS) | Yes (Python) | Yes (CPU/GPU, Docker) | Yes (Docker), needs Postgres, Redis, S3-compatible storage and LLM/VLM keys |
| License | Code Apache-2.0; weights modified OpenRAIL-M | MIT | MinerU Open Source License (Apache-2.0 + extra terms) | Apache-2.0 core; SDK/CLI/MCP MIT |
| Hosted price | Datalab API $4–$10 / 1,000 pages | Free (OSS) | Official hosted API available | $0.015/page ($15 / 1,000 pages), $5 free credit |
Facts as of October 2026, taken from each project’s README and pricing pages.
How Knowhere works
1. Dual-track parsing, one memory schema
Knowhere 2.0 (September 2026) parses documents along two tracks:
- Text Track “preserves precise text and native structure where they are reliable.” It is used for Word, Excel, Markdown, HTML, images and other text-first formats. Tables become HTML with summaries, and scanned text goes through local OCR.
- Vision Track (VISION-MAP) “uses frontier vision models to understand a page or slide as a whole.” PDF and
.pptxuploads through the V2 Jobs API use this track. The original page image is kept as first-class evidence, and Knowhere builds a chapter–section–page map with light per-page notes.
Why keep the page? The team’s example is a scanned contract where “the amount is 53 million. After extraction, it becomes 5.3 million.” If the agent can open the original page, it can catch that kind of error.
Both tracks produce the same chunk and metadata schema. Each document becomes a tree of sections addressed by section_path (for example AI_Security_Report.docx / 1. Overview / 1.1 Key Findings). Tables and images are separate chunks linked back to their section. Typed entities and keywords connect related documents across a namespace.
2. Agent-native retrieval tools
Instead of returning a fixed top-k list, Knowhere exposes the corpus as tools that an agent can combine. In the open-source server these are:
- List:
corpus.list_documents - Outline:
corpus.outlineandcorpus.node_filter, a deterministic filter over section titles and summaries - Search:
corpus.grepfor exact strings andcorpus.recallfor fuzzy ranked candidates - Read:
corpus.readfor full sections or chunks - Tables and images:
corpus.assetsandcorpus.query_table, which runs a read-only SELECT on a table’s grid - Related docs:
corpus.neighbors, built on the persisted document graph
Under the hood this is keyword search plus navigation. The classic mode uses BM25 with rank fusion across path, content and term channels. Agentic mode lets an LLM decide how to search, traverse and read. You don’t need a vector database or embeddings. Knowhere returns evidence rather than synthesized answers, plus a decision_trace for agentic runs, so your own model writes the final answer.
3. Page-level citations
Every result carries references that resolve to the source document, section path, page numbers and, where available, the rendered page image or asset URL. Agents can show their sources, and reviewers can check them.
Getting started
Hosted API (REST). Sign up at knowhereto.ai ($5 in free credits), then:
curl -X POST https://api.knowhereto.ai/v1/jobs \
-H "Authorization: Bearer $KNOWHERE_API_KEY" -H "Content-Type: application/json" \
-d '{"source_type": "url", "source_url": "https://example.com/document.pdf"}'
curl https://api.knowhereto.ai/v1/jobs/JOB_ID -H "Authorization: Bearer $KNOWHERE_API_KEY"
Python SDK (pip install knowhere-python-sdk):
import knowhere
client = knowhere.Knowhere(api_key="sk_...")
# Parse one document
result = client.parse(url="https://example.com/report.pdf")
print(result.statistics.total_chunks)
print(result.full_markdown[:200])
# Publish into a namespace, then retrieve evidence
job = client.jobs.create(source_type="url", source_url="https://example.com/manual.pdf",
namespace="support-center", document_metadata={"title": "Support manual"})
client.jobs.wait(job.job_id)
response = client.retrieval.query(namespace="support-center", query="How do I reset Bluetooth pairing?",
chunk_types=["page"], top_k=5, channels=["path", "term"], filter_mode="keep",
signal_paths=["Bluetooth", "Pairing"])
for reference in response.referenced_chunks:
print(reference.chunk_id, reference.chunk_type, reference.content_source)
MCP for Cursor, Claude Code or Codex (Node.js 20.19+):
npx -y @ontos-ai/knowhere-mcp login
# Claude Code
claude mcp add --transport stdio knowhere -- npx -y @ontos-ai/knowhere-mcp
# Codex
codex mcp add knowhere -- npx -y @ontos-ai/knowhere-mcp
For Cursor, add this to ~/.cursor/mcp.json:
{ "mcpServers": { "knowhere": { "command": "npx", "args": ["-y", "@ontos-ai/knowhere-mcp"] } } }
The same library you manage in the hosted Notebook is then available to every MCP client on your account. Note that the published MCP package is a local stdio wrapper around the hosted API, so clients that only support remote HTTP servers can’t use it yet.
Self-host: use knowhere-self-hosted (docker compose up -d) or build from the main repo.
When Knowhere is not the right choice
- You only need Markdown. If your pipeline just feeds converted text to a model, MarkItDown is lighter, and Marker or MinerU give you higher-fidelity conversion with published or community-tested quality.
- You need independent benchmarks. Knowhere’s benchmark is the team’s own internal evaluation on agentic RAG tasks, with no published dataset or reproduction script. In that chart Knowhere leads MinerU only narrowly on first-pass accuracy (68% vs 66%) and recall (82% vs 78%). The bigger gap is accuracy after feedback (79% vs 64%). MarkItDown used slightly fewer tokens and was marginally faster. There are no Knowhere scores on standard parsing benchmarks such as olmOCR-bench.
- Price per page is the deciding factor. At $15 per 1,000 pages, Knowhere Cloud is not the cheapest parser. Datalab ($4–$10), Reducto Parse ($10) and Unstructured ($15 after 10,000 free pages) are comparable or lower. The price covers parsing plus hierarchy, published memory and retrieval, but if you only need parsing you may pay for things you won’t use.
- You need a fully offline, lightweight install. Self-hosting needs Postgres, Redis, S3-compatible storage and an LLM/VLM API key. The V1 PDF path uses a MinerU API key, and the vision track sends page images to a model endpoint unless you run your own OpenAI-compatible one. Self-hosted installs send anonymous telemetry by default (opt out with
TELEMETRY_ENABLED=false). Also note that the PDF worker depends on PyMuPDF, which is AGPL or commercially licensed, so commercial self-hosters should do a license review. - You need compliance paperwork today. Reducto, LlamaParse, Unstructured and Datalab publish compliance programs such as SOC 2. We haven’t published equivalent certifications.
FAQ
What is the best PDF to Markdown converter for RAG?
It depends on the documents. For quick, lightweight conversion, MarkItDown is the simplest. For complex PDFs with tables, math and multi-column layouts, Marker v2 and MinerU are stronger. Marker publishes olmOCR-bench results, and MinerU offers fully local OCR/VLM parsing. If agents must navigate and cite many documents, add a memory layer such as Knowhere on top.
Is Markdown enough for agentic RAG?
Usually not on its own. Markdown keeps headings and tables but loses page provenance, links between tables or images and the text that cites them, and relationships across documents. Agents work better with a section hierarchy, linked assets and citations that resolve to a page.
Does Knowhere use a vector database?
No. Knowhere’s retrieval is keyword search (BM25 with rank fusion) plus agent-driven navigation over a section tree and document graph. It needs no embeddings or vector store. The trade-off is that pure paraphrase matches rely on the agent’s search strategy rather than semantic similarity.
How is Knowhere different from MinerU 4.0?
MinerU 4.0 is a local parser with a document library, search and stable page/block citation locators for agent reading. Knowhere focuses on what happens after parsing: a published, multi-document memory with section hierarchy, linked tables and images, cross-document relationships, page-image evidence and hosted MCP access. The two can work together.
Can I self-host Knowhere?
Yes. The core is Apache-2.0 and runs with Docker Compose. You’ll need Postgres, Redis, S3-compatible storage and LLM/VLM API keys (any OpenAI-compatible provider, such as DeepSeek or Qwen).
Conclusion
Converting PDF to Markdown for RAG is now table stakes, and Marker v2, MarkItDown and MinerU all do it well for different needs. The open question for agent builders is what happens after conversion: can the agent find the right section, open the table that backs the claim, and cite the page? That’s the gap Knowhere is built to fill.
If that matches your use case, try it with the $5 free credit at knowhereto.ai, read the docs, or star and inspect the code on GitHub. And if a converter alone covers what you need, we hope this comparison helped you pick the right one.