VISION-MAP: Knowhere Doesn’t Just Read Text. It Sees the Page.
Most document systems sketch a file into words. VISION-MAP keeps the original page, maps it, and lets the agent look at it.
Here’s a simple question: if a document system is actually built for agents, how should it understand a file?
The usual answer is extract first. Parse the document into text, then let the agent read the extraction.
That path is useful. Clean PDFs, Word files, web pages, and ordinary tables convert well. Search is fast. Citations are easy. Downstream processing is cheap. PAGE-INDEX, MinerU, and most OCR pipelines have been on this road for about twenty years.
Knowhere can do this too.
The problem shows up on the files people actually use: manuals, proposals, research reports, drawings, dense tables. Crooked scans. Stamps sitting on numbers. Two-column layouts. Handwritten notes. Tables that run across pages. Even a strong extractor will get something wrong. Once that wrong text lands in the knowledge base, the agent keeps consuming it. You can keep training better parsers. On real, messy files, the returns are already close to zero.
Most extractors are trying to sketch the document. A sketch is always a little wrong.
So we built a second path. Instead of extracting text, VISION-MAP understands, organizes, and annotates pages. The agent sees the original page, not a lossy stand-in.
We like to think of it as FSD for this space. Extractors sketch. VISION-MAP photographs the file and writes notes on it. High fidelity, on purpose.

What VISION-MAP actually does
The page is a first-class citizen.
We do not assume a page should be rewritten into a text double. The original page stays in the system as the most complete, closest-to-source record. Knowhere then organizes those pages by the document’s own chapters, topics, and relationships, into a hierarchy and a network the agent can navigate across files.
To make pages findable, the system lightly extracts the core text and uses a vision model to add chapter, topic, and content labels. Tables, charts, and important figures can still be pulled out on demand, as assets the agent can search, cite, and reuse.
When the agent needs the material, it first searches and navigates the page map. After it finds the relevant pages, it hands the originals to a vision model, which reads amounts, clauses, charts, or layout relationships for the question at hand, and traces the answer back to a specific page and source file.
Knowhere is not choosing between text parsing and visual understanding. It runs both tracks.
The text track parses documents into paragraphs, tables, and other structured content. It is a good fit for clean files where the words come out easily: DOCX, XLSX, Markdown, JSON, and other non-page formats.
The vision track is VISION-MAP. It treats the full page as the object of understanding. That is the better fit for scans, drawings, dense tables, and any file where the layout itself carries meaning: PDFs, PowerPoint, and other page-native formats.
The knowledge-base schema is designed so both tracks land on the same document map. The agent’s way of finding its way does not change. What changes is what it gets when it arrives: a passage of text, or the original page.

VISION-MAP does more than patch the gaps in text extraction. It widens what Knowhere can understand.
The system used to be strongest on documents where text is the main event. With the vision track, photos, scans, engineering drawings, chart-heavy reports, and visually dense decks can all become part of document memory. More kinds of files. More kinds of work.
Why a vision-only path?
In industries that cannot afford a sloppy parse — manufacturing, engineering, finance, legal, healthcare — a result with no source page to go back to is not acceptable.
Take a scanned contract. The amount is 53 million. After extraction, it becomes 5.3 million.
One missing zero.
Once that number is in the knowledge base, every search, calculation, and answer that follows may be built on 5.3 million. The agent can even return something that looks complete: clean logic, tidy citations, a confident tone. The error was already in the source.
You cannot get this to zero.
It is not always because the parser is weak, and it is not something a stronger model automatically fixes.
Real files are too messy. Crooked scans. Stamps over text. Two-column layouts. Tables that span pages. Handwritten notes. Engineering drawings. There is no closed answer key for these. Accuracy can keep going up. But if the system treats the extraction as the only ground truth, one small error can travel the whole chain. That is a cascading error.
In ordinary material, that might mean a slightly wrong answer. In engineering, finance, or law, it is much worse.
On a drawing, one line or one elevation can change the next decision. In a financial statement, the unit might be thousands or billions. Where the decimal sits matters.
These industries do not just want an answer. They want an answer that can be checked, audited, and traced.
If the original page is gone, or has degraded into an attachment nobody can find, there is nothing to go back to. You only see the machine’s secondhand transcript, not the page that actually entered the system.
You never copied it completely. And you cannot undo the rewrite.

Keeping the source file is insurance.
If traditional extraction is sketching the document, then every new recognition model is trying to make the sketch look more like the original.
VISION-MAP takes a different bet. If a sketch can always miss a detail, keep the original page in the system the way a photographer keeps the negative.
Later, when the agent answers a question, or when a person checks the answer, the citation can go back to the original file and the specific page. You can see where the number was written, how the table was laid out, and what the stamp or the handwritten note actually covered. Keeping the source does not make the model infallible. It makes errors findable, checkable, and auditable.
That does not make text parsing worthless. The opposite: anything that parses cleanly into text should stay on the text track. VISION-MAP is for the parts that should not be understood only as a text double.

Why now?
“Just look at the page” was easy to imagine and hard to run.
Early vision models were unreliable on complex pages. They could only take a few images at a time. Speed and cost were not in the range for large document collections. Handing every page to a model, live, was a bad experience.
Multimodal models can now read several images together, and they can handle longer sequences of pages. Tables, figures, layout, and text do not have to be split into isolated pieces first. The model can judge them as one page.
Work like ColPali, which retrieves against whole-page images, also showed that “find the page, then understand the page” is a viable route.
The stack is finally good enough to turn visual understanding from a demo into a real track inside Knowhere’s document system.
How VISION-MAP works
First, keep the source and annotate the page.
When a document enters Knowhere, the original is not replaced by a Markdown file. Each page is stored as it appeared. Its core content and visual assets have already been understood and labeled. The text, table lines, images, stamps, and relative positions are still there, so later audit and traceback have a complete entry point. What we keep is page annotation, not a second copy of the page as an image.
Second, build the map.
A pile of pages is not enough. A few hundred pages with no structure still forces the agent to flip from the beginning.
VISION-MAP organizes pages by chapter into a chapter–section–page map, and adds a light note on each page. Those notes tell the agent what the page is roughly about, and where it belongs. They do not replace the original page.
Third, find the route, then look at the page.
When a user asks a question, the agent does not reread the whole file. It first uses the chapter map to decide direction: which chapters matter, whether to go deeper, which pages are worth opening.
Once it has candidate pages, a vision model reads the originals, checks amounts, clauses, charts, or layout, and cites the answer back to a specific page.

This is the same three-step idea Knowhere has used all along: discover what might be relevant, navigate along the chapter structure, then return evidence you can trace.
What this gets you
The information stays more complete.
Text extraction is good at “what was written on the page.” Visual understanding also sees “how those things sat together.” That matters most on dense tables, mixed layouts, scans, and drawings.
The boundary of files and use cases gets wider.
Image packets, drawings, scanned archives, and visual reports that used to be awkward in a knowledge base can now be organized, retrieved, and understood. For Knowhere, that is a move from “documents that are mostly text” to “complete materials that include text, tables, images, and layout.”
Errors are easier to correct, and easier to trace.
VISION-MAP does not promise it will never be wrong. A chapter can be attached to the wrong place. The agent can open the wrong page first. Because the source file and the original page are still there, it can check a neighboring page, take another route, or reread the page. The error is not frozen into the knowledge base at ingest. The traceback is also easier for people. Instead of a pile of text that may even be formatted wrong, you get the page, with the answer highlighted. That lowers the work, and the anxiety, of reviewing what the AI produced.
The two tracks cover for each other.
Clean digital files can stay on the text track for faster, easier retrieval. Many formats (Markdown, DOCX, XLSX, and others) also support non-page parsing. Pages that will not parse reliably can go on the vision track and keep their full information. When it helps, the same document can use both. You do not have to force every file through one pipeline.
It is a better fit for agents.
An agent is not searching a keyword and bringing back a few passages. It has to understand the task, pick a direction, enter the right chapter, decide whether to read text or look at the original page, and come back with evidence and a source.
That is close to how a person drives:
The chapter map is the road network. It tells the agent where to go. The page is the camera. It shows what is actually there. The agent is the driver. It chooses the route from the user’s question.
The map answers “where.” Vision answers “what did I see.” Together, that is VISION-MAP.

How to use it in Knowhere
When you upload a document, pick the reading method that matches the file.
If it is a clean, text-first digital file, use text parsing. The agent can search and cite passages quickly.
If it is a scan, a complex report, a drawing, a mixed layout, or anything where the original layout and the numbers matter, use VISION-MAP. The agent follows the chapter map to the relevant pages, then reads the originals.
You do not need to think about how many model calls sit underneath. Ask the way you normally would:
- “Check this year’s contract amounts and payment terms.”
- “On this drawing, what is the relationship between equipment A and pipe B?”
- “Does this report’s conclusion match the charts earlier in the file?”
Knowhere chooses a path from the document and the question, and hands the relevant text or pages to the agent. When you need to verify, you can go from the answer back to the chapter, the specific page, and the file you originally uploaded.
The text track is for reading efficiently. The vision track is for seeing completely.
Those are Knowhere’s two ways of understanding a document. That is also why we built VISION-MAP.
VISION-MAP is still improving.
We are still working on chapter detection, page localization, cross-page understanding, and links between files.
If you are dealing with scanned contracts, complex reports, engineering drawings, or any file that feels like something went missing after it was turned into text, try VISION-MAP in Knowhere. Send us the files that are hardest to read, and easiest to get wrong.
We will keep sharpening these eyes.


