You give it a picture of a document page, a scan, a PDF page, or a photo, and it writes out the whole page as Markdown in natural reading order. Tables come back as proper HTML tables instead of scrambled text, formulas come back as LaTeX, and figures are marked with their exact position on the page so you can crop and keep them. The output is ready to paste into notes, docs, or a knowledge base without reflowing paragraphs by hand.
The model is a 0.8B build on Qwen3.5 from Alibaba's Ovis team, and it scored 96.58 on OmniDocBench v1.6, the widest document parsing test suite. That made it the first single end-to-end model to lead that board, which pipeline systems had owned until now. For you that means one model call replaces a whole parsing stack, with less to install and fewer places for errors to creep in.