You hand it an image and type what you want to know: what does this chart say, read this receipt, which items on this shelf are on sale. It answers in text, and it is unusually good at dense pages for its size: it can transcribe a full page while keeping the layout, so a two-column scan does not come back as one scrambled paragraph. Ask it to find something, "the red car", "the signature line", and it points to where that thing sits in the image. It follows questions in 16 languages, English and Chinese among them, and can compare several images in one conversation.
The unusual part is where it runs. The model fits in about 3 GB of memory, small enough for laptops and phones, which is why the demo can run it inside the browser itself instead of on a server. For anything private, receipts, medical letters, ID scans, that locality is the feature: the image never leaves your machine.