You type an instruction like "export this spreadsheet as HTML and open it in Chrome". The model looks at a live screenshot of your desktop, reasons about what has to happen next, and answers with a concrete mouse or keyboard action. A small runner executes that action, takes a fresh screenshot, and the loop repeats until the task is complete. Because it reads the actual pixels on screen, it is not tied to one fixed layout: when a dialog pops up or a menu moves, it re-plans from what it sees.
It handles long jobs that span several applications, works on both Ubuntu and Windows, and posts strong scores on the standard obstacle courses for desktop agents. Tencent ships it as a 27B model with weights free for commercial use, alongside a lighter 9B version and a "democua" variant that adapts a workflow from a single recorded demonstration you show it.