In plain terms
Most systems talk to AI through an API. Many business applications have none: an old ERP screen, a supplier portal, a desktop program. Computer use lets the model work those anyway, by reading a screenshot and deciding where to click, one step after another.
Why it matters
It extends automation to the long tail of software that was never built to be automated, without an integration project. It is slower and less reliable than an API call, and an agent that reads whatever is on screen can be misled by it, so it suits supervised, well-scoped tasks first.
Example
Each morning an agent opens a supplier's web portal, logs in through a company-managed session, downloads yesterday's invoices and enters the totals into an accounting screen that has no import function.
Most often confused with
Computer Use vs. RPA
RPA follows recorded steps: click here, type there. It is fast and breaks when a button moves. Computer use looks at the screen each time and works out what to do, so it survives layout changes but is slower, costs more per run and can make judgment errors.
Under the hood
The model runs in a loop: it receives a screenshot (sometimes with the accessibility tree), returns an action such as click at coordinates, type text or press a key, the harness executes it and sends the next screenshot. Accuracy depends on resolution, UI density and the length of the task. Risks are specific: on-screen text is untrusted input (prompt injection through a web page or a document), and actions happen with the logged-in user's rights. Standard precautions are a dedicated browser profile or virtual machine, allowlisted sites, and human confirmation before purchases, sends and deletions.