In plain terms
Picture handing a laptop to an assistant with the words “renew the parking permit” or “find the three cheapest quotes”. The assistant opens the sites, reads what is on them, types into the forms and comes back with the result. A browser agent is that assistant in software. It works through the same pages a person uses, so the website does not have to offer any special connection for it.
Why it matters
Much everyday work happens on websites that offer no API: supplier portals, government forms, booking systems. A browser agent can take over that work without an integration project. Two limits belong in the decision. It is slower and less reliable than a direct connection, and it fails more often as tasks get longer. It also reads whatever a page contains while acting through the user's logged-in accounts, so a hostile page can try to steer it. Vendors say openly that this risk has not been eliminated. Start with tasks where a mistake is cheap and can be undone.
Example
A purchasing team asks a browser agent to collect current prices for 40 items from six supplier websites and to fill a basket at the cheapest one. It finishes in 25 minutes; a buyer used to need three hours. On one site, text hidden in a product page tells the agent to change the delivery address. The agent may only fill the basket, and checkout requires a person's confirmation. The buyer sees the altered address and rejects the order.
Most often confused with
Browser Agent vs. Computer use
Computer use is the broader capability: a model operating any software through screenshots, mouse and keyboard. A browser agent is confined to the browser and can usually read the structure of a page as well as its picture, which makes it faster and more accurate on websites. For a task that lives entirely on the web, a browser agent is the narrower and safer tool; a desktop program needs computer use.
Under the hood
The agent runs a loop: observe the page, choose an action, execute it, observe again. Observation comes from the DOM or the accessibility tree, from screenshots, or from both; actions are navigate, click, type, scroll and read. Three forms exist: a browser with an agent built in, an extension inside the user's existing browser, and a remote or headless browser driven through automation libraries such as Playwright. Products include Perplexity's Comet, Claude in Chrome and Gemini in Chrome. The main risk is indirect prompt injection: page content is untrusted input, while the session holds the user's cookies and logged-in accounts, a combination that can complete the lethal trifecta. Defences are layered: a separate profile with no sensitive logins, site allowlists and blocked categories, confirmation before purchases, sends and deletions, classifiers that screen page content and proposed actions, action logs and step limits. Reliability falls with task length and with obstacles such as CAPTCHAs and pop-ups.