Lesson 045 covered agents that operate websites.
What if the target is a desktop application or a legacy system with no useful API?
That is where computer use appears.
Computer use operates the GUI
A graphical user interface (GUI) is the visible layer of windows, buttons, menus and text fields that people normally operate.
A computer-use agent can follow a loop like:
Inspect screen
→ understand current state
→ choose next action
→ click / type / scroll
→ inspect the new screen
→ continue
This is a perception-action loop.
Screenshots can act as the agent’s eyes
Many systems provide a screenshot or similar visual representation. A vision-capable model identifies buttons, fields, errors and current application state.
It can then request actions conceptually similar to:
click(x, y)
type("hello")
scroll(0, 600)
The exact tool format differs by platform, but the idea is the same: translate visual understanding into mouse and keyboard operations.
APIs remain better when they exist
Lesson 018 introduced structured APIs.
An API might create a calendar event with one structured request. A GUI agent may need to open the app, find “Create,” choose a date, enter text and click Save.
APIs are generally:
- faster,
- more stable,
- easier to permission,
- less sensitive to layout changes.
Computer use wins on generality: it can work with systems that were never designed for AI integration.
GUI automation is surprisingly fragile
The visible interface changes constantly:
- popups cover controls,
- loading becomes slower,
- resolution changes,
- a button moves,
- focus lands in the wrong field,
- a drag-and-drop action behaves differently.
A single incorrect step can put the agent in a completely different state.
The agent must re-observe after actions
A robust loop should not assume:
I clicked Save → therefore it succeeded
It should inspect the next screen and verify that a success state actually appeared.
That feedback loop is what lets an agent recover from some interface surprises instead of blindly replaying a fixed macro.
Permissions become more dangerous at the GUI layer
Lesson 035 matters even more when the agent can click, upload, download, send or delete.
Limit:
- which sites or applications it can operate,
- which files it can read,
- which data it can type,
- which actions require confirmation,
- which environment is sandboxed.
Do not equate “desktop access” with “all access”
Safer patterns include isolated virtual machines or cloud environments, low-privilege accounts, only mounting necessary files, logging actions and requiring human approval for high-impact operations.
That is the same defense-in-depth idea from Lesson 036.
One thing to remember
Computer use lets AI inspect a GUI and repeatedly control it with mouse and keyboard actions. It can reach software without APIs, but its generality comes with greater fragility and security risk, so isolation, least privilege and verification matter.
Comments
Questions, reactions and useful additions are welcome here.
No comments yet. Be the 1F.