The desktop is now part of the agent route.
GitHub released computer use in public preview for Copilot CLI and the Copilot app on macOS and Windows. It can read accessible app content and screenshots, click controls, edit text, press keys, scroll, drag, and move through workflows across applications.[1]
This reaches the work that APIs and command lines leave behind. A presentation, an expense form, or a legacy admin screen can become part of one agent session. GitHub's own documentation still prefers an API, MCP server, terminal command, filesystem tool, or dedicated browser tool when one can do the job. Those routes return more structured information and tend to behave more predictably.[2]
Use pixels for the gap. Do not turn every structured door into a screenshot.
Approval has a time axis.
Computer use starts disabled. When an app requests control, a person can allow it for the current session, save approval for later sessions, or deny it. A saved approval is local and shared by Copilot CLI and the Copilot app on that computer. Deny rules still take priority.[2]
Here is the detail worth putting on the dashboard: removing a saved app approval blocks future sessions, but it does not revoke access already granted to a running session. Stopping the current operation is a separate action. In the CLI, GitHub documents pressing Esc twice. In the app, use Stop or Esc.[2]
That means one generic lock icon lies by omission. A useful control surface needs two indicators. One shows whether the application will ask again next time. The other shows whether this session can act now.
The screen is both input and cargo.
On macOS, the feature asks for Accessibility permission to operate controls and Screen Recording permission when visual context is needed. The screen may include personal, financial, enterprise, or third-party information. GitHub tells users to limit computer use to apps and tasks whose visible content they are willing to provide as context.[2]
Older computer-use research from Anthropic explains the same exposure in mechanical terms. A model looks at screenshots, estimates cursor movement, and acts on the next view. Anthropic also identifies on-screen prompt injection as a risk and describes the screenshot sequence as a flipbook that can miss short-lived changes. Its early system could click the wrong control or interrupt the wrong process.[3]
The specific GitHub product should be judged by GitHub's current documentation. The older Anthropic report is not evidence about Copilot's implementation or reliability. It is useful because it shows why visible content, action authority, and verification remain separate concerns across computer-use systems.
Write a desktop route before granting one.
Name the target application and the intended result. List the app regions the task may inspect. Keep credentials, unrelated windows, notifications, and other people's data outside that view. Choose session approval unless repetition has earned a saved rule.
Then write both brakes. One stops the active operation. The other removes future approval. Finish with a result check that does not depend on the same visual guess that drove the click. Read the saved file, query the destination, reopen the record, or ask a person to inspect the high-impact result.
VS Code 1.140 is also moving agent behavior into a dedicated host process that can connect the same session across windows.[4] The wider product direction is clear: sessions now travel across surfaces and applications. Permission status has to travel with equal clarity.