Display controls are available on macOS Devboxes only. On a Linux Devbox, they reject with
DevboxDisplayUnavailableError. See Create a Devbox.Take screenshots
Capture the full screen as a PNG. The screenshot is 2560 by 1600 pixels. A display call starts a stopped Devbox and keeps its connection open for later screenshots and clicks. Seedisplay.screenshot() for its options and return value.
- TS SDK
png, the screen size in width and height, and the desktopName reported by the display server. Pass the PNG to your model as an image.

Click the mouse
Click at a point on the screen. Coordinates are in screenshot pixels, with(0, 0) at the top-left corner, so a point read from a screenshot can be passed as is.
See display.click() for its options, including the mouse button.
- TS SDK
click() returns as soon as the click is sent, before the app reacts. Wait briefly, then take a new screenshot to confirm the result before choosing the next action.

Agent loop
Screenshots and clicks combine into a loop that lets an agent work on an app until a task is done:- Start the app with
devbox.exec(), for exampleopen -a Safari https://namespace.so. - Take a screenshot to capture what the screen shows.
- Send the screenshot and the task to a vision model, and ask it for the next action.
- Run the action the model returns, such as a click, then wait for the app to react.
- Repeat from step 2 until the model reports that the task is done.
decideNextAction() stands for your model call, for example to the Claude or OpenAI computer use tools. It receives the screenshot and returns the next action: a click, an app or URL to open, or done.
- TS SDK
open action shows how to mix commands into the loop. For anything other than clicks, such as launching apps or opening URLs, run a command with devbox.exec(). Commands are usually faster and more reliable than working through the screen.