Skip to main content
Use display controls to capture screenshots of a macOS Devbox and click anywhere on its screen. Together they let an agent see what an app shows and act on it, for example to check that a page renders or to click through a flow that has no command-line equivalent.
Display controls are available on macOS Devboxes only. On a Linux Devbox, they reject with DevboxDisplayUnavailableError. See Create a Devbox.

Take screenshots

Capture the full screen as a PNG. The screenshot is 2560 by 1600 pixels. A display call starts a stopped Devbox and keeps its connection open for later screenshots and clicks. See display.screenshot() for its options and return value.
The result holds the PNG bytes in png, the screen size in width and height, and the desktopName reported by the display server. Pass the PNG to your model as an image.
A screenshot of a macOS Devbox desktop with namespace.so open in Safari
Each screenshot takes about a second. open returns before an app has finished loading, so wait briefly, or take screenshots until the screen stops changing, before acting on what you see.

Click the mouse

Click at a point on the screen. Coordinates are in screenshot pixels, with (0, 0) at the top-left corner, so a point read from a screenshot can be passed as is. See display.click() for its options, including the mouse button.
click() returns as soon as the click is sent, before the app reacts. Wait briefly, then take a new screenshot to confirm the result before choosing the next action.
The screenshot after the click, with the namespace.so contact page open in a new Safari tab

Agent loop

Screenshots and clicks combine into a loop that lets an agent work on an app until a task is done:
  1. Start the app with devbox.exec(), for example open -a Safari https://namespace.so.
  2. Take a screenshot to capture what the screen shows.
  3. Send the screenshot and the task to a vision model, and ask it for the next action.
  4. Run the action the model returns, such as a click, then wait for the app to react.
  5. Repeat from step 2 until the model reports that the task is done.
The loop below follows that pattern. decideNextAction() stands for your model call, for example to the Claude or OpenAI computer use tools. It receives the screenshot and returns the next action: a click, an app or URL to open, or done.
Tell the model the screen size, 2560 by 1600 pixels, so that the coordinates it returns match the screenshot. The open action shows how to mix commands into the loop. For anything other than clicks, such as launching apps or opening URLs, run a command with devbox.exec(). Commands are usually faster and more reliable than working through the screen.
Last modified on October 2, 2026