Given a screen and a closed set of element and action options, pick the right fill, check, click, or skip.
Computer-use agents make many small decisions per screen. CUA-Bench S1 measures whether a decision model picks the right element and action from a closed list, using the page structure or a screenshot.
Each task gives the model the screen and a closed set of (element, action) options. We report accuracy on the general-decision set as a percentage.
1 model(s) with undisclosed parameter counts not shown. Most closed-source labs do not publish model size.
Not enough scored models yet.
Cua’s own CUA-S1 4B scores 88.7% on the general-decision set, ahead of Jev at 66.7%.
Based on score correlations across our database.