ScreenSpot Pro

A GUI-grounding benchmark for 1,581 instructions in full-screen, high-resolution professional interfaces. It tests where a target is, not whether an agent can finish the surrounding workflow.

ModelsModels with a result.18
Top resultBest result on this test.92.7%GPT-6 Astra
Top-3 spreadPoints from first to third.7.3 pts
YearYear of release.2025

Result and price

020406080100$1$3$10$30$100RESULTOUTPUT PRICE PER 1M TOKENS · LOG SCALE

Results

18 results
#ModelResultThe score from the source.
03GPT-5.485.4%
06Muse Spark84.1%
15Holo2-8B58.9%
17Holo2-4B57.2%