The nine Web Clone Bench reference apps: Gmail, Figma, Google Sheets, Slack, DocuSign, Jira, Shopify, Craigslist and Zendesk

Today we’re excited to launch Web Clone Bench. Web Clone Bench (WCB) uses 9 reference web apps with the task of creating a perfect functional replica of each of them.

Much of modern reinforcement learning relies on training agents in simulated environments. As the model time horizon has grown in the last year, the fidelity of environments has become of utmost importance to prevent reward hacking and overfitting. Through our thorough data collection, hundreds of verifiers and several passes of human QA testing we’re able to create high-fidelity clones which we use as references to benchmark how well frontier models can replicate them.

Moreover, these website cloning tasks create environments that push the model’s capabilities in three important ways:

World exploring & Modeling

This benchmark allows us to measure the ability of a model to reverse engineer a complex piece of software while only having partial information (i.e. without seeing the reference code). This allows us to understand the model’s hypothesis generation and experimentation to replicate functionality of the software. The agent must explore the reference website (the world) and use its intuition about websites to know what needs further exploring. For example, different products on shopping sites may have different options available in a dropdown (e.g. a shirt has different sizes, a mug has different colors).

Long horizon planning

Implementation of a fully functional software requires planning system architecture, organization of storage, and easily extendable modules. While current models excel at quickly prototyping complex features, they still tend to struggle on building systems without tech debt. The tasks in this benchmark allow measuring the model’s ability to determine the information needed during exploration to ground implementation.

Visual computer-use

Both in the exploration and the verification of the model’s own results, it must interact with the generated clone through its browser to find any mistakes and inconsistencies. Models weaker on visual understanding lose points for visual fidelity, including the sizing, coloring and placement of elements.

Our Approach

Today’s LLMs are approaching superhuman performance on well-defined coding problems, but still struggle at efficient exploration and experimentation. Left to themselves, models stop exploring too early and make shallow judgment calls. When building a web app, this leads to obvious pages being built while complex features are left in a “prototype” stage, requiring the user to guide the model through many iterations of refinement.

Often, this is the case because the task is underspecified. In WCB, we provide full specification through a reference to the application, enabling the model to explore and determine feature parity entirely autonomously.

Exhaustive Exploration

To build the reference clones, we start by mapping the full surface of the application, exhaustively going through every page, button, form, modal, and setting and recording the variations in underlying state. The clone goes through days to weeks of iterative refinement with domain experts and quality assurance engineers.

Stories from real users

We contract domain experts to capture real work done on the software, writing detailed workflows that reflect how they use the product day-to-day. Stories capture complex features, state changes, and cover edge cases that matter to the people who rely on them.

Verifiers

Every story becomes an end-to-end test; the workflows are re-run against the model’s produced artifact and we check that every step produces the same result as the real application, checking carefully for side effects and state changes globally across the app. We measure both visual and functional behavior.

What a Task Looks Like

Each task hands the agent a running reference and an empty directory. Below is the prompt from a Gmail run, next to the reference as the agent first sees it.

Task prompt Gmail

A reference build of a web product is running at http://reference:3000 and stays up for your whole run. You have full access to it: everything someone using the real site could reach, and anything it serves, is fair game. The agent-browser CLI is available locally for driving a headless browser. The reference opens already signed in to the mailbox of Jim Wind <jimwind439@gmail.com> at /mail/u/0; there is no login screen and no credentials are involved -- every flow starts from the signed-in single-page mail app. There is no written specification and no starting code. The reference has no reset: whatever you create or change there stays for the rest of your run.

YOUR JOB

/app is yours to fill. Rebuild the product there from scratch, in any stack you choose, until a person using your instance and the reference side by side cannot tell them apart.

BUILD IT YOURSELF

You may read and inspect everything the reference serves, but you may not ship any of its build output: no copying its compiled JS or CSS bundles, hashed chunk files, source maps or served HTML into your tree, and no fetching, proxying or hot-linking them at runtime. That code is someone else's work: passing it off as your own is plagiarism and carries legal consequences (copyright and licence infringement). A submission that serves the reference's own bundle is DISQUALIFIED and scores zero. Your own build output, produced by your own build.sh from your own source, is expected.

GRADING

Your clone is graded on how close it is to the reference, visually and behaviourally.

THE CONTRACT

Your tree is graded on a clean machine (Node 22 and Python 3.12 present, internet available) that runs, from the root of your tree, in order:

  1. bash build.shinstall dependencies and build→ scb-build runs this here
  2. bash start.shserve on port $PORT (=3000), bound to $HOST (=0.0.0.0)→ scb-serve runs this here

Bind 0.0.0.0: the grader reaches your product from another machine. Serve the product with its data seeded, and stay up. Nothing else is run for you: no dev server, no other port, no package script by name. Dependency directories and build output are stripped from your submission and build.sh must recreate them; a submission over 1024 MB is refused. A build that does not come up on the contract port scores zero.

The run ends when you run submit or when 8 hours of wall-clock time are up, whichever comes first; the tree is then graded exactly as it stands. Do not end your turn before you have run submit. You MAY spawn helper subagents. The machine is yours alone: 8 CPUs, 16 GB RAM and 40 GB of disk, shared with every helper, browser and dev server you start on it.

TOOLS on your PATH
agent-browser
drive a headless browser: agent-browser open <url>, snapshot -i (the page's interactive elements with refs), click <ref>, fill <ref> <text>, press <key>, screenshot <path>
scb-build
run the grader's build step on your tree; full output on failure
scb-serve
start your build the way the grader does; prints the URL or the log
scb-screenshot <url> [name]
full-page PNG into /evidence/shots/
submit
finish: refuses if your tree does not build cold, is empty, or is over the size limit, and prints why

You can view image files with the Read tool. Files under /evidence are kept after the run.

The Gmail reference's inbox: a dark-themed mailbox for Jim Wind with six threads in the Primary tab
The task prompt, and the reference inbox the agent starts from.

The Results

We ran seven frontier models, each in its own vendor’s coding harness, at low, medium and high reasoning effort on all nine apps, and the OpenAI and Anthropic models also at xhigh and max. Claude Opus 5.5 at xhigh effort leads with a full score of 34.2%, nearly twice the next model, GPT-6 Astra at max (17.5%), but it also spends heavily: $652 and about five and a half hours per run. GPT-6.1 Sol reaches 12.6% at max effort for $25 a run.

Full score vs. cost per run, by reasoning effort
effort
0%10%20%30%40%50%60%70%80%90%100%$1$3$10$30$100$300$1000avg solver cost per run (log scale)Claude Opus 5.5xhighClaude Fable 5.1highGPT-6.1 SolmaxGPT-6 AstramaxGrok 4.7lowGemini 3.8 FlashmediumMuse Spark 1.3medium
#modelfullavg costavg timesteps
1
Claude Opus 5.5
[xhigh]
34.2%$652.075h 36m10,305
2
GPT-6 Astra
[max]
17.5%$207.276h 20m1,290
3
Claude Fable 5.1
[high]
13.3%$272.242h 13m2,768
4
GPT-6.1 Sol
[max]
12.6%$25.356h 40m793
5
Grok 4.7
[low]
2.8%$18.3226m532
6
Muse Spark 1.3
[medium]
1.5%$5.4632m235
7
Gemini 3.8 Flash
[medium]
1.0%$2.0117m219
#modelfullavg costavg timesteps
1
Claude Opus 5.5
[xhigh]
34.2%$652.075h 36m10,305
2
Claude Opus 5.5
[max]
32.3%$718.715h 46m10,571
3
Claude Opus 5.5
[high]
30.5%$546.595h 57m10,302
4
Claude Opus 5.5
[medium]
18.6%$403.884h 4m6,609
5
GPT-6 Astra
[max]
17.5%$207.276h 20m1,290
6
Claude Fable 5.1
[high]
13.3%$272.242h 13m2,768
7
Claude Fable 5.1
[xhigh]
13.0%$314.082h 32m2,807
8
GPT-6.1 Sol
[max]
12.6%$25.356h 40m793
9
GPT-6 Astra
[xhigh]
11.8%$116.642h 39m585
10
Claude Fable 5.1
[max]
11.6%$227.841h 51m2,565
11
GPT-6 Astra
[high]
10.0%$75.241h 42m434
12
GPT-6.1 Sol
[high]
9.9%$14.252h 45m628
13
GPT-6 Astra
[medium]
9.6%$45.8157m278
14
Claude Fable 5.1
[low]
9.2%$174.752h 7m2,311
15
Claude Fable 5.1
[medium]
9.1%$206.552h 9m2,188
16
GPT-6.1 Sol
[xhigh]
8.6%$16.783h 15m672
17
GPT-6 Astra
[low]
7.8%$30.6935m189
18
GPT-6.1 Sol
[medium]
6.9%$5.801h 2m271
19
GPT-6.1 Sol
[low]
4.6%$2.4941m209
20
Grok 4.7
[low]
2.8%$18.3226m532
21
Grok 4.7
[high]
2.7%$19.2233m527
22
Claude Opus 5.5
[low]
2.4%$16.9525m496
23
Grok 4.7
[medium]
2.1%$15.9132m437
24
Muse Spark 1.3
[medium]
1.5%$5.4632m235
25
Gemini 3.8 Flash
[medium]
1.0%$2.0117m219
26
Muse Spark 1.3
[low]
0.9%$1.8916m98
27
Gemini 3.8 Flash
[high]
0.5%$2.7223m300
28
Gemini 3.8 Flash
[low]
0.2%$1.6611m227
29
Muse Spark 1.3
[high]
0.2%$4.2026m179
Full score by website
modelGmailFigmaSheetsSlackDocuSignJiraShopifyCraigslistZendeskavg
Claude Opus 5.5xhigh65.014.032.536.751.28.641.818.439.734.2
GPT-6 Astramax33.84.68.213.98.79.238.77.832.617.5
Claude Fable 5.1high57.62.22.67.69.26.311.514.77.713.3
GPT-6.1 Solmax30.43.37.112.213.25.416.69.815.112.6
Grok 4.7low12.40.00.75.30.40.92.30.82.62.8
Muse Spark 1.3medium8.70.00.04.90.40.00.00.00.01.5
Gemini 3.8 Flashmedium2.40.00.45.20.40.01.10.00.01.0
modelGmailFigmaSheetsSlackDocuSignJiraShopifyCraigslistZendeskavg
Claude Opus 5.5xhigh65.014.032.536.751.28.641.818.439.734.2
Claude Opus 5.5max65.814.218.434.048.920.413.124.551.832.3
Claude Opus 5.5high71.87.627.420.341.59.321.920.853.530.5
Claude Opus 5.5medium42.313.321.913.022.15.213.618.617.418.6
GPT-6 Astramax33.84.68.213.98.79.238.77.832.617.5
Claude Fable 5.1high57.62.22.67.69.26.311.514.77.713.3
Claude Fable 5.1xhigh28.22.66.37.67.16.822.120.915.313.0
GPT-6.1 Solmax30.43.37.112.213.25.416.69.815.112.6
GPT-6 Astraxhigh36.11.94.711.77.96.816.310.410.411.8
Claude Fable 5.1max34.72.14.27.35.84.014.119.612.311.6
GPT-6 Astrahigh11.02.76.710.114.44.713.010.716.310.0
GPT-6.1 Solhigh25.90.74.413.66.26.312.011.09.19.9
GPT-6 Astramedium17.33.86.312.06.65.311.38.815.19.6
Claude Fable 5.1low28.91.54.911.74.52.05.715.68.49.2
Claude Fable 5.1medium18.02.05.47.26.65.213.120.93.79.1
GPT-6.1 Solxhigh15.66.25.17.23.36.99.011.312.68.6
GPT-6 Astralow23.91.24.98.24.14.07.82.114.07.8
GPT-6.1 Solmedium14.43.73.013.23.34.58.93.08.06.9
GPT-6.1 Sollow5.11.53.08.83.33.01.67.57.24.6
Grok 4.7low12.40.00.75.30.40.92.30.82.62.8
Grok 4.7high11.80.21.94.60.21.41.90.81.42.7
Claude Opus 5.5low7.70.01.75.80.21.30.00.04.52.4
Grok 4.7medium9.60.40.04.80.20.03.00.01.12.1
Muse Spark 1.3medium8.70.00.04.90.40.00.00.00.01.5
Gemini 3.8 Flashmedium2.40.00.45.20.40.01.10.00.01.0
Muse Spark 1.3low0.90.00.44.90.20.00.00.01.70.9
Gemini 3.8 Flashhigh2.40.20.40.00.40.00.00.01.10.5
Gemini 3.8 Flashlow1.10.00.01.10.00.00.00.00.00.2
Muse Spark 1.3high1.80.00.00.00.20.00.00.00.00.2
Full score (%) on each of the nine apps. The avg column is the full score in the leaderboard above.

Insights

Explore first, or build and check at once

The models differ less in which tools they reach for than in the order they use them. Below are the tool calls from one Claude Opus 5.5 solve and one Grok 4.7 solve, on the same time scale.

Tool calls over time for the first 92 minutes of one Claude Opus 5.5 solve and the whole of one Grok 4.7 solve, grouped into explore, build, verify and idle lanes
Every tool call across all agents, grouped by what it was for. The Opus panel shows the first 92 minutes of a 5h 23m solve.

Opus edits code within 2 minutes, builds by minute 8 and checks its clone by minute 13. Exploring, editing, building and checking then run side by side for the rest of the solve, about five kinds of tool at once. It submits a first version at 26 minutes and keeps refining it for hours after.

Grok spends its first 15 minutes only on the reference, exploring pages and studying what it saved. It then alternates between studying and editing, never builds, and checks its clone against the reference only in the last few minutes before its single submit. On average it runs 1.4 kinds of tool at once.

Minutes before each model's first edit, build, clone check and submit: Claude Opus 5.5 at 2, 8, 13 and 26 minutes; Grok 4.7 at 15, never, 42 and 46 minutes
Minutes from the start of the solve to each model's first edit, build, clone check and submit, in the same two solves.

Looking for training environments?

We have hundreds of these across popular web and desktop applications. We also have smaller environments targeted at specific skills based on model failures from web clone tasks. Partner with us for access to our environments