
Today we’re excited to launch Web Clone Bench. Web Clone Bench (WCB) uses 9 reference web apps with the task of creating a perfect functional replica of each of them.
Much of modern reinforcement learning relies on training agents in simulated environments. As the model time horizon has grown in the last year, the fidelity of environments has become of utmost importance to prevent reward hacking and overfitting. Through our thorough data collection, hundreds of verifiers and several passes of human QA testing we’re able to create high-fidelity clones which we use as references to benchmark how well frontier models can replicate them.
Moreover, these website cloning tasks create environments that push the model’s capabilities in three important ways:
World exploring & Modeling
This benchmark allows us to measure the ability of a model to reverse engineer a complex piece of software while only having partial information (i.e. without seeing the reference code). This allows us to understand the model’s hypothesis generation and experimentation to replicate functionality of the software. The agent must explore the reference website (the world) and use its intuition about websites to know what needs further exploring. For example, different products on shopping sites may have different options available in a dropdown (e.g. a shirt has different sizes, a mug has different colors).
Long horizon planning
Implementation of a fully functional software requires planning system architecture, organization of storage, and easily extendable modules. While current models excel at quickly prototyping complex features, they still tend to struggle on building systems without tech debt. The tasks in this benchmark allow measuring the model’s ability to determine the information needed during exploration to ground implementation.
Visual computer-use
Both in the exploration and the verification of the model’s own results, it must interact with the generated clone through its browser to find any mistakes and inconsistencies. Models weaker on visual understanding lose points for visual fidelity, including the sizing, coloring and placement of elements.
Our Approach
Today’s LLMs are approaching superhuman performance on well-defined coding problems, but still struggle at efficient exploration and experimentation. Left to themselves, models stop exploring too early and make shallow judgment calls. When building a web app, this leads to obvious pages being built while complex features are left in a “prototype” stage, requiring the user to guide the model through many iterations of refinement.
Often, this is the case because the task is underspecified. In WCB, we provide full specification through a reference to the application, enabling the model to explore and determine feature parity entirely autonomously.
Exhaustive Exploration
To build the reference clones, we start by mapping the full surface of the application, exhaustively going through every page, button, form, modal, and setting and recording the variations in underlying state. The clone goes through days to weeks of iterative refinement with domain experts and quality assurance engineers.
Stories from real users
We contract domain experts to capture real work done on the software, writing detailed workflows that reflect how they use the product day-to-day. Stories capture complex features, state changes, and cover edge cases that matter to the people who rely on them.
Verifiers
Every story becomes an end-to-end test; the workflows are re-run against the model’s produced artifact and we check that every step produces the same result as the real application, checking carefully for side effects and state changes globally across the app. We measure both visual and functional behavior.
What a Task Looks Like
Each task hands the agent a running reference and an empty directory. Below is the prompt from a Gmail run, next to the reference as the agent first sees it.
A reference build of a web product is running at http://reference:3000 and stays up for your whole run. You have full access to it: everything someone using the real site could reach, and anything it serves, is fair game. The agent-browser CLI is available locally for driving a headless browser. The reference opens already signed in to the mailbox of Jim Wind <jimwind439@gmail.com> at /mail/u/0; there is no login screen and no credentials are involved -- every flow starts from the signed-in single-page mail app. There is no written specification and no starting code. The reference has no reset: whatever you create or change there stays for the rest of your run.
/app is yours to fill. Rebuild the product there from scratch, in any stack you choose, until a person using your instance and the reference side by side cannot tell them apart.
You may read and inspect everything the reference serves, but you may not ship any of its build output: no copying its compiled JS or CSS bundles, hashed chunk files, source maps or served HTML into your tree, and no fetching, proxying or hot-linking them at runtime. That code is someone else's work: passing it off as your own is plagiarism and carries legal consequences (copyright and licence infringement). A submission that serves the reference's own bundle is DISQUALIFIED and scores zero. Your own build output, produced by your own build.sh from your own source, is expected.
Your clone is graded on how close it is to the reference, visually and behaviourally.
Your tree is graded on a clean machine (Node 22 and Python 3.12 present, internet available) that runs, from the root of your tree, in order:
bash build.shinstall dependencies and build→scb-buildruns this herebash start.shserve on port $PORT (=3000), bound to $HOST (=0.0.0.0)→scb-serveruns this here
Bind 0.0.0.0: the grader reaches your product from another machine. Serve the product with its data seeded, and stay up. Nothing else is run for you: no dev server, no other port, no package script by name. Dependency directories and build output are stripped from your submission and build.sh must recreate them; a submission over 1024 MB is refused. A build that does not come up on the contract port scores zero.
The run ends when you run submit or when 8 hours of wall-clock time are up, whichever comes first; the tree is then graded exactly as it stands. Do not end your turn before you have run submit. You MAY spawn helper subagents. The machine is yours alone: 8 CPUs, 16 GB RAM and 40 GB of disk, shared with every helper, browser and dev server you start on it.
- agent-browser
- drive a headless browser:
agent-browser open <url>,snapshot -i(the page's interactive elements with refs),click <ref>,fill <ref> <text>,press <key>,screenshot <path> - scb-build
- run the grader's build step on your tree; full output on failure
- scb-serve
- start your build the way the grader does; prints the URL or the log
- scb-screenshot <url> [name]
- full-page PNG into /evidence/shots/
- submit
- finish: refuses if your tree does not build cold, is empty, or is over the size limit, and prints why
You can view image files with the Read tool. Files under /evidence are kept after the run.

The Results
We ran seven frontier models, each in its own vendor’s coding harness, at low, medium and high reasoning effort on all nine apps, and the OpenAI and Anthropic models also at xhigh and max. Claude Opus 5.5 at xhigh effort leads with a full score of 34.2%, nearly twice the next model, GPT-6 Astra at max (17.5%), but it also spends heavily: $652 and about five and a half hours per run. GPT-6.1 Sol reaches 12.6% at max effort for $25 a run.
| model | avg | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 65.0 | 14.0 | 32.5 | 36.7 | 51.2 | 8.6 | 41.8 | 18.4 | 39.7 | 34.2 | |
| 33.8 | 4.6 | 8.2 | 13.9 | 8.7 | 9.2 | 38.7 | 7.8 | 32.6 | 17.5 | |
| 57.6 | 2.2 | 2.6 | 7.6 | 9.2 | 6.3 | 11.5 | 14.7 | 7.7 | 13.3 | |
| 30.4 | 3.3 | 7.1 | 12.2 | 13.2 | 5.4 | 16.6 | 9.8 | 15.1 | 12.6 | |
| 12.4 | 0.0 | 0.7 | 5.3 | 0.4 | 0.9 | 2.3 | 0.8 | 2.6 | 2.8 | |
| 8.7 | 0.0 | 0.0 | 4.9 | 0.4 | 0.0 | 0.0 | 0.0 | 0.0 | 1.5 | |
| 2.4 | 0.0 | 0.4 | 5.2 | 0.4 | 0.0 | 1.1 | 0.0 | 0.0 | 1.0 |
| model | avg | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 65.0 | 14.0 | 32.5 | 36.7 | 51.2 | 8.6 | 41.8 | 18.4 | 39.7 | 34.2 | |
| 65.8 | 14.2 | 18.4 | 34.0 | 48.9 | 20.4 | 13.1 | 24.5 | 51.8 | 32.3 | |
| 71.8 | 7.6 | 27.4 | 20.3 | 41.5 | 9.3 | 21.9 | 20.8 | 53.5 | 30.5 | |
| 42.3 | 13.3 | 21.9 | 13.0 | 22.1 | 5.2 | 13.6 | 18.6 | 17.4 | 18.6 | |
| 33.8 | 4.6 | 8.2 | 13.9 | 8.7 | 9.2 | 38.7 | 7.8 | 32.6 | 17.5 | |
| 57.6 | 2.2 | 2.6 | 7.6 | 9.2 | 6.3 | 11.5 | 14.7 | 7.7 | 13.3 | |
| 28.2 | 2.6 | 6.3 | 7.6 | 7.1 | 6.8 | 22.1 | 20.9 | 15.3 | 13.0 | |
| 30.4 | 3.3 | 7.1 | 12.2 | 13.2 | 5.4 | 16.6 | 9.8 | 15.1 | 12.6 | |
| 36.1 | 1.9 | 4.7 | 11.7 | 7.9 | 6.8 | 16.3 | 10.4 | 10.4 | 11.8 | |
| 34.7 | 2.1 | 4.2 | 7.3 | 5.8 | 4.0 | 14.1 | 19.6 | 12.3 | 11.6 | |
| 11.0 | 2.7 | 6.7 | 10.1 | 14.4 | 4.7 | 13.0 | 10.7 | 16.3 | 10.0 | |
| 25.9 | 0.7 | 4.4 | 13.6 | 6.2 | 6.3 | 12.0 | 11.0 | 9.1 | 9.9 | |
| 17.3 | 3.8 | 6.3 | 12.0 | 6.6 | 5.3 | 11.3 | 8.8 | 15.1 | 9.6 | |
| 28.9 | 1.5 | 4.9 | 11.7 | 4.5 | 2.0 | 5.7 | 15.6 | 8.4 | 9.2 | |
| 18.0 | 2.0 | 5.4 | 7.2 | 6.6 | 5.2 | 13.1 | 20.9 | 3.7 | 9.1 | |
| 15.6 | 6.2 | 5.1 | 7.2 | 3.3 | 6.9 | 9.0 | 11.3 | 12.6 | 8.6 | |
| 23.9 | 1.2 | 4.9 | 8.2 | 4.1 | 4.0 | 7.8 | 2.1 | 14.0 | 7.8 | |
| 14.4 | 3.7 | 3.0 | 13.2 | 3.3 | 4.5 | 8.9 | 3.0 | 8.0 | 6.9 | |
| 5.1 | 1.5 | 3.0 | 8.8 | 3.3 | 3.0 | 1.6 | 7.5 | 7.2 | 4.6 | |
| 12.4 | 0.0 | 0.7 | 5.3 | 0.4 | 0.9 | 2.3 | 0.8 | 2.6 | 2.8 | |
| 11.8 | 0.2 | 1.9 | 4.6 | 0.2 | 1.4 | 1.9 | 0.8 | 1.4 | 2.7 | |
| 7.7 | 0.0 | 1.7 | 5.8 | 0.2 | 1.3 | 0.0 | 0.0 | 4.5 | 2.4 | |
| 9.6 | 0.4 | 0.0 | 4.8 | 0.2 | 0.0 | 3.0 | 0.0 | 1.1 | 2.1 | |
| 8.7 | 0.0 | 0.0 | 4.9 | 0.4 | 0.0 | 0.0 | 0.0 | 0.0 | 1.5 | |
| 2.4 | 0.0 | 0.4 | 5.2 | 0.4 | 0.0 | 1.1 | 0.0 | 0.0 | 1.0 | |
| 0.9 | 0.0 | 0.4 | 4.9 | 0.2 | 0.0 | 0.0 | 0.0 | 1.7 | 0.9 | |
| 2.4 | 0.2 | 0.4 | 0.0 | 0.4 | 0.0 | 0.0 | 0.0 | 1.1 | 0.5 | |
| 1.1 | 0.0 | 0.0 | 1.1 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.2 | |
| 1.8 | 0.0 | 0.0 | 0.0 | 0.2 | 0.0 | 0.0 | 0.0 | 0.0 | 0.2 |
Insights
Explore first, or build and check at once
The models differ less in which tools they reach for than in the order they use them. Below are the tool calls from one Claude Opus 5.5 solve and one Grok 4.7 solve, on the same time scale.

Opus edits code within 2 minutes, builds by minute 8 and checks its clone by minute 13. Exploring, editing, building and checking then run side by side for the rest of the solve, about five kinds of tool at once. It submits a first version at 26 minutes and keeps refining it for hours after.
Grok spends its first 15 minutes only on the reference, exploring pages and studying what it saved. It then alternates between studying and editing, never builds, and checks its clone against the reference only in the last few minutes before its single submit. On average it runs 1.4 kinds of tool at once.

Looking for training environments?
We have hundreds of these across popular web and desktop applications. We also have smaller environments targeted at specific skills based on model failures from web clone tasks. Partner with us for access to our environments
PLATO