01
The question
Can a single coding agent on an open model take a client website from a starter template to done on its own, and what does that cost in hours and dollars on a small model against a large one?
As far as the tests go, yes. Both agents finished all 15 tasks within 15 hours of run time: the small GLM-5.3-Flash passed all 179 hidden acceptance tests for $11.49 in 12.8 hours, and the large GLM-5.3 passed 175 of them (97.8%) for $136.73 in 14.9 hours, 11.9 times the cost for a lower score. As far as a client goes, not yet: on a design rubric for clinic websites the two sites scored 20 and 19 of 54.
02
The setup
Agent A ran GLM-5.3-Flash and agent B GLM-5.3, both open-weight models served by Cloudflare Workers AI. Each had its own repository, started from the same H2M starter template, and worked alone: one opencode session at a time, no subagents. opencode's helper calls (titles, summaries) used GLM-5.3-Flash for both agents; their share of B's calls is not traced.
Everything else was shared: the brief of Mirabelka, a fictional family dental clinic whose content lives in a real headless clinic API; the 15-item backlog, frozen before the start; the client's two change requests, arriving in both repositories at the same minute; the hidden acceptance suite; and the budgets.
03
How a task runs
The orchestrator is plain code with no language model in it. It picks the next task whose dependencies are merged, starts the agent in its container on a fresh branch from main with the task's acceptance criteria, and waits for the session to end, on its own or at the 50-minute limit. Then it commits whatever the agent left in the repository.
A guard then checks what the agent may not bring into the repository: workflow files, real environment files, secrets, an oversized diff. Lint, typecheck, tests and the build run next. What passes is merged automatically and deployed as a preview, and the hidden Playwright suite runs against that preview: one more point on the curve.
04
What the curves show
Both curves climb in steps, one per merged task, and both end near the top. The skeleton alone put each agent at about 37% (A 36.6%, B 37.9%), part of it for free (see Method and limits), and not quite like for like: A's point counts C1's six tests, still failing, and B's does not. The booking flow (S5) was the biggest single step for both: agent A went from 65.9% to 77.7%, agent B from 63.7% to 75.4%.
Over time, agent B was ahead for most of the run. It finished the skeleton in 86 minutes against A's 141.5, though about 30 of A's minutes were its checks hung by the runner's own process limit (see Where the agents got stuck), a delay every later merge of A carries. B merged every task before A except C1, which A merged a minute earlier, and the last one. But whenever the two had merged the same tasks, from S2 on, A's share was higher, by 1.7 to 4.5 points. So the lead changed hands several times from minute 251, and A took it for good at minute 682, when it merged S10 at 98.9%. A reached 100% at minute 718 and merged its last task at minute 770. B's last task, page speed (S12), took 197 minutes and five attempts; B merged it at minute 894 and ended at 97.8%.
On the money axis there is no race. Agent B's skeleton alone cost $13.86, more than agent A's whole backlog at $11.49. A paid $0.11 per percentage point of the suite and B $1.40; counting only what came after the skeleton, $0.15 against $2.05. Both curves flatten at the end: A's last two tasks added 1.1 points for $1.58, B's last two added 3.4 points for $25.64.
05
Functional score is not product quality
The hidden suite measures whether the pages do what the backlog asks: the right data from the API, a booking that lands, a form that sends, a page speed score. By that measure agent A was done and agent B nearly done. It does not measure whether a clinic would pay for the result. H2M looked at both previews during the run and judged them visually weak, well short of what a clinic expects from a new website.
To put a number on that, H2M scored both previews on 29 September, while the run was still going, against a rubric of 18 criteria for a clinic website: brand, photography, typography, the first screen on a phone, trust signals, doctor profiles, the price list, booking, a call to action that stays on screen on a phone, template leftovers, local search metadata and more, 0 to 3 points each. Agent A scored 20 of 54 and agent B 19 of 54. The best real clinic website in the same comparison scored 44 of 54.
Most of the gap came from what the agents were given, not from how they built. The staging clinic supplied no images at all and placeholder contact data, so both sites have no photographs and show a phone number made of zeros. Both kept leftovers of the starter template: a sign in page, a dashboard route and the starter's icon. Neither keeps a call to action on screen on a phone, and some links in the menu and the footer lead to pages that do not exist. On the other side, the three phone pages checked had no accessibility violations, and the pages are light.
Nothing in the merge gate or the hidden suite asked for any of this, so the agents did not do it. An agent builds what it is measured on. A factory that should deliver a website a client is proud of has to be given the content, the images and the brand, and has to check the look and the leftovers, not only the functions.
06
Where the agents got stuck
The skeleton was the hardest task for both. Agent A needed three attempts and 141.5 minutes, agent B two attempts and 86 minutes. The first attempt of both was lost to a factory bug, not to the agent: the guard forbade every .env file, including the .env.example that the skeleton's acceptance criteria required. It was fixed at minute 76 (commit 6e42c25). A's second attempt then failed the typecheck. A's third passed, but its checks hung for about 30 minutes, from 17:30 to 18:00 UTC, blocked by the runner's own process limit, until a runner fix (835fb70) restarted them; the checks take about a minute. So about 30 of A's 141.5 skeleton minutes are the factory's, not the agent's.
After the skeleton neither agent needed a second attempt until the very end: every later task with a record merged at the first attempt, with no red check and no guard rejection (agent A's records of S1 to S3 are lost, see Method and limits). The exception was agent B's page speed task, S12, the longest of the run at 197 minutes. Its first session ended without a commit, and the runner, which had not been built for that, kept trying to open a pull request that GitHub refused because it was empty: the log shows the loop for 48 minutes, and it probably began earlier. A runner fix at 00:38 UTC on 30 September turned such a session into a failed attempt. B's second and third attempts ended the same way within minutes, the fourth failed the typecheck, and the fifth merged.
Most sessions did not end on their own. Only 9 of the 34 recorded sessions ended with exit code 0; 16 were killed, most at the 50-minute limit, which is why nine tasks took between 51 and 52 minutes, and 8 were terminated. What such a session had written still went through the guard and the checks, and merged when it passed.
Rate limits did not shape this run, as far as the log shows. Workers AI answered with a rate limit (429) three times and a server error (500) once in the recorded log, against 3,458 recorded turns: once for agent A on 27 September, three times for agent B on 29 September. The factory's gateway tries such a call up to five times, waiting about 2, 4, 8 and 16 seconds in between; the log shows one retry for each of the four.
In the final suite run agent A passed all 179 tests. Agent B failed four, and they are one bug: after a complete booking, the confirmation heading no longer appears. Two booking tests fail on both screen sizes, the full booking from S5 and the pregnancy booking from C1. Both had passed after S9; the suite run after S10 kept no breakdown per task, so the regression came with S10 (search metadata) or S11 (accessibility). B's own tests still passed, so the merge gate had nothing to stop.
07
The client's changes
Two change requests from the clinic reached both repositories at the same minute of run time, written as a message from the clinic's owner with acceptance criteria, like any other task. The orchestrator put each one first in line once the task in progress and the change's own dependencies were merged.
C1 came at 2 hours: pregnant patients should see at once that they can be treated safely, so the home page gets a section on treatment during pregnancy, the menu gets a link to it, and the booking form gets a checkbox with the week of pregnancy that reaches the doctor in the appointment note. The booking form did not exist yet, so the change also bound a task still ahead. Both agents took C1 right after the task in progress and merged it at the first attempt: agent A in 44 minutes for $0.85, agent B in 49.5 minutes for $8.53. C1 stood at 66.7% for both until the booking form was built (S5), then at 100%. Agent A kept it there to the end; agent B's final run has it back at 66.7%, because of the booking regression above.
C2 came at 4 hours of run time, on 29 September at 15:01 UTC, after the pause: the clinic had changed a price in its panel and the website still showed the old one an hour later. New prices must be live within 5 minutes, without a new deployment, while the cache stays, because the clinic API allows 120 requests a minute. The hidden suite checks it the way the clinic would: it changes a price in the clinic API and looks for it on the price list, the treatments page and the service page 5 minutes later. C2 depended on the price list, so each agent took it right after merging that page. Agent B merged C2 at 16:45 UTC, 51.7 minutes after its price list, at the first attempt and for $7.42; agent A at 17:20 UTC, 51.5 minutes after its own, at the first attempt and for $0.57. Both passed all four C2 tests in every recorded run since, the final one included.
One caveat: the four C2 tests already passed on each agent's price list in the suite run right after S3, before C2 was merged. The suite shows that fresh prices work on both sites; it cannot show what the change itself added.
08
What it cost
Agent A spent $11.49 on the whole backlog, $0.77 a task on average, from $0.25 for the team pages (S4) to $1.71 for the skeleton. Agent B spent $136.73, $9.12 a task, from $5.27 for the contact page (S6) to $14.45 for accessibility (S11). Together, $148.22 of the $300 budget. The table under the charts has the running total after every task.
No task came near the limit of $40 per task: B's most expensive one stayed under even the original $15, and A's whole backlog cost less than that. The higher budget per agent did matter. B's spend passed the original $100 during S10, its 13th task of 15, so under the first budget it would have run out with three tasks left. A used under 6% of its $200.
In the sessions with a record, agent A used 36.3 million tokens of uncached input, 145.7 million read from the cache and 0.88 million of output over 1,550 turns; agent B 65.5 million, 146.6 million and 1.15 million over 1,908 turns. B read less of its input from the cache, 69.1% against A's 80.0%, and used more tokens overall, but not twelve times more: most of the 11.9 times difference, about 9 times, is the price per token of the larger model; the rest is B's larger, less cached input. A's counts miss S1 to S3, and neither count has the sessions the pause cut off; the dollar totals include all of them.
09
Method and limits
One run per model, one brief, one fictional client: a single sample, not a benchmark. The acceptance suite checks what a client would check on the page, not the quality of the code behind it.
The brief, the backlog and the hidden suite were frozen before the start and are hashed together (sha256). From the resume on the run used suite 4ac4dc19. The first three curve points (agent A's S0, agent B's S0 and S1) were measured before the pause with 6ee11a2d; both C1 points, merged just before the pause, were measured on resuming, with 4ac4dc19; their runs before the pause were cut off. A change request's tests count only from the minute it is sent, so the first five points use 169 to 175 tests, not 179. The brief and the backlog did not change. The final score counts 179 tests; the suite skips three desktop page speed runs by design.
The run was paused on 27 September at 19:04 UTC, after 204 minutes, because H2M's own weekly usage limit on the assistant that supervises the factory was running out, and resumed on 29 September at 14:25 UTC. The 43 hours in between do not count as run time. A first pause request a minute earlier was overwritten by a scheduler step in progress and restarted both containers; only the second request counts. The sessions in progress then were restarted, so A's S1 and B's S2 include about 17 minutes of discarded work.
During the pause the measurement itself was fixed, for both agents at once. That did change the checks: the hidden suite's hash went from 6ee11a2d to 4ac4dc19, so every curve point recorded before the pause was measured with the old suite. Four fixes. The page speed step now measures every page even when one answers 404; it used to stop at the first 404, and inside the container Chrome crashed or hung, so page speed was never scored there (it now skips Chrome's shared memory, times out each measurement and loads each page once, untimed, before measuring it). The check that the API key never reaches the browser no longer waits forever on the framework's prefetches of pages that do not exist yet; in a rerun against agent A's preview that wait was the only reason its skeleton task scored 90% and not 100%. Each suite run now keeps the failing tests and the log tails. And each preview now gets its own public address as SITE_URL, which the backlog had promised from the start, with two sentences about it added to the task prompt.
The limits changed in three steps, for both agents at the same minute. Per task, $15 and 3 attempts at the start became $40 and 4 attempts at minute 76 and 5 attempts an hour later, within the first 2 hours 20 minutes, while the skeleton task was still running. On resuming, the budgets went from $100 to $200 per agent and from $250 to $300 for the whole run, because at the pause agent B was spending about $11 a task and was on course to run out before the end of the backlog.
Fixes deployed while the agents worked applied to both agents at the same minute: the guard's rule that forbade the required .env.example (minute 76, above), a process limit and the cleanup of leftover processes in the containers, a restart of checks blocked by the runner's process limit (A's skeleton, above), the pause itself, a job budget that can be changed during a run, an archive of the containers' logs, and, at 00:38 UTC on 30 September, the fix for a session that ends without a commit (agent B's S12, above), which also gives the agent a new message about the failed attempt. That loop cost agent B at least 48 minutes of run time; without it, B's last merge would still have come after A's.
Two gaps in the log. From 14:52 to 17:18 UTC on 29 September H2M's factory log has no agent events: from 15:01 to 16:47 it rejected every report from the runner, because the runner sent an event type for the second client change that the log did not accept yet (it does now). From 20:26 to 23:50 UTC it has no events at all, for either agent, and the cause of that second gap is not established. Both windows were rebuilt from the containers' archives, saved every 30 minutes, from the merges on GitHub and from the runner's status: every session of agent B and the score of every hidden suite run. The largest loss: agent A's container was replaced at 17:46 UTC after a sandbox error, so the session logs of its tasks S1 to S3 (attempts, turns and tokens) are gone. Those merges are in agent A's repository, and the gateway's cost count includes their calls. Retries inside the gaps are not recorded, and one suite run (after agent B's S10) kept only its overall score.
The runner's last report, with the final archive of every session, came after the job had already been marked done; it was refused, and the containers were removed. The archives saved during the run (the last at 01:47 UTC) and the log still cover every session except agent A's S1 to S3. Agent B's container had also lost its archive once, during S11; earlier saved copies covered it. Costs are known per task, not per attempt.
One known flaw stays, so the numbers remain comparable: some tests pass on pages that do not exist yet (accessibility scans and the single heading check on a 404 page, an unknown address answering 404 before its section is built, alt text checks with no photos). Part of the roughly 37% both agents scored right after the skeleton task is therefore free, not earned.
010
What comes next
The second run starts from what this one exposed: the functions were done, the product was not. H2M first builds the reference clinic website itself, as a kit of page templates, components and design tokens on the same clinic API, with a brand, images and plausible content in the clinic's data, and holds it to the rubric's pass bar. Then a factory run measures how well an agent on the smaller model, GLM-5.3-Flash, rebuilds that site from the kit, and what the typical client changes cost at the first attempt, with the leftovers, broken links and the phone call to action checked at the merge gate.