Articles/Arch Studio 1.5 Improved Cold-Start Project Records
Arch Studio 1.5 Improved Cold-Start Project Records
In 27 cold-start runs per arm, Arch Studio produced valid project records 55.6% of the time against vanilla Claude's 14.8%.
Arch Studio 1.5 produced valid project records more often than vanilla Claude when a project did not contain a prior example of the requested operation.
The cold-start condition covered 81 runs: 27 with Arch Studio, 27 with vanilla Claude and 27 with a short rules paragraph. Arch Studio passed 15/27 (55.6%); vanilla Claude passed 4/27 (14.8%), a difference of 40.7 percentage points calculated from the unrounded rates. On Opus 5.5, Arch Studio passed all nine runs; vanilla Claude passed two.
Arch Studio also improved on version 1.4.5. On checks shared by both versions, 1.5 completed 81.5% of requests correctly in one turn. Version 1.4.5 completed 40.7%.
The advantage was specific. In a separate evaluation of isolated product-page tasks, Arch Studio scored 0.86 against vanilla Claude’s 0.87. The product improved project-record work, not every task given to the model.
Results at a glance
Arch Studio 1.5 against vanilla Claude and the rules paragraph
| Project condition | Arch Studio 1.5 | Vanilla Claude | Rules paragraph | Difference vs vanilla |
|---|---|---|---|---|
| No prior examples | 15/27, 55.6% | 4/27, 14.8% | 9/27, 33.3% | +40.7 points |
| Prior examples present | 16/27, 59.3% | 12/27, 44.4% | 17/27, 63.0% | +14.8 points |
Arch Studio 1.5 against 1.4.5
| Test | Arch Studio 1.5 | Arch Studio 1.4.5 | Difference |
|---|---|---|---|
| Project-record operations completed correctly in one turn | 44/54, 81.5% | 22/54, 40.7% | +40.7 points |
Isolated product-page tasks
| Arch Studio 1.5 | Vanilla Claude | Difference |
|---|---|---|
| 0.86 | 0.87 | −0.01 |
Experimental design
The evaluation program contained 277 valid model runs at a recorded model cost of $87.26.
| Evaluation | Runs | Cost | Question |
|---|---|---|---|
| Live product-page tasks | 42 | $27.68 | Does Arch Studio improve isolated product research and extraction? |
| Project records with prior examples | 81 | $22.62 | Does Arch Studio preserve record contracts when examples already exist? |
| Cold-start project records | 81 | $21.32 | Does Arch Studio preserve the same contracts without an example to copy? |
| New 1.4.5 comparison runs; 1.5 reused the 54 project-record runs above | 54 | $8.28 | Did version 1.5 improve single-turn completion? |
| Two-turn confirmation protocol | 19 | $7.36 | Does 1.4.5 complete the work after one approval reply? |
| Total | 277 | $87.26 |
The follow-up cost includes $1.12 spent on two discarded smoke tests. Those tests are not included in the 277 valid runs.
Project-record design
| Factor | Levels |
|---|---|
| Project condition | No prior example of the requested operation; prior examples present |
| Evaluation arm | Arch Studio 1.5; vanilla Claude with the same file tools; vanilla Claude with a 167-word rules paragraph |
| Model | Claude Haiku 4.5; Sonnet 5; Opus 5.5 |
| Operation | Task-register update; superseding project decision; FF&E schedule revision |
| Repetitions | Three runs per model, operation and arm |
| Primary outcome | Pass only if every scored record-contract check passes |
| Runs | 81 per project condition; 162 across the two conditions |
The project-record cases used three ordinary operations:
- Close one task, cancel another and add two tasks without reusing IDs or rewriting history.
- Record that one project decision supersedes an earlier decision.
- Revise the fabric of one scheduled item without changing another location or breaking the revision record.
The cold-start and example-present conditions were analyzed separately. They were not pooled into one headline result.
For probability comparisons, each arm’s pass rate used an independent Beta(1,1) prior and 100,000 draws. P(A better than B) is the share of draws in which arm A’s pass rate exceeded arm B’s. The prespecified decision rule was better at P ≥ 0.90, worse at P ≤ 0.10 and no detected difference otherwise. Probabilities shown as 1.00 are rounded to two decimal places; they do not mean literal certainty.
Primary result: cold-start project records
The cold-start projects contained the existing project records but no example of the operation being requested. Tasks had not previously been completed or cancelled, no decision used the supersession format and the schedule had no earlier revision to copy.
| Model | Arch Studio 1.5 | Rules paragraph | Vanilla Claude |
|---|---|---|---|
| Haiku 4.5 | 1/9, 11.1% | 1/9, 11.1% | 0/9, 0.0% |
| Sonnet 5 | 5/9, 55.6% | 2/9, 22.2% | 2/9, 22.2% |
| Opus 5.5 | 9/9, 100.0% | 6/9, 66.7% | 2/9, 22.2% |
| All models | 15/27, 55.6% | 9/27, 33.3% | 4/27, 14.8% |
Arch Studio improved the aggregate pass rate over vanilla Claude by 40.7 percentage points, calculated from the unrounded rates.
On Opus, the probability that Arch Studio was better than vanilla Claude was 1.00. The probability that it was better than the rules paragraph was 0.96.
The failures were record-contract failures, not differences in writing style.
For the task register, vanilla Claude used values such as done, completed and cancelled. The register required the status completed and the history events complete and cancel. Arch Studio used the accepted values in all three Opus runs.
In the decision case, vanilla Opus refused to write the record in all three cold-start runs because the project instructions required an Arch Studio operation that was unavailable. Arch Studio created the new decision and preserved the earlier decision in all three runs.
In the schedule case, vanilla Claude sometimes failed to create a revision or wrote the revision in an encoding that Arch Studio’s own reader rejects. Arch Studio passed every Opus run.
Secondary result: project records with prior examples
The warm condition contained prior examples of each requested operation. Claude could inspect a completed task, an existing superseding decision and an earlier schedule revision.
| Model | Arch Studio 1.5 | Rules paragraph | Vanilla Claude |
|---|---|---|---|
| Haiku 4.5 | 0/9, 0.0% | 3/9, 33.3% | 2/9, 22.2% |
| Sonnet 5 | 8/9, 88.9% | 5/9, 55.6% | 4/9, 44.4% |
| Opus 5.5 | 8/9, 88.9% | 9/9, 100.0% | 6/9, 66.7% |
| All models | 16/27, 59.3% | 17/27, 63.0% | 12/27, 44.4% |
Arch Studio’s observed pass rate exceeded vanilla Claude’s by 14.8 percentage points, but the probability of superiority was 0.86, below the evaluation’s 0.90 decision threshold. The evaluation therefore detected no difference in the example-present aggregate. Arch Studio also did not outperform the rules paragraph.
The warm and cold conditions answer different questions. The cold condition measures whether Arch Studio can establish a valid record when the project contains no precedent. The warm condition measures how much of that advantage remains after the project contains examples.
For Opus, removing the examples changed the results as follows:
| Condition | Arch Studio 1.5 | Rules paragraph | Vanilla Claude |
|---|---|---|---|
| Prior examples present | 8/9, 88.9% | 9/9, 100.0% | 6/9, 66.7% |
| No prior examples | 9/9, 100.0% | 6/9, 66.7% | 2/9, 22.2% |
| Change | +11.1 points | −33.3 points | −44.4 points |
Removing prior examples reduced vanilla Claude’s pass rate by 44.4 points and the rules condition by 33.3 points. Arch Studio’s pass rate increased by 11.1 points.
The valid records already present in the warm project gave Claude examples to reproduce. Arch Studio was most differentiated before those examples existed.
Version analysis: 1.5 completed twice as many requests in one turn
We also ran the same record operations against Arch Studio 1.5 and 1.4.5 using project fixtures native to each version. The score used the checks required by both record formats.
| Version | Passed | Pass rate |
|---|---|---|
| Arch Studio 1.5 | 44/54 | 81.5% |
| Arch Studio 1.4.5 | 22/54 | 40.7% |
| Difference | +22 runs | +40.7 points |
The probability that 1.5 was better was 1.00.
Version 1.4.5 stopped and asked for confirmation in 22 of 54 runs:
| Model | 1.4.5 confirmation stops |
|---|---|
| Haiku 4.5 | 0/18, 0.0% |
| Sonnet 5 | 10/18, 55.6% |
| Opus 5.5 | 12/18, 66.7% |
| All models | 22/54, 40.7% |
The 1.4.5 skills instructed the model to preview every write and wait for confirmation. Sonnet and Opus usually followed that instruction. Haiku did not.
We reran the affected conditions under a two-turn protocol. The second message, supplied only when a run stopped for confirmation, was exactly “Yes, go ahead.”
| Follow-up result | Measure |
|---|---|
| Measured 1.4.5 reruns passed | 16/16, 100.0% |
| 1.4.5 runs that stopped, received the reply and passed | 15/15, 100.0% |
| 1.4.5 runs that proceeded without stopping and passed | 1/1, 100.0% |
| Arch Studio 1.5 control reruns passed without stopping | 3/3, 100.0% |
| Planned 1.4.5 runs not measured because of the cost cap | 8 |
All 19 measured reruns passed, but they did not all receive a confirmation reply: 16 used version 1.4.5 and three were 1.5 controls. Of the 1.4.5 runs, 15 stopped and then passed after the scripted reply; one completed without stopping. The eight unmeasured 1.4.5 runs were six Opus cold-start runs and two Sonnet cold-start task-register runs.
When 1.4.5 proceeded, its measured records were as correct as 1.5’s. Version 1.5’s demonstrated improvement was completing already-authorized work in one turn. The follow-up does not establish an all-cells result for 1.4.5 because the cost cap left eight planned runs unmeasured.
That difference applies directly to delegated and scheduled work. A process that stops for another message has not completed, even if it could complete correctly after receiving one.
Exploratory result: isolated product-page tasks
Before the project-record evaluation, we tested seven live product workflows:
- capture one product URL;
- fetch several product sources;
- download and process a selected product image;
- create a product sheet;
- audit a saved product record;
- find comparable products;
- prepare a research brief.
Each case ran three times with Arch Studio and three times with vanilla Claude. The reported mean combines deterministic graders with JEv’s probability of a pass on text rubrics; it is not a simple fraction of successful runs. JEv’s median probability for its selected verdict was 0.65, and 18 of 34 verdicts were below 0.70. The comparison between arms is therefore more reliable than either absolute score.
| Measure | Arch Studio 1.5 | Vanilla Claude |
|---|---|---|
| Mean score | 0.86 | 0.87 |
| Selected image workflow | 1.00 | 1.00 |
| Probability Arch Studio was better | 0.19 | — |
| Cost per run | $0.67 | $0.59 |
No case differed by more than 0.12, which was within the run-to-run variation at three runs per condition.
Both conditions had the same main failure: product pages often stored the selected variant’s SKU, price and image in structured markup that the standard web-fetch tool did not return. Some runs recovered the data by fetching the raw page; the Arch Studio skills did not consistently require that fallback.
The product-page evaluation therefore found no measurable advantage from the plugin. It also identified a concrete implementation requirement: selected-variant extraction needs a deterministic parser or an enforced raw-page fallback.
What the evaluations establish
| Claim | Result |
|---|---|
| Arch Studio improves cold-start project-record work over vanilla Claude | Higher aggregate pass rate: 55.6% against 14.8%; the Opus comparison met the decision threshold at P = 1.00 after rounding |
| Arch Studio improves cold-start project-record work over a short rules paragraph | Higher aggregate pass rate: 55.6% against 33.3%; the Opus comparison met the decision threshold at P = 0.96 |
| Arch Studio remains better after prior examples exist | Higher observed pass rate against vanilla: 59.3% against 44.4%, but below the evaluation’s decision threshold (P = 0.86); not supported against the rules paragraph |
| Arch Studio 1.5 improves single-turn completion over 1.4.5 | Supported: 81.5% against 40.7%; P = 1.00 after rounding |
| 1.5 writes more correct records than 1.4.5 after confirmation | Not established by the measured follow-up |
| Arch Studio improves isolated product-page tasks | Not supported: 0.86 against 0.87 |
| These results apply to the hosted MCP | Not tested |
| These results apply to non-Claude models | Not tested |
Interpretation
Measured result. Arch Studio passed 15/27 cold-start runs against vanilla Claude’s 4/27. With prior examples, the observed rates were 16/27 and 12/27, below the evaluation’s decision threshold for a difference. The isolated product-page scores were 0.86 and 0.87.
Interpretation. The strongest measured use of Arch Studio is establishing valid project records before the project contains an example to copy. Skills that only restated a well-formed isolated task did not improve the tested outcomes. Skills that supplied otherwise absent project-record requirements did.
Product consequence
The next releases should continue to concentrate on:
- explicit record contracts;
- deterministic conformance checks;
- stable IDs, revisions and append-only history;
- completion of already-authorized work;
- parsers and adapters for information that the model’s normal tools cannot see.
Limits
The project-record evaluation used one synthetic office project, three record operations and three runs per model and condition.
It simulated work continuing in a fresh session by staging the project as an earlier session had left it. It did not run a long sequence of dependent sessions or measure whether copied records degrade over time.
The evaluation covered Claude Haiku 4.5, Sonnet 5 and Opus 5.5 through the local plugin. It did not test the hosted Arch Studio connection or another model family.
The 1.5 and 1.4.5 projects were native-format twins, not identical files. The 1.4.5 schedule case had no revisions or pins, and its decision case had no document-register checks, making those operations structurally easier than their 1.5 counterparts. The comparison used only checks shared by both contracts, but it was not a byte-identical fixture experiment.
The confirmation follow-up stopped at the cost cap. Six Opus cold-start runs and two Sonnet cold-start task-register runs at 1.4.5 remain unmeasured. The measured cells therefore do not establish a complete two-turn pass rate for 1.4.5.
The product-page result includes probabilistic JEv scoring with low-confidence verdicts. Its near-zero difference between arms is more defensible than either absolute mean score.
The cold-start result is direct: under the tested record contracts, Arch Studio 1.5 passed 15 of 27 runs and vanilla Claude passed 4.
Reproducibility
The project-record runs used Claude Code 2.1.280. The example-present subject was the private Arch Studio source at 6afebd8; the cold-start fixture revision was e1e3296. The version comparison used public Arch Studio 1.4.5 at tag v1.4.5, commit e7e3644, and the shared-check suite at a3c8908. The live product-page runs used Claude Code 2.1.278 and the 1.5 source based at dc28ce3.
The public plugin source is available, but the evaluation fixtures, raw transcripts, model-cost records and scoring artifacts remain in the private product repository. The article reports the recorded results; it does not claim that the full run corpus is independently reproducible from public artifacts alone.
See the Arch Studio product page for ALPA’s product overview and how to spot a wrapper for the broader evaluation criteria behind this work. Arch Studio is available as a hosted connection and as an open-source plugin. The public repository contains the project-record skills evaluated here.