Skip to content

Articles/Arch Studio 1.5 Improved Cold-Start Project Records

Arch Studio 1.5 Improved Cold-Start Project Records

In 27 cold-start runs per arm, Arch Studio produced valid project records 55.6% of the time against vanilla Claude's 14.8%.

Arch Studio 1.5 produced valid project records more often than vanilla Claude when a project did not contain a prior example of the requested operation.

The cold-start condition covered 81 runs: 27 with Arch Studio, 27 with vanilla Claude and 27 with a short rules paragraph. Arch Studio passed 15/27 (55.6%); vanilla Claude passed 4/27 (14.8%), a difference of 40.7 percentage points calculated from the unrounded rates. On Opus 5.5, Arch Studio passed all nine runs; vanilla Claude passed two.

Arch Studio also improved on version 1.4.5. On checks shared by both versions, 1.5 completed 81.5% of requests correctly in one turn. Version 1.4.5 completed 40.7%.

The advantage was specific. In a separate evaluation of isolated product-page tasks, Arch Studio scored 0.86 against vanilla Claude’s 0.87. The product improved project-record work, not every task given to the model.

Results at a glance

Arch Studio 1.5 against vanilla Claude and the rules paragraph

Project conditionArch Studio 1.5Vanilla ClaudeRules paragraphDifference vs vanilla
No prior examples15/27, 55.6%4/27, 14.8%9/27, 33.3%+40.7 points
Prior examples present16/27, 59.3%12/27, 44.4%17/27, 63.0%+14.8 points

Arch Studio 1.5 against 1.4.5

TestArch Studio 1.5Arch Studio 1.4.5Difference
Project-record operations completed correctly in one turn44/54, 81.5%22/54, 40.7%+40.7 points

Isolated product-page tasks

Arch Studio 1.5Vanilla ClaudeDifference
0.860.87−0.01

Experimental design

The evaluation program contained 277 valid model runs at a recorded model cost of $87.26.

EvaluationRunsCostQuestion
Live product-page tasks42$27.68Does Arch Studio improve isolated product research and extraction?
Project records with prior examples81$22.62Does Arch Studio preserve record contracts when examples already exist?
Cold-start project records81$21.32Does Arch Studio preserve the same contracts without an example to copy?
New 1.4.5 comparison runs; 1.5 reused the 54 project-record runs above54$8.28Did version 1.5 improve single-turn completion?
Two-turn confirmation protocol19$7.36Does 1.4.5 complete the work after one approval reply?
Total277$87.26

The follow-up cost includes $1.12 spent on two discarded smoke tests. Those tests are not included in the 277 valid runs.

Project-record design

FactorLevels
Project conditionNo prior example of the requested operation; prior examples present
Evaluation armArch Studio 1.5; vanilla Claude with the same file tools; vanilla Claude with a 167-word rules paragraph
ModelClaude Haiku 4.5; Sonnet 5; Opus 5.5
OperationTask-register update; superseding project decision; FF&E schedule revision
RepetitionsThree runs per model, operation and arm
Primary outcomePass only if every scored record-contract check passes
Runs81 per project condition; 162 across the two conditions

The project-record cases used three ordinary operations:

  1. Close one task, cancel another and add two tasks without reusing IDs or rewriting history.
  2. Record that one project decision supersedes an earlier decision.
  3. Revise the fabric of one scheduled item without changing another location or breaking the revision record.

The cold-start and example-present conditions were analyzed separately. They were not pooled into one headline result.

For probability comparisons, each arm’s pass rate used an independent Beta(1,1) prior and 100,000 draws. P(A better than B) is the share of draws in which arm A’s pass rate exceeded arm B’s. The prespecified decision rule was better at P ≥ 0.90, worse at P ≤ 0.10 and no detected difference otherwise. Probabilities shown as 1.00 are rounded to two decimal places; they do not mean literal certainty.

Primary result: cold-start project records

The cold-start projects contained the existing project records but no example of the operation being requested. Tasks had not previously been completed or cancelled, no decision used the supersession format and the schedule had no earlier revision to copy.

ModelArch Studio 1.5Rules paragraphVanilla Claude
Haiku 4.51/9, 11.1%1/9, 11.1%0/9, 0.0%
Sonnet 55/9, 55.6%2/9, 22.2%2/9, 22.2%
Opus 5.59/9, 100.0%6/9, 66.7%2/9, 22.2%
All models15/27, 55.6%9/27, 33.3%4/27, 14.8%

Arch Studio improved the aggregate pass rate over vanilla Claude by 40.7 percentage points, calculated from the unrounded rates.

On Opus, the probability that Arch Studio was better than vanilla Claude was 1.00. The probability that it was better than the rules paragraph was 0.96.

The failures were record-contract failures, not differences in writing style.

For the task register, vanilla Claude used values such as done, completed and cancelled. The register required the status completed and the history events complete and cancel. Arch Studio used the accepted values in all three Opus runs.

In the decision case, vanilla Opus refused to write the record in all three cold-start runs because the project instructions required an Arch Studio operation that was unavailable. Arch Studio created the new decision and preserved the earlier decision in all three runs.

In the schedule case, vanilla Claude sometimes failed to create a revision or wrote the revision in an encoding that Arch Studio’s own reader rejects. Arch Studio passed every Opus run.

Secondary result: project records with prior examples

The warm condition contained prior examples of each requested operation. Claude could inspect a completed task, an existing superseding decision and an earlier schedule revision.

ModelArch Studio 1.5Rules paragraphVanilla Claude
Haiku 4.50/9, 0.0%3/9, 33.3%2/9, 22.2%
Sonnet 58/9, 88.9%5/9, 55.6%4/9, 44.4%
Opus 5.58/9, 88.9%9/9, 100.0%6/9, 66.7%
All models16/27, 59.3%17/27, 63.0%12/27, 44.4%

Arch Studio’s observed pass rate exceeded vanilla Claude’s by 14.8 percentage points, but the probability of superiority was 0.86, below the evaluation’s 0.90 decision threshold. The evaluation therefore detected no difference in the example-present aggregate. Arch Studio also did not outperform the rules paragraph.

The warm and cold conditions answer different questions. The cold condition measures whether Arch Studio can establish a valid record when the project contains no precedent. The warm condition measures how much of that advantage remains after the project contains examples.

For Opus, removing the examples changed the results as follows:

ConditionArch Studio 1.5Rules paragraphVanilla Claude
Prior examples present8/9, 88.9%9/9, 100.0%6/9, 66.7%
No prior examples9/9, 100.0%6/9, 66.7%2/9, 22.2%
Change+11.1 points−33.3 points−44.4 points

Removing prior examples reduced vanilla Claude’s pass rate by 44.4 points and the rules condition by 33.3 points. Arch Studio’s pass rate increased by 11.1 points.

The valid records already present in the warm project gave Claude examples to reproduce. Arch Studio was most differentiated before those examples existed.

Version analysis: 1.5 completed twice as many requests in one turn

We also ran the same record operations against Arch Studio 1.5 and 1.4.5 using project fixtures native to each version. The score used the checks required by both record formats.

VersionPassedPass rate
Arch Studio 1.544/5481.5%
Arch Studio 1.4.522/5440.7%
Difference+22 runs+40.7 points

The probability that 1.5 was better was 1.00.

Version 1.4.5 stopped and asked for confirmation in 22 of 54 runs:

Model1.4.5 confirmation stops
Haiku 4.50/18, 0.0%
Sonnet 510/18, 55.6%
Opus 5.512/18, 66.7%
All models22/54, 40.7%

The 1.4.5 skills instructed the model to preview every write and wait for confirmation. Sonnet and Opus usually followed that instruction. Haiku did not.

We reran the affected conditions under a two-turn protocol. The second message, supplied only when a run stopped for confirmation, was exactly “Yes, go ahead.”

Follow-up resultMeasure
Measured 1.4.5 reruns passed16/16, 100.0%
1.4.5 runs that stopped, received the reply and passed15/15, 100.0%
1.4.5 runs that proceeded without stopping and passed1/1, 100.0%
Arch Studio 1.5 control reruns passed without stopping3/3, 100.0%
Planned 1.4.5 runs not measured because of the cost cap8

All 19 measured reruns passed, but they did not all receive a confirmation reply: 16 used version 1.4.5 and three were 1.5 controls. Of the 1.4.5 runs, 15 stopped and then passed after the scripted reply; one completed without stopping. The eight unmeasured 1.4.5 runs were six Opus cold-start runs and two Sonnet cold-start task-register runs.

When 1.4.5 proceeded, its measured records were as correct as 1.5’s. Version 1.5’s demonstrated improvement was completing already-authorized work in one turn. The follow-up does not establish an all-cells result for 1.4.5 because the cost cap left eight planned runs unmeasured.

That difference applies directly to delegated and scheduled work. A process that stops for another message has not completed, even if it could complete correctly after receiving one.

Exploratory result: isolated product-page tasks

Before the project-record evaluation, we tested seven live product workflows:

  • capture one product URL;
  • fetch several product sources;
  • download and process a selected product image;
  • create a product sheet;
  • audit a saved product record;
  • find comparable products;
  • prepare a research brief.

Each case ran three times with Arch Studio and three times with vanilla Claude. The reported mean combines deterministic graders with JEv’s probability of a pass on text rubrics; it is not a simple fraction of successful runs. JEv’s median probability for its selected verdict was 0.65, and 18 of 34 verdicts were below 0.70. The comparison between arms is therefore more reliable than either absolute score.

MeasureArch Studio 1.5Vanilla Claude
Mean score0.860.87
Selected image workflow1.001.00
Probability Arch Studio was better0.19—
Cost per run$0.67$0.59

No case differed by more than 0.12, which was within the run-to-run variation at three runs per condition.

Both conditions had the same main failure: product pages often stored the selected variant’s SKU, price and image in structured markup that the standard web-fetch tool did not return. Some runs recovered the data by fetching the raw page; the Arch Studio skills did not consistently require that fallback.

The product-page evaluation therefore found no measurable advantage from the plugin. It also identified a concrete implementation requirement: selected-variant extraction needs a deterministic parser or an enforced raw-page fallback.

What the evaluations establish

ClaimResult
Arch Studio improves cold-start project-record work over vanilla ClaudeHigher aggregate pass rate: 55.6% against 14.8%; the Opus comparison met the decision threshold at P = 1.00 after rounding
Arch Studio improves cold-start project-record work over a short rules paragraphHigher aggregate pass rate: 55.6% against 33.3%; the Opus comparison met the decision threshold at P = 0.96
Arch Studio remains better after prior examples existHigher observed pass rate against vanilla: 59.3% against 44.4%, but below the evaluation’s decision threshold (P = 0.86); not supported against the rules paragraph
Arch Studio 1.5 improves single-turn completion over 1.4.5Supported: 81.5% against 40.7%; P = 1.00 after rounding
1.5 writes more correct records than 1.4.5 after confirmationNot established by the measured follow-up
Arch Studio improves isolated product-page tasksNot supported: 0.86 against 0.87
These results apply to the hosted MCPNot tested
These results apply to non-Claude modelsNot tested

Interpretation

Measured result. Arch Studio passed 15/27 cold-start runs against vanilla Claude’s 4/27. With prior examples, the observed rates were 16/27 and 12/27, below the evaluation’s decision threshold for a difference. The isolated product-page scores were 0.86 and 0.87.

Interpretation. The strongest measured use of Arch Studio is establishing valid project records before the project contains an example to copy. Skills that only restated a well-formed isolated task did not improve the tested outcomes. Skills that supplied otherwise absent project-record requirements did.

Product consequence

The next releases should continue to concentrate on:

  • explicit record contracts;
  • deterministic conformance checks;
  • stable IDs, revisions and append-only history;
  • completion of already-authorized work;
  • parsers and adapters for information that the model’s normal tools cannot see.

Limits

The project-record evaluation used one synthetic office project, three record operations and three runs per model and condition.

It simulated work continuing in a fresh session by staging the project as an earlier session had left it. It did not run a long sequence of dependent sessions or measure whether copied records degrade over time.

The evaluation covered Claude Haiku 4.5, Sonnet 5 and Opus 5.5 through the local plugin. It did not test the hosted Arch Studio connection or another model family.

The 1.5 and 1.4.5 projects were native-format twins, not identical files. The 1.4.5 schedule case had no revisions or pins, and its decision case had no document-register checks, making those operations structurally easier than their 1.5 counterparts. The comparison used only checks shared by both contracts, but it was not a byte-identical fixture experiment.

The confirmation follow-up stopped at the cost cap. Six Opus cold-start runs and two Sonnet cold-start task-register runs at 1.4.5 remain unmeasured. The measured cells therefore do not establish a complete two-turn pass rate for 1.4.5.

The product-page result includes probabilistic JEv scoring with low-confidence verdicts. Its near-zero difference between arms is more defensible than either absolute mean score.

The cold-start result is direct: under the tested record contracts, Arch Studio 1.5 passed 15 of 27 runs and vanilla Claude passed 4.

Reproducibility

The project-record runs used Claude Code 2.1.280. The example-present subject was the private Arch Studio source at 6afebd8; the cold-start fixture revision was e1e3296. The version comparison used public Arch Studio 1.4.5 at tag v1.4.5, commit e7e3644, and the shared-check suite at a3c8908. The live product-page runs used Claude Code 2.1.278 and the 1.5 source based at dc28ce3.

The public plugin source is available, but the evaluation fixtures, raw transcripts, model-cost records and scoring artifacts remain in the private product repository. The article reports the recorded results; it does not claim that the full run corpus is independently reproducible from public artifacts alone.

See the Arch Studio product page for ALPA’s product overview and how to spot a wrapper for the broader evaluation criteria behind this work. Arch Studio is available as a hosted connection and as an open-source plugin. The public repository contains the project-record skills evaluated here.

Next article

Arch Studio 1.5: The Four Layers a Firm Builds On →