Appearance
CoN4Cl Control-Matrix Autonomous Study-Design Benchmark
Status: Software qualified; physical qualification open Updated: 2026-07-21 Owner: ASCEND / Composer / PACE / PRISM
Use this page to understand what the benchmark proves, how its historical XPS evidence is mapped, and where software qualification stops. It is a scientific study-design example, not a general operating guide or evidence that a physical furnace run has been qualified.
Purpose
This benchmark tests a product claim, not merely a data-import path: after one user action, RIGOR must analyze three real historical control arms as separate scientific Iterations and have ASCEND propose a useful fourth experiment from the committed evidence and a generic factorial study-design description.
The benchmark must not put the expected fourth sample name, factor combination, or a benchmark-specific decision branch in the Agent prompt, preset, Skill, or module identifier. Deterministic code may reject an illegal or duplicate arm; it must not select the answer.
Supplied evidence and mapping
The source package remains at /home/flare/data and is not committed. It contains three legacy XPS workbooks, one thesis, and a short experimental description. Only a deterministic, sanitized derivative is allowed under qualification/fixtures/con4cl-xps/v1.
| Source label | Scientific sample | Co factor | NH4Cl factor | Role |
|---|---|---|---|---|
| workbook 3 | N-CSs | absent | absent | common-carbon control |
| workbook 2 | CoN4O | present | absent | cobalt-only control |
| workbook 1 | CoN4Cl | present | present | joint-treatment sample |
Peak-table atomic percentages used as import checks are:
| Sample | C | N | O | Co | Cl |
|---|---|---|---|---|---|
| N-CSs | 91.28 | 4.56 | 4.16 | not detected | not detected |
| CoN4O | 91.86 | 4.13 | 3.71 | 0.30 | not detected |
| CoN4Cl | 91.25 | 5.21 | 3.14 | 0.25 | 0.15 |
The converter records original workbook SHA-256 values, sheet names, source labels, row counts, and energy-column corrections. Workbook 2 includes a +0.6 eV corrected-energy column in its O1s and C1s sheets; that transform must remain visible and must not be silently applied to unrelated scans.
Scientific interpretation
The three measurements are reasonable human-designed controls, but they are not three historical adaptive optimization loops. Together they cover three cells of a two-factor, two-level comparison. A useful next experiment should resolve the remaining confounding between NH4Cl treatment of the carbon support and a Co-Cl coordination interpretation.
The benchmark therefore expects the model to identify a non-duplicate fourth arm with allowed factor levels and to explain the comparison it unlocks. The semantic scorer checks the factor assignment and reasoning/evidence boundaries, not an exact sample label or prose string.
This design cannot by itself prove an axial Co-Cl structure or a catalytic mechanism. XPS composition is supporting evidence; structure or mechanism claims require the relevant additional characterization. ASCEND and PRISM must retain that limitation in their summaries.
Round and lineage contract
The first three evidence rounds each own one Composer package, one PACE Loop Run, one delivery, and one committed PRISM result. The fourth round is a study design record: it does not bind equipment or claim a physical result.
text
one preset Run
-> Iteration 1: historical arm, workbook 3
-> Iteration 2: historical arm, workbook 2
-> Iteration 3: historical arm, workbook 1
-> ASCEND evidence decision
-> Iteration 4: prospective arm selected by the Agent
-> benchmark completes with the proposed design saved
-> optional separate physical Experiment, created only after the researcher
selects a compatible registered real furnaceThe Iteration chain is decision lineage. All four materials are siblings derived from the same PR-NSs source. Candidates must use an explicit common source_sample_id and must not inherit the previous arm as parent_sample_id. The raw source number remains in provenance even though the scientific execution order is 3, 2, 1.
Preset and Agent contract
The versioned preset contains:
- localized goal, description, and success criteria;
- localized, researcher-facing names for each factor and level, kept separate from the stable execution fields such as
cobalt_source=absent; - autonomous policy with exactly four maximum Iterations;
- two categorical factors and their allowed levels;
- the common precursor identity and sibling-lineage rule;
- the ordered three historical arms and their derived fixture locations;
- the generic historical-XPS module and generic prospective-factorial module;
- required unknown-SOP fields and physical-risk boundaries.
After each committed result, the decision context contains the factor schema, historical-arm catalog, prior candidates, and bounded PRISM result views. It does not contain a computed “missing arm” field. Before all historical arms are used, the Agent should select an unused declared historical arm. Afterwards it should propose a non-duplicate allowed factor combination and explain the information gain.
Physical boundary
The prospective synthesis inherits only facts present in the supplied method: 300 mg PR-NSs, a 1:2 PR-NSs:NH4Cl mass relation when NH4Cl is selected, separated downstream/upstream placement, Ar atmosphere, 900 °C, and one-hour dwell. Unknown ramp rate, Ar flow, tube geometry, boat positions, cooling procedure, and local safe operating limits remain unknown.
Composer must not invent these values. The benchmark ends after ASCEND saves the fourth design, so missing furnace facts do not turn the design benchmark itself into Needs you.
Physical work is an explicit follow-on Experiment. Its preparation screen lists only online, registered, non-simulated tube-furnace capabilities whose input contract supports this module's sample, atmosphere, gas-flow, temperature-step, ramp, dwell, and cooling fields. ASCEND repeats the same checks server-side and rejects simulated or incompatible equipment. Only after a researcher selects a compatible real furnace does Composer prepare the physical workflow. Missing approved-method, risk-assessment, geometry, position, and operating values can then become a schema-driven Needs you form for that physical Experiment. The submitted facts are append-only and do not rewrite the Agent's proposed factor combination.
Qualification
Software qualification passes only when all of the following hold:
- the preset parses and one user Run request creates an autonomous four-round Coordination from its accepted Intake definition without requiring technical lifecycle clicks;
- rounds 1-3 use one parameterized module but retain independent immutable packages, loop IDs, analysis IDs, evidence, and source-workbook provenance;
- normalized fixture checksums and the three peak tables match the source;
- PRISM reports composition and correction metadata without choosing round 4;
- the decision Agent cites all three committed analysis IDs when proposing round 4;
- the candidate uses allowed factor levels, the common source, and no previously executed factor combination;
- deleting or changing one historical result changes the Agent context and can prevent a qualified fourth-round decision;
- exact-name prompt leakage and benchmark-specific fourth-arm branches are absent;
- round 4 completes as a saved design without binding or starting equipment;
- a physical follow-on rejects simulation and any real furnace whose declared capability cannot accept the module's heat-program fields;
- simulated qualification, if used, is visibly labeled and does not count as physical qualification.
Production and physical qualification remain separate from this software benchmark.
Live-stack software qualification
The local four-deployment stack passed the software benchmark on 2026-07-20 through the same direct preset start API used by the ASCEND Home action:
- Iterations 1-3 independently composed, executed the read-only XPS fixture import, embedded three checksummed evidence artifacts per delivery, and committed PRISM analyses using the
xps-composition-controlSkill; - the reported peak-table compositions matched the fixture checks above and retained the single-series uncertainty and energy-correction limitations;
- ASCEND cited all three new analysis runs and selected the previously untested
cobalt_source=absent,nh4cl_treatment=presentcombination without a benchmark-specific fourth-arm field; - the decision created Iteration 4 with the common source lineage and all five unknown local-SOP markers. The current product saves that Iteration as a completed design without asking Composer to bind equipment;
- physical preparation is now a separate Experiment. Its equipment picker and server checks exclude SimFleet and incompatible furnace contracts before Composer can create a package.
Qualification also covered retry safety, fixture-controller calls, sample lineage, operation evidence, artifact resolution, independent Agent decision sessions, unknown-SOP fields, and policy projection. Exact test counts belong in dated release evidence because they change as the suites evolve.
This result qualifies the local software workflow only. It does not qualify a real furnace, local SOP, gas handling, tube/boat geometry, operator work, or a physical fourth synthesis.