Appearance
RIGOR Capabilities and Vision
This page is the reader-oriented inventory of RIGOR. It groups the product's major user, Agent, execution, laboratory, evidence, and operational capabilities without replacing the contracts that define exact fields or the progress page that records qualification. Most importantly, it keeps implemented software, open deployment gates, and long-term vision separate.
Status: Maintained overview
Updated: 2026-08-29
Scope: Current software and simulated laboratory, governed human-assisted work, open production gates, and explicitly labeled future vision
The accepted Agent-First System Design remains normative. Current RIGOR Progress is the source for implementation and qualification status. Typed schemas, code, and tests are authoritative for exact interfaces.
Why RIGOR
Scientific discovery requires reproducibility. It can also begin with a new combination of experimental capabilities—or with a result that no one expected. RIGOR is built for this kind of experimentation. LabBridge connects instruments through a common interface and breaks their functions into atomic capabilities: independently callable actions with defined checks and stop behavior. This allows the experimental path to be recomposed as evidence emerges instead of being locked into a fixed workflow. Agent-first changes the division of labor: researchers set the goals and boundaries and decide what is worth pursuing, while the Agent keeps the experiment moving. Human time should be respected rather than spent babysitting a platform. When a person's hands or judgment are genuinely needed, the system hands off an explicit task; otherwise, the Agent continues the work, leaving researchers more time for scientific judgment and creative inquiry. Every adjustment and outcome remains traceable, while Rehearsal establishes a safety boundary before physical execution, so flexibility does not come at the cost of control. RIGOR is not trying to build a human-free laboratory. It aims to let human creativity and intuition set the direction while Agents accelerate validation, so scientific inquiry can deepen as evidence accumulates.
Status Legend
| Label | Meaning on this page |
|---|---|
| Available | Implemented in the current software/simulated scope |
| Qualified baseline | Covered by the dated software qualification; not automatically a production or scientific claim |
| Open gate | Implemented foundation exists, but deployment evidence or operational provisioning is incomplete |
| Future study | A concrete extension worth studying, but not current product behavior |
| Long-term vision | Strategic direction rather than a committed roadmap or evaluated outcome |
| Outside current scope | Deliberately excluded from the current delivery order |
The Current End-to-End Loop
text
Goal + constraints + autonomy policy
-> ASCEND coordination, policy, Attention, and decisions
-> Composer planning and immutable RunSpec/package when needed
-> mandatory isolated Rehearsal of exact experimental intent
-> PACE and LabFlow deterministic execution
-> LabBridge semantic Capability and Operation boundary
-> simulated, algorithmic, or governed human-assisted work
-> committed evidence and archive
-> PRISM isolated analysis
-> ASCEND next-round decisionThe normal user sees Working, Needs you, Paused, Completed, or Problem. Compose, Sync, Prepare, Start, Refresh, polling, analysis dispatch, and continuation are internal activities rather than a sequence of required buttons.
Capabilities Available Today
The inventory below describes major product capabilities, not every endpoint or internal state.
Researcher Experience and Lifecycle
| Capability | What it provides |
|---|---|
| Goal-first workspace | Durable scientific goals, constraints, success criteria, autonomy policy, iterations, evidence, and decisions |
| Agent-first advancement | Routine planning, composition, rehearsal, execution, delivery, analysis, and continuation advance automatically when policy permits |
| Five truthful phases | Backend-owned user phases, blocking reasons, experiment start/end time, and available_actions |
One Needs you channel | Actionable policy decisions and live Human Tasks are unified instead of replicated across modules |
| Governed controls | Pause, Stop, Take over, reply-only stop, and device reset retain policy, audit, cancellation, and safe-state semantics |
| Live conversation feedback | Durable queue/activity phases, elapsed time, bounded Tool summaries, partial response streaming, reconnect, and reduced-motion support |
| Long-task continuation | Context-window pressure and cumulative task spend use separate ledgers; saved checkpoints allow bounded continuation without repeating completed work |
| Observe → Inspect → Act laboratory view | One typed current-lab projection prioritizes reality, source age, current work, exceptions, selection, and governed actions |
| Experiment-owned history | Experiment is the normal data-management root for records, artifacts, retention, and durable owner-confirmed deletion |
| Engineering Console disclosure | Composer, PACE, LabFlow, LabBridge, PRISM, service status, and recovery controls remain available without becoming the ordinary user journey |
| Multilingual operating surfaces | Primary operational surfaces support maintained localization while identifiers, evidence, and audit remain stable |
Agent, Planning, and Self-Awareness
| Capability | What it provides |
|---|---|
| Shared Agent Runtime | Scientist, Planner, and Analyst use versioned profiles on one runtime while retaining module authority, prompts, Skills, tools, context, budgets, and sandboxes |
| Typed Tools | Each Tool declares input/result schemas, side effects, idempotency, policy requirements, and trace correlation |
| Bounded Specialist delegation | Parent/child tasks use a versioned Specialist envelope and cannot commit domain state without deterministic validation |
| Versioned PlanRevision | Plans may be revised by appending a new immutable revision; executed Operations, evidence, decisions, and iterations are not overwritten |
| Direct semantic action | A bounded low-risk atomic intent can avoid artificial Workflow/package ceremony while retaining policy, Rehearsal, Operation, audit, timeout, verification, evidence, and trace |
| Reproducible multi-step planning | Composer selects semantic capabilities and Procedures, validates bounded plans, and constructs immutable RunSpec/packages |
| Curated research presets | Versioned goal/rehearsal presets and factorial study guardrails support controlled study design and truthful demonstrations |
| Source self-awareness | ASCEND may receive a filtered, versioned, read-only view of relevant RIGOR source and project context with a recorded snapshot and policy |
| Replayable Agent decisions | AgentRun records link profile, prompt, Skill, tool schema, context snapshot, policy, tool calls, Operations, evidence, and structured outcomes |
Here, self-aware means aware of bounded, versioned project context and its own runtime/tool constraints. It does not mean consciousness, unrestricted self-modification, or permission to treat mutable source files as evidence.
Rehearsal and Digital-Twin Behavior
| Capability | What it provides |
|---|---|
| Mandatory Rehearsal gate | Every side-effecting experimental Capability input must match a completed selected branch tied to an immutable PlanRevision and ExecutionSnapshot |
| Isolated environments | Rehearsal branches can fork, reset, accelerate, compare, and inject faults without changing formal execution state |
| Exact-intent coverage | ASCEND, PACE, signed LabFlow context, and LabBridge preserve and validate the same capability input lineage |
| Staleness detection | Changed execution state or capability catalog causes rejection or bounded automatic re-snapshot/rehearsal, never a bypass |
| Plan-only promotion | Promotion carries the selected plan and lineage, never predicted world state, findings, artifacts, or measurements |
| Formal surrogate execution | SimFleet devices use durable Operation, lock, timeout, cancellation, safe-stop, verification, and evidence behavior |
| Branch comparison and findings | Feasibility, conflicts, duration, failures, and candidate outcomes can inform plan revision while remaining predicted |
| Frozen scene and replay | Historical scenes retain their map, device state, recorded path, source mode, and authority instead of drifting with the live laboratory |
| Shared 2D/3D projection | Both renderers consume one read-only scene/event contract; animation cannot issue commands or create evidence |
| Rehearsal-only showcase | Curated multi-device scenarios can demonstrate planning and replay while explicitly skipping Composer, PACE, LabFlow, PRISM, formal Operations, and evidence |
Rehearsal is a digital-twin execution rehearsal, not a general physics, chemistry, or molecular simulator. Its strongest current claim is that RIGOR can test bounded operational intent and preserve the semantic gap between predicted behavior and formal evidence.
Deterministic Execution and Safety
| Capability | What it provides |
|---|---|
| Semantic Capability execution | Agents select laboratory meaning such as weighing or image capture; LabBridge/controllers own MQTT topics, serial commands, GPIO, poses, trajectories, and SDK details |
| Governed experimental primitives | Small typed semantic actions can run directly or compose into reproducible Procedures while retaining the same policy, Rehearsal, idempotency, Operation, verification, evidence, and trace contract |
| Immutable Operations | Side effects become correlated, auditable Operation records rather than informal Tool success messages |
| Policy outcomes | ALLOW, ASK, and DENY determine automatic progress, human intervention, or refusal |
| Procedure/workflow scheduling | LabFlow owns validation, deterministic scheduling, waits, locks, retries, cancellation, safe-stop, and run export |
| Loop integrity and delivery | PACE owns immutable RunSpec execution, result completeness, archives, diagnostics, and evidence delivery |
| Verification separation | Dispatch, Operation completion, and observed effect verification are separate; missing or unknown verification is unverified |
| Durable liveness and recovery | Internal stages have progress time, deadlines/leases, recovery budgets, reconciliation, and typed failures rather than indefinite Working |
| Long-running experiment support | EXECUTING has no generic elapsed-time deadline; ASCEND polls authoritative state without repeating a device action after a read failure |
| Governed stop and reset | Stop freezes scientific scope before cancellation; device reset proceeds independently through automatic Operations or Human Tasks with no-resume semantics |
| Fail-safe exceptions | Cancellation, emergency response, stop, and safe-stop are never blocked by Rehearsal |
LabBridge is the semantic boundary above MQTT, ESPHome, vendor SDKs, and other device transports. This is an architectural capability; it does not by itself qualify a real instrument or prove a physical effect.
The corresponding design rule is atomic in composition, governed at execution. The same upstream semantics can serve Rehearsal, simulated execution, algorithmic work, or governed Human Tasks; a later physical binding would still require its own controller and qualification.
Human Participation and Human Awareness
| Capability | Status | What it provides |
|---|---|---|
| Typed Attention | Available | Important decisions include reason, alternatives, evidence, input schema, allowed decisions, consequences, and recovery |
| Typed Human Task | Available | Required manual input, review, or supplied evidence has instructions, claim/submit/reject lifecycle, expiration, actor, artifacts, and provenance |
| Human evidence authority | Available | Accepted schema-valid human submissions remain explicitly human_evidence rather than being relabeled as device evidence |
| Immutable human review | Available | Questions, annotations, review state, and clarification requests append ASCEND-owned records without rewriting PRISM analysis |
| Shared-workspace human gate | Available in software/simulation | Privacy-bounded observations of presence, route obstruction, or equipment interaction produce deterministic PROCEED, WAIT, REROUTE, or COMMUNICATE decisions before Operation creation |
| Physical perception and human-subject qualification | Outside current scope | A future physical path requires a qualified controller, consent, independent safety, and human-subject evaluation |
| Availability-aware scheduling | Future study | Scheduling could use explicitly declared availability windows and handoff rules; it is not implemented as learned monitoring of a person's work habits |
Human-aware coordination intentionally does not collect images, identity, fatigue, workload, or learned schedules. Human Task means a person performs part of an Operation; Human-aware means a shared-workspace Operation reacts to a short-lived bounded observation. They are related but different capabilities.
Evidence, Provenance, and Analysis
| Capability | What it provides |
|---|---|
| Explicit authority classes | Predicted, simulated, physical, human, algorithmic, and unverified outputs remain distinguishable |
| Automatic experimental Provenance | Every committed round projects Goal → Plan → Rehearsal → Execution → Operations → Evidence → Analysis → Conclusion |
| Independent trust dimensions | Completeness, evidence authority, verification/data quality, and gaps are displayed separately rather than collapsed into one score |
| Evidence page | Conclusion, status badges, optional constrained explanation, source chain, warnings, and bounded source details form the normal review surface |
| Isolated PRISM analysis | Versioned Analysis Skills consume committed evidence and produce structured results, reports, archives, and callbacks without controlling experiments |
| Automatically captured analysis trace | Inputs, Tool calls/results, errors, files, artifacts, and validators are recorded without relying on an Analyst to log them manually |
| Typed decision checkpoints | Questions, hypotheses, findings, decisions, limitations, rejected material alternatives, uncertainty, and conclusions form a reviewable record |
| Decision Story, Graph, and Timeline | Scientists can read a concise conclusion-first story, inspect the scientific support graph, or replay grouped runtime events |
| Non-fabricating repair and explanation | One bounded repair may connect existing records; explainers cannot invent nodes, support, authority, verification, or missing reasons |
| Immutable scientific review | Review and clarification append records or start a new Analysis Run; original evidence and analysis are never edited |
| Scoped artifact storage | Immutable object writes, checksums, deterministic keys, scoped identities, retention, and backup inventories preserve large evidence outside domain databases |
RIGOR therefore preserves two complementary lineages: the cross-module lineage from research intent to committed conclusion, and the PRISM decision graph that shows how one analysis reached that conclusion.
Laboratory, Spatial, and Device Model
| Capability | What it provides |
|---|---|
| Capability catalog | Versioned semantic capability, device, environment, reality, and provenance projection |
| Controller availability | Reachability and state freshness are separate; SimFleet heartbeat updates the complete fleet atomically |
| LayoutDraft workspace | One mutable autosaved draft supports optimistic concurrency and continuous qualification without becoming runtime truth |
| Immutable MapRevision | Review & use revalidates and atomically publishes/activates a checksummed laboratory layout |
| Unified equipment editing | The 2D canvas adds, selects, moves, and removes fixed equipment, marked movable equipment, footprints, restricted zones, and walls |
| Role-aware Access Points | A fixed device may expose device-relative delivery, operation, and pickup positions for compatible mobile units |
| Private target materialization | Plans name semantic equipment and work roles; derived targets, station-like IDs, routes, and navigation internals remain private |
| Deterministic path planning | Frozen geometry, footprints, clearance, other mobile poses, directional A*, turn/comfort costs, and collision-checked smoothing produce a recorded path |
| ExecutionSnapshot | Registry, capabilities, DeviceModels, state, spatial facts, custody, and active Operations are frozen behind a checksum |
| Context-pure rendering | Current observation, draft/revision editing, rehearsal, execution, historical replay, and Docs showcase cannot mix state or authority |
| Truthful Showcase Mode | Guided camera, Agent activity, capability pulses, handoffs, and evidence-return effects narrate typed records without controlling lifecycle or creating evidence |
Platform Operations
| Capability | What it provides |
|---|---|
| Four deployment units | Control, Execution, LabBridge, and Analysis Worker co-deploy existing domain services without merging their APIs, stores, or authority |
| Supervisor and Docs infrastructure | Required startup, health, status/control, logs, configured URLs, and public read-only documentation remain outside domain authority |
| Centralized configuration | Ignored YAML plus a separate secret store provides validation, redaction, revisions, impact reporting, generated service projections, apply, and rollback |
| Engineering Console workspace | One responsive ASCEND workspace opens module consoles while preserving service and deep-link context |
| MCP transport gateway | ASCEND exposes allowlisted authenticated Streamable HTTP and Xiaozhi-compatible WebSocket routes while module tool/policy/evidence ownership remains local |
| Durable background jobs | Shared job mechanics support restart-safe worker activity without creating a new domain state machine |
| Schema and RunSpec runtimes | Versioned migrations, immutable RunSpec validation/hashing, and generated OpenAPI clients reduce contract drift |
| Sandboxed Agent execution | Immutable-root Podman baseline, dropped capabilities, bounded writable storage, cleanup reconciliation, and owner-specific policies isolate Agent work |
| Backup and storage qualification | SQLite inventory, native PostgreSQL restore mechanics, object-store checksum behavior, and scoped storage identities have software evidence |
| Change-aware validation | Quick and deep test plans derive from the Git diff and fail closed for unmapped production code |
| Release qualification | Static, contract, module, Web, browser, recovery, Agent, storage, deliberate-failure, backup/restore, and simulated soak evidence is retained for dated baselines |
Module Responsibility Map
| Domain module | Primary responsibility |
|---|---|
| ASCEND | Goals, policy, Coordination, PlanRevision, Rehearsal orchestration, Attention, decisions, researcher projection, and next-round advancement |
| Composer | Bounded planning, capability/Procedure selection, validation, and immutable RunSpec/package construction |
| PACE | Immutable loop execution, result completeness, archive, diagnostics, and evidence delivery |
| LabFlow | Procedure/workflow contract, deterministic scheduling, waits, locks, cancellation, safe-stop, and export |
| LabBridge | Semantic capabilities, devices, Operations, Rehearsal environments/gate, Human Tasks, artifacts, spatial state, controllers, and audit |
| PRISM | Isolated analysis, Analysis Skills, structured results, reports, archives, callbacks, and reviewable decision graph |
Rehearsal is a first-class subsystem coordinated by ASCEND and enforced by LabBridge, not a seventh module. Supervisor and Docs are operational infrastructure, not additional domain modules.
Qualified Baseline and Open Gates
The selected software tree passed the dated 2026-08-06 software qualification, including simulated-laboratory and one-hour soak evidence. That result does not automatically qualify another commit or installation.
| Area | Current boundary |
|---|---|
| Software and simulated laboratory | Implemented and requalified for the selected software baseline |
| Governed human-assisted work | Implemented for typed tasks, decisions, evidence, and software/simulated shared-workspace gating |
| Institution identity | Open production gate: trusted proxy headers and organization account lifecycle remain to be integrated |
| Online system of record | Open production gate: select and qualify the target PostgreSQL deployment and restore path |
| Production object storage | Open production gate: redundant managed storage and independently monitored off-host backup remain to be provisioned |
| Horizontal ASCEND scaling | Add shared event fan-out only if multiple API instances create the need |
| Maintainability | Twenty-three registered legacy hotspots must continue to shrink and cannot grow |
| Real devices and physical effects | Outside current product scope, release gates, backlog order, and current claims |
What Makes RIGOR Different
In one sentence, RIGOR does not put a language model in direct control of lab equipment. It gives Agents a laboratory operating system that can keep work moving, stop safely when something is unclear, and preserve a record of what actually happened.
People set the goal; the system follows through
Most automation software asks people to define a workflow and then keep pushing it through each stage. RIGOR starts from a goal and a set of constraints. It plans, prepares, runs, checks the result, and arranges the next step on its own. It asks for help only when a decision matters, a task genuinely requires a person, or the available information is not enough to proceed.
One laboratory language can serve many devices and experiments
RIGOR breaks laboratory work into actions such as heating, weighing, dispensing, moving a sample, taking an image, waiting, or handing work to a person. These actions can be used on their own or combined into different experiments. An Agent works with meanings such as “heat the sample to this temperature,” not MQTT topics, serial commands, robot joints, or vendor SDKs. A new instrument still needs an adapter, but the experiment plan and its safety rules do not need to be rewritten around that instrument's protocol.
The checked plan is the plan that runs
Rehearsal is more than a simulation that happens to succeed. RIGOR checks whether the plan can complete, which constraints it encounters, and whether the action about to run is still the action that was checked. A change to the parameters, available equipment, or execution environment calls for a new rehearsal rather than borrowing approval from an old one.
An unknown device state is not a reason to try again
If a pump does not reply, it may still have dispensed liquid. If a robot loses its connection, it may not be where the software expects. Repeating the command could double a dose, cause a collision, or ruin a sample. RIGOR first determines what the device actually did, then continues, stops, or asks for help. How the system behaves after a failure matters at least as much as whether the happy path works.
The system works around people, not the other way around
Manual work, human decisions, and evidence supplied by a researcher remain part of the experiment record instead of disappearing into a chat log. Over time, RIGOR can also arrange handoffs and manual tasks around availability that people choose to share. It does not need to identify people or infer their working hours through surveillance.
Every conclusion keeps its evidence
RIGOR connects the original goal to the actions that ran, the data they produced, and the analysis used to reach a conclusion. A plan can change, but the history of an experiment cannot be overwritten. Simulated evidence also remains simulated; putting it in a polished report does not turn it into physical data.
None of these ideas is necessarily unprecedented on its own. RIGOR brings them together in one runtime: Agents can keep different experiments moving, while each equipment action remains rehearsed, checked, and recorded. When the state is uncertain, the system stops to establish what happened instead of silently guessing or running the action twice.
Explaining RIGOR Without the Jargon
The easiest way to introduce RIGOR is to start with the work it takes off a researcher's plate. Experiments involve many steps, different devices, waiting, handoffs, and occasional failures. RIGOR handles as much of that coordination as possible, leaving people to set the goal, change the constraints, and make the decisions that genuinely require judgment.
A homepage version could read:
Researchers set the goal. RIGOR keeps the experiment moving.
It tries the plan first in a digital laboratory. When device state is unclear, it checks what happened instead of blindly repeating an action. When the work is done, each conclusion remains linked to the operations and evidence behind it.
RIGOR is not tied to one instrument or one type of experiment. LabBridge presents device functions as a common set of experimental capabilities, so planning and safety logic can be reused across instruments, automation, and manual steps. The claim is versatility across many experiments, not a universal system that can run every experiment without integration work.
The same practical language works for recovery: if a device stops responding, RIGOR does not assume that nothing happened and simply try again. It first reconciles the device's current state, then continues, stops, or asks for help. This behavior is currently validated in software and simulation; guarantees on physical equipment also require qualified controllers, calibration, and safe-stop mechanisms.
For scientific traceability, RIGOR keeps the goal, plan, actual operations, evidence, and analysis connected. Simulated results remain labeled as simulated even when they appear in a report. That record makes a conclusion easier to inspect, but does not by itself make the science correct.
People remain part of the system. Routine work can advance in the background; RIGOR asks for attention when a material decision, a genuinely manual task, or an unresolved ambiguity needs a person. In the longer term, it should schedule work around explicitly shared availability without monitoring people or inferring their routines.
For a short presentation, this becomes four plain messages:
- Try the plan before acting: rehearse it in a digital laboratory.
- Connect many kinds of equipment: let Agents use capabilities, not drivers and protocols.
- Check before retrying: do not turn an unknown device state into a duplicate action.
- Keep the evidence attached: preserve the path from goal to operation, evidence, and analysis.
These statements describe RIGOR's implemented or actively evaluated software foundation. Physical instrument qualification, long unattended operation, and specific scientific discoveries remain future work and should be presented as such.
From Laboratory Automation to Self-Driving Discovery
RIGOR ultimately aims to make the result of one experiment a natural starting point for the next. Data, theory, models, and experiments should no longer live in separate tools. They become one loop: ask a question, design an experiment, run and analyze it, then use the result to decide what to do next.
Today: make the experimental process dependable
The current work establishes this foundation in software and simulated laboratories. RIGOR can turn a goal into a plan, coordinate the steps, record what actually ran, preserve the evidence, and carry the result into analysis. A researcher does not have to push every internal stage forward, but important decisions and genuinely manual work still go to a person.
Next: bring in physical equipment one instrument at a time
Real instruments will be added gradually, with each one receiving its own driver, calibration, safe-stop behavior, and operating tests. The Agent will continue to ask for laboratory actions such as heating a sample, moving it, or taking an image. A qualified device controller will decide how a pump, robot, or sensor carries out that request. Adding an instrument should not require the whole experiment to be redesigned around its protocol.
Manual steps can also be arranged around availability that researchers choose to share. The purpose is to reduce waiting and handoff delays, not to monitor people or infer their working habits.
Later: learn from each experimental cycle
Once physical equipment and long-running operation have been properly tested, RIGOR can learn which plans tend to work, which recovery strategies are useful, and which analyses are worth reusing. It can then help researchers propose the next experiment, search for better candidates, and identify design patterns. Those scientific results still have to be established through real experiments and review.
Some principles remain the same at every stage: simulated data does not pass as physical data, an Agent does not bypass a device controller to send low-level commands, and a new plan cannot overwrite the history of work already done.
Keeping This Inventory Current
Update this page when a major user outcome or cross-module capability changes. Update the normative PRD when accepted behavior changes, exact contracts when fields or enums change, and Current RIGOR Progress when implementation or qualification changes. Visual concepts and paper figures may illustrate this inventory, but they do not promote a future item to an available capability.
For concrete extension contracts and current limitations, see Process and Analysis Extensions.