CryoDAQ is acquisition, control, and operator-support software for a cryogenics laboratory. It reads sensors, drives configured instruments, records measurements, and reports conditions requiring attention. Hazardous-output protection is fail-closed on the named, tested paths; real-instrument, independent final-element, and laboratory acceptance remain separate open gates.
A cryostat cools test hardware to a few degrees above absolute zero and holds it there for days. While it does, a dozen thermometers, pressure gauges and a laser interferometer report what is happening inside — and a programmable current source can put real power into hardware sitting in a vacuum at 4 K. Getting that wrong damages equipment that took months to build. CryoDAQ is the software in the middle.
When the software does not know something, it says so.
That sounds obvious. It is not what most instrument software does. The usual failure is quiet: a
sensor link drops, and the display keeps showing the last number it saw, or falls back to 0.0, or
reports a power source as OFF because nothing told it otherwise. An operator reads OFF and walks
up to hardware that may still be live.
CryoDAQ's named unavailable-state paths are built the other way round. In the
screenshots below, a channel with no data reads —, and a source with a dead
link reads UNKNOWN, not OFF; controls on that panel disable because there
is no transport to command over. Those screenshots were taken with no
instruments connected. This is demonstrated UI behavior, not a claim that every
current or future producer/view has complete stale-state revocation coverage;
that product-wide gate remains open in docs/OPEN_CELLS.md.
The candidate adds canonical reading descriptors and exact binding checks on
named acquisition and safety paths. It does not yet make cross-device identity
confusion impossible across every GUI, archive, interlock, Telegram, and report
path; those remaining spelling- and descriptor-authority gaps are listed in
docs/OPEN_CELLS.md. Safety configuration errors fail startup on the paths
documented below; optional physical-alarm configuration still has explicitly
documented fallback behavior.
Configured sensor channels across LakeShore, Keithley, Thyracont and Etalon hardware · scripted experiment campaigns · automated sensor calibration with multi-format export · auto-generated reports · Telegram alerting with time-based escalation · anomaly detection and an alarm pipeline · a question-answering layer over the lab's own documentation, running locally · replay of historical data through the live interface · interferometric length metrology · and a cross-platform regression suite of several thousand tests.
It replaced a three-year-old LabVIEW program that drove the instruments and sent email.
CryoDAQ was written for the Astro Space Center of the Lebedev Physical Institute, for cryogenic materials testing on the Millimetron space observatory — measuring how mirror structures deform as they cool toward the temperatures they will meet in orbit.
Every screenshot below is the real interface, rendered from this repository's own code. All of them show the system with no engine connected — which is deliberate, because that is where the design decision that matters is visible.
The shift summary answers the questions an operator actually asks — can we continue?, what needs attention? — and when it cannot answer, it says so. Every card carries an explicit НЕТ СВЯЗИ badge and names the next safe step. It never renders a plausible-looking zero.
The main window: sixteen sensor channels, temperature and pressure history, and an operator log.
Unavailable channels read — / «Нет данных», never 0.0.
Hazardous source control. With the link down each channel reads НЕИЗВЕСТНО — unknown — not
OFF, and the controls disable because there is no transport to command over. An operator who reads
"OFF" walks up to hardware that may be live; this panel refuses to tell that lie.
Sensor calibration across three LakeShore instruments, grouped by the hardware they belong to.
- Latest release: v0.64.1 (2026-07-08)
- Released baseline: the latest locally available release tag is
v0.64.1(2026-07-08). - Qualification campaign: the large software-side laboratory-readiness
refactor was merged into
masteron 2026-07-31 after independent review. This README does not assert a current branch or candidate SHA;PROJECT_STATUS.mddefines the live boundary. The merge is a software checkpoint only — it is not accepted merely because mock tests or ordinary CI pass. - Evidence boundary: physical instrument, dummy-load, independent final-element,
and laboratory acceptance gates remain open until the procedures in
docs/lab_verification_checklist.mdare executed and recorded. - Current truth: see
PROJECT_STATUS.md. A working before/after narrative, metrics, design decisions, and architecture maps are indocs/MONTANA_REFACTOR_REPORT.md; the status document defines the current acceptance boundary.
The qualification refactor (merged 2026-07-31; campaign records in
docs/campaigns/) moved CryoDAQ toward narrower ownership, explicit
evidence, and visible failure boundaries while preserving the information-dense
operator workflow. Several boundaries remain open below, so this is a design
direction with open acceptance gates, not a finished system property.
The most important changes are:
- Narrower ownership, with an open exception. Acquisition, safety state,
recording, periodic delivery, archive rotation, and operator snapshots have
explicit owners. The current assistant still instantiates a second
SQLiteWriterfor operator-log output, so the one-writer boundary is not yet complete. - Persistence before ordinary data publication. In the production
acquisition path, the archive-selected batch commits before
DataBrokerpublication.SafetyBrokerseparately receives the full raw batch for safety evaluation after the archive-selection/persistence branch; adaptive-throttle omissions and synthetic broker events are not claimed to be durably stored. - Fail-closed hazardous output. Source authority is currently restricted by a reviewed one-model roster and driver-type equality; capability-derived authority remains open under OC-007. Verified OFF, emergency shutdown, interlocks, and bounded process cleanup were strengthened; passive extension mechanisms cannot acquire actuator authority.
- Descriptor-qualified channel identity, incomplete end to end. Canonical descriptors are carried through named acquisition, SQLite, archive/replay, report, and GUI paths. Archive-finalisation rows and multiple GUI, interlock, assistant, and notification routes still fall back to spelling or lack descriptor authority (OC-008, OC-011, OC-023–OC-025). Startup validation fails closed for the exact safety bindings it resolves; this is not a claim that every safety/alarm/interlock pattern is descriptor-bound.
- Process isolation and recovery. The launcher supervises the engine, GUI bridge, assistant, and bounded report children with explicit lifecycle machinery. The assistant is process-isolated but does not yet satisfy the repository's strict observational boundary because it still holds delivery credentials and mutation/write paths.
- Periodic reporting without control authority. Rendering and delivery are observational. Durable state and receipts distinguish rendering, delivery, acknowledgement, ambiguity, retry, and terminal success across crashes.
- Operator-centred GUI governance. The panoramic dashboard remains primary and summaries are additive. The design contract requires current values, stale/disconnected state, provenance, active hazards, and acknowledgement state to remain reachable; shared freshness/provenance/lifecycle wiring and operator-scenario acceptance are still open.
- Reproducible evidence. CI is split across Windows and Ubuntu; installer and runtime SQLite policy share a tested contract; Windows ONEDIR and Linux soak procedures bind evidence to exact commits. Mock evidence never claims to be physical-hardware evidence.
The refactor is large, but its governing idea is simple: make authority narrow, state explicit, failure visible, and every acceptance claim traceable to the environment that actually produced it.
The refactor merged into master on 2026-07-31 after independent review. The
merge is a software checkpoint, not acceptance: it establishes no release
readiness, physical-OFF proof, real-instrument, packaged-Windows, or
laboratory-acceptance claim. PROJECT_STATUS.md defines the live boundary;
docs/OPEN_CELLS.md is the live software-defect disclosure; deferred
directions with trigger conditions are registered in docs/OBLIGATIONS.md.
Historical checkpoint note: commit
503c8bf8d884654256ede4f08a9e44ab7b382242 is associated with reported
eight-job GitHub Actions run 29662599972; that evidence covers only that
commit.
A pre-merge review record stood in this section — a dirty-tree status, a
REJECTED verdict on a then-local SafetyManager shutdown/HOLD slice, and a
pre-merge work list. It described the pre-merge candidate, not the merged
tree, and is superseded here; the text remains in this file's own pre-merge
git history, and the current status of every such item is read from
docs/OPEN_CELLS.md, never from this README.
Physical instrument, dummy-load, host-death, independent final-element, long-duration soak, and laboratory operator acceptance remain separate and open until their prescribed evidence is recorded. The future 100+ sensor / 4K projector and semantic-zoom view is deferred and does not block ordinary lab readiness.
Use this section when an agent must interview a maintainer, reviewer, operator, or cryogenic engineer about CryoDAQ. Start with plain operational questions; only then descend into modules and protocols. Do not use internal campaign labels as substitutes for explaining behavior.
Interview the CryoDAQ maintainer and laboratory stakeholders to produce a factual, operator-centred account of the system and the laboratory-readiness refactor. Establish the real experiment workflow first, then trace authority, data, failure recovery, GUI truth, and acceptance evidence. Separate released behavior, current checkpoint behavior, planned work, and physical claims that remain unverified. Challenge vague answers with concrete scenarios and ask for the owning process, persisted record, operator-visible state, relevant source or test, and exact evidence for every important claim. Do not treat CI, simulation, mocks, screenshots, or documentation as proof of hardware behavior. Do not recommend weakening fail-closed behavior or hiding operator information to simplify the design. End with: (1) a plain-language system summary; (2) a before/after comparison; (3) an authority and data-flow map; (4) unresolved safety and operability questions; (5) open software, Windows, WSL, packaging, dummy-load, and physical-lab gates; and (6) contradictions between interviews, code, tests, and documentation.
Interview people separately where practical. Operators describe actual work and failure visibility; cryogenic engineers describe the apparatus and hazards; maintainers describe implementation and ownership; reviewers challenge the evidence. Record who supplied each operational claim, but do not put private personal data or raw private transcripts in the repository.
Read these sources first, in order:
- This README for the product and refactor overview.
PROJECT_STATUS.mdfor the exact current evidence and open gates.docs/MONTANA_REFACTOR_REPORT.mdfor the full before/after narrative, metrics, decisions, and architecture diagrams.docs/architecture.mdfor runtime ownership and data flow.docs/lab_verification_checklist.mdfor what software tests cannot prove.docs/design-system/README.mdbefore asking about GUI changes.
Recommended interview questions:
- What experiment does CryoDAQ support from preparation through cooldown, measurement, warmup, reporting, and archive?
- Which information must an operator see continuously, and which information is acceptable in an overlay or drill-down?
- What did the old LabVIEW workflow do well, and which operational habits must CryoDAQ preserve?
- Which failures have actually occurred in the lab, and which are currently only anticipated by tests or hazard analysis?
- Which process owns instruments and safety state? What happens if the GUI, assistant, report renderer, or launcher dies?
- What evidence is required before the software may say a hazardous source is OFF? What happens when OFF readback is unavailable or contradictory?
- How do alarm acknowledgement, alarm clearing, safety recovery, and interlock reset differ? Which of them can change physical authority?
- Why can a passive driver plugin never become a source driver through duck typing or configuration alone?
- Which physical claims remain impossible to close with mocks, CI, replay, or a screenshot?
- Where is the persistence-before-publication order enforced?
- What is the canonical identity of a channel, and how does it survive hot SQLite data, cold Parquet rotation, replay, reports, and GUI display?
- How are partial writes, storage backpressure, corrupt metadata, archive races, and cancellation represented without creating a second writer?
- What does the GUI show when a reading is stale, disconnected, unavailable, or from a different timestamp than a comparison channel?
- Which long-running and ephemeral processes exist, who starts them, and who is responsible for reaping their descendants?
- How does CryoDAQ distinguish a rendered periodic report from a durably delivered and acknowledged one?
- What prevents a restarted assistant from duplicating delivery or accepting an old owner token, slot, process identity, or acknowledgement?
- How are blocking LibreOffice/report operations kept off the engine loop, and what useful artifact remains when PDF conversion times out?
- Why is the panoramic dashboard still primary, and why must a summary display remain additive?
- Which colors are reserved for safety meaning, and how are state differences communicated without color alone?
- What becomes better and what becomes worse for the operator with each proposed GUI change?
- Can any responsive breakpoint, clipping rule, filter, acknowledgement, or auto-ranging choice hide current value, status, provenance, or an active hazard?
- How should the future 100+ sensor / 4K projector view aggregate information without turning the application into a black box?
- Which exact commit, operating system, Python, SQLite, dependency lock, and artifact hashes produced the claimed result?
- Did the test exercise real loopback/process/filesystem behavior, or was the property mocked away?
- What do the Windows ONEDIR, WSL short soak, hosted CI, long soak, dummy-load, and physical-lab gates each prove—and explicitly not prove?
- What evidence would make you reject the candidate even if every unit test were green?
Ask for concrete examples, file paths, state transitions, and evidence records. An answer such as “the tests pass” is incomplete unless it identifies the exact candidate, environment, gate, pass/skip counts, and the claims that remain open.
CryoDAQ exposes four primary operator deployment surfaces/modes:
cryodaq— the full cross-platform operator launcher. Its process hosts the Qt GUI and supervises the engine, GUI bridge, and optional assistant candidate. Bounded report children are owned by the engine or assistant component that requested them, not directly by the launcher.cryodaq-engine— the standalone headless asyncio runtime. It drives the instruments, runs the safety-manager FSM, evaluates alarm rules and interlocks, persists data, and serves GUI commands over ZMQ.cryodaq-gui— a reduced standalone Qt client for an already-running engine. It can restart without stopping data acquisition. A standalone engine still supports on-demand reports; this reduced path does not provide the launcher's assistant/periodic-delivery lifecycle.cryodaq.web.server:app— the optional FastAPI monitoring dashboard.
Only the engine owns instruments and safety authority. Report workers have no control authority. The assistant is intended to be observational, but its current second operator-log writer, RAG mutation path, and Telegram credential must be removed or reassigned before that boundary is complete. GUI paths remain clients of backend truth rather than control owners.
Data flow:
Instrument → allowlisted Driver/Capability → Scheduler
→ SQLiteWriter (archive-selected batch)
→ DataBroker → {GUI, alarms, analytics}
↘ SafetyBroker (full raw batch; safety path)
IPC: ZeroMQ PUB/SUB :5555 (msgpack) + REP/REQ :5556 (JSON commands).
- 3× LakeShore 218S (GPIB) — 24 temperature channels
- Keithley 2604B (USB-TMC) — dual-channel SMU (
smua+smub) - Thyracont VSP63D (RS-232) — 1 pressure channel
- Etalon MultiLine (TCP/IP) — interferometric length metrology; averaged and continuous modes, vibration burst capture to Parquet
The list below describes the active tree: released v0.64.1 workflows together with the merged checkpoint behavior that still awaits physical acceptance. Checkpoint defaults and hardening are not release or physical-acceptance claims; see Status above.
- Knowledge base (RAG): local semantic search over the experiment archive,
vault notes, the operator log, and the
data/knowledge/corpus (equipment_manuals— instrument PDFs via pypdf;procedures— Markdown;reference— operator manual / README / CHANGELOG). The RAG module: loader -> LanceDB indexer -> top-K searcher; embeddingsqwen3-embedding:0.6b(1024-dim) via Ollama. Indexing is offline only: runcryodaq-rag-indexto build or rebuild the index, then usecryodaq-rag-searchto query it. The running assistant exposes no rebuild command, update button, or startup bootstrap, but a separate current RAG mutation path keeps the assistant's overall observational boundary open. - Local operator-query service: a local Ollama service (no external model APIs) classifies operator intent (IntentClassifier), routes the query (QueryRouter), and answers from live data (BrokerSnapshot) and the knowledge base (KNOWLEDGE_QUERY). The query adapters are intended to be read-only; the service-level write/credential exceptions above remain open. Model-call audit records are produced on the configured paths; this is not a completeness guarantee for every failure mode.
- Historical-data replay: replays records through the DataBroker; a predictor
runs on top of the replay stream with a decoupled clock for accelerated
playback;
cryodaq-replay-curvefor curve transforms; a legacy channel map for pre-2025 records. - Experiment FSM: six canonical phases: preparation → vacuum → cooldown → measurement → warmup → teardown. Abort/fault is an outcome or state, not a seventh experiment phase. Templated scripted runs.
- Calibration v2: continuous SRDG acquisition during calibration experiments;
post-processing (extract → downsample → Chebyshev fit per zone); export to
.cof(raw coefficients) /.340/ JSON / CSV; import from.340/ JSON; runtime application with a global / per-channel policy. - Auto-generated reports: templated sections; a guaranteed
report_editable.docx; best-effort PDF viasoffice/ LibreOffice. - Telegram alerts: one configured default chat target. Time-based escalation sends the same message to each configured escalation chat after its respective delay; there is no role-based filtering.
- Sensor diagnostics → alarm pipeline: MAD-outlier + cross-channel correlation-drift detection. A persistent anomaly publishes an alarm: warning after 5 min, critical after 15 min, auto-clear on recovery. Concurrent events are aggregated into a single Telegram message; configurable cooldown.
- Alarm engine v2: threshold / rate / composite / phase-dependent rules; a hysteresis deadband; in-place severity escalation (WARNING→CRITICAL); an ack/clear publish path.
- Interlocks: 3 hard-protection rules (cryostat / compressor / detector).
overheat_cryostatandoverheat_compressoruseemergency_off, which latchesFAULT_LATCHED.detector_warmupusesstop_source: a soft stop that turns outputs off and transitions toSAFE_OFFwithout a fault latch. - Fail-closed safety discipline: Keithley output OFF is readback-verified;
unverified OFF becomes a fault or a blocking RUN precondition instead of a
false SAFE_OFF.
config/physical_alarms.yamlexplicitly setsvacuum.escalate_to_safety: true; the built-in missing-file default is alarm-only (false), while invalid existing configuration fails safer totrue. - Operator log: SQLite-backed; accessed via the GUI + ZMQ.
- Experiment templates, lifecycle metadata, artifact archiving: a
data/experiments/<id>/directory withmetadata.json,reports/, and an optional Parquet archive. - Plugin architecture: ABC isolation; if a callback raises, the loader logs the exception, skips that batch for the plugin, and retries it on the next batch. No degraded state is recorded.
- Housekeeping: adaptive throttle + retention + compression.
- Cold-storage rotation (F17): enabled by default
(
cold_rotation.enabled: true).ColdRotationServiceis wired into the engine and runs daily atschedule_time(03:00): daily SQLite files older than 30 days rotate into Parquet/Zstd. Every reader goes throughArchiveReader(hot SQLite ∪ cold Parquet) — GUI history, live operator journal, reports, CSV/XLSX/HDF5/Parquet exports, replay, and calibration all see rotated days. The only kill-switch iscold_rotation.enabled; rotation is idempotent and the stranded-DB sweep deletes only a byte-identical original (source_md5). - SQLite fail-closed runtime: all runtime DB connections go through
storage/_sqlite.py. The supported Windows/Linux environment pins a safe SQLite version; an unsafe selected implementation blocks startup. - Leak-rate estimation (F13):
LeakRateEstimator— a rolling window, OLS regression without numpy, history indata/leak_rate_history.json. Commands:leak_rate_start/leak_rate_stop(ZMQ). Requireschamber.volume_lininstruments.local.yaml.
Primary shell: MainWindowV2 — Phase III completed in v0.40.0.
Layout — an ambient information radiator for week-long experiments:
- TopWatchBar — engine indicator, experiment status, time-window echo
- ToolRail — navigation across overlay panels
- DashboardView — 5 live zones:
- Sensor grid (temperature + pressure overview)
- Temperature plot (multi-channel, clickable legend, window selection)
- Pressure plot (compact log-Y)
- Phase widget (experiment-phase indicator + transition)
- Quick log (inline operator-log view)
- BottomStatusBar — safety-state indicator
- OverlayContainer — host for analytics and archive
Overlay panels (from the ToolRail):
- Analytics — phase-aware widgets: temperature trajectory, cooldown history, experiment summary (channel statistics, top alarm, artifact links), cooldown forecast (cooldown predictor, a progress-variable ensemble with ETA), and steady-state temperature forecast (T∞ via an exponential fit).
- Archive — past experiments + reports + Parquet exports
- Calibration — the capture / fit / export workflow
- Knowledge base — RAG search + embedded operator chat
- MultiLine — interferometric metrology + "Vibration capture" (burst)
- Operator log
- Other overlays via ToolRail icons
MainWindowV2 is the sole operator shell. The legacy tab-based MainWindow and
all tab-era overlays were removed in Phase II.13; cryodaq-gui has used
MainWindowV2 since Phase I.1.
The launcher tray is deliberately coarse. With current wiring, a known safety fault can produce red; alarm count remains unknown, so alarms alone cannot drive red and green is unreachable. All other connected/disconnected/unknown, stale-data, or reporting-fault cases resolve to the amber caution shape. Shape and Russian tooltip duplicate color. The tray is not an authoritative alarm or readiness summary and must not replace the dashboard or alarm surface.
- Windows 10/11 or Linux
- Python
>=3.12(must link SQLite>=3.51.3, or a backport-safe 3.44.6 / 3.50.7 — see "Known limitations") - Git
- A VISA backend / instrument drivers as needed
conda env create --file environment.yml
conda activate cryodaq
pip install -r requirements-lock.txt
pip install -e . --no-deps --no-build-isolation
pip checkThe lock includes the PEP 517 build-backend closure. The supported path disables a second runtime or build resolver when installing the project itself.
The tracked environment pins the supported Python/SQLite runtime. The pip lock pins resolved Python package versions but is not a hashed, bit-for-bit artifact lock. A bare editable extras install is a developer convenience only and is supported only inside an independently verified safe SQLite runtime:
pip install -e ".[dev,web]"Supported workflow: install from the repository root into the active cryodaq
environment.
Running pytest without pip install -e ... is not supported.
Key runtime dependencies: PySide6, pyqtgraph, pyvisa, pyserial-asyncio,
pyzmq, python-docx, scipy, matplotlib, openpyxl, pyarrow.
cryodaq-engine # headless engine (real instruments)
cryodaq-gui # GUI only (connects to a running engine)
cryodaq # full cross-platform operator launcher
cryodaq-engine --mock # mock mode (simulated instruments)
uvicorn cryodaq.web.server:app --host 127.0.0.1 --port 8080 # optional web (loopback)The web dashboard's GET surface has no authentication — bind it to 127.0.0.1
only; public access requires a reverse proxy with authorization (or an SSH
tunnel). The two /api/v1 write endpoints (POST /log, POST /alarms/{id}/ack)
require a bearer token from the gitignored config/web.local.yaml.
POST /api/v1/log may include the exact experiment_id; if it is omitted the
entry is explicitly experiment_unbound and is never attached implicitly to
the current experiment. The caller supplies Idempotency-Key as an exactly
32-character lowercase hexadecimal value and reuses that same key for retries; the
server copies it to request_id. Clients do not put request_id in JSON. The
server owns author/source. Public live
readings encode NaN and infinities as JSON null while retaining their
identity and status, so unavailable data cannot masquerade as a valid number.
For POST /api/v1/log, HTTP 409 is definite non-commit (obey retry_safe),
502 is unknown (do not blindly retry or replace the key), and 503 is committed
with publication pending; follow the normative settlement and retry
contract.
Helper CLIs:
cryodaq-cooldown build --help # cooldown ML: training options
cryodaq-cooldown predict --help # cooldown ML: ETA-prediction options
cryodaq-trends scan --help # cross-experiment feature-table options
cryodaq-trends drift --help # cross-experiment drift-check options
cryodaq-replay-curve # curve transforms for replay
cryodaq-rag-index # build the knowledge-base index
cryodaq-rag-search # semantic search over the knowledge baseActive configuration files in the current checkpoint:
config/instruments.yaml— GPIB/serial/USB addresses, LakeShore channels,chamber.volume_lfor the F13 leak rateconfig/instruments.local.yaml.example— template for machine-specific instrument overrides (instruments.local.yamlis gitignored)config/channel_descriptors.yaml— tracked declared descriptor/binding roster for acquired channels; physical reconciliation and the remaining consumer migrations are open gatesconfig/channel_descriptors.local.yaml.example— machine-specific whole-file replacement, never a partial merge; reconcile it with the physical roster before real-hardware useconfig/safety.yaml— FSM timeouts, rate limits, drain timeoutconfig/alarms_v3.yaml— alarm engine rules (threshold/rate/composite/phase)config/interlocks.yaml— interlock conditions + actionsconfig/physical_alarms.yaml— tunables for the cold-cryostat physical guards; its trackedvacuum.escalate_to_safetyvalue istrue, while the built-in missing-file default is alarm-only (false)config/channels.yaml— display names, visibility, groupingconfig/notifications.yaml— tracked placeholder/schema; real credentials belong only in gitignoredconfig/notifications.local.yaml. Engine and periodic loaders prefer the local file, but the current assistant Telegram sender still reads only the tracked base file; that candidate wiring is open.config/notifications.local.yaml.example— template for local Telegram credentials (notifications.local.yamlis gitignored)config/housekeeping.yaml— throttle, retention, compression,cold_rotationconfig/plugins.yaml— sensor_diagnostics + vacuum_trend;aggregation_threshold+escalation_cooldown_sconfig/cooldown.yaml— cooldown-predictor parametersconfig/analytics_layout.yaml— phase-aware analytics widget layoutconfig/agent.yaml— local operator-query service (Ollama model, triggers, rate limit)config/rag.yaml.example— knowledge base / RAG (embedding model, corpus)config/rag_categories.yaml— KnowledgeBasePanel sidebar query presetsconfig/sinks.yaml.example— sinks (vault notes, webhook) on finalizeconfig/web.local.yaml.example— template for the FastAPI write-token (web.local.yamlis gitignored)config/themes/*.yaml— bundled GUI theme packs; selected via gitignoredconfig/settings.local.yamlconfig/experiment_templates/*.yaml— experiment-type templates
Most *.local.yaml files override base settings. Channel descriptors are the
deliberate exception and are selected as a pair with the instrument authority:
when instruments.local.yaml is selected, channel_descriptors.local.yaml is
required and completely replaces the base manifest; with base
instruments.yaml, the base descriptor manifest is used even if a local
descriptor file happens to exist. Manifest/schema errors, ambiguous bindings, a
missing required local manifest, and instrument-set mismatch block startup.
Each accepted reading must then resolve exactly one
(instrument_id, emitted_channel) binding to a stable channel_id; an
undeclared emitted channel is rejected when it is first bound. A descriptor
grants identity only, never hazardous-source capability.
data/experiments/<experiment_id>/
metadata.json
reports/
report_editable.docx
report_raw.pdf # optional, best-effort (soffice/LibreOffice)
report_raw.docx
assets/
data/calibration/sessions/<session_id>/
data/calibration/curves/<sensor_id>/<curve_id>/
data/archive/year=YYYY/month=MM/ # Parquet cold storage (F17)
data/leak_rate_history.json # leak-measurement history (F13)
Templated sections: title_page, cooldown_section, thermal_section,
pressure_section, operator_log_section, alarms_section, config_section.
Guaranteed artifact: report_editable.docx. Optional: report_raw.pdf
(best-effort, requires soffice / LibreOffice).
tsp/cryodaq_wdog.lua — a TSP software late-pet checker below the host
SafetyManager. P=const still runs host-side in keithley_2604b.py.
The watchdog is operator-selectable via config/instruments.yaml →
keithley.watchdog.mode: off (driver default — script not loaded, host is the
sole authority), best_effort (activate on connect, fall back to host-only on
failure), required (fail-closed — requires the explicit autonomous bit and
makes connect() raise while it is absent, so SAFE_OFF holds). Version 3
explicitly reports cryodaq_wdog_autonomous=0:
it covers only stall-then-recover when a later pet arrives and has zero
full-host-death coverage. The previous timer implementation was removed because
it used commands and action values that the 2600B reference manual does not
document as valid. A true host-death OFF path requires a documented redesign
and physical proof; an independent latching cutout/interlock is preferred.
watchdog.timeout_s must be a finite number from 1 to 300 seconds; the TSP
clock has one-second granularity and uses a strict elapsed > timeout test.
src/cryodaq/
agents/ # local query service + RAG knowledge base
analytics/ # calibration fitter, cooldown predictor, plugins, vacuum trend,
# leak_rate estimator (F13)
core/ # safety FSM, scheduler, broker, alarms v2, interlocks,
# sensor_diagnostics, experiments, zmq_bridge
drivers/ # LakeShore, Keithley, Thyracont, Etalon MultiLine + transports
gui/ # MainWindowV2, dashboard, overlays
notifications/ # Telegram alerts + interactive bot + escalation
replay/ # historical-data replay + curve transforms
replay_engine/ # ZMQ-compatible replay engine (accelerated playback)
reporting/ # template-driven DOCX generator
sinks/ # vault notes + webhook on experiment finalize
storage/ # SQLite, Parquet, CSV, HDF5, XLSX,
# cold_rotation (F17), archive_reader (F17)
tools/ # CLI utilities (cooldown_cli)
utils/ # shared helpers
web/ # FastAPI monitoring
tsp/ # Keithley TSP watchdog (cryodaq_wdog.lua; loaded per watchdog.mode)
tests/ # cross-platform unit, integration, GUI, process, and evidence tests
config/ # YAML configuration
python -m pytest tests/core -q
python -m pytest tests/storage -q
python -m pytest tests/drivers -q
python -m pytest tests/analytics -q
python -m pytest tests/gui -q
python -m pytest tests/reporting -qRun after the supported installation above. GUI tests require PySide6 +
pyqtgraph. Do not use CRYODAQ_ALLOW_BROKEN_SQLITE=1 to claim storage-test or
deployment evidence; an unsafe runtime must be replaced.
CryoDAQ runs a local text-generation service (current brand: Gemma, default model gemma4:e4b via Ollama; downgraded to gemma4:e2b on low-VRAM dev machines). It uses no external model API; Telegram delivery is a separate external service and remains part of the open credential/authority boundary.
Subscribes to engine events (alarms, phase transitions, finalize, sensor anomalies, shift handovers). When an alarm fires or an experiment finalizes, it generates a human-readable summary for the operator in:
- Telegram (the bot chat)
- Operator log
- GUI insight panel (an overlay in MainWindowV2)
It also generates diagnostic suggestions (alarms + sensor_anomaly_critical) and intro paragraphs for campaign DOCX reports.
Answers operator queries. IntentClassifier determines intent, QueryRouter routes the query to adapters: live data via BrokerSnapshot and semantic search over the knowledge base (KNOWLEDGE_QUERY → RAG). Available from the embedded chat in the "Knowledge base" overlay and via the Telegram bot.
The intended boundary is read-only engine queries plus operator text. The current candidate does not yet satisfy it: the assistant still has a second operator-log writer, a RAG mutation path, and a Telegram delivery credential. It must not be described as wholly observational or read-only until those paths are removed or reassigned and the gate is reviewed.
See config/agent.yaml. Key parameters:
agent.enabled: enable/disable the serviceagent.brand_name: the operator-facing name (can change when migrating to another model)agent.ollama.default_model: the Ollama modelagent.triggers.*: which events activate the serviceagent.rate_limit: limits (60 calls/hour by default)
ollama pull <new_model>- Edit
config/agent.yaml:agent: brand_name: "New name" brand_emoji: "🦉" ollama: default_model: <new_model>
- Restart the engine
- Smoke test: trigger an alarm in mock mode
No code changes.
Two lanes: a live observer (subscribes to engine events, emits summaries) and a
query router (classifies operator intent, routes to read-only adapters). See
docs/architecture.md for the system architecture.
Every model call is recorded under data/agents/.../audit/<YYYY-MM-DD>/.
Full context, prompt, response, tokens, latency, output targets. A verifiable
trail for post-hoc review.
These limitations apply at the current v0.64.1/checkpoint boundary. The
software and laboratory checks are collected as a turnkey protocol in
docs/lab_verification_checklist.md.
- SQLite WAL gate: the engine hard-fails on startup on SQLite versions in the
range
[3.7.0, 3.51.3)(F25). Backport-safe: 3.44.6, 3.50.7 (pass without the variable). The supported Windows/Linux installation usesenvironment.ymlto pin a safe SQLite version. No fallback package is installed by default; an unsafe or absent implementation remains a startup failure. The bypass is emergency acknowledgement only and is not acceptable deployment evidence. - Lab Ubuntu PC verification: CI and WSL exercise the H5 ZMQ idle-death contract, but they do not close the physical laboratory-Ubuntu gate. Run and record the checklist procedure on the actual lab PC.
- Engine shutdown warning: one
Unclosed client sessionERROR can appear at engine shutdown because anaiohttpsession is not closed on that exit path. This is cosmetic on shutdown; data and safety state are not affected. - PDF reports: best-effort. The guaranteed artifact is DOCX.
- Runtime calibration policy: global on/off + per-channel KRDG/SRDG+curve. A conservative fallback to KRDG when a curve / SRDG is missing or a computation fails. Real LakeShore behavior requires lab verification.
- Leak rate (F13):
chamber.volume_lmust be set inconfig/instruments.local.yamlbefore the first measurement;finalize()raisesValueErrorwhenvolume_l == 0.0. - Keithley host-death protection: the TSP v3 script is intentionally
non-autonomous and covers only a late pet after a host stall.
requiredmode therefore refuses v3;best_effortlogs a CRITICAL degraded warning and uses the late-pet check. No software status bit proves physical terminal OFF. Host-death removal of terminal energy and any external interlock remain lab gates measured with independent instruments.
This repository is built to be worked by an agent, and it distinguishes two roles that are easy to conflate:
- The building role. Arranges existing blocks and creates new blocks from
instructions. This is ordinary implementation work, and it is what most of
AGENTS.mdaddresses. A mid-tier model should be able to do it. - The governing role. Decides how the blocks lie against each other and whether they are well made: whether a premise is true, whether an invariant is the right one, whether a guard can actually fail.
Only the first of those is agent-agnostic, and that asymmetry is deliberate rather than an omission. The building role must stay within reach of a mid-tier model, because that is what lets a lab adapt this repository to its own hardware with whatever agent it has; any instruction such a model cannot follow is a defect in the instruction, not in the model. The governing role is capability-gated and cannot be made agent-agnostic by wishing it. A model too weak to hold a subsystem in view, construct a mutation that tries to disprove a fix, or notice an invariant applied in one place and omitted in another does not review badly -- it reviews invisibly, producing a review-shaped artifact with nothing behind it. That is worse than having no second reviewer, because it manufactures confidence. Name the model used for each review; the constraint is capability and independence, never a particular vendor.
Today that adaptation promise covers configuration and passive-driver work, not
end-to-end adoption of an arbitrary hazardous actuator. The loader can admit a
rostered source, but SafetyManager remains Keithley-shaped
(src/cryodaq/engine.py:2082-2158, src/cryodaq/core/safety_manager.py:140-158,
:2203, :2434, :2605); the generic actuator contract is an exposed gap.
The distinction matters because a guard is governing work wearing implementation clothes. A test that merely exercises code is building work. A guard is a claim about what could go wrong -- and that claim can only be made by someone who knows how the whole thing is wired.
When one agent writes the code and then writes its own guard, in one sitting, from one mental model, the guard inherits the code's blind spots. It passes, it reads as coverage, and it cannot catch the defect it names. During the qualification-campaign review this happened five separate times, each caught by an independent reviewer rather than by the guard:
- a summariser guard fed itself tidy synthetic input, so four successive versions of its parser shipped broken;
- an evidence rule was verified with a single-partition fixture, so it missed every multi-partition record;
- a structurally correct pytest plugin was proven by in-process tests, while in production it was silently stripped from the sealed candidate's environment and never loaded at all.
None of those were careless. Each was the building role being asked to do the governing role's job.
The operating consequence: what a guard must falsify is specified by the
governing layer; the guard is then implemented by the building role. Whoever
authored a correction does not get to decide, alone, what would prove it wrong.
This is the same separation the registry already applies to dispositions --
disposition_owner exists so that no author closes its own record -- extended
to the point where it is actually load-bearing.
Two corollaries worth stating plainly:
- A guard must be exercised against the real invocation path, not a convenient stand-in. For any mechanism that only demonstrates itself on failure, follow the environment and arguments all the way into the process that really runs it. In-process and structural tests cannot see a transport-level break.
- Documenting a trap is not preventing a trap. Every hazard hit during the qualification-campaign review was already written down somewhere in this repository. The goal is not more prose telling an agent what to avoid; it is a governing layer that makes the failure unreachable.
Repository-wide engineering and safety rules live in AGENTS.md.
CLAUDE.md contains subordinate ecosystem-specific convenience
guidance; it cannot override AGENTS.md or the detailed, tool-neutral workflow
in docs/ORCHESTRATION.md.
Historical prompts, handoffs, generated memory, and agent-run artifacts are not
current policy unless an active task explicitly selects them.
See LICENSE. Third-party notices: THIRD_PARTY_NOTICES.md.




