KYMA v3 Symbolic Composition Probe — Preregistration¶
Date: 2026-09-29
Status: frozen before any model is trained. The freeze is the commit that adds
this file and pushes it to the public remote; the code it names
(src/scpn_quantum_control/benchmarks/kyma_v3/, scripts/run_kyma_v3_probe.py)
is in the same commit. No substrate or baseline has been trained on this task.
The result will be appended below a RESULT heading; nothing above it changes
after training.
0 QPU. This is a classical oscillator-substrate probe.
Scientific question¶
Does the gated-coupling oscillator substrate still generalise to a held-out combination of learned operations when the ground truth is produced by a symbolic program rather than by oscillator dynamics? (KYMA Part B assumption A2, MS1 criterion family.)
v2 (docs/campaigns/kyma_v2_composition_probe_2026-07-21.md, PASS) used an
oscillator teacher, so its task lived inside the substrate's hypothesis class.
v3 removes that objection: no oscillator, integrator or model is in the label
path.
Authority: the KYMA v3 probe specification (seven requirements);
owner decision A9 (probe GO; a negative is reported with its diagnosis) and
owner decision B3 of 2026-09-29 (option (a): evaluate the held-out pair on
query a only and state that restriction here).
Ground truth and split¶
Three registers (a, b, c) with values in Z4. Operations:
R0:(a, b, c) → (b, c, a);R1:a ← (a + b) mod 4;R2:b ← (b + 1) mod 4.
A configuration is one operation or an ordered pair applied left to right (3 singles + 9 ordered pairs). An item is (initial state, configuration, queried register); its label is the final value of the queried register.
- Training: every configuration except the held-out pair, all 64 states, all three queries: 11 × 64 × 3 = 2,112 items.
- Held-out test: the ordered pair
(R0, R1), all 64 states, queryaonly: 64 items per seed. The pair's queriesbandcare neither trained nor evaluated. - Restriction (owner decision B3): query
ais the only query of the pair whose answer is not a function of the single-operation answers. For queriesbandcthe held-out answer equals theR0-alone answer, so they cannot test composition and are excluded.
Every query register appears in every trained configuration.
Teacher-free design checks (recorded before training)¶
Computed by kyma_v3.task.design_report() and kyma_v3.probe.realisability_accuracy()
from the symbolic program and hand-set gates only:
| check | value |
|---|---|
every configuration is a bijection on Z4^3 |
yes |
| label counts per (configuration, query) | exactly 16 per class (all 36) |
| held-out states whose answer the single-op answers do not fix: query a / b / c | 1.00 / 0.00 / 0.00 |
minimum Hamming distance of the held-out answer vector (b + c) to any trained (configuration, query) answer vector |
48 of 64 |
| measured training-marginal chance floor on the test items | 0.25 |
| hand-set substrate, all 12 configurations × 64 states × 3 queries | 2,304 / 2,304 exact; worst phase error 0.026 rad against a 0.785 rad margin |
The last row shows the task is realisable inside the substrate class. It is a validity check only: training never sees the hand-set gates.
Substrate (frozen)¶
Each register is one oscillator; its phase relative to a fixed reference
(phase 0) encodes the value, v ↦ v·π/2. Two banks of three oscillators. One
operation is one write stage: the destination bank is reset to the
off-lattice phase π/4, the source bank receives no coupling (it is held), and
the destination integrates
dθ_i/dt = Σ_j K[o,i,j] sin(x_j − θ_i + α[o,i,j])
+ Σ_p T[o,i,p] sin(x_{p1} + x_{p2} − θ_i + β[o,i,p])
with directed pairwise couplings K and lags α (source register j → target
i), and triadic couplings T and lags β over the three source-register
pairs p ∈ {(a,b), (a,c), (b,c)}. The operation code o gates which couplings
act. The banks then swap roles; an ordered pair is two write stages with the
gates of each operation in program order. Readout: the queried register's final
phase, rounded to the nearest lattice value.
- Integrator: fixed-step RK4,
dt = 0.05, 60 steps per stage (T = 3). - Trainable parameters:
K, α, T, βfor 3 operations × 3 targets × 3 sources or pairs = 108. - Initialisation: every parameter drawn from
N(0, 0.3²)with the run seed. - Loss: mean
1 − cos(φ_q − label·π/2)of the queried register's final phase. Only the final queried value is supervised.
Stated limitation. Two architectural choices were made for realisability, not
from any model's performance: the staged schedule (one write stage per
operation) and the triadic term (pairwise phase coupling cannot add phases,
which R1 requires). Because the stages are applied in program order, a
substrate that learns each operation exactly composes the held-out pair by
construction. The probe therefore tests whether gradient descent learns
reusable operation gates from end-of-program labels alone, and the diagnostic
baselines below test whether staging alone, without oscillator dynamics, does
as well.
Baselines (frozen)¶
Contract baselines, parameter count within ±10 % of the substrate (108):
| baseline | architecture | width | parameters |
|---|---|---|---|
| MLP | one tanh hidden layer over sin/cos of the three register phases, one-hot first op, one-hot second op (with "none") and one-hot query (16 inputs) |
5 | 109 (+0.9 %) |
| staged GNN | register nodes; per-operation learned 3×3 adjacency and node bias; embedding, then one message-passing round per operation in program order; read at the queried node | 4 | 103 (−4.6 %) |
| transformer | one single-head attention layer with residual over tokens [first op, second op, query, a, b, c]; learned embeddings and positions; read at the query token |
3 | 100 (−7.4 %) |
Diagnostic baselines (reported; they never change the verdict):
| diagnostic | purpose | width | parameters |
|---|---|---|---|
| sequential MLP | one small MLP per operation maps sin/cos phases to sin/cos phases, applied in program order; label = queried angle rounded to the lattice; loss as the substrate's |
2 | 96 (−11.1 %, outside ±10 %; diagnostic only) |
| MLP, large | capacity control | 64 | 1,348 |
| staged GNN, large | capacity control | 16 | 703 |
| transformer, large | capacity control | 16 | 1,348 |
Chance floor: always predicting the most frequent training label (ties to the smallest), measured on the test items (0.25 by the design check).
No baseline receives intermediate states or any privileged information.
Training (frozen)¶
Full-batch Adam on the 2,112 training items for every model; 3,000 epochs for every model; learning rates fixed in advance without search: substrate 0.05, MLP 0.02, staged GNN 0.01, transformer 0.01, sequential MLP 0.02. Seeds 0, 1, 2, 3, 4. No early stopping, no model selection and no tuning on the held-out items. Training accuracy is reported for every model and seed.
Decision procedure (the contract)¶
Primary statistic: held-out accuracy (fraction of the 64 test items classified correctly), per model and seed, reported as mean ± population standard deviation over the five seeds.
PASS iff both hold:
- substrate mean ≥ (best contract-baseline mean) + 10 percentage points; and
- substrate mean − substrate sd > the measured chance floor.
Otherwise NEGATIVE, reported with a diagnosis (training accuracy of each operation's single-op items, per-seed results, where composition failed).
Design-selection seed: the design checks use no random draws, so no seed was used for design; all five seeds count and the "excluding the design seed" robustness check is vacuous. It is stated here so it is not read as omitted.
Pre-declared attribution rule. If the staged GNN or the sequential MLP reaches within 10 percentage points of the substrate mean, any Part B wording must attribute the generalisation to staged operator application, which non-oscillator staged models share, and not to oscillator dynamics specifically. If a large-capacity diagnostic reaches within 10 points, the matched-budget margin is stated as budget-dependent. These rules apply whatever the verdict.
Exploratory (labelled as such if reported): anything beyond the per-seed accuracies, training accuracies and the attribution flags above.
Energy¶
For every model: J/task = declared nominal package power of the host CPU × the measured wall-clock seconds per test item of a warm inference pass. This is an energy proxy, not a measurement; no power meter is read, and it is not evidence for oscillator-hardware frugality. The declared power and the host are recorded in the artefact.
Compute and environment¶
Host: ML350 (approved compute host). Before the run: health check, uptime,
other projects' running units; at most 12 threads; a systemd-run --user unit;
outputs under ~/, then copied into the repository. Environment: the pinned
requirements-ci-py312-linux.txt plus requirements-ci-jax-py312-linux.txt
(jax 0.10.1, CPU).
Analysis and data¶
- Analysis:
scripts/run_kyma_v3_probe.py, which callsscpn_quantum_control.benchmarks.kyma_v3.probe.run_probeand applies the contract inprobe.verdict, frozen at the freeze commit. The artefact records the source commit. - Data:
data/kyma_v3_symbolic_composition/kyma_v3_symbolic_composition.json; archived with the next Zenodo release.
Amendments¶
Any change before training is a new commit labelled as an amendment in this file, pushed before the run. No change after training starts.
Reporting¶
The measured result is reported whichever way it falls, with the attribution flags. The CEO receives this file's path and SHA-256 at freeze, and the result or truthful "in progress" wording by 12 October 2026. Part B quotes no partial or assumed number.
Seat: 90ad
RESULT¶
Appended 2026-09-30 after the run; nothing above this heading has changed since the freeze commit ddae107dc
(file SHA-256 before this section: 9248cf8a4fe9bfaf4d4807133f86dd9da51fac821fb71f70a3e8932eb232d1a5).
Verdict under the frozen contract: PASS.
- Run: ML350 (hostname
god-of-the-math),systemd-run --userunit pinned to cores 0–11, JAX CPU; source commitdbcb0b0778221f85723ab7f60118c27e6369260f, a descendant of the freeze commit with zero changes undersrc/scpn_quantum_control/benchmarks/kyma_v3/,benchmarks/kyma_v2/andscripts/run_kyma_v3_probe.py(it adds only CI registrations and evidence digests). Started 2026-09-29T19:01:33Z (freeze pushed 17:11Z), finished 2026-09-30T01:25:17Z; 9 h 08 min CPU time. Artefactdata/kyma_v3_symbolic_composition/kyma_v3_symbolic_composition.json, SHA-256dd27ae8953d660626abbd23f23a7ea9c5c13123fa93280bef092982104f90cb8(identical on the host and in the repository). - Design checks in the artefact match the table above: 2,112 training items, 64 test items, ambiguity by query 1.00 / 0.00 / 0.00, minimum distance 48, uniform label counts, hand-set realisability 1.000.
Held-out accuracy on the pair (R0, R1), query a, mean ± population SD over seeds 0–4 (training accuracy in
brackets):
| model | params | held-out | training |
|---|---|---|---|
| substrate | 108 | 1.000 ± 0.000 | 0.9995 ± 0.0009 |
| MLP (contract) | 109 | 0.256 ± 0.041 | 0.677 ± 0.012 |
| staged GNN (contract) | 103 | 0.441 ± 0.099 | 0.943 ± 0.027 |
| transformer (contract) | 100 | 0.253 ± 0.015 | 0.504 ± 0.002 |
| sequential MLP (diagnostic) | 96 | 0.247 ± 0.025 | 0.495 ± 0.025 |
| MLP, large (diagnostic) | 1,348 | 0.234 ± 0.043 | 1.000 ± 0.000 |
| staged GNN, large (diagnostic) | 703 | 0.306 ± 0.021 | 1.000 ± 0.000 |
| transformer, large (diagnostic) | 1,348 | 0.250 ± 0.000 | 0.496 ± 0.008 |
Every seed of the substrate classified all 64 held-out items correctly (seed 2 missed 5 of 2,112 training items).
Contract. (1) 1.000 ≥ 0.441 + 0.10: margin over the best contract baseline (staged GNN) +55.9 percentage points. (2) 1.000 − 0.000 > chance floor 0.25. Both hold → PASS.
Attribution rule (pre-declared). Staged GNN within 10 points of the substrate: no (0.441). Sequential MLP:
no (0.247). Large-capacity diagnostics within 10 points: no (best gnn_large 0.306). The generalisation is
therefore not attributed to staged operator application alone, and the matched-budget margin is not stated as
budget-dependent within the budgets tested.
What the result supports and what it does not. As stated in the frozen limitation, the staged schedule makes a substrate that learns each operation exactly compose the held-out pair by construction. The result shows that gradient descent learned the three operation gates from end-of-program labels alone, and that parameter-matched and larger non-oscillator models, including two with the same staging, did not generalise to the pair. It does not separate the oscillator dynamics from the other architectural priors chosen for realisability (phase-lattice value encoding, reset phase, triadic coupling that can add phases); that separation would need an ablation not pre-registered here.
Energy proxy (not a measurement). Declared nominal package power 190 W × measured warm wall time per test item, on cores shared with other projects' jobs on the host: substrate 1.68 J/item (8.8 ms/item, RK4 simulation on CPU); contract baselines 0.003–0.043 J/item; all baselines ≤ 0.050 J/item. The simulated substrate costs 33–536 times more per item than the baselines on this CPU. This says nothing about oscillator hardware.
Exploratory (labelled; not part of the contract). The matched-budget contract baselines under-fit the training set (MLP 0.68, transformer 0.50), while the large MLP and large staged GNN fit it exactly and still stay near chance on the held-out pair (0.23, 0.31): the baselines' held-out failure is not explained by capacity alone.