Skip to content

KYMA v2.1 supplementary rigor — ablations, baselines, convergence, LOO — 2026-07-21

Status: complete — all four analyses run. Strengthens and corrects the landed v2 PASS (db7c5c47). Code: src/scpn_quantum_control/benchmarks/kyma_v2/{ablations,baselines,rigor}.py, runner scripts/run_kyma_v2_rigor.py, tests tests/test_kyma_v2_{ablations,baselines,rigor}.py, artifact data/kyma_v2_composition_probe/kyma_v2_1_rigor.json. Pre-registration (predictions frozen before the run): .coordination/planning/CEO/KYMA_V2_1_SUPPLEMENTARY_RIGOR_PREREGISTRATION_7f6b_2026-07-21.md. 0 QPU. All numbers below are at the frozen v2 budget (5 seeds, 1500/2000 epochs).

Why v2.1

v2 landed a PASS (substrate 80.1 % vs param-matched MLP 36.9 % on a held-out conjunction, +43.1 pp). v2.1 pre-registers four analyses to (a) attribute the gap causally and (b) pre-empt the standard objections — capacity, undertraining, "any relational bias", one-lucky-split — and to test whether each of the two v2 fixes is load-bearing. Every prediction was frozen before the run; a failed prediction is reported as an honest finding.

#1 Ablations — is each fix load-bearing?

ablation result pre-registered prediction verdict
A1 — no coupling gating (shared-K + code-drive) shared 0.331 vs MLP 0.369, margin −0.039 shared ≤ 0.55 AND margin < 0.10 MET
A2 — separable readout (bridge to R1 cluster only) substrate 0.895 vs MLP 0.385, margin +0.510 margin < 0.10 REFUTED

A1 confirms fix 1. With one shared coupling (v1's architecture) the student collapses to 0.331 — below the MLP and below any useful level. A single symmetric K cannot be attractive (in-phase) and frustrated (anti-phase) at once; per-relation gating is necessary.

A2 refutes its prediction — and this is the most informative result. Even with a separable, single-relation readout, the substrate still beats the MLP by 51 pp. So the MLP's failure is not primarily about non-separability or "forcing the relations to interact". The load-bearing difficulty is that the label is a θ0-dependent achieved phase (the circular-mean lock of a cluster's initial phases) — a hard nonlinear function of the input the MLP cannot approximate even for one relation, while the substrate integrates the dynamics. This corrects the v2 campaign doc's mechanism language (a correction note is now recorded there): the empirical v2 result and the compositional-generalisation claim stand; what changes is why the baselines fail — dynamics-computation, of which composition is one instance.

#2 Stronger baselines — capacity? or something only oscillators have?

model accuracy params substrate − model prediction verdict
gated substrate 0.801 336
param-matched MLP (v2) 0.369 361 +0.432
over-parameterised MLP 0.387 1330 (≈4×) +0.413 < 0.55 AND ≥ 0.20 gap MET
deep 2-layer MLP 0.411 1828 +0.389 (context)
code-conditioned GNN 0.507 548 +0.294 ≥ 0.15 gap MET

Capacity does not close the gap: a ~4× MLP and a 2-hidden-layer MLP both stay near the param-matched MLP. The GNN is the load-bearing control — it is given the same relational structure as the substrate (a code-conditioned adjacency telling it which oscillators are in the active pairs) but aggregates with learned message passing instead of oscillator dynamics. It does better than a plain MLP (0.507 vs 0.369 — relational bias helps) yet still loses to the substrate by 29 pp. So the advantage is specific to the oscillator dynamics, not capacity and not a generic relational inductive bias.

#3 MLP convergence — is the failure real, or undertraining?

The param-matched MLP reaches training accuracy 1.000 with a plateaued loss (0.023 → 0.001, final 0.0005) — prediction MET. It memorises the training conjunctions perfectly; its ~37 % held-out accuracy is therefore a genuine compositional-generalisation failure, not undertraining. This rules out the simplest objection to the whole comparison.

#4 Leave-one-out over held-out conjunctions

Each of the six disjoint conjunctions is held out in turn (v2 used only (R1-AB, R2-CD)); the readout bridge and the §5 design constants (notably k_bridge, which ranged 0.6–2.0 across splits) are re-derived per split from teacher dynamics only. 3 seeds per split.

held-out (r1,r2) substrate MLP margin
(0, 5) 0.787 0.361 +0.426
(1, 4) 0.836 0.386 +0.450
(2, 3) 0.840 0.376 +0.464
(3, 2) 0.803 0.408 +0.396
(4, 1) 0.838 0.373 +0.464
(5, 0) 0.786 0.387 +0.399
mean 0.815 0.382 +0.433

Substrate > MLP on 6 of 6 splits, mean 81.5 % vs 38.2 %, mean margin +43.3 ppprediction MET (≥ 70 % AND ≥ 20 pp AND ≥ 5/6). The advantage is robust to which conjunction is held out, not an artefact of one lucky split.

What v2.1 establishes (honest, bounded)

The substrate's held-out advantage is (i) not capacity (over-parameterised and deep MLPs fail), (ii) not undertraining (MLP fits training at 100 %), (iii) not a generic relational bias (a code-conditioned GNN loses by 29 pp), (iv) dependent on coupling gating (A1), and (v) robust across all six held-out conjunctions (LOO, 6/6, mean +43.3 pp). The decisive property is that the readout is a θ0-dependent achieved phase the substrate computes by integrating dynamics and the baselines cannot approximate — a sharper mechanism than v2's original "non-separability" framing (corrected there). The teacher–student limitation from v2 is unchanged: the ground truth is oscillator-generated, so the claim remains bounded to compositional phase-locking, not arbitrary tasks.

Authored by Anulum Fortis & Arcane Sapience (protoscience@anulum.li) Seat: 7f6b