Skip to content

ANN-to-SNN Conversion

Convert trained PyTorch ANNs to rate-coded spiking neural networks.

Contract

Training and ANN extraction require PyTorch. The exported ConvertedSNN runtime uses NumPy, optionally dispatches replay to configured Rust, Go or Julia, and can run without PyTorch; convert and ReLU replacement report a missing dependency when called without Torch, and resolving QCFSActivation requires Torch.

  • convert(model, calibration_data=None, T=None, percentile=99.9, max_working_bytes=268435456) captures an independent inference graph and lowers supported dense operations in actual forward order. Shared modules retain each invocation, unused registrations do not define topology, and consecutive affine maps are composed without inserting IF nonlinearities.
  • Each ReLU invocation, including function and tensor-method forms, receives its own calibrated scale. Without calibration, ReLU uses unit scale and ReLU6 retains its saturation scale of six. ReLU IF neurons start from rest.
  • Each QCFS invocation retains its learned threshold and a half-threshold IF preload. Mixed ReLU/QCFS networks keep separate scales and preloads. T=None adopts a common trained QCFS budget, or 16 without QCFS; differing trained budgets require an explicit simulation T. Matching T does not guarantee zero conversion loss: validate source/target loss for the declared encoding.
  • ConvertedSNN.run(x) rate-codes NumPy input with a fixed RNG seed and returns final-layer spike counts or signed integrated linear output for one vector or a batch. Dense runtime weights must be two-dimensional; convolution execution is not part of this runtime profile. initial_membrane_fraction supplies a global default preload; layer_membrane_fractions overrides it for individual stages (0.0 ReLU rest, 0.5 QCFS shift). A linear readout starts from zero regardless of these fractions.
  • ConvertedSNN.rates(x) divides the response by T and restores the source activation scale. classify(x) selects the first greatest response index.
  • QCFSActivation replaces ReLU during conversion-aware training by clipping activations to [0, theta] and quantising them to T + 1 spike-rate levels with a straight-through gradient. T is an integer in [1, 2**32 - 1], the step domain every native counterpart shares. Only interior elements carry a gradient, so a saturated infinite or NaN neighbour never turns a batch's threshold gradient into NaN; upper saturation contributes one and lower saturation zero, also for infinite inputs.
  • replace_relu_with_qcfs(model, T=8, theta=1.0, learn_theta=True) substitutes every ReLU/ReLU6 in a model (recursing through submodules) with a QCFSActivation, preparing the network for QCFS conversion-aware fine-tuning.

The dense target admits Linear, module/function/method ReLU and ReLU6, QCFS, inference identity/dropout and batch-preserving module Flatten operations. Residual arithmetic, unrepresented operators, multiple outputs and input-dependent Python control flow fail compatibility admission. Source dtype rounding and affine-composition reduction order remain part of measured source/target loss. Source graph capture uses a model copy. Tracing, source calibration and calibrate_activation_thresholds restore the Python, NumPy and PyTorch CPU generators and the generator of every device of the current accelerator type, also when user code raises; custom user forward code can still perform other external effects. These sections are serialised within the process, so concurrent conversions cannot restore one another's states; code that draws from the global generators concurrently outside a conversion is not isolated. Accelerator custody was verified on a CUDA device. Source copying, inserted identity/fused matrices and source calibration buffers have separate metadata checks against the supplied budget.

Replay and buffer budgets

ConvertedSNN.replay(frames, initial_state=None, trace=False, binary_inputs=True, max_working_bytes=268435456, backend="auto") consumes explicit (steps, batch, input_neurons) frames. Binary drive requires exact zero/one values; binary_inputs=False accepts finite currents in [0, 1]. The result owns its output, final states and optional complete state/event traces. Supplying previous final states continues the trajectory without modifying caller buffers. Spiking output counts events in this call; linear output retains the cumulative readout integral.

The dense-if-f64-sequential-v1 profile accumulates input columns in ascending order with separate float64 multiply and add operations. Bias is added every step. IF events fire at equality, propagate within the same timestep and reset by subtracting one threshold. Each neuron emits at most one event per step. The final linear layer integrates signed current without firing or reset.

Constructor snapshots and replay, run, rates, and classify accept an operator-selected max_working_bytes; each call defaults to 256 MiB. Invalid budgets raise ValueError; reservations above the budget raise MemoryError before the relevant allocation: coefficient storage is checked before snapshots, and the complete replay reservation before frame or state copies. The reservation is 8 * (2P + 2F + 2S + 2O + H + 5M + I) bytes, where P counts coefficients, F all frames, S all membranes, O final outputs, H requested state/event traces, M the largest layer and I one input frame. The bound covers numeric temporaries; it excludes caller-owned storage, interpreter and allocator overhead, and is not a process-RSS guarantee. Array-like coercion of Python containers precedes metadata admission; existing array metadata is inspected without materializing broadcast views. run generates frames in blocks of at most 64 timesteps. An empty batch performs no timestep iteration, including at T=2**53. Native encoded runs longer than 64 steps additionally reserve 8 * batch * (sum(layer_widths) + final_width) for the preceding opaque result owner retained during the next block. This reservation is admitted before encoding begins. Explicit replay treats supplied initial-state storage as caller-owned storage.

Shared registered ReLU/ReLU6 children remain shared when replace_relu_with_qcfs prepares a source for fine-tuning. Every registration alias points to one QCFS activation and one learned threshold, including aliases through separate parents. The outer model and each replaced activation retain their training modes. The root module itself is returned unchanged.

Python native replay

All four runtime methods accept backend="auto", "numpy", "rust", "go", "mojo", or "julia". Auto tries SC_NEUROCORE_IF_RUST_LIB, then SC_NEUROCORE_IF_GO_LIB, then SC_NEUROCORE_IF_MOJO_LIB, then Julia when SC_NEUROCORE_IF_JULIA_ENABLED=1, and otherwise selects NumPy. Without a comparison this order is a configuration preference. To use measured native ordering, set SC_NEUROCORE_IF_BENCHMARK to the complete JSON emitted by benchmarks/bench_ann_to_snn_replay.py in a source checkout, or by the sc-neurocore-if-benchmark command an installed wheel provides. Both bind the source bytes of the imported sc_neurocore package and of the comparison scripts that ran, so a report captured from byte-identical sources is admitted in either location. Auto verifies the same CPU, Python/NumPy versions, uninstrumented measurement, current source bytes and every configured native artifact. All five providers must have the same complete 20-workload response digests; workload medians and the aggregate are recomputed from positive raw samples. Stale, malformed or incompatible explicit reports raise RuntimeError; a different CPU retains the static order. The measured native order never displaces the NumPy floor. Explicit backend selection bypasses the report. This records a local warm public-call ranking, not trained accuracy or energy. Explicit NumPy always uses the reference runtime. Explicit native selection requires its corresponding configuration. A configured library that cannot load or has incompatible entry points raises RuntimeError; native errors propagate without silent fallback. Selection is checked even for empty encoded batches.

The wheel ships the dense IF Rust, Go, Mojo and Julia sources, so the native libraries can also be built from an installed package: replace src/sc_neurocore below with the installed package directory. Build Go with -buildvcs=false there, because an installed package is not a module checkout.

Build the actual dependency-free Rust ownership ABI from the source checkout:

Bash
cargo build --offline --release \
  --manifest-path src/sc_neurocore/accel/rust/safety/if_native/Cargo.toml \
  --target-dir /your/build/directory
export SC_NEUROCORE_IF_RUST_LIB=/your/build/directory/release/libsc_neurocore_if_replay.so
Python
from sc_neurocore.conversion import ConvertedSNN

snn = ConvertedSNN([[[1.0]]], [None], [1.0], T=129)
result = snn.replay([[[1.0]]], trace=True, backend="rust")
counts = snn.run([1.0], backend="rust")

Rust replays the same ordered float64 profile, with complete owned state and event traces. NumPy views retain their native result owner and supplying library until the last view expires; slicing also retains that lifetime. Final states and traces occupy independent buffers. Invalid domain, numeric reservation refusal, and arithmetic overflow map to ValueError, MemoryError, and FloatingPointError. Libraries are loaded from explicit configuration; calling the runtime does not build or download them. Go exposes the same ABI-one result contract through its C shared library. Julia uses managed JuliaCall entry points; Mojo provides the same C ownership ABI through an explicitly built library. Normal garbage collection releases an owner after its last view expires. Forced finalization at interpreter exit is disabled, preserving retained views during other exit hooks; process teardown reclaims any remaining storage.

Python Julia replay

Prepare an installed Julia executable and an existing locked Julia project whose Project.toml pins PythonCall to =VERSION, exactly matching the installed Python juliacall version. Its Manifest.toml must resolve that same version. The adapter checks these files and refuses drift; it does not install packages or resolve the project. Conflicting python -X overrides for Julia's executable, project, worker count, signal handling or initialization are refused before JuliaCall imports. An existing runtime must actually have one default-pool worker and managed signal handling enabled; changing environment variables after startup cannot make an incompatible runtime eligible.

Bash
export SC_NEUROCORE_IF_JULIA_ENABLED=1
export PYTHON_JULIACALL_EXE=/absolute/path/to/julia
export PYTHON_JULIACALL_PROJECT=/absolute/path/to/locked/project
export PYTHON_JULIACALL_THREADS=1
export PYTHON_JULIACALL_HANDLE_SIGNALS=yes
export JULIA_CONDAPKG_BACKEND=Null
export JULIA_PKG_OFFLINE=true

On Linux, Julia 1.13 bundles an OpenSSL whose libssl needs OPENSSL_3.3.0 symbols. A Python that links the system libcrypto (distribution packages, actions/setup-python) has already loaded an older libcrypto.so.3 by the time JuliaCall starts, and the runtime fails with version `OPENSSL_3.3.0' not found. Preload the runtime's own library for that process (uv's standalone Pythons carry OpenSSL statically and do not need it; Julia 1.11 does not need it):

Bash
export LD_PRELOAD="$(dirname "$PYTHON_JULIACALL_EXE")/../lib/julia/libcrypto.so.3${LD_PRELOAD:+:$LD_PRELOAD}"

Initialize the configured runtime before importing PyTorch or using conversion factories. Resolving ConvertedSNN itself only loads the NumPy runtime.

Python
from sc_neurocore.conversion.if_julia import load_julia
load_julia()

from sc_neurocore.conversion import ConvertedSNN
snn = ConvertedSNN([[[1.0]]], [None], [1.0], T=129)
result = snn.replay([[[1.0]]], trace=True, backend="julia")
counts = snn.run([1.0], backend="julia")

The packaged Julia module uses the same owned replay kernel and ABI-one layouts. Its borrowed LayerParameters entry admits the complete reservation before copying each layer once. Results stay rooted in Julia until the last Python view expires; a C-allocated metadata slot identifies each result. Complete numeric traces are exposed without copying them into another output buffer. Managed calls reject runtime shutdown and avoid entering Julia after its exit hook. Raw C callbacks require a registered Julia runtime thread and the same live-pointer and exactly-once-release contract as the other native providers.

The adapter retains JuliaCall's GIL serialization. JuliaCall documents managed Python-thread support as experimental and recommends explicit signal handling: JuliaCall threading. With this signal setting, Ctrl-C does not raise Python KeyboardInterrupt. The dedicated public integration corpus exercises both installed Julia runtime families, full state/event parity, continued states, retained slices across Python and Julia garbage collection, 16 worker calls and process exit refusal. The direct callback corpus independently exercises native layouts and refusal slot custody. These checks do not establish support for every host platform.

Verification

The public conversion files are covered by the scoped NumPy-docstring policy:

  • src/sc_neurocore/conversion/__init__.py
  • src/sc_neurocore/conversion/ann_to_snn.py
  • src/sc_neurocore/conversion/qcfs.py
  • src/sc_neurocore/conversion/calibration.py
  • src/sc_neurocore/conversion/converted_snn.py
  • src/sc_neurocore/conversion/if_parameters.py
  • src/sc_neurocore/conversion/if_replay.py
  • src/sc_neurocore/conversion/if_encoding.py
  • src/sc_neurocore/conversion/if_resources.py
  • src/sc_neurocore/conversion/if_inputs.py
  • src/sc_neurocore/conversion/if_dispatch.py
  • src/sc_neurocore/conversion/if_native.py
  • src/sc_neurocore/conversion/if_native_types.py
  • src/sc_neurocore/conversion/if_julia.py
  • src/sc_neurocore/conversion/if_julia_types.py
  • src/sc_neurocore/conversion/if_julia_configuration.py
  • src/sc_neurocore/conversion/model_trace.py
  • src/sc_neurocore/conversion/source_graph.py
  • src/sc_neurocore/conversion/source_calibration.py

Focused production tests live in the split tests/test_conversion_*.py suites. They exercise real PyTorch modules, the threshold-balancing and QCFS conversion routes, ConvertedSNN.run, ConvertedSNN.classify, the membrane shift, the ReLU→QCFS substitution helper, QCFS range and gradient behaviour, and the layer-extraction contract.

Additional tests/test_conversion_*.py suites exercise full state/event trajectories against independent scalar arithmetic, state continuation, parameter custody, signed readouts, calibration mode/hook/RNG preservation, actual execution without Torch and exact buffer-budget boundaries. Source graph tests additionally cover actual forward order, repeated modules, functional activation calibration, affine composition, mixed-route preloads, source model and CPU RNG custody, bfloat16 observations and genuine compatibility refusals. tests/test_conversion_random_custody.py converts models whose forward draws from every global generator, in two threads at once and through a failing forward, and requires every generator state to be unchanged.

Julia explicit-frame replay

The native AnnToSnnAccel module exposes owned DenseLayer, ConvertedSNN, ReplayResult, replay and zero-based classify. Its coefficient, frame, state and trace vectors use row-major ordering, matching the Python replay profile. Source graph tracing and stochastic input encoding belong to the frontend.

Julia
include("src/sc_neurocore/accel/julia/conversion/ann_to_snn.jl")
using .AnnToSnnAccel

layer = DenseLayer(1, 1, [1.0], [0.25], 1.0; initial_fraction=0.5)
model = ConvertedSNN([layer]; output_mode=:spikes)
result = replay(model, [0.5, 0.5], (2, 1); trace=true, binary_inputs=false)
continued = replay(model, [0.0], (1, 1); initial_state=result.final_state)

The shape tuple is (steps, batch). Bias applies every timestep. The last stage can instead use output_mode=:linear for cumulative signed integration. max_working_bytes defaults to 256 MiB; admission uses the same numeric-buffer formula as Python. Domain/shape refusals raise ArgumentError, finite arithmetic overflow raises OverflowError, and budget refusal raises OutOfMemoryError. Parameters and returned arrays are independently owned; public coefficient mutations are checked again at replay. Direct Julia calls do not automatically select a Python backend.

tests/test_accel_julia_if_replay.py compares every output, state and event bit against Python for both installed Julia runtime families, checks actual chunked continuation and requires inferred concrete replay returns. Julia line coverage and Python branch coverage describe different measurements.

Go explicit-frame replay

The github.com/anulum/sc-neurocore/accel/conversion package exposes owned DenseLayer, ConvertedSNN, ReplayOptions and ReplayResult values. Constructors require a numeric byte budget, positive finite thresholds and finite per-layer preloads. Coefficients are private and copied independently. Replay takes flat row-major frames and explicit step/batch counts; Classify returns zero-based first-maximum labels. The resource formula and response semantics match Python.

Go
layer, err := conversion.NewDenseLayer(1, 1, []float64{1}, []float64{0.25}, 1, 0.5, 256<<20)
if err != nil { return err }
model, err := conversion.NewConvertedSNN([]conversion.DenseLayer{layer}, conversion.Spikes, 256<<20)
if err != nil { return err }
result, err := model.Replay([]float64{0.5, 0.5}, 2, 1, conversion.ReplayOptions{
    Trace: true, BinaryInputs: false, MaxWorkingBytes: 256 << 20,
})
if err != nil { return err }
continued, err := model.Replay([]float64{0}, 1, 1, conversion.ReplayOptions{
    InitialState: result.FinalState, MaxWorkingBytes: 256 << 20,
})
if err != nil { return err }
fmt.Println(continued.Output)

Use conversion.Linear for cumulative signed readout. Refusals return nil results and ErrInvalidInput, ErrOverflow or ErrResourceLimit. Separate explicit float64 rounding follows the Go floating-point specification, so multiply/add fusion cannot discard the profile's required intermediate rounding. Concurrent replays use independent states and traces. NewConvertedSNNFromParameters admits the whole borrowed coefficient stack before copying each layer once, without retaining caller storage. The C ABI uses this constructor to avoid a redundant coefficient snapshot.

Build the Go ownership boundary from the source checkout:

Bash
cd src/sc_neurocore/accel/go
go build -buildmode=c-shared -o /your/build/directory/if-go.so ./conversion/cshared
export SC_NEUROCORE_IF_GO_LIB=/your/build/directory/if-go.so

ConvertedSNN.replay, run, rates, and classify accept backend="go". The Go owner retains its result through a runtime/cgo.Handle stored in an opaque C-allocated metadata slot. Numeric result vectors are pinned with runtime.Pinner while any Python view retains the owner. The final release unpins the vectors, deletes the handle and frees the metadata slot. No complete trace copy crosses the boundary; empty vectors use a null data pointer with zero length. The unsafe C caller contract still requires live handles, exactly-once free, and unchanged live input arrays throughout a call. The Python adapter manages these lifetimes.

tests/test_accel_go_if_abi.py builds with GOEXPERIMENT=cgocheck2 and exercises complete C buffers and refusal-slot custody. tests/test_conversion_go_native.py exercises all public Python methods, continuation, retained array views and concurrent results. Its native race profile runs the same corpus in a separate Linux x86-64 loader process with LD_PREFER_MAP_32BIT_EXEC=1; the host-default instrumented library failed ThreadSanitizer shadow allocation. Both ordinary and race-instrumented profiles retain strict cgo pointer checks.

tests/test_accel_go_if_replay.py compares full output/state/event bits and actual chunked continuation through exported Go methods. Native package tests exercise admission, ownership, signed classification and concurrent replay with the race detector. Go statement coverage is distinct from Python branch coverage.

Python Mojo replay

Build the shared ownership boundary using Mojo 1.0.0 with contraction disabled:

Bash
mojo build --fp-mode contract=off --diagnose-missing-doc-strings --Werror \
  -I src/sc_neurocore/accel/mojo/kernels --emit shared-lib \
  -o /your/build/directory/if-mojo.so \
  src/sc_neurocore/accel/mojo/kernels/ann_to_snn_native.mojo
export SC_NEUROCORE_IF_MOJO_LIB=/your/build/directory/if-mojo.so

ConvertedSNN.replay, run, rates, and classify accept backend="mojo". The boundary admits all descriptor extents and the whole numeric reservation before copying caller doubles once. Consuming constructors and the shared ann_to_snn_compute.mojo kernel transfer those owners without additional coefficient/frame copies. The native allocation retains every result vector until the Python adapter releases its last view; slices retain the same owner. Empty buffers have a null address and zero length. The process runtime is initialized before replay or buffer access and retained for process lifetime.

The C request/buffer/release contract targets 64-bit Linux. Raw C callers must provide complete live aligned arrays unchanged during each call, exclusive writable destination slots, live result handles and exactly-once release without outstanding or concurrent borrows. Address checks cannot establish OS accessibility. Refusals leave destination slots unchanged. The Python adapter manages result lifetimes and maps domain, resource and arithmetic refusals to the same exceptions as the other native backends.

tests/test_accel_mojo_if_abi.py runs the shared complete C-buffer and refusal contract. tests/test_conversion_mojo_native.py exercises all public runtime methods, encoded continuation, retained slices and concurrent independent results. These acceptance tests establish behavior; native quantitative coverage and comparison benchmarks remain separate gates.

Mojo explicit-frame replay

ann_to_snn_parameters.mojo provides owned DenseLayer and ConvertedSNN values; ann_to_snn.mojo exposes ReplayResult, replay and zero-based classify. Flat vectors follow the same row-major coefficient, frame, state and trace order as Python. Compile this numerical profile with --fp-mode contract=off; the default compiler mode can fuse operations across statements.

Mojo
from ann_to_snn import replay
from ann_to_snn_parameters import DenseLayer, ConvertedSNN

var weights: List[Float64] = [1.0]
var bias: List[Float64] = [0.25]
var layers = List[DenseLayer]()
layers.append(DenseLayer(1, 1, weights, bias, 1.0, 0.5))
var model = ConvertedSNN(layers)
var frames: List[Float64] = [0.5, 0.5]
var result = replay(model, frames, 2, 1, trace=True, binary_inputs=False)
var next_frames: List[Float64] = [0.0]
var continued = replay(model, next_frames, 1, 1,
    initial_state=result.final_state, use_initial_state=True)
print(continued.output[0])

Use ConvertedSNN(layers, linear=True) for cumulative signed output. Bias applies every timestep; IF thresholds are inclusive with single-event subtractive reset. use_initial_state=True explicitly selects supplied states, including empty arrays for zero-batch replay. Public parameter edits are snapshotted and checked again. max_working_bytes defaults to 256 MiB and uses the shared numeric-buffer reservation. Refusals raise Error with IF invalid input, IF resource limit or IF overflow; no partial result or changed caller state is returned.

tests/test_accel_mojo_if_replay.py compiles real native calls and compares all output/state/event bits with Python, including actual chunk continuation. tests/test_accel_mojo_if_admission.py exercises constructor/replay refusals, exact memory limits, overflow and independent ownership. Missing-docstring and warning diagnostics are checked separately on each production module, because checking only an importing caller does not check every imported docstring. Direct native use does not automatically change Python backend selection.

Converter

sc_neurocore.conversion.ann_to_snn

Convert actual source forward invocations to a normalized dense IF network.

The source is captured on an independent inference copy. Repeated module calls remain distinct; unused registrations do not define topology. Consecutive affine maps are composed without adding an IF nonlinearity. ReLU scales come from each actual module/function/method invocation, or defaults (unit ReLU, six for ReLU6). Each QCFS invocation keeps its learned theta and half-threshold membrane shift; ReLU stages start from rest, including in mixed networks. The final affine-only readout integrates signed current.

Source dtype rounding, finite-timestep quantization and spike timing can produce conversion loss. Matching QCFS budgets does not guarantee lossless conversion; source/target loss must be measured using the declared input encoding.

References

Diehl et al. 2015 — "Fast-classifying, high-accuracy spiking deep networks through weight and threshold balancing". Bu et al. 2022 — "Optimal ANN-SNN Conversion for High-accuracy and Ultra-low-latency Spiking Neural Networks" (ICLR).

ConvertedSNN dataclass

A dense IF stack with deterministic input encoding and replayable state.

Parameters

weights : sequence of array_like Output-by-input matrices. Constructor inputs are copied to float64. biases : sequence of array_like or None Constant per-step currents in each layer's normalized threshold units. thresholds : sequence of float Positive finite thresholds, one per layer. T : int Positive timestep budget, at most 2**53 for exact count arithmetic. initial_membrane_fraction : float Default IF membrane preload in threshold units; QCFS uses 0.5. output_scale : float Positive finite source activation units per unit of decoded rate. output_mode : {'spikes', 'linear'} IF spike-count output or integrated signed linear readout. A linear final layer has no threshold/reset events and starts at zero. max_working_bytes : int Numeric buffer budget for constructor coefficient snapshots. Runtime calls accept their own budget; each defaults to 256 MiB.

sequence of float, optional

Per-layer preloads for mixed activation routes. None uses the global fraction; the final linear integrator always starts at zero.

Notes

Public coefficient arrays are owned by this object. Replays snapshot and validate them again, so caller edits cannot bypass shape/domain admission. The dense-if-f64-sequential-v1 profile orders input-column reductions and separates multiplication/addition instead of using BLAS reductions.

Source code in src/sc_neurocore/conversion/converted_snn.py
Python
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
@dataclass(init=False)
class ConvertedSNN:
    """A dense IF stack with deterministic input encoding and replayable state.

    Parameters
    ----------
    weights : sequence of array_like
        Output-by-input matrices. Constructor inputs are copied to float64.
    biases : sequence of array_like or None
        Constant per-step currents in each layer's normalized threshold units.
    thresholds : sequence of float
        Positive finite thresholds, one per layer.
    T : int
        Positive timestep budget, at most ``2**53`` for exact count arithmetic.
    initial_membrane_fraction : float
        Default IF membrane preload in threshold units; QCFS uses ``0.5``.
    output_scale : float
        Positive finite source activation units per unit of decoded rate.
    output_mode : {'spikes', 'linear'}
        IF spike-count output or integrated signed linear readout. A linear
        final layer has no threshold/reset events and starts at zero.
    max_working_bytes : int
        Numeric buffer budget for constructor coefficient snapshots. Runtime
        calls accept their own budget; each defaults to 256 MiB.

    layer_membrane_fractions : sequence of float, optional
        Per-layer preloads for mixed activation routes. None uses the global
        fraction; the final linear integrator always starts at zero.

    Notes
    -----
    Public coefficient arrays are owned by this object. Replays snapshot and
    validate them again, so caller edits cannot bypass shape/domain admission.
    The ``dense-if-f64-sequential-v1`` profile orders input-column reductions
    and separates multiplication/addition instead of using BLAS reductions.
    """

    weights: list[FloatArray]
    biases: list[FloatArray | None]
    thresholds: list[float]
    T: int
    initial_membrane_fraction: float
    output_scale: float
    output_mode: OutputMode
    layer_membrane_fractions: list[float] | None

    def __init__(
        self,
        weights: Sequence[npt.ArrayLike],
        biases: Sequence[npt.ArrayLike | None],
        thresholds: Sequence[float],
        T: int,
        initial_membrane_fraction: float = 0.0,
        output_scale: float = 1.0,
        output_mode: OutputMode = "spikes",
        *,
        max_working_bytes: int = DEFAULT_WORKING_BYTES,
        layer_membrane_fractions: Sequence[float] | None = None,
    ) -> None:
        """Copy and validate all coupled parameters before exposing the network."""
        if type(T) is not int or not 0 < T <= 2**53:
            raise ValueError("T must be a positive integer no greater than 2**53")
        output_scale = float(output_scale)
        if not np.isfinite(output_scale) or output_scale <= 0:
            raise ValueError("output scale must be finite and positive")
        parameters = parameter_snapshot(
            weights,
            biases,
            thresholds,
            initial_membrane_fraction,
            output_mode,
            max_working_bytes=max_working_bytes,
            layer_membrane_fractions=layer_membrane_fractions,
        )
        self.weights = [weight.copy() for weight in parameters.weights]
        self.biases = [None if bias is None else bias.copy() for bias in parameters.biases]
        self.thresholds = list(parameters.thresholds)
        self.T = T
        self.initial_membrane_fraction = parameters.initial_membrane_fraction
        self.output_scale = float(output_scale)
        self.output_mode = output_mode
        self.layer_membrane_fractions = (
            None if layer_membrane_fractions is None else list(parameters.layer_membrane_fractions)
        )

    @property
    def n_layers(self) -> int:
        """Return the current number of connected weighted layers."""
        return len(self.weights)

    def _parameters(self, max_working_bytes: int) -> IFParameters:
        """Freeze the current public coefficients for one complete operation."""
        return parameter_snapshot(
            self.weights,
            self.biases,
            self.thresholds,
            self.initial_membrane_fraction,
            self.output_mode,
            max_working_bytes=max_working_bytes,
            layer_membrane_fractions=self.layer_membrane_fractions,
        )

    def replay(
        self,
        inputs: npt.ArrayLike,
        *,
        initial_state: Sequence[npt.ArrayLike] | None = None,
        trace: bool = False,
        binary_inputs: bool = True,
        max_working_bytes: int = DEFAULT_WORKING_BYTES,
        backend: ReplayBackend = "auto",
    ) -> IFReplayResult:
        """Replay explicit frames, optionally continuing independently owned states.

        Parameters
        ----------
        inputs : array_like
            ``(steps, batch, input_neurons)`` frames in ``[0, 1]``.
        initial_state : sequence of array_like, optional
            ``(batch, output_neurons)`` state per layer, copied before execution.
        trace : bool
            Capture every post-step state and every IF spike event.
        binary_inputs : bool
            Require exact zero/one events; False admits bounded current drive.

        max_working_bytes : int
            Positive numeric buffer budget; excludes caller storage and runtime overhead.
        backend : {"auto", "numpy", "rust", "go", "mojo", "julia"}
            Explicit replay runtime; auto orders configured Rust, Go, Mojo and Julia runtimes
            by a validated measured comparison, else statically, before the NumPy floor.

        Returns
        -------
        IFReplayResult
            Incremental spike counts or the cumulative signed linear readout,
            final states and requested traces, all in independently owned buffers.
        """
        return replay_backend(
            self._parameters(max_working_bytes),
            inputs,
            initial_state=initial_state,
            trace=trace,
            binary_inputs=binary_inputs,
            max_working_bytes=max_working_bytes,
            backend=backend,
        )

    def run(
        self,
        x: npt.ArrayLike,
        *,
        input_mode: Literal["poisson", "constant"] = "poisson",
        seed: int = 42,
        max_working_bytes: int = DEFAULT_WORKING_BYTES,
        backend: ReplayBackend = "auto",
    ) -> FloatArray:
        """Encode and simulate one input vector or a batch for the stored budget.

        Parameters
        ----------
        x : array_like
            ``(input_neurons,)`` or ``(batch, input_neurons)`` values in ``[0, 1]``.
        input_mode : {'poisson', 'constant'}
            NumPy MT19937 Bernoulli event encoding or direct constant current.
            The legacy default uses seed 42 and the same row-major uniform stream.
        seed : int
            Unsigned 32-bit MT19937 seed used by Poisson event encoding.

        max_working_bytes : int
            Positive numeric buffer budget; excludes caller storage and runtime overhead.
        backend : {"auto", "numpy", "rust", "go", "mojo", "julia"}
            Explicit replay runtime; auto orders configured Rust, Go, Mojo and Julia runtimes
            by a validated measured comparison, else statically, before the NumPy floor.

        Returns
        -------
        ndarray
            Accumulated final-layer response, retaining the input batch shape.
            Spiking output is a count; linear output is an integrated current.

        Raises
        ------
        ValueError
            If encoding, seed, input geometry or domains are invalid.
        """
        return simulate_encoded(
            self._parameters(max_working_bytes),
            x,
            self.T,
            input_mode,
            seed,
            max_working_bytes,
            backend,
        )

    def rates(
        self,
        x: npt.ArrayLike,
        *,
        input_mode: Literal["poisson", "constant"] = "poisson",
        seed: int = 42,
        max_working_bytes: int = DEFAULT_WORKING_BYTES,
        backend: ReplayBackend = "auto",
    ) -> FloatArray:
        """Decode accumulated responses into the source activation's units.

        Parameters
        ----------
        x : array_like
            Input vector or batch in ``[0, 1]``.
        input_mode : {'poisson', 'constant'}
            Declared input encoding passed to :meth:`run`.
        seed : int
            MT19937 seed for Poisson input encoding.

        max_working_bytes : int
            Positive numeric buffer budget; excludes caller storage and runtime overhead.
        backend : {"auto", "numpy", "rust", "go", "mojo", "julia"}
            Explicit replay runtime; auto orders configured Rust, Go, Mojo and Julia runtimes
            by a validated measured comparison, else statically, before the NumPy floor.

        Returns
        -------
        ndarray
            Mean output response rescaled to the source ANN's activation units.
        """
        if not np.isfinite(self.output_scale) or self.output_scale <= 0:
            raise ValueError("output scale must be finite and positive")
        return (
            self.run(
                x,
                input_mode=input_mode,
                seed=seed,
                max_working_bytes=max_working_bytes,
                backend=backend,
            )
            / self.T
            * self.output_scale
        )

    def classify(
        self,
        x: npt.ArrayLike,
        *,
        max_working_bytes: int = DEFAULT_WORKING_BYTES,
        backend: ReplayBackend = "auto",
    ) -> npt.NDArray[np.intp] | np.intp:
        """Return the first maximal response index for a vector or each batch row.

        Parameters
        ----------
        x : array_like
            Input vector or batch in ``[0, 1]``.

        max_working_bytes : int
            Positive numeric buffer budget; excludes caller storage and runtime overhead.
        backend : {"auto", "numpy", "rust", "go", "mojo", "julia"}
            Explicit replay runtime; auto orders configured Rust, Go, Mojo and Julia runtimes
            by a validated measured comparison, else statically, before the NumPy floor.

        Returns
        -------
        integer or ndarray
            Index of the greatest accumulated output; ties select the first.
        """
        predictions: npt.NDArray[np.intp] | np.intp = np.argmax(
            self.run(x, max_working_bytes=max_working_bytes, backend=backend), axis=-1
        )
        return predictions

n_layers property

Return the current number of connected weighted layers.

__init__(weights, biases, thresholds, T, initial_membrane_fraction=0.0, output_scale=1.0, output_mode='spikes', *, max_working_bytes=DEFAULT_WORKING_BYTES, layer_membrane_fractions=None)

Copy and validate all coupled parameters before exposing the network.

Source code in src/sc_neurocore/conversion/converted_snn.py
Python
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
def __init__(
    self,
    weights: Sequence[npt.ArrayLike],
    biases: Sequence[npt.ArrayLike | None],
    thresholds: Sequence[float],
    T: int,
    initial_membrane_fraction: float = 0.0,
    output_scale: float = 1.0,
    output_mode: OutputMode = "spikes",
    *,
    max_working_bytes: int = DEFAULT_WORKING_BYTES,
    layer_membrane_fractions: Sequence[float] | None = None,
) -> None:
    """Copy and validate all coupled parameters before exposing the network."""
    if type(T) is not int or not 0 < T <= 2**53:
        raise ValueError("T must be a positive integer no greater than 2**53")
    output_scale = float(output_scale)
    if not np.isfinite(output_scale) or output_scale <= 0:
        raise ValueError("output scale must be finite and positive")
    parameters = parameter_snapshot(
        weights,
        biases,
        thresholds,
        initial_membrane_fraction,
        output_mode,
        max_working_bytes=max_working_bytes,
        layer_membrane_fractions=layer_membrane_fractions,
    )
    self.weights = [weight.copy() for weight in parameters.weights]
    self.biases = [None if bias is None else bias.copy() for bias in parameters.biases]
    self.thresholds = list(parameters.thresholds)
    self.T = T
    self.initial_membrane_fraction = parameters.initial_membrane_fraction
    self.output_scale = float(output_scale)
    self.output_mode = output_mode
    self.layer_membrane_fractions = (
        None if layer_membrane_fractions is None else list(parameters.layer_membrane_fractions)
    )

replay(inputs, *, initial_state=None, trace=False, binary_inputs=True, max_working_bytes=DEFAULT_WORKING_BYTES, backend='auto')

Replay explicit frames, optionally continuing independently owned states.

Parameters

inputs : array_like (steps, batch, input_neurons) frames in [0, 1]. initial_state : sequence of array_like, optional (batch, output_neurons) state per layer, copied before execution. trace : bool Capture every post-step state and every IF spike event. binary_inputs : bool Require exact zero/one events; False admits bounded current drive.

int

Positive numeric buffer budget; excludes caller storage and runtime overhead.

backend : {"auto", "numpy", "rust", "go", "mojo", "julia"} Explicit replay runtime; auto orders configured Rust, Go, Mojo and Julia runtimes by a validated measured comparison, else statically, before the NumPy floor.

Returns

IFReplayResult Incremental spike counts or the cumulative signed linear readout, final states and requested traces, all in independently owned buffers.

Source code in src/sc_neurocore/conversion/converted_snn.py
Python
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
def replay(
    self,
    inputs: npt.ArrayLike,
    *,
    initial_state: Sequence[npt.ArrayLike] | None = None,
    trace: bool = False,
    binary_inputs: bool = True,
    max_working_bytes: int = DEFAULT_WORKING_BYTES,
    backend: ReplayBackend = "auto",
) -> IFReplayResult:
    """Replay explicit frames, optionally continuing independently owned states.

    Parameters
    ----------
    inputs : array_like
        ``(steps, batch, input_neurons)`` frames in ``[0, 1]``.
    initial_state : sequence of array_like, optional
        ``(batch, output_neurons)`` state per layer, copied before execution.
    trace : bool
        Capture every post-step state and every IF spike event.
    binary_inputs : bool
        Require exact zero/one events; False admits bounded current drive.

    max_working_bytes : int
        Positive numeric buffer budget; excludes caller storage and runtime overhead.
    backend : {"auto", "numpy", "rust", "go", "mojo", "julia"}
        Explicit replay runtime; auto orders configured Rust, Go, Mojo and Julia runtimes
        by a validated measured comparison, else statically, before the NumPy floor.

    Returns
    -------
    IFReplayResult
        Incremental spike counts or the cumulative signed linear readout,
        final states and requested traces, all in independently owned buffers.
    """
    return replay_backend(
        self._parameters(max_working_bytes),
        inputs,
        initial_state=initial_state,
        trace=trace,
        binary_inputs=binary_inputs,
        max_working_bytes=max_working_bytes,
        backend=backend,
    )

run(x, *, input_mode='poisson', seed=42, max_working_bytes=DEFAULT_WORKING_BYTES, backend='auto')

Encode and simulate one input vector or a batch for the stored budget.

Parameters

x : array_like (input_neurons,) or (batch, input_neurons) values in [0, 1]. input_mode : {'poisson', 'constant'} NumPy MT19937 Bernoulli event encoding or direct constant current. The legacy default uses seed 42 and the same row-major uniform stream. seed : int Unsigned 32-bit MT19937 seed used by Poisson event encoding.

int

Positive numeric buffer budget; excludes caller storage and runtime overhead.

backend : {"auto", "numpy", "rust", "go", "mojo", "julia"} Explicit replay runtime; auto orders configured Rust, Go, Mojo and Julia runtimes by a validated measured comparison, else statically, before the NumPy floor.

Returns

ndarray Accumulated final-layer response, retaining the input batch shape. Spiking output is a count; linear output is an integrated current.

Raises

ValueError If encoding, seed, input geometry or domains are invalid.

Source code in src/sc_neurocore/conversion/converted_snn.py
Python
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
def run(
    self,
    x: npt.ArrayLike,
    *,
    input_mode: Literal["poisson", "constant"] = "poisson",
    seed: int = 42,
    max_working_bytes: int = DEFAULT_WORKING_BYTES,
    backend: ReplayBackend = "auto",
) -> FloatArray:
    """Encode and simulate one input vector or a batch for the stored budget.

    Parameters
    ----------
    x : array_like
        ``(input_neurons,)`` or ``(batch, input_neurons)`` values in ``[0, 1]``.
    input_mode : {'poisson', 'constant'}
        NumPy MT19937 Bernoulli event encoding or direct constant current.
        The legacy default uses seed 42 and the same row-major uniform stream.
    seed : int
        Unsigned 32-bit MT19937 seed used by Poisson event encoding.

    max_working_bytes : int
        Positive numeric buffer budget; excludes caller storage and runtime overhead.
    backend : {"auto", "numpy", "rust", "go", "mojo", "julia"}
        Explicit replay runtime; auto orders configured Rust, Go, Mojo and Julia runtimes
        by a validated measured comparison, else statically, before the NumPy floor.

    Returns
    -------
    ndarray
        Accumulated final-layer response, retaining the input batch shape.
        Spiking output is a count; linear output is an integrated current.

    Raises
    ------
    ValueError
        If encoding, seed, input geometry or domains are invalid.
    """
    return simulate_encoded(
        self._parameters(max_working_bytes),
        x,
        self.T,
        input_mode,
        seed,
        max_working_bytes,
        backend,
    )

rates(x, *, input_mode='poisson', seed=42, max_working_bytes=DEFAULT_WORKING_BYTES, backend='auto')

Decode accumulated responses into the source activation's units.

Parameters

x : array_like Input vector or batch in [0, 1]. input_mode : {'poisson', 'constant'} Declared input encoding passed to :meth:run. seed : int MT19937 seed for Poisson input encoding.

int

Positive numeric buffer budget; excludes caller storage and runtime overhead.

backend : {"auto", "numpy", "rust", "go", "mojo", "julia"} Explicit replay runtime; auto orders configured Rust, Go, Mojo and Julia runtimes by a validated measured comparison, else statically, before the NumPy floor.

Returns

ndarray Mean output response rescaled to the source ANN's activation units.

Source code in src/sc_neurocore/conversion/converted_snn.py
Python
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
def rates(
    self,
    x: npt.ArrayLike,
    *,
    input_mode: Literal["poisson", "constant"] = "poisson",
    seed: int = 42,
    max_working_bytes: int = DEFAULT_WORKING_BYTES,
    backend: ReplayBackend = "auto",
) -> FloatArray:
    """Decode accumulated responses into the source activation's units.

    Parameters
    ----------
    x : array_like
        Input vector or batch in ``[0, 1]``.
    input_mode : {'poisson', 'constant'}
        Declared input encoding passed to :meth:`run`.
    seed : int
        MT19937 seed for Poisson input encoding.

    max_working_bytes : int
        Positive numeric buffer budget; excludes caller storage and runtime overhead.
    backend : {"auto", "numpy", "rust", "go", "mojo", "julia"}
        Explicit replay runtime; auto orders configured Rust, Go, Mojo and Julia runtimes
        by a validated measured comparison, else statically, before the NumPy floor.

    Returns
    -------
    ndarray
        Mean output response rescaled to the source ANN's activation units.
    """
    if not np.isfinite(self.output_scale) or self.output_scale <= 0:
        raise ValueError("output scale must be finite and positive")
    return (
        self.run(
            x,
            input_mode=input_mode,
            seed=seed,
            max_working_bytes=max_working_bytes,
            backend=backend,
        )
        / self.T
        * self.output_scale
    )

classify(x, *, max_working_bytes=DEFAULT_WORKING_BYTES, backend='auto')

Return the first maximal response index for a vector or each batch row.

Parameters

x : array_like Input vector or batch in [0, 1].

int

Positive numeric buffer budget; excludes caller storage and runtime overhead.

backend : {"auto", "numpy", "rust", "go", "mojo", "julia"} Explicit replay runtime; auto orders configured Rust, Go, Mojo and Julia runtimes by a validated measured comparison, else statically, before the NumPy floor.

Returns

integer or ndarray Index of the greatest accumulated output; ties select the first.

Source code in src/sc_neurocore/conversion/converted_snn.py
Python
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
def classify(
    self,
    x: npt.ArrayLike,
    *,
    max_working_bytes: int = DEFAULT_WORKING_BYTES,
    backend: ReplayBackend = "auto",
) -> npt.NDArray[np.intp] | np.intp:
    """Return the first maximal response index for a vector or each batch row.

    Parameters
    ----------
    x : array_like
        Input vector or batch in ``[0, 1]``.

    max_working_bytes : int
        Positive numeric buffer budget; excludes caller storage and runtime overhead.
    backend : {"auto", "numpy", "rust", "go", "mojo", "julia"}
        Explicit replay runtime; auto orders configured Rust, Go, Mojo and Julia runtimes
        by a validated measured comparison, else statically, before the NumPy floor.

    Returns
    -------
    integer or ndarray
        Index of the greatest accumulated output; ties select the first.
    """
    predictions: npt.NDArray[np.intp] | np.intp = np.argmax(
        self.run(x, max_working_bytes=max_working_bytes, backend=backend), axis=-1
    )
    return predictions

convert(model, calibration_data=None, T=None, percentile=99.9, *, max_working_bytes=DEFAULT_WORKING_BYTES)

Convert a trained PyTorch ANN to a rate-coded SNN.

Activations are associated with their actual preceding source operations. ReLU and QCFS stages retain separate calibration scales and preloads in mixed networks. Unsupported dense-target operators fail compatibility admission.

Parameters

model : nn.Module Trained PyTorch model representable by a single-input dense forward path with Linear, ReLU/ReLU6 and QCFS operations. Actual invocations, including shared modules and functional ReLUs, determine the exported topology. calibration_data : Tensor, optional Sample source-format input for every ReLU invocation, including functional activations in mixed ReLU/QCFS networks. None uses unit ReLU scales; QCFS invocations always retain their learned thresholds. T : int, optional Number of simulation timesteps (higher = more accurate, slower). If None, the QCFS route adopts the layers' trained step budget and the ReLU route defaults to 16. percentile : float Activation percentile for threshold normalization on the ReLU route.

int

Numeric storage budget for source copying and exported snapshots.

Returns

ConvertedSNN Converted spiking network ready to run.

Source code in src/sc_neurocore/conversion/ann_to_snn.py
Python
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
def convert(
    model: object,
    calibration_data: object = None,
    T: int | None = None,
    percentile: float = 99.9,
    *,
    max_working_bytes: int = DEFAULT_WORKING_BYTES,
) -> ConvertedSNN:
    """Convert a trained PyTorch ANN to a rate-coded SNN.

    Activations are associated with their actual preceding source operations.
    ReLU and QCFS stages retain separate calibration scales and preloads in mixed
    networks. Unsupported dense-target operators fail compatibility admission.

    Parameters
    ----------
    model : nn.Module
        Trained PyTorch model representable by a single-input dense forward path
        with Linear, ReLU/ReLU6 and QCFS operations. Actual invocations, including
        shared modules and functional ReLUs, determine the exported topology.
    calibration_data : Tensor, optional
        Sample source-format input for every ReLU invocation, including functional
        activations in mixed ReLU/QCFS networks. None uses unit ReLU scales;
        QCFS invocations always retain their learned thresholds.
    T : int, optional
        Number of simulation timesteps (higher = more accurate, slower). If
        None, the QCFS route adopts the layers' trained step budget and the
        ReLU route defaults to 16.
    percentile : float
        Activation percentile for threshold normalization on the ReLU route.

    max_working_bytes : int
        Numeric storage budget for source copying and exported snapshots.

    Returns
    -------
    ConvertedSNN
        Converted spiking network ready to run.
    """
    if not HAS_TORCH:
        raise ImportError("PyTorch required for ANN-to-SNN conversion")

    if not isinstance(model, nn.Module):
        raise TypeError("model must be a PyTorch Module")
    plan = compile_source_graph(model, max_working_bytes)
    layers = _extract_layers(model, plan)
    qcfs_layers = _extract_qcfs_layers(model, plan)
    calibrated: dict[str, float] = {}
    has_relu = any(layer.activation.kind == "relu" for layer in plan.layers)
    if has_relu and calibration_data is not None:
        if not isinstance(calibration_data, torch.Tensor):
            raise TypeError("calibration_data must be a PyTorch Tensor")
        calibrated = calibrate_source_nodes(plan, calibration_data, percentile)
    thresholds = [
        layer.activation.theta
        if layer.activation.kind == "qcfs"
        else calibrated.get(layer.activation.node_name, layer.activation.theta)
        for layer in plan.layers
    ]
    if T is None:
        budgets = {budget for _, budget in qcfs_layers}
        if len(budgets) > 1:
            raise ValueError("mixed QCFS timestep budgets require an explicit T")
        T = qcfs_layers[0][1] if qcfs_layers else 16
    normalized_weights: list[FloatArray] = []
    normalized_biases: list[FloatArray | None] = []
    previous_scale = 1.0
    for (weight, bias), threshold in zip(layers, thresholds):
        normalized_weights.append(weight * previous_scale / threshold)
        normalized_biases.append(None if bias is None else bias / threshold)
        previous_scale = threshold
    output_mode: OutputMode = "linear" if plan.layers[-1].activation.kind == "linear" else "spikes"
    return ConvertedSNN(
        weights=normalized_weights,
        biases=normalized_biases,
        thresholds=[1.0] * len(layers),
        T=T,
        initial_membrane_fraction=0.5 if qcfs_layers else 0.0,
        output_scale=thresholds[-1],
        output_mode=output_mode,
        layer_membrane_fractions=[
            0.5 if layer.activation.kind == "qcfs" else 0.0 for layer in plan.layers
        ],
        max_working_bytes=max_working_bytes,
    )

replace_relu_with_qcfs(model, T=8, theta=1.0, learn_theta=True)

Swap every ReLU/ReLU6 in a model for a QCFS activation, in place.

This prepares a trained or fresh ANN for conversion-aware fine-tuning: after substitution the network is retrained for a few epochs so the QCFS thresholds settle. Conversion loss is then measured against the source ANN; QCFS does not guarantee lossless conversion for arbitrary spike timing.

Parameters

model : nn.Module Model whose registered ReLU/ReLU6 children are replaced in place. All aliases of one original activation retain one shared QCFS module and learned threshold. The container and activation training modes are preserved; the root module itself is returned unchanged. T : int Quantisation step budget for each inserted QCFS layer. theta : float Initial firing threshold for each inserted QCFS layer. learn_theta : bool Whether each inserted threshold is a trainable parameter (the QCFS fine-tuning default).

Returns

nn.Module The same model instance, returned for chaining.

Source code in src/sc_neurocore/conversion/ann_to_snn.py
Python
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
def replace_relu_with_qcfs(
    model: object,
    T: int = 8,
    theta: float = 1.0,
    learn_theta: bool = True,
) -> nn.Module:
    """Swap every ReLU/ReLU6 in a model for a QCFS activation, in place.

    This prepares a trained or fresh ANN for conversion-aware fine-tuning:
    after substitution the network is retrained for a few epochs so the QCFS
    thresholds settle. Conversion loss is then measured against the source ANN;
    QCFS does not guarantee lossless conversion for arbitrary spike timing.

    Parameters
    ----------
    model : nn.Module
        Model whose registered ReLU/ReLU6 children are replaced in place.
        All aliases of one original activation retain one shared QCFS module
        and learned threshold. The container and activation training modes
        are preserved; the root module itself is returned unchanged.
    T : int
        Quantisation step budget for each inserted QCFS layer.
    theta : float
        Initial firing threshold for each inserted QCFS layer.
    learn_theta : bool
        Whether each inserted threshold is a trainable parameter (the QCFS
        fine-tuning default).

    Returns
    -------
    nn.Module
        The same ``model`` instance, returned for chaining.
    """
    if not HAS_TORCH:
        raise ImportError("PyTorch required for ANN-to-SNN conversion")

    if not isinstance(model, nn.Module):
        raise TypeError("model must be a PyTorch Module")

    replacements: dict[int, QCFSActivation] = {}
    # Snapshot unique parents; named_children omits duplicate registrations.
    for parent in tuple(model.modules()):
        for name, child in tuple(parent._modules.items()):
            if isinstance(child, (nn.ReLU, nn.ReLU6)):
                identity = id(child)
                if identity not in replacements:
                    replacements[identity] = QCFSActivation(
                        T=T, theta=theta, learn_theta=learn_theta
                    ).train(child.training)
                setattr(parent, name, replacements[identity])
    return model

QCFS Activation

sc_neurocore.conversion.qcfs

QCFS (Quantization-Clip-Floor-Shift) activation function.

Replaces ReLU in the ANN during conversion-aware training or post-hoc conversion. QCFS approximates the rate-coded SNN firing rate as a quantized step function, minimizing conversion error.

Reference: Bu et al. 2022 — "Optimal ANN-SNN Conversion for High-accuracy and Ultra-low-latency Spiking Neural Networks"

QCFSActivation

Bases: Module

QCFS activation: quantized clip-floor-shift ReLU replacement.

For T timesteps and threshold theta

QCFS(x) = clip(floor(x * T / theta + 0.5), 0, T) * theta / T

This quantizes activations to T+1 levels in [0, theta], matching the achievable spike rates of an IF neuron over T timesteps.

Parameters

T : int Number of simulation timesteps, 1 <= T <= 2**32 - 1: the step domain every native counterpart shares. theta : float Firing threshold (default 1.0). learn_theta : bool Make threshold trainable (default False).

Source code in src/sc_neurocore/conversion/qcfs.py
Python
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
class QCFSActivation(nn.Module):
    """QCFS activation: quantized clip-floor-shift ReLU replacement.

    For T timesteps and threshold theta:
        QCFS(x) = clip(floor(x * T / theta + 0.5), 0, T) * theta / T

    This quantizes activations to T+1 levels in [0, theta], matching
    the achievable spike rates of an IF neuron over T timesteps.

    Parameters
    ----------
    T : int
        Number of simulation timesteps, ``1 <= T <= 2**32 - 1``: the step
        domain every native counterpart shares.
    theta : float
        Firing threshold (default 1.0).
    learn_theta : bool
        Make threshold trainable (default False).
    """

    def __init__(self, T: int = 8, theta: float = 1.0, learn_theta: bool = False) -> None:
        """Create a positive finite threshold and positive integer rate grid."""
        super().__init__()
        if type(T) is not int or not 1 <= T <= QCFS_STEP_LIMIT:
            raise ValueError("QCFS T must be a positive integer no larger than 2**32 - 1")
        if not math.isfinite(theta) or theta <= 0:
            raise ValueError("QCFS theta must be finite and positive")
        self.T = T
        if learn_theta:
            self.theta = nn.Parameter(torch.tensor(float(theta)))
        else:
            self.register_buffer("theta", torch.tensor(float(theta)))

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        """Quantise activations to the spike-rate grid with a straight-through gradient.

        Parameters
        ----------
        x : torch.Tensor
            ANN activation tensor to clip and quantise into ``T + 1`` rate levels.

        Returns
        -------
        torch.Tensor
            Tensor with values clipped to ``[0, theta]`` and quantised to the
            finite-timestep spike-rate lattice. The backward pass uses the
            identity floor surrogate on the continuously clipped shifted input;
            The two clipping endpoints use zero input derivative. Saturated values
            have zero input derivative, and upper saturation differentiates to
            one with respect to the learned threshold.
        """
        if not bool(torch.isfinite(self.theta) & (self.theta > 0)):
            raise ValueError("QCFS theta must be finite and positive")
        scaled = (x * self.T / self.theta + 0.5).detach()
        interior = (scaled > 0) & (scaled < self.T)
        # Only interior elements carry a gradient. Rebuilding their coordinate
        # from the element alone keeps a saturated infinite or NaN neighbour
        # out of the threshold derivative, where a zero upstream times its
        # infinite quotient would otherwise turn the whole batch into NaN.
        carrier = torch.where(interior, x, torch.zeros_like(x)) * self.T / self.theta + 0.5
        clipped = torch.where(interior, carrier, scaled).clamp(0, self.T)
        # Bu et al., Eq. 17: floor has an identity surrogate derivative.
        quantized = clipped + (clipped.floor() - clipped).detach()
        out: torch.Tensor = quantized * self.theta / self.T
        return out

    def extra_repr(self) -> str:
        """Return the compact PyTorch module representation."""
        return f"T={self.T}, theta={self.theta.item():.2f}"

__init__(T=8, theta=1.0, learn_theta=False)

Create a positive finite threshold and positive integer rate grid.

Source code in src/sc_neurocore/conversion/qcfs.py
Python
49
50
51
52
53
54
55
56
57
58
59
60
def __init__(self, T: int = 8, theta: float = 1.0, learn_theta: bool = False) -> None:
    """Create a positive finite threshold and positive integer rate grid."""
    super().__init__()
    if type(T) is not int or not 1 <= T <= QCFS_STEP_LIMIT:
        raise ValueError("QCFS T must be a positive integer no larger than 2**32 - 1")
    if not math.isfinite(theta) or theta <= 0:
        raise ValueError("QCFS theta must be finite and positive")
    self.T = T
    if learn_theta:
        self.theta = nn.Parameter(torch.tensor(float(theta)))
    else:
        self.register_buffer("theta", torch.tensor(float(theta)))

forward(x)

Quantise activations to the spike-rate grid with a straight-through gradient.

Parameters

x : torch.Tensor ANN activation tensor to clip and quantise into T + 1 rate levels.

Returns

torch.Tensor Tensor with values clipped to [0, theta] and quantised to the finite-timestep spike-rate lattice. The backward pass uses the identity floor surrogate on the continuously clipped shifted input; The two clipping endpoints use zero input derivative. Saturated values have zero input derivative, and upper saturation differentiates to one with respect to the learned threshold.

Source code in src/sc_neurocore/conversion/qcfs.py
Python
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
def forward(self, x: torch.Tensor) -> torch.Tensor:
    """Quantise activations to the spike-rate grid with a straight-through gradient.

    Parameters
    ----------
    x : torch.Tensor
        ANN activation tensor to clip and quantise into ``T + 1`` rate levels.

    Returns
    -------
    torch.Tensor
        Tensor with values clipped to ``[0, theta]`` and quantised to the
        finite-timestep spike-rate lattice. The backward pass uses the
        identity floor surrogate on the continuously clipped shifted input;
        The two clipping endpoints use zero input derivative. Saturated values
        have zero input derivative, and upper saturation differentiates to
        one with respect to the learned threshold.
    """
    if not bool(torch.isfinite(self.theta) & (self.theta > 0)):
        raise ValueError("QCFS theta must be finite and positive")
    scaled = (x * self.T / self.theta + 0.5).detach()
    interior = (scaled > 0) & (scaled < self.T)
    # Only interior elements carry a gradient. Rebuilding their coordinate
    # from the element alone keeps a saturated infinite or NaN neighbour
    # out of the threshold derivative, where a zero upstream times its
    # infinite quotient would otherwise turn the whole batch into NaN.
    carrier = torch.where(interior, x, torch.zeros_like(x)) * self.T / self.theta + 0.5
    clipped = torch.where(interior, carrier, scaled).clamp(0, self.T)
    # Bu et al., Eq. 17: floor has an identity surrogate derivative.
    quantized = clipped + (clipped.floor() - clipped).detach()
    out: torch.Tensor = quantized * self.theta / self.T
    return out

extra_repr()

Return the compact PyTorch module representation.

Source code in src/sc_neurocore/conversion/qcfs.py
Python
95
96
97
def extra_repr(self) -> str:
    """Return the compact PyTorch module representation."""
    return f"T={self.T}, theta={self.theta.item():.2f}"

Torch-free QCFS evaluation

qcfs_forward(x, steps=8, theta=1.0, backend="auto") and qcfs_backward(x, upstream, steps=8, theta=1.0, backend="auto") evaluate the QCFS lattice and its straight-through derivatives in float64 without PyTorch. They follow the operation order of QCFSActivation and its autograd, so the values, the input derivative and each element's threshold derivative equal the float64 activation bit for bit, including NaN payloads and the sign of zero. The threshold array holds the gradient a one-element batch receives; its sum is a shared threshold's gradient up to summation order. upstream must have the shape of x; boolean, complex, object and text inputs are refused.

backend accepts "auto", "numpy", "rust", "go", "mojo" or "julia", and every runtime returns the NumPy bits. Rust, Go and Mojo load the libraries named by SC_NEUROCORE_QCFS_RUST_LIB, SC_NEUROCORE_QCFS_GO_LIB and SC_NEUROCORE_QCFS_MOJO_LIB; Julia runs in the configured JuliaCall runtime (the same PYTHON_JULIACALL_* settings as Julia replay) when SC_NEUROCORE_QCFS_JULIA_ENABLED=1. Auto tries Mojo, Rust, Go, then Julia over what is configured and otherwise uses NumPy: the order of the local comparison below, a preference rather than a guarantee. An explicit runtime without its configuration raises RuntimeError instead of falling back.

Each library exports sc_qcfs_abi_version, sc_qcfs_forward and sc_qcfs_backward over float64 arrays, validates the grid, threshold and every span before writing, and returns 0 or -1 with outputs unchanged. Build them from a source checkout, or from an installed package by replacing src/sc_neurocore with the package directory:

Bash
cargo build --offline --release \
  --manifest-path src/sc_neurocore/accel/rust/safety/qcfs_native/Cargo.toml \
  --target-dir /your/build/directory
(cd src/sc_neurocore/accel/go && \
  go build -buildvcs=false -buildmode=c-shared -o /your/qcfs-go.so ./conversion/qcfscshared)
mojo build --fp-mode contract=off -I src/sc_neurocore/accel/mojo/kernels \
  --emit shared-lib -o /your/qcfs-mojo.so src/sc_neurocore/accel/mojo/kernels/qcfs.mojo
export SC_NEUROCORE_QCFS_RUST_LIB=/your/build/directory/release/libsc_neurocore_qcfs.so

benchmarks/bench_qcfs_runtimes.py --output report.json --cpu 0 times public forward and backward calls for 1, 64, 4096 and 262144 elements in all five runtimes, each in a fresh process pinned to one CPU. Every response must equal the NumPy bits before its sample counts, and the report binds the QCFS source and library bytes; it is published atomically only after all five pass. Host load is recorded, not isolated.

sc_neurocore.conversion.qcfs_dispatch

Quantise with QCFS and evaluate its surrogate derivatives through a chosen runtime.

NumPy is always available. Rust, Go and Mojo run from owner-built libraries named by SC_NEUROCORE_QCFS_RUST_LIB, SC_NEUROCORE_QCFS_GO_LIB and SC_NEUROCORE_QCFS_MOJO_LIB; Julia runs in the configured JuliaCall runtime when SC_NEUROCORE_QCFS_JULIA_ENABLED=1. Every runtime returns the same float64 bits as NumPy and QCFSActivation.

qcfs_forward(x, steps=8, theta=1.0, *, backend='auto')

Quantise activations onto the QCFS rate lattice without PyTorch.

Parameters

x : array_like Real activations of any shape. steps : int Simulation steps, 1 <= steps <= 2**32 - 1. theta : float Finite positive firing threshold. backend : {'auto', 'numpy', 'rust', 'go', 'mojo', 'julia'} Runtime selection; see select_qcfs_native.

Returns

numpy.ndarray float64 floor(clip(x * T / theta + 0.5, 0, T)) * theta / T with the shape of x; infinities saturate and NaN stays NaN.

Raises

ValueError Step count, threshold or backend name invalid. TypeError Non-real activations. RuntimeError Selected runtime unavailable or a provider refused admitted input.

Source code in src/sc_neurocore/conversion/qcfs_dispatch.py
Python
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
def qcfs_forward(
    x: npt.ArrayLike, steps: int = 8, theta: float = 1.0, *, backend: QCFSBackend = "auto"
) -> FloatArray:
    """Quantise activations onto the QCFS rate lattice without PyTorch.

    Parameters
    ----------
    x : array_like
        Real activations of any shape.
    steps : int
        Simulation steps, ``1 <= steps <= 2**32 - 1``.
    theta : float
        Finite positive firing threshold.
    backend : {'auto', 'numpy', 'rust', 'go', 'mojo', 'julia'}
        Runtime selection; see ``select_qcfs_native``.

    Returns
    -------
    numpy.ndarray
        float64 ``floor(clip(x * T / theta + 0.5, 0, T)) * theta / T`` with the
        shape of ``x``; infinities saturate and NaN stays NaN.

    Raises
    ------
    ValueError
        Step count, threshold or backend name invalid.
    TypeError
        Non-real activations.
    RuntimeError
        Selected runtime unavailable or a provider refused admitted input.
    """
    steps, theta = checked_qcfs_parameters(steps, theta)
    values = qcfs_values(x, "x")
    api = select_qcfs_native(backend)
    if api is None:
        return reference_forward(values, steps, theta)
    output = np.empty_like(values)
    status = api.forward(steps, theta, values.ctypes.data, values.size, output.ctypes.data)
    if status != 0:
        raise RuntimeError("native QCFS provider refused admitted input")
    return output

qcfs_backward(x, upstream, steps=8, theta=1.0, *, backend='auto')

Evaluate QCFS straight-through derivatives without PyTorch.

Parameters

x, upstream : array_like Real activations and upstream gradients of one shape. steps : int Simulation steps, 1 <= steps <= 2**32 - 1. theta : float Finite positive firing threshold. backend : {'auto', 'numpy', 'rust', 'go', 'mojo', 'julia'} Runtime selection; see select_qcfs_native.

Returns

tuple of numpy.ndarray Input derivative and each element's threshold derivative, as QCFSActivation autograd yields for a one-element batch; summing the second array gives a shared threshold's gradient up to summation order.

Raises

ValueError Step count, threshold, backend name or shape mismatch. TypeError Non-real activations or gradients. RuntimeError Selected runtime unavailable or a provider refused admitted input.

Source code in src/sc_neurocore/conversion/qcfs_dispatch.py
Python
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
def qcfs_backward(
    x: npt.ArrayLike,
    upstream: npt.ArrayLike,
    steps: int = 8,
    theta: float = 1.0,
    *,
    backend: QCFSBackend = "auto",
) -> tuple[FloatArray, FloatArray]:
    """Evaluate QCFS straight-through derivatives without PyTorch.

    Parameters
    ----------
    x, upstream : array_like
        Real activations and upstream gradients of one shape.
    steps : int
        Simulation steps, ``1 <= steps <= 2**32 - 1``.
    theta : float
        Finite positive firing threshold.
    backend : {'auto', 'numpy', 'rust', 'go', 'mojo', 'julia'}
        Runtime selection; see ``select_qcfs_native``.

    Returns
    -------
    tuple of numpy.ndarray
        Input derivative and each element's threshold derivative, as
        ``QCFSActivation`` autograd yields for a one-element batch; summing the
        second array gives a shared threshold's gradient up to summation order.

    Raises
    ------
    ValueError
        Step count, threshold, backend name or shape mismatch.
    TypeError
        Non-real activations or gradients.
    RuntimeError
        Selected runtime unavailable or a provider refused admitted input.
    """
    steps, theta = checked_qcfs_parameters(steps, theta)
    values = qcfs_values(x, "x")
    gradients = qcfs_values(upstream, "upstream")
    if gradients.shape != values.shape:
        raise ValueError("QCFS upstream gradients must have the shape of x")
    api = select_qcfs_native(backend)
    if api is None:
        return reference_backward(values, gradients, steps, theta)
    inputs, thresholds = np.empty_like(values), np.empty_like(values)
    status = api.backward(
        steps,
        theta,
        values.ctypes.data,
        gradients.ctypes.data,
        values.size,
        inputs.ctypes.data,
        thresholds.ctypes.data,
    )
    if status != 0:
        raise RuntimeError("native QCFS provider refused admitted input")
    return inputs, thresholds

select_qcfs_native(backend)

Resolve the requested QCFS runtime; None selects NumPy.

Parameters

backend : {'auto', 'numpy', 'rust', 'go', 'mojo', 'julia'} Auto takes the first configured provider of QCFS_AUTO_ORDER (Mojo, Rust, Go, Julia) and otherwise NumPy. An explicit native backend requires its configuration; explicit NumPy never reads any.

Returns

QCFSNativeAPI or None Loaded provider, or None for NumPy.

Raises

ValueError Backend name is unsupported. RuntimeError Explicit configuration is absent, or a configured provider cannot load.

Source code in src/sc_neurocore/conversion/qcfs_dispatch.py
Python
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
def select_qcfs_native(backend: QCFSBackend) -> QCFSNativeAPI | None:
    """Resolve the requested QCFS runtime; None selects NumPy.

    Parameters
    ----------
    backend : {'auto', 'numpy', 'rust', 'go', 'mojo', 'julia'}
        Auto takes the first configured provider of ``QCFS_AUTO_ORDER`` (Mojo,
        Rust, Go, Julia) and otherwise NumPy. An explicit native backend
        requires its configuration; explicit NumPy never reads any.

    Returns
    -------
    QCFSNativeAPI or None
        Loaded provider, or None for NumPy.

    Raises
    ------
    ValueError
        Backend name is unsupported.
    RuntimeError
        Explicit configuration is absent, or a configured provider cannot load.
    """
    if backend not in ("auto", "numpy", "rust", "go", "mojo", "julia"):
        raise ValueError("unsupported QCFS backend")
    if backend == "numpy":
        return None
    enabled = os.environ.get("SC_NEUROCORE_QCFS_JULIA_ENABLED", "0")
    if enabled not in ("0", "1"):
        raise RuntimeError("Julia QCFS opt-in must be 0 or 1")
    if backend == "julia":
        if enabled != "1":
            raise RuntimeError("Julia QCFS backend requires SC_NEUROCORE_QCFS_JULIA_ENABLED=1")
        return load_qcfs_julia()
    if backend != "auto":
        variable = f"SC_NEUROCORE_QCFS_{backend.upper()}_LIB"
        configured = os.environ.get(variable)
        if not configured:
            raise RuntimeError(f"{backend} QCFS backend requires {variable}")
        return load_qcfs_library(configured)
    for name in QCFS_AUTO_ORDER:
        if (
            (enabled == "1")
            if name == "julia"
            else bool(os.environ.get(f"SC_NEUROCORE_QCFS_{name.upper()}_LIB"))
        ):
            return select_qcfs_native(cast(QCFSBackend, name))
    return None

Measured conversion loss

measure_conversion_loss(model, snn, inputs, labels, *, input_mode="constant", seed=42, batch_size=256, max_working_bytes=..., backend="auto") classifies the same labelled samples with a PyTorch source and the network converted from it, and returns a ConversionLossReport of what both did. Nothing is estimated:

Field Meaning
source_accuracy, converted_accuracy fraction of samples each network labels correctly
accuracy_drop source_accuracy - converted_accuracy; positive is a loss
agreement fraction of samples on which both predict the same class
rate_mean_abs_error, rate_max_abs_error decoded SNN rates against the source outputs
timesteps, input_mode, seed, batch_size how the converted network was driven
backend the replay runtime every converted batch ran on
source_sha256, converted_sha256, data_sha256 digests of the source parameters and buffers, the converted coefficients and semantics, and the inputs with their shape and labels

inputs are (samples, *source_shape) values in [0, 1]; the converted network receives each sample flattened in C order. The source runs under no_grad in inference mode on its first parameter's device and dtype, and every module's training flag is restored afterwards. Poisson batch k uses seed (seed + k) mod 2**32. backend="auto" is resolved once, by resolve_replay_backend, before the first batch, so the report names the runtime that executed rather than the one requested. to_public_dict() gives the JSON form the Studio seals.

sc_neurocore.conversion.loss_report

Measure what conversion cost on labelled data, bound to the exact models and inputs.

The source ANN and its converted SNN classify the same samples. The report states both accuracies, their agreement, the decoded-rate error against the source output, the declared input encoding and timestep budget, the replay runtime that actually executed, and digests of the source parameters, the converted coefficients and the evaluated data. Nothing is estimated: every number comes from running both networks on the given samples.

ConversionLossReport dataclass

Measured agreement between a source ANN and its converted SNN.

Attributes

schema_version : str sc-neurocore.conversion-loss-report.v1. samples : int Number of labelled samples both networks classified. classes : int Width of the output both networks produce. timesteps : int The converted network's timestep budget. input_mode : {'constant', 'poisson'} Declared input encoding of the converted network. seed : int Poisson seed of the first batch; batch k uses (seed + k) mod 2**32. batch_size : int Samples per converted-network call. backend : {'numpy', 'rust', 'go', 'mojo', 'julia'} Replay runtime every converted batch executed on. source_accuracy : float Fraction of samples the source ANN labels correctly. converted_accuracy : float Fraction of samples the converted SNN labels correctly. accuracy_drop : float source_accuracy - converted_accuracy; positive is a loss. agreement : float Fraction of samples on which both networks predict the same class. rate_mean_abs_error : float Mean absolute difference between decoded SNN rates and source outputs. rate_max_abs_error : float Largest such difference. source_sha256 : str Digest of every source parameter and buffer, by name, dtype and shape. converted_sha256 : str Digest of the converted coefficients and replay semantics. data_sha256 : str Digest of the evaluated inputs, their shape and their labels. numerical_profile : str Arithmetic profile of the converted replay.

Source code in src/sc_neurocore/conversion/loss_report.py
Python
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
@dataclass(frozen=True)
class ConversionLossReport:
    """Measured agreement between a source ANN and its converted SNN.

    Attributes
    ----------
    schema_version : str
        ``sc-neurocore.conversion-loss-report.v1``.
    samples : int
        Number of labelled samples both networks classified.
    classes : int
        Width of the output both networks produce.
    timesteps : int
        The converted network's timestep budget.
    input_mode : {'constant', 'poisson'}
        Declared input encoding of the converted network.
    seed : int
        Poisson seed of the first batch; batch ``k`` uses ``(seed + k) mod 2**32``.
    batch_size : int
        Samples per converted-network call.
    backend : {'numpy', 'rust', 'go', 'mojo', 'julia'}
        Replay runtime every converted batch executed on.
    source_accuracy : float
        Fraction of samples the source ANN labels correctly.
    converted_accuracy : float
        Fraction of samples the converted SNN labels correctly.
    accuracy_drop : float
        ``source_accuracy - converted_accuracy``; positive is a loss.
    agreement : float
        Fraction of samples on which both networks predict the same class.
    rate_mean_abs_error : float
        Mean absolute difference between decoded SNN rates and source outputs.
    rate_max_abs_error : float
        Largest such difference.
    source_sha256 : str
        Digest of every source parameter and buffer, by name, dtype and shape.
    converted_sha256 : str
        Digest of the converted coefficients and replay semantics.
    data_sha256 : str
        Digest of the evaluated inputs, their shape and their labels.
    numerical_profile : str
        Arithmetic profile of the converted replay.
    """

    schema_version: str
    samples: int
    classes: int
    timesteps: int
    input_mode: Literal["constant", "poisson"]
    seed: int
    batch_size: int
    backend: Literal["numpy", "rust", "go", "mojo", "julia"]
    source_accuracy: float
    converted_accuracy: float
    accuracy_drop: float
    agreement: float
    rate_mean_abs_error: float
    rate_max_abs_error: float
    source_sha256: str
    converted_sha256: str
    data_sha256: str
    numerical_profile: str = "dense-if-f64-sequential-v1"

    def to_public_dict(self) -> dict[str, Any]:
        """Return the report as JSON-ready fields.

        Returns
        -------
        dict
            Every attribute under its own name.
        """
        return asdict(self)

to_public_dict()

Return the report as JSON-ready fields.

Returns

dict Every attribute under its own name.

Source code in src/sc_neurocore/conversion/loss_report.py
Python
101
102
103
104
105
106
107
108
109
def to_public_dict(self) -> dict[str, Any]:
    """Return the report as JSON-ready fields.

    Returns
    -------
    dict
        Every attribute under its own name.
    """
    return asdict(self)

measure_conversion_loss(model, snn, inputs, labels, *, input_mode='constant', seed=42, batch_size=256, max_working_bytes=DEFAULT_WORKING_BYTES, backend='auto')

Classify labelled samples with a source ANN and its converted SNN and compare.

Parameters

model : torch.nn.Module Source network. It runs under no_grad in inference mode on the device and dtype of its first parameter; every module's training flag is restored afterwards. snn : ConvertedSNN The network converted from model. inputs : array_like (samples, *source_shape) values in [0, 1]. The converted network receives each sample flattened in C order. labels : array_like (samples,) integer classes in [0, classes). input_mode : {'constant', 'poisson'} Declared encoding for the converted network. seed : int Unsigned 32-bit Poisson seed of the first batch. batch_size : int Positive number of samples per call of either network. max_working_bytes : int Numeric buffer budget of each converted-network call. backend : {'auto', 'numpy', 'rust', 'go', 'mojo', 'julia'} Replay runtime. Auto is resolved once, and every batch runs on the runtime the report names.

Returns

ConversionLossReport Accuracies, agreement, rate error, executed runtime and digests.

Raises

ValueError Empty, non-finite or malformed inputs, invalid labels, an invalid batch size or seed, or outputs of different widths.

Source code in src/sc_neurocore/conversion/loss_report.py
Python
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
def measure_conversion_loss(
    model: Any,
    snn: ConvertedSNN,
    inputs: npt.ArrayLike,
    labels: npt.ArrayLike,
    *,
    input_mode: Literal["constant", "poisson"] = "constant",
    seed: int = 42,
    batch_size: int = 256,
    max_working_bytes: int = DEFAULT_WORKING_BYTES,
    backend: ReplayBackend = "auto",
) -> ConversionLossReport:
    """Classify labelled samples with a source ANN and its converted SNN and compare.

    Parameters
    ----------
    model : torch.nn.Module
        Source network. It runs under ``no_grad`` in inference mode on the
        device and dtype of its first parameter; every module's training flag
        is restored afterwards.
    snn : ConvertedSNN
        The network converted from ``model``.
    inputs : array_like
        ``(samples, *source_shape)`` values in ``[0, 1]``. The converted
        network receives each sample flattened in C order.
    labels : array_like
        ``(samples,)`` integer classes in ``[0, classes)``.
    input_mode : {'constant', 'poisson'}
        Declared encoding for the converted network.
    seed : int
        Unsigned 32-bit Poisson seed of the first batch.
    batch_size : int
        Positive number of samples per call of either network.
    max_working_bytes : int
        Numeric buffer budget of each converted-network call.
    backend : {'auto', 'numpy', 'rust', 'go', 'mojo', 'julia'}
        Replay runtime. Auto is resolved once, and every batch runs on the
        runtime the report names.

    Returns
    -------
    ConversionLossReport
        Accuracies, agreement, rate error, executed runtime and digests.

    Raises
    ------
    ValueError
        Empty, non-finite or malformed inputs, invalid labels, an invalid
        batch size or seed, or outputs of different widths.
    """
    if type(batch_size) is not int or batch_size < 1:
        raise ValueError("batch_size must be a positive integer")
    if type(seed) is not int or not 0 <= seed < _SEED_LIMIT:
        raise ValueError("seed must be an unsigned 32-bit integer")
    if input_mode not in ("constant", "poisson"):
        raise ValueError("input_mode must be 'constant' or 'poisson'")
    samples = np.array(inputs, dtype=np.float64)
    if samples.ndim < 2 or samples.shape[0] == 0:
        raise ValueError("inputs must hold at least one sample with its source shape")
    if not np.all(np.isfinite(samples)):
        raise ValueError("inputs must be finite")
    targets = np.array(labels)
    if targets.dtype.kind not in "iu" or targets.shape != (samples.shape[0],):
        raise ValueError("labels must be one integer class per sample")
    targets = targets.astype(np.int64)
    executed = resolve_replay_backend(backend)
    source = _source_outputs(model, samples, batch_size)
    if source.ndim != 2 or source.shape[0] != samples.shape[0]:
        raise ValueError("the source must produce one output vector per sample")
    classes = int(source.shape[1])
    if targets.min() < 0 or targets.max() >= classes:
        raise ValueError("labels must lie in [0, classes)")
    flat = samples.reshape(samples.shape[0], -1)
    rates = np.empty_like(source)
    responses = np.empty_like(source)
    for index, part in enumerate(_batches(len(flat), batch_size)):
        response = snn.run(
            flat[part],
            input_mode=input_mode,
            seed=(seed + index) % _SEED_LIMIT,
            max_working_bytes=max_working_bytes,
            backend=executed,
        )
        if response.shape != source[part].shape:
            raise ValueError("the converted network's output width differs from the source")
        responses[part] = response
        rates[part] = response / snn.T * snn.output_scale
    source_predictions = np.argmax(source, axis=1)
    converted_predictions = np.argmax(responses, axis=1)
    source_accuracy = float(np.mean(source_predictions == targets))
    converted_accuracy = float(np.mean(converted_predictions == targets))
    error = np.abs(rates - source)
    return ConversionLossReport(
        schema_version=LOSS_REPORT_SCHEMA_VERSION,
        samples=int(samples.shape[0]),
        classes=classes,
        timesteps=snn.T,
        input_mode=input_mode,
        seed=seed,
        batch_size=batch_size,
        backend=executed,
        source_accuracy=source_accuracy,
        converted_accuracy=converted_accuracy,
        accuracy_drop=source_accuracy - converted_accuracy,
        agreement=float(np.mean(source_predictions == converted_predictions)),
        rate_mean_abs_error=float(np.mean(error)),
        rate_max_abs_error=float(np.max(error)),
        source_sha256=source_sha256(model),
        converted_sha256=converted_sha256(snn),
        data_sha256=data_sha256(samples, targets),
    )

Target fixed-point calibration

calibrate_for_target(snn, profile, inputs, labels=None, *, batch_size=256, max_working_bytes=..., backend="auto") fits a converted network into one hardware profile's fixed-point format and measures what that costs on the given samples, driven as constant currents. Integrate-and-fire dynamics with subtractive reset are unchanged when one layer's weights, bias, threshold and membrane are scaled together, so each layer gets the finest power-of-two scale at which its coefficients, threshold and measured membrane peak fit the format; its coefficients are then rounded half to even onto the format's grid.

The rounded network is replayed on the same samples and compared with the unrounded one: agreement, the largest decoded-output difference and, with labels, both accuracies. Each layer reports its scale, threshold code, measured peak and headroom, rounding errors, weights rounded to zero, whether its preload lies on the grid and every sample-step at which the rounded network's membrane left the range. compatible is false, with named refusals, for a threshold below one grid step, negative coefficients in an unsigned format or any measured overflow.

A layer driven by spikes accumulates exact multiples of the grid step, so its replay equals integer fixed-point accumulation when exact_accumulation is true (format at most 53 bits wide, preloads on the grid). The first layer's analog products are not rounded as the target would round them, the profile's overflow and rounding modes are recorded rather than emulated, and only the numeric format is checked: neuron counts, fan-in, core mapping, routing, latency and energy are not. A range measured on calibration samples can be exceeded by other inputs.

sc_neurocore.conversion.target_report

Fit a converted network into one target's fixed-point format and measure the cost.

Integrate-and-fire dynamics with subtractive reset are invariant under scaling one layer's weights, bias, threshold and membrane by the same positive factor: the same neurons spike at the same steps. Each layer is therefore given the finest power-of-two scale at which its coefficients, its threshold and the membrane peak measured on the calibration samples fit the target's representable range, and its coefficients are rounded (half to even) onto the target's grid at that scale.

The network with rounded coefficients is then replayed on the same samples and compared with the unrounded one. A layer driven by spikes accumulates exact multiples of the grid step, so its float64 replay equals integer fixed-point accumulation while the width is at most 53 bits and the registers do not overflow; the replay counts every step at which a measured membrane leaves the range. A layer driven by analog input multiplies it in float64, which the target would round; that layer is not emulated and says so. Nothing here runs on, or estimates anything about, the target's latency or energy.

TargetReport dataclass

A converted network fitted into one target format and measured there.

Attributes

schema_version : str sc-neurocore.conversion-target-report.v1. profile : dict The target profile's identity and numeric format. compatible : bool Every layer fits with a threshold of at least one grid step and no measured overflow. refusals : list of str Why the network is not compatible; empty when it is. layers : list of LayerCalibration One entry per layer. samples, timesteps : int Calibration samples and the network's timestep budget. input_mode : str constant: every sample drives the first layer as a current. backend : str Replay runtime both networks ran on. agreement : float Fraction of samples on which both networks predict the same class. output_max_abs_difference : float Largest decoded-output difference between the two networks. float_accuracy, quantized_accuracy, accuracy_drop : float or None Present when labels were given. exact_accumulation : bool Whether spike-driven layers' replay equals integer accumulation: the format is at most 53 bits wide and their preloads lie on the grid. converted_sha256, data_sha256 : str Digests of the unrounded network and of the calibration samples. arithmetic : str What the measurement emulates and what it does not.

Source code in src/sc_neurocore/conversion/target_report.py
Python
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
@dataclass(frozen=True)
class TargetReport:
    """A converted network fitted into one target format and measured there.

    Attributes
    ----------
    schema_version : str
        ``sc-neurocore.conversion-target-report.v1``.
    profile : dict
        The target profile's identity and numeric format.
    compatible : bool
        Every layer fits with a threshold of at least one grid step and no
        measured overflow.
    refusals : list of str
        Why the network is not compatible; empty when it is.
    layers : list of LayerCalibration
        One entry per layer.
    samples, timesteps : int
        Calibration samples and the network's timestep budget.
    input_mode : str
        ``constant``: every sample drives the first layer as a current.
    backend : str
        Replay runtime both networks ran on.
    agreement : float
        Fraction of samples on which both networks predict the same class.
    output_max_abs_difference : float
        Largest decoded-output difference between the two networks.
    float_accuracy, quantized_accuracy, accuracy_drop : float or None
        Present when labels were given.
    exact_accumulation : bool
        Whether spike-driven layers' replay equals integer accumulation: the
        format is at most 53 bits wide and their preloads lie on the grid.
    converted_sha256, data_sha256 : str
        Digests of the unrounded network and of the calibration samples.
    arithmetic : str
        What the measurement emulates and what it does not.
    """

    schema_version: str
    profile: dict[str, Any]
    compatible: bool
    refusals: list[str]
    layers: list[LayerCalibration]
    samples: int
    timesteps: int
    input_mode: str
    backend: str
    agreement: float
    output_max_abs_difference: float
    float_accuracy: float | None
    quantized_accuracy: float | None
    accuracy_drop: float | None
    exact_accumulation: bool
    converted_sha256: str
    data_sha256: str
    arithmetic: str

    def to_public_dict(self) -> dict[str, Any]:
        """Return the report as JSON-ready fields.

        Returns
        -------
        dict
            Every attribute under its own name, layers as objects.
        """
        return asdict(self)

to_public_dict()

Return the report as JSON-ready fields.

Returns

dict Every attribute under its own name, layers as objects.

Source code in src/sc_neurocore/conversion/target_report.py
Python
161
162
163
164
165
166
167
168
169
def to_public_dict(self) -> dict[str, Any]:
    """Return the report as JSON-ready fields.

    Returns
    -------
    dict
        Every attribute under its own name, layers as objects.
    """
    return asdict(self)

LayerCalibration dataclass

How one layer was fitted into the target format, and what that cost.

Attributes

index : int Zero-based layer position. drive : {'analog', 'spikes'} What the layer integrates; only a spike-driven layer's replay is exact. readout : {'spiking', 'linear'} Whether the layer fires or integrates its current as the readout. scale_exponent : int e of the scale 2**e applied to the layer's normalised values. threshold_code : int The threshold as a stored integer; zero for a linear readout. measured_peak : float Largest pre-reset membrane magnitude on the calibration samples, in normalised threshold units. headroom_bits : float or None log2 of the representable maximum over the scaled peak; None when the membrane never moved. weights : int Number of weights. zeroed_weights : int Non-zero weights that round to zero. weight_max_abs_error, weight_rms_error : float Rounding error of the weights in normalised units. bias_max_abs_error : float Rounding error of the bias; zero without a bias. preload_exact : bool Whether the initial membrane preload lies on the grid. overflow_steps : int Sample-steps at which the rounded network's membrane left the range.

Source code in src/sc_neurocore/conversion/target_report.py
Python
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
@dataclass(frozen=True)
class LayerCalibration:
    """How one layer was fitted into the target format, and what that cost.

    Attributes
    ----------
    index : int
        Zero-based layer position.
    drive : {'analog', 'spikes'}
        What the layer integrates; only a spike-driven layer's replay is exact.
    readout : {'spiking', 'linear'}
        Whether the layer fires or integrates its current as the readout.
    scale_exponent : int
        ``e`` of the scale ``2**e`` applied to the layer's normalised values.
    threshold_code : int
        The threshold as a stored integer; zero for a linear readout.
    measured_peak : float
        Largest pre-reset membrane magnitude on the calibration samples, in
        normalised threshold units.
    headroom_bits : float or None
        ``log2`` of the representable maximum over the scaled peak; ``None``
        when the membrane never moved.
    weights : int
        Number of weights.
    zeroed_weights : int
        Non-zero weights that round to zero.
    weight_max_abs_error, weight_rms_error : float
        Rounding error of the weights in normalised units.
    bias_max_abs_error : float
        Rounding error of the bias; zero without a bias.
    preload_exact : bool
        Whether the initial membrane preload lies on the grid.
    overflow_steps : int
        Sample-steps at which the rounded network's membrane left the range.
    """

    index: int
    drive: Literal["analog", "spikes"]
    readout: Literal["spiking", "linear"]
    scale_exponent: int
    threshold_code: int
    measured_peak: float
    headroom_bits: float | None
    weights: int
    zeroed_weights: int
    weight_max_abs_error: float
    weight_rms_error: float
    bias_max_abs_error: float
    preload_exact: bool
    overflow_steps: int

calibrate_for_target(snn, profile, inputs, labels=None, *, batch_size=256, max_working_bytes=DEFAULT_WORKING_BYTES, backend='auto')

Fit snn into profile's fixed-point format and measure the rounded network.

Parameters

snn : ConvertedSNN Converted network in normalised threshold units. profile : HardwareProfile Target format from :mod:sc_neurocore.compiler.platforms. inputs : array_like (samples, input_neurons) calibration values in [0, 1], driven as constant currents for the network's timestep budget. labels : array_like, optional (samples,) integer classes; accuracies are reported when given. batch_size : int Positive samples per replay. max_working_bytes : int Numeric buffer budget of each replay chunk. backend : {'auto', 'numpy', 'rust', 'go', 'mojo', 'julia'} Replay runtime, resolved once and named in the report.

Returns

TargetReport Per-layer scales and rounding costs, measured ranges, both networks' agreement and, with labels, accuracies.

Raises

ValueError Empty or out-of-range inputs, invalid labels or batch size.

Source code in src/sc_neurocore/conversion/target_report.py
Python
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
def calibrate_for_target(
    snn: ConvertedSNN,
    profile: HardwareProfile,
    inputs: npt.ArrayLike,
    labels: npt.ArrayLike | None = None,
    *,
    batch_size: int = 256,
    max_working_bytes: int = DEFAULT_WORKING_BYTES,
    backend: ReplayBackend = "auto",
) -> TargetReport:
    """Fit ``snn`` into ``profile``'s fixed-point format and measure the rounded network.

    Parameters
    ----------
    snn : ConvertedSNN
        Converted network in normalised threshold units.
    profile : HardwareProfile
        Target format from :mod:`sc_neurocore.compiler.platforms`.
    inputs : array_like
        ``(samples, input_neurons)`` calibration values in ``[0, 1]``, driven as
        constant currents for the network's timestep budget.
    labels : array_like, optional
        ``(samples,)`` integer classes; accuracies are reported when given.
    batch_size : int
        Positive samples per replay.
    max_working_bytes : int
        Numeric buffer budget of each replay chunk.
    backend : {'auto', 'numpy', 'rust', 'go', 'mojo', 'julia'}
        Replay runtime, resolved once and named in the report.

    Returns
    -------
    TargetReport
        Per-layer scales and rounding costs, measured ranges, both networks'
        agreement and, with labels, accuracies.

    Raises
    ------
    ValueError
        Empty or out-of-range inputs, invalid labels or batch size.
    """
    if type(batch_size) is not int or batch_size < 1:
        raise ValueError("batch_size must be a positive integer")
    samples = np.array(inputs, dtype=np.float64)
    if samples.ndim != 2 or samples.shape[0] == 0:
        raise ValueError("inputs must be a non-empty (samples, input_neurons) array")
    if not np.all(np.isfinite(samples)) or samples.min() < 0 or samples.max() > 1:
        raise ValueError("inputs must be finite values in [0, 1]")
    targets = None
    if labels is not None:
        targets = np.array(labels)
        if targets.dtype.kind not in "iu" or targets.shape != (samples.shape[0],):
            raise ValueError("labels must be one integer class per sample")
    executed = resolve_replay_backend(backend)
    float_out, peaks, _ = _replay_with_peaks(
        snn, samples, batch_size, executed, max_working_bytes, None
    )
    low, high = _code_bounds(profile)
    grid = 2.0**-profile.fraction
    top = high * grid
    fractions = snn.layer_membrane_fractions or [snn.initial_membrane_fraction] * snn.n_layers
    linear_index = snn.n_layers - 1 if snn.output_mode == "linear" else -1
    refusals: list[str] = []
    layers: list[LayerCalibration] = []
    weights: list[npt.NDArray[np.float64]] = []
    biases: list[npt.NDArray[np.float64] | None] = []
    thresholds: list[float] = []
    limits: list[tuple[float, float]] = []
    for index in range(snn.n_layers):
        weight = np.asarray(snn.weights[index], dtype=np.float64)
        bias = snn.biases[index]
        threshold = float(snn.thresholds[index])
        linear = index == linear_index
        extent = max(
            float(np.abs(weight).max(initial=0.0)),
            0.0 if bias is None else float(np.abs(bias).max(initial=0.0)),
            0.0 if linear else threshold,
            peaks[index],
        )
        exponent = math.floor(math.log2(top / extent)) if extent > 0 and top > 0 else 0
        scale = 2.0**exponent
        if not profile.signed and (
            weight.min(initial=0.0) < 0 or (bias is not None and bias.min() < 0)
        ):
            refusals.append(f"layer {index}: negative coefficients in an unsigned format")
        threshold_code = 0 if linear else int(np.rint(threshold * scale / grid))
        if not linear and threshold_code < 1:
            refusals.append(f"layer {index}: its dynamic range leaves the threshold below one step")
        weight_q = _on_grid(weight, scale, grid, low, high)
        error = np.abs(weight_q - weight)
        bias_q = None
        bias_error = 0.0
        if bias is not None:
            bias_q = _on_grid(np.asarray(bias, dtype=np.float64), scale, grid, low, high)
            bias_error = float(np.abs(bias_q - bias).max(initial=0.0))
        weights.append(weight_q)
        biases.append(bias_q)
        stored = threshold if linear else max(threshold_code, 1) * grid / scale
        thresholds.append(stored)
        stored_preload = 0.0 if linear else float(fractions[index]) * stored
        limits.append((low * grid / scale, top / scale))
        layers.append(
            LayerCalibration(
                index=index,
                drive="analog" if index == 0 else "spikes",
                readout="linear" if linear else "spiking",
                scale_exponent=exponent,
                threshold_code=threshold_code,
                measured_peak=peaks[index],
                headroom_bits=math.log2(top / (peaks[index] * scale)) if peaks[index] > 0 else None,
                weights=int(weight.size),
                zeroed_weights=int(np.count_nonzero((weight != 0) & (weight_q == 0))),
                weight_max_abs_error=float(error.max(initial=0.0)),
                weight_rms_error=float(np.sqrt(np.mean(error**2))) if error.size else 0.0,
                bias_max_abs_error=bias_error,
                preload_exact=bool(
                    np.rint(stored_preload * scale / grid) * grid / scale == stored_preload
                ),
                overflow_steps=0,
            )
        )
    quantized = ConvertedSNN(
        weights,
        biases,
        thresholds,
        snn.T,
        snn.initial_membrane_fraction,
        snn.output_scale,
        snn.output_mode,
        max_working_bytes=max_working_bytes,
        layer_membrane_fractions=snn.layer_membrane_fractions,
    )
    quant_out, _, exits = _replay_with_peaks(
        quantized, samples, batch_size, executed, max_working_bytes, limits
    )
    for index, count in enumerate(exits):
        layers[index] = LayerCalibration(**{**asdict(layers[index]), "overflow_steps": count})
        if count:
            refusals.append(f"layer {index}: membrane left the range at {count} sample-steps")
    float_classes = float_out.argmax(axis=1)
    quant_classes = quant_out.argmax(axis=1)
    decode = snn.output_scale / snn.T
    float_accuracy = quantized_accuracy = drop = None
    data_labels = np.zeros(0, dtype=np.int64) if targets is None else targets.astype(np.int64)
    if targets is not None:
        float_accuracy = float(np.mean(float_classes == targets))
        quantized_accuracy = float(np.mean(quant_classes == targets))
        drop = float_accuracy - quantized_accuracy
    return TargetReport(
        schema_version=TARGET_REPORT_SCHEMA_VERSION,
        profile=_profile_identity(profile),
        compatible=not refusals,
        refusals=refusals,
        layers=layers,
        samples=int(samples.shape[0]),
        timesteps=snn.T,
        input_mode="constant",
        backend=executed,
        agreement=float(np.mean(float_classes == quant_classes)),
        output_max_abs_difference=float(np.abs(quant_out - float_out).max() * decode),
        float_accuracy=float_accuracy,
        quantized_accuracy=quantized_accuracy,
        accuracy_drop=drop,
        exact_accumulation=profile.data_width <= _EXACT_WIDTH
        and all(layer.preload_exact for layer in layers[1:]),
        converted_sha256=converted_sha256(snn),
        data_sha256=data_sha256(samples, data_labels),
        arithmetic=(
            "Coefficients rounded half to even onto the target grid at a per-layer "
            "power-of-two scale; replayed in float64. Spike-driven layers equal integer "
            "accumulation when exact_accumulation is true; the analog-driven first layer's "
            "products are not rounded as the target would round them. Target overflow and "
            "rounding modes are recorded, not emulated. Only the numeric format is checked: "
            "neuron counts, fan-in, core mapping and routing are not, and latency and energy "
            "are not measured."
        ),
    )

Reproducible runtime comparisons

After explicitly building and configuring all four native providers as above, run the complete public-call comparison on one allowed Linux CPU:

Bash
PYTHONPATH=src python benchmarks/bench_ann_to_snn_replay.py \
  --output /your/evidence/dense-if-comparison.json --cpu 0 --samples 7 --warmup 3

From an installed wheel, the same bundled comparison runs as:

Bash
sc-neurocore-if-benchmark \
  --output /your/evidence/dense-if-comparison.json --cpu 0 --samples 7 --warmup 3

The command runs NumPy, Rust, Go, Mojo and Julia sequentially in fresh equally pinned processes. Twenty fixed seeded workloads cover single/dense stacks, full state/event traces, both output modes, empty timesteps and 129-step constant/Poisson encoders. Every response bit must match NumPy before a sample is accepted. Each record binds shaped input/output digests, raw timings, current source bytes and installed native artifact bytes. The report is published atomically only after all five runtimes pass and sources/artifacts remain unchanged; worker stdout/stderr are retained even when a comparison fails.

Latency includes public-call admission, encoding where applicable, native transport and collection of the owned result vectors. Final returned-vector release, interpreter startup and builds are excluded. The first selected provider call follows the NumPy reference and includes lazy native startup; warm samples follow the declared exact-workload repetitions. The aggregate is an equal-weight geometric mean of the twenty warm workload medians. Single-CPU affinity does not isolate host load, and results are specific to that capture. These seeded weights are untrained runtime workloads. The comparison is separate from trained source/target loss acceptance and has no energy estimate or physical instrument/calibration claim.