Skip to content

Neuromorphic Datasets

Module: sc_neurocore.datasets Source: src/sc_neurocore/datasets/ — 6 files, 1593 lines Status (v3.16.0): 18 public symbols; 98 tests across the dataset and dataset CLI test files; NumPy and h5py I/O with optional native decoding of N-MNIST binary records through Go, Rust, Mojo or Julia, no synaptic kinetics. The "Poisson" encoder is actually Bernoulli (§3.1, same wording issue as network/stimulus.PoissonInput).

This page covers the two encoders (poisson_encode, latency_encode), the three event-camera / cochlear loaders (load_nmnist, load_shd, load_dvs_cifar10), each of which can fall back to synthetic data when the real archive is not on disk, and what a study needs to say which data it used and how: manifests of the files on disk (§2.5), splits that keep every speaker or recording on one side (§2.6) and encoders declared completely enough to rebuild them (§3.3).


1. Public surface

sc_neurocore.datasets.__init__ re-exports 18 symbols:

Symbol Source file Role
poisson_encode encoding.py Per-neuron Bernoulli draw → spike train
latency_encode encoding.py Continuous value → first-spike-time
load_nmnist loaders.py N-MNIST (Orchard 2015), 34×34 DVS
load_shd loaders.py Spiking Heidelberg Digits (Cramer 2022), 700 channels
load_dvs_cifar10 loaders.py DVS-CIFAR10 (Li 2017), 128×128 DVS
EventBinning, PoissonRates, FirstSpikeLatency encoders.py Encoders with a complete declaration
encoder_from_declaration encoders.py Rebuild an encoder from its declaration
EventDatasetManifest, build_manifest manifest.py Files by SHA-256, samples by split, label and group
verify_manifest, manifest_from_dict manifest.py Compare a directory with a manifest; read one back
SplitPlan, group_split, split_plan_from_dict splits.py Divide a published split by whole groups
group_overlap, leaked_groups splits.py Groups shared by published splits; groups a plan leaks

Each loader accepts synthetic=True to bypass disk reads — useful for unit tests and for CI where the real archives are not stored.


2. Loaders

2.1 load_nmnist

Python
def load_nmnist(
    root: str | Path = "data/nmnist",
    train: bool = True,
    dt_ms: float = 1.0,
    T: int = 300,
    synthetic: bool = False,
    n_samples: int = 100,
    seed: int = 42,
) -> tuple[list[np.ndarray], np.ndarray]:

Loads the N-MNIST dataset:

Orchard G., Cohen G., Jayawant A., Thakor N. "Converting Static Image Datasets to Spiking Neuromorphic Datasets Using Saccades." Front Neurosci 9:437 (2015).

34×34 ATIS DVS recordings of MNIST digits moved across the sensor by saccadic eye movements. 10 classes.

Returns (samples, labels): - samples: list of (N_events, 4) arrays with columns [x, y, polarity, timestamp_ms]; real N-MNIST recordings use float64 to retain timestamp precision before binning. Synthetic samples use float32. - labels: int64 array of length len(samples)

The real-data path expects the directory layout root/{Train,Test}/<class_id>/<sample>.bin. Each .bin file is a sequence of 40-bit events in the format published with the dataset, parsed by the helper _parse_nmnist_bin:

Bits Meaning
39–32 (byte 0) x address in pixels
31–24 (byte 1) y address in pixels
23 polarity (0 OFF, 1 ON)
22–0 timestamp in microseconds

Incomplete 40-bit records are refused instead of silently discarded. Timestamp conversion retains float64: narrowing to float32 can move events across fractional bin and window boundaries.

Timestamps are returned in milliseconds (microseconds / 1000); dt_ms does not affect the real path. The 34 × 34 sensor needs six address bits per axis, which is why the addresses are whole bytes.

Native indexed SHD readers

sc_neurocore.accel.shd_recordings.read_shd_recording(path, index) returns one float64 event matrix and its integer label. The manifest sample reader uses this interface and refuses a recording whose actual label differs from the manifest. Numeric variable-length time/channel vectors and the integer labels must have matching recording counts. Seconds are widened before conversion to milliseconds; an empty recording stays empty.

The Python/h5py reader is always available. Build the optional Rust reader using the existing system HDF5 library and pkg-config; this standalone crate has no Cargo dependency downloads:

Bash
cargo build --offline --locked --release --manifest-path src/sc_neurocore/accel/rust/safety/shd_native/Cargo.toml --target-dir /absolute/operator/path/shd-target

Set SC_NEUROCORE_SHD_RUST_LIBRARY to the generated /absolute/operator/path/shd-target/release/libsc_neurocore_shd.so. The public Rust API is sc_neurocore_shd::read_shd_recording; the shared C ABI returns an owning result with shd_read_c and releases it with shd_free_c. Rust owns selected HDF5 ids, native variable-length reclaim and result buffers; its mutex serialises this reader's HDF5 calls.

Build the optional Go reader from src/sc_neurocore/accel/go with system HDF5 headers/library and pkg-config:

Bash
go build -tags=hdf5 -buildmode=c-shared -o /absolute/operator/path/libshd.so ./services/loaders/cshared

Set SC_NEUROCORE_SHD_GO_LIBRARY to that absolute file before reading. This is separate from the N-MNIST library setting. Julia uses the existing absolute executable selected by SC_NEUROCORE_SHD_JULIA_EXE, with backend="julia". Its standalone standard-library CLI calls system HDF5 directly; no JuliaCall project is required. On Linux, SC_NEUROCORE_SHD_JULIA_HDF5_LIBRARY can select an existing absolute HDF5 library; otherwise the library name is libhdf5_serial.so. Explicit backend="rust", backend="go" or backend="numpy" selects those readers; auto considers measured shd-recording backend order, otherwise Rust, Go, configured Julia and configured Mojo before Python. A configured native failure refuses the read without substitution.

Rust and Go execute in fresh Python processes loading their actual HDF5 C ABIs. Julia executes in its own runtime process. This keeps Julia's already-loaded dependencies separate from system HDF5; direct same-process library coexistence is not required. Each native read has a 30-second lifetime, inherits the caller's process group for supervised cancellation, and exits if its parent dies. The caller reaps it. Julia installs its Linux parent-death guard and 30-second alarm after runtime startup, before loading and compiling the reader, and checks the expected parent before opening HDF5. The Linux signal tracks the creating parent thread. Startup still relies on the caller deadline and supervisor process-group ownership; it is not guarded before Julia begins executing the CLI. Actual blocked-open tests separately exercise parent death, the native alarm, stale-parent refusal and public-call timeout/reaping. Library code is operator-owned, never supplied by an experiment request.

maximum_bytes defaults to 64 MiB and limits the returned event matrix. It is not an aggregate memory cap: HDF5 vectors, native copies and IPC buffers use additional transient memory. This bound is distinct from the Studio encoded-input admission budget. Cross-language SHD comparison benchmarks remain under implementation; these native paths do not establish whole-chain performance parity.

Mojo native SHD API

From a source checkout, the shd module in src/sc_neurocore/accel/mojo/kernels provides read_shd_recording(path, index, maximum_bytes=67108864, hdf5_library="libhdf5_serial.so"). It returns owned row-major double values and an int64 label, using direct system HDF5 calls through Mojo's dynamic FFI. Compile the caller with that kernel directory on the Mojo import path, on a 64-bit host with system HDF5. No Python decoder is called. The owning context closes identifiers in reverse order and reclaims each selected VLEN buffer; errors propagate after cleanup. This API does not serialise foreign HDF5 callers and should run in an isolated process.

Compile the supervised standalone CLI from the checkout root:

Bash
mojo build --Werror --fp-mode contract=off src/sc_neurocore/accel/mojo/kernels/shd_cli.mojo -o /absolute/operator/path/shd-reader

Set SC_NEUROCORE_SHD_MOJO_EXE to that existing absolute executable and select backend="mojo". SC_NEUROCORE_SHD_MOJO_HDF5_LIBRARY optionally selects an existing absolute system HDF5 file. This Linux CLI emits the same bounded SHD protocol, arms a parent-death guard and native 30-second alarm before reading, and runs in the caller's process group. The caller checks the whole result and reaps the process. A declared incompatible library or executable refuses without Python substitution. Actual blocked-FIFO tests independently verify parent-death SIGKILL, process-group cancellation, the native SIGALRM deadline and stale-parent refusal; the contained subreaper reaps its owned reader. The guard is installed when the CLI begins execution, so runtime startup remains under the caller deadline and supervisor ownership. Linux parent-death signals follow the creating parent thread. These checks do not establish full coverage or cross-language performance parity.

Optional native decoders

Both load_nmnist and the manifest-bound read_event_sample use the same decoder. By default NumPy is available without a native build. To compile the Go implementation from a checkout, run from src/sc_neurocore/accel/go:

Bash
go build -buildmode=c-shared -o /absolute/operator/path/libloaders.so ./services/loaders/cshared
export SC_NEUROCORE_DATASET_GO_LIBRARY=/absolute/operator/path/libloaders.so

The same C interface is available from Rust. From the checkout root:

Bash
cargo build --release --locked --manifest-path src/sc_neurocore/accel/rust/safety/nmnist_native/Cargo.toml --target-dir /absolute/operator/path/nmnist-target
export SC_NEUROCORE_DATASET_RUST_LIBRARY=/absolute/operator/path/nmnist-target/release/libsc_neurocore_nmnist.so

Automatic decoding uses the host's recorded nmnist-recording benchmark order. Without matching measurements, Rust precedes Mojo, Julia and Go; NumPy remains the final floor. Mojo uses the same C interface; build from the checkout root:

Bash
mojo build --Werror --fp-mode contract=off --emit shared-lib -o /absolute/operator/path/libnmnist.so src/sc_neurocore/accel/mojo/kernels/nmnist.mojo
export SC_NEUROCORE_DATASET_MOJO_LIBRARY=/absolute/operator/path/libnmnist.so

Julia uses an existing JuliaCall/PythonCall environment with matching package versions. Configure it before starting Studio; reads do not install packages:

Bash
export SC_NEUROCORE_DATASET_JULIA_ENABLED=1
export PYTHON_JULIACALL_EXE=/absolute/operator/path/julia
export PYTHON_JULIACALL_PROJECT=/absolute/operator/path/julia-project
export PYTHON_JULIACALL_THREADS=1
export PYTHON_JULIACALL_HANDLE_SIGNALS=yes
export PYTHON_JULIACALL_STARTUP_FILE=no

The project must already contain PythonCall matching the installed JuliaCall. JuliaCall performs JIT compilation on first use; warmup excludes that startup from decoder timing. Disabled Julia is skipped by automatic selection, while explicit Julia selection requires opt-in. Invalid runtime settings refuse the read, including after a kernel has already been loaded. Runtime settings are process-wide and cannot be changed after JuliaCall starts.

The library path is an operator setting, never part of an imported dataset or workspace. Python does not build or download native libraries during reads. An explicitly selected missing, unreadable or incompatible native library refuses the read; it does not silently switch backends. The native decoder retains float64 timestamps and refuses incomplete records before output mutation. sc_neurocore.accel.event_recordings.decode_nmnist_recording also offers explicit backend="numpy", backend="rust", backend="mojo", backend="julia" or backend="go" selection for parity measurements. This native path decodes the raw bytes; manifest verification, file selection and encoder validation still use the same public dataset contracts.

benchmarks/bench_event_recordings.py compares the executable Rust, Mojo, Julia, Go and NumPy decoders on a caller-supplied .bin recording. It checks exact output parity before timing and records the file digest, CPU, warmup and repetitions. All three native libraries and the configured Julia runtime must be available; missing implementations are refused. File I/O is outside the timed region; output allocation and the native call are included. A format fixture can exercise the command, but its timings do not qualify publisher datasets, training performance, latency or power.

2.2 load_shd

Loads the Spiking Heidelberg Digits dataset:

Cramer B., Stradmann Y., Schemmel J., Zenke F. "The Heidelberg Spiking Data Sets for the Systematic Evaluation of Spiking Neural Networks." IEEE Transactions on Neural Networks and Learning Systems 33(7):2744-2757 (2022).

20-class English/German digit utterances (0–9 in two languages) spike-encoded through Lauscher's artificial cochlea model. 700 input channels.

Returns (samples, labels): - samples: list of (T_per_sample, 700) bool arrays (binned spike rasters). T_per_sample is min(ceil(times.max() / (dt_ms/1000)) + 1, T). - labels: int64 array

The real-data path requires h5py (declared in extras) and reads root/shd_{train,test}.h5. The H5 layout is the standard SHD release: /spikes/times[i] in seconds, /spikes/units[i], /labels, and /extra/speaker. Spikes after the T-step window are dropped rather than merged into the last bin.

Stored seconds are widened to float64 before conversion to milliseconds and binning, matching the manifest-bound training reader. Arithmetic in the source float16 dtype can shift events to an earlier bin even when the recorded time itself is exact. Widening retains stored values; it cannot recover precision lost when recordings were written.

Publisher integrity and release metadata

The publisher's resource page licenses SHD under CC BY 4.0 and supplies the dataset citation. Its checksum list identifies the compressed shd_train.h5.gz and shd_test.h5.gz files. Verify those archives against the publisher list, then compare the SHA-256 of their complete decompressed streams with the HDF5 files used by the manifest. A matching archive does not prove that a separately cached HDF5 file is unchanged. Publisher MD5 is an integrity reference; it is not a signed SHA-256 receipt.

The publisher README names release 1.0 but lists 8,332 training and 2,088 test recordings; the resource page lists 8,156 and 2,264. Record the actual file digests and sample counts, and retain this metadata disagreement when declaring a release. Do not infer an unpublished revision from the counts. Published train/test partitions also share speakers: use the manifest's group-overlap report to distinguish the publisher benchmark from a split that holds out every evaluation speaker. A correct reader and encoder do not establish classifier accuracy or hardware performance.

2.3 load_dvs_cifar10

Loads the DVS-CIFAR10 dataset:

Li H., Liu H., Ji X., Li G., Shi L. "CIFAR10-DVS: An Event-Stream Dataset for Object Classification." Front Neurosci 11:309 (2017).

CIFAR-10 images displayed on a monitor and recorded by a 128×128 DVS camera. 10 classes.

Returns (samples, labels) in the same shape as load_nmnist. The real-data path expects .npy files (one per sample) under root/{train,test}/<class_id>/. Each .npy must be an array with columns [x, y, polarity, timestamp_ms]. Raw .aedat / .mat conversion is left to the caller.

Both the eager loader and the manifest-bound single-recording reader refuse arrays without four columns, complex/text/object values, malformed or ambiguous headers, truncated payloads and additional file content. Pickle is never read. NPY versions 1.0/2.0/3.0, C/Fortran order and either byte order are accepted. Real integer, Boolean and floating arrays become writable owned float64 matrices, preserving the NumPy conversion of their stored values before binning; synthetic recordings remain float32. Conversion cannot recover precision already lost when the source file was written.

The public sc_neurocore.accel.dvs_recordings.read_dvs_recording(path, maximum_bytes=67108864) reference reader checks shape and a 64 MiB returned matrix budget before reading the payload. Its header limit is 10,000 bytes; input bytes and temporary copies consume additional memory. Each converted file must contain exactly one event array. Geometry, finite values and time ordering remain the encoder's responsibility. Native Go, Rust, Julia and Mojo readers share this format contract. The DVS benchmark below measures actual public reader calls across all five paths; format checks are not publisher-data or target-hardware acceptance.

The native Go API loaders.ReadDVSRecording(path, maximumBytes) returns an owned row-major []float64 with the same four columns. Its standalone command is built from services/loaders/dvscli in the Go module. Arguments are the NPY path, returned matrix byte budget and optional expected parent PID. The command requires Linux, arms parent-death termination and exits with status 124 after 30 seconds, including blocked file or output I/O. Supplying the creating parent's PID also refuses stale startup. The caller must reap the process and reject incomplete output frames. The 16-byte extended floating representation is qualified for an x86-64 host.

Python read_dvs_recording(..., backend="auto") selects the explicitly declared SC_NEUROCORE_DVS_GO_EXE, SC_NEUROCORE_DVS_RUST_EXE, SC_NEUROCORE_DVS_JULIA_EXE or SC_NEUROCORE_DVS_MOJO_EXE; without a declaration it uses NumPy. Multiple declarations refuse automatic selection. backend="go", backend="rust" or backend="julia" or backend="mojo" requires its own setting, while backend="numpy" selects the reference. An attempted native read never falls back after failure. Eager DVS loading and manifest-bound lazy samples use the same automatic selection. Isolated Studio workers receive the setting only from the operator's event_input.dvs_go_executable or event_input.dvs_rust_executable, event_input.dvs_julia_executable or event_input.dvs_mojo_executable configuration, validated as an existing absolute executable. Job requests do not supply executable paths. The parent checks the complete binary result, bounds its wait and kills/reaps the child on timeout or interruption. Input/native buffers and pipe copies consume additional memory beyond the returned matrix.

Rust sc_neurocore_dvs::read_dvs_recording(path, maximum_bytes) returns an owned row-major Vec<f64> under the same recording contract. Build the command from the repository root:

Bash
cargo build --offline --locked --release --manifest-path src/sc_neurocore/accel/rust/safety/dvs_native/Cargo.toml

It accepts the same arguments and emits the same little-endian frame as Go. Linux parent death is SIGKILL; the independent 30-second deadline terminates with SIGALRM. The caller must reap it and reject partial frames. Extended scalars use the host C long double conversion and require a 16-byte ABI; representation follows that host. Geometry, timestamp ordering and manifest validation remain caller duties.

Julia DVSRecordings.read_dvs_recording(path; maximum_bytes=64*1024*1024) in accel/julia/datasets/dvs.jl returns an independent Matrix{Float64} with four event columns. The decoder uses Julia Base and supports the same NPY versions, byte orders and C/Fortran input layouts. Its extended conversion interprets the padded x87 encoding on Linux x86-64 and rounds to float64 through BigFloat. The direct command is:

Bash
julia --startup-file=no --check-bounds=yes --depwarn=error src/sc_neurocore/accel/julia/datasets/dvs_cli.jl PATH BYTE_BUDGET [EXPECTED_PARENT]

It emits the same row-major little-endian DVS1 frame and arms Linux parent-death termination and a 30-second input/output alarm before loading and compiling the recording reader. Julia runtime startup is additional time; callers must bound the complete process lifetime and reap it. Python selects the explicitly declared installed Julia runtime and supplies the packaged script, --startup-file=no, --check-bounds=yes, --depwarn=error and --threads=1. Its 30-second communication deadline includes runtime startup; timeout and interruption kill and reap the child. No JuliaCall project or package installation is involved.

The Mojo read_dvs_recording(path, maximum_bytes) API in accel/mojo/kernels/dvs_recordings.mojo returns an owned row-major List[Float64]. It parses inert NPY metadata and converts stored scalars natively; Python is not used to decode the recording. Build its standalone command with the same host C extended-precision conversion used by the Rust reader:

Bash
cc -fPIC -c src/sc_neurocore/accel/rust/safety/dvs_native/src/extended.c -o /tmp/sc-neurocore-dvs-extended.o
mojo build src/sc_neurocore/accel/mojo/kernels/dvs_cli.mojo --Werror --fp-mode contract=off -j 2 -Xlinker /tmp/sc-neurocore-dvs-extended.o -o /tmp/sc-neurocore-dvs-mojo

Set SC_NEUROCORE_DVS_MOJO_EXE to the compiled command's absolute path, or use the operator's event_input.dvs_mojo_executable setting for isolated Studio workers. Arguments and DVS1 output framing match the other native commands. Linux parent-death termination and a 30-second alarm cover input and output blocking. The command completes decoding before emitting its frame; native refusal never selects NumPy. This implementation targets little-endian x86-64 Linux with a 16-byte host C long-double ABI.

2.4 Common contracts

All three loaders: - Raise FileNotFoundError with the dataset's download URL embedded in the message when root does not exist (_check_root, loaders.py:73). - Raise FileNotFoundError when root exists but the train/test subdirectory is missing. - Accept synthetic=True to bypass disk reads entirely; the synthetic path uses _synthetic_event_dataset (event-based loaders) or _synthetic_shd (binned-raster loader). Both pin their RNG to seed for reproducibility.

The synthetic generators draw class-conditional rate templates from U(0, 0.3) (event loaders) or U(0, 0.1) (SHD), then expand them through poisson_encode to per-sample spike trains. Polarities for event loaders are randint(0, 2).

2.5 Manifests of the files on disk

build_manifest(name, root, version=...) records one dataset directory in the layout its loader reads: every file with its size and SHA-256, and every sample with its published split, label and group. The dataset description it carries is the publisher's: citation and DOI, where the files are distributed, the data licence, the sensor geometry and the unit of time the files store.

Dataset Licence Sensor File times Group
nmnist CC-BY-SA-4.0 DVS 34 × 34, 2 polarities microseconds recording file
shd CC-BY-4.0 cochlea, 700 channels seconds speaker
dvs_cifar10 CC-BY-4.0 DVS 128 × 128, 2 polarities ms, in user-converted .npy recording file

The group is the unit that must stay on one side of a split. SHD records a speaker for every sample; N-MNIST and CIFAR10-DVS name no identity finer than the recording, so each file is its own group. CIFAR10-DVS publishes no train/test split: the train/test folders the loader reads are the user's.

The version is stated by the user, because the files do not carry it. Nothing is downloaded. verify_manifest(manifest, root) lists files that are missing, changed in size or bytes, or present in the layout but not in the manifest; ok is true only when there are none. manifest.digest identifies the exact manifest and can be recorded with a result. manifest_from_dict reads the JSON form back and refuses another schema, an unknown field, or a dataset description that differs from the one this version ships.

2.6 Group splits

group_split(manifest, fractions={"train": 0.8, "validation": 0.2}, source_split="train", seed=0) divides one published split by whole groups. Groups are taken in an order fixed by the seed, and each goes to the part furthest below its requested share of samples, so the shares are met as closely as whole groups allow. Every part receives at least one group; a request for more parts than there are groups is refused. The other published splits are untouched.

The plan records the manifest digest it was drawn from. leaked_groups(manifest, plan) returns the groups with samples in more than one part — empty for a plan group_split made, and the check to run on a plan read back with split_plan_from_dict. group_overlap(manifest) reports groups the publisher's own splits share, which matters before comparing a score with published ones. The SHD release states that two of its twelve speakers appear only in the test set; which training speakers the test set also holds is what group_overlap reports for the files on disk.


3. Encoders

3.1 poisson_encode (actually Bernoulli)

Python
def poisson_encode(
    rates: npt.ArrayLike,
    T: int,
    dt_ms: float = 1.0,
    seed: int | None = None,
) -> np.ndarray:  # shape (T, N), bool

Returns (T, N) boolean spike train: each cell is rng.random() < min(rate * dt_ms, 1). The function name says "Poisson" but the per-step sample is Bernoulli, not a true Poisson draw. For low rate * dt_ms (< 0.1) the Bernoulli / Poisson distinction is < 5 % — the two distributions agree to first order. For high rate * dt_ms (> 0.5) Bernoulli under-counts because it cannot emit more than one spike per timestep; a true Poisson would.

Same wording issue as PoissonInput in network/stimulus.py. Either rename to bernoulli_encode or replace the < scaled line with rng.poisson(scaled, size) and accept fractional spike counts. Tracked as task #26.

3.2 latency_encode (first-spike-time, FIXED by task #27)

Python
def latency_encode(
    values: npt.ArrayLike,
    T: int,
    tau: float = 5.0,
    strict: bool = True,
) -> np.ndarray:  # shape (T, N), bool

Each value v ∈ [0, 1] produces exactly one spike at timestep int(tau * (1 - v)), clamped to [0, T-1]. Higher value → earlier spike.

Input range guard (strict=True default): the function now raises ValueError when any element of values is outside [0, 1]. The error message reports the offending min/max and suggests strict=False for the legacy silent-clip behaviour. This closes the contract gap that the original docstring claimed but did not enforce.

strict=False keeps the v3.14.0 behaviour: values=1.5 clips to spike-time 0, values=-0.5 clips toward T-1.

tau = 5.0 (default) means the latest possible spike (for v=0) is at timestep 5. For larger T, most timesteps are silent.

3.3 Declared encoders

How events or values become the tensor a network sees changes every result downstream. EventBinning, PoissonRates and FirstSpikeLatency are frozen values whose declaration() states everything they do, and encoder_from_declaration rebuilds the identical encoder from it; digest identifies the declaration.

  • EventBinning(dt_ms, n_steps, width, height, polarity="separate") puts an event at t ms in step floor(t / dt_ms) and drops events at or after n_steps * dt_ms. With "separate" ON and OFF events have their own channels (OFF first); with "merge" they share one. Events outside the sensor, with a polarity other than 0 or 1, or at a negative or non-finite time are refused, not clipped.
  • PoissonRates(n_steps, dt_ms, seed) wraps poisson_encode with its seed.
  • FirstSpikeLatency(n_steps, tau) wraps latency_encode with strict=True.

A declaration the running version would not reproduce exactly — an unknown encoder, another schema, or a changed rule such as late_events — is refused rather than rebuilt approximately.


4. Performance — measured (this workstation)

Hardware: Intel i5-11600K, 32 GB DDR4, Python 3.12.3, NumPy 2.2.6.

4.1 Encoder throughput (mean of 20 calls)

Encoder N T Per-call wall Spike-cells/s
poisson_encode 100 300 0.37 ms 81.1 M
poisson_encode 1 000 300 3.33 ms 90.0 M
poisson_encode 10 000 300 37.38 ms 80.3 M
latency_encode 100 300 0.06 ms —
latency_encode 1 000 300 0.05 ms —
latency_encode 10 000 300 0.63 ms —

poisson_encode is dominated by the rng.random((T, N)) call (uniform draw of T*N floats). Throughput is ~80 M spike-cells/s across all sizes, which matches NumPy's PRNG cost (~10 ns/element).

latency_encode is much faster because it draws no random numbers — just one fancy-indexed write per call. The (T=300, N=10000) call still runs in under 1 ms.

4.2 Synthetic loader cost

Loading N-MNIST in synthetic mode at T = 300, single-threaded:

n_samples Wall Total events generated
10 170.7 ms 515 780
100 1 739.8 ms 5 168 651
500 7 566.2 ms 25 894 833

Linear in n_samples: ~17 ms per sample, ~50 k events per sample. The cost is split between poisson_encode and the per-event np.column_stack + dtype cast inside _synthetic_event_dataset.

4.3 Native recording path and NumPy encoders

N-MNIST recording decoding has Rust, Mojo, Julia and Go implementations as specified above. Indexed SHD reads have Rust and Go implementations using separate native processes. Representative SHD timing comparisons have not yet been recorded; process startup and file I/O belong in that measurement. The historical encoder and synthetic tables above do not qualify SHD throughput. Event encoders and synthetic generation continue to use the NumPy implementations. Decoder-call measurements exclude file I/O and cannot be used to infer full loader, encoder or training performance.


4.4 DVS public reader calls

benchmarks/bench_dvs_recordings.py requires one existing converted NPY recording and all four explicitly configured native executables. It selects each backend explicitly; regular automatic reading still requires one native declaration. It refuses an unavailable backend or any difference in returned float64 storage bits. Example after setting SC_NEUROCORE_DVS_{GO,RUST,JULIA,MOJO}_EXE separately to the actual commands:

Bash
python benchmarks/bench_dvs_recordings.py /operator/recording.npy --corpus-kind operator-converted --repetitions 25 --warmup 3 --output dvs-reader-measurement.json

The timed region includes file I/O, output allocation, process startup and transport. Every Julia call starts a new runtime and includes compilation; warmup warms host caches. Output retains individual samples, affinity, load, source and native artifact SHA-256 values, recording digest and declared provenance. Exact output comparison occurs outside each timed region. These numbers measure reader calls and establish neither training speed nor physical target latency or energy. A generated-format fixture must use that provenance label and cannot qualify publisher acceptance.


5. Pipeline wiring

Surface How it's wired Verifier
from sc_neurocore.datasets import load_nmnist, ... __init__.py:8-9 re-export tests/test_datasets.py
Synthetic fallback path each loader checks synthetic first TestSyntheticLoaders
Real-data path _check_root raises with download URL TestNMNISTRealLoader::test_load_nmnist_real_path, etc.
_synthetic_event_dataset calls poisson_encode loaders.py:56 covered transitively
H5 path imports h5py lazily inside load_shd body works without h5py if synthetic=True

No orphan helpers; _parse_nmnist_bin and _check_root are private but reachable from public loaders.


6. Audit (7-point checklist)

# Dimension Status Detail
1 Pipeline wiring ✅ PASS All 18 symbols wired; loaders → encoders → synthetic fallbacks; manifests → group splits; sc-neurocore dataset
2 Multi-angle tests ✅ PASS 23 tests across 6 classes (TestCheckRoot, TestSyntheticLoaders, TestEncoding, TestNMNISTRealLoader, TestSHDRealLoader, TestDVSCIFAR10RealLoader); covers shape, reproducibility, file-not-found, real-data parse, encoder rate correlation
3 Native recording decoding N-MNIST and indexed SHD N-MNIST: Rust/Mojo/Julia/Go; SHD: Rust/Go; encoders and synthetic paths remain NumPy
4 Benchmarks Historical encoder/synthetic baseline §4.1 + §4.2 do not measure the indexed SHD process path
5 Performance docs Partial SHD representative timing comparisons remain unmeasured
6 Documentation page ✅ PASS This page
7 Rules followed ⚠️ WARN SPDX header on every file ✅. poisson_encode is misnamed — it is Bernoulli, not Poisson (§3.1). latency_encode has an unenforced [0, 1] input contract (§3.2). British English consistent.

The encoder caveats below apply independently of native recording parity. The historical tables do not establish complete loader, SHD process, encoder or training performance.


7. Known issues

7.1 poisson_encode is Bernoulli (task #26)

For low rates this is fine; for high rates it under-counts. Either rename to bernoulli_encode (preferred — the function does not implement what the name claims) or replace the body with an actual Poisson draw and accept fractional spike counts (wider behaviour change).

7.2 latency_encode silently clips out-of-range input (FIXED by task #27)

The function now raises ValueError by default when any value is outside [0, 1]. Pass strict=False to keep the legacy silent-clip behaviour. Regression tests: tests/test_datasets.py::TestLatencyEncodeStrict (5 cases — above-1 raises, negative raises, strict=False keeps clip, boundary values 0.0 / 1.0 accepted, interior values correctly ordered).

7.3 N-MNIST _NMNIST_RES constant is unused on the real path

loaders.py declares _NMNIST_RES = 34 but the real-data parser (_parse_nmnist_bin) decodes coordinates from whole address bytes (0–255) without referring to the constant, which is used only on the synthetic path. A corrupt file can therefore yield addresses beyond the sensor; EventBinning refuses such events rather than placing them.

7.4 load_dvs_cifar10 real path requires .npy not raw

The docstring says "DVS-CIFAR10 event-camera dataset" but the loader expects pre-converted .npy files, not the raw .aedat/.mat released by Li et al. 2017. The error message at line 329-332 makes this clear, but the docstring at line 271-300 does not. Either add a one-line "Note: requires .npy-converted input" to the docstring, or ship a convert_dvs_cifar10_to_npy utility.

7.5 Synthetic SHD differs from real SHD distributionally

_synthetic_shd draws class templates from U(0, 0.1) independently per channel, then Poisson-encodes them. Real SHD has rich temporal structure (cochlear filter banks, formants). The synthetic data produces correct shapes and labels for unit testing but trains a classifier to chance if used for actual learning. Document this constraint in the loader docstring.


8. Tests

Bash
PYTHONPATH=src python3 -m pytest tests/test_datasets.py -q
# 23 passed in 10.04s (verified 2026-04-17)

Coverage breakdown:

  • TestCheckRoot (2): _check_root returns Path on existing dir, raises FileNotFoundError with URL on missing dir.
  • TestSyntheticLoaders (7): synthetic-shape correctness for all 3 loaders, missing-root paths raise even with bad inputs, reproducibility across two same-seed calls.
  • TestEncoding (6): poisson_encode shape + rate correlation + zero/ones edge cases; latency_encode shape + monotonic earlier-fire-for-higher-value.
  • TestNMNISTRealLoader (3): _parse_nmnist_bin decodes a hand-crafted 5-byte event correctly; load_nmnist real path with a synthesised directory tree; missing-split raises.
  • TestSHDRealLoader (2): real-path with synthesised H5 file; missing-h5 raises.
  • TestDVSCIFAR10RealLoader (3): real-path with synthesised npy tree; missing-split raises; empty-dir raises.

Not covered:

  • High-rate Poisson distinction — no test asserts that poisson_encode(rates=1.5, T=10) saturates at 1 spike/step (the Bernoulli ceiling). A test would document the §3.1 issue.
  • Latency input range — no test asserts behaviour for value > 1 or value < 0.
  • Published files — the SHD tests write real HDF5 files in the release layout with h5py, and the N-MNIST tests pack real 40-bit records, but the published archives themselves are not downloaded in CI.

Manifests, splits, declared encoders and the dataset command are held by tests/test_datasets_manifest.py, tests/test_datasets_splits.py, tests/test_datasets_encoders.py and tests/test_cli_dataset.py, on files written in the published formats by tests/event_dataset_support.py.


9. References

Datasets (cited by source):

  • Orchard G. et al. "Converting Static Image Datasets to Spiking Neuromorphic Datasets Using Saccades." Front Neurosci 9:437 (2015). N-MNIST.
  • Cramer B., Stradmann Y., Schemmel J., Zenke F. "The Heidelberg Spiking Data Sets for the Systematic Evaluation of Spiking Neural Networks." IEEE TNNLS 33(7):2744-2757 (2022). SHD.
  • Li H. et al. "CIFAR10-DVS: An Event-Stream Dataset for Object Classification." Front Neurosci 11:309 (2017). DVS-CIFAR10.

Data licences, as the publishers state them: N-MNIST under CC BY-SA 4.0 (garrickorchard.com/datasets/n-mnist), SHD under CC BY 4.0 (zenkelab.org), CIFAR10-DVS under CC BY 4.0 (figshare 4724671, version 2).

Encoders (background):

  • Gerstner W., Kistler W. M. Spiking Neuron Models: Single Neurons, Populations, Plasticity. Cambridge UP (2002). Chapters on rate vs latency coding.
  • Thorpe S., Fize D., Marlot C. "Speed of processing in the human visual system." Nature 381:520-522 (1996). The original motivation for first-spike-time / latency coding.

Internal:


10. Auto-rendered API

sc_neurocore.datasets

Expose event-dataset loaders, manifests, group splits and declared encoders.

EventBinning dataclass

Bin camera events into a binary spike tensor.

An event at t ms lands in step floor(t / dt_ms); events at or after n_steps * dt_ms are dropped, never merged into the last step. With polarity="separate" ON and OFF events have their own channels (OFF first), with "merge" they share one. An event outside the sensor or before time zero is refused rather than clipped.

Attributes

dt_ms: Step length in milliseconds. n_steps: Number of steps in the window. width, height: Sensor geometry in pixels. polarity: "separate" or "merge".

Source code in src/sc_neurocore/datasets/encoders.py
Python
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
@dataclass(frozen=True, slots=True)
class EventBinning:
    """Bin camera events into a binary spike tensor.

    An event at ``t`` ms lands in step ``floor(t / dt_ms)``; events at or after
    ``n_steps * dt_ms`` are dropped, never merged into the last step. With
    ``polarity="separate"`` ON and OFF events have their own channels (OFF
    first), with ``"merge"`` they share one. An event outside the sensor or
    before time zero is refused rather than clipped.

    Attributes
    ----------
    dt_ms:
        Step length in milliseconds.
    n_steps:
        Number of steps in the window.
    width, height:
        Sensor geometry in pixels.
    polarity:
        ``"separate"`` or ``"merge"``.
    """

    dt_ms: float
    n_steps: int
    width: int
    height: int
    polarity: Literal["separate", "merge"] = "separate"

    def __post_init__(self) -> None:
        """Refuse a setting the declaration could not state faithfully."""
        _positive_finite("dt_ms", self.dt_ms)
        _positive_int("n_steps", self.n_steps)
        _positive_int("width", self.width)
        _positive_int("height", self.height)
        if self.polarity not in ("separate", "merge"):
            raise ValueError(f"polarity must be 'separate' or 'merge'; got {self.polarity!r}")

    @property
    def channels(self) -> int:
        """Channels per step: pixels, twice over when polarities are separate."""
        pixels = self.width * self.height
        return 2 * pixels if self.polarity == "separate" else pixels

    def declaration(self) -> dict[str, Any]:
        """Return the full description, enough to rebuild this encoder."""
        return {
            "schema": ENCODER_SCHEMA,
            "encoder": "event-binning",
            "dt_ms": self.dt_ms,
            "n_steps": self.n_steps,
            "width": self.width,
            "height": self.height,
            "polarity": self.polarity,
            "step": "floor(t_ms / dt_ms)",
            "late_events": "dropped",
            "channel": "polarity * width * height + y * width + x"
            if self.polarity == "separate"
            else "y * width + x",
            "output": "bool (n_steps, channels)",
        }

    @property
    def digest(self) -> str:
        """``sha256:`` over the declaration."""
        return _digest(self.declaration())

    def encode(self, events: npt.ArrayLike) -> np.ndarray[Any, Any]:
        """Bin ``(N, 4)`` events with columns ``x, y, polarity, t_ms``.

        Parameters
        ----------
        events:
            Events as the event loaders return them.

        Returns
        -------
        numpy.ndarray
            ``bool`` array of shape ``(n_steps, channels)``.

        Raises
        ------
        ValueError
            On a malformed array, an event outside the sensor, a polarity
            other than 0 or 1, or a negative or non-finite time.
        """
        array = np.asarray(events, dtype=np.float64)
        if array.ndim != 2 or array.shape[1] != 4:
            raise ValueError(f"events must have shape (N, 4); got {array.shape}")
        spikes = np.zeros((self.n_steps, self.channels), dtype=bool)
        if array.shape[0] == 0:
            return spikes
        x, y, polarity, time_ms = array.T
        if not np.all(np.isfinite(array)):
            raise ValueError("events hold a non-finite value")
        if np.any(x != np.floor(x)) or np.any(y != np.floor(y)):
            raise ValueError("pixel addresses must be whole numbers")
        if np.any((x < 0) | (x >= self.width) | (y < 0) | (y >= self.height)):
            raise ValueError(f"an event lies outside the {self.width} x {self.height} sensor")
        if np.any((polarity != 0) & (polarity != 1)):
            raise ValueError("polarity must be 0 or 1")
        if np.any(time_ms < 0):
            raise ValueError("an event has a negative time")
        step = np.floor(time_ms / self.dt_ms).astype(np.int64)
        inside = step < self.n_steps
        channel = y.astype(np.int64) * self.width + x.astype(np.int64)
        if self.polarity == "separate":
            channel = channel + polarity.astype(np.int64) * self.width * self.height
        spikes[step[inside], channel[inside]] = True
        return spikes

channels property

Channels per step: pixels, twice over when polarities are separate.

digest property

sha256: over the declaration.

__post_init__()

Refuse a setting the declaration could not state faithfully.

Source code in src/sc_neurocore/datasets/encoders.py
Python
79
80
81
82
83
84
85
86
def __post_init__(self) -> None:
    """Refuse a setting the declaration could not state faithfully."""
    _positive_finite("dt_ms", self.dt_ms)
    _positive_int("n_steps", self.n_steps)
    _positive_int("width", self.width)
    _positive_int("height", self.height)
    if self.polarity not in ("separate", "merge"):
        raise ValueError(f"polarity must be 'separate' or 'merge'; got {self.polarity!r}")

declaration()

Return the full description, enough to rebuild this encoder.

Source code in src/sc_neurocore/datasets/encoders.py
Python
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
def declaration(self) -> dict[str, Any]:
    """Return the full description, enough to rebuild this encoder."""
    return {
        "schema": ENCODER_SCHEMA,
        "encoder": "event-binning",
        "dt_ms": self.dt_ms,
        "n_steps": self.n_steps,
        "width": self.width,
        "height": self.height,
        "polarity": self.polarity,
        "step": "floor(t_ms / dt_ms)",
        "late_events": "dropped",
        "channel": "polarity * width * height + y * width + x"
        if self.polarity == "separate"
        else "y * width + x",
        "output": "bool (n_steps, channels)",
    }

encode(events)

Bin (N, 4) events with columns x, y, polarity, t_ms.

Parameters

events: Events as the event loaders return them.

Returns

numpy.ndarray bool array of shape (n_steps, channels).

Raises

ValueError On a malformed array, an event outside the sensor, a polarity other than 0 or 1, or a negative or non-finite time.

Source code in src/sc_neurocore/datasets/encoders.py
Python
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
def encode(self, events: npt.ArrayLike) -> np.ndarray[Any, Any]:
    """Bin ``(N, 4)`` events with columns ``x, y, polarity, t_ms``.

    Parameters
    ----------
    events:
        Events as the event loaders return them.

    Returns
    -------
    numpy.ndarray
        ``bool`` array of shape ``(n_steps, channels)``.

    Raises
    ------
    ValueError
        On a malformed array, an event outside the sensor, a polarity
        other than 0 or 1, or a negative or non-finite time.
    """
    array = np.asarray(events, dtype=np.float64)
    if array.ndim != 2 or array.shape[1] != 4:
        raise ValueError(f"events must have shape (N, 4); got {array.shape}")
    spikes = np.zeros((self.n_steps, self.channels), dtype=bool)
    if array.shape[0] == 0:
        return spikes
    x, y, polarity, time_ms = array.T
    if not np.all(np.isfinite(array)):
        raise ValueError("events hold a non-finite value")
    if np.any(x != np.floor(x)) or np.any(y != np.floor(y)):
        raise ValueError("pixel addresses must be whole numbers")
    if np.any((x < 0) | (x >= self.width) | (y < 0) | (y >= self.height)):
        raise ValueError(f"an event lies outside the {self.width} x {self.height} sensor")
    if np.any((polarity != 0) & (polarity != 1)):
        raise ValueError("polarity must be 0 or 1")
    if np.any(time_ms < 0):
        raise ValueError("an event has a negative time")
    step = np.floor(time_ms / self.dt_ms).astype(np.int64)
    inside = step < self.n_steps
    channel = y.astype(np.int64) * self.width + x.astype(np.int64)
    if self.polarity == "separate":
        channel = channel + polarity.astype(np.int64) * self.width * self.height
    spikes[step[inside], channel[inside]] = True
    return spikes

PoissonRates dataclass

Encode per-step firing probabilities as seeded Bernoulli spike trains.

Attributes

n_steps: Number of steps. dt_ms: Step length; the probability per step is rate * dt_ms, clipped to [0, 1]. seed: Generator seed; the same seed and input give the same spikes.

Source code in src/sc_neurocore/datasets/encoders.py
Python
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
@dataclass(frozen=True, slots=True)
class PoissonRates:
    """Encode per-step firing probabilities as seeded Bernoulli spike trains.

    Attributes
    ----------
    n_steps:
        Number of steps.
    dt_ms:
        Step length; the probability per step is ``rate * dt_ms``, clipped
        to ``[0, 1]``.
    seed:
        Generator seed; the same seed and input give the same spikes.
    """

    n_steps: int
    dt_ms: float = 1.0
    seed: int = 0

    def __post_init__(self) -> None:
        """Refuse a setting the declaration could not state faithfully."""
        _positive_int("n_steps", self.n_steps)
        _positive_finite("dt_ms", self.dt_ms)

    def declaration(self) -> dict[str, Any]:
        """Return the full description, enough to rebuild this encoder."""
        return {
            "schema": ENCODER_SCHEMA,
            "encoder": "poisson-rates",
            "n_steps": self.n_steps,
            "dt_ms": self.dt_ms,
            "seed": self.seed,
            "generator": "numpy.random.default_rng(seed).random((n_steps, N))",
            "probability": "clip(rate * dt_ms, 0, 1)",
            "output": "bool (n_steps, N)",
        }

    @property
    def digest(self) -> str:
        """``sha256:`` over the declaration."""
        return _digest(self.declaration())

    def encode(self, rates: npt.ArrayLike) -> np.ndarray[Any, Any]:
        """Return the spike trains for a vector of rates."""
        return poisson_encode(rates, self.n_steps, dt_ms=self.dt_ms, seed=self.seed)

digest property

sha256: over the declaration.

__post_init__()

Refuse a setting the declaration could not state faithfully.

Source code in src/sc_neurocore/datasets/encoders.py
Python
181
182
183
184
def __post_init__(self) -> None:
    """Refuse a setting the declaration could not state faithfully."""
    _positive_int("n_steps", self.n_steps)
    _positive_finite("dt_ms", self.dt_ms)

declaration()

Return the full description, enough to rebuild this encoder.

Source code in src/sc_neurocore/datasets/encoders.py
Python
186
187
188
189
190
191
192
193
194
195
196
197
def declaration(self) -> dict[str, Any]:
    """Return the full description, enough to rebuild this encoder."""
    return {
        "schema": ENCODER_SCHEMA,
        "encoder": "poisson-rates",
        "n_steps": self.n_steps,
        "dt_ms": self.dt_ms,
        "seed": self.seed,
        "generator": "numpy.random.default_rng(seed).random((n_steps, N))",
        "probability": "clip(rate * dt_ms, 0, 1)",
        "output": "bool (n_steps, N)",
    }

encode(rates)

Return the spike trains for a vector of rates.

Source code in src/sc_neurocore/datasets/encoders.py
Python
204
205
206
def encode(self, rates: npt.ArrayLike) -> np.ndarray[Any, Any]:
    """Return the spike trains for a vector of rates."""
    return poisson_encode(rates, self.n_steps, dt_ms=self.dt_ms, seed=self.seed)

FirstSpikeLatency dataclass

Encode values in [0, 1] as one spike each, larger values earlier.

Attributes

n_steps: Number of steps. tau: The spike of value v falls in step int(tau * (1 - v)), limited to the window.

Source code in src/sc_neurocore/datasets/encoders.py
Python
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
@dataclass(frozen=True, slots=True)
class FirstSpikeLatency:
    """Encode values in ``[0, 1]`` as one spike each, larger values earlier.

    Attributes
    ----------
    n_steps:
        Number of steps.
    tau:
        The spike of value ``v`` falls in step ``int(tau * (1 - v))``,
        limited to the window.
    """

    n_steps: int
    tau: float = 5.0

    def __post_init__(self) -> None:
        """Refuse a setting the declaration could not state faithfully."""
        _positive_int("n_steps", self.n_steps)
        _positive_finite("tau", self.tau)

    def declaration(self) -> dict[str, Any]:
        """Return the full description, enough to rebuild this encoder."""
        return {
            "schema": ENCODER_SCHEMA,
            "encoder": "first-spike-latency",
            "n_steps": self.n_steps,
            "tau": self.tau,
            "step": "int(clip(tau * (1 - v), 0, n_steps - 1))",
            "values_outside_unit_interval": "refused",
            "output": "bool (n_steps, N)",
        }

    @property
    def digest(self) -> str:
        """``sha256:`` over the declaration."""
        return _digest(self.declaration())

    def encode(self, values: npt.ArrayLike) -> np.ndarray[Any, Any]:
        """Return one spike per value; a value outside ``[0, 1]`` is refused."""
        return latency_encode(values, self.n_steps, tau=self.tau, strict=True)

digest property

sha256: over the declaration.

__post_init__()

Refuse a setting the declaration could not state faithfully.

Source code in src/sc_neurocore/datasets/encoders.py
Python
225
226
227
228
def __post_init__(self) -> None:
    """Refuse a setting the declaration could not state faithfully."""
    _positive_int("n_steps", self.n_steps)
    _positive_finite("tau", self.tau)

declaration()

Return the full description, enough to rebuild this encoder.

Source code in src/sc_neurocore/datasets/encoders.py
Python
230
231
232
233
234
235
236
237
238
239
240
def declaration(self) -> dict[str, Any]:
    """Return the full description, enough to rebuild this encoder."""
    return {
        "schema": ENCODER_SCHEMA,
        "encoder": "first-spike-latency",
        "n_steps": self.n_steps,
        "tau": self.tau,
        "step": "int(clip(tau * (1 - v), 0, n_steps - 1))",
        "values_outside_unit_interval": "refused",
        "output": "bool (n_steps, N)",
    }

encode(values)

Return one spike per value; a value outside [0, 1] is refused.

Source code in src/sc_neurocore/datasets/encoders.py
Python
247
248
249
def encode(self, values: npt.ArrayLike) -> np.ndarray[Any, Any]:
    """Return one spike per value; a value outside ``[0, 1]`` is refused."""
    return latency_encode(values, self.n_steps, tau=self.tau, strict=True)

EventDatasetManifest dataclass

Every file and sample of one event dataset as it lies on disk.

Attributes

dataset: The dataset description. version: The release of the data the user holds, as its publisher names it. files: Every file read, sorted by path. samples: Every sample, in loader order.

Source code in src/sc_neurocore/datasets/manifest.py
Python
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
@dataclass(frozen=True, slots=True)
class EventDatasetManifest:
    """Every file and sample of one event dataset as it lies on disk.

    Attributes
    ----------
    dataset:
        The dataset description.
    version:
        The release of the data the user holds, as its publisher names it.
    files:
        Every file read, sorted by path.
    samples:
        Every sample, in loader order.
    """

    dataset: DatasetDescription
    version: str
    files: tuple[FileRecord, ...]
    samples: tuple[SampleRecord, ...]

    def to_dict(self) -> dict[str, Any]:
        """Return the JSON form, with the schema identifier."""
        return {
            "schema": MANIFEST_SCHEMA,
            "dataset": self.dataset.to_dict(),
            "version": self.version,
            "files": [
                {"path": record.path, "bytes": record.bytes, "sha256": record.sha256}
                for record in self.files
            ],
            "samples": [
                {
                    "split": sample.split,
                    "file": sample.file,
                    "index": sample.index,
                    "label": sample.label,
                    "group": sample.group,
                }
                for sample in self.samples
            ],
        }

    @property
    def digest(self) -> str:
        """``sha256:`` over the canonical JSON form; identifies this exact manifest."""
        canonical = json.dumps(self.to_dict(), sort_keys=True, separators=(",", ":"))
        return "sha256:" + hashlib.sha256(canonical.encode("utf-8")).hexdigest()

    def splits(self) -> tuple[str, ...]:
        """Return the published split names, in first-seen order."""
        return tuple(dict.fromkeys(sample.split for sample in self.samples))

digest property

sha256: over the canonical JSON form; identifies this exact manifest.

to_dict()

Return the JSON form, with the schema identifier.

Source code in src/sc_neurocore/datasets/manifest.py
Python
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
def to_dict(self) -> dict[str, Any]:
    """Return the JSON form, with the schema identifier."""
    return {
        "schema": MANIFEST_SCHEMA,
        "dataset": self.dataset.to_dict(),
        "version": self.version,
        "files": [
            {"path": record.path, "bytes": record.bytes, "sha256": record.sha256}
            for record in self.files
        ],
        "samples": [
            {
                "split": sample.split,
                "file": sample.file,
                "index": sample.index,
                "label": sample.label,
                "group": sample.group,
            }
            for sample in self.samples
        ],
    }

splits()

Return the published split names, in first-seen order.

Source code in src/sc_neurocore/datasets/manifest.py
Python
249
250
251
def splits(self) -> tuple[str, ...]:
    """Return the published split names, in first-seen order."""
    return tuple(dict.fromkeys(sample.split for sample in self.samples))

SplitPlan dataclass

Which samples of a manifest go to which split, by whole groups.

Attributes

manifest_digest: Digest of the manifest the positions refer to. source_split: The published split that was divided. seed: Seed of the group order. fractions: Requested share of samples per split, in declaration order. assignment: Split name to the positions of its samples in manifest.samples. groups: Split name to the groups it holds.

Source code in src/sc_neurocore/datasets/splits.py
Python
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
@dataclass(frozen=True, slots=True)
class SplitPlan:
    """Which samples of a manifest go to which split, by whole groups.

    Attributes
    ----------
    manifest_digest:
        Digest of the manifest the positions refer to.
    source_split:
        The published split that was divided.
    seed:
        Seed of the group order.
    fractions:
        Requested share of samples per split, in declaration order.
    assignment:
        Split name to the positions of its samples in ``manifest.samples``.
    groups:
        Split name to the groups it holds.
    """

    manifest_digest: str
    source_split: str
    seed: int
    fractions: tuple[tuple[str, float], ...]
    assignment: Mapping[str, tuple[int, ...]]
    groups: Mapping[str, tuple[str, ...]]

    def to_dict(self) -> dict[str, Any]:
        """Return the JSON form, with the schema identifier."""
        return {
            "schema": SPLIT_SCHEMA,
            "manifest_digest": self.manifest_digest,
            "source_split": self.source_split,
            "seed": self.seed,
            "fractions": [[name, share] for name, share in self.fractions],
            "assignment": {name: list(positions) for name, positions in self.assignment.items()},
            "groups": {name: list(groups) for name, groups in self.groups.items()},
        }

    @property
    def digest(self) -> str:
        """``sha256:`` over the canonical JSON form."""
        canonical = json.dumps(self.to_dict(), sort_keys=True, separators=(",", ":"))
        return "sha256:" + hashlib.sha256(canonical.encode("utf-8")).hexdigest()

digest property

sha256: over the canonical JSON form.

to_dict()

Return the JSON form, with the schema identifier.

Source code in src/sc_neurocore/datasets/splits.py
Python
61
62
63
64
65
66
67
68
69
70
71
def to_dict(self) -> dict[str, Any]:
    """Return the JSON form, with the schema identifier."""
    return {
        "schema": SPLIT_SCHEMA,
        "manifest_digest": self.manifest_digest,
        "source_split": self.source_split,
        "seed": self.seed,
        "fractions": [[name, share] for name, share in self.fractions],
        "assignment": {name: list(positions) for name, positions in self.assignment.items()},
        "groups": {name: list(groups) for name, groups in self.groups.items()},
    }

load_nmnist(root='data/nmnist', train=True, dt_ms=1.0, T=300, synthetic=False, n_samples=100, seed=42)

Load N-MNIST spiking vision dataset.

Neuromorphic-MNIST: 34x34 DVS recordings of MNIST digits moved on an ATIS sensor via saccadic eye movements. 10 classes.

Orchard et al., "Converting Static Image Datasets to Spiking Neuromorphic Datasets Using Saccades", Front. Neurosci. 2015.

Parameters

root : path Directory containing the extracted dataset. train : bool Load training split if True, test split otherwise. dt_ms : float Temporal resolution for synthetic fallback. T : int Number of timesteps for synthetic fallback. synthetic : bool Force synthetic data generation. n_samples : int Number of synthetic samples to generate. seed : int RNG seed for reproducible synthetic data.

Returns

samples : list of ndarray, each shape (N_events, 4) Real recordings use float64 columns [x, y, polarity, timestamp_ms] to avoid float32 timestamp rounding before temporal binning. labels : ndarray of int

Source code in src/sc_neurocore/datasets/loaders.py
Python
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
def load_nmnist(
    root: str | Path = "data/nmnist",
    train: bool = True,
    dt_ms: float = 1.0,
    T: int = 300,
    synthetic: bool = False,
    n_samples: int = 100,
    seed: int = 42,
) -> tuple[list[np.ndarray[Any, Any]], np.ndarray[Any, Any]]:
    """Load N-MNIST spiking vision dataset.

    Neuromorphic-MNIST: 34x34 DVS recordings of MNIST digits moved on
    an ATIS sensor via saccadic eye movements. 10 classes.

    Orchard et al., "Converting Static Image Datasets to Spiking
    Neuromorphic Datasets Using Saccades", Front. Neurosci. 2015.

    Parameters
    ----------
    root : path
        Directory containing the extracted dataset.
    train : bool
        Load training split if True, test split otherwise.
    dt_ms : float
        Temporal resolution for synthetic fallback.
    T : int
        Number of timesteps for synthetic fallback.
    synthetic : bool
        Force synthetic data generation.
    n_samples : int
        Number of synthetic samples to generate.
    seed : int
        RNG seed for reproducible synthetic data.

    Returns
    -------
    samples : list of ndarray, each shape (N_events, 4)
        Real recordings use float64 columns [x, y, polarity, timestamp_ms]
        to avoid float32 timestamp rounding before temporal binning.
    labels : ndarray of int
    """
    if synthetic:
        return _synthetic_event_dataset(
            n_samples,
            _NMNIST_RES,
            10,
            T,
            dt_ms,
            seed,
        )
    _check_root(root, "N-MNIST", _NMNIST_URL)
    split_dir = Path(root) / ("Train" if train else "Test")
    if not split_dir.exists():
        raise FileNotFoundError(
            f"Expected split directory {split_dir.resolve()}. Download from {_NMNIST_URL}"
        )
    # Real loader: N-MNIST uses .bin files, one per sample, grouped by class
    samples: list[np.ndarray[Any, Any]] = []
    label_list: list[int] = []
    for class_dir in sorted(split_dir.iterdir()):
        if not class_dir.is_dir():
            continue
        class_label = int(class_dir.name)
        for bin_file in sorted(class_dir.glob("*.bin")):
            events = _parse_nmnist_bin(bin_file)
            samples.append(events)
            label_list.append(class_label)
    return samples, np.array(label_list, dtype=np.int64)

load_shd(root='data/shd', train=True, dt_ms=1.0, T=1000, synthetic=False, n_samples=100, seed=42)

Load Spiking Heidelberg Digits (SHD) dataset.

Audio digits 0-9 in English and German, spike-encoded through an artificial cochlea model. 700 input channels, 20 classes.

Cramer et al., "The Heidelberg Spiking Data Sets for the Systematic Evaluation of Spiking Neural Networks", IEEE TNNLS 2022.

Parameters

root : path Directory containing shd_train.h5 / shd_test.h5. train : bool Load training split if True, test split otherwise. dt_ms : float Temporal resolution for binning spikes. T : int Number of timesteps for synthetic fallback. synthetic : bool Force synthetic data generation. n_samples : int Number of synthetic samples to generate. seed : int RNG seed for reproducible synthetic data.

Returns

samples : list of ndarray, each shape (T, 700) dtype bool Binned spike trains. labels : ndarray of int

Source code in src/sc_neurocore/datasets/loaders.py
Python
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
def load_shd(
    root: str | Path = "data/shd",
    train: bool = True,
    dt_ms: float = 1.0,
    T: int = 1000,
    synthetic: bool = False,
    n_samples: int = 100,
    seed: int = 42,
) -> tuple[list[np.ndarray[Any, Any]], np.ndarray[Any, Any]]:
    """Load Spiking Heidelberg Digits (SHD) dataset.

    Audio digits 0-9 in English and German, spike-encoded through an
    artificial cochlea model. 700 input channels, 20 classes.

    Cramer et al., "The Heidelberg Spiking Data Sets for the Systematic
    Evaluation of Spiking Neural Networks", IEEE TNNLS 2022.

    Parameters
    ----------
    root : path
        Directory containing shd_train.h5 / shd_test.h5.
    train : bool
        Load training split if True, test split otherwise.
    dt_ms : float
        Temporal resolution for binning spikes.
    T : int
        Number of timesteps for synthetic fallback.
    synthetic : bool
        Force synthetic data generation.
    n_samples : int
        Number of synthetic samples to generate.
    seed : int
        RNG seed for reproducible synthetic data.

    Returns
    -------
    samples : list of ndarray, each shape (T, 700) dtype bool
        Binned spike trains.
    labels : ndarray of int
    """
    if synthetic:
        return _synthetic_shd(n_samples, T, dt_ms, seed)

    _check_root(root, "SHD", _SHD_URL)
    fname = "shd_train.h5" if train else "shd_test.h5"
    h5_path = Path(root) / fname
    if not h5_path.exists():
        raise FileNotFoundError(f"{h5_path.resolve()} not found. Download from {_SHD_URL}")
    import h5py

    samples: list[np.ndarray[Any, Any]] = []
    with h5py.File(h5_path, "r") as f:
        spike_times = f["spikes"]["times"]
        spike_units = f["spikes"]["units"]
        raw_labels = f["labels"][:]
        for i in range(len(raw_labels)):
            times = np.asarray(spike_times[i], dtype=np.float64) * 1000.0
            units = np.asarray(spike_units[i])
            if len(times) > 0:
                n_bins = min(int(np.ceil(times.max() / dt_ms)) + 1, T)
            else:
                n_bins = T
            train_arr = np.zeros((n_bins, _SHD_CHANNELS), dtype=bool)
            if len(times) > 0:
                bin_idx = (times / dt_ms).astype(int)
                unit_idx = np.clip(units.astype(int), 0, _SHD_CHANNELS - 1)
                # A spike after the T-step window is dropped: merging it into the
                # last bin would invent a burst at the end of every long sample.
                inside = bin_idx < n_bins
                train_arr[bin_idx[inside], unit_idx[inside]] = True
            samples.append(train_arr)

    return samples, raw_labels.astype(np.int64)

load_dvs_cifar10(root='data/dvs_cifar10', train=True, dt_ms=1.0, T=300, synthetic=False, n_samples=100, seed=42)

Load DVS-CIFAR10 event-camera dataset.

CIFAR-10 images displayed on a monitor and recorded by a DVS camera at 128x128 resolution. 10 classes.

Li et al., "CIFAR10-DVS: An Event-Stream Dataset for Object Classification", Front. Neurosci. 2017.

Parameters

root : path Directory containing the extracted dataset. train : bool Load training split if True, test split otherwise. dt_ms : float Temporal resolution for synthetic fallback. T : int Number of timesteps for synthetic fallback. synthetic : bool Force synthetic data generation. n_samples : int Number of synthetic samples to generate. seed : int RNG seed for reproducible synthetic data.

Returns

samples : list of ndarray, each shape (N_events, 4) Real recordings use float64 columns [x, y, polarity, timestamp_ms] to avoid float32 timestamp rounding before temporal binning. labels : ndarray of int

Source code in src/sc_neurocore/datasets/loaders.py
Python
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
def load_dvs_cifar10(
    root: str | Path = "data/dvs_cifar10",
    train: bool = True,
    dt_ms: float = 1.0,
    T: int = 300,
    synthetic: bool = False,
    n_samples: int = 100,
    seed: int = 42,
) -> tuple[list[np.ndarray[Any, Any]], np.ndarray[Any, Any]]:
    """Load DVS-CIFAR10 event-camera dataset.

    CIFAR-10 images displayed on a monitor and recorded by a DVS camera
    at 128x128 resolution. 10 classes.

    Li et al., "CIFAR10-DVS: An Event-Stream Dataset for Object
    Classification", Front. Neurosci. 2017.

    Parameters
    ----------
    root : path
        Directory containing the extracted dataset.
    train : bool
        Load training split if True, test split otherwise.
    dt_ms : float
        Temporal resolution for synthetic fallback.
    T : int
        Number of timesteps for synthetic fallback.
    synthetic : bool
        Force synthetic data generation.
    n_samples : int
        Number of synthetic samples to generate.
    seed : int
        RNG seed for reproducible synthetic data.

    Returns
    -------
    samples : list of ndarray, each shape (N_events, 4)
        Real recordings use float64 columns [x, y, polarity, timestamp_ms]
        to avoid float32 timestamp rounding before temporal binning.
    labels : ndarray of int
    """
    if synthetic:
        return _synthetic_event_dataset(
            n_samples,
            _DVS_CIFAR10_RES,
            10,
            T,
            dt_ms,
            seed,
        )
    _check_root(root, "DVS-CIFAR10", _DVS_CIFAR10_URL)
    split_dir = Path(root) / ("train" if train else "test")
    if not split_dir.exists():
        raise FileNotFoundError(
            f"Expected split directory {split_dir.resolve()}. Download from {_DVS_CIFAR10_URL}"
        )
    # User-converted .npy recordings are grouped by class.
    samples: list[np.ndarray[Any, Any]] = []
    label_list: list[int] = []
    for class_dir in sorted(split_dir.iterdir()):
        if not class_dir.is_dir():
            continue
        class_label = int(class_dir.name)
        for event_file in sorted(class_dir.glob("*.npy")):
            events = _parse_dvs_npy(event_file)
            samples.append(events)
            label_list.append(class_label)
    if not samples:
        raise FileNotFoundError(
            f"No .npy event files found in {split_dir.resolve()}. "
            f"Convert raw data to .npy arrays with columns [x, y, pol, ts_ms]."
        )
    return samples, np.array(label_list, dtype=np.int64)

poisson_encode(rates, T, dt_ms=1.0, seed=None)

Convert firing-rate array to Poisson spike trains.

Parameters

rates : array_like, shape (N,) Firing probabilities per timestep, clipped to [0, 1]. T : int Number of timesteps. dt_ms : float Timestep duration in ms (scales rates linearly). seed : int or None RNG seed for reproducibility.

Returns

spikes : ndarray, shape (T, N), dtype bool

Source code in src/sc_neurocore/datasets/encoding.py
Python
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
def poisson_encode(
    rates: npt.ArrayLike,
    T: int,
    dt_ms: float = 1.0,
    seed: int | None = None,
) -> np.ndarray[Any, Any]:
    """Convert firing-rate array to Poisson spike trains.

    Parameters
    ----------
    rates : array_like, shape (N,)
        Firing probabilities per timestep, clipped to [0, 1].
    T : int
        Number of timesteps.
    dt_ms : float
        Timestep duration in ms (scales rates linearly).
    seed : int or None
        RNG seed for reproducibility.

    Returns
    -------
    spikes : ndarray, shape (T, N), dtype bool
    """
    rng = np.random.default_rng(seed)
    rates = np.asarray(rates, dtype=np.float64)
    scaled = np.clip(rates * (dt_ms / 1.0), 0.0, 1.0)
    return rng.random((T, rates.shape[0])) < scaled

latency_encode(values, T, tau=5.0, strict=True)

Convert normalised values in [0, 1] to first-spike-time trains.

Higher values spike earlier. Each neuron fires exactly once.

Parameters

values : array_like, shape (N,) Input values, expected in [0, 1]. T : int Number of timesteps. tau : float Time constant controlling the spike-time spread. strict : bool If True (default), raise ValueError when any value lies outside [0, 1]. If False, silently clip the resulting spike times to [0, T-1] (the legacy behaviour). The clip happens regardless of strict; this flag controls only whether the function raises before clipping.

Returns

spikes : ndarray, shape (T, N), dtype bool

Raises

ValueError If strict=True (default) and any element of values is outside [0, 1].

Source code in src/sc_neurocore/datasets/encoding.py
Python
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
def latency_encode(
    values: npt.ArrayLike,
    T: int,
    tau: float = 5.0,
    strict: bool = True,
) -> np.ndarray[Any, Any]:
    """Convert normalised values in [0, 1] to first-spike-time trains.

    Higher values spike earlier. Each neuron fires exactly once.

    Parameters
    ----------
    values : array_like, shape (N,)
        Input values, expected in ``[0, 1]``.
    T : int
        Number of timesteps.
    tau : float
        Time constant controlling the spike-time spread.
    strict : bool
        If True (default), raise ``ValueError`` when any value lies
        outside ``[0, 1]``. If False, silently clip the resulting
        spike times to ``[0, T-1]`` (the legacy behaviour). The
        clip happens regardless of ``strict``; this flag controls
        only whether the function raises before clipping.

    Returns
    -------
    spikes : ndarray, shape (T, N), dtype bool

    Raises
    ------
    ValueError
        If ``strict=True`` (default) and any element of ``values``
        is outside ``[0, 1]``.
    """
    values = np.asarray(values, dtype=np.float64)
    if strict and (values.min() < 0.0 or values.max() > 1.0):
        bad_min = float(values.min())
        bad_max = float(values.max())
        raise ValueError(
            f"latency_encode: values must be in [0, 1] when strict=True; "
            f"got min={bad_min}, max={bad_max}. Pass strict=False to "
            f"accept the legacy silent-clip behaviour."
        )
    # spike_time = tau * (1 - value); higher value => earlier spike
    spike_times = np.clip(tau * (1.0 - values), 0, T - 1).astype(int)
    spikes = np.zeros((T, values.shape[0]), dtype=bool)
    neuron_idx = np.arange(values.shape[0])
    spikes[spike_times, neuron_idx] = True
    return spikes

encoder_from_declaration(declaration)

Rebuild the encoder a declaration describes.

Parameters

declaration: A declaration as :meth:EventBinning.declaration and its siblings return it.

Returns

EventBinning or PoissonRates or FirstSpikeLatency The encoder; its own declaration equals the one given.

Raises

ValueError On another schema, an unknown encoder, or a declaration that is not exactly what the rebuilt encoder declares.

Source code in src/sc_neurocore/datasets/encoders.py
Python
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
def encoder_from_declaration(declaration: dict[str, Any]) -> InputEncoder:
    """Rebuild the encoder a declaration describes.

    Parameters
    ----------
    declaration:
        A declaration as :meth:`EventBinning.declaration` and its siblings
        return it.

    Returns
    -------
    EventBinning or PoissonRates or FirstSpikeLatency
        The encoder; its own declaration equals the one given.

    Raises
    ------
    ValueError
        On another schema, an unknown encoder, or a declaration that is not
        exactly what the rebuilt encoder declares.
    """
    if declaration.get("schema") != ENCODER_SCHEMA:
        raise ValueError(f"encoder schema {declaration.get('schema')!r} is not {ENCODER_SCHEMA!r}")
    kind = declaration.get("encoder")
    encoder: InputEncoder
    if kind == "event-binning":
        encoder = EventBinning(
            dt_ms=float(declaration["dt_ms"]),
            n_steps=int(declaration["n_steps"]),
            width=int(declaration["width"]),
            height=int(declaration["height"]),
            polarity=declaration["polarity"],
        )
    elif kind == "poisson-rates":
        encoder = PoissonRates(
            n_steps=int(declaration["n_steps"]),
            dt_ms=float(declaration["dt_ms"]),
            seed=int(declaration["seed"]),
        )
    elif kind == "first-spike-latency":
        encoder = FirstSpikeLatency(
            n_steps=int(declaration["n_steps"]), tau=float(declaration["tau"])
        )
    else:
        raise ValueError(f"unknown encoder {kind!r}")
    if encoder.declaration() != declaration:
        raise ValueError(
            "the declaration does not match what this version of the encoder does; "
            "it cannot be rebuilt faithfully"
        )
    return encoder

build_manifest(name, root, *, version)

Scan a dataset directory and record its files and samples.

Parameters

name: "nmnist", "shd" or "dvs_cifar10". root: Directory in the layout the matching loader reads. version: The release the files come from, as the publisher names it; it is recorded, not inferred, because the files do not carry it.

Returns

EventDatasetManifest The manifest.

Raises

ValueError On an unknown dataset, an empty version, a directory holding none of the dataset's files, or a label outside the dataset's classes.

Source code in src/sc_neurocore/datasets/manifest.py
Python
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
def build_manifest(name: str, root: str | Path, *, version: str) -> EventDatasetManifest:
    """Scan a dataset directory and record its files and samples.

    Parameters
    ----------
    name:
        ``"nmnist"``, ``"shd"`` or ``"dvs_cifar10"``.
    root:
        Directory in the layout the matching loader reads.
    version:
        The release the files come from, as the publisher names it; it is
        recorded, not inferred, because the files do not carry it.

    Returns
    -------
    EventDatasetManifest
        The manifest.

    Raises
    ------
    ValueError
        On an unknown dataset, an empty version, a directory holding none of
        the dataset's files, or a label outside the dataset's classes.
    """
    description = _description(name)
    if not version.strip():
        raise ValueError("the dataset version must be stated")
    base = Path(root)
    paths = _layout_files(name, base)
    if not paths:
        raise ValueError(f"{base} holds no {description.title} files in the expected layout")
    files = tuple(
        FileRecord(path=_relative(base, path), bytes=path.stat().st_size, sha256=_sha256(path))
        for path in sorted(paths, key=lambda item: _relative(base, item))
    )
    if name == "shd":
        samples = _shd_samples(base)
    else:
        split_dirs, suffix = _CLASS_LAYOUT[name]
        samples = []
        for split, path, label in _class_files(base, split_dirs, suffix):
            relative = _relative(base, path)
            samples.append(
                SampleRecord(split=split, file=relative, index=0, label=label, group=relative)
            )
    bad = sorted({s.label for s in samples if not 0 <= s.label < description.classes})
    if bad:
        raise ValueError(f"labels {bad} are outside the {description.classes} classes")
    return EventDatasetManifest(
        dataset=description, version=version, files=files, samples=tuple(samples)
    )

verify_manifest(manifest, root)

Compare the files under root with a manifest, byte for byte.

Parameters

manifest: The manifest an experiment recorded. root: Directory holding the dataset now.

Returns

ManifestVerification Missing, changed and unexpected files; ok when there are none.

Source code in src/sc_neurocore/datasets/manifest.py
Python
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
def verify_manifest(manifest: EventDatasetManifest, root: str | Path) -> ManifestVerification:
    """Compare the files under ``root`` with a manifest, byte for byte.

    Parameters
    ----------
    manifest:
        The manifest an experiment recorded.
    root:
        Directory holding the dataset now.

    Returns
    -------
    ManifestVerification
        Missing, changed and unexpected files; ``ok`` when there are none.
    """
    base = Path(root)
    listed = {record.path: record for record in manifest.files}
    on_disk = {_relative(base, path) for path in _layout_files(manifest.dataset.name, base)}
    missing = sorted(path for path in listed if path not in on_disk)
    changed = sorted(
        path
        for path in listed
        if path in on_disk
        and (
            (base / path).stat().st_size != listed[path].bytes
            or _sha256(base / path) != listed[path].sha256
        )
    )
    unexpected = sorted(on_disk - set(listed))
    return ManifestVerification(
        missing=tuple(missing), changed=tuple(changed), unexpected=tuple(unexpected)
    )

manifest_from_dict(data)

Read a manifest's JSON form, refusing anything it does not define.

Parameters

data: The parsed JSON.

Returns

EventDatasetManifest The manifest.

Raises

ValueError On another schema, a missing or unknown field, or a dataset description that differs from the one this version supports.

Source code in src/sc_neurocore/datasets/manifest.py
Python
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
def manifest_from_dict(data: Mapping[str, Any]) -> EventDatasetManifest:
    """Read a manifest's JSON form, refusing anything it does not define.

    Parameters
    ----------
    data:
        The parsed JSON.

    Returns
    -------
    EventDatasetManifest
        The manifest.

    Raises
    ------
    ValueError
        On another schema, a missing or unknown field, or a dataset
        description that differs from the one this version supports.
    """
    top = _require_keys(data, {"schema", "dataset", "version", "files", "samples"}, "manifest")
    if top["schema"] != MANIFEST_SCHEMA:
        raise ValueError(f"manifest schema {top['schema']!r} is not {MANIFEST_SCHEMA!r}")
    dataset_data = top["dataset"]
    name = dataset_data.get("name") if isinstance(dataset_data, Mapping) else None
    description = _description(str(name))
    if dataset_data != description.to_dict():
        raise ValueError(
            f"the manifest describes {name!r} differently from this version of SC-NeuroCore"
        )
    files = tuple(
        FileRecord(
            path=str(record["path"]), bytes=int(record["bytes"]), sha256=str(record["sha256"])
        )
        for record in (
            _require_keys(item, {"path", "bytes", "sha256"}, "file record") for item in top["files"]
        )
    )
    samples = tuple(
        SampleRecord(
            split=str(record["split"]),
            file=str(record["file"]),
            index=int(record["index"]),
            label=int(record["label"]),
            group=str(record["group"]),
        )
        for record in (
            _require_keys(item, {"split", "file", "index", "label", "group"}, "sample record")
            for item in top["samples"]
        )
    )
    return EventDatasetManifest(
        dataset=description, version=str(top["version"]), files=files, samples=samples
    )

group_split(manifest, *, fractions, source_split='train', seed=0)

Divide one published split into new splits made of whole groups.

Groups are taken in an order fixed by seed; each goes to the split whose sample count is furthest below its requested share. The shares are therefore met as closely as whole groups allow, never by cutting a group.

Parameters

manifest: The dataset manifest. fractions: Share of the source split's samples per new split; positive, summing to one. source_split: The published split to divide; other published splits are untouched. seed: Seed of the group order.

Returns

SplitPlan The plan, tied to the manifest's digest.

Raises

ValueError On shares that are not positive or do not sum to one, an unknown source split, or fewer groups than requested splits.

Source code in src/sc_neurocore/datasets/splits.py
Python
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
def group_split(
    manifest: EventDatasetManifest,
    *,
    fractions: Mapping[str, float],
    source_split: str = "train",
    seed: int = 0,
) -> SplitPlan:
    """Divide one published split into new splits made of whole groups.

    Groups are taken in an order fixed by ``seed``; each goes to the split
    whose sample count is furthest below its requested share. The shares are
    therefore met as closely as whole groups allow, never by cutting a group.

    Parameters
    ----------
    manifest:
        The dataset manifest.
    fractions:
        Share of the source split's samples per new split; positive, summing
        to one.
    source_split:
        The published split to divide; other published splits are untouched.
    seed:
        Seed of the group order.

    Returns
    -------
    SplitPlan
        The plan, tied to the manifest's digest.

    Raises
    ------
    ValueError
        On shares that are not positive or do not sum to one, an unknown
        source split, or fewer groups than requested splits.
    """
    names = list(fractions)
    shares = [float(fractions[name]) for name in names]
    if len(names) < 2:
        raise ValueError("a split needs at least two parts")
    if not all(math.isfinite(share) and share > 0 for share in shares):
        raise ValueError(f"every share must be positive and finite; got {dict(fractions)}")
    if not math.isclose(sum(shares), 1.0, rel_tol=0.0, abs_tol=1e-9):
        raise ValueError(f"the shares sum to {sum(shares)}, not 1")
    members: dict[str, list[int]] = {}
    for position, sample in enumerate(manifest.samples):
        if sample.split == source_split:
            members.setdefault(sample.group, []).append(position)
    if not members:
        raise ValueError(
            f"the manifest has no {source_split!r} samples; its splits are {manifest.splits()}"
        )
    if len(members) < len(names):
        raise ValueError(
            f"{len(members)} groups cannot fill {len(names)} splits without cutting a group"
        )
    total = sum(len(positions) for positions in members.values())
    counts = dict.fromkeys(names, 0)
    placed: dict[str, list[str]] = {name: [] for name in names}
    order = _group_order(list(members), seed)
    for index, group in enumerate(order):
        remaining = len(order) - index
        empty = [name for name in names if not placed[name]]
        if len(empty) == remaining:
            # Every split still without a group takes one of the last groups.
            target = empty[0]
        else:
            target = max(
                names,
                key=lambda name: (fractions[name] * total - counts[name], -names.index(name)),
            )
        placed[target].append(group)
        counts[target] += len(members[group])
    assignment = {
        name: tuple(sorted(position for group in placed[name] for position in members[group]))
        for name in names
    }
    return SplitPlan(
        manifest_digest=manifest.digest,
        source_split=source_split,
        seed=seed,
        fractions=tuple((name, float(fractions[name])) for name in names),
        assignment=assignment,
        groups={name: tuple(sorted(placed[name])) for name in names},
    )

group_overlap(manifest)

Report groups that the publisher's own splits share.

Parameters

manifest: The dataset manifest.

Returns

dict Each shared group to the published splits it appears in; empty when the published splits keep every group apart.

Source code in src/sc_neurocore/datasets/splits.py
Python
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
def group_overlap(manifest: EventDatasetManifest) -> dict[str, tuple[str, ...]]:
    """Report groups that the publisher's own splits share.

    Parameters
    ----------
    manifest:
        The dataset manifest.

    Returns
    -------
    dict
        Each shared group to the published splits it appears in; empty when
        the published splits keep every group apart.
    """
    seen: dict[str, dict[str, None]] = {}
    for sample in manifest.samples:
        seen.setdefault(sample.group, {})[sample.split] = None
    return {group: tuple(splits) for group, splits in sorted(seen.items()) if len(splits) > 1}

leaked_groups(manifest, plan)

Return the groups that have samples in more than one split of a plan.

Parameters

manifest: The manifest the plan was drawn from. plan: The plan to check.

Returns

tuple of str Leaked groups, sorted; empty for a sound plan.

Raises

ValueError When the plan was drawn from another manifest.

Source code in src/sc_neurocore/datasets/splits.py
Python
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
def leaked_groups(manifest: EventDatasetManifest, plan: SplitPlan) -> tuple[str, ...]:
    """Return the groups that have samples in more than one split of a plan.

    Parameters
    ----------
    manifest:
        The manifest the plan was drawn from.
    plan:
        The plan to check.

    Returns
    -------
    tuple of str
        Leaked groups, sorted; empty for a sound plan.

    Raises
    ------
    ValueError
        When the plan was drawn from another manifest.
    """
    if plan.manifest_digest != manifest.digest:
        raise ValueError("the plan was drawn from another manifest")
    seen: dict[str, set[str]] = {}
    for name, positions in plan.assignment.items():
        for position in positions:
            seen.setdefault(manifest.samples[position].group, set()).add(name)
    return tuple(sorted(group for group, splits in seen.items() if len(splits) > 1))

split_plan_from_dict(data)

Read a plan's JSON form, refusing anything it does not define.

A plan read back is not trusted to be sound: check it against its manifest with :func:validate_split_plan before training on it.

Parameters

data: The parsed JSON.

Returns

SplitPlan The plan.

Raises

ValueError On another schema, missing or unknown fields, invalid types, duplicate positions, or inconsistent split names. Values are never coerced.

Source code in src/sc_neurocore/datasets/splits.py
Python
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
def split_plan_from_dict(data: Mapping[str, Any]) -> SplitPlan:
    """Read a plan's JSON form, refusing anything it does not define.

    A plan read back is not trusted to be sound: check it against its
    manifest with :func:`validate_split_plan` before training on it.

    Parameters
    ----------
    data:
        The parsed JSON.

    Returns
    -------
    SplitPlan
        The plan.

    Raises
    ------
    ValueError
        On another schema, missing or unknown fields, invalid types, duplicate
        positions, or inconsistent split names. Values are never coerced.
    """
    keys = {
        "schema",
        "manifest_digest",
        "source_split",
        "seed",
        "fractions",
        "assignment",
        "groups",
    }
    if not isinstance(data, Mapping) or set(data) != keys:
        raise ValueError(f"a split plan has exactly the fields {sorted(keys)}")
    if data["schema"] != SPLIT_SCHEMA:
        raise ValueError(f"split schema {data['schema']!r} is not {SPLIT_SCHEMA!r}")
    digest = _split_text(data["manifest_digest"], "manifest_digest")
    if not digest.startswith("sha256:") or len(digest) != 71:
        raise ValueError("manifest_digest must be a sha256 digest")
    if any(character not in "0123456789abcdef" for character in digest[7:]):
        raise ValueError("manifest_digest must be a sha256 digest")
    source = _split_text(data["source_split"], "source_split")
    seed = data["seed"]
    if isinstance(seed, bool) or not isinstance(seed, int) or not 0 <= seed < 2**32:
        raise ValueError("seed must be an integer in [0, 2**32)")
    fractions = _split_fractions(data["fractions"])
    names = {name for name, _ in fractions}
    assignments = _split_mapping(data["assignment"], names, "assignment")
    declared_groups = _split_mapping(data["groups"], names, "groups")
    assignment: dict[str, tuple[int, ...]] = {}
    groups: dict[str, tuple[str, ...]] = {}
    seen: set[int] = set()
    for name, _ in fractions:
        positions: list[int] = []
        for position in assignments[name]:
            if isinstance(position, bool) or not isinstance(position, int) or position < 0:
                raise ValueError("sample positions must be non-negative integers")
            if position in seen:
                raise ValueError("each sample position must appear exactly once")
            seen.add(position)
            positions.append(position)
        assignment[name] = tuple(positions)
        labels = tuple(_split_text(group, "group") for group in declared_groups[name])
        if len(set(labels)) != len(labels):
            raise ValueError("group declarations must not contain duplicates")
        groups[name] = labels
    return SplitPlan(digest, source, seed, fractions, assignment, groups)