Skip to content

Hyper-Dimensional Computing — Binary Vector Algebra

High-dimensional binary vector algebra for symbolic reasoning in spiking networks. HDC maps naturally to stochastic computing hardware: bind = XOR gate, bundle = popcount tree, similarity = Hamming distance.

Theory

HDC represents symbols as random binary vectors of dimension D (typically D >= 10,000). At high D, random vectors are quasi-orthogonal with high probability: E[d_H(a,b)] = D/2. Three operations form an algebra:

Operation Implementation Property
Bind (⊗) XOR Self-inverse: a ⊗ a = 0, a ⊗ b ⊗ b = a
Bundle (⊕) Majority vote Preserves similarity to all inputs
Permute (ρ) Cyclic shift Breaks commutativity for ordered structures

The full reference-locked semantics — representation, seed contract, tie policies, permutation direction, distance, clean-up-memory ties, and the executed enforcement map — live in the HDC/VSA semantic contract.

Components

  • HDCEncoder — Generate random D-dimensional binary vectors and perform algebraic operations.
Parameter Default Meaning
dim 10000 Hypervector dimension
seed None Seeds the encoder's own generator; a seeded encoder is fully deterministic for the same call order
tie_policy "zeros" Even-count bundle ties: "zeros" clears tied bits (historical strict majority), "ones" sets them, "random" decides them from a fresh seeded tie-break vector

Methods: generate_random_vector(), item(name) (cached named item memory), bind(v1, v2), bundle(vectors), majority(sum_vec, count) (shared bundle kernel), permute(v, shifts), level_vectors(low, high, levels) and encode_level(value, low, high, levels=16) (linear level encoding whose Hamming distance grows linearly with level separation, for scalar features).

  • AssociativeMemory — Clean-up memory via Hamming distance nearest-neighbor lookup. Store labeled vectors, retrieve by similarity. Tolerates up to ~35% bit noise.

  • CentroidHDClassifier — Nearest-centroid classifier over binary hypervectors with mistake-driven retraining. Each class keeps a bipolar accumulator; the centroid is its sign with exact zeros resolved by the encoder's tie policy. fit(vectors, labels) accumulates, predict(vector) returns the nearest centroid by Hamming distance, and retrain(vectors, labels, epochs) applies the standard mistake-driven update (add the misclassified example to its true class, subtract it from the predicted class), returning the misclassification count per epoch. Deterministic for a seeded encoder. Rejects non-binary or wrongly shaped vectors, unknown retrain labels, and non-positive epochs with typed ValueErrors; the whole surface is enforced at 100% statement and branch coverage by the hosted HDC exact coverage lane.

Usage

Python
from sc_neurocore.hdc import HDCEncoder, AssociativeMemory
import numpy as np

np.random.seed(42)
enc = HDCEncoder(dim=10000)

# Create symbols
country = enc.generate_random_vector()
capital = enc.generate_random_vector()
usa = enc.generate_random_vector()
washington = enc.generate_random_vector()

# Encode: USA_record = bind(country, usa) ⊕ bind(capital, washington)
record = enc.bundle([
    enc.bind(country, usa),
    enc.bind(capital, washington),
])

# Query: "What is the capital of USA?" → bind(record, capital)
query = enc.bind(record, capital)

# Store in associative memory and retrieve
mem = AssociativeMemory()
mem.store("washington", washington)
mem.store("usa", usa)
print(mem.query(query))  # → "washington"

See Tutorial 4: Hyper-Dimensional Computing.

sc_neurocore.hdc.base

Hyperdimensional computing encoder and associative clean-up memory.

Binary {0, 1} hypervector algebra (bind = XOR, bundle = majority, permute = cyclic shift) with a seeded generator, a named item memory, an explicit bundle tie policy, and linear level (thermometer) encoding for scalars. All randomness flows through one numpy generator owned by the encoder, so a seeded encoder is fully deterministic given the same call sequence.

HDCEncoder dataclass

Hyperdimensional computing encoder.

Dimension D is usually >= 10,000. seed makes every draw deterministic (given the same call order); tie_policy states what an even-count bundle does on exactly tied bit positions: "zeros" clears them (historical strict-majority behaviour), "ones" sets them, and "random" decides each tied position from a fresh seeded tie-break hypervector (the unbiased Kanerva convention).

Source code in src/sc_neurocore/hdc/base.py
Python
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
@dataclass
class HDCEncoder:
    """Hyperdimensional computing encoder.

    Dimension D is usually >= 10,000. ``seed`` makes every draw
    deterministic (given the same call order); ``tie_policy`` states
    what an even-count bundle does on exactly tied bit positions:
    ``"zeros"`` clears them (historical strict-majority behaviour),
    ``"ones"`` sets them, and ``"random"`` decides each tied position
    from a fresh seeded tie-break hypervector (the unbiased Kanerva
    convention).
    """

    dim: int = 10000
    seed: int | None = None
    tie_policy: str = "zeros"
    _rng: np.random.Generator = field(init=False, repr=False)
    _item_memory: dict[str, np.ndarray[Any, Any]] = field(
        init=False, repr=False, default_factory=dict
    )
    _level_memory: dict[tuple[float, float, int], np.ndarray[Any, Any]] = field(
        init=False, repr=False, default_factory=dict
    )

    def __post_init__(self) -> None:
        """Validate configuration and initialise the seeded generator."""
        if not isinstance(self.dim, int) or isinstance(self.dim, bool) or self.dim < 1:
            raise ValueError("dim must be a positive integer")
        if self.tie_policy not in _TIE_POLICIES:
            raise ValueError(f"tie_policy must be one of {list(_TIE_POLICIES)}")
        self._rng = np.random.default_rng(self.seed)

    def generate_random_vector(self) -> np.ndarray[Any, Any]:
        """Generate a random D-dimensional binary vector in {0, 1}."""
        # We use {0, 1} for compatibility with our SC
        vector: np.ndarray[Any, Any] = self._rng.integers(0, 2, self.dim, dtype=np.uint8)
        return vector

    def item(self, name: str) -> np.ndarray[Any, Any]:
        """Return the named item hypervector, drawing it on first use.

        The same encoder always returns the identical vector for the
        same name; a seeded encoder reproduces the whole item memory
        when the names are first requested in the same order.
        """
        if name not in self._item_memory:
            self._item_memory[name] = self.generate_random_vector()
        return self._item_memory[name].copy()

    def bind(self, v1: np.ndarray[Any, Any], v2: np.ndarray[Any, Any]) -> np.ndarray[Any, Any]:
        """Bind two hypervectors via XOR."""
        bound: np.ndarray[Any, Any] = np.bitwise_xor(v1, v2)
        return bound

    def bundle(self, vectors: list[np.ndarray[Any, Any]]) -> np.ndarray[Any, Any]:
        """Bundle hypervectors by majority superposition.

        Bit positions with a strict majority of ones become one and a
        strict majority of zeros become zero; exactly tied positions
        (possible only for an even vector count) follow ``tie_policy``.
        """
        if not vectors:
            return np.zeros(self.dim, dtype=np.uint8)
        sum_vec = np.sum(vectors, axis=0)
        return self.majority(sum_vec, len(vectors))

    def majority(self, sum_vec: np.ndarray[Any, Any], count: int) -> np.ndarray[Any, Any]:
        """Return the majority vector of ``count`` bundled binary vectors.

        ``sum_vec`` holds the per-position count of ones. This is the
        bundle kernel, shared with the centroid classifier so both
        apply the identical tie policy.
        """
        if not isinstance(count, int) or isinstance(count, bool) or count < 1:
            raise ValueError("count must be a positive integer")
        doubled = 2 * np.asarray(sum_vec, dtype=np.int64)
        majority = (doubled > count).astype(np.uint8)
        if count % 2 == 0:
            tied = doubled == count
            if bool(np.any(tied)):
                if self.tie_policy == "ones":
                    majority[tied] = 1
                elif self.tie_policy == "random":
                    tie_break = self.generate_random_vector()
                    majority[tied] = tie_break[tied]
        bundled: np.ndarray[Any, Any] = majority
        return bundled

    def permute(self, v: np.ndarray[Any, Any], shifts: int = 1) -> np.ndarray[Any, Any]:
        """Permute a hypervector by a cyclic shift."""
        shifted: np.ndarray[Any, Any] = np.roll(v, shifts)
        return shifted

    def level_vectors(self, low: float, high: float, levels: int) -> np.ndarray[Any, Any]:
        """Return the ``levels`` linear level hypervectors for [low, high].

        Level 0 is a fresh random hypervector; each subsequent level
        flips the next ``(dim // 2) // (levels - 1)`` positions of a
        fixed seeded permutation, so the Hamming distance between two
        levels grows linearly with their separation and the endpoints
        differ in ``(dim // 2) // (levels - 1) * (levels - 1)`` bits
        (approaching orthogonality). The family is drawn once per
        ``(low, high, levels)`` triple and cached.
        """
        if not isinstance(levels, int) or isinstance(levels, bool) or levels < 2:
            raise ValueError("levels must be an integer >= 2")
        if not (math.isfinite(low) and math.isfinite(high)) or not low < high:
            raise ValueError("low and high must be finite with low < high")
        key = (float(low), float(high), levels)
        if key not in self._level_memory:
            base = self.generate_random_vector()
            order = self._rng.permutation(self.dim)
            per_gap = (self.dim // 2) // (levels - 1)
            family = np.empty((levels, self.dim), dtype=np.uint8)
            family[0] = base
            for level in range(1, levels):
                vector = family[level - 1].copy()
                flips = order[(level - 1) * per_gap : level * per_gap]
                vector[flips] ^= 1
                family[level] = vector
            self._level_memory[key] = family
        return self._level_memory[key].copy()

    def encode_level(
        self, value: float, low: float, high: float, levels: int = 16
    ) -> np.ndarray[Any, Any]:
        """Encode a scalar as its nearest linear level hypervector.

        ``value`` is clipped into [low, high] and mapped to the closest
        of the ``levels`` cached level vectors for that range.
        """
        if not math.isfinite(value):
            raise ValueError("value must be finite")
        family = self.level_vectors(low, high, levels)
        clipped = min(max(value, low), high)
        index = round((clipped - low) / (high - low) * (levels - 1))
        encoded: np.ndarray[Any, Any] = family[index]
        return encoded

__post_init__()

Validate configuration and initialise the seeded generator.

Source code in src/sc_neurocore/hdc/base.py
Python
54
55
56
57
58
59
60
def __post_init__(self) -> None:
    """Validate configuration and initialise the seeded generator."""
    if not isinstance(self.dim, int) or isinstance(self.dim, bool) or self.dim < 1:
        raise ValueError("dim must be a positive integer")
    if self.tie_policy not in _TIE_POLICIES:
        raise ValueError(f"tie_policy must be one of {list(_TIE_POLICIES)}")
    self._rng = np.random.default_rng(self.seed)

generate_random_vector()

Generate a random D-dimensional binary vector in {0, 1}.

Source code in src/sc_neurocore/hdc/base.py
Python
62
63
64
65
66
def generate_random_vector(self) -> np.ndarray[Any, Any]:
    """Generate a random D-dimensional binary vector in {0, 1}."""
    # We use {0, 1} for compatibility with our SC
    vector: np.ndarray[Any, Any] = self._rng.integers(0, 2, self.dim, dtype=np.uint8)
    return vector

item(name)

Return the named item hypervector, drawing it on first use.

The same encoder always returns the identical vector for the same name; a seeded encoder reproduces the whole item memory when the names are first requested in the same order.

Source code in src/sc_neurocore/hdc/base.py
Python
68
69
70
71
72
73
74
75
76
77
def item(self, name: str) -> np.ndarray[Any, Any]:
    """Return the named item hypervector, drawing it on first use.

    The same encoder always returns the identical vector for the
    same name; a seeded encoder reproduces the whole item memory
    when the names are first requested in the same order.
    """
    if name not in self._item_memory:
        self._item_memory[name] = self.generate_random_vector()
    return self._item_memory[name].copy()

bind(v1, v2)

Bind two hypervectors via XOR.

Source code in src/sc_neurocore/hdc/base.py
Python
79
80
81
82
def bind(self, v1: np.ndarray[Any, Any], v2: np.ndarray[Any, Any]) -> np.ndarray[Any, Any]:
    """Bind two hypervectors via XOR."""
    bound: np.ndarray[Any, Any] = np.bitwise_xor(v1, v2)
    return bound

bundle(vectors)

Bundle hypervectors by majority superposition.

Bit positions with a strict majority of ones become one and a strict majority of zeros become zero; exactly tied positions (possible only for an even vector count) follow tie_policy.

Source code in src/sc_neurocore/hdc/base.py
Python
84
85
86
87
88
89
90
91
92
93
94
def bundle(self, vectors: list[np.ndarray[Any, Any]]) -> np.ndarray[Any, Any]:
    """Bundle hypervectors by majority superposition.

    Bit positions with a strict majority of ones become one and a
    strict majority of zeros become zero; exactly tied positions
    (possible only for an even vector count) follow ``tie_policy``.
    """
    if not vectors:
        return np.zeros(self.dim, dtype=np.uint8)
    sum_vec = np.sum(vectors, axis=0)
    return self.majority(sum_vec, len(vectors))

majority(sum_vec, count)

Return the majority vector of count bundled binary vectors.

sum_vec holds the per-position count of ones. This is the bundle kernel, shared with the centroid classifier so both apply the identical tie policy.

Source code in src/sc_neurocore/hdc/base.py
Python
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
def majority(self, sum_vec: np.ndarray[Any, Any], count: int) -> np.ndarray[Any, Any]:
    """Return the majority vector of ``count`` bundled binary vectors.

    ``sum_vec`` holds the per-position count of ones. This is the
    bundle kernel, shared with the centroid classifier so both
    apply the identical tie policy.
    """
    if not isinstance(count, int) or isinstance(count, bool) or count < 1:
        raise ValueError("count must be a positive integer")
    doubled = 2 * np.asarray(sum_vec, dtype=np.int64)
    majority = (doubled > count).astype(np.uint8)
    if count % 2 == 0:
        tied = doubled == count
        if bool(np.any(tied)):
            if self.tie_policy == "ones":
                majority[tied] = 1
            elif self.tie_policy == "random":
                tie_break = self.generate_random_vector()
                majority[tied] = tie_break[tied]
    bundled: np.ndarray[Any, Any] = majority
    return bundled

permute(v, shifts=1)

Permute a hypervector by a cyclic shift.

Source code in src/sc_neurocore/hdc/base.py
Python
118
119
120
121
def permute(self, v: np.ndarray[Any, Any], shifts: int = 1) -> np.ndarray[Any, Any]:
    """Permute a hypervector by a cyclic shift."""
    shifted: np.ndarray[Any, Any] = np.roll(v, shifts)
    return shifted

level_vectors(low, high, levels)

Return the levels linear level hypervectors for [low, high].

Level 0 is a fresh random hypervector; each subsequent level flips the next (dim // 2) // (levels - 1) positions of a fixed seeded permutation, so the Hamming distance between two levels grows linearly with their separation and the endpoints differ in (dim // 2) // (levels - 1) * (levels - 1) bits (approaching orthogonality). The family is drawn once per (low, high, levels) triple and cached.

Source code in src/sc_neurocore/hdc/base.py
Python
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
def level_vectors(self, low: float, high: float, levels: int) -> np.ndarray[Any, Any]:
    """Return the ``levels`` linear level hypervectors for [low, high].

    Level 0 is a fresh random hypervector; each subsequent level
    flips the next ``(dim // 2) // (levels - 1)`` positions of a
    fixed seeded permutation, so the Hamming distance between two
    levels grows linearly with their separation and the endpoints
    differ in ``(dim // 2) // (levels - 1) * (levels - 1)`` bits
    (approaching orthogonality). The family is drawn once per
    ``(low, high, levels)`` triple and cached.
    """
    if not isinstance(levels, int) or isinstance(levels, bool) or levels < 2:
        raise ValueError("levels must be an integer >= 2")
    if not (math.isfinite(low) and math.isfinite(high)) or not low < high:
        raise ValueError("low and high must be finite with low < high")
    key = (float(low), float(high), levels)
    if key not in self._level_memory:
        base = self.generate_random_vector()
        order = self._rng.permutation(self.dim)
        per_gap = (self.dim // 2) // (levels - 1)
        family = np.empty((levels, self.dim), dtype=np.uint8)
        family[0] = base
        for level in range(1, levels):
            vector = family[level - 1].copy()
            flips = order[(level - 1) * per_gap : level * per_gap]
            vector[flips] ^= 1
            family[level] = vector
        self._level_memory[key] = family
    return self._level_memory[key].copy()

encode_level(value, low, high, levels=16)

Encode a scalar as its nearest linear level hypervector.

value is clipped into [low, high] and mapped to the closest of the levels cached level vectors for that range.

Source code in src/sc_neurocore/hdc/base.py
Python
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
def encode_level(
    self, value: float, low: float, high: float, levels: int = 16
) -> np.ndarray[Any, Any]:
    """Encode a scalar as its nearest linear level hypervector.

    ``value`` is clipped into [low, high] and mapped to the closest
    of the ``levels`` cached level vectors for that range.
    """
    if not math.isfinite(value):
        raise ValueError("value must be finite")
    family = self.level_vectors(low, high, levels)
    clipped = min(max(value, low), high)
    index = round((clipped - low) / (high - low) * (levels - 1))
    encoded: np.ndarray[Any, Any] = family[index]
    return encoded

AssociativeMemory dataclass

Simple HDC associative clean-up memory.

Stores (key, value) pairs or bare prototypes for nearest-match retrieval.

Source code in src/sc_neurocore/hdc/base.py
Python
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
@dataclass
class AssociativeMemory:
    """Simple HDC associative clean-up memory.

    Stores (key, value) pairs or bare prototypes for nearest-match retrieval.
    """

    memory: dict[str, Any] = field(default_factory=dict)

    def store(self, label: str, vector: np.ndarray[Any, Any]) -> None:
        """Store a labelled hypervector in the clean-up memory."""
        self.memory[label] = vector

    def query(self, query_vec: np.ndarray[Any, Any]) -> str | None:
        """Return the label of the closest stored vector by Hamming distance."""
        best_label = None
        min_dist = float("inf")

        for label, mem_vec in self.memory.items():
            # Hamming distance = count(XOR)
            dist = float(np.count_nonzero(np.bitwise_xor(query_vec, mem_vec)))
            if dist < min_dist:
                min_dist = dist
                best_label = label

        return best_label

store(label, vector)

Store a labelled hypervector in the clean-up memory.

Source code in src/sc_neurocore/hdc/base.py
Python
179
180
181
def store(self, label: str, vector: np.ndarray[Any, Any]) -> None:
    """Store a labelled hypervector in the clean-up memory."""
    self.memory[label] = vector

query(query_vec)

Return the label of the closest stored vector by Hamming distance.

Source code in src/sc_neurocore/hdc/base.py
Python
183
184
185
186
187
188
189
190
191
192
193
194
195
def query(self, query_vec: np.ndarray[Any, Any]) -> str | None:
    """Return the label of the closest stored vector by Hamming distance."""
    best_label = None
    min_dist = float("inf")

    for label, mem_vec in self.memory.items():
        # Hamming distance = count(XOR)
        dist = float(np.count_nonzero(np.bitwise_xor(query_vec, mem_vec)))
        if dist < min_dist:
            min_dist = dist
            best_label = label

    return best_label