Skip to content

Gate a rebuild

Goal: decide whether a rebuilt embedding corpus changed more than a rebuild changes on its own, using a floor measured from null rebuilds.

Prerequisites: the Python package installed (it provides the semq command), a reference state, at least three null rebuilds of the same corpus, and the candidate to judge. All must use the same CodecConfig and id kind.

1. Produce the states

Encode each corpus with the same configuration and ids, declaring the encoder in the manifest:

import numpy as np
from semq import Codec

rng = np.random.default_rng(0)
vectors = rng.standard_normal((100, 16)).astype(np.float64)
vectors /= np.linalg.norm(vectors, axis=1, keepdims=True)
vectors = vectors.astype(np.float32)
ids = np.arange(100, dtype=np.uint64)
manifest = {"encoder": "example-encoder", "encoder_revision": "1"}

codec = Codec.quant(dim=16, bins=4)
codec.encode(vectors, ids=ids, manifest=manifest).save("reference.semq")

# Null rebuilds: the same corpus through the same encoder, with the noise a
# real rebuild has. Here one coordinate of one row moves slightly in each.
for k, row in enumerate((7, 21, 42), start=1):
    null = vectors.copy()
    null[row, 3] += np.float32(0.02)
    null[row] /= np.linalg.norm(null[row])
    codec.encode(null, ids=ids, manifest=manifest).save(f"null-{k}.semq")

# The candidate to judge: rows 0 and 1 changed direction.
candidate = vectors.copy()
candidate[0] = -candidate[0]
candidate[1] = -candidate[1]
codec.encode(candidate, ids=ids, manifest=manifest).save("candidate.semq")

The coordinate perturbations above only make this example runnable. For a CI gate, produce each null state by independently rebuilding the unchanged corpus with the same encoder and production settings. Synthetic perturbations do not establish the rebuild noise floor.

2. Measure the floor

semq floor reference.semq null-1.semq null-2.semq null-3.semq > floor.json
cat floor.json

floor.json records the three counts (changed_rows, total_rows, hamming: the floor takes the worst ratio and the largest p99 hamming over the nulls) and where they were measured: the config, the id kind, the reference's state_id and how many nulls went in. The floor applies only to diffs against that reference; measure a new one when the reference changes. semq floor requires three nulls by default because one null only shows what that rebuild happened to do; --min-nulls N lowers the requirement explicitly. It exits 2 and prints nothing on stdout when an input is not a valid null diff (rows added or removed, the encoder keys changed, or a null of another reference).

3. Gate the candidate

semq diff reference.semq candidate.semq --floor floor.json
echo "exit $?"

Exit 0 means the candidate is within the floor. Exit 1 means it is not; a changed encoder or encoder_revision is never within. Exit 2 means the gate could not be evaluated (invalid file, invalid floor, or a floor measured against another reference, config or id kind). Changed manifest keys are listed on stderr; the verdict is computed by the core. Add --json for the complete report; without --floor, semq diff is a report and always exits 0.

The same verdict is available in code:

from semq import Encoding, Floor

reference = Encoding.load("reference.semq")
floor = Floor.measure([reference.diff(Encoding.load(f"null-{k}.semq")) for k in (1, 2, 3)])
floor.save("floor.json")  # the same file `semq floor` writes
print(floor)
diff = reference.diff(Encoding.load("candidate.semq"))
print(diff)
print(diff.within(floor))
print(diff.evaluate(floor).reasons)  # every failed check, not only the first
print(diff.units(0)[:3])  # the first units of row 0 that moved: (unit, before, after)

4. Check every row

The p99 check ignores the most changed 1% of rows, so a few rows with a large change can pass. The per-row check also fails the verdict when any changed row moved more than any row of any null did. It needs a floor measured for it:

semq floor reference.semq null-1.semq null-2.semq null-3.semq --per-row > floor-per-row.json
semq diff reference.semq candidate.semq --floor floor-per-row.json --per-row
echo "exit $?"

With --per-row, semq floor also records max_hamming, the largest hamming of any changed row of any null, and distinct_nulls, how many of the nulls were different states (two nulls are the same state when their candidates have the same content_digest). When the gate fails, stderr lists the rows above max_hamming. semq diff --per-row exits 2 on a floor without these keys, such as one from semq floor without --per-row or one saved by SEMQ 1.0. SEMQ 1.0 cannot read a floor written with --per-row; a floor written without it reads in every version.

In code:

from semq import Encoding, Floor

reference = Encoding.load("reference.semq")
nulls = [reference.diff(Encoding.load(f"null-{k}.semq")) for k in (1, 2, 3)]
floor = Floor.measure(nulls, per_row=True)
print(floor.max_hamming, floor.distinct_nulls)
verdict = reference.diff(Encoding.load("candidate.semq")).evaluate(floor, per_row=True)
print(verdict.passed, verdict.reasons, verdict.rows)  # the rows above max_hamming

How many nulls. Each check compares one statistic of the candidate with the largest value of that statistic among the nulls. If an unchanged rebuild and the N nulls are exchangeable (produced the same way, so that any order of the N + 1 is equally likely), the unchanged rebuild is above all N nulls with probability at most 1/(N+1), so each statistic rejects it with probability at most 1/(N+1). within uses two statistics, the changed ratio and the p99 hamming, so it rejects an unchanged rebuild with probability at most 2/(N+1). The per-row check adds a third, for at most 3/(N+1): 75% with 3 nulls, about 14% with 20. In practice it adds less, because hamming distances are integers and a tie with the largest null does not fail. When every null rebuild gave the same rows (distinct_nulls is 1), the rebuild is deterministic: an unchanged rebuild gives those rows again, so the per-row check adds no false alarm. When the nulls vary, measure the floor from at least 20 nulls; semq diff --per-row prints a warning on stderr below that, without changing the exit code.

What the verdict means

within is true only if the candidate removed no rows, shares at least one row with the reference, changed no more than the floor's ratio of the shared rows, its p99 hamming is at most the floor's, and neither encoder nor encoder_revision changed. Added rows do not affect it; they are listed so you can judge them. The floor is an envelope of what you observed, taken component-wise over the nulls; it does not estimate the probability of the next rebuild. The bound in step 4 holds only if the nulls and the candidate are exchangeable: it comes from how you produce them, not from the floor.

Encode a large corpus in batches

For a corpus that does not fit in memory as one float32 matrix, encode bounded batches and join the resulting states once with concat. Every batch must use the same config, id kind and manifest, with distinct ids:

import numpy as np
from semq import Codec

codec = Codec.quant(dim=16, bins=4)
manifest = {"encoder": "example-encoder", "encoder_revision": "1"}
ids = np.arange(100, dtype=np.uint64)
vectors = np.zeros((100, 16), dtype=np.float32)
vectors[np.arange(100), np.arange(100) % 16] = 1.0
source_batches = (
    (ids[start:start + 20], vectors[start:start + 20])
    for start in range(0, len(ids), 20)
)
parts = [
    codec.encode(batch_vectors, ids=batch_ids, manifest=manifest)
    for batch_ids, batch_vectors in source_batches
]
state = parts[0].concat(*parts[1:])
state.save("reference.semq")
assert state == codec.encode(vectors, ids=ids, manifest=manifest)

The compressed parts still occupy memory until concat finishes. The host may encode batches concurrently with a shared Codec; keep the number of in-flight batches bounded. One concat call avoids rebuilding the state and rehashing it for every batch. The same ids, vectors and manifest produce the same state as a single encode call.

Next steps