Core contracts¶
The rules that decide bytes and verdicts live in the C core, and every
binding is a thin layer over it. This page states those rules exactly: what
encode accepts, how each operator maps a float to a symbol, what diff
reports, and when a candidate is within a floor. They are pinned by the
conformance vectors,
which CI regenerates on every architecture and every binding runs. The byte
layouts they produce are on the file format page.
What is byte-identical. For the same float32 input, the same ids, the
same manifest and the same config, encode, unpack, diff, the two
digests and save/load produce identical bytes on every CPU, operating
system and binding, as long as construction and encoding run with
round-to-nearest-even. Codec construction rejects other modes before
computing cached quantities; the core never changes the mode. No
floating-point value takes part in an identity or a verdict. decode is
outside this promise: its representatives may differ across platforms by
one unit in the last place for quant and orbit and by four for phase.
Input rows¶
encode takes float32 vectors of shape [n, dim], contiguous and
row-major. A float64 array is rejected with InvalidInput, never cast. The
core then applies four steps to every row:
- Subnormals. Every coordinate with
|x| < 2^-126, including both zeros, becomes+0.0. The result therefore does not depend on the flush-to-zero state of the process. - Finiteness. A NaN or an infinity is
InvalidInputwith the row and the coordinate. - Norm.
s = Σ (double)x_i × (double)x_i, accumulated in binary64 in index order (each square of a float32 is exact in binary64). The row is accepted when|s - 1| <= 2^-10; otherwiseInvalidInputwith the row. No square root is taken. - Rounding mode. Once per call, the process must be in round-to-nearest
mode; otherwise
Unsupported. The core never changes the mode.
Operators¶
Every operator turns a unit (one coordinate, or one coordinate pair for
phase) into one symbol. decode returns a representative per unit; it is
not normalized, and re-encoding it without the norm check gives the same
row back.
orbit¶
scale in 1..2^30, default 50; the alphabet is fixed at 19 symbols.
p = (double)x × (double)scale, thenm = round_half_even(p), computed without relying on the rounding mode:f = floor(p),r = p - f;m = f + 1whenr > 0.5,m = fwhenr < 0.5, and on a tiem = fwhenfis even, elsef + 1. With an admitted row andscale <= 2^30,|m| < 2^31always holds.- Digital root: for
a = |m| > 0,d = 1 + ((a - 1) mod 9), in1..9. - Symbol:
0whenm = 0;dwhenm > 0;9 + dwhenm < 0.
Representative: sign × d / scale in binary64, cast to float32 (0.0 for
symbol 0).
phase¶
dim even, sectors in 2..256, and dim a multiple of 4 when
sectors <= 16. A unit is the pair (x, y) = (x_2j, x_2j+1).
- Angle. With
ax = |x|anday = |y|in binary64: a zero pair has angle0. Otherwise the first-octant angle isP(ay / ax)whenax >= ayandπ/2 - P(ax / ay)whenax < ay, whereP(t) = ((c5 × t² + c3) × t² + c1) × tis evaluated in binary64 with the two inner steps as fused multiply-adds and the final product unfused, withc1 = 0.9947660466480732,c3 = -0.28543420102605926,c5 = 0.07606631777543432,π = 3.141592653589793238462643383279502884. Thenθ = π - φwhenx < 0, andθ = -θwheny < 0. - Sector:
raw = (θ + π) × (sectors / (2π))in binary64, the quotient computed first;rawis clamped below at0, truncated to an integer, and clamped above atsectors - 1.
P is the minimax odd quintic for atan on [0, 1] under the constraint
P(1) = π/4, which holds to the bit in binary64: the two octants meet
without a jump and the angle is monotone across the diagonal. Its error
against atan is at most 7.04 × 10⁻⁴ rad (0.04°), reached at three
interior points, and its slope stays above 0.518. P is the rule, not the
exact angle: an input within that error of a sector boundary may land in
the neighbouring sector of the exact atan2, and every backend reproduces
these bits.
Representative: the unit direction whose encoder angle is the midpoint of
the sector, -π + (sector + 0.5) × 2π / sectors. It is found by inverting
P on [0, 1] by bisection, reflecting the quadrant as above and
normalizing (ax, ay), which is the one place the platform libm is used.
quant¶
bins in 2..64. The range is M = max_magnitude = (float)(2.0 /
sqrt((double)dim)).
sign = 1whenx >= +0.0, else0(after step 1 there is no-0.0).m = |x|in float32.- When
m >= M,bin = bins - 1. Otherwiset = (m × (float)bins) / M, both operations in float32 without fused contraction, andbin = min(floor(t), bins - 1). - Symbol:
sign × bins + bin.
Representative: c = (bin + 0.5) × M / bins in binary64, with the sign of
the unit, cast to float32.
Encoding¶
An Encoding is immutable: config, id_kind, sorted ids, one canonical
row per id, manifest, content_digest and state_id.
Codec.encodebuilds one from vectors.id_kindis inferred from the first id and must be given when there are no rows.Encoding(ids, rows, config, manifest, id_kind)builds one from rows you already hold: it sorts the pairs by id, rejects a duplicate id (naming the first repeated row), checks the row width and canonicity, validates the manifest, copies the rows once and computes both digests.loadalways copies and validates in the order the file format page describes.concat(*others)needs the same config and id kind (elseIncompatible) and the same manifest with pairwise-disjoint ids (elseInvalidInput). It is commutative and associative, and the empty Encoding of the same config, kind and manifest is its neutral element.diff(candidate)needs the same config and id kind (elseIncompatible). When bothcontent_digestare equal the row lists are empty without a scan;manifest_changesis still computed.- Two Encodings are equal when their
state_idis equal.
Diff¶
added: ids only in the candidate;removed: ids only in the reference; both in id order.changed: the ids present in both with different rows, each with its hamming distance: the number of units whose symbol differs, compared per symbol and never per byte.n_unchanged: the ids present in both with identical rows.manifest_changes: each key whose value differs or exists on one side only, with its value before and after (absent isnull).units(id): for a changed id, the units that moved with both symbols;InvalidInputwhen the id is absent from either side.
A Diff keeps the rows it needs until it is released, even if the caller drops its Encodings.
Report schema¶
{
"reference_id": hex64,
"candidate_id": hex64,
"id_kind": "u64" | "utf8",
"config": { "operator": "quant" | "phase" | "orbit", "dim": int,
"bins" | "sectors" | "scale": int, "rule_revision": int },
"added": [id, ...],
"removed": [id, ...],
"changed": [[id, hamming], ...],
"n_unchanged": int,
"manifest_changes": { key: [before | null, after | null], ... }
}
u64 ids are decimal strings, because JavaScript cannot hold integers above
2^53 exactly; utf8 ids are JSON strings; counts are JSON integers;
digests are lowercase hex; there are no floats. The keys and values are the
contract, not the JSON text: each binding serializes with its standard
library, and vector 13 compares structure and values.
Floor¶
A floor is the envelope of variation seen in rebuilds that changed nothing
on purpose, bound to where it was measured: config, id_kind,
reference_id (the state_id of the reference every null was taken
against), nulls (how many null diffs went in), changed_rows,
total_rows and hamming. A floor measured for the per-row check also
records the per-row data, max_hamming and distinct_nulls. A floor
without them, such as one from measure or saved by SEMQ 1.0, does not
record them; its accessors return none (SEMQ_NONE in C, None, nil
or undefined in the bindings).
Construction. Every rule is checked when a floor is built or loaded, not
when it is applied: the config is valid, id_kind is u64 or utf8,
nulls >= 1, total_rows >= 1, changed_rows <= total_rows,
hamming <= units_per_row of the config, and, when present,
hamming <= max_hamming <= units_per_row and
1 <= distinct_nulls <= nulls. Otherwise InvalidInput. The constructors
build a floor without per-row data; it comes only from measure_for or
from the JSON form.
Floor.measure(null_diffs). Every null must share the config, the id
kind and the reference of the first (else Incompatible, naming the index),
and must be a valid null: rows in common, nothing added or removed, no
change to encoder or encoder_revision (else InvalidInput, naming the
index). Then (changed_rows, total_rows) is the pair
(len(changed), n_common) with the largest ratio among the nulls, compared
exactly by cross-multiplication; hamming is the largest per-null p99 of
the hamming distances; nulls is the count of nulls. The floor records no
per-row data, so its JSON form is the one SEMQ 1.0 writes and reads.
Floor.measure_for(null_diffs, checks). measure, and also the data
the checks in checks need. In C checks is a set of semq_check_t
flags (SEMQ_CHECK_PER_ROW = 1; any other bit is InvalidInput), and
checks = 0 is measure. The bindings take their evaluate options:
Python Floor.measure(nulls, per_row=True), Rust
Floor::measure_for(&nulls, &GateOptions::new().per_row(true)), Go
MeasureFloorFor(nulls, GateOptions{PerRow: true}), TypeScript
Floor.measure(nulls, { perRow: true }). The per-row check records:
max_hamming: the largest hamming distance of any changed row of any null (0when no null changed a row);distinct_nulls: how many distinct null states went in. Two nulls are the same state when their candidates have the samecontent_digest(rows, ids and config; the manifest does not count). It is1when every null rebuild gave the same rows.
p99 of m integers is 0 when m = 0, and otherwise the k-th smallest
with k = m - floor(m / 100): nearest rank, with no product that can
overflow.
diff.within(floor). The floor must match the diff's config, id kind
and reference; otherwise Incompatible. With n_common = n_unchanged +
len(changed), the diff is within the floor when all of these hold:
n_common > 0;removedis empty;len(changed) × floor.total_rows <= floor.changed_rows × n_common, compared exactly, never in floating point;- the p99 of the hamming distances over
changedis at mostfloor.hamming(0when nothing changed); manifest_changescontains neitherencodernorencoder_revision.
Added rows and other manifest changes do not affect the verdict. Every null
used to measure a floor is within it. The envelope is taken component by
component, so the ratio may come from one null and the hamming from
another, and the floor can be looser than any single null observed. It
describes what was observed and assumes no model of the rebuild noise. The
core accepts a single null; the semq command asks for three by default,
because one null only shows what that rebuild happened to do.
False alarms. Each check compares one statistic of the candidate with
the largest value among the nulls. If the N nulls and an unchanged
candidate are exchangeable, the candidate exceeds all N with probability
at most 1/(N+1), so each check rejects it with at most that probability:
2/(N+1) for within (checks 3 and 4) and 3/(N+1) with the per-row
check. This bound comes from how the nulls and the candidate are produced,
not from the floor, which promises nothing about a rebuild it has not seen.
With distinct_nulls = 1 an unchanged candidate equals every null, and the
per-row check adds no false alarm.
diff.evaluate(floor, options). The same checks as within, and a
verdict that names every one that failed, not only the first. With no
options, passed equals within. The options select further checks; each
is off by default, so a new one never changes an existing verdict:
- per row (
SEMQ_CHECK_PER_ROW): the verdict also fails when any changed row has a hamming distance abovefloor.max_hamming, androwslists those ids in canonical order. It needs a floor withmax_hamming; otherwiseIncompatible. The p99 of check 4 ignores thefloor(m / 100)most changed rows; this check does not ignore any.
In C the checks are the checks flags of semq_diff_evaluate, the same
semq_check_t flags as measure_for; checks = 0 is within, and an
unknown bit is InvalidInput.
The verdict is passed, reasons and rows. reasons uses these names,
in this order: no_common_rows, removed_rows, changed_ratio,
hamming, encoder (checks 1 to 5), row_above_max (per row).
Floor schema¶
{
"version": "semq-floor/1",
"config": { "operator": "quant" | "phase" | "orbit", "dim": int,
"bins" | "sectors" | "scale": int, "rule_revision": int },
"id_kind": "u64" | "utf8",
"reference_id": hex64,
"nulls": int,
"changed_rows": int,
"total_rows": int,
"hamming": int,
"max_hamming": int, (optional, per-row data)
"distinct_nulls": int (optional, per-row data)
}
The core writes and reads this form (semq_floor_save, semq_floor_load);
every binding calls it, so the rules below are the same everywhere.
Writing. One object, keys in this order, no whitespace, the config as in
the report schema, reference_id as lowercase hex. max_hamming and
distinct_nulls are written only when the floor records them, so a floor
without per-row data is written byte for byte as SEMQ 1.0 writes it.
Reading. Valid JSON in valid UTF-8, one object, nesting at most 64
levels deep. The keys above are read strictly: each at most once, version
exactly "semq-floor/1", counts as JSON integers in [0, 2^64) without
sign, fraction or exponent (no booleans, strings or null), reference_id
as 64 hex characters in either case, and the config with exactly its
operator's parameter. max_hamming and distinct_nulls are optional and
independent of each other. 18446744073709551615 (2^64 - 1) is reserved
in both: it is SEMQ_NONE, which marks an absent key, so a file that
carries it is InvalidInput. Any other key is ignored, at the top level
and inside config. Then the construction rules apply, including
hamming <= max_hamming <= units_per_row and
1 <= distinct_nulls <= nulls. Any violation is InvalidInput.
Evolution. A key added later is optional, so older readers ignore it and
newer readers accept files without it. version changes only when the
meaning of an existing key changes. SEMQ 1.0 read this schema with no other
keys allowed, so it rejects a file that carries max_hamming or
distinct_nulls: upgrade the readers before writing floors with per-row
data. A floor without it reads in every version.