Skip to content

Python API

The root exports sixteen names: Operator, CodecConfig, Codec, Encoding, Diff, Verdict, Floor, BuildInfo, build_info, the six errors and __version__. Start with the quickstart for a complete program.

Conventions

  • Codec.encode takes the vectors first and the rest by name: encode(vectors, ids=..., manifest=None, id_kind=None). vectors must be a float32 array of shape [n, dim] (float64 is rejected, a non-contiguous view is copied); ids a sequence of ints (u64) or strs (utf8); id_kind is required when n = 0.
  • Encoding.rows and get(id) are read-only zero-copy views that keep the native memory alive; ids is a read-only uint64 view for u64 and a tuple of str for utf8. Iteration yields (id, row) in canonical order.
  • save(target) takes a path (written atomically) or a binary stream; atomic replacement does not promise durability of the directory entry after a system crash. Partial stream writes are completed; a stream that cannot advance raises BlockingIOError. Encoding.load(source) takes a path, a binary stream or bytes.
  • Errors are the six classes below; out of memory raises MemoryError and file I/O raises OSError. Objects are released by the garbage collector.
  • Native calls release the GIL.
  • str() of an Encoding, a Diff or a Floor is a one-line summary, the same text in every binding; repr() is the debugging form.

Operator

Bases: IntEnum

The three operators. Values are pinned in the canonical form.

CodecConfig

CodecConfig(_c: Any)

An immutable value: operator, dimension and the operator's parameter.

Build one with :meth:quant, :meth:phase or :meth:orbit. Equality is equality of the 13 canonical bytes.

operator property

operator: Operator

Which rule applies: Operator.QUANT, PHASE or ORBIT.

dim property

dim: int

Coordinates per vector; every row encoded under this config has this length.

rule_revision property

rule_revision: int

Revision of the operator's symbol mapping (p2 of the canonical form), currently 0.

It increments only when an operator rule changes; configs with different revisions are incompatible.

bins property

bins: int

quant only: magnitude bins per sign, in [2, 64]. AttributeError for other operators.

sectors property

sectors: int

phase only: angular sectors per coordinate pair, in [2, 256]. AttributeError otherwise.

scale property

scale: int

orbit only: the scale applied before taking the symbol, in [1, 2^30]. AttributeError otherwise.

bytes_per_vector property

bytes_per_vector: int

Packed size of one row in bytes, computed by the core; the second dimension of Encoding.rows.

units_per_row property

units_per_row: int

Symbols per row: dim for quant and orbit, dim / 2 for phase. The upper bound of a hamming distance.

max_magnitude property

max_magnitude: float

quant only: (float)(2.0 / sqrt(dim)), computed by the core.

parameter_name property

parameter_name: str

The operator's parameter key as used in reports: "bins", "sectors" or "scale".

quant classmethod

quant(dim: int, bins: int) -> CodecConfig

quant: sign and magnitude bin per coordinate; bins in [2, 64].

phase classmethod

phase(dim: int, sectors: int) -> CodecConfig

phase: angular sector per coordinate pair; sectors in [2, 256].

orbit classmethod

orbit(dim: int, scale: int = 50) -> CodecConfig

orbit: digital-root symbol per coordinate; scale in [1, 2^30].

from_bytes classmethod

from_bytes(data: bytes) -> CodecConfig

Parse the 13-byte canonical form.

to_bytes

to_bytes() -> bytes

The 13-byte canonical form.

as_dict

as_dict() -> dict[str, Any]

{"operator", "dim", <parameter>, "rule_revision"} as in reports.

Codec

Codec(config: CodecConfig)

An immutable encoder for one :class:CodecConfig, safe to share between threads.

config property

config: CodecConfig

The rule this codec applies; every Encoding it produces carries the same config.

backend property

backend: str

The kernel the core runs for this operator on this host.

quant classmethod

quant(dim: int, bins: int) -> Codec

A codec for CodecConfig.quant(dim, bins): sign and magnitude bin per coordinate.

phase classmethod

phase(dim: int, sectors: int) -> Codec

A codec for CodecConfig.phase(dim, sectors): angular sector per coordinate pair.

orbit classmethod

orbit(dim: int, scale: int = 50) -> Codec

A codec for CodecConfig.orbit(dim, scale): one discrete symbol per coordinate.

encode

encode(
    vectors: NDArray[float32] | None,
    *,
    ids: IdsInput,
    manifest: dict[str, str] | None = None,
    id_kind: str | None = None
) -> Encoding

Encode vectors (float32, [n, dim], unit-norm) under ids.

id_kind is inferred from the first id and required when n = 0. Row errors carry the input row index.

decode

decode(encoding: Encoding) -> NDArray[np.float32]

Representatives, float32 [n, dim], row i for encoding.ids[i]. Not normalized.

unpack

unpack(encoding: Encoding) -> NDArray[np.uint8]

Symbols, uint8 [n, units_per_row].

Encoding

Encoding(
    ids: IdsInput,
    rows: NDArray[uint8] | None,
    config: CodecConfig,
    manifest: dict[str, str] | None = None,
    id_kind: str | None = None,
)

An immutable set of ids with one canonical row each, a manifest and two identities.

content_digest identifies the rows under the rule; state_id adds the manifest. Equality of two Encodings is equality of state_id.

Build from rows the caller already holds (uint8 [n, bytes_per_vector]).

config property

config: CodecConfig

The rule the rows were encoded under; diff and concat require equal configs.

id_kind property

id_kind: str

"u64" or "utf8": the kind of every id in this Encoding (never mixed).

rows property

rows: NDArray[uint8]

Canonical rows in id order: a read-only uint8 view [n, bytes_per_vector].

ids property

ids: IdsView

Sorted ids: a read-only uint64 view for u64, a tuple of str for utf8.

manifest property

manifest: dict[str, str]

The declared string pairs, stored verbatim and never verified.

encoder and encoder_revision are conventions: a diff reports changes to them and the gate fails when they change.

content_digest property

content_digest: bytes

SHA-256 (32 bytes) of the config, the id kind, the count, the ids and the rows.

Answers "same rows under the same rule", whatever the manifest says.

state_id property

state_id: bytes

SHA-256 (32 bytes) of content_digest and the manifest: the identity of the state as declared.

Two Encodings are equal when their state_id is equal.

load classmethod

load(source: Source) -> Encoding

Parse a file image from a path, a binary stream or bytes. Always copies.

save

save(target: str | PathLike[str] | BinaryWriter) -> None

Write to a path (atomic replacement) or binary stream.

Streams report bytes written; no progress raises BlockingIOError. Atomic replacement does not promise durability of the directory entry.

get

get(id: InputId) -> NDArray[np.uint8]

The row of id (a read-only view). KeyError when absent.

concat

concat(*others: Encoding) -> Encoding

Merge with Encodings of the same config, kind and manifest and disjoint ids.

diff

diff(candidate: Encoding) -> Diff

self is the reference, candidate the state under review.

Diff

Diff()

The result of reference.diff(candidate).

Keeps the rows it needs alive until it is released, even if the caller drops its Encodings.

config property

config: CodecConfig

The config shared by the reference and the candidate.

id_kind property

id_kind: str

"u64" or "utf8", shared by both sides.

reference_id property

reference_id: bytes

state_id of the reference Encoding (32 bytes); a floor applies only to diffs of the same reference.

candidate_id property

candidate_id: bytes

state_id of the candidate Encoding (32 bytes).

added property

added: list[Id]

Ids only in the candidate, canonical order.

removed property

removed: list[Id]

Ids only in the reference, canonical order.

changed property

changed: list[tuple[Id, int]]

(id, hamming) for ids on both sides with different rows, canonical order.

n_unchanged property

n_unchanged: int

Ids on both sides whose rows are byte-identical. n_unchanged + len(changed) is the number of shared rows.

manifest_changes property

manifest_changes: dict[str, tuple[str | None, str | None]]

key -> (before | None, after | None) for keys that differ or exist on one side.

units

units(id: InputId) -> list[tuple[int, int, int]]

(unit, symbol_reference, symbol_candidate) for every unit of id that differs.

Empty for an identical row; InvalidInput when id is absent from either side.

within

within(floor: Floor) -> bool

True iff the candidate is within floor (exact integer arithmetic in the core).

The floor must have been measured with this diff's config, id kind and reference; otherwise Incompatible. A change to encoder or encoder_revision is never within.

evaluate

evaluate(floor: Floor, *, per_row: bool = False) -> Verdict

The verdict of floor on this diff, with every check that failed.

With no options, evaluate(floor).passed == within(floor). With per_row=True the verdict also fails when any changed row has a hamming above floor.max_hamming, and rows lists those ids; a floor without per-row data (one from Floor.measure without per_row=True, or saved by SEMQ 1.0) raises Incompatible.

as_dict

as_dict() -> dict[str, Any]

The report schema: digests as hex, u64 ids as decimal strings, no floats.

Verdict

Verdict(
    passed: bool, reasons: tuple[str, ...], rows: list[Id]
)

The result of Diff.evaluate: passed, the checks that failed, and the rows above max_hamming.

reasons names every failed check, in this order: no_common_rows, removed_rows, changed_ratio, hamming, encoder, row_above_max. rows is empty unless the per-row check ran.

passed instance-attribute

passed: bool = passed

reasons instance-attribute

reasons: tuple[str, ...] = reasons

rows instance-attribute

rows: list[Id] = rows

as_dict

as_dict() -> dict[str, Any]

{"passed", "reasons", "rows"} with ids as strings, as in reports.

Floor

Floor(
    config: CodecConfig,
    *,
    id_kind: str,
    reference_id: bytes,
    nulls: int,
    changed_rows: int,
    total_rows: int,
    hamming: int
)

An envelope of observed variation: (changed_rows, total_rows, hamming) measured from null rebuilds, bound to the config, the id kind and the reference state those nulls were taken against.

A diff is within the floor when it removes no rows, shares at least one row with its reference, changes at most changed_rows / total_rows of the shared rows, the nearest-rank p99 of its changed-row hamming distances does not exceed hamming, and it changes neither encoder nor encoder_revision. Added rows do not affect the verdict. Applying a floor to a diff of another config, id kind or reference raises Incompatible. The floor describes what was observed; it claims no probabilistic coverage of the next rebuild.

Floor.measure(nulls, per_row=True) also records the per-row data for the per-row check of Diff.evaluate: max_hamming, the largest hamming of any changed row of any null, and distinct_nulls. Without it, or for a floor saved by SEMQ 1.0, both are None and the JSON form is the one SEMQ 1.0 reads. The constructor never records them.

The JSON form is written and read by the core: save/load and as_dict/from_dict apply the same rules in every binding.

config property

config: CodecConfig

The config the nulls were encoded under; a diff of another config is Incompatible with this floor.

id_kind property

id_kind: str

"u64" or "utf8": the id kind the floor was measured on.

reference_id property

reference_id: bytes

state_id (32 bytes) of the reference every null was diffed against. The floor applies only to diffs of that reference.

nulls property

nulls: int

How many null diffs the floor was measured from (at least 1).

changed_rows property

changed_rows: int

Numerator of the floor's ratio: changed rows in the null with the largest changed_rows / total_rows.

total_rows property

total_rows: int

Denominator of the floor's ratio: rows shared with the reference in that same null (at least 1).

hamming property

hamming: int

The largest per-null p99 of changed-row hamming distances (0 when no null changed a row).

p99 is nearest-rank: of m distances, the m - m // 100-th smallest, so it equals the maximum below 100 changed rows. A candidate's p99 must not exceed it. The ratio and this bound may come from different nulls.

max_hamming property

max_hamming: int | None

The largest hamming of any changed row of any null (0 when no null changed a row).

Diff.evaluate(floor, per_row=True) flags every changed row above it. None for a floor without per-row data: one from the constructor, from measure without per_row=True, or read from JSON without it.

distinct_nulls property

distinct_nulls: int | None

How many of the nulls were distinct states (1 to nulls); None without per-row data.

Two nulls are the same state when their candidates have the same content_digest. With 1, every null rebuild gave the same rows and the per-row check adds no false alarm; when the nulls vary, each check can reject an unchanged rebuild with probability up to 1 / (nulls + 1).

measure classmethod

measure(
    null_diffs: Iterable[Diff], *, per_row: bool = False
) -> Floor

The envelope of one or more null diffs of one reference; every input is within the result.

With per_row=True the floor also records max_hamming and distinct_nulls, which Diff.evaluate(floor, per_row=True) needs. SEMQ 1.0 cannot read a floor saved with them; without them the saved form is the one SEMQ 1.0 writes.

as_dict

as_dict() -> dict[str, Any]

The floor schema as a dict: the config as in reports, the reference id as hex.

max_hamming and distinct_nulls are present only when the floor records them.

from_dict classmethod

from_dict(data: object) -> Floor

The inverse of as_dict, by the core's rules for the floor schema.

Unknown keys are ignored; known keys are checked strictly (integers only, no booleans or floats). Anything else is InvalidInput.

save

save(target: str | PathLike[str] | TextWriter) -> None

Write the floor's JSON form, and a newline, to a path or a text stream.

load classmethod

load(source: Source) -> Floor

Read the floor's JSON form from a path, a text or binary stream, or bytes.

BuildInfo dataclass

BuildInfo(
    sdk_version: str,
    core_version: str,
    backend: dict[str, str],
    build_id: str,
)

Versions and kernel selection of the loaded core.

backend maps each operator name to the kernel the core runs for it on this host. build_id identifies a reproducible build recipe; it is not a provenance proof.

sdk_version instance-attribute

sdk_version: str

core_version instance-attribute

core_version: str

backend instance-attribute

backend: dict[str, str]

build_id instance-attribute

build_id: str

as_dict

as_dict() -> dict[str, Any]

The four fields as plain JSON-ready values, as semq version --json prints them.

build_info

build_info() -> BuildInfo

Query the loaded core.

Next steps