Retrieval quality in detail¶
Explore frozen evaluations on SciFact and FiQA. Select a dataset, model and scoring profile. Each point represents one measured format at its total bits per dimension, including stored scales, codebooks and rotations. The source report for the selected run contains every exact value and method configuration.
nDCG@10 measures relevance against available judgments; Overlap@10 measures preservation of FP32 neighbors. The Δ nDCG interval is a paired query-bootstrap interval against FP32. An interval crossing zero does not establish equivalence; these intervals exclude dataset, model and training-seed variation and have no multiple-comparison correction.
The scoring profiles answer different questions. Normalized direction uses cosine on decoded rows, raw reconstruction uses inner products without renormalization, and native estimators use a backend's own scoring. Do not combine their rankings. No method uses rescoring.
Loading verified reports…
Exact measurements and provenance in the frozen reports:
fiqa · snowflake-arctic-embed-m-v1.5 · fiqa · e5-small-v2 · fiqa · mxbai-embed-large-v1 · scifact · snowflake-arctic-embed-m-v1.5 · scifact · e5-small-v2 · scifact · mxbai-embed-large-v1
What these runs show¶
- Sensitivity varies with the inputs: 2-bit SEMQ quant's mean nDCG difference from FP32 ranges from -0.0681 to -0.0040. These runs do not isolate a cause.
FP32 baseline check¶
This sanity check compares FP32 nDCG@10 with pinned model-card values. The model cards do not pin dataset revisions, so a difference is not evidence of a reproduction failure.
Scope and verification¶
FP32 per-query nDCG was checked against pytrec-eval-terrier on identical rankings in both common profiles (absolute tolerance 1e-12). This checks the shared metric and stored FP32 values, not each codec or the tie policy. Compact reports retain the source revision, input hashes, verification results, storage breakdowns and paired comparisons.
The benchmark overview covers speed, scale, rebuild variation and cross-platform identity. The reproduction guide describes the datasets, preprocessing and evaluation commands.