Benchmarks¶
How SEMQ's operators compare with the vector formats a retrieval pipeline already uses: how much search quality survives, how many bytes each vector takes, and how fast it encodes. Every number comes from a frozen result file with its inputs, commit and hardware.
Search quality kept¶
For each format, the line spans the lowest and the highest nDCG@10, relative to FP32, across three embedding models and two datasets (SciFact and FiQA). Formats are ordered by storage, fewest bits per dimension first. SEMQ is in blue.
Table view
| Format | Bits per dimension | Lowest | Highest |
|---|---|---|---|
| Binary, sentence-transformers | 1 | 38.4% | 98.7% |
| SEMQ phase · 16 sectors | 2 | 75.5% | 99.6% |
| SEMQ quant · 2 bits | 2 | 81.8% | 99.4% |
| Faiss PQ · 2 bits | 2.0–2.1 | 94.9% | 100.5% |
| TurboQuant · 2 bits | 2.3–8.4 | 81.9% | 99.9% |
| SEMQ quant · 3 bits | 3 | 94.7% | 100.5% |
| SEMQ quant · 4 bits | 4 | 98.8% | 100.0% |
| RaBitQ · 4 bits | 4.6–10.5 | 99.5% | 101.0% |
| SEMQ orbit · scale 50 | 8 | 97.4% | 100.2% |
| INT8, sentence-transformers | 8 | 99.8% | 100.4% |
A value above 100% means the ranking scored slightly better against the available relevance judgments; it is not a quality improvement in general. Queries search decoded rows by cosine, with no rescoring.
Explore one run¶
Pick a model and a dataset to see every measured format at its storage cost.
Loading measured results…
Storage per vector¶
Bytes per 768-dimension vector, codes only. A .semq file adds an 8-byte id per row and a fixed header; Faiss PQ also stores a codebook once.
Table view
| Format | Bytes per vector | Smaller than FP32 | GB for 100M vectors |
|---|---|---|---|
| Binary | 96 | 32.0× | 9.6 |
| SEMQ quant · 2 bits | 192 | 16.0× | 19.2 |
| SEMQ phase · 16 sectors | 192 | 16.0× | 19.2 |
| Faiss PQ | 192 | 16.0× | 19.2 |
| SEMQ quant · 3 bits | 288 | 10.7× | 28.8 |
| SEMQ quant · 4 bits | 384 | 8.0× | 38.4 |
| SEMQ orbit · scale 50 | 768 | 4.0× | 76.8 |
| INT8 | 772 | 4.0× | 77.2 |
| FP16 | 1,536 | 2.0× | 153.6 |
| FP32 | 3,072 | 1.0× | 307.2 |
Encode and decode speed¶
Vectors per second for 768-dimension rows, measured from Python on one arm64 CPU with the neon backend, median of five runs. Encode includes the unit-norm check and the state's digests.
Table view
| Format | Encode, vectors/s | Decode, vectors/s |
|---|---|---|
| SEMQ quant · 2 bits | 875,934 | 1,165,874 |
| SEMQ quant · 3 bits | 798,074 | 902,340 |
| SEMQ quant · 4 bits | 795,248 | 712,510 |
| SEMQ phase · 16 sectors | 691,664 | 6,473,382 |
| SEMQ orbit · scale 50 | 618,763 | 2,319,688 |
| NumPy binary | 6,536,037 | 3,060,038 |
| NumPy fp16 | 22,920,225 | 25,037,820 |
| NumPy int8 | 1,472,129 | 3,311,098 |
The table also lists plain NumPy casts to FP16, INT8 and binary for scale: they convert arrays and compute no identity.
Choose a configuration¶
| Configuration | Bits per dimension | Bytes per vector (768) | Search quality kept | Encode, vectors/s |
|---|---|---|---|---|
| SEMQ quant · 2 bits | 2 | 192 | 81.8–99.4% | 876k |
| SEMQ quant · 3 bits | 3 | 288 | 94.7–100.5% | 798k |
| SEMQ quant · 4 bits | 4 | 384 | 98.8–100.0% | 795k |
| SEMQ phase · 16 sectors | 2 | 192 | 75.5–99.6% | 692k |
| SEMQ orbit · scale 50 | 8 | 768 | 97.4–100.2% | 619k |
Quality data · Storage data · Throughput data · Extended evaluation with uncertainty intervals · Methodology
Other measurements¶
- Rebuild detection: does the gate pass a rebuild that changed nothing, and fail one that did?
- Scale and portability: a million rows, and the same bytes on every platform.