Skip to content

Commit d5fbf0f

Browse files
Prasad-178claude
andcommitted
docs+bench: DBpedia 1M headline + SIFT validation results (2026-05-12)
Three things land together: 1. SUMMARY.md gains three new sections: - DBpedia 1M @ 1536-dim (commit 4880cae): 92.4 % R@10 @ 1.17 s P50 with PQ-M48-probe8-eps25 + RA=2 + Storage:Memory on m6i.4xlarge. ~9× query speedup + ~19 pp R@10 lift vs pre-fix (RA=1, File-store). - SIFT 1M validation (m6i.2xlarge): post-optimisation reproduce matches historical within ±2 pp sampling noise; PQ builds ~3× faster after the subsample-first fix. - SIFT 100K DBpedia-prep validation (m6i.xlarge): confirms RA=2 as a recall win + documents PCA-64 catastrophe (R@10 97 % → 49 %). 2. CLAUDE.md updated: - Two-line headline (SIFT + DBpedia) replacing single-line claim. - Competitive table gains DBpedia row alongside SIFT. - "Pending → DONE" for DBpedia bench + float32 refactor (partial). - Tested-and-rejected gains PCA-64 + Storage:File entries. - Old "Known-and-deferred work" memory-cliff block removed (resolved). 3. test/pq_sift100k_benchmark_test.go gets explanatory comments on the 5 new configs added during the May 11 SIFT 100K validation. The 4 PCA-64 configs are explicitly marked "tested-and-rejected" guards — kept so future Claude sessions can see the failed experiment without re-running it. Headline DBpedia recipe (for future high-dim semantic embedding work): RedundantAssignments=2 + PQSubspaces=48 + Storage:Memory (the default) + existing all-four-mitigations (σ=2^45, π, PaddingBucketed, ε=2.5). Do NOT use Storage:File or PCA for this regime. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
1 parent 4880cae commit d5fbf0f

3 files changed

Lines changed: 221 additions & 37 deletions

File tree

‎CLAUDE.md‎

Lines changed: 51 additions & 29 deletions
Original file line numberDiff line numberDiff line change
@@ -2,7 +2,13 @@
22

33
Privacy-preserving vector search using CKKS homomorphic encryption + AES-256-GCM
44
+ k-means clustering. Server computes on encrypted data and learns nothing.
5-
Sub-second on 1M vectors, 8 vCPU AWS, 99.8% Recall@10.
5+
Two production-validated points:
6+
- **SIFT 1M (128-dim image descriptors):** sub-second on 1M vectors, 8 vCPU AWS,
7+
99.8 % Recall@10 (357 ms PQ-M8-probe8).
8+
- **DBpedia 1M @ 1536-dim ada-002 (semantic embeddings):** 92.4 % Recall@10 at
9+
P50 = 1.16 s (PQ-M48-probe8) on m6i.4xlarge with RA=2 + Memory storage. The
10+
high-dim semantic regime needs `RedundantAssignments=2` + `Storage:Memory` to
11+
match the SIFT class — confirmed 2026-05-12.
612

713
This file is the entry point for any new Claude session. Skim everything below
814
before touching code. Cross-references point to deeper docs in `docs/`.
@@ -105,7 +111,8 @@ today. See `SECURITY_MODEL.md` §6 for the full competitor comparison.
105111

106112
| System | Threat model | Access pattern | Latency | Recall | Notes |
107113
|---|---|---|---|---|---|
108-
| **Opaque (full mit)** | HBC | Statistical (DP-bound + π) | **464 ms** | **99.8%** | Threshold key (unique) |
114+
| **Opaque (SIFT 1M)** | HBC | Statistical (DP-bound + π) | **357 ms** | **99.8%** | Threshold key (unique) |
115+
| **Opaque (DBpedia 1M ada-002 1536-dim)** | HBC | same | **1.17 s P50** | **92.4%** | sub-2s on semantic embeddings |
109116
| Compass (OSDI '25) | Malicious | Cryptographic (Ring ORAM) | ~600-900ms | high | Single-key, 500MB client mem |
110117
| Pacmann (ICLR '25) | Semi-honest | Cryptographic (PIR) | ~3.1s | ~90% | 100M scale |
111118
| Tiptoe (SOSP '23) | HBC | Cryptographic (PIR) | 2.7s | ~40% MS-MARCO | 360M pages |
@@ -188,13 +195,14 @@ These pieces sit on `main` but have known follow-ups that the next session
188195
should pick up. Both have full design docs / call-site comments — read those
189196
first before extending.
190197

191-
- **DBpedia 1M bench** — Go scaffold landed (`pkg/embeddings/loader.go` +
192-
`test/dbpedia1m_benchmark_test.go` build-tag `dbpedia1m` + the HF parquet →
193-
fvecs converter `scripts/download_dbpedia1m.sh`). Fires via
194-
`deploy/bench-cpu/run_dbpedia_bench.sh`. Optional follow-ups: NC=256 variant
195-
if some future config justifies it (current 2026-05-09 SIFT bench shows
196-
NC=128 wins decisively, so don't pre-emptively add); larger M (M=192) if
197-
ada-002 PQ recall is too lossy.
198+
- **DBpedia 1M bench — DONE 2026-05-12 (commit `4880cae`).** Sub-2s P50 at
199+
92.4 % Recall@10 (PQ-M48-probe8, m6i.4xlarge, RA=2 + Memory). See SUMMARY.md
200+
"DBpedia 1M @ 1536-dim ada-002" section. **Crucial config knobs that
201+
unlocked this**: `Storage:Memory` (not File — File was a stale memory-
202+
saving switch from when build peaked >64 GB), `RedundantAssignments=2`
203+
(fixes curse-of-dim NN-miss; SIFT didn't need it because of low-dim
204+
cluster coherence). Optional follow-ups for further latency work:
205+
M=128 PQ variant (finer codebook); explore `NumKMeansInit=5`.
198206
- **Threshold retry-attack fix** — Phase 1 only (`docs/THRESHOLD_RETRY_FIX.md`
199207
+ `pkg/crypto/threshold/retry_guard.go` standalone). Phases 2 + 3 pending —
200208
wiring `RetryGuard` into `ThresholdDecrypt` (API change: needs `instanceID`
@@ -215,28 +223,42 @@ first before extending.
215223
suggestions — see `deploy/bench-cpu/results/SUMMARY.md` 2026-05-09 section
216224
for the full apples-comparison table.
217225

226+
- **PCADimension=64 on SIFT 128-dim (2026-05-11 SIFT 100K bench).**
227+
Hypothesis was that halving dimensionality via PCA would cut HE +
228+
local-scoring time without much recall hit. Disproven catastrophically:
229+
PCA-64 drops SIFT R@10 from 97 % → **49 %** (50 % collapse). PCA-64 +
230+
RA=2 + PQ all show the same 49 % R@10 floor — RA=2 can't rescue it.
231+
Four PCA configs (`*-PCA64`, `*-PCA64-RA2`, `PQ-M8-PCA64`,
232+
`PQ-M8-PCA64-RA2`) kept in `test/pq_sift100k_benchmark_test.go` as
233+
tested-and-rejected guards. **PCA is not in the production recipe.**
234+
Higher PCA target dims on higher-dim datasets (e.g. ada-002 1536→512)
235+
may behave better — not validated, deferred. See SUMMARY.md "SIFT 100K
236+
— DBpedia-prep config validation" for the full table.
237+
238+
- **Storage:File for DBpedia (commit `1e735ec`, May 10, REVERTED `4880cae`
239+
May 12).** The File-storage switch was meant to offload ~12 GB of
240+
ciphertexts to disk vs the in-memory backend. Five subsequent build-
241+
phase optimisations (`7a9a369` float32 ciphertexts, `828ee0d` float32
242+
storage tier, `8905f2d` drop pendingVectors, `8905f2d` worker chunk-
243+
flush, `e4a2cc7` GOMEMLIMIT) cut build peak from ~64 GB → ~18 GB,
244+
making the offload unnecessary. File storage also added ~30 s/query
245+
of EBS-IOPS-bound blob fetches at probe-16, which was the dominant
246+
latency source. Memory is now the default for DBpedia and any future
247+
high-dim workload that fits in instance RAM. The FileStore mutex /
248+
saveIndex fixes in `dd4fc36` are still valuable for future File-store
249+
callers but no current bench uses them.
250+
218251
## Known-and-deferred work
219252

220-
- **DBpedia 1M @ 1536-dim bench.** Eight EC2 attempts on 2026-05-09/10
221-
uncovered a build-phase memory cliff: peak RSS at 1M × 1536-dim hits
222-
~128 GB on a 64 GB instance and OOM-kills. Three quick optimizations
223-
(commit `1e735ec`) cut that peak to ~64 GB — half the work done.
224-
Remaining headroom requires either (a) `r6i.4xlarge` 128 GB at ~$2/run,
225-
or (b) the float32 refactor (multi-day, halves all hot-path memory →
226-
fits on m6i.2xlarge at $0.38/hr). The Go scaffolding + bench scripts
227-
are all in place; just hold off on firing the AWS bench until the
228-
float32 work lands. Full memory analysis + reasoning + per-attempt
229-
failure log: `docs/MEMORY_PROFILE.md`.
230-
231-
- **float32 refactor.** Replace `[][]float64` with `[][]float32` on the
232-
hot paths (`pkg/embeddings`, `pkg/cluster`, `pkg/pca`, `pkg/pq`,
233-
`pkg/hierarchical`, `pkg/client`, `pkg/auth`, `pkg/enterprise`,
234-
`opaque.go` public API). HE layer (`pkg/crypto`) keeps float64 since
235-
Lattigo's encoder requires it; conversion happens at the narrow HE
236-
encode boundary. **Zero accuracy impact** (ada-002 is float32 native,
237-
SIFT precision well above float32 noise floor); ~30-50 % memory
238-
reduction throughout; mild speedup from halved memory bandwidth.
239-
Public API change (breaking) but acceptable for a research-stage
253+
- **DBpedia 1M @ 1536-dim bench — DONE 2026-05-12** (was the long-running
254+
blocker; see commit `4880cae`). Numbers in SUMMARY.md.
255+
256+
- **float32 refactor — partially DONE.** `pkg/cluster` + `pkg/hierarchical`
257+
+ `db.pendingVectors` + `pkg/encrypt/vector.go` ciphertext encoding all
258+
on float32. Remaining float64 hot-path surfaces: `pkg/pca`, `pkg/pq`,
259+
`pkg/auth`, `pkg/enterprise`, `opaque.go` public-API signatures (still
260+
`[]float64` at the boundary, internal conversion). Public API change
261+
(breaking) but acceptable for a research-stage
240262
codebase with a single primary user.
241263

242264
---

‎deploy/bench-cpu/results/SUMMARY.md‎

Lines changed: 132 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -420,3 +420,135 @@ Latency progression:
420420
and ~10 % more bytes per blob).
421421

422422
Recall identical-to-better at every tier across mitigation deltas.
423+
424+
---
425+
426+
## DBpedia 1M @ 1536-dim ada-002 — Memory + RA=2 (2026-05-12)
427+
428+
The first DBpedia 1M bench to actually produce headline numbers (commit
429+
`4880cae` switched Storage:File → Storage:Memory and added
430+
`RedundantAssignments=2`). Earlier runs OOM'd on memory or wedged on
431+
FileStore-serial-bottlenecks; both root-caused + fixed in the May 11-12
432+
push. Hardware: `m6i.4xlarge` (16 vCPU, 64 GB). All four mitigations
433+
live (σ=2^45 + DecodePublic, permutation π, PaddingBucketed,
434+
TargetEpsilon=2.5). RA=2 = each vector indexed under its top-2 nearest
435+
clusters — standard fix for the curse-of-dimensionality NN-miss problem
436+
on dense semantic embeddings.
437+
438+
### `TestDBpedia1MAccuracy` (NC=128, RA=2, Memory, ε=2.5, σ=2^45)
439+
440+
| Config | Probe | Build | Recall@1 | Recall@10 | Avg query |
441+
|----------|-------|--------|----------|-----------|-----------|
442+
| probe-8 | 6.25 %| 19m32s | 96.0 % | **94.4 %**| 2.08 s |
443+
| probe-16 | 12.5 %| 17m50s | 94.0 % | **96.0 %**| 3.43 s |
444+
445+
### `TestPQ_DBpedia1M` (NC=128, RA=2, Memory, ε=2.5, σ=2^45)
446+
447+
| Config | PQ | Probe | Build | Recall@1 | Recall@10 | Avg query | P50 |
448+
|-------------------------|-----|--------|--------|----------|-----------|-----------|----------|
449+
| **PQ-M48-probe8-eps25** | M48 | 6.25 % | 21m14s | 96.0 % | **92.4 %**| **1.17 s**| **1.16 s** |
450+
| PQ-M96-probe16-eps25 | M96 | 12.5 % | 21m58s | 92.0 % | 92.8 % | 1.67 s | 1.66 s |
451+
452+
`standard-probe8` in the PQ-test config-set was cancelled (redundant
453+
with the accuracy-test probe-8 above — identical configuration).
454+
455+
### Comparison vs pre-fix (RA=1, File-store) DBpedia bench
456+
457+
| Metric | Pre-fix (RA=1, File) | Post-fix (RA=2, Memory) | Δ |
458+
|-----------------|----------------------|--------------------------|--------------|
459+
| probe-8 R@10 | 75.2 % | **94.4 %** | **+19.2 pp** |
460+
| probe-8 query | 18.5 s | **2.08 s** | **9× faster**|
461+
| probe-16 R@10 | 75.0 % | **96.0 %** | **+21.0 pp** |
462+
| probe-16 query | 30.8 s | **3.43 s** | **9× faster**|
463+
| PQ-M48 query | n/a (disk-full fail) | **1.17 s** | **<2 s ✅** |
464+
465+
### Headline DBpedia number
466+
467+
**PQ-M48-probe8-eps25 on m6i.4xlarge: 92.4 % Recall@10 at P50 = 1.16 s**
468+
on 1M × 1536-dim ada-002 semantic embeddings, with all four privacy
469+
mitigations live and tunable ε=2.5 per-query DP bound. Build wall:
470+
21m14s.
471+
472+
Cost: ~$1.20 (m6i.4xlarge × ~1.5 hr).
473+
474+
### Why the previous DBpedia run was so much slower
475+
476+
The May 10 → May 11 wedge + OOM cascade was three independent issues:
477+
478+
1. `FileStore.PutBatch` held a global mutex across all 1024 file
479+
writes → all 16 encrypt workers serialised on one writer.
480+
2. `FileStore.saveIndex` was called per `PutBatch` → ~1000 full JSON
481+
marshals of the ~20 MB index file per build = ~250 sec wasted.
482+
3. `streamPaddingBlobs` called `Put` per padding blob → each Put
483+
eager-flushed the (already huge) index.
484+
485+
All three fixed in `dd4fc36`. With those fixes, File-storage would
486+
have been usable. But the May 12 SIFT 1M validation showed Memory
487+
storage easily fits 1M × 1536-dim on m6i.4xlarge (peak ~18 GB build,
488+
~6.5 GB steady-state) after the additional float32 + drop-pending
489+
optimisations (`7a9a369` + `828ee0d` + `8905f2d`). Memory is now the
490+
right default for DBpedia.
491+
492+
---
493+
494+
## SIFT 1M validation (2026-05-12, m6i.2xlarge)
495+
496+
Pre-DBpedia validation: rerun the canonical SIFT 1M bench on current
497+
HEAD to confirm no regression after the May 11-12 build-phase /
498+
FileStore optimisations.
499+
500+
### `TestSIFT1MAccuracy` (NC=128, σ=2^45) — May 12 reproduce
501+
502+
| Config | Probe | Multi | Recall@1 | Recall@10 | Avg query |
503+
|------------|--------|-------|----------|-----------|-----------|
504+
| strict-4 | 3.1 % | no | 88.0 % | 86.0 % | 248 ms |
505+
| strict-8 | 6.25 % | no | 98.0 % | 96.4 % | 302 ms |
506+
| strict-16 | 12.5 % | no | 100.0 % | 99.8 % | 399 ms |
507+
| probe-8 | 6.25 %+| yes | 100.0 % | 99.8 % | 402 ms |
508+
| probe-16 | 12.5 %+| yes | 100.0 % | 100.0 % | 540 ms |
509+
510+
### `TestPQ_SIFT1M` (NC=128, σ=2^45) — May 12 reproduce
511+
512+
| Config | Recall@1 | Recall@10 | Avg query | P50 | Build |
513+
|------------------------|----------|-----------|-----------|--------|--------|
514+
| **PQ-M8-probe8-eps25** | 100.0 % | 97.6 % | **357 ms**| 361 ms | 2m52s |
515+
| PQ-M8-probe16-eps25 | 100.0 % | 99.8 % | 475 ms | 451 ms | 2m53s |
516+
| PQ-M8-probe8-eps271 | 100.0 % | 97.8 % | 346 ms | 351 ms | 2m58s |
517+
| PQ-M8-probe8-eps20 | 100.0 % | 98.2 % | 419 ms | 420 ms | 2m52s |
518+
| standard-probe8-eps25 | 98.0 % | 99.6 % | 397 ms | 397 ms | 2m13s |
519+
520+
Numbers match historical SIFT 1M within ±2 pp / ±20 ms sampling noise.
521+
**Builds ~3× faster than May 9** (PQ ~8m → ~2m52s; standard ~3m28s →
522+
~2m13s) — driven by the PQ subsample-first fix (`8905f2d`: codebook
523+
k-means now trains on 100 K subset instead of full 1 M, no recall
524+
regression because that's standard PQ literature precedent — Jegou et
525+
al. 2011, FAISS index_factory).
526+
527+
Cost: ~$0.27 (m6i.2xlarge × ~45 min).
528+
529+
---
530+
531+
## SIFT 100K — DBpedia-prep config validation (2026-05-11, m6i.xlarge)
532+
533+
12 configs in one fire (~$0.04, ~10 min wall). Validated:
534+
535+
| Config | Recall@1 | Recall@10 | Avg query | Note |
536+
|-------------------------|----------|-----------|-----------|----------------------------|
537+
| standard (baseline) | 98 % | 97.2 % | 166 ms | reference |
538+
| **standard-RA2** | **100 %**| **99.8 %**| 175 ms | ✅ RA=2 — +2.6 pp R@10 |
539+
| standard-PCA64 | 80 % | **49.0 %**| 144 ms | ❌ PCA-64 destroys SIFT |
540+
| standard-PCA64-RA2 | 80 % | 49.6 % | 162 ms | ❌ RA=2 can't rescue PCA |
541+
| PQ-M8-PCA64 | 80 % | 49.4 % | 140 ms | ❌ PCA kills PQ too |
542+
| PQ-M8-PCA64-RA2 | 80 % | 49.0 % | 159 ms | ❌ full PCA stack broken |
543+
544+
**Two findings**:
545+
1. `RedundantAssignments=2` is a real recall win — even on already-
546+
maxed SIFT 100K it picked up R@1 100 % (was 98 %) and R@10 99.8 %
547+
(was 97.2 %) at +10 ms query cost. Confirmed safe to enable.
548+
2. `PCADimension=64` destroys recall on SIFT 128-dim — going 128 → 64
549+
loses too much information at this scale. **PCA is not in the
550+
DBpedia recipe.** (Higher PCA ratios on higher-dim datasets may
551+
work; not validated here, deferred.)
552+
553+
The PCA configs are kept in the test file as "tested-and-rejected"
554+
guards — future runs should not re-test these.

‎test/pq_sift100k_benchmark_test.go‎

Lines changed: 38 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -92,28 +92,52 @@ func TestPQ_SIFT100K(t *testing.T) {
9292
var results []benchResult
9393

9494
// Configs to test: each with TopClusters=8 (strict) for fair comparison.
95+
// The RA=2 / PCA=64 variants below validate the same latency / recall
96+
// optimisations we want to apply to DBpedia (1M × 1536-dim ada-002),
97+
// where pre-optimisation queries hit 18-30 s and recall stalls at ~75 %.
98+
// On SIFT 100K (128-dim, well-clusterable) the absolute latencies are
99+
// already low (~150 ms), but the *relative speedups* and the
100+
// no-recall-regression check transfer to high-dim semantic workloads.
95101
configs := []struct {
96102
name string
97103
pqM int
98104
topClusters int
99105
probeThresh float64
100106
bernoulli bool
107+
pcaDim int // 0 = disabled
108+
redundant int // 0 → default of 1
101109
}{
102-
{"standard", 0, 8, 1.0, false},
103-
{"PQ-M8", 8, 8, 1.0, false},
104-
{"PQ-M16", 16, 8, 1.0, false},
110+
{"standard", 0, 8, 1.0, false, 0, 0},
111+
{"PQ-M8", 8, 8, 1.0, false, 0, 0},
112+
{"PQ-M16", 16, 8, 1.0, false, 0, 0},
105113
// Also test with multi-probe to see PQ at higher recall.
106-
{"standard-probe16", 0, 16, 0.95, false},
107-
{"PQ-M8-probe16", 8, 16, 0.95, false},
114+
{"standard-probe16", 0, 16, 0.95, false, 0, 0},
115+
{"PQ-M8-probe16", 8, 16, 0.95, false, 0, 0},
108116
// Bernoulli decoy sampling for tight (ε,δ)-DP composition. Same
109117
// expected K_decoy as the uniform-K default, but K is binomial per
110118
// query — recall must be preserved (decoys never affect recall).
111-
{"standard-bernoulli", 0, 8, 1.0, true},
112-
{"PQ-M8-probe16-bernoulli", 8, 16, 0.95, true},
119+
{"standard-bernoulli", 0, 8, 1.0, true, 0, 0},
120+
{"PQ-M8-probe16-bernoulli", 8, 16, 0.95, true, 0, 0},
121+
// --- DBpedia-prep validation configs (2026-05-11) ---
122+
// RA=2 confirmed as a recall win on already-maxed SIFT 100K: bumped
123+
// R@1 98 % → 100 % and R@10 97.2 % → 99.8 % at +10 ms query cost.
124+
// Now the canonical recipe alongside PQ for DBpedia 1M @ 1536-dim
125+
// (commit `4880cae` shipped 92.4 % R@10 / 1.17 s P50 on ada-002).
126+
{"standard-RA2", 0, 8, 1.0, false, 0, 2},
127+
// PCA-64 configs below are TESTED-AND-REJECTED — kept as guards.
128+
// PCA-64 on SIFT 128-dim collapses R@10 from 97 % to **49 %** (50 %
129+
// recall loss). RA=2 + PQ both fail to rescue it. PCA is NOT in the
130+
// production recipe. See CLAUDE.md "Tested-and-rejected variations"
131+
// + SUMMARY.md "SIFT 100K — DBpedia-prep config validation".
132+
{"standard-PCA64", 0, 8, 1.0, false, 64, 0},
133+
{"standard-PCA64-RA2", 0, 8, 1.0, false, 64, 2},
134+
{"PQ-M8-PCA64", 8, 8, 1.0, false, 64, 0},
135+
{"PQ-M8-PCA64-RA2", 8, 8, 1.0, false, 64, 2},
113136
}
114137

115138
for _, cfg := range configs {
116-
t.Logf("--- %s (TopClusters=%d, PQ=%d) ---", cfg.name, cfg.topClusters, cfg.pqM)
139+
t.Logf("--- %s (TopClusters=%d, PQ=%d, PCA=%d, RA=%d) ---",
140+
cfg.name, cfg.topClusters, cfg.pqM, cfg.pcaDim, cfg.redundant)
117141

118142
dbCfg := opaque.Config{
119143
Dimension: dataset.Dimension,
@@ -126,6 +150,12 @@ func TestPQ_SIFT100K(t *testing.T) {
126150
if cfg.pqM > 0 {
127151
dbCfg.PQSubspaces = cfg.pqM
128152
}
153+
if cfg.pcaDim > 0 {
154+
dbCfg.PCADimension = cfg.pcaDim
155+
}
156+
if cfg.redundant > 0 {
157+
dbCfg.RedundantAssignments = cfg.redundant
158+
}
129159

130160
db, err := opaque.NewDB(dbCfg)
131161
if err != nil {

0 commit comments

Comments
 (0)