Skip to content

Commit 41bca06

Browse files
Add wizard step 7: interactive Duplicate Verification viewer
After saving the duplicate file (step 6 now auto-advances on a successful save), every extracted row is re-validated: - Row Key = the rule's duplicate key columns plus every cell to their right, under the scan's trim/case semantics. - Condition A: the Row Key appears multiple times in the original file; the viewer lists the matching original row indices with their first cells (capped at 20). - Condition B: it appears exactly once AND is absent from the reference file (over the mapped columns, reusing the extraction's exact key set - the ref-key db now survives the extraction for this). - Anything else is INVALID: the audit trail for the extraction. Occurrence lookups ride the scan store's dup_entries index; results live in their own on-disk SQLite, so verification streams at any size. The viewer offers per-page size (preset or custom), Previous/Next, go-to row/page, and live search over the first two cells (debounced LIKE). Runs in its own subprocess (VerifyController); new scans/extractions discard stale verification results. New i18n keys in all five locales. Focused adversarial review applied: the right-hand key is now padding-insensitive (ragged source rows audit VALID instead of falsely INVALID - trailing empty cells equal missing cells everywhere), the verification reads the duplicate file without trailing-empty suppression so fully-empty extracted rows are audited too, and unparseable dup-file records surface as malformed_rows in the summary. Local tmp/ data files are excluded and gitignored - user data never belongs in the repository. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017es7LRoYVn9aCwZ8pKwffS
1 parent d6af8ef commit 41bca06

19 files changed

Lines changed: 1136 additions & 19 deletions

‎.gitignore‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -12,3 +12,4 @@ build/
1212
htmlcov/
1313
uv.lock
1414
tests/_fixtures_big/
15+
tmp/

‎README.md‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -40,6 +40,14 @@ Nothing ever leaves your machine.
4040
unique right-hand cells are kept only when an identical row (over the
4141
mapped columns) exists in the reference file. The extraction streams in
4242
its own process and the result is saved wherever you choose.
43+
- **Verification viewer** (step 7, entered automatically after saving):
44+
every extracted row is re-checked against the original file and the
45+
reference — condition A (its content appears several times in the
46+
original, with the matching row indices and first cells listed) or
47+
condition B (appears exactly once and is missing from the reference);
48+
anything else is flagged INVALID. The viewer paginates (selectable or
49+
custom page size), jumps to a row or page, and live-filters on the
50+
first two columns.
4351
- **CLI**: the same engine headless, for pipelines
4452
(exit code `0` = no matches, `1` = matches found, `2` = error).
4553
- **Report**: per-rule match counts with lazily loaded sample rows

‎src/lsa/core/csv_stream.py‎

Lines changed: 11 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -234,9 +234,17 @@ def _iter_raw(self) -> Iterator[Row]:
234234
yield Row(number, cells, offset)
235235
number += 1
236236

237-
def rows(self) -> Iterator[Row]:
238-
"""Yield data rows; the trailing run of blank lines is dropped."""
239-
return suppress_trailing_empty(self._iter_raw())
237+
def rows(self, *, keep_trailing_empty: bool = False) -> Iterator[Row]:
238+
"""Yield data rows; the trailing run of blank lines is dropped.
239+
240+
``keep_trailing_empty`` disables that suppression — used when the
241+
file is machine-written and every record is meaningful (e.g. the
242+
verification step auditing a generated duplicate file).
243+
"""
244+
iterator = self._iter_raw()
245+
if keep_trailing_empty:
246+
return iterator
247+
return suppress_trailing_empty(iterator)
240248

241249
def bytes_read(self) -> int:
242250
if self._lines is not None:

‎src/lsa/core/evaluate.py‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -137,6 +137,11 @@ def right_key_of(cells: Sequence[str], _start: int = right_start) -> str:
137137
if trim:
138138
value = value.strip()
139139
parts.append(value if fold is None else fold(value))
140+
# Padding-insensitive: a missing trailing cell and an empty one
141+
# are the same content (a ragged source row and its padded copy
142+
# in the duplicate file must produce identical keys).
143+
while parts and parts[-1] == "":
144+
parts.pop()
140145
return _KEY_SEPARATOR.join(parts)
141146

142147
return DuplicateCondition(f"{rule_id}{_KEY_SEPARATOR}{index}", key_of, right_key_of)

‎src/lsa/core/extract.py‎

Lines changed: 5 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -175,6 +175,7 @@ def extract_duplicates(
175175
transform = _value_transform(settings)
176176
source_indexes = mapping.source_indexes()
177177

178+
owns_ref_db = ref_db_path is None
178179
if ref_db_path is None:
179180
ref_db_path = make_temp_store_path(prefix="lsa-refkeys-")
180181
ref_db_path = Path(ref_db_path)
@@ -274,7 +275,10 @@ def flush_sub() -> None:
274275
source.close()
275276
store.close()
276277
ref_db.close()
277-
ref_db_path.unlink(missing_ok=True)
278+
if owns_ref_db:
279+
# A caller-provided ref db is kept: the verification step reuses
280+
# the exact key set the extraction ran against.
281+
ref_db_path.unlink(missing_ok=True)
278282

279283
return ExtractStats(
280284
reference_rows=reference_rows,

‎src/lsa/core/locales/de.json‎

Lines changed: 18 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -103,5 +103,22 @@
103103
"dupfile.extracting.groups": "Gruppen werden verarbeitet... {done} / {total}",
104104
"dupfile.done": "{written} Duplikatzeile(n) extrahiert; {kept} Zeile(n) behalten; {groups} Gruppe(n) verarbeitet; Referenz: {reference} Zeile(n).",
105105
"dupfile.cancelled": "Extraktion abgebrochen - die Datei ist unvollständig.",
106-
"dupfile.save": "Duplikatdatei speichern..."
106+
"dupfile.save": "Duplikatdatei speichern...",
107+
"step.verify.title": "Überprüfung",
108+
"step.verify.help": "Jede Zeile der erzeugten Duplikatdatei wird erneut geprüft: Bedingung A bedeutet, ihr Inhalt (Schlüsselspalten plus alles rechts davon) kommt mehrfach in der Originaldatei vor; Bedingung B bedeutet, er kommt genau einmal vor und fehlt in der Referenzdatei. UNGÜLTIGE Zeilen hätten nicht extrahiert werden dürfen.",
109+
"verify.no_data": "Extrahieren und speichern Sie zuerst eine Duplikatdatei - es gibt noch nichts zu überprüfen.",
110+
"verify.running": "Duplikatzeilen werden überprüft... {done}",
111+
"verify.summary": "{total} Zeile(n) überprüft: {valid_a} gültig (Bedingung A), {valid_b} gültig (Bedingung B), {invalid} UNGÜLTIG.",
112+
"verify.status": "Status",
113+
"verify.valid": "GÜLTIG",
114+
"verify.invalid": "UNGÜLTIG",
115+
"verify.condition": "Bedingung",
116+
"verify.occurrences": "Vorkommen",
117+
"verify.matches": "Originalzeilen (Zeile:erste Zelle)",
118+
"verify.filter": "Filter (erste zwei Spalten)",
119+
"verify.page_size": "Pro Seite",
120+
"verify.goto": "Gehe zu",
121+
"verify.goto.row": "Zeile",
122+
"verify.goto.page": "Seite",
123+
"verify.go": "Los"
107124
}

‎src/lsa/core/locales/en.json‎

Lines changed: 18 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -103,5 +103,22 @@
103103
"dupfile.extracting.groups": "Processing groups... {done} / {total}",
104104
"dupfile.done": "{written} duplicate row(s) extracted; {kept} row(s) kept; {groups} group(s) processed; reference: {reference} row(s).",
105105
"dupfile.cancelled": "Extraction cancelled - the file is incomplete.",
106-
"dupfile.save": "Save duplicate file..."
106+
"dupfile.save": "Save duplicate file...",
107+
"step.verify.title": "Verification",
108+
"step.verify.help": "Every row of the generated duplicate file is checked again: condition A means its content (key columns plus everything to their right) appears several times in the original file; condition B means it appears exactly once and is missing from the reference file. INVALID rows should not have been extracted.",
109+
"verify.no_data": "Extract and save a duplicate file first - there is nothing to verify yet.",
110+
"verify.running": "Verifying duplicate rows... {done}",
111+
"verify.summary": "{total} row(s) verified: {valid_a} valid (condition A), {valid_b} valid (condition B), {invalid} INVALID.",
112+
"verify.status": "Status",
113+
"verify.valid": "VALID",
114+
"verify.invalid": "INVALID",
115+
"verify.condition": "Condition",
116+
"verify.occurrences": "Occurrences",
117+
"verify.matches": "Original rows (row:first cell)",
118+
"verify.filter": "Filter (first two columns)",
119+
"verify.page_size": "Per page",
120+
"verify.goto": "Go to",
121+
"verify.goto.row": "Row",
122+
"verify.goto.page": "Page",
123+
"verify.go": "Go"
107124
}

‎src/lsa/core/locales/es.json‎

Lines changed: 18 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -103,5 +103,22 @@
103103
"dupfile.extracting.groups": "Procesando grupos... {done} / {total}",
104104
"dupfile.done": "{written} fila(s) duplicada(s) extraída(s); {kept} fila(s) conservada(s); {groups} grupo(s) procesado(s); referencia: {reference} fila(s).",
105105
"dupfile.cancelled": "Extracción cancelada: el archivo está incompleto.",
106-
"dupfile.save": "Guardar archivo de duplicados..."
106+
"dupfile.save": "Guardar archivo de duplicados...",
107+
"step.verify.title": "Verificación",
108+
"step.verify.help": "Cada fila del archivo de duplicados se comprueba de nuevo: la condición A significa que su contenido (columnas clave más todo lo que está a su derecha) aparece varias veces en el archivo original; la condición B significa que aparece exactamente una vez y falta en el archivo de referencia. Las filas NO VÁLIDAS no deberían haberse extraído.",
109+
"verify.no_data": "Primero extraiga y guarde un archivo de duplicados: aún no hay nada que verificar.",
110+
"verify.running": "Verificando filas duplicadas... {done}",
111+
"verify.summary": "{total} fila(s) verificada(s): {valid_a} válida(s) (condición A), {valid_b} válida(s) (condición B), {invalid} NO VÁLIDA(S).",
112+
"verify.status": "Estado",
113+
"verify.valid": "VÁLIDA",
114+
"verify.invalid": "NO VÁLIDA",
115+
"verify.condition": "Condición",
116+
"verify.occurrences": "Apariciones",
117+
"verify.matches": "Filas originales (fila:primera celda)",
118+
"verify.filter": "Filtro (dos primeras columnas)",
119+
"verify.page_size": "Por página",
120+
"verify.goto": "Ir a",
121+
"verify.goto.row": "Fila",
122+
"verify.goto.page": "Página",
123+
"verify.go": "Ir"
107124
}

‎src/lsa/core/locales/fr.json‎

Lines changed: 18 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -103,5 +103,22 @@
103103
"dupfile.extracting.groups": "Traitement des groupes... {done} / {total}",
104104
"dupfile.done": "{written} ligne(s) en double extraite(s) ; {kept} ligne(s) conservée(s) ; {groups} groupe(s) traité(s) ; référence : {reference} ligne(s).",
105105
"dupfile.cancelled": "Extraction annulée - le fichier est incomplet.",
106-
"dupfile.save": "Enregistrer le fichier des doublons..."
106+
"dupfile.save": "Enregistrer le fichier des doublons...",
107+
"step.verify.title": "Vérification",
108+
"step.verify.help": "Chaque ligne du fichier des doublons est recontrôlée : la condition A signifie que son contenu (colonnes clés et tout ce qui est à leur droite) apparaît plusieurs fois dans le fichier d'origine ; la condition B signifie qu'il apparaît exactement une fois et est absent du fichier de référence. Les lignes INVALIDES n'auraient pas dû être extraites.",
109+
"verify.no_data": "Extrayez et enregistrez d'abord un fichier des doublons - il n'y a encore rien à vérifier.",
110+
"verify.running": "Vérification des lignes en double... {done}",
111+
"verify.summary": "{total} ligne(s) vérifiée(s) : {valid_a} valide(s) (condition A), {valid_b} valide(s) (condition B), {invalid} INVALIDE(S).",
112+
"verify.status": "Statut",
113+
"verify.valid": "VALIDE",
114+
"verify.invalid": "INVALIDE",
115+
"verify.condition": "Condition",
116+
"verify.occurrences": "Occurrences",
117+
"verify.matches": "Lignes d'origine (ligne:première cellule)",
118+
"verify.filter": "Filtre (deux premières colonnes)",
119+
"verify.page_size": "Par page",
120+
"verify.goto": "Aller à",
121+
"verify.goto.row": "Ligne",
122+
"verify.goto.page": "Page",
123+
"verify.go": "OK"
107124
}

‎src/lsa/core/locales/pt.json‎

Lines changed: 18 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -103,5 +103,22 @@
103103
"dupfile.extracting.groups": "A processar grupos... {done} / {total}",
104104
"dupfile.done": "{written} linha(s) duplicada(s) extraída(s); {kept} linha(s) mantida(s); {groups} grupo(s) processado(s); referência: {reference} linha(s).",
105105
"dupfile.cancelled": "Extração cancelada - o ficheiro está incompleto.",
106-
"dupfile.save": "Guardar ficheiro de duplicados..."
106+
"dupfile.save": "Guardar ficheiro de duplicados...",
107+
"step.verify.title": "Verificação",
108+
"step.verify.help": "Cada linha do ficheiro de duplicados é verificada novamente: a condição A significa que o seu conteúdo (colunas-chave mais tudo o que está à direita) aparece várias vezes no ficheiro original; a condição B significa que aparece exatamente uma vez e está ausente do ficheiro de referência. As linhas INVÁLIDAS não deveriam ter sido extraídas.",
109+
"verify.no_data": "Extraia e guarde primeiro um ficheiro de duplicados - ainda não há nada para verificar.",
110+
"verify.running": "A verificar linhas duplicadas... {done}",
111+
"verify.summary": "{total} linha(s) verificada(s): {valid_a} válida(s) (condição A), {valid_b} válida(s) (condição B), {invalid} INVÁLIDA(S).",
112+
"verify.status": "Estado",
113+
"verify.valid": "VÁLIDA",
114+
"verify.invalid": "INVÁLIDA",
115+
"verify.condition": "Condição",
116+
"verify.occurrences": "Ocorrências",
117+
"verify.matches": "Linhas originais (linha:primeira célula)",
118+
"verify.filter": "Filtro (duas primeiras colunas)",
119+
"verify.page_size": "Por página",
120+
"verify.goto": "Ir para",
121+
"verify.goto.row": "Linha",
122+
"verify.goto.page": "Página",
123+
"verify.go": "Ir"
107124
}

0 commit comments

Comments
 (0)