Design and Implementation of an AI-Driven Multilingual Crime Behaviour and Digital-Forensic Evidence Analysis Platform
Research console: published NCRB Crime in India tables (53 metros, national heads 2021–2024), a 32-agency Indian institutional register (CBI, NIA, ED, I4C, CERT-In, NFSU, CCTNS…), UNODC homicide comparators, a paraphrastic Indic crime-text corpus, a hybrid NLP / digital-forensic engine, and Grok for grounded interpretation.
This is not an operational police system. Synthetic narratives contain no real complainants. Official counts should be verified against NCRB PDFs before journal citation.
- No public multilingual FIR corpus — template synthetics leak 99% TF-IDF accuracy.
- NCRB aggregates hide modus operandi and digital exhibits.
- English-centric DFIR drops Hinglish / Indic IOCs.
- Incomplete IPC/BNS/IT Act ↔ ICCS mapping.
- No chain-of-custody protocol for LLM-derived exhibits.
- Uneven state forensic capacity (NFSU/ORF 2024–25).
- I4C complaints, NCRB cyber cases, and RBI fraud are three unjoined thermometers.
- Low-resource languages (Odia, Assamese, Kashmiri, Manipuri) dropped from “multilingual” papers.
- Gendered cybercrime flattened into “cyber”.
- Hallucination / groundedness unmeasured on Indic evidence.
- TanStack Start, React 19, Tailwind v4, Recharts
- Postgres (Neon in production, PGLite in preview)
- Symbolic NLP (script LID, gazetteer NER, NCRB-head classifier)
- xAI Grok (
grok-4.5) for user-initiated analysis, translation, behaviour narrative, paper assistant
Earlier scikit-learn / DistilBERT experiments and raw NCRB XLSX live at adityaankana1807/india-crime-forensic-platform. NyayaLens is the paper-facing product: honest evaluation notes, city-level official tables, and a harder corpus.
Research code: use freely with attribution. Do not deploy as a decision-making tool for arrest, charge, or bail.