This project builds a domain-adapted, fine-tuned LLM to predict outcomes of Supreme Court cases in India, using real legal texts from 2023–2024. It combines web scraping, case summarization, retrieval-augmented generation (RAG), continuous pretraining, fine-tuning, and robust evaluation to build a specialized legal reasoning model.
| Step | Description |
|---|---|
| Data Extraction | Scraped Supreme Court case links and case texts from Indian Kanoon |
| Case Summarization | Summarized cases into structured sections using a local LLaMA3.1-8B model |
| Law Retrieval (RAG) | Retrieved relevant Indian Penal Code laws using vector search (Qdrant) |
| Pretraining & Fine-tuning | Domain-adapted the LLM and trained it to predict judgments |
| Evaluation | Evaluated predictions using faithfulness, ROUGE, BLEU, and BERT-based metrics |
.
├── links.py
├── case-extraction.py
├── data-preprocessing.ipynb
├── rag.ipynb
├── pretraining+finetuning.ipynb
├── evaluation.ipynb
├── data/
│ ├── 2024_cases/
│ └── 2023_cases/
└── models/
└── llama3.1-8b/- Website: Indian Kanoon
- Court: Supreme Court of India
- Training Data: Cases from 2024
- Testing Data: Cases from 2023
- Base Model: LLaMA 3.1 8B
- Techniques:
- Summarization into
FACTS,ARGUMENTS,OBSERVATIONS,JUDGMENT - Retrieval-Augmented Generation (RAG) with IPC
- Instruction–Input–Output formatting
- Continuous pretraining and task-specific fine-tuning
- Summarization into
| Metric | Score |
|---|---|
| Faithfulness | 0.7064 |
| ROUGE-1 | 0.4797 |
| ROUGE-2 | 0.2797 |
| ROUGE-L | 0.3613 |
| BLEU | 0.2139 |
| BERTScore | 0.8862 |
| Sentence-BERT Cosine | 0.7064 |
These scores indicate strong semantic and factual alignment with true outcomes.
- Python 3.10+
- LLaMA 3.1–8B (local inference)
- Qdrant (vector DB)
- PyPDF (for IPC PDF parsing)
- HuggingFace metrics: ROUGE, BLEU, BERTScore, Sentence-BERT
# Clone the repository
git clone https://github.com/yourusername/indian-case-outcome-predictor.git
cd indian-case-outcome-predictor
# Install dependencies
pip install -r requirements.txt
# Run steps
python links.py
python case-extraction.py
# Run notebooks in order
1. data-preprocessing.ipynb
2. rag.ipynb
3. pretraining+finetuning.ipynb
4. evaluation.ipynb{
"instruction": "Predict the outcome of the case.",
"input": "[Summarized case text + relevant IPC sections]",
"output": "Appeal Dismissed"
}- Support more courts (High Courts, Tribunals)
- Legal precedent tracing & citation support
- Web interface for legal researchers
This project is licensed under the MIT License.