Skip to content

Latest commit

Β 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

🧬 Protein Stability Prediction Pipeline

Python 3.11+ License: MIT Open Source

A production-grade machine learning pipeline for predicting and optimizing protein thermal stability through rational mutation design. Combines biophysical feature engineering, multi-engine mutation generation (RAG, TRIZ-LLM, MSA consensus), and evolutionary multi-objective optimization.

Note: This repository includes a Lite Mode optimized for 16GB RAM devices. The full pipeline is designed for HPC/cloud environments with significantly larger datasets and computation.


πŸ—οΈ Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                              ORCHESTRATOR (run_all.py)                      β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                       β”‚
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
    β–Ό              β–Ό                        β–Ό                      β–Ό             β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚Phase 1 β”‚    β”‚ Phase 2   β”‚    β”‚        Phase 3           β”‚    β”‚ Phase 4 β”‚   β”‚Phase 5 β”‚
β”‚ Data   │──▢│ Feature   β”‚ ──▢│ β”Œβ”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β” │──▢│ NSGA-II │──▢│Validateβ”‚
β”‚Ingest  β”‚    β”‚Engineeringβ”‚    β”‚ β”‚Engine β”‚Engine β”‚Engineβ”‚ β”‚    β”‚Optimize β”‚   β”‚& SHAP  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β”‚ β”‚A: RAG β”‚B: TRIZβ”‚C: MSAβ”‚ β”‚    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β””β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”˜ β”‚
                               β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Pipeline Components

Phase Component Description
1 Data Ingestion FireProtDB loader + Saboteur synthetic data generator
2 Feature Engineering 14D biophysical vectors (RSA, B-factor, BLOSUM, Ξ”Volume, etc.)
3A Engine A: RAG Literature mining via ChromaDB + sentence-transformers
3B Engine B: TRIZ Inventive problem-solving + local LLM (Ollama)
3C Engine C: MSA Consensus mutations from multiple sequence alignments
4 Optimization XGBoost surrogate + NSGA-II multi-objective evolution
5 Validation Steric clash filter, Domain of Applicability, SHAP explainability

⚑ Full Mode vs Lite Mode

Parameter Full Mode Lite Mode Impact
PDB structures 50+ 20 Dataset size
Mutants/protein 10 5 Training data
NSGA-II generations 50 20 Optimization depth
Population size 100 30 Search space
Engine B (TRIZ) LLM-based Skipped / Mocked Resource usage
Typical runtime 2-4 hours ~2 minutes -
RAM required 32GB+ 16GB -

The Lite Mode demonstrates the full pipeline architecture on reduced data. For production use with novel proteins, run Full Mode on appropriate hardware.


πŸ“Š Sample Results (Lite Mode Demo)

Pipeline output for Lipase A (PDB: 1ISP) target:

Rank Mutations Stability Score Conservation Confidence
1 G14P; G30P; G111P; G158P; G13A; G176P; G45V; G145P; G67P; A132I 12.16 0.72 HIGH
2 G111P; G14P; G158P; A132I; G67P; G30A; G176P; G145P 11.48 0.73 HIGH
3 G176A; G145P; G13P; G30A; G111P; G158P; G14P; G67P; A132I 11.15 0.73 HIGH

Key insight: Pipeline correctly identifies glycine→proline substitutions for rigidifying flexible loops, a well-established thermostabilization strategy.


πŸ’» Requirements

  • OS: Windows 10/11, Linux, macOS
  • Python: 3.11+
  • RAM: 16GB (Lite) / 32GB+ (Full)
  • Optional: Ollama for Phase 3B LLM features

πŸš€ Quick Start

1. Setup Environment

Windows:

setup_env.bat

Linux/macOS:

chmod +x setup_env.sh && ./setup_env.sh

Manual:

python -m venv venv
source venv/bin/activate  # Windows: venv\Scripts\activate
pip install -r requirements.txt

2. (Optional) Enable LLM Features

# Install Ollama from https://ollama.ai, then:
ollama pull qwen3:4b

3. Run Pipeline

python run_all.py

Or run individual phases:

python main_phase1.py   # Data ingestion
python main_phase2.py   # Feature engineering
python main_phase3_a.py # RAG engine
python main_phase3_b.py # TRIZ engine (requires Ollama)
python main_phase3_c.py # Consensus engine
python main_phase4.py   # Evolutionary optimization
python main_phase5.py   # Validation & explainability

πŸ“ Project Structure

β”œβ”€β”€ config_lite.py          # 16GB RAM settings (edit for Full Mode)
β”œβ”€β”€ run_all.py              # Master orchestrator
β”œβ”€β”€ main_phase*.py          # Phase entry points
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ ingestion.py        # FireProtDB + PDB fetching
β”‚   β”œβ”€β”€ saboteur.py         # Synthetic destabilizing mutations
β”‚   β”œβ”€β”€ physics_engine.py   # SASA/RSA calculations
β”‚   β”œβ”€β”€ features/           # Biophysical feature extractors
β”‚   β”œβ”€β”€ engines/            # 3 mutation generation engines
β”‚   β”‚   β”œβ”€β”€ engine_a/       # RAG (ChromaDB)
β”‚   β”‚   β”œβ”€β”€ engine_b/       # TRIZ + LLM
β”‚   β”‚   └── engine_c/       # MSA consensus
β”‚   β”œβ”€β”€ optimization/       # XGBoost + NSGA-II
β”‚   └── validation/         # Steric checks + SHAP
β”œβ”€β”€ data/                   # Input data (auto-populated)
β”œβ”€β”€ results/                # Output candidates + visualizations
└── docs/                   # Additional documentation

πŸ“€ Output Files

File Description
results/FINAL_CANDIDATES_VALIDATED.csv Top mutation candidates with confidence scores
results/phase4_candidates.csv All Pareto-optimal candidates
results/viz/*.pml PyMOL visualization scripts
results/plots/*.png SHAP explanation plots
results/pipeline.log Full execution log

βš™οΈ Configuration

Edit config_lite.py to customize:

# Target protein
TARGET_PDB_ID = "1isp"
TARGET_CHAIN_ID = "A"
PROTECTED_RESIDUES = [57, 102, 195]  # Active site - never mutate

# Scale up for Full Mode
MAX_PDB_DOWNLOADS = 50      # Lite: 20
NSGA2_GENERATIONS = 50      # Lite: 20
XGBOOST_ESTIMATORS = 500    # Lite: 200

πŸ“œ License

MIT License - See LICENSE

🀝 Contributing

See CONTRIBUTING.md for guidelines.

πŸ”’ Security

See SECURITY.md for security policy.

πŸ“š References

About

Production-grade ML pipeline for predicting protein thermostability mutations using XGBoost, NSGA-II optimization, RAG, and TRIZ-based reasoning. Combines biophysical feature engineering with multi-objective evolutionary optimization.

Topics

Resources

Contributing

Security policy

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages