I'm an AI Engineer at Baylor College of Medicine working at the intersection of machine learning, LLMs, and computational biology.
My interests include:
- 🤖 LLM post-training and reasoning
- 🧠 AI for biology and healthcare
- 🧬 Single-cell genomics & computational biology
- ⚙️ Building production ML systems
- 🔬 ML interpretability and representation learning
Research into detecting and intervening on failing reasoning trajectories inside large language models using hidden-state probes and reinforcement-learning inspired methods.
Topics:
- hidden-state interpretability
- reasoning models
- RLVR / GRPO
- evaluation methodology
- efficient inference
A transformer foundation model for longitudinal electronic health records that predicts future ICD diagnoses from patient histories.
Highlights:
- transformer pretraining
- large-scale clinical tokenization
- evaluation framework
- healthcare AI
An agentic retrieval system for matching trainees with research mentors using semantic search, reranking, and LLM-assisted workflows.
Stack:
- FastAPI
- PostgreSQL
- LangChain
- Docker
- AWS
I'm particularly interested in problems involving
- reasoning models
- post-training
- representation learning
- AI agents
- computational biology
- foundation models
- biomedical machine learning
- healthcare AI
Languages
Python • R • SQL • Bash
ML
PyTorch • Hugging Face • TensorFlow • CUDA
LLMs
SFT • RLVR • GRPO • RAG • Cross-Encoder Reranking • LangChain
Infrastructure
Docker • AWS • PostgreSQL • FastAPI • Slurm • Git
Computational Biology
single-cell RNA-seq • scATAC-seq • Bioconductor • Scanpy • Seurat
- Cell Patterns (2023)
- Genes (2021)
- Frontiers in Neuroscience (Accepted)
💼 LinkedIn: https://www.linkedin.com/in/johnathan-jia-a01315170/



