βThe soul is neither born, and nor does it die.β β Bhagavad Gita source: zenquotes
I am interested in Mechanism-guided Trustworthy AI Systems β building AI systems that can be not only powerful, but also understood, calibrated, monitored, corrected, and controlled when they operate in changing real-world environments.
My current work starts from a practical question:
How can an AI system know when it is becoming unreliable β and what should happen next?
I study this question through the connection between model internals, uncertainty, causal mechanisms, dynamic environments, and system-level intervention.
My near-term focus is to build reliability pipelines for AI models under distribution shift, noisy observations, incomplete data, and fragile deployment workflows. My long-term goal is broader: to understand how AI systems can move from black-box prediction toward mechanism-grounded, corrigible, and controllable intelligence.
From black-box models to reliable, interpretable, and controllable intelligence.
Not a linear ladder, but a spiral: each turn deepens understanding and strengthens control.
βNature operates through the principle of least action β intelligence, perhaps, should too.β
- π Incoming / M.S. in Data Science @ UC San Diego, HalΔ±cΔ±oΔlu Data Science Institute (HDSI)
- π Background: Management Information Systems β Machine Learning / AI Systems / Trustworthy AI
- π§ Research-oriented transition:
- Model User β Model Builder β System Thinker β Reliability-oriented Researcher
- π Currently building foundations in:
- Machine Learning, Deep Learning, and NLP
- Spatiotemporal Graph Learning and Complex Dynamic Systems
- Causal Discovery, Uncertainty Quantification, and Trustworthy Evaluation
- AI Reliability, Model Monitoring, and ML Systems
- Signal Processing, Dynamical Systems, and Mechanism Discovery
I currently think about trustworthy AI through four connected research paths.
| Path | Core Idea | My View |
|---|---|---|
| Agentic AI | Give language models a body through tools, memory, planning, and interaction. | This is an important application frontier, but not my only research identity. I am more interested in how agentic systems fail, how tool use should be monitored, and when human review or intervention should be triggered. |
| Model Mechanisms | Open and repair the model body: representations, circuits, concepts, world models, and causal structure. | This is the depth layer of my long-term interest. I want to understand whether internal representations can support correction, steering, and reliable behavior rather than remaining opaque high-dimensional artifacts. |
| External Control Harness | Keep AI systems inside evaluation, monitoring, guardrail, audit, fallback, and human-in-the-loop control loops. | This is my near-term anchor. Before fully understanding every internal mechanism, we can still build systems that detect risk, expose failure signals, abstain, defer, or roll back. |
| Human and Scientific Understanding | Study how human concepts, machine representations, and world mechanisms can align. | This is the philosophical and scientific motivation behind my interest in interpretability. I care about when an explanation is not just plausible, but connected to stable, human-understandable, and scientifically meaningful structure. |
My current strategy is:
near-term anchor: external reliability harness and monitoring
research depth: model mechanisms, concepts, and causal structure
application frontier: agentic AI and scientific AI systems
long-term foundation: human concepts, machine representations, and world mechanisms
In this view, interpretability is not only about explaining a model after it makes a prediction. It can also become part of a broader control interface:
model behavior
β uncertainty and calibration
β mechanism / concept signals
β failure detection
β human review, fallback, correction, or intervention
This is the direction I am gradually building toward: AI systems whose reliability is grounded in both internal mechanisms and external control loops.
I focus on mechanism-guided trustworthy AI systems: models and evaluation pipelines that remain understandable, calibratable, monitorable, and actionable under real-world uncertainty and distribution shift.
My long-term research question is:
How can mechanistic understanding make AI systems more reliable, corrigible, and controllable in complex dynamic environments?
I am especially interested in AI systems where prediction alone is not enough. In domains such as traffic, environment, energy, scientific modeling, healthcare, and AI agents, models must expose uncertainty, reveal failure signals, support intervention, and remain reliable when the environment changes.
- Spatiotemporal graph learning for complex dynamic systems
- Reliable forecasting under missingness, noise, regime changes, and long-horizon uncertainty
- Moving beyond accuracy-only evaluation toward calibration, robustness, and decision usefulness
- Understanding whether models learn stable dynamic mechanisms or dataset-specific correlations
- Causal discovery and invariant prediction under distribution shift
- Separating stable mechanisms from spurious correlations
- Applying causal and mechanistic thinking to environmental, traffic, energy, and scientific prediction
- Studying whether learned graphs, latent states, and representations correspond to meaningful system structure
- Uncertainty quantification, calibration, and predictive intervals
- Conformal prediction under temporal shift, change points, and nonstationarity
- Turning model confidence into decision-relevant reliability signals
- Asking when a model should predict, abstain, defer to human review, or trigger fallback
- Representation analysis, CKA, probing, and concept stability
- From post-hoc explanations to mechanism-level understanding
- Evaluating whether explanations and representations survive sanity checks
- Exploring concept bottlenecks, sparse features, and mechanistic signals as possible interfaces for model correction and control
- Model monitoring under data drift, prediction drift, calibration drift, and representation drift
- RAG / LLM evaluation pipelines where reliability and failure diagnosis matter
- Human-in-the-loop review, audit trails, fallback mechanisms, and intervention triggers
- Toward AI systems that can be audited, corrected, and controlled after deployment
I currently organize trustworthy AI around three connected layers:
| Layer | Core Question | Methods / Signals |
|---|---|---|
| Mechanism Layer | What internal concepts, representations, causal structures, or dynamic mechanisms does the model rely on? | Mechanistic interpretability, concept probing, causal discovery, representation analysis |
| Reliability Layer | When is the model uncertain, miscalibrated, out-of-distribution, or about to fail? | Calibration, conformal prediction, uncertainty quantification, drift detection |
| Control Layer | How should the system respond when risk is detected? | Abstention, human review, fallback, rollback, monitoring, audit trails |
A practical research loop I want to build is:
Model prediction
β uncertainty and calibration check
β concept / mechanism reliability check
β drift and failure monitoring
β risk-aware intervention
β human review, fallback, or correction
This is why I am particularly interested in dynamic distribution shift: it forces a model to reveal whether it has learned stable mechanisms or only fragile correlations.
I am interested in interpretability not only as a tool for explaining model outputs, but as a way to study the relationship between:
world mechanisms
β human concepts
β machine representations
β model behavior
β system-level decisions
Human concepts and machine representations are both compressed descriptions of the world. A central challenge is to understand when high-dimensional learned representations can be aligned with stable, human-understandable, and scientifically meaningful mechanisms.
This motivates my interest in:
- mechanism discovery
- causal representation learning
- concept-based interpretability
- scientific machine learning
- trustworthy evaluation
- AI systems that can be corrected and controlled after deployment
I try to keep this broader question grounded in measurable reliability, concrete systems, and reproducible experiments.
Flagship research lab for:
- reliable spatiotemporal forecasting under dynamic distribution shift
- graph construction validation
- uncertainty quantification and conformal calibration
- risk-aware decision evaluation
- mechanism-guided reliability experiments
- monitoring triggers for abstention, human review, and fallback
Research-level reading system for:
- spatiotemporal graph learning
- uncertainty and calibration
- conformal prediction
- mechanism discovery and causal time-series analysis
- interpretable representation and trustworthy evaluation
- AI systems, monitoring, and technical debt in ML systems
Exploratory lab for:
- concept representations
- CKA and representation similarity
- probing and hidden-state diagnostics
- sanity checks for explanations and representations
- representation drift under noise, missingness, and distribution shift
Foundation implementations from scratch:
- classical ML, MLPs, optimization, and backpropagation
- deep learning and attention mechanisms
- tiny transformer and NLP foundations
- mathematical intuition and implementation discipline
- baseline models for reliability and shift experiments
Earlier RAG / LLM evaluation work remains part of my broader AI reliability interest, especially when evaluation, monitoring, and failure diagnosis are central.
| Domain | Focus |
|---|---|
| π§ AI | Mechanism-guided trustworthy AI, reliability, interpretability, causality |
| π¬ Science | Complex dynamic systems, uncertainty-aware prediction, physics-inspired modeling |
| π Systems | Evaluation pipelines, monitoring, human-in-the-loop control, ML systems for reliable deployment |
| π‘ Signals | Noise, sampling, filtering, graph signals, state estimation, dynamic observations |
| π Humanities | Epistemology, pragmatism, causality, interpretability, philosophy of science |
| Project | Direction | Core Question |
|---|---|---|
| Reliable Spatiotemporal Forecasting under Dynamic Shift | STGNN / UQ / Conformal Prediction | How can forecasts remain calibrated and decision-useful when graph structure, sensors, and environments shift? |
| Mechanism Discovery for Complex Dynamic Systems | Causal Discovery / Dynamic Systems | Can models recover stable mechanisms rather than exploiting unstable correlations? |
| Graph Construction and Reliability Validation | Spatiotemporal Graph Learning | When is a learned or designed graph a valid representation of physical, statistical, or causal influence? |
| Concept and Representation Stability | Interpretability / Representation Analysis | Do model representations encode meaningful concepts, and do they remain stable under shift? |
| AI Evaluation and Monitoring Pipelines | AI Reliability / Model Monitoring | How can deployed AI systems be audited, calibrated, monitored, corrected, and controlled over time? |
| Mechanism-guided Reliability Harness | Trustworthy AI Systems | Can uncertainty signals, mechanism checks, and monitoring triggers form a practical control loop for AI failure detection and intervention? |
-
π‘ Topics:
- Mechanism-guided Trustworthy AI
- Spatiotemporal Graph Learning
- Uncertainty and Calibration
- Causal Discovery
- AI Evaluation and Monitoring
- Interpretability and Representation Analysis
- Human-in-the-loop AI Reliability
-
π« Open to:
- Research collaborations
- ML / AI internships
- Applied AI reliability and evaluation projects
- Mechanism-guided AI systems research
