Data Scientist | Machine Learning at Terabyte Scale | Python, TensorFlow & Spark
I am a Data Scientist and astrophysicist with 9+ years of experience turning complex, high-volume data into reliable analytical results. My work combines machine learning, statistical modeling, and reproducible data engineering to solve rare-event and high-noise detection problems.
My research has contributed to 130+ publications, including work published in Nature and Science. I bring that same rigor to data products: clear problem framing, robust pipelines, careful model validation, and results that can be trusted.
- Machine learning: predictive modeling, classification, feature engineering, model evaluation, deep learning
- Statistics: likelihood-based inference, hypothesis testing, uncertainty quantification, quantitative analysis
- Data engineering: ETL pipelines, data orchestration, distributed processing, reproducible workflows
- Tools: Python, pandas, NumPy, scikit-learn, TensorFlow, PyTorch, Spark/PySpark, SQL, Docker, GitHub Actions, Jupyter
Built and optimized end-to-end data pipelines for multi-terabyte scientific datasets. Applied machine-learning classifiers, statistical inference, and simulation-based validation to identify rare signals in highly imbalanced, high-noise data.
Focus: Python, Spark/HPC, ETL, feature engineering, classification, model evaluation, reproducibility.
Developed containerized analytical environments and automated workflows for data processing, model validation, visualisation, and team collaboration.
Focus: Docker, Conda, CI/CD, Jupyter, Matplotlib, version control, documentation.
Created an open-source collection of prompt decorators that makes AI interactions more structured, transparent, and reusable for researchers, students, and practitioners.
Focus: AI tooling, prompt engineering, YAML, documentation, open source.
I am interested in Data Scientist opportunities where machine learning, statistical reasoning, and large-scale data can create measurable impact.
- LinkedIn: linkedin.com/in/nayerhoda
- GitHub: github.com/Amidn
I value rigorous analysis, reproducible work, and clear communication of results.