Skip to content
View Amidn's full-sized avatar

Block or report Amidn

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Amidn/README.md

Hi, I'm Amid Nayerhoda

Data Scientist | Machine Learning at Terabyte Scale | Python, TensorFlow & Spark

I am a Data Scientist and astrophysicist with 9+ years of experience turning complex, high-volume data into reliable analytical results. My work combines machine learning, statistical modeling, and reproducible data engineering to solve rare-event and high-noise detection problems.

My research has contributed to 130+ publications, including work published in Nature and Science. I bring that same rigor to data products: clear problem framing, robust pipelines, careful model validation, and results that can be trusted.

What I work with

  • Machine learning: predictive modeling, classification, feature engineering, model evaluation, deep learning
  • Statistics: likelihood-based inference, hypothesis testing, uncertainty quantification, quantitative analysis
  • Data engineering: ETL pipelines, data orchestration, distributed processing, reproducible workflows
  • Tools: Python, pandas, NumPy, scikit-learn, TensorFlow, PyTorch, Spark/PySpark, SQL, Docker, GitHub Actions, Jupyter

Featured work

Large-scale rare-event detection

Built and optimized end-to-end data pipelines for multi-terabyte scientific datasets. Applied machine-learning classifiers, statistical inference, and simulation-based validation to identify rare signals in highly imbalanced, high-noise data.

Focus: Python, Spark/HPC, ETL, feature engineering, classification, model evaluation, reproducibility.

Reproducible data science workflows

Developed containerized analytical environments and automated workflows for data processing, model validation, visualisation, and team collaboration.

Focus: Docker, Conda, CI/CD, Jupyter, Matplotlib, version control, documentation.

Structured AI interaction tools

Created an open-source collection of prompt decorators that makes AI interactions more structured, transparent, and reusable for researchers, students, and practitioners.

Focus: AI tooling, prompt engineering, YAML, documentation, open source.

Current focus

I am interested in Data Scientist opportunities where machine learning, statistical reasoning, and large-scale data can create measurable impact.

Connect


I value rigorous analysis, reproducible work, and clear communication of results.

Pinned Loading

  1. JTBrix JTBrix Public

    HTML

  2. ChatGPT-decorators ChatGPT-decorators Public

    A collection of customizable ChatGPT prompt decorators that adapt responses to specific needs. They enable dynamic control over style, structure, and content, making conversations more flexible and…

  3. rag-document-qa rag-document-qa Public

    Python RAG pipeline to chat with your documents using LLMs — built with LangChain, ChromaDB, Groq, and HuggingFace embeddings

    Python

  4. SignalSeeker SignalSeeker Public

    SignalSeeker: production quality ML pipeline for imbalanced classification

    Python

  5. SynerG SynerG Public

    Python