Skip to content
View ChinmayA301's full-sized avatar

Block or report ChinmayA301

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
ChinmayA301/README.md

Hi, I'm Chinmay Arora

Data Scientist · Applied AI Engineer · Enterprise AI Solutions

I turn ambiguous operational problems into practical data and AI systems—connecting stakeholder discovery, solution architecture, implementation, evaluation, and adoption.

One candid note: this GitHub is deliberately AI-polished. I use AI tools to help structure documentation, accelerate selected scaffolding, and pressure-test how the work is presented. The builds, datasets, experiments, ambitions, decisions, and results are real. I review, test, and own what I publish, and I state maturity and limitations where they matter.

Portfolio · Data Science · Analytics · Applied AI · Responsible AI

Current focus

  • Building citation-grounded knowledge and document-intelligence systems.
  • Designing evaluation gates, human review, audit evidence, and monitoring for high-trust AI workflows.
  • Applying data science, analytics engineering, and solution architecture to public-sector, healthcare, operations, finance, and real-estate problems.
  • Targeting forward-deployed, applied AI, enterprise solutions, and AI transformation roles.

Featured work

Status: Completed analysis. Evaluated 101,766 real de-identified encounters with leakage control, calibration, subgroup errors, SHAP, and a model card. A measured 0.44 age-group false-negative-rate gap supports the documented recommendation not to deploy as-is.

Status: Completed analysis. Designed a config-driven benchmark of five pipelines across three real domains and random/shifted splits. Preprocessing added up to 0.19 AUC while tuning and AutoML added almost no mean lift under the tested conditions.

Status: Completed analysis. Built a DuckDB star schema, data-quality gates, supplier-risk logic, and five-page Looker Studio decision flow over 98,666 real orders. Late orders were 6.5× more likely to receive a 1–2-star review.

Status: Deployed demo. Built a citation-grounded RAG workflow that gives multiple models the same retrieved evidence, separating retrieval quality from generation behavior. Includes ingestion, OCR fallback, FAISS retrieval, evaluation, FastAPI, Docker, and a live demo.

Capabilities, with evidence

  • Data and analytics: Python, SQL, statistical analysis, KPI design, dimensional modeling, BI, and geospatial analysis.
  • Machine learning: calibration, feature engineering, subgroup analysis, SHAP, drift evaluation, and temporal validation.
  • Applied AI: RAG, embeddings, vector retrieval, document ingestion, OCR, FastAPI, Docker, grounding, and citation evaluation.
  • Responsible AI: fairness evaluation, model cards, human review, auditability, risk registers, and deployment gates.
  • Solution delivery: stakeholder discovery, problem framing, architecture, prototyping, implementation planning, and technical communication.

How I work

Discover → Architect → Implement → Evaluate → Adopt

I start with the user, workflow, value, data, and constraints. I then design and build the core system, test quality and failure modes, and translate the result into a responsible deployment or do-not-deploy decision.

Portfolio intelligence

My 'What Excites You To Wake Up'

My longer-term trajectory is technical product building and venture exploration grounded in data science and AI. Check out Product and Venture Lab if interested. It is organized around one thesis: AI becomes genuinely useful when context, authority, evidence, and human decision-making are connected. Aegis is the current independent product direction; StepLens is the longer-term interface opportunity. The remaining explorations are preserved as an honestly staged archive, not presented as simultaneous companies.

Some blabberring pieces over at Blog

Contact

Portfolio · LinkedIn · Email · Resume

Pinned Loading

  1. Operations-Control-Tower Operations-Control-Tower Public

    Operations analytics control tower: DuckDB star schema, supplier risk scoring, Great Expectations checks, and Looker Studio dashboard on real Olist data.

    Jupyter Notebook

  2. Risk-Prediction-with-Fairness-Geography Risk-Prediction-with-Fairness-Geography Public

    Responsible ML risk modeling with leakage control, calibration, subgroup fairness, SHAP, and a model-card style deployment recommendation.

    Python

  3. Signal-Graph Signal-Graph Public

    Open-source diligence tool for credibility-adjusted traction, adoption durability, builder quality, and manipulation-risk signals.

    Python 1

  4. wax-seal wax-seal Public

    Trust-provenance framework for digital artifacts: cryptographic seal envelopes, lifecycle state machines, CLI verification, and React components.

    TypeScript 1

  5. Decision-Intelligence Decision-Intelligence Public

    AI decision-support app that retrieves historical decision patterns and produces structured strategy briefs with lenses and pre-mortems.

    Python

  6. skm-football skm-football Public

    Open-source football analytics pipeline for process-based player valuation using StatsBomb data, VAEP-style scoring, validation, and Streamlit.

    Python