Skip to content
View mekala27-45's full-sized avatar

Block or report mekala27-45

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
mekala27-45/README.md

Ajay Mekala

I'm an AI/ML engineer at Walmart, where I work on the machine learning platform behind promotion and clearance pricing. Since 2022 I've also done contract model evaluation for Handshake AI, Snorkel AI, Mercor and Outlier: writing benchmark tasks, scoring coding-agent runs and building rubrics.

MS in Data Science, Montclair State University (2026). I live in New Jersey and I'm authorized to work in the U.S. without sponsorship.

Portfolio · LinkedIn · Email

At work

The pricing platform runs on Azure Databricks, Delta Lake and MLflow, with the infrastructure in Terraform, and serves more than 200,000 predictions a month. Two systems run on it. PromotionsAI picks the SKUs that go on promotion and has driven $7.8M in incremental revenue. ClearanceAI sets markdown prices from demand elasticity models plus stock-age rules, and it lifted sell-through by 4.6%.

A lot of my time goes into the parts around the models: Dockerized inference, MLflow lineage, blue-green releases through Azure Pipelines, an A/B testing harness, and connector libraries the rest of the team builds on. Research-to-production lead time is down 40%, experiment turnaround is down 60%, and incident MTTR is about half what it was.

The evaluation work adds up to 200+ golden-solution tasks for frontier coding benchmarks, 50+ accepted Terminal-Bench environments, and 3,000+ side-by-side preference comparisons at 98%+ agreement with senior reviewers. It's under NDA, so my public projects are where I use the same ideas on problems I can share.

Projects

Personal projects on public or synthetic data. None of them use employer code or data.

  • trajectory: scores coding agents on every step of a run instead of only on whether the hidden tests pass at the end. Leaderboard
  • pricepoint: price elasticity and markdown decisions on the UCI Online Retail II data, with a temporal backtest, promotion gates, shadow serving and drift monitoring. The trailing-mean baseline beat the LightGBM candidate, so the baseline is what it serves. Demo
  • cityflow: a dbt warehouse for NYC taxi and for-hire trip records, with a dashboard that queries parquet in the browser through DuckDB-WASM. The published figures come from a seeded generator that copies the TLC schema and its data problems. Dashboard
  • frontdesk: a WhatsApp booking agent for a fictional clinic. Most of the work went into concurrent booking, webhook retries, and refusals that are recorded as tool calls. Evidence explorer
  • readout: an A/B testing platform with frozen designs, health checks that run before any metric, and a sequential test checked against simulated null experiments. Demo
  • groundwork: upload a PDF and ask questions about it. Answers cite their sources, the citations are checked, and there are tests for prompt injection and workspace isolation. Demo
  • northstar: analytics on the public Olist e-commerce data, with a dbt warehouse, forecasting, a SQL assistant scored against a test set, and a written memo. Dashboard
  • ledger: the analytics work of a fictional retail bank. A fraud desk that sets its threshold by cost and waits for late labels, a credit scorecard with adverse action reasons, fair lending on real New Jersey HMDA data, money laundering rules with an alert triage model, and a validation report for every model. Site
  • shortlist: the pipeline I use for my own job search. It pulls postings from 16 sources, scores them and pre-fills applications, but it never presses submit.

Older and smaller: intent-sentinel (purchase intent model with drift monitoring), edge-vision (INT8 MobileNetV2 running in the browser), grounded-rag (hybrid retrieval evaluated on SciFact) and nanogpt-lab (a small Llama-style transformer with ablations).

Tools

  • Languages: Python, SQL, TypeScript, Go, Bash
  • ML: PyTorch, TensorFlow, scikit-learn, XGBoost, LightGBM, Hugging Face Transformers
  • MLOps: MLflow, Weights & Biases, Databricks, Docker, Kubernetes, Terraform, FastAPI
  • Data: Spark and PySpark, Delta Lake, Airflow, Kafka, dbt, DuckDB, Postgres
  • Cloud: Azure, AWS, GCP

I'm looking for ML engineering, MLOps and model evaluation roles, remote or around New York. Email is the fastest way to reach me.

Pinned Loading

  1. trajectory trajectory Public

    Evaluation harness for coding agents: containerized tasks, hidden tests and step-level trajectory scoring

    Python

  2. pricepoint pricepoint Public

    Weekly demand forecasting and constrained markdown pricing on UCI Online Retail II, with a FastAPI service

    Python

  3. cityflow cityflow Public

    DuckDB and dbt warehouse and in-browser dashboard for NYC taxi trip records. Published figures are synthetic.

    Python

  4. frontdesk frontdesk Public

    WhatsApp booking agent for a fictional clinic: Postgres slot locking, webhook retries, auditable refusals

    Python

  5. readout readout Public

    A/B testing platform with frozen designs, deterministic assignment, health checks and rendered decision docs

    Python

  6. ledger ledger Public

    The analytics function of a fictional retail bank: a live fraud desk with late labels, a credit scorecard studio, SR 11-7 validation reports for every model, and a general ledger that reconciles to…

    Python