| Project | Focus | Technical Highlights |
|---|---|---|
| AgentMage | Local-first AI agent harness Pre-alpha / In development |
Developing a Rust-based, local-first AI agent harness with explicit authorization contracts and bounded workspace tools. Current work includes isolated read-only worker components and adversarial validation tests; integrated model workflows and cross-platform releases remain in development. |
| CodingMage | Rust-based multi-agent coding coordinator In development |
Developing a coordinator for supervised Claude Code implementation and independent Codex review. Includes isolated Git worktrees, deterministic checks, commit-bound review records, durable checkpoints, and an initial serial-campaign workflow. Parallel execution and broader interruption recovery remain in development. |
| Labels On Tap | ML-assisted alcohol-label verification Treasury interview project |
Built and publicly deployed a local-first alcohol-label verification prototype for a U.S. Department of the Treasury interview challenge focused on the Alcohol and Tobacco Tax and Trade Bureau (TTB). Integrated local docTR OCR, optional DistilRoBERTa evidence scoring, deterministic validation rules, and visual checks of government-warning typography. Implemented a FastAPI reviewer interface, automated tests, filesystem-backed batch processing with restart recovery, and Docker Compose deployment on AWS Lightsail. |
| Agent Gatsby | Local-first GenAI pipeline | Built for the U.S. Treasury's supplemental Great Gatsby job application requirement. Uses a deterministic state machine and open-weight models to coordinate evidence extraction, source-backed quotation and citation validation, bounded expansion passes, and configurable word-count checks, followed by Spanish and Mandarin translations. |
| AdaptiveLASSO | Statistical computing Published on PyPI |
Developed and published a Python package implementing adaptive LASSO regression through a Statsmodels-style API. Includes automated tests for variable selection, penalty handling, and API behavior; documentation and a demonstration notebook; and a GitHub Actions workflow for testing, building, and publishing releases. |
| LexiChess | LLM benchmarking and interactive chess Backend MVP; web interface in development |
A chess tournament and evaluation platform for benchmarking LLMs, chess agents, and custom engines: reproducible match pipeline with deterministic legal-move validation, SQLite-backed game storage, tournament scheduling, Elo-style ratings, PGN exports, Stockfish analysis, and pluggable providers for local models. Includes a FastAPI live-broadcast interface with showmatch controls, moderation workflows, referee commentary, and human-vs-agent play. |
| TopicMiner | Classical NLP email classification | Developed a modular Python toolkit for email preprocessing, topic discovery, and classification. Combines LDA and embedding-based clustering with Random Forest and XGBoost training, tuning, and evaluation utilities. Designed for narrow-taxonomy classification workflows using classical NLP and machine learning. |
| iSignify | Bioinformatics prototype Developed for Google's Gemma 3n Impact Challenge |
Built a Python/FastAPI pipeline that compares target and background genomes to identify candidate DNA signature regions without a precomputed database. Includes FASTA preprocessing, k-mer comparison, and locally generated Gemma summaries of analysis counts and parameters. |
| USTE | Universal Spatial-Temporal Engine - Deterministic simulation architecture Design stage |
Designed the architecture and requirements for a planned Rust simulation kernel using procedural generation, event-sourced deviations, and reproducible replay. Defined numerical-reproducibility, persistence, and multiscale simulation contracts; implementation has not started. |
Local-first AI agent harness · Pre-alpha / In development
A Rust-based, local-first AI agent harness built around explicit authorization contracts and bounded workspace tools. The design question it exists to answer is what an agent should be allowed to do, and how that permission is proven rather than assumed. Current work includes isolated read-only worker components and adversarial validation tests. Integrated model workflows and cross-platform releases remain in development. Source available under the Business Source License 1.1.
Key themes: local-first AI, open-weight models, explicit authorization contracts, bounded tools, Rust, adversarial validation.
Rust-based multi-agent coding coordinator · In development
A coordinator for supervised Claude Code implementation and independent Codex review, so that the agent writing the code is never the agent approving it. Includes isolated Git worktrees, deterministic checks, commit-bound review records, durable checkpoints, and an initial serial-campaign workflow. Parallel execution and broader interruption recovery remain in development.
Key themes: multi-agent orchestration, independent review, isolated worktrees, deterministic gates, durable checkpoints, Rust.
ML-assisted alcohol-label verification · Treasury interview project
A local-first alcohol-label verification prototype, built and publicly deployed for a U.S. Department of the Treasury interview challenge focused on the Alcohol and Tobacco Tax and Trade Bureau. It integrates local docTR OCR, optional DistilRoBERTa evidence scoring, deterministic validation rules, and visual checks of government-warning typography. Delivery included a FastAPI reviewer interface, automated tests, filesystem-backed batch processing with restart recovery, and Docker Compose deployment on AWS Lightsail.
Key themes: regulatory triage, OCR and NLP, deterministic validation, FastAPI, Docker, human-in-the-loop review, cloud deployment.
Local-first GenAI pipeline
Built for the U.S. Treasury's supplemental Great Gatsby job application requirement. A deterministic state machine drives open-weight models through evidence extraction, source-backed quotation and citation validation, bounded expansion passes, and configurable word-count checks, followed by Spanish and Mandarin translations. The point is that the control flow is deterministic and the model fills bounded slots inside it.
Key themes: GenAI orchestration, deterministic state machines, citation validation, bounded generation, multilingual output.
Statistical computing · Published on PyPI
A Python package implementing adaptive LASSO regression through a Statsmodels-style API, developed and published to PyPI. Includes automated tests for variable selection, penalty handling, and API behavior; documentation and a demonstration notebook; and a GitHub Actions workflow for testing, building, and publishing releases.
Key themes: regression, shrinkage, statistical computing, Python packaging, CI/CD, reproducible modeling.
LLM benchmarking and interactive chess · Backend MVP; web interface in development
A chess tournament and evaluation platform for benchmarking LLMs, chess agents, and custom engines. The match pipeline is reproducible, with deterministic legal-move validation, SQLite-backed game storage, tournament scheduling, Elo-style ratings, PGN exports, Stockfish analysis, and pluggable providers for local models. A FastAPI live-broadcast interface adds showmatch controls, moderation workflows, referee commentary, and human-versus-agent play.
Key themes: LLM evaluation, agent benchmarking, reproducible pipelines, SQLite, FastAPI, live broadcast tooling.
Classical NLP email classification
A modular Python toolkit for email preprocessing, topic discovery, and classification. It combines LDA and embedding-based clustering with Random Forest and XGBoost training, tuning, and evaluation utilities. Designed for narrow-taxonomy classification workflows where fine-tuned classical methods remain competitive with LLM zero-shot at a fraction of the inference cost and latency.
Key themes: LDA, Word2Vec, Doc2Vec, K-Means, Random Forest, XGBoost, auditable classification.
Bioinformatics prototype · Developed for Google's Gemma 3n Impact Challenge
A Python and FastAPI pipeline that compares target and background genomes to identify candidate DNA signature regions without a precomputed database. Includes FASTA preprocessing, a custom k-mer comparison, and locally generated Gemma summaries of analysis counts and parameters, so interpretation stays on the device.
Key themes: bioinformatics, k-mer algorithms, on-device inference, data privacy, FastAPI.
Deterministic simulation architecture · Design stage
Architecture and requirements for a planned Rust simulation kernel using procedural generation, event-sourced deviations, and reproducible replay: compute only the present moment, store only what deviates from the rules, and replay any moment exactly. Numerical-reproducibility, persistence, and multiscale simulation contracts are defined. Implementation has not started.
Key themes: deterministic simulation, event sourcing, procedural generation, replay verification, multiscale coordinates, Rust.
AI engineering for Treasury business units.
October 2025 to June 2026 · Houston, TX
Built and deployed production Generative AI systems, standardizing Python LLM and RAG pipelines for internal clients on LangGraph and MCP architectures, and validated PwC's internal LangGraph wrapper product. Engineered an ETL process for a large multinational telecommunications client to parse and validate multi-sheet, human-entered service-order and renewal workbooks, enforcing business logic and strict verification rules with Pydantic. Analyzed client Swagger and OpenAPI documentation to build automated, read-only API connections from a SOX-compliant Snowflake database into the ingestion pipeline. Designed modular, object-oriented pipeline components for a hybrid-infrastructure ingestion system, wrote unit tests, took part in peer code review under enterprise CI/CD standards, and resolved defects using bug reports and Datadog traces.
Data Scientist (Statistician), GS-14 — Internal Revenue Service, Research, Applied Analytics & Statistics
November 2021 to September 2025 · Houston, TX
Led design and build of an LLM email-classification pipeline: before access to the email data was granted, built a local LLM agent to generate a synthetic training set, then combined API calls to internally hosted LLMs with classical ML classifiers, evaluated using stratified train, validation, and test sampling. Co-led, with SB/SE Revenue Agents, an application automating financial data preparation for audit analysis; pitched and secured unanimous sponsorship from RAAS Directors and SB/SE executives, briefed them throughout, and engineered the Python and SQL ETL pipelines behind it. Supported an internal RAG chatbot giving IRS employees access to over one million pages of the Internal Revenue Manual and the U.S. Tax Code. Conducted an independent technical review of a proposed Gaussian Mixture Model for return selection and delivered findings to leadership. Developed SARIMAX and LOWESS forecasts with prediction intervals for taxpayer interaction volume, including a novel in-time parameter-tuning approach. Assisted in evaluating IBM Cloud Pak for Data, represented the IRS at industry conferences, and created internal Jupyter lessons on machine learning, statistical modeling, and CRISP-DM/SEMMA.
August 2019 to November 2021 · Houston, TX
Pair-programmed a workforce management and planning model showing when employees would mobilize and demobilize across project sites, yards, fabrication facilities, and offices worldwide, and built the automated ETL that scraped monthly operational submissions and loaded them into the data warehouse. Produced the company's weekly COVID-19 forecast and long-term scenarios for twelve countries to monitor impacts on multi-billion-dollar active projects, automating extraction from the Johns Hopkins repository, engineering predictions with a custom auto-ARIMA and statistical-learning methods as alternatives to early IHME models, and delivering output directly to Word and PowerPoint for executives. Built the team's Python regression toolset with Statsmodels and scikit-learn. Taught a weekly graduate-level applied statistics and Python course modeled on MITx 6.00.1x; every participating team member earned the certificate.
November 2018 to July 2019 · Houston, TX
Coded, trained, validated, and tested convolutional neural networks for an image-classification model identifying recycled materials. Built a statistical model predicting waste-compactor capacity from discrete, high-noise pressure readings taken during ram cycles, fitting a smooth trend that adjusted in real time as new sensor data arrived, and deployed a model estimating container fullness by normalizing continuous pressure readings pulled via API into percentage-full metrics for client optimization. Evaluated relative machine performance with ANOVA and Tukey range tests, and built SQL queries over JD Edwards warehouse data plus a time-series model for inventory forecasting.
Founder, The Confidential Advisor LLC, a recruiting agency in Houston, 2010 to 2018. Executive Recruiter and Business Development, Qualitec Technical Services, Houston, 2005 to 2008.
United States Army, 2000 to 2004. Airborne Ranger, 2nd Ranger Battalion; Infantry, 1-23 Stryker Brigade. Honorable discharge, Good Conduct Medal.
MS, Analytics / Data Science — Texas A&M University, Department of Statistics, 2018 MS, Statistical Data Science — Texas A&M University, Department of Statistics, in progress BS, Economics — University of Houston, 2010 BS, Political Science — Texas A&M University, 2000
LLMs · RAG · Agentic Workflows · LangGraph / MCP · Local-First Inference with Open-Weight Models · Synthetic Data Generation · Hallucination Controls · Human-in-the-Loop Review
Scikit-Learn · PyTorch · XGBoost · Random Forests · CNNs · NLP (LDA, Word2Vec, Doc2Vec) · Gaussian Mixture Models
SARIMAX · auto-ARIMA · LOWESS · Survival Analysis · Regression and Shrinkage (adaptive LASSO) · ANOVA / Tukey · Experimental Design · Stratified Sampling
Python ETL · Pydantic Validation · SQL · Snowflake · Swagger / OpenAPI Integration · API Ingestion · FastAPI · Docker · SQLite · AWS
Python · Rust · SQL · Linux · Unit Testing · Peer Code Review · Git-Based CI/CD · GitHub / GitLab · Datadog · Agile Delivery · VS Code · JupyterLab · Codex · Claude Code · Gemini Code Assist · R, SAS, JMP (academic)
I am especially interested in AI and data systems that are:
- Useful in real operational settings
- Reviewable and explainable
- Reliable under messy real-world data
- Statistically defensible
- Testable and maintainable
- Production-ready rather than demo-only
My preferred approach is pragmatic: use GenAI where it adds value, use classical ML where it is stronger, and use deterministic validation wherever reliability matters.
Please use the contact and social links in my GitHub profile sidebar.
GitHub: github.com/AaronNHorvitz
LinkedIn: linkedin.com/in/aaron-horvitz-a666215
Resume: Download Resume
