Skip to content
View Amaankaa's full-sized avatar

Highlights

  • Pro

Block or report Amaankaa

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Amaankaa/README.md

Amanuel Merara Gutu

Software Engineer · LLM Evaluation & Alignment · Backend / RAG Systems

LinkedIn Email


AI-focused software engineer building production systems with hands-on frontier-LLM evaluation experience. I've been contracted to Anthropic (via Revelo) to evaluate Claude Code's production PRs — generating the preference data behind RLHF and reward-model training — and I now author expert coding trajectories at AfterQuery that teach frontier models to reason.

I like working where alignment research meets real engineering: not just labeling data, but building the RAG platforms and evaluation harnesses that turn "is this answer good?" into a reproducible, regression-tested metric.

🧠 AI Engineering

Hands-on experience shaping how frontier models learn — from generating the preference data behind RLHF to building the eval infrastructure that measures whether a model is actually right.

  • Frontier-model evaluation @ Anthropic (via Revelo). Evaluated & ranked 50+ AI-generated PRs across real-world Python repos — scoring correctness, test coverage, and edge-case handling at an 80% first-pass acceptance rate — producing the preference labels that fed RLHF and reward-model training.
  • Training-data authorship @ AfterQuery. Author expert coding trajectories (reference solutions + stepwise reasoning) across 12+ repositories, 4 languages, and 6+ problem types, plus agentic coding environments that teach models how an expert resolves multi-file problems.
  • RAG evaluation harnesses. Build eval systems that measure hit rate, MRR, precision@k, and LLM-as-judge groundedness/relevance on auto-generated test sets — making answer quality a reproducible, regression-tested metric instead of a vibe.
  • Applied LLM systems. Embeddings & vector retrieval (pgvector), Claude API & Gemini, LangChain, Socratic RAG tutoring, and token-by-token SSE streaming with source citations.

⚙️ Backend Engineering

Shipping production backends with an eye for latency, resilience, and clean service boundaries.

  • Performance. Rebuilt chat-history storage as a standalone Go microservice over MongoDB, cutting response latency 263ms → 41ms (~84%) as an independently scalable service with its own datastore.
  • Resilience & caching. Integrated 3 external FX APIs behind a Redis write-through cache with scheduled refresh and provider failover — cutting redundant upstream calls ~80%.
  • Event-driven systems. Shipped an event-driven price-drop alert feature (Firebase Cloud Messaging) on a product ranked #1 of 13 company-wide; comfortable with async workers (Celery), WebSockets, and SSE.
  • Architecture & quality. Multi-tenant APIs, JWT/RBAC auth, SSRF-guarded ingestion, encrypted BYOK, Dockerized services, and CI/CD — always backed by real test suites (115- and 70+-test pytest suites on recent projects).

What I'm doing now

  • 🔬 Authoring expert coding trajectories & agentic environments that train frontier LLMs — @ AfterQuery
  • 🏗️ Building RAG systems and scalable backends with FastAPI, Go, and Next.js
  • 📚 Head of Education — mentoring 40+ engineers in DSA, algorithms & system design
  • 🌱 Going deeper on system design, cloud deployment workflows, and Go internals

Selected work

AlgoMentorGraph-Guided RAG Learning Platform Multi-tenant full-stack SaaS (FastAPI + Next.js, 46 endpoints / 19 Postgres tables) with top-12 pgvector retrieval over Gemini embeddings, async Celery + Redis ingestion, and token-by-token SSE streaming. Includes a RAG evaluation harness (hit rate, MRR, precision@k, LLM-as-judge) and a prerequisite-graph engine that recommends what to study next. Shipped with a 115-test suite and CI/CD to DigitalOcean.

SafeSight AnalyticsReal-time Video Surveillance Platform Async FastAPI backend for real-time monitoring — REST + WebSocket alerting, a configurable rules engine, role-based access, and a DeepFace vision integration. Backed by a 70+ test pytest suite and Docker Compose.

Tech I work with

Languages — Python · Go · TypeScript · JavaScript · C++ · Java Backend — FastAPI · Gin (Go) · Node.js · REST · WebSockets · SSE · Celery AI / LLM — RAG · Embeddings · pgvector · LangChain · Claude API · Gemini · RLHF & Evaluation Data / Infra — PostgreSQL · MongoDB · Redis · MySQL · Docker · GitHub Actions · Linux Frontend — React · Next.js · Tailwind CSS

Beyond the code

Google-backed (A2SV) competitive programmer — 1,000+ problems solved across LeetCode & Codeforces, 26+ rated contests. I care about clean architecture, correctness, and code that's a pleasure for the next person to read.

Open to roles in AI/LLM engineering, alignment & evaluation, and backend systems.

Popular repositories Loading

  1. AlgoMentor AlgoMentor Public

    Graph-guided RAG interview prep — FastAPI, pgvector, Next.js, eval harness

    TypeScript 5 1

  2. ShareSpace ShareSpace Public

    React Native / Expo app for student community content sharing

    JavaScript 2

  3. Amaankaa Amaankaa Public

    2

  4. studypal-backend studypal-backend Public

    StudyPal API (Django REST) — notebooks, notes, flashcards & quizzes + Gemini

    Python 2

  5. StudyPal_Frontend StudyPal_Frontend Public

    Study management app — notebooks, flashcards & quizzes (React + TypeScript)

    TypeScript 2

  6. The-matrix-rain-effect The-matrix-rain-effect Public archive

    JavaScript 1