Skip to content
View akrisanov's full-sized avatar

Block or report akrisanov

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
akrisanov/README.md

Software Engineer, LLM Inference & AI Infrastructure

I design and build production LLM inference infrastructure on Kubernetes and NVIDIA GPUs. My current work covers vLLM-based model serving, traffic routing, performance and reliability, capacity planning, and multi-data-center resilience.

Previously, I built and scaled production backend and platform systems across SaaS, fintech, data privacy, and high-traffic consumer products.

Current focus

  • Production LLM serving with vLLM on NVIDIA GPUs
  • Multi-model routing, quotas, fallbacks, and admission control
  • Inference performance, benchmarking, and GPU capacity planning
  • Kubernetes-based model deployment, observability, and reliability
  • Multi-data-center inference architecture and failure recovery

Pinned Loading

  1. llm-serving-lab llm-serving-lab Public

    Reproducible vLLM serving environment with GPU provisioning, workload telemetry, ClickHouse analytics, and Grafana observability.

    Python

  2. webclip webclip Public

    CLI for saving web pages (content, comments, and images) into a local archive or an Obsidian vault.

    Python

  3. docstring-verifier docstring-verifier Public

    A proof-of-concept of AI-assisted VS Code extension that validates Python docstrings against actual code and uses GitHub Copilot through the VS Code Language Model API to generate context-aware Qui…

    TypeScript

  4. talks talks Public

    Slides and materials from my public talks