Skip to content
View tatavishnurao's full-sized avatar
💭
Inferencemaxxxxing
💭
Inferencemaxxxxing

Block or report tatavishnurao

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
tatavishnurao/README.md

About Me:

ML Systems Engineer · LLM Inference & Performance

Building efficient, correctness-tested AI systems—from GPU memory and KV-cache mechanics to retrieval infrastructure and real-time serving.

Rust · Python · CUDA · C++ · PyTorch

Socials:

LinkedIn Medium X YouTube email

Pinned Loading

  1. LatentPagedAttention-rs LatentPagedAttention-rs Public

    Rust + cuTile research prototype for paged latent-cache LLM decode attention, validated on an RTX 4060.

    Rust 9

  2. DubPatch DubPatch Public

    Evidence-first review and selective repair for Sarvam-dubbed English Shorts in Indian languages.

    Python

  3. signalflow-rs signalflow-rs Public

    Rust-based Real-time audio DSP frontend for log-Mel feature extraction, streaming, WAV input, stress reports, and its Criterion benchmarks.

    Rust

  4. ml-systems-engineering ml-systems-engineering Public

    Illustrated implementations of core ML/LLM concepts

    Python 1

  5. GlyphPortrait GlyphPortrait Public

    GlyphPortrait is a semantic digital micrography engine for portrait reconstruction. It takes a person image and a user-provided word prompt, extracts/vectorizes visual regions, generates coherent t…

    Python