Skip to content
View nareshns2004's full-sized avatar

Block or report nareshns2004

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
nareshns2004/README.md

💫 Naresh Kumar

AI-ML Infrastructure Engineering

AI Infrastructure & ML Systems Engineer specializing in distributed ML and LLM inference platforms across Kubernetes, GPU-accelerated compute and high-performance networking

I design and operate production-grade systems in Linux, C++ and Python with deep focus on resource scheduling, workload reliability and cluster-scale performance optimization

Currently building distributed inference platforms and GPU orchestration systems, with strong emphasis on observability, fault tolerance and efficient resource utilization


💻 Skills

Programming

C++ Java Python Go C Git Linux Bash

Environment

AWS Docker Kubernetes Jenkins Terraform Ansible Elastic Search OpenStack

Frameworks

Kafka Hadoop Redis GraphQL TensorFlow pytorch OpenCV Keras

💻 Platforms

nareshns2004 nareshns2004

💻 Statistics

Visitor Count

Pinned Loading

  1. ai-nic-performance-profiler ai-nic-performance-profiler Public

    An observability primitive that closes the attribution gap between NIC hardware counters and distributed training throughput degradation enabling data-driven decisions on fabric topology, RDMA tuni…

    Python

  2. kernel-level-ai-traffic-shaper kernel-level-ai-traffic-shaper Public

    high-performance traffic management system for AI inference workloads built using eBPF and Linux networking primitives

    C 1

  3. kernel-performance-toolkit kernel-performance-toolkit Public

    A Linux kernel performance analysis toolkit for profiling CPU scheduling, memory behavior, NUMA locality, cache efficiency, page faults and Huge Pages

    Python

  4. distributed-training-framework-nccl distributed-training-framework-nccl Public

    Mini Distributed Training Framework using NCCL

    C++

  5. high-performance-llm-inference-engine high-performance-llm-inference-engine Public

    Inference server supporting continuous batching, KV cache management, speculative decoding and INT4/INT8 quantization

    Python

  6. custom-cuda-fused-attention-triton custom-cuda-fused-attention-triton Public

    Building high-performance GPU kernels from first principles by progressively implementing and optimizing deep learning operators in CUDA and Triton

    Python