Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

V100 Deep Dive: Understanding GPU Capabilities for ML/AI

A hands-on tutorial exploring NVIDIA Tesla V100 architecture, performance characteristics, and real-world capabilities through systematic benchmarking and analysis.

What You'll Learn:

  • How to properly benchmark GPU compute performance
  • Understanding FP32 vs FP16 precision and Tensor Cores
  • Memory bandwidth analysis and optimization
  • Thermal and power characteristics under load
  • Real quantization performance for LLM deployment
  • Real-world LLM inference benchmarking
  • KV cache behavior and memory optimization
  • Interpreting GPU specifications vs actual performance

Phase 1: GPU Deep Validation - Know Your Hardware

Before running any ML models, you need to understand what your GPU can actually do. This phase teaches you how to validate GPU capabilities and interpret the results.

What We're Testing

  1. Compute Performance - How fast can it multiply matrices? (The core ML operation)
  2. Memory Bandwidth - How quickly can data move in/out of GPU memory?
  3. Thermal Behavior - Does it throttle under sustained load?
  4. Efficiency - How close to theoretical peak can we get?

Results Dashboard

V100 Comprehensive Analysis


Understanding the Results

Compute Performance: FP32 vs FP16

What are FLOPS?
FLOPS (Floating Point Operations Per Second) measures how many calculations the GPU can perform. More FLOPS = faster training/inference.

Precision Result Theoretical Max What This Means
FP32 13.78 TFLOPS 15.0 TFLOPS Standard precision, 92% efficiency - excellent
FP16 89.97 TFLOPS 30.0 TFLOPS* Mixed precision with Tensor Cores - 6.5× faster!

Why FP16 exceeds the spec?
The V100 has two FP16 modes:

  • CUDA Cores: 30 TFLOPS (standard)
  • Tensor Cores: 125 TFLOPS (specialized for matrix math)

Our 89.97 TFLOPS means Tensor Cores activated automatically, giving us 72% of their theoretical peak. This is why modern deep learning is so fast on V100.

Key Takeaway: Always use FP16/mixed precision for training - you get 6.5× speedup with minimal accuracy loss.


Matrix Size Matters

Performance scales with problem size:

Matrix Size FP32 Performance FP16 Performance
512×512 7.8 TFLOPS 9.4 TFLOPS
1024×1024 10.5 TFLOPS 54.8 TFLOPS
2048×2048 12.6 TFLOPS 74.2 TFLOPS
4096×4096 13.8 TFLOPS 89.9 TFLOPS

Why? Larger matrices = better parallelization = more cores active simultaneously.

Practical Implication: Batch your data. Larger batches (within memory limits) = better GPU utilization.


Memory Bandwidth: The Hidden Bottleneck

Many operations are memory-bound, not compute-bound. If data can't reach the cores fast enough, they sit idle.

V100 Memory Bandwidth Test:

Transfer Size Bandwidth Efficiency
128 MB 326.9 GB/s 36%
512 MB 812.4 GB/s 90%
2048 MB 832.8 GB/s 92.5%

What This Shows:

  • Small transfers waste bandwidth (overhead dominates)
  • Large transfers saturate HBM2 memory (900 GB/s theoretical)
  • 92.5% efficiency is excellent - we're getting almost all available bandwidth

Practical Implication:

  • Minimize small memory operations
  • Fuse operations when possible
  • Use larger batch sizes to amortize memory overhead

Thermal Analysis: Can It Sustain Performance?

30-Second Sustained Load Results:

  • Starting temp: 36°C
  • Peak temp: 42°C
  • Average: 37.7°C
  • Power: 204W average (68% of 300W limit)

What This Means:

  • No thermal throttling - GPU stays cool
  • Significant headroom before hitting 300W power limit
  • Can run at peak performance indefinitely
  • Datacenter cooling is effective

Why This Matters: Some GPUs throttle under sustained load. The V100 doesn't - critical for long training runs.


Tutorial: Running Phase 1 Benchmarks

Step 1: Setup Environment

# Clone repository
git clone 
cd v100-deep-dive

# Build Docker container
docker build -t v100-benchmark .

Step 2: Run Deep Analysis

# Launch container with GPU access
docker run -d \
  --gpus '"device=0"' \
  --name v100-benchmark \
  --shm-size=16g \
  -v $(pwd)/scripts:/workspace/scripts \
  -v $(pwd)/results:/workspace/results \
  -it v100-benchmark:latest bash

# Execute validation (takes ~2 minutes)
docker exec -it v100-benchmark python3 /workspace/scripts/01_validate_gpu.py

What happens during the test:

  1. Matrix multiplication at multiple sizes (512 to 4096)
  2. FP32 and FP16 precision comparison
  3. Memory bandwidth test (128MB to 2GB transfers)
  4. 30-second thermal stress test
  5. Efficiency calculations vs theoretical peak

Step 3: Generate Visualization

docker exec -it v100-benchmark python3 /workspace/scripts/01_visualize_validation.py

Output: results/v100_analysis.png


Key Capabilities Discovered

1. Tensor Core Acceleration

Discovery: Matrix multiplications automatically use Tensor Cores when:

  • Using FP16 precision
  • Matrix dimensions are multiples of 8
  • Operation is torch.matmul() or similar

Impact: 6.52× speedup over FP32
Use Case: Training large language models, transformers, CNNs

2. Memory Bandwidth Saturation

Discovery: Can achieve 92.5% of theoretical memory bandwidth
Impact: Memory-bound operations (embeddings, activations) run at near-optimal speed
Use Case: Large batch inference, data-heavy models

3. Thermal Stability

Discovery: Maintains peak performance indefinitely without throttling
Impact: Reliable for multi-day training runs
Use Case: Research experiments, hyperparameter sweeps

4. Power Efficiency

Discovery: Runs at 68% TDP under sustained load
Impact: Can push harder if needed, or save power in multi-GPU setups
Use Case: Cost optimization in cloud deployments


Phase 2: Quantization Analysis - Choosing the Right Precision for LLMs

After understanding raw GPU capabilities, the next critical question is: what precision should you use for LLM inference? This isn't just about speed - it's about finding the sweet spot between performance, memory usage, and quality.

Why Quantization Matters for LLMs

Large language models are massive. A 7B parameter model in FP32 needs 28GB of memory just for weights. On a 32GB V100, that leaves almost nothing for activations, KV cache, or batching. Quantization solves this by representing numbers with fewer bits, but there's a catch - not all quantization methods are created equal, and hardware support varies dramatically.

Results Dashboard

V100 Quantization Analysis


The Math Behind Quantization

Floating Point Representations

FP32 (32-bit Float):

x = (-1)^s × 2^(e-127) × (1 + m/2^23)
where: s = sign (1 bit), e = exponent (8 bits), m = mantissa (23 bits)

FP16 (16-bit Float):

x = (-1)^s × 2^(e-15) × (1 + m/2^10)
where: s = sign (1 bit), e = exponent (5 bits), m = mantissa (10 bits)

BF16 (Brain Float 16):

x = (-1)^s × 2^(e-127) × (1 + m/2^7)
where: s = sign (1 bit), e = exponent (8 bits), m = mantissa (7 bits)

Integer Quantization

INT8 Symmetric Quantization:

Q(x) = clamp(round(x/s), -128, 127)
x̂ = Q(x) × s
where: s = max(|x|) / 127

INT4 Symmetric Quantization:

Q(x) = clamp(round(x/s), -8, 7)
x̂ = Q(x) × s
where: s = max(|x|) / 7

Error Metrics

Mean Squared Error (MSE):

MSE = (1/n) × Σ(xᵢ - Q(xᵢ))²

Signal-to-Noise Ratio (SNR):

SNR = 10 × log₁₀(σ²_signal / MSE)
where: σ²_signal = variance of original signal

Relative Error:

ε_rel = MAE / mean(|x|)
where: MAE = mean absolute error

Real Performance Results

We tested actual matrix multiplications (GEMM operations) on 4096×4096 matrices - the same operations your LLM does thousands of times per forward pass. Here's what we found:

Compute Performance

Precision TFLOPS Speedup vs FP32 Hardware Support
FP32 13.3 1.00× (baseline) CUDA Cores
FP16 54.0 4.05× ✓ Tensor Cores
BF16 10.0 0.75× ✗ Falls back to FP32
INT8 11.3 0.85× ◐ Software only
INT4 12.5 0.94× ◐ Software only

The Big Surprise: BF16 is actually slower than FP32 on V100. Why? The V100's Volta architecture has no hardware support for BF16. It falls back to FP32 execution paths, adding overhead with no benefit. This is why checking hardware support matters.

Precision Quality

Precision SNR (dB) Relative Error Quality Assessment
FP16 73.6 0.02% Excellent - unnoticeable loss
BF16 55.6 0.14% Good - wider range, less precision
INT8 38.5 1.31% Acceptable - noticeable but usable
INT4 13.3 23.83% Poor - significant degradation

Higher SNR is better - it means less noise introduced by quantization.

Memory Footprint

For a 4096×4096 matrix (representing a layer in your LLM):

Precision Memory (MB) Compression Ratio
FP32 64.0 1.0×
FP16 32.0 2.0×
BF16 32.0 2.0×
INT8 16.0 4.0×
INT4 8.0 8.0×

Real-world impact: A 7B parameter LLM goes from 28GB (FP32) → 14GB (FP16) → 7GB (INT8) → 3.5GB (INT4)


Understanding the Trade-offs

FP16: The Sweet Spot

Why it wins on V100:

  • 4× faster than FP32 (hardware Tensor Cores)
  • 2× memory savings
  • 73.6 dB SNR means quality loss is negligible
  • Native hardware support

Real-world example: Running a 7B LLM in FP16 gives you ~500-800 tokens/sec on V100, fits comfortably in 32GB, and maintains near-identical quality to FP32.

BF16: Don't Use on V100

Why it fails:

  • 0.75× (slower than FP32!)
  • No hardware acceleration on Volta
  • Only useful on Ampere+ GPUs (A100, H100)

The lesson: Always check hardware support. BF16 is great on newer GPUs but actively harmful on V100.

INT8: When Memory is Tight

When to use:

  • You need to fit a larger model (10B+)
  • Can tolerate ~1-2% accuracy loss
  • Inference only (training in INT8 is tricky)

The reality: You get 4× memory savings but only 0.85× speed due to software quantization overhead. Worth it if you're memory-constrained, not worth it if you have room for FP16.

INT4: Maximum Compression

When to use:

  • Extreme memory constraints
  • Quality isn't critical (chatbots, creative writing)
  • You've tested and validated acceptable degradation

The reality: 8× memory savings, 23% quality loss, minimal speed benefit. Use cautiously and always validate output quality.


How We Measured This

Test Setup

# Real GEMM operations, not just type conversions
def measure_performance(size=4096, dtype=torch.float16):
    a = torch.randn(size, size, device='cuda', dtype=dtype)
    b = torch.randn(size, size, device='cuda', dtype=dtype)
    
    # Warmup
    for _ in range(10):
        c = torch.matmul(a, b)
    
    torch.cuda.synchronize()
    start = time.perf_counter()
    
    # Actual measurement
    for _ in range(50):
        c = torch.matmul(a, b)
        torch.cuda.synchronize()
    
    end = time.perf_counter()
    
    # FLOPS = 2 × M × N × K for matrix multiplication
    flops = 2 * size * size * size
    avg_time = (end - start) / 50
    tflops = (flops / avg_time) / 1e12
    
    return tflops

Why This Method?

  1. Real operations: We do actual matrix multiplications, not just memory copies
  2. Warmup included: First few iterations are slower (compilation, caching)
  3. Multiple iterations: Average over 50 runs to reduce noise
  4. CUDA sync: Ensures GPU finishes before we stop timing
  5. Large matrices: 4096×4096 saturates GPU, shows real-world performance

Running Phase 2 Tests Yourself

# Inside your container
docker exec -it v100-benchmark bash

# Navigate to notebooks
cd /workspace/notebooks

# Run quantization tests (~3-4 minutes)
python3 02_validate_quant.py

# Generate visualization
python3 02_visualize_validation.py

Output:

  • data/quantization_results.json - Full numerical results
  • results/phase_2.png - Visual dashboard

Practical Recommendations for LLM Deployment

For V100 Specifically:

✓ Use FP16:

  • 4× faster than FP32
  • 2× memory savings
  • No quality loss
  • Hardware accelerated

✗ Avoid BF16:

  • Slower than FP32
  • No hardware support
  • Only useful on Ampere+ GPUs

◐ Consider INT8 if:

  • You need to run 10B+ models
  • Memory is your bottleneck
  • You can afford ~1% quality loss

◐ Consider INT4 if:

  • Extreme memory constraints
  • Quality isn't critical
  • You've validated output

Memory Budgets

With 32GB V100 VRAM:

Model Size FP32 FP16 INT8 INT4
7B params ✗ Too big ✓ Fits ✓ Plenty of room ✓ Overkill
13B params ✗ Way too big ✗ Tight ✓ Fits ✓ Fits
30B params ✗ ✗ ✗ Tight ✓ Fits
70B params ✗ ✗ ✗ ✗ Need multiple GPUs

Phase 3: Real LLM Inference Benchmarks

Moving from theoretical performance to real-world LLM inference, Phase 3 measures actual token generation speed, memory behavior, and optimization strategies using Gemma 3 1B from Unsloth and Google.

What We're Testing

Real inference performance differs from raw compute benchmarks. This phase reveals:

  1. Latency Analysis - Time per token generation (what users actually experience)
  2. Throughput Testing - Batch processing capabilities
  3. KV Cache Behavior - Memory growth during generation
  4. Batch Optimization - Finding the sweet spot for maximum efficiency

Understanding Key Concepts

Latency (ms/token):
The time it takes to generate each token. Lower is better. This is what determines how "responsive" your model feels to users. At 65ms/token, users see text appearing character by character in real-time.

Throughput (tokens/sec):
How many tokens the system can generate per second. With batching, you can process multiple requests simultaneously, dramatically increasing total throughput even if individual latency stays constant.

KV Cache:
During autoregressive generation, transformers cache key-value pairs from previous tokens to avoid recomputing them. This cache grows linearly with sequence length and is often the memory bottleneck in long-context generation. Understanding KV cache growth is critical for memory planning.

Results Dashboards

Gemma 3 1B Latency & Throughput Analysis

Gemma 3 1B Memory & Optimization


Phase 3 Results

Single Request Performance

Testing with vanilla transformers on FP16:

Metric Value What This Means
Average Latency 65.05 ms/token Fast enough for real-time chat
Average Throughput 15.37 tokens/sec Typical single-request speed
Generation Speed ~1 token every 65ms Smooth user experience

Real-world context: At this speed, generating a 200-token response takes ~13 seconds - acceptable for most applications.

Batch Processing Capabilities

Batch Size Throughput Efficiency
1 15 tokens/sec 100% (baseline)
2 30 tokens/sec 98%
4 61 tokens/sec 103%
8 118 tokens/sec 98%
16 238 tokens/sec 77%
32 475 tokens/sec 97%

Key Discovery: Batch size 32 is optimal, delivering 31× throughput improvement while maintaining 97% efficiency. Beyond this, memory constraints limit further scaling.

KV Cache Memory Growth

Sequence Length KV Cache Size Total Memory
128 tokens 7.3 MB 1.88 GB
256 tokens 12.2 MB 1.88 GB
512 tokens 21.9 MB 1.89 GB
1024 tokens 30.8 MB 1.90 GB
2048 tokens 35.3 MB 1.91 GB

Analysis: KV cache grows sub-linearly with sequence length due to efficient caching mechanisms. Model weights (1.86 GB) dominate memory usage, with KV cache adding minimal overhead for Gemma 3 1B.

Latency by Configuration

Prompt Length Generation Length Latency (ms/token)
32 tokens 50 tokens 65.16 ms
128 tokens 50 tokens 64.94 ms
512 tokens 100 tokens 64.77 ms
1024 tokens 200 tokens 65.01 ms

Insight: Latency remains consistent (~65ms/token) regardless of prompt or generation length, indicating compute-bound rather than memory-bound operation at these scales.


Key Discoveries

1. Batch Size is Critical

Finding: Throughput scales nearly linearly up to batch size 32, then efficiency drops.

Why: At batch 32, we hit the optimal balance between:

  • GPU utilization (keeping all SMs busy)
  • Memory bandwidth (minimizing overhead)
  • Cache efficiency (data fits in L2 cache)

Practical Impact: For production serving, use batch size 32 for maximum throughput without wasting resources.

2. KV Cache is Not the Bottleneck

Finding: KV cache uses only 35 MB even at 2048 tokens.

Why: Gemma 3 1B is relatively small with fewer layers/heads, making KV cache minimal compared to model weights.

Practical Impact: Memory planning should focus on batch size and activation memory, not KV cache for models under 3B parameters.

3. Consistent Latency Profile

Finding: Latency stays flat across different prompt lengths and generation lengths.

Why: V100's Tensor Cores provide consistent FP16 performance. The autoregressive generation pattern ensures each step takes similar time.

Practical Impact: Predictable performance makes capacity planning straightforward.

4. Memory Efficiency

Finding: Can run batch size 64 before OOM, using only 6GB peak memory.

Why: Efficient memory management plus small model size (1B params = 1.86GB FP16).

Practical Impact: V100's 32GB allows running much larger models or significantly larger batches.


About vLLM Optimization

Important Note: We attempted to benchmark vLLM (which typically provides 3-5× speedup through PagedAttention and continuous batching) but encountered a compatibility issue.

The Problem: Gemma 3 architecture requires BF16 precision in vLLM for numerical stability. However, as Phase 2 demonstrated, V100 has no hardware BF16 support - it falls back to FP32, making vLLM slower than vanilla transformers for this specific model.

The Solution: vLLM works excellently on V100 with models that support native FP16 inference (Llama, Mistral, Phi, etc.). The PagedAttention and continuous batching optimizations provide significant speedups for these architectures.

Lesson Learned: Always verify model architecture requirements match your hardware capabilities. This reinforces Phase 2's finding: hardware-software compatibility is critical for optimal performance.


Running Phase 3 Benchmarks

Download Model

cd /workspace/notebooks
python3 download_gemma3.py

This downloads Gemma 3 1B and converts from BF16 to FP16 for V100 optimization.

Run Inference Benchmarks

# Full inference benchmark suite (~5-8 minutes)
python3 03_validate_inference.py

What it tests:

  • Latency across 12 configurations (4 prompt lengths × 3 generation lengths)
  • Throughput scaling (batch sizes 1, 2, 4, 8, 16, 32)
  • KV cache growth (128 to 2048 tokens)
  • Batch optimization (binary search for optimal batch size)

Generate Visualizations

# Create professional dashboards (~30 seconds)
python3 03_visualize_inference.py

Output:

  • data/phase3_inference_results.json - Complete numerical results
  • results/phase3_latency_throughput.png - Latency and throughput analysis
  • results/phase3_memory_optimization.png - Memory and optimization dashboard

Practical Recommendations

For Production Deployment on V100

Optimal Configuration:

  • Use FP16 precision (hardware accelerated)
  • Batch size: 32 for maximum throughput
  • Expected performance: 475 tokens/sec aggregate, 65ms/token per request
  • Memory budget: ~2GB per batch of 32 requests

Scaling Strategy:

  • Single V100: Up to 32 concurrent users with good responsiveness
  • Multi-V100: Simple data parallelism, each GPU handles 32 users
  • Long context (2K+ tokens): Monitor KV cache, but it's minimal for small models

When to Use Vanilla Transformers vs vLLM:

  • Vanilla: Guaranteed compatibility, straightforward implementation
  • vLLM: 3-5× better throughput for compatible models (Llama, Mistral, etc.)
  • Check model architecture requirements before choosing framework

Acknowledgments

Special thanks to:

  • Unsloth AI and Google for the excellent Gemma 3 1B model
  • The model provides an ideal testbed for benchmarking - small enough for fast iteration, yet representative of modern transformer architectures

Hardware Specifications Reference

Component Specification Real-World Capability
Architecture Volta (GV100) Compute 7.0 - supports all modern ML frameworks
CUDA Cores 5,120 Peak: 13.8 TFLOPS FP32
Tensor Cores 640 Peak: 90 TFLOPS FP16 (practical)
Memory 32 GB HBM2 Can fit 7B parameter models in FP16
Bandwidth 900 GB/s 832 GB/s achievable (92.5%)
Power 300W TDP 200W average under training load
Cooling SXM2 module Excellent - no throttling observed

What's Next?

Phase 4: Fine-tuning Performance

  • LoRA training speed
  • Full fine-tuning memory requirements
  • Gradient checkpointing impact
  • Optimal batch size discovery

Phase 5: Multi-GPU Scaling

  • NVLink bandwidth utilization
  • Model parallelism strategies
  • Pipeline parallelism efficiency
  • Communication overhead analysis

Understanding Your Results

If your FP16 is slower than expected:

  • Check if Tensor Cores are activating
  • Matrix dimensions must be multiples of 8
  • Verify CUDA capability 7.0+ features enabled

If BF16 is fast:

  • You might be on a newer GPU (A100, H100)
  • V100 specifically has no BF16 support

If INT8/INT4 are very slow:

  • This is expected on V100
  • Software quantization has overhead
  • Consider bitsandbytes library optimizations

If throughput doesn't scale:

  • Check batch size (optimal: 32 for 1B models)
  • Monitor GPU utilization (should be >90%)
  • Verify no CPU bottlenecks in data loading

Environment Details

CUDA: 12.8
PyTorch: 2.8.0+cu128
Python: 3.10
Container: Ubuntu 22.04
Driver: NVIDIA 550.90.07
Transformers: 4.45.0

Repository Structure

v100-deep-dive/
├── notebooks/
│   ├── 01_validate_gpu.py          # Phase 1: GPU validation
│   ├── 01_visualize_validation.py  # Phase 1: Visualization
│   ├── 02_validate_quant.py        # Phase 2: Quantization tests
│   ├── 02_visualize_validation.py  # Phase 2: Visualization
│   ├── 03_validate_inference.py    # Phase 3: Inference benchmarks
│   ├── 03_visualize_inference.py   # Phase 3: Visualization
│   └── download_gemma3.py          # Download and setup Gemma 3 1B
├── results/
│   ├── v100_analysis.png           # Phase 1 results
│   ├── phase_2.png                 # Phase 2 results
│   ├── phase3_latency_throughput.png   # Phase 3 latency analysis
│   └── phase3_memory_optimization.png  # Phase 3 memory analysis
├── data/
│   ├── deep_validation.json        # Phase 1 data
│   ├── quantization_results.json   # Phase 2 data
│   └── phase3_inference_results.json   # Phase 3 data
├── Dockerfile                       # Container setup
├── docker-compose.yml              # Easy deployment
├── requirements.txt                # Python dependencies
└── README.md                       # This guide

Platform: NVIDIA Tesla V100-SXM2-32GB
Purpose: Educational resource for understanding GPU capabilities, quantization trade-offs, and real-world LLM inference performance

About

Educational benchmark suite for NVIDIA V100: FP16/FP32 compute analysis, precision quantization testing, and transformer inference profiling.

Topics

Resources

Stars

2 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages