A hands-on tutorial exploring NVIDIA Tesla V100 architecture, performance characteristics, and real-world capabilities through systematic benchmarking and analysis.
What You'll Learn:
- How to properly benchmark GPU compute performance
- Understanding FP32 vs FP16 precision and Tensor Cores
- Memory bandwidth analysis and optimization
- Thermal and power characteristics under load
- Real quantization performance for LLM deployment
- Real-world LLM inference benchmarking
- KV cache behavior and memory optimization
- Interpreting GPU specifications vs actual performance
Before running any ML models, you need to understand what your GPU can actually do. This phase teaches you how to validate GPU capabilities and interpret the results.
- Compute Performance - How fast can it multiply matrices? (The core ML operation)
- Memory Bandwidth - How quickly can data move in/out of GPU memory?
- Thermal Behavior - Does it throttle under sustained load?
- Efficiency - How close to theoretical peak can we get?
What are FLOPS?
FLOPS (Floating Point Operations Per Second) measures how many calculations the GPU can perform. More FLOPS = faster training/inference.
| Precision | Result | Theoretical Max | What This Means |
|---|---|---|---|
| FP32 | 13.78 TFLOPS | 15.0 TFLOPS | Standard precision, 92% efficiency - excellent |
| FP16 | 89.97 TFLOPS | 30.0 TFLOPS* | Mixed precision with Tensor Cores - 6.5× faster! |
Why FP16 exceeds the spec?
The V100 has two FP16 modes:
- CUDA Cores: 30 TFLOPS (standard)
- Tensor Cores: 125 TFLOPS (specialized for matrix math)
Our 89.97 TFLOPS means Tensor Cores activated automatically, giving us 72% of their theoretical peak. This is why modern deep learning is so fast on V100.
Key Takeaway: Always use FP16/mixed precision for training - you get 6.5× speedup with minimal accuracy loss.
Performance scales with problem size:
| Matrix Size | FP32 Performance | FP16 Performance |
|---|---|---|
| 512×512 | 7.8 TFLOPS | 9.4 TFLOPS |
| 1024×1024 | 10.5 TFLOPS | 54.8 TFLOPS |
| 2048×2048 | 12.6 TFLOPS | 74.2 TFLOPS |
| 4096×4096 | 13.8 TFLOPS | 89.9 TFLOPS |
Why? Larger matrices = better parallelization = more cores active simultaneously.
Practical Implication: Batch your data. Larger batches (within memory limits) = better GPU utilization.
Memory Bandwidth: The Hidden Bottleneck
Many operations are memory-bound, not compute-bound. If data can't reach the cores fast enough, they sit idle.
V100 Memory Bandwidth Test:
| Transfer Size | Bandwidth | Efficiency |
|---|---|---|
| 128 MB | 326.9 GB/s | 36% |
| 512 MB | 812.4 GB/s | 90% |
| 2048 MB | 832.8 GB/s | 92.5% |
What This Shows:
- Small transfers waste bandwidth (overhead dominates)
- Large transfers saturate HBM2 memory (900 GB/s theoretical)
- 92.5% efficiency is excellent - we're getting almost all available bandwidth
Practical Implication:
- Minimize small memory operations
- Fuse operations when possible
- Use larger batch sizes to amortize memory overhead
30-Second Sustained Load Results:
- Starting temp: 36°C
- Peak temp: 42°C
- Average: 37.7°C
- Power: 204W average (68% of 300W limit)
What This Means:
- No thermal throttling - GPU stays cool
- Significant headroom before hitting 300W power limit
- Can run at peak performance indefinitely
- Datacenter cooling is effective
Why This Matters: Some GPUs throttle under sustained load. The V100 doesn't - critical for long training runs.
# Clone repository
git clone
cd v100-deep-dive
# Build Docker container
docker build -t v100-benchmark .# Launch container with GPU access
docker run -d \
--gpus '"device=0"' \
--name v100-benchmark \
--shm-size=16g \
-v $(pwd)/scripts:/workspace/scripts \
-v $(pwd)/results:/workspace/results \
-it v100-benchmark:latest bash
# Execute validation (takes ~2 minutes)
docker exec -it v100-benchmark python3 /workspace/scripts/01_validate_gpu.pyWhat happens during the test:
- Matrix multiplication at multiple sizes (512 to 4096)
- FP32 and FP16 precision comparison
- Memory bandwidth test (128MB to 2GB transfers)
- 30-second thermal stress test
- Efficiency calculations vs theoretical peak
docker exec -it v100-benchmark python3 /workspace/scripts/01_visualize_validation.pyOutput: results/v100_analysis.png
Discovery: Matrix multiplications automatically use Tensor Cores when:
- Using FP16 precision
- Matrix dimensions are multiples of 8
- Operation is
torch.matmul()or similar
Impact: 6.52× speedup over FP32
Use Case: Training large language models, transformers, CNNs
Discovery: Can achieve 92.5% of theoretical memory bandwidth
Impact: Memory-bound operations (embeddings, activations) run at near-optimal speed
Use Case: Large batch inference, data-heavy models
Discovery: Maintains peak performance indefinitely without throttling
Impact: Reliable for multi-day training runs
Use Case: Research experiments, hyperparameter sweeps
Discovery: Runs at 68% TDP under sustained load
Impact: Can push harder if needed, or save power in multi-GPU setups
Use Case: Cost optimization in cloud deployments
After understanding raw GPU capabilities, the next critical question is: what precision should you use for LLM inference? This isn't just about speed - it's about finding the sweet spot between performance, memory usage, and quality.
Large language models are massive. A 7B parameter model in FP32 needs 28GB of memory just for weights. On a 32GB V100, that leaves almost nothing for activations, KV cache, or batching. Quantization solves this by representing numbers with fewer bits, but there's a catch - not all quantization methods are created equal, and hardware support varies dramatically.
FP32 (32-bit Float):
x = (-1)^s × 2^(e-127) × (1 + m/2^23)
where: s = sign (1 bit), e = exponent (8 bits), m = mantissa (23 bits)
FP16 (16-bit Float):
x = (-1)^s × 2^(e-15) × (1 + m/2^10)
where: s = sign (1 bit), e = exponent (5 bits), m = mantissa (10 bits)
BF16 (Brain Float 16):
x = (-1)^s × 2^(e-127) × (1 + m/2^7)
where: s = sign (1 bit), e = exponent (8 bits), m = mantissa (7 bits)
INT8 Symmetric Quantization:
Q(x) = clamp(round(x/s), -128, 127)
x̂ = Q(x) × s
where: s = max(|x|) / 127
INT4 Symmetric Quantization:
Q(x) = clamp(round(x/s), -8, 7)
x̂ = Q(x) × s
where: s = max(|x|) / 7
Mean Squared Error (MSE):
MSE = (1/n) × Σ(xᵢ - Q(xᵢ))²
Signal-to-Noise Ratio (SNR):
SNR = 10 × log₁₀(σ²_signal / MSE)
where: σ²_signal = variance of original signal
Relative Error:
ε_rel = MAE / mean(|x|)
where: MAE = mean absolute error
We tested actual matrix multiplications (GEMM operations) on 4096×4096 matrices - the same operations your LLM does thousands of times per forward pass. Here's what we found:
| Precision | TFLOPS | Speedup vs FP32 | Hardware Support |
|---|---|---|---|
| FP32 | 13.3 | 1.00× (baseline) | CUDA Cores |
| FP16 | 54.0 | 4.05× | ✓ Tensor Cores |
| BF16 | 10.0 | 0.75× | ✗ Falls back to FP32 |
| INT8 | 11.3 | 0.85× | ◐ Software only |
| INT4 | 12.5 | 0.94× | ◐ Software only |
The Big Surprise: BF16 is actually slower than FP32 on V100. Why? The V100's Volta architecture has no hardware support for BF16. It falls back to FP32 execution paths, adding overhead with no benefit. This is why checking hardware support matters.
| Precision | SNR (dB) | Relative Error | Quality Assessment |
|---|---|---|---|
| FP16 | 73.6 | 0.02% | Excellent - unnoticeable loss |
| BF16 | 55.6 | 0.14% | Good - wider range, less precision |
| INT8 | 38.5 | 1.31% | Acceptable - noticeable but usable |
| INT4 | 13.3 | 23.83% | Poor - significant degradation |
Higher SNR is better - it means less noise introduced by quantization.
For a 4096×4096 matrix (representing a layer in your LLM):
| Precision | Memory (MB) | Compression Ratio |
|---|---|---|
| FP32 | 64.0 | 1.0× |
| FP16 | 32.0 | 2.0× |
| BF16 | 32.0 | 2.0× |
| INT8 | 16.0 | 4.0× |
| INT4 | 8.0 | 8.0× |
Real-world impact: A 7B parameter LLM goes from 28GB (FP32) → 14GB (FP16) → 7GB (INT8) → 3.5GB (INT4)
Why it wins on V100:
- 4× faster than FP32 (hardware Tensor Cores)
- 2× memory savings
- 73.6 dB SNR means quality loss is negligible
- Native hardware support
Real-world example: Running a 7B LLM in FP16 gives you ~500-800 tokens/sec on V100, fits comfortably in 32GB, and maintains near-identical quality to FP32.
Why it fails:
- 0.75× (slower than FP32!)
- No hardware acceleration on Volta
- Only useful on Ampere+ GPUs (A100, H100)
The lesson: Always check hardware support. BF16 is great on newer GPUs but actively harmful on V100.
When to use:
- You need to fit a larger model (10B+)
- Can tolerate ~1-2% accuracy loss
- Inference only (training in INT8 is tricky)
The reality: You get 4× memory savings but only 0.85× speed due to software quantization overhead. Worth it if you're memory-constrained, not worth it if you have room for FP16.
When to use:
- Extreme memory constraints
- Quality isn't critical (chatbots, creative writing)
- You've tested and validated acceptable degradation
The reality: 8× memory savings, 23% quality loss, minimal speed benefit. Use cautiously and always validate output quality.
# Real GEMM operations, not just type conversions
def measure_performance(size=4096, dtype=torch.float16):
a = torch.randn(size, size, device='cuda', dtype=dtype)
b = torch.randn(size, size, device='cuda', dtype=dtype)
# Warmup
for _ in range(10):
c = torch.matmul(a, b)
torch.cuda.synchronize()
start = time.perf_counter()
# Actual measurement
for _ in range(50):
c = torch.matmul(a, b)
torch.cuda.synchronize()
end = time.perf_counter()
# FLOPS = 2 × M × N × K for matrix multiplication
flops = 2 * size * size * size
avg_time = (end - start) / 50
tflops = (flops / avg_time) / 1e12
return tflops- Real operations: We do actual matrix multiplications, not just memory copies
- Warmup included: First few iterations are slower (compilation, caching)
- Multiple iterations: Average over 50 runs to reduce noise
- CUDA sync: Ensures GPU finishes before we stop timing
- Large matrices: 4096×4096 saturates GPU, shows real-world performance
# Inside your container
docker exec -it v100-benchmark bash
# Navigate to notebooks
cd /workspace/notebooks
# Run quantization tests (~3-4 minutes)
python3 02_validate_quant.py
# Generate visualization
python3 02_visualize_validation.pyOutput:
data/quantization_results.json- Full numerical resultsresults/phase_2.png- Visual dashboard
✓ Use FP16:
- 4× faster than FP32
- 2× memory savings
- No quality loss
- Hardware accelerated
✗ Avoid BF16:
- Slower than FP32
- No hardware support
- Only useful on Ampere+ GPUs
◐ Consider INT8 if:
- You need to run 10B+ models
- Memory is your bottleneck
- You can afford ~1% quality loss
◐ Consider INT4 if:
- Extreme memory constraints
- Quality isn't critical
- You've validated output
With 32GB V100 VRAM:
| Model Size | FP32 | FP16 | INT8 | INT4 |
|---|---|---|---|---|
| 7B params | ✗ Too big | ✓ Fits | ✓ Plenty of room | ✓ Overkill |
| 13B params | ✗ Way too big | ✗ Tight | ✓ Fits | ✓ Fits |
| 30B params | ✗ | ✗ | ✗ Tight | ✓ Fits |
| 70B params | ✗ | ✗ | ✗ | ✗ Need multiple GPUs |
Moving from theoretical performance to real-world LLM inference, Phase 3 measures actual token generation speed, memory behavior, and optimization strategies using Gemma 3 1B from Unsloth and Google.
Real inference performance differs from raw compute benchmarks. This phase reveals:
- Latency Analysis - Time per token generation (what users actually experience)
- Throughput Testing - Batch processing capabilities
- KV Cache Behavior - Memory growth during generation
- Batch Optimization - Finding the sweet spot for maximum efficiency
Latency (ms/token):
The time it takes to generate each token. Lower is better. This is what determines how "responsive" your model feels to users. At 65ms/token, users see text appearing character by character in real-time.
Throughput (tokens/sec):
How many tokens the system can generate per second. With batching, you can process multiple requests simultaneously, dramatically increasing total throughput even if individual latency stays constant.
KV Cache:
During autoregressive generation, transformers cache key-value pairs from previous tokens to avoid recomputing them. This cache grows linearly with sequence length and is often the memory bottleneck in long-context generation. Understanding KV cache growth is critical for memory planning.
Testing with vanilla transformers on FP16:
| Metric | Value | What This Means |
|---|---|---|
| Average Latency | 65.05 ms/token | Fast enough for real-time chat |
| Average Throughput | 15.37 tokens/sec | Typical single-request speed |
| Generation Speed | ~1 token every 65ms | Smooth user experience |
Real-world context: At this speed, generating a 200-token response takes ~13 seconds - acceptable for most applications.
| Batch Size | Throughput | Efficiency |
|---|---|---|
| 1 | 15 tokens/sec | 100% (baseline) |
| 2 | 30 tokens/sec | 98% |
| 4 | 61 tokens/sec | 103% |
| 8 | 118 tokens/sec | 98% |
| 16 | 238 tokens/sec | 77% |
| 32 | 475 tokens/sec | 97% |
Key Discovery: Batch size 32 is optimal, delivering 31× throughput improvement while maintaining 97% efficiency. Beyond this, memory constraints limit further scaling.
| Sequence Length | KV Cache Size | Total Memory |
|---|---|---|
| 128 tokens | 7.3 MB | 1.88 GB |
| 256 tokens | 12.2 MB | 1.88 GB |
| 512 tokens | 21.9 MB | 1.89 GB |
| 1024 tokens | 30.8 MB | 1.90 GB |
| 2048 tokens | 35.3 MB | 1.91 GB |
Analysis: KV cache grows sub-linearly with sequence length due to efficient caching mechanisms. Model weights (1.86 GB) dominate memory usage, with KV cache adding minimal overhead for Gemma 3 1B.
| Prompt Length | Generation Length | Latency (ms/token) |
|---|---|---|
| 32 tokens | 50 tokens | 65.16 ms |
| 128 tokens | 50 tokens | 64.94 ms |
| 512 tokens | 100 tokens | 64.77 ms |
| 1024 tokens | 200 tokens | 65.01 ms |
Insight: Latency remains consistent (~65ms/token) regardless of prompt or generation length, indicating compute-bound rather than memory-bound operation at these scales.
Finding: Throughput scales nearly linearly up to batch size 32, then efficiency drops.
Why: At batch 32, we hit the optimal balance between:
- GPU utilization (keeping all SMs busy)
- Memory bandwidth (minimizing overhead)
- Cache efficiency (data fits in L2 cache)
Practical Impact: For production serving, use batch size 32 for maximum throughput without wasting resources.
Finding: KV cache uses only 35 MB even at 2048 tokens.
Why: Gemma 3 1B is relatively small with fewer layers/heads, making KV cache minimal compared to model weights.
Practical Impact: Memory planning should focus on batch size and activation memory, not KV cache for models under 3B parameters.
Finding: Latency stays flat across different prompt lengths and generation lengths.
Why: V100's Tensor Cores provide consistent FP16 performance. The autoregressive generation pattern ensures each step takes similar time.
Practical Impact: Predictable performance makes capacity planning straightforward.
Finding: Can run batch size 64 before OOM, using only 6GB peak memory.
Why: Efficient memory management plus small model size (1B params = 1.86GB FP16).
Practical Impact: V100's 32GB allows running much larger models or significantly larger batches.
Important Note: We attempted to benchmark vLLM (which typically provides 3-5× speedup through PagedAttention and continuous batching) but encountered a compatibility issue.
The Problem: Gemma 3 architecture requires BF16 precision in vLLM for numerical stability. However, as Phase 2 demonstrated, V100 has no hardware BF16 support - it falls back to FP32, making vLLM slower than vanilla transformers for this specific model.
The Solution: vLLM works excellently on V100 with models that support native FP16 inference (Llama, Mistral, Phi, etc.). The PagedAttention and continuous batching optimizations provide significant speedups for these architectures.
Lesson Learned: Always verify model architecture requirements match your hardware capabilities. This reinforces Phase 2's finding: hardware-software compatibility is critical for optimal performance.
cd /workspace/notebooks
python3 download_gemma3.pyThis downloads Gemma 3 1B and converts from BF16 to FP16 for V100 optimization.
# Full inference benchmark suite (~5-8 minutes)
python3 03_validate_inference.pyWhat it tests:
- Latency across 12 configurations (4 prompt lengths × 3 generation lengths)
- Throughput scaling (batch sizes 1, 2, 4, 8, 16, 32)
- KV cache growth (128 to 2048 tokens)
- Batch optimization (binary search for optimal batch size)
# Create professional dashboards (~30 seconds)
python3 03_visualize_inference.pyOutput:
data/phase3_inference_results.json- Complete numerical resultsresults/phase3_latency_throughput.png- Latency and throughput analysisresults/phase3_memory_optimization.png- Memory and optimization dashboard
Optimal Configuration:
- Use FP16 precision (hardware accelerated)
- Batch size: 32 for maximum throughput
- Expected performance: 475 tokens/sec aggregate, 65ms/token per request
- Memory budget: ~2GB per batch of 32 requests
Scaling Strategy:
- Single V100: Up to 32 concurrent users with good responsiveness
- Multi-V100: Simple data parallelism, each GPU handles 32 users
- Long context (2K+ tokens): Monitor KV cache, but it's minimal for small models
When to Use Vanilla Transformers vs vLLM:
- Vanilla: Guaranteed compatibility, straightforward implementation
- vLLM: 3-5× better throughput for compatible models (Llama, Mistral, etc.)
- Check model architecture requirements before choosing framework
Special thanks to:
- Unsloth AI and Google for the excellent Gemma 3 1B model
- The model provides an ideal testbed for benchmarking - small enough for fast iteration, yet representative of modern transformer architectures
| Component | Specification | Real-World Capability |
|---|---|---|
| Architecture | Volta (GV100) | Compute 7.0 - supports all modern ML frameworks |
| CUDA Cores | 5,120 | Peak: 13.8 TFLOPS FP32 |
| Tensor Cores | 640 | Peak: 90 TFLOPS FP16 (practical) |
| Memory | 32 GB HBM2 | Can fit 7B parameter models in FP16 |
| Bandwidth | 900 GB/s | 832 GB/s achievable (92.5%) |
| Power | 300W TDP | 200W average under training load |
| Cooling | SXM2 module | Excellent - no throttling observed |
- LoRA training speed
- Full fine-tuning memory requirements
- Gradient checkpointing impact
- Optimal batch size discovery
- NVLink bandwidth utilization
- Model parallelism strategies
- Pipeline parallelism efficiency
- Communication overhead analysis
If your FP16 is slower than expected:
- Check if Tensor Cores are activating
- Matrix dimensions must be multiples of 8
- Verify CUDA capability 7.0+ features enabled
If BF16 is fast:
- You might be on a newer GPU (A100, H100)
- V100 specifically has no BF16 support
If INT8/INT4 are very slow:
- This is expected on V100
- Software quantization has overhead
- Consider bitsandbytes library optimizations
If throughput doesn't scale:
- Check batch size (optimal: 32 for 1B models)
- Monitor GPU utilization (should be >90%)
- Verify no CPU bottlenecks in data loading
CUDA: 12.8
PyTorch: 2.8.0+cu128
Python: 3.10
Container: Ubuntu 22.04
Driver: NVIDIA 550.90.07
Transformers: 4.45.0
v100-deep-dive/
├── notebooks/
│ ├── 01_validate_gpu.py # Phase 1: GPU validation
│ ├── 01_visualize_validation.py # Phase 1: Visualization
│ ├── 02_validate_quant.py # Phase 2: Quantization tests
│ ├── 02_visualize_validation.py # Phase 2: Visualization
│ ├── 03_validate_inference.py # Phase 3: Inference benchmarks
│ ├── 03_visualize_inference.py # Phase 3: Visualization
│ └── download_gemma3.py # Download and setup Gemma 3 1B
├── results/
│ ├── v100_analysis.png # Phase 1 results
│ ├── phase_2.png # Phase 2 results
│ ├── phase3_latency_throughput.png # Phase 3 latency analysis
│ └── phase3_memory_optimization.png # Phase 3 memory analysis
├── data/
│ ├── deep_validation.json # Phase 1 data
│ ├── quantization_results.json # Phase 2 data
│ └── phase3_inference_results.json # Phase 3 data
├── Dockerfile # Container setup
├── docker-compose.yml # Easy deployment
├── requirements.txt # Python dependencies
└── README.md # This guide
Platform: NVIDIA Tesla V100-SXM2-32GB
Purpose: Educational resource for understanding GPU capabilities, quantization trade-offs, and real-world LLM inference performance



