Skip to content

ethanlee928/fp8-multiplier-adder

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

4 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

FP8 Multiplier & Adder

Implementation of NVIDIA FP8 E4M3 Multiplier and Adder in Verilog.

  • 1 bit sign
  • 4 bits exponent
  • 3 bits mantissa
  • Range: ±448
  • Ability to represent NaN.

NVIDIA FP8 Format:

Property E4M3 E5M2
Exponent bias 7 15
Infinities N/A S.11111.00₂
NaN S.1111.111₂ S.11111.{01, 10, 11}₂
Zeros S.0000.000₂ S.00000.00₂
Max normal S.1111.110₂ = 1.75 × 2⁸ = 448 S.11110.11₂ = 1.75 × 2¹⁵ = 57,344
Min normal S.0001.000₂ = 2⁻⁶ S.00001.00₂ = 2⁻¹⁴
Max subnorm S.0000.111₂ = 0.875 × 2⁻⁶ S.00000.11₂ = 0.75 × 2⁻¹⁴
Min subnorm S.0000.001₂ = 2⁻⁹ S.00000.01₂ = 2⁻¹⁶

IEEE 754-2019: IEEE Standard for Floating-Point Arithmetic:

Exponent Fraction Object Value
0 0 0 0
0 Nonzero Denormalized number (-1)^S × (0.F) × 2^(E-B)
Nonzero Anything Floating-point number (-1)^S × (1.F) × 2^(E-B)
All “1” 0 infinity
All “1” Nonzero NaN (not a number)

Reference:

Simulation with Testbench

# FP8 Multiplier
make compile-mul

# FP8 Adder
make compile-add

# Run simulation
make sim

# Cleanup artifacts
make clean
  • Click on Signal List -> Select signals -> Right Click -> New Wave View
  • Click on Green right arrow -> simulation.

Results

FP8 Adder

Figure 1: Testbench Waveform for the FP8 Adder.

FP8 Multiplier

Figure 2: Testbench Waveform for the FP8 Multiplier.

About

Verilog design of FP8 E4M3 Multiplier and Adder

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages