# Measured on the rented A100-SXM4-80GB (massedcompute_A100_sxm4_80G), 2026-07-09 # via peak.cu (fp64 FMA microbenchmark + STREAM-triad), CUDA 12.4, driver 580.126.09 device = NVIDIA A100-SXM4-80GB, SMs = 108 fp64_peak_TFLOPs 7.737 # achievable; nominal vector fp64 is 9.7 -> 80% triad_bandwidth_GBs 1751.7 # achievable; nominal HBM2e is 2039 -> 86% measured_ridge_flop_per_byte 4.42 # 7.737e12 / 1751.7e9 (model assumed 4.76 nominal)