Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer
NVIDIA's Nemotron 3.5 Lightning NVFP4 delivers 4x faster throughput compressed from 66GB to 22GB
“preserves accuracy while unlocking up to 4x faster throughput”
NVIDIA released a new quantized checkpoint of Nemotron 3.5 Lightning using NVFP4 precision and Quantization-Aware Distillation (QAD), shrinking the model from 66GB to 22GB while achieving up to 4x throughput gains. This matters because it demonstrates practical path to deploying large models with significantly reduced memory footprint and latency. The technique is part of NVIDIA's broader Model Optimizer toolchain for production inference optimization.