Creepybits

Quantization vs Quality Degradation

45 min · 22. mar. 2026
episode Quantization vs Quality Degradation cover

Description

What really happens when we compress AI models? In this episode, we break down the mechanics of quantization and quality degradation. We explore why FP32 is essential for training but complete overkill for inference, and unravel the paradox of why smaller GGUF files run significantly slower than FP8 and NVFP4 on modern GPUs. Finally, we put GGUF and NVFP4 to the ultimate test in video generation, wrapping up with a look at pushing 0.4 megapixel video to 1080p in real-time using NVIDIA's RTX upscaler. Is the loss of precision just a myth?

Comments

0

Be the first to comment

Sign up now and become a member of the Creepybits community!

Get Started

2 months for 19 kr.

Then 99 kr. / month · Cancel anytime.

  • Podcasts kun på Podimo
  • 20 lydbogstimer pr. måned
  • Gratis podcasts

All episodes

14 episodes