Creepybits

Quantization vs Quality Degradation

45 min · 22 mrt 2026
aflevering Quantization vs Quality Degradation artwork

Beschrijving

What really happens when we compress AI models? In this episode, we break down the mechanics of quantization and quality degradation. We explore why FP32 is essential for training but complete overkill for inference, and unravel the paradox of why smaller GGUF files run significantly slower than FP8 and NVFP4 on modern GPUs. Finally, we put GGUF and NVFP4 to the ultimate test in video generation, wrapping up with a look at pushing 0.4 megapixel video to 1080p in real-time using NVIDIA's RTX upscaler. Is the loss of precision just a myth?

Reacties

0

Wees de eerste die een reactie plaatst

Meld je nu aan en word lid van de Creepybits community!

Probeer gratis

Probeer 14 dagen gratis

€ 9,99 / maand na proefperiode. · Elk moment opzegbaar.

  • Podcasts die je alleen op Podimo hoort
  • 20 uur luisterboeken / maand
  • Gratis podcasts

Alle afleveringen

14 afleveringen