Creepybits

Creepybits

Quantization vs Quality Degradation

45 min · 22 de mar de 2026
Portada del episodio Quantization vs Quality Degradation

Descripción

What really happens when we compress AI models? In this episode, we break down the mechanics of quantization and quality degradation. We explore why FP32 is essential for training but complete overkill for inference, and unravel the paradox of why smaller GGUF files run significantly slower than FP8 and NVFP4 on modern GPUs. Finally, we put GGUF and NVFP4 to the ultimate test in video generation, wrapping up with a look at pushing 0.4 megapixel video to 1080p in real-time using NVIDIA's RTX upscaler. Is the loss of precision just a myth?

Comentarios

0

Sé la primera persona en comentar

¡Regístrate ahora y únete a la comunidad de Creepybits!

Prueba gratis

Empieza 7 días de prueba

$99 / mes después de la prueba. · Cancela cuando quieras.

  • Podcasts solo en Podimo
  • 20 horas de audiolibros al mes
  • Podcast gratuitos

Todos los episodios

14 episodios