Krina Vaghela
← Blog

Day 32

What is quantization?

An LLM stores billions of numerical weights.

Quantization represents them with fewer bits for example, 16-bit values become 8-bit or 4-bit values.

Moving from 16-bit to 4-bit can make raw weight storage roughly four times smaller.

This helps models fit on limited hardware and may improve speed.

But lower precision creates rounding errors, so quality can drop.

Speed also depends on hardware and software support.

My takeaway: quantization makes a model cheaper to run, not smarter.

Day 32 of 1% better.

LinkedIn post: https://lnkd.in/p/gBimZvff (opens in a new tab)