Skip to content

Tag

#q4

Every story tagged q4, newest first.

Quantization Explained: What Q4, Q8, and FP16 Actually Do to a Local Model
Article · aiDeep read

Quantization Explained: What Q4, Q8, and FP16 Actually Do to a Local Model

Quantization shrinks AI model weights from 16-bit floats to 4- or 8-bit. Here is exactly what that trade-off costs you, and how to pick between Q4_K_M, Q8, and FP16 for the model you actually want to run, with an interactive memory calculator.

BitByteCore Silicon Desk · Aug 6, 2026 · 8 min read