Quantizing MedGemma to INT4 (GPTQ/W4A16): Everything That Broke Along the Way
Quantized Google's MedGemma-1.5-4B (a medical vision-language model) to INT4 (W4A16) via llm-compressor's GPTQModifier, for self-hosted deployment. 8.6 GB in BF16 -> 5.2 GB quantized. Full step-by-ste
Jul 14, 20266 min read3
