Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Qwen-2.5 32B Coder & Math quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
Uncompressed weights alone consume 16.2 GB. In addition, the KV cache scales with context tokens and concurrency batch size, plus ~1.8 GB CUDA driver overhead.
If total weights + KV cache exceeds the 48 GB boundary, Tensor Parallelism (TP) or vLLM PagedAttention multi-GPU sharding across NVLink is required.
Modern AWQ and GPTQ retain >98% perplexity compared to FP16 while halving memory footprint and doubling memory-bandwidth-bound token generation speed.