Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Llama-3.2 1B Ultra-Compact quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
Uncompressed weights alone consume 2.4 GB. In addition, the KV cache scales with context tokens and concurrency batch size, plus ~1.8 GB CUDA driver overhead.
If total weights + KV cache exceeds the 141 GB boundary, Tensor Parallelism (TP) or vLLM PagedAttention multi-GPU sharding across NVLink is required.
Modern AWQ and GPTQ retain >98% perplexity compared to FP16 while halving memory footprint and doubling memory-bandwidth-bound token generation speed.