AI inference lab · v1
GPU Memory Planner
Budget model weights, KV cache, runtime workspace, and operational reserve at the per-GPU allocation boundary—then find the modeled concurrency ceiling.
Model and serving inputs
Required / GPU
KV / request total
Headroom / GPU
Modeled max concurrency