StorageCraftManmeet Nain@manmeetnainRepository ↗
AI inference lab · v1

GPU Memory Planner

Budget model weights, KV cache, runtime workspace, and operational reserve at the per-GPU allocation boundary—then find the modeled concurrency ceiling.

Model and serving inputs

Required / GPU
KV / request total
Headroom / GPU
Modeled max concurrency

Per-GPU allocation