News — 48GB VRAM
What Can You Actually Run on a 48GB NVIDIA A40? Measured Model Sizes and Tokens Per Second (2026)
One A40 48GB runs 27–34B models at 23–29 tokens/sec, fits a 70B at Q4 at 12–13.5 tokens/sec, and serves 1,705 tokens/sec of aggregate output on an 8B model at 100 concurrent requests. The measured numbers, the KV cache math behind the 72B cliff, and the three things an Ampere card cannot do.