The resident generalist model measures 47.8 tokens per second on our own bench.
VerifiedSource: internal engine bakeoff, 2026-07
llama.cpp vs. the stock serving engine, same 120B checkpoint, same hardware. llama.cpp: 47.5 to 47.8 tok/s. Stock engine: 32.1 tok/s. Figure is per Spark unit, measured, not a combined or clustered number.
Verified. Measured on our own hardware, not a vendor spec sheet.