The inference API is too slow
The inference API is too slow, for glm 5.2. They claim a tps of 100+ but the output is snail's pace, far far from 100 tps. More like 10tps.
Edit in response to the reply: Your models (/models) list advertises GLM-5.2 with "191 toks/s" in the model card.
As for the average provider claim, I disagree. TensorX, ollama, NeuralWatt all provide 100+ t/s.

Reply from LLMBase





