Home / Models & Releases / Article
Models & Releases

llm performance community metric

R/

r/LocalLLaMA

September 13, 2026 at 08:52 AM

๐Ÿ“Œ my question about LLM performance We see a lot of posts about token prediction, token generation per second, etc. But is it really the metric? I can see that DeepSeek V4 Flash 0731 (with DSPark; mac studio + llama.cpp) produces about 22โ€“28 TPS, but I also see that the LLM does a lot of reasoning. And this relates to others. So maybe the correct way is not to check TPS or other metrics, but to check execution: task complexity/second. I don't know if such a metric already exists and if it exis...

my question about LLM performance

We see a lot of posts about token prediction, token generation per second, etc.

But is it really the metric? I can see that DeepSeek V4 Flash 0731 (with DSPark; mac studio + llama.cpp) produces about 22โ€“28 TPS, but I also see that the LLM does a lot of reasoning.

And this relates to others. So maybe the correct way is not to check TPS or other metrics, but to check execution: task complexity/second.

I don't know if such a metric already exists

and if it exists why community doesn't use it by default

submitted by /u/AleksandrNikitin
[link] [comments]

Read the full article at

r/LocalLLaMA

Visit Source โ†—