llm performance community metric
r/LocalLLaMA
September 13, 2026 at 08:52 AM
๐ my question about LLM performance We see a lot of posts about token prediction, token generation per second, etc. But is it really the metric? I can see that DeepSeek V4 Flash 0731 (with DSPark; mac studio + llama.cpp) produces about 22โ28 TPS, but I also see that the LLM does a lot of reasoning. And this relates to others. So maybe the correct way is not to check TPS or other metrics, but to check execution: task complexity/second. I don't know if such a metric already exists and if it exis...
my question about LLM performance
We see a lot of posts about token prediction, token generation per second, etc.
But is it really the metric? I can see that DeepSeek V4 Flash 0731 (with DSPark; mac studio + llama.cpp) produces about 22โ28 TPS, but I also see that the LLM does a lot of reasoning.
And this relates to others. So maybe the correct way is not to check TPS or other metrics, but to check execution: task complexity/second.
I don't know if such a metric already exists
and if it exists why community doesn't use it by default
[link] [comments]
Read the full article at
r/LocalLLaMA
More in Models & Releases
Models & Releases
NVIDIA Unveils RTX PRO 5500 "Blackwell" Workstation GPU with 84 GB GDDR7 Memory
submitted by /u/Lumpy_Phase_9539 [link] [comments]
What happened to AI browsers?
AI browsers were supposedly going to kill Chrome/Safari, and help people quickly buy stuff, book tickets etc. at reasonable prices while avoiding ads. That was such a great sales pitch, but it has been just radio silence from then on. Ar...
[Bi-Weekly Megathread] Project Showcase
Do you have something you'd like to share with the r/LocalLLaMA community. This is the place for it! Recommendation on presentation: Please share plain english description of what your project does and why people should care about it Ho...