Home / Models & Releases / Article
Models & Releases

I have just moved from MacBook M5 pro 48 GB to RTX3090

R/

r/LocalLLaMA

September 13, 2026 at 02:25 PM

๐Ÿ“Œ Using the RTX3090 on a linux machine I built for it and running Qwen 3.8 27B getting average 100t/s compared to my MacBook 20t/s I think I can finally get rid of my Claude subscription, this is good enough for me. I am a software dev and I can get what I need from this set up and be more productive. I am using the linux machine serving the model over an Open AI endpoint and using a custom build desktop app with pi behind it all. Regarding speeds average 20t/s on Mac with llama.cpp with mtp Un...

Using the RTX3090 on a linux machine I built for it and running Qwen 3.8 27B getting average 100t/s compared to my MacBook 20t/s

I think I can finally get rid of my Claude subscription, this is good enough for me. I am a software dev and I can get what I need from this set up and be more productive. I am using the linux machine serving the model over an Open AI endpoint and using a custom build desktop app with pi behind it all.

Regarding speeds average 20t/s on Mac with llama.cpp with mtp Unsloth Q6, Q4 on Mac didnt make much difference in speed for me.

Linux running https://github.com/syv-ai/qwen38-27b-rtx3090 which is vLLM and a Q4 model I believe. Tool calls definitely fail more but qwen3.8 seems smart enough to fix and correct itself

submitted by /u/gutard
[link] [comments]

Read the full article at

r/LocalLLaMA

Visit Source โ†—