Home / Models & Releases / Article
Models & Releases

Micron's memory wall chart. Compute up ~3x every two years, HBM bandwidth under 2x

R/

r/LocalLLaMA

September 13, 2026 at 10:56 PM

Micron's memory wall chart. Compute up ~3x every two years, HBM bandwidth under 2x

๐Ÿ“Œ From Raghu Sreeramaneni's memory tutorial at hot chips 2026. top line is normalised tflops for tpu v3 through r200, bottom line is hbm2e through hbm4, both log scale, so the distance between them is a lot wider than it looks. The three boxes down the right are the fixes people are actually building. Memory beside the compute, memory closer on a shorter link, then multiply units inside the memory itself. Samsung has that last one shipping in lpddr5x and measured 3.01x tokens a second on llama ...

Micron's memory wall chart. Compute up ~3x every two years, HBM bandwidth under 2x

From Raghu Sreeramaneni's memory tutorial at hot chips 2026. top line is normalised tflops for tpu v3 through r200, bottom line is hbm2e through hbm4, both log scale, so the distance between them is a lot wider than it looks.

The three boxes down the right are the fixes people are actually building. Memory beside the compute, memory closer on a shorter link, then multiply units inside the memory itself. Samsung has that last one shipping in lpddr5x and measured 3.01x tokens a second on llama 3.1 8B.

Full analysis (this slide sits in the memory chapter): https://allaboutchips.com/#memory

submitted by /u/Summit-Star001
[link] [comments]

Read the full article at

r/LocalLLaMA

Visit Source โ†—