Every guide on local AI hardware buries this under specification tables, so here it is upfront: token generation speed is approximately equal to memory bandwidth divided by model size in memory. 

Full Story