OpenaiJalapeño custom inference chip cuts latency and power use – practical implications for engineers4 min read·1mo ago·via The New Stack