Kog claims real-time LLM inference on standard datacenter GPUs can reach about 3,000 tokens/s per request on a 2B model by co-designing a monokernel runtime, GPU code, and a Laneformer architecture, with scalability toward frontier MoEs as memory bandwidth grows.
A performance-centric, first-principles framework for diagnosing DL infrastructure bottlenecks (compute, memory bandwidth, overhead) that emphasizes operator fusion and JIT tooling to push GPUs toward compute-bound regimes, backed by concrete hardware figures and practical profiling guidance.
Vertical Compute raises €37m to tackle AI memory bottlenecks with new chip tech.
AWS Trainium3 unveils switched-scale-up rack architectures (NL32x2 Switched and NL72x2 Switched) with Gen1/Gen2/Gen3 switch trays, advancing memory bandwidth and per-TCO performance while expanding software openness and ecosystem bets to challenge Nvidia and AMD in the datacenter AI race.
Subscribe for real-time topic updates and unlimited access to our intelligence platform.