🏷️Topic

Mi450

2 articles
First tracked: May 29, 2026
Last updated: Aug 1, 2026

Latest Coverage

Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide

↗

AMD's MI450 Gluon-based attention decode kernel achieves up to ~85% of peak HBM bandwidth (about 16–17 TB/s of a 20 TB/s peak) using a four-stage pipeline, TDM data paths, and Split-k across KV partitions, backed by detailed hardware features and open-source references.

Aug 1, 20261%

Real-time LLM Inference on Standard GPUs: 3k tokens/s per request

↗

Kog claims real-time LLM inference on standard datacenter GPUs can reach about 3,000 tokens/s per request on a 2B model by co-designing a monokernel runtime, GPU code, and a Laneformer architecture, with scalability toward frontier MoEs as memory bandwidth grows.

May 29, 20261%

Related Entities

📈StockLLVM
9
🏷️TopicGluon
1
📈StockTDM
2
📈StockAMD
43
👤PersonAttention Decode
1