H100 is a topic tracked in our intelligence system with 5 linked articles.
H100, backed by Adam Back, completes an acquisition paid with 790.5 million new shares, boosting its bitcoin treasury to 3,506 BTC.
A technical deep-dive into vLLM's high-throughput inference stack, covering core architecture (engine core, KV cache, paged attention), advanced features (chunked prefill, prefix caching, guided and speculative decoding, disaggregated P/D), multi-GPU scaling (Uniproc/MultiProc, TP/PP/DP), distributed serving, and latency/throughput benchmarking with concrete configs and code sketches.
HBM memory is scaling to multi-layer stacks with high capacities (HBM4E up to ~1 TB per GPU) and complex packaging, while regulatory headwinds and China subsidies are shaping supply dynamics.
GB200 NVL72 lags in economics and reliability versus H100 despite signaling potential QE for frontier training; capex is ~1.6x higher per GPU with ~1.6x higher TCO driven mainly by power draw, while H100 software improvements continuously reduce per-token costs; GB200 NVL72 software ramp and diagnostics lag, with no large-scale frontier runs yet, complicating deployment decisions.
Kog claims real-time LLM inference on standard datacenter GPUs can reach about 3,000 tokens/s per request on a 2B model by co-designing a monokernel runtime, GPU code, and a Laneformer architecture, with scalability toward frontier MoEs as memory bandwidth grows.
A multi-article TechCrunch digest highlights mega AI funding, open-social/content interoperability, emerging AI-token futures markets, startup milestone metrics, and notable cybersecurity/regulatory risk signals.
Subscribe for real-time topic updates and unlimited access to our intelligence platform.