H200 is a topic tracked in our intelligence system with 5 linked articles.
Kog claims real-time LLM inference on standard datacenter GPUs can reach about 3,000 tokens/s per request on a 2B model by co-designing a monokernel runtime, GPU code, and a Laneformer architecture, with scalability toward frontier MoEs as memory bandwidth grows.
A multi-article TechCrunch digest highlights mega AI funding, open-social/content interoperability, emerging AI-token futures markets, startup milestone metrics, and notable cybersecurity/regulatory risk signals.
Anthropic inks a compute deal with SpaceXAI to access Colossus 1’s ~220,000 Nvidia GPUs and ~300 MW capacity in Memphis, as SpaceXAI eyes an IPO and orbital compute ambitions, all amid regulatory and environmental scrutiny and large cloud-spend implications.
Freeform raises $67M in a Series B to scale laser AI manufacturing, touting on-site H200 AI hardware clusters.
China approved the sale of hundreds of thousands of Nvidia H200 AI chips to Chinese firms, signaling a notable shift in US export policy.
China approves imports of Nvidia H200 AI chips, with over 400,000 units slated for Chinese tech giants, underscoring demand for AI compute amid a push for self-reliance.
Subscribe for real-time topic updates and unlimited access to our intelligence platform.