Throughput is a topic tracked in our intelligence system with 5 linked articles.
A technical deep-dive into vLLM's high-throughput inference stack, covering core architecture (engine core, KV cache, paged attention), advanced features (chunked prefill, prefix caching, guided and speculative decoding, disaggregated P/D), multi-GPU scaling (Uniproc/MultiProc, TP/PP/DP), distributed serving, and latency/throughput benchmarking with concrete configs and code sketches.
World Chain becomes the first production L2 to stream full EIP-7928 Block Access Lists via Flashblocks on mainnet (Aug 17), targeting up to 1 Ggas/s throughput with BAL updates every ~200 ms and enabling parallel block validation without a hard fork.
RTX 5080 and RTX 3090 reportedly achieve 80 Tok/s on Qwen 3.6 27B Q8, according to a Hacker News discussion.
A kernel optimization claims 2.2x speedup but causes the training loop to slow by 3x, illustrating end-to-end performance trade-offs.
A technical survey of MAC protocols from early ALOHA to modern WiFi/Bluetooth, highlighting throughput metrics, collision-avoidance techniques, and regulatory channel allocations that shape network design.
A research paper demonstrates Rotary GPU enabling local execution of large Mixture-of-Experts models on consumer hardware (8 GB VRAM), achieving 2048 tokens at ~6.3 GB VRAM and ~21 tokens/sec, signaling edge-deployment viability under VRAM constraints.
Subscribe for real-time topic updates and unlimited access to our intelligence platform.