Inference is a topic tracked in our intelligence system with 5 linked articles.
AMD acquires Taalas, a Canada-based AI inference chip startup known for SRAM-based, high-token-speed inference (HC1 for Llama3.1-8B) with integration planned into AMD’s Instinct GPUs; deal terms undisclosed and closing contingent on regulatory approvals.
AI memory demands are driving a multi-tier memory hierarchy (HBM, LPDDR, SOCAMM, DRAM) with a notable real-world efficiency edge from SOCAMM, exemplified by Micron’s 256GB SOCAMM and its one-third power draw versus DDR5 RDIMM.
Sequoia Capital leads Etched's $300 million Series C at a $10 billion pre-money valuation, validating the startup's specialized inference silicon and cluster-scale architecture as critical infrastructure for the AI boom.
Infinity raises $15M at a $100M valuation to build a CUDA-style, chip-agnostic kernel software stack for universal AI inference, backed by Touring Capital, Principal VC, and researchers from OpenAI/Anthropic, with 26 employees and a notable customer (D-Matrix).
TechCrunch publishes a living AI glossary with concise, practical definitions of key terms (e.g., AGI, LLM, RLHF) and notes its ongoing updates, plus a small event promo embedded in the page.
SambaNova raises $1B at an $11B valuation in a first close of its Series F, five months after a mega round, with JPMorgan as a customer and Intel partnership, signaling continued demand for AI inference hardware and potential exit paths.
Subscribe for real-time topic updates and unlimited access to our intelligence platform.