Inference is a topic tracked in our intelligence system with 5 linked articles.
Infinity raises $15M at a $100M valuation to build a CUDA-style, chip-agnostic kernel software stack for universal AI inference, backed by Touring Capital, Principal VC, and researchers from OpenAI/Anthropic, with 26 employees and a notable customer (D-Matrix).
TechCrunch publishes a living AI glossary with concise, practical definitions of key terms (e.g., AGI, LLM, RLHF) and notes its ongoing updates, plus a small event promo embedded in the page.
SambaNova raises $1B at an $11B valuation in a first close of its Series F, five months after a mega round, with JPMorgan as a customer and Intel partnership, signaling continued demand for AI inference hardware and potential exit paths.
Baseten reportedly near a $1.5B funding round valuing the company around $13B, signaling ongoing investor appetite for AI inference platforms.
DIY install of a Tesla V100 SXM2 datacenter GPU into a gaming PC with an SXM2-to-PCIe adapter yields 32GB VRAM total and ~32 tokens/sec local LLM inference for ~£200, plus caveats on fan noise and software compatibility.
Liquid AI unveils LFM2.5-8B-A1B, an 8B parameter MoE edge model with 128K context, 38T pretraining, expanded tokenizer, and strong on-device benchmarking and tool-calling capabilities.
Subscribe for real-time topic updates and unlimited access to our intelligence platform.