AI Engineering Signal #70
Cerebras capacity absorbed by OpenAI deal, killing waitlist access for independent users
Signals
Cerebras capacity absorbed by OpenAI deal, killing waitlist access for independent users
teams relying on Cerebras inference for fast throughput need alternative routing now.
DeepSeek V4 support merged into llama.cpp
local inference deployments can now run V4 without custom forks; update your serving stack.
vLLM Micro-Agent beats frontier models via intra-API collaboration
multi-agent routing inside a single model API call changes latency and cost assumptions for agentic pipelines.
Web
RAM prices forecast to rise sharply through Q3 and Q4 2026
procurement for inference nodes and training clusters should be pulled forward; spot pricing will not improve this year.
Web
Anthropic Claude now generally available in Microsoft Azure AI Foundry
teams on Azure can route to Claude (current family: Sonnet 4.6 / Opus 4.6) without separate Anthropic contracts.
Web
Apple Neural Engine architecture paper published on arXiv
first detailed public analysis of ANE internals; relevant for on-device inference optimization and model quantization targeting Apple silicon.
ArXiv
Nvidia AI chip sales stalling in China as Huawei gains share
supply chain diversification away from Nvidia for China-deployed inference is accelerating faster than expected.
Web
The Take
Inference infrastructure is fracturing along three axes simultaneously: capacity lock-up at the fast-inference layer (Cerebras), hardware cost spikes in memory, and chip supply bifurcation in China. Teams that assumed stable, open access to any single provider or component are now carrying unpriced risk in their cost models.
Subscribe
Related Signals