DeepSeek V4 Flash cuts model depth to 43 layers, bucking the trend of deeper-is-better. Architecture innovation over scale brute force.
Texas just froze new data center approvals, putting 20% of the US pipeline at risk. This is what compute centralization looks like as a single point o
Alibaba ships Qwen3.8 Max with 2.4 trillion parameters and calls it the end of the fully free open source era. A 2.4T open-weight model from a non-US
20% of the US project pipeline faces delays while regulators audit power infrastructure.
Centralized compute at this scale demands billions in capital, years of construction, and entirely new energy infrastructure. Those economics support
Training kernel optimization is an underappreciated bottleneck in distributed training.
We are postponing the launch of the SOMA SN114 Conviction Program. 🧵 1/2
Anthropic is building a custom chip design team. The motivation is clear: there is simply not enough hardware to train the models they want at the sca
Anthropic commits $10B in compute capacity with Volta, a cloud startup backed by a $4.7B lease at a Bitdeer facility.
The White House AI safety framework mandates 30-day pre-release review for closed models and explicitly exempts open source.
DeepSeek V4 Flash at $0.14 per million tokens. A production-grade coding model at a price that makes local inference competitive with cloud APIs.