Quantitative research agents that write their own experiments can corrupt the evidence they later learn from. A leaky feature that […]
Category: Software Engineering
Keenable AI Open-Sources NEEDLE: A Live Search Benchmark That Rebuilds Its Query Set Every Hour
How do you benchmark a web search API when the thing being tested can read the answer key? A search […]
OpenClaw Releases OpenClaw 2.0: Guided Model Setup, 575 ms Control UI Startup, and One Trust Boundary Per Gateway
The OpenClaw team has just released OpenClaw 2.0. The team shipped nothing for nearly seven weeks, after 106 releases in […]
Google AI Introduces EnvHarness: A Programmable Layer That Turns Static Agent Environments Into Adaptive Training Worlds
A team of researchers from Google Cloud AI Research, Washington University in St. Louis and UNC Chapel Hill has released […]
Anthropic Opens a Research Preview of the Model Hardware Standard (MHS): A Shared Specification for AI Agents to Safely Operate Physical Devices
Anthropic has opened a research preview of the Model Hardware Standard (MHS), a shared specification that lets AI agents discover […]
Hugging Face Unveils Microduck: A $399 Open-Source 25 cm Biped You Train with Reinforcement Learning
Most robotics launches ask you to trust a demo video. Pollen Robotics, the Bordeaux robotics team at Hugging Face, is […]
Vercel AI Open-Sources vgpu: A TypeScript WebGPU Library for AI Agent Shaders
Shaders are still the hardest thing to ship on a normal web team. WebGPU gives you the hardware, then hands […]
Best Agent Sandboxes in 2026: Cold Start, Per-Second Pricing, and Network Policy Across E2B, Daytona, Modal, Cloudflare, and Vercel
Every agent that writes code needs somewhere to run it. That “somewhere” is now a product category with at least […]
What Would Have to Be True for Agentic Coding to Replace Junior Engineers
I read every major model release. Most of them ship a coding number. The number goes up. The conclusion everyone […]
Liquid AI Open-Sources Pipette: A Reproducible Benchmarking Suite That Measures On-Device Models, Quantization, Runtime and Hardware Together
Model cards report quality under server-class, full-precision conditions. Those numbers rarely predict how the same model behaves on a phone. […]
