Platform teams running AI on Kubernetes rarely run one thing. They run a queueing system, a distributed runtime, GPU node […]
Category: Staff
Anthropic Launches Claude Code Projects in Beta: Parallel Cloud Sessions That Keep Running After You Close Your Laptop
Anthropic redesigned Projects in Claude Code. The old project was a folder: some files plus one chat. The new one […]
OpenAI Releases a Model Misalignment Disclosure Framework With 3 Review Tracks and 6 Incident Reports From RL Training
OpenAI has released a new framework for tracking, investigating, and disclosing misalignment in its own models. The OpenAI team announced […]
Google Research Introduces Retrieve-for-Train (R4T): An RL-Compiled Diffusion Retriever for 12× to 20× Faster Query Fan-Out
Search and recommendation systems increasingly need to return a set of results, not one best match. A query like ‘camping […]
Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers
Nunchux AI has released VC-Attention, a training-free low-bit attention kernel built for video Diffusion Transformers (DiTs). It targets 2 problems […]
Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data
Computational papers ship code that readers must clone, install, configure and debug. That cost keeps useful methods locked inside PDFs. […]
Knowledgator Releases GLiFormer: A 575M-Parameter Encoder That Hits 91.10 F1 on Nested JSON Extraction Without Generating Tokens
Knowledgator Engineering has released GLiFormer, a schema-conditioned encoder framework for information extraction. One model handles named-entity recognition (NER), text classification, […]
Prior Labs Releases TabPFN-3.5: A Tabular Foundation Model That Beats the Winning Otto Kaggle Solution With Default Settings
Prior Labs has released TabPFN-3.5, the newest version of its tabular foundation model. It predicts on a table in a […]
Inside NVIDIA’s cuDNN Graph API: Fusion, Autotuning, and Plan Reuse with cuDNN Frontend
In this tutorial, we work through the cuDNN Frontend‘s graph API from below the framework: we describe a computation as […]
Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents
Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue models to date. […]
