Cohere has released North Small Translate, an open-weight machine translation model from Cohere and Cohere Labs. It is a sparse […]
Category: Large Language Model
DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse
Long-horizon agents have turned LLM serving into an input-heavy workload. Repeated prefills and million-token contexts leave KV caches that strain […]
Google DeepMind Releases AlphaGenome Atlas With Precomputed Molecular Effect Predictions and AVI Scores for 9 Billion Human DNA Variants
Google DeepMind has released AlphaGenome Atlas, a catalogue of precomputed predictions for the molecular effects of every possible single-nucleotide variant […]
Reducto Releases r-1: A Single Pass Document Parsing Model That Cuts Errors 20% at 1 Cent Per Page
Last week, Reducto announced r-1. It is the first model in a new parsing family built on a rewritten architecture, […]
OpenBMB Releases MiniCPM5-2B: A 2.52B Dense Model Averaging 53.9 Across 34 Benchmarks and Built to Run On Device
OpenBMB has released MiniCPM5-2B, the second checkpoint in the MiniCPM5 series and the follow-up to MiniCPM5-1B. It is a dense […]
Axis Robotics Releases AXIS: A Browser-Based Data Engine With 207 Robot Manipulation Tasks and 50,129 Trajectories
Robot manipulation datasets have grown far slower than the models trained on them, mostly because collection stays closed and centralized. […]
IFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B
Most open model launches release one checkpoint and a benchmark table. The Institute of Foundation Models (IFM) released something wider […]
H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder
Most visual document retrievers in production today are hand-me-downs. ColPali and the models that followed it take a generative vision-language […]
Meta FAIR Introduces AI Research Preference Models (RPMs): Ranking ML Experiments Before Spending GPU Hours
AI research agents can already propose, implement and score their own machine learning experiments. Idea generation is cheap; verification is […]
NVIDIA Releases Personal AI Router (PAIR): An Open Source Virtual Inference Router that Distributes Local AI Requests Across RTX, DGX Spark, and Mac Nodes
Multi-agent workflows have changed the shape of local inference. A lead agent decomposes a task and spawns subagents. What looked […]
