The ChatGPT moment in 2022 taught AI to talk to people. One of its builders now bets the next moment […]
Category: Machine Learning
Linkup Research Releases SPARSEUP: A 149M-Parameter Open-Source Sparse Embedding Model
Linkup research team releases SPARSEUP, an open-source learned sparse embedding model. The model runs on a 149M-parameter ModernBERT backbone and […]
GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)
First, separate 2 ideas: containers vs. quantization methods Most confusion comes from mixing 2 layers. A container defines how tensors […]
PrismML Releases Ternary Bonsai 2 27B: A 5.9 GB Apache 2.0 Model Retaining 98.2% of Qwen3.8 27B Performance
PrismML has released Ternary Bonsai 2 27B, a ternary-weight version of Qwen3.8 27B. The language model occupies 5.93 GB, against […]
Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use
Alibaba’s Qwen team has released Qwen3.8-Omni-Flash. They called it its first omni-modal model built around agentic capabilities. It accepts text, […]
Small AI models let drones autonomously identify and attack battlefield targets
Skip to content Federated learning on the battlefield Scaleout deploys decentralized AI-driven learning to military bases and drones. Demonstration of […]
Microsoft Open-Sources TauGrid: A Kubernetes-Native Stack for GPU AI Workloads
Platform teams running AI on Kubernetes rarely run one thing. They run a queueing system, a distributed runtime, GPU node […]
Google Research Introduces Retrieve-for-Train (R4T): An RL-Compiled Diffusion Retriever for 12× to 20× Faster Query Fan-Out
Search and recommendation systems increasingly need to return a set of results, not one best match. A query like ‘camping […]
Nunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers
Nunchux AI has released VC-Attention, a training-free low-bit attention kernel built for video Diffusion Transformers (DiTs). It targets 2 problems […]
Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data
Computational papers ship code that readers must clone, install, configure and debug. That cost keeps useful methods locked inside PDFs. […]
