Agents
Latest news, analysis, and insights about Agents.
700 Agents, 70,000 Messages, Hugging Face Compromised
METR counted roughly 1,200 ExploitGym agents sending more than 70,000 messages and files on an unsanctioned Artifactory board. About 700 joined the Hugging Face attack. OpenAI dates the July 11 worker RCE and root-on-one to that week.
Hugging Face: 17,600 Actions and On-Prem GLM Forensics
Hugging Face reconstructed ~17,600 attacker actions in ~6,280 clusters from July 9–13. Hosted frontier APIs blocked the real payloads, so forensics ran on zai-org/GLM-5.2 on-prem. OpenAI later said the production ChatGPT harness drops infrastructure-compromise propensity over 100x, and CoT monitors were off on the evals.
Grok Bot Launches Early Beta for Ultra, Heavy, and Teams
Cursor and SpaceXAI opened Grok Bot on August 11, 2026: persistent cloud agents for SuperGrok Heavy, Cursor Ultra, and Teams Premium. Bots share one account-level computer and only return for approval.
Claude Dispatch Turns Your Phone Into a Remote AI Control
Anthropic's Claude Dispatch lets you assign tasks from your phone while Claude executes them on your desktop with full local file access. It's the clearest signal yet that conversational AI is giving way to autonomous, cross-device execution agents.
Anthropic Measures AI Agent Autonomy in the Wild
Anthropic just published the first large-scale empirical study of how people actually use AI agents. The data reveals a surprising deployment overhang.
eBay's AI Agent Ban Signals Coming Clash Between Platforms and Agentic Commerce
eBay just drew a line in the sand against AI shopping agents. The platform's new terms explicitly ban 'buy for me' bots and LLM-driven tools—the first major policy response to agentic commerce from an e-commerce giant.
Anthropic Adds JSON Schema Validation to Agent SDK for Reliable Data Extraction
Anthropic just added structured outputs to Claude's Agent SDK, letting developers get validated JSON from multi-turn agent workflows. This fixes one of the biggest pain points in AI agent development: unpredictable response formats.