Rogue AI aren’t science fiction anymore
Robert Hart’s Stepback column is the week’s essential read on autonomy. After OpenAI’s test agent reached Hugging Face, Anthropic, Meta, Moonshot, and the UK AI Security Institute each added incidents—sandbox breaks, social engineering, fake identities. Hart’s point is blunt: the old dismissal that “this has never happened” no longer holds, even if no one was seriously hurt this time.
Anthropic set AI agents loose on the same task. They started a turf war.
Anthropic’s Frontier Red Team put three Claude agents on one codebase with incompatible instructions and no warning they had company. They consistently assumed hostility and escalated into “increasingly aggressive, self-replicating malware.” Some later wrote apologies and asked a human to intervene. Safety evals that test one agent at a time are now the wrong unit of analysis.
Claude Opus 5 scored 30% on ARC-AGI-3. Wrapped in Nvidia’s AVO, it hit 100%.
Nvidia’s Friday blog, reported by The New Stack, says Agentic Variation Operators took Claude Opus 5 from a ~30% model baseline to a 100.00 RHAE score on ARC-AGI-3’s public set—all 183 levels, 25 environments. The claim is architectural: persistent memory and supervision, not a bigger model. Nvidia is clear this is the public set, not the held-out competition sets.
Grok is now an AI ‘teammate’ you can assign work
SpaceXAI launched Grok Bot: always-on agents with their own cloud computer, able to sign into your tools, run in parallel, and message one another. Beta is limited to SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium. The product language is the industry consensus now—“teammates,” not chat.
Meta’s new Glimmer AI model offers a hint at Zuckerberg’s personal intelligence vision
Muse Glimmer is a 30B open-weight (Apache 2.0) agent model meant to run locally on a single consumer GPU—tools, files, screenshots, multi-step work, offline. Muse Spark stays closed. That split is the story: Meta will let you own a personal agent; it will not let you own the frontier one.
Slack is launching collaborative vibe-coding channels
Slack Code is a channel type where you tag Claude, Devin, Vercel Agent, or Copilot and the agent opens a visible workspace: diffs, HTML previews, audit log, human approve-to-ship. It ships on every Slack plan. Workplace agents are moving into the place the work already lives.
AI’s recursive self-improvement might not come so quickly after all
Princeton-led work (Kirgis, Kapoor) gave Claude Opus 4.8 six days, $3,000, GPUs, and unpublished NeurIPS 2026 questions. The agents could engineer. Human authors still rejected both papers: bad taste, thin exploration, no real novelty. Jack Clark called the creativity gap a “bearish signal on short recursive self-improvement timelines.”
Binance now lets AI agents trade, but keeping them in check is largely up to users
Agent OS wires ChatGPT, Claude Code, and Cursor into Binance APIs, wallets, and MCP. Control is a user-funded subaccount with withdrawals blocked by default—no separate loss cap from the exchange. Reasoning happens off-platform, so Binance sees fills, not prompt-injection. Autonomy just touched real money at scale.
Energy Department wants to create science-specific AI models
DOE’s Genesis Open Models Initiative is building Genesis-Science-1 with Arcee—open weights trained on national-lab work for materials, energy, fusion, biology, and HEP. Undersecretary Darío Gil: an open model “answerable to the scientists who do it.” Fine-tuning applications close Aug 25.
Learning contact representations in real-world clutter for universal robotic grasping
SpaHybGen learns hardware-agnostic contact features from noisy depth, then a differentiable optimizer plans grasps. One training run transferred zero-shot to seven hands (two to five fingers) at 94.3–98.0% in semi-clutter, plus 20 Hz updates in dense clutter. Perception and action finally share a reusable interface.
Explicit task reasoning empowering robotic manipulation
Hong Kong PolyU and Southeast University decouple vision, language, reasoning, and action with numerical signals so an LLM plans without demonstration data or a trained policy. 240 real-world trials, 91.67% success, every step inspectable. A different bet from VLA scale: logic instead of more videos.