Topic: Agents
114 stories
- T1 lifts Terminal-Bench 2.1 score from 43.8% to 64.0% with RL
- Shopify ditches React Native for native apps, crediting AI coding agents
- PARSER splits reading from reasoning in long-context LLM agents, cuts latency by up to 11x
- OpenAI ships Agents API with a managed Codex harness
- Thoughtworks engineers rediscover the blackboard pattern coordinating AI agents
- Meta launches Muse, a personal AI agent built on Secure VM
- GPT-5.6 Sol runs routine qubit calibration measurements at MIT
- TradingAgents open-sources multi-agent LLM trading framework
- NiCE COO Arun Chandra on scaling agentic AI past the pilot stage
- MaxKernel uses AI agents to generate optimized TPU kernels
- Iris-mini and Iris-pro top open-source search agent scores
- TERMy skips LLMs, matching patterns for terminal commands instead
- OpenAI agents used an obscure public wiki to collude and bypass sandboxes
- Anthropic ships blueprints for Claude shopping agents that buy on your behalf
- Zed's Delta reframes Xanadu's vision for AI coding agents
- Hugging Face's funes gives coding agents a memory they own
- Claude Fable 5 ports 1993 Amiga 68000 assembly into Godot
- AI coding agents are reshaping frontend development education
- Claude reverse-engineers Direct2D for Paint.NET on WINE
- UI-Venus-2 scales GUI agent to 170+ apps and desktop OSes
- OpenAI: frontier firms now generate 8.3x more AI output
- A harness design for autonomous coding agents that can't grade itself
- OpenAI says agents that hacked Hugging Face were trained to cheat
- Analytics system flips AI chat from question-first to analyst-first
- OpenClaw ships 2.0, the largest update in its history
- OpenClaw deletes Meta researcher's inbox after losing instruction
- ContextPilot teaches agents proactive context management via RL
- ChatGPT Work runs 223 tools and 44 skills, author finds
- A developer's domain-driven manifest stops AI coding agents from guessing in legacy code
- PILOT gives long-horizon AI agents live self-improvement
- Lemmalog turns LLM memory into a Datalog analysis engine
- Google DeepMind's Co-Scientist now plans and runs lab experiments
- Snowboard Kids N64 decompilation finished in 84 days with AI agents
- Maintainer closes AI-slop PRs used to game GitHub profiles
- Gravitee CEO argues agent complexity, not autonomy, is enterprise AI's real risk
- Tata Communications argues CX needs orchestration, not just automation
- SenteLabsAI releases Open Executive, an AI virtual executive team
- Thinkingbox benchmark: top AI agent reliable only 25% of the time
- Perplexity launches Portable Computer, a local AI agent for Nvidia DGX Spark
- Paul Dix: AI wrote 1M lines of code, then spent months refining it
- LLM agents perform controlled experiments using simulation models
- CarWatch runs Qwen3.6-35B-A3B on a Raspberry Pi 5 for offline car AI
- Prime Agent harness lifts ARC-AGI-3 score from 30% to 95.5%
- OpenAI Codex cuts Asana 5-year testing migration to 2 weeks
- OpenAI brings GPT-5.6 to AWS coding agent Kiro
- Laude open-sources Headlong, a persistent-agent microharness
- Apodex 1.1 claims leading agentic performance from a smaller model
- OpenAI's ChatGPT Work targets non-engineers, but usage stays thin
- Slack launches Slack Code for AI coding agents like Claude, Devin
- Simon Willison: line by line review isn't the best way to verify AI code
- Munder Difflin launches open source harness for encrypted agent clones
- MCP publishes new roadmap covering agent identity and HTTP transport
- Inherent's Faraday agent outperforms Claude Opus 4.8 and GPT-5.5 at replicating research
- Autolith debuts as a terminal coding agent with a live Lisp runtime
- Stampli cuts launch hours 68% using OpenAI's Codex
- Nvidia shows AI agent harness matters more than the model
- AI coding agents cut performance optimization costs by orders of magnitude, blogger argues
- Claude Fable 5 finds no /dev/kvm, routes sandbox tests through GitHub Actions
- AI coding agents make modularity worth designing for, a blog post argues
- Simon Willison: coding agents undermine conceptual integrity
- OneCLI launches open-source sandboxed agent harness for teams
- fx debuts as a 6MB coding agent CLI written in Zig
- StateM lifts GPT-5.6 to 95.3% accuracy on Terminal-Bench 2.1
- IBM Research: ALTK-Evolve calibrates agent memory by model
- GxP-Agent reaches 100% structural match on clinical trial benchmark
- Dan Luu's FRE shows how easily AI agents game benchmarks
- AMD says AI already lifted software productivity 30 percent, eyes agent swarms next
- MathCode turns plain-language math problems into Lean 4 proofs
- Software engineering fundamentals matter more with AI coding agents
- Anthropic finds Claude agents collude on price and sabotage each other
- Mole enforces a hard spending budget on AI research runs
- Mixedbread ships Toast 1, a search subagent 10x cheaper than frontier models
- DeepSeek's new agent harness makes every component a plugin
- Understanding is the new bottleneck in coding with AI agents
- Pi details how it compacts long coding conversations
- OpenAI's GPT-5.6 guide shows agents can match frontier results at a fraction of the cost
- DeepSeek ships open-source Harness with plugin-based design
- OpenAI: frontier firms now use AI agents 8.3x more than typical firms
- IBM Research's ALTK-Evolve matches ACE's agent accuracy at a fraction of the cost
- Discovered Materials benchmarks AI agents on synthesizable chip materials
- Claude and Antithesis catch SQLite's WAL-Reset bug in 15 minutes
- AI agents are erasing software engineering's middle class, essay argues
- Adapt engineer: build the whole feature, split into PRs later
- xAI launches Grok Bot, AI agents that log into your apps
- Agent Memory Distillation lifts small LLM agents by up to 27.2%p on tool use
- A few weeks away from AI reveals eleven paused agent sessions
- Ouroboros self-developing agent sets SOTA on Terminal-Bench, OSWorld
- Humanising LLM output should happen at the boundary, not mid-task
- Mistral AI patents code-block method for pausing and resuming tool calls
- GPT-5.6 Sol cuts token use, lifts pass rates in Model ML's finance decks
- Eric Schmidt: AI agents will outpace AlphaFold in science
- Claude Code adds cross-session messaging between agents
- Vercel drafts Agent Plugins 1.0, AAIF adopts spec for agent skills
- Herdr joins Y Combinator's F26 batch, keeps runtime open source
- Gemini becomes an FPV drone flight coach for one hobbyist
- CopilotKit ships Channels SDK to bring AI agents into Slack and Teams
- Browser Company CEO says almost no one actually uses AI agents
- Univé rolls out ChatGPT Enterprise across its workforce
- Prime Intellect launches Prime Agent, a self-improving coding harness
- Meta launches Muse Code, a coding agent, with Muse Spark 1.2
- Cloudflare open-sources Cloudflare OS, its internal agent platform
- Castform claims RL post-trained open models beat GPT-5.6 Sol on retrieval
- Pi's minimalist harness beats Claude Code and Codex on cost
- Why AI coding agents mean devtools must be open source
- Microsoft releases Orchard, an open framework for training AI agents
- MemoryForge builds lifelong memory to make LLM agents act less generic
- Hoplite launches cloud platform for coding agents
- Claude Opus 4.7 breaks Steve Yegge's Gas Town coding agent
- Circles cuts churn 9% and lifts ARPU 22% with OpenAI tech
- OpenAI launches Presence for production-ready enterprise AI agents
- YC ships qm, a multiplayer agent harness for Slack and web
- Refactoring cut Claude Code's token cost for the same task by 83%
- Qwen-UI-Agent hits SOTA on mobile GUI benchmarks
- Meta plans a big push into personal AI agents, Zuckerberg says