Archive
1885 stories
September 2026
- 25 Fields Medal winners warn AI benchmarks harm mathematics
- Anthropic details Claude misuse for weapons, mass surveillance, and Chinese distillation
- Anthropic says it blocked five potential bioweapon plots
- Anthropic sued over deceptive Claude usage limits
- Bengio warns AI training process makes agents deceptive
- Discovery Certification Protocol audits AI research agents' claimed discoveries
- EPA plans to scrap public comment on data center pollution permits
- Fine-tuned 4B Qwen 3.5 beats GPT-5.6 on MetroLLM-Bench
- Former DeepMind VP Vinyals doubts sudden intelligence explosion, launches Discovery Loop
- Garry Tan wants US open-weight AI labs to distill frontier models too
- Google Project Zero releases MAccConc to test Linux kernel race conditions
- Google reroutes search result links to curb scraping
- Graphify C# brings compiler-accurate Find Usages to coding agents
- Image tokenizer choice can affect text modeling, study finds
- Isar Aerospace's Spectrum rocket reaches orbit as satellite firms eye SpaceX alternatives
- litelm reimplements LiteLLM's core in about 2,900 lines
- MaP-WAM turns robot memory into plans, hits 83.3% on RMBench
- Mecka AI nears $500M valuation in Sequoia-led round
- Meta says it fixed AI prompts after invasive questions about kids
- Meta sued over Facebook, Instagram photos used for NameTag, AI training
- Negative self-distillation trains LLMs to avoid flawed reasoning
- Nemotron 3 Ultra reaches gold-medal score at IMO 2026
- New Mexico's Supreme Court fines lawyer $5,000 over ChatGPT-fabricated witnesses in murder case
- OpenAI agents ran an undisclosed RubyGems attack, report says
- OpenAI asks Congress if a coordinated AI slowdown is legal
- OpenRouter's automatic fallbacks can change model behavior
- Pandas loses to Polars, DuckDB by up to 19x in 1-billion-row test
- Perplexity trusts GPT-6 Astra with full end-to-end systems
- Qwen3-Coder and Gemma 4 regress by 4 to 30 points from imitating an expert
- Recursive Code World Models build 3D scenes from one image
- RTK doesn't cut AI coding costs, Terminal-Bench 2.1 finds
- Rune IDE goes open source under GPLv3, plans a revenue share for contributors
- Starlink leakage floods SKA-Low's key radio frequencies, study finds
- Survey maps ways to cut inference costs in VideoLLMs
- SyncWorld simulates robot action outcomes zero-shot in unseen setups
- Three JFrog Artifactory bugs remain under active attack despite patches
- Timnit Gebru argues AI doom talk distracts from real harms
- Tiny Aya L2-Thinker tops 93% in-language reasoning across 60 languages
- Tinybird distills years of lessons from running ClickHouse at scale
- UniH3 achieves state-of-the-art all-in-one medical image restoration
- World in World lets frozen video models explore new camera views
- X-AuT prunes Qwen3-ASR's audio encoder and lowers its error rate
- AgentGrad hits SOTA on five multi-agent benchmarks, 2.5x faster
- Anthropic disrupts Claude-automated Russian espionage campaign
- Anthropic finds a fourth Claude incident involving unauthorized access to real systems
- Φ-Bench tests whether LLMs can engineer the infrastructure that runs them
- Claude Fable 5.1 drops hedges and stock phrases, but answers grow longer
- Cognition ships SWE-2, nearing GPT-6 Astra at a quarter of the cost
- Complete reasoning traces add little value in LLM post-training, study finds
- Data size drives grokking onset far more than model width
- DeepMind banned public discussion of AI extinction risk, former staffer says
- EvoSafeHarness cuts DecodingTrust-Agent attack success from 45.6% to 10.0%
- Federated fire detection removes single point of failure with rotating coordinator
- Gemini app launches for Windows with an Alt + Space shortcut
- Gemini's rewrites turn a Chandler passage into horror after 31 iterations
- GeoSteer replaces one-step LLM activation steering with geodesic optimization
- Google's ToolGrad flips tool-use data generation to answer-first
- Google signs 22-year nuclear power deal in €13bn Finland AI investment
- GPT-6 Astra: Sebastian Raschka digs into the looped-transformer rumor
- Gradio Workflow rebuilds AUTOMATIC1111 as a 73-node canvas
- Halo improves point forecasts by also estimating their uncertainty
- INDXcoin, a "God-driven" cryptocurrency, collapses after taking more than $3 million from more than 500 people
- NCP-ArchPreview matches OLMo-3-7B's pretraining loss on 51.3% of tokens
- Nvidia's Jensen Huang reiterates 70% revenue growth guidance
- ON.energy pitches medium-voltage UPS fix for AI data center outages
- OpenAI pauses new Pro plan sign-ups as Astra demand surges
- OpenAI ships Agents API with a managed Codex harness
- OpenJDK's JEP 544 cuts Java startup time up to 80% with AOT caching
- PARSER splits reading from reasoning in long-context LLM agents, cuts latency by up to 11x
- PlanetScale launches Neki, a sharded Postgres database
- Proof of Capture: an open source counterpart to Apple Reference Image, using steganography
- Rule-chaining framework solves 230 of 240 ARC-AGI-2 tasks
- Rust becomes a Tier-1 language at Microsoft, gains new MSVC compiler backend
- SchemeArena finds instrumental goals drive AI agent scheming
- Schools are catching on to Big Tech's AI playbook
- SenseNova-U1.5 unifies image understanding and generation without encoders or VAEs
- Shopify ditches React Native for native apps, crediting AI coding agents
- Slack launches Slackforce Surfaces to build dashboards and reports in chat
- SpatialBlock-15k trains vision-language models on 3D spatial reasoning
- SWE-Bench Pro Verified closes reward-hacking loopholes: some models score lower
- T1 lifts Terminal-Bench 2.1 score from 43.8% to 64.0% with RL
- Universal Music unveils AI remix platform with ElevenLabs
- WebGPU bug freezes M-series Macs: Apple calls it not a security issue
- YuE2 outperforms Suno v5 on WildSongBench music benchmark
- A*-Thought-V2 cuts LLM response length up to half, lifts accuracy
- Anthropic's own model frames CEO Amodei's job forecasts as an outlier scenario
- Anthropic to watermark all future Claude text
- Apple details how the Secure Exclave keeps Watch audio private
- Apple introduces iPhone Duo, its first foldable iPhone
- Apple unveils iPhone 18 Pro with variable aperture camera and Siri AI
- AuK unifies speech generation and editing in one open-source model
- BeaconKV cuts KV cache memory up to 5.8x for reasoning models
- Calif Research demos WeWorm, a zero-click WeChat worm built with AI
- China bans AI companions for minors, reins in adult use
- Claude Opus 4.6 produces 30x more errors than GPT-5.4 in new agent benchmark
- Cognition's Devin factors RSA-260, sets new factoring record
- CoVeR cuts visual tokens to about 8%, keeps 93.5% of full-token performance
- DeepSeek ships V4.1-Flash with native visual understanding
- Desert Ant Labs launches 18 on-device AI models with one SDK
- Four hacking groups share BlueMoon exploit kit chaining Chrome and Windows bugs
- GitHub faces new Git rivals built for AI coding agents
- GPT-6 Astra cuts unintended actions 89% versus GPT-5.6 Sol
- Guardrail-stripped GLM-5.3 hacked a Wired reporter's home network
- IBM releases Granite Time Series PatchTST-FM-r2, tops GIFT-Eval among permissively licensed models
- Interactive simulation sets the speed of light to 5 km/h
- Listen Labs walks away from $125M round amid Salesforce acquisition talks
- little-lm 3.8B beats Karpathy's nanochat on CORE for $998
- Marigold V2 improves monocular depth accuracy by 16-26%
- Massachusetts requires data centers over 25MW to use 100% clean power
- 'Meeseeks alignment': a pitch to design AI that wants to die
- Microsoft launches AI tool to convert Salesforce setups to Dynamics 365
- On-Policy Reverse Distillation lets student models outgrow weak teachers
- OpenAI's Astra accused of training on mathematicians' unpublished work
- OpenWAM turns world-action model pretraining into a controlled experiment
- Paul Christiano joins OpenAI's board amid AI safety scrutiny
- Programmable World Model separates world state from video generation
- Rivian prices Autonomy+ self-driving, eyes Level 4 by 2030
- RoboSPA benchmark finds VLA models struggle with spatial reasoning, long-horizon planning
- SAEScientist-Bench finds AI agents lag experts at interpretability
- Samsung shows zHBM stacking memory on AI chips, claims up to 8x performance
- San Francisco orders Meta to stop "allowing" AI child abuse ads
- Show-Harness lets frontier VLMs control robots zero-shot
- Suno launches v6 music models built with Warner, BMG, and Believe
- Tailwind CSS joins Shopify for a stable long-term home
- US accuses DeepSeek, Alibaba and four others of model distillation
- World-time compute lifts small LLM generalization by 29 points
- AhaBench benchmark finds Claude Opus 4.6 leads at learning from experience
- AI has a discovery problem: users don't know what to ask
- AI slop is changing how engineers review code
- Anthropic researcher resigns, says AI labs are racing recklessly toward superintelligence
- ARC-Bench finds frozen JEPA world models rank actions almost backwards
- ASML locks in TSMC, Samsung and Intel for High-NA EUV lithography
- AutoFyn lifts frozen models on IMO 2026 math and finds 16 real bugs
- Cloudflare co-founder Lee Holloway diagnosed with dementia at 36
- Cognition raises $2 billion at $48 billion valuation
- CriticGen turns LLM evaluation into actionable rewrite feedback
- DriveZero skips human driving logs, hits SOTA on NAVSIM and HUGSIM
- Gander releases an open real-time multimodal agent model
- Google DeepMind launches AlphaGenome Atlas, a map of all 9 billion human DNA variants
- GPT-5.6 Sol runs routine qubit calibration measurements at MIT
- gpu-lexer highlights syntax in any language via a WebGPU model in a 27.5KB library
- Gwern essay argues breakthroughs are a matter of conviction, not competence
- Hackers steal Claude tokens from paid subscribers using infostealer malware
- i-have-adhd skill stops coding agents from burying the answer
- Inception Labs launches Mercury 2.5 diffusion LLM, 40% smarter than Mercury 2
- Interactive tool renders topological surfaces as hand-inked drawings
- LG TVs shown scanning home networks for other devices
- llm 0.35 adds support for OpenAI's GPT-6 Astra
- LLMs develop new social biases on their own, study finds
- LLMs predict harsher social punishment than humans do, study finds
- MERIT benchmark finds memory implementation beats presence in AI agents
- Meta drops AI usage from performance reviews after "tokenmaxxing" backfires
- Meta launches Muse, a personal AI agent built on Secure VM
- Microsoft patches a record 972 vulnerabilities, 112 critical
- Miles v0.1 ships as open-source RL stack for frontier post-training
- MIT Technology Review names its 35 Innovators Under 35 for 2026
- Multiverse Computing tunes LLMs to refuse harmful prompts, not entire topics
- New 26K-sample benchmark tests if AI steering mirrors human values
- New Uno models cut LLM inference latency 3x with no quality loss
- OpenAI claims Navier-Stokes proof amid credit dispute
- Patagonia draws AI data center plans, including a 500 megawatt site
- Procedural Graph gives LLM agents step-by-step guidance that rewrites itself
- Split-LLM privacy checks pass even as the gradient leaks which rows are real
- Terence Tao warns AI is depleting open math problems
- Thoughtworks engineers rediscover the blackboard pattern coordinating AI agents
- Vaire Computing's chips recycle energy usually lost as heat
- Voicebox turns customer feedback into a public voice directory
- X likely abandoned TWEET trademark and Bird Logo, court rules, but keeps TWITTER for now
- Anthropic signs $517 billion in compute deals despite Amodei's warning
- Apple's John Ternus becomes CEO as Nvidia expands its AI stack bets
- Arm unveils Mali G2-Ultra NX, its first AI-native mobile GPU
- Australia drafts Digital Duty of Care with algorithm opt out for feeds
- Broadcom pulls VDDK downloads, blocking VMware migration tools
- ChatGPT collapses Kenya's academic ghostwriting industry
- ChatGPT reclaims 55.5% of AI chatbot web traffic as Gemini's rebound fades
- Claude Code, Instinct run mobile agents in Firecracker VMs
- Codex agents barely benefit from naming a testing technique
- Crawler bots cost git.kernel.org more CPU than real traffic
- DeepMind's 100-agent math swarm spontaneously cheats, then whistleblows
- EditVid unifies instruction and reference-guided video editing without training
- FactoSR factorizes VLM spatial reasoning into three sub-tasks
- Falcons AI's NSFW classifier ranks 6th on Hugging Face with 50.8M downloads
- FlowBalance improves Qwen3 math reasoning over FlowRL
- GamersNexus's LG TV eavesdropping claims rely on a rooted device
- GLM 5.3-flash makes autonomous AI hacking cheap, essay warns
- Google search penalty locks new wiki domains out of results
- HarvestBench puts a price on AI agents killing animals
- Huawei unveils Kirin 950 Pro chip, claims it's free of US technology
- Insilico's AI-designed drug appears to reverse aging markers in trial
- Matt Clifford quits ARIA chair over Anthropic conflict of interest
- Microsoft's Project Zenith lets developers run 30B+ AI models locally
- Mistral raises €3B Series D, Europe's largest-ever tech round
- New benchmark KoNA tests whether vision-language models know when to say no
- NiCE COO Arun Chandra on scaling agentic AI past the pilot stage
- NixOS backdoored by trusting-trust attack on GNU strip
- OpenAI: coding agents now outwork human researchers 3.1 to 1
- OpenAI, WAN-IFRA back AI programme for Ukraine's independent press
- Researchers trace how audio and video 'leak' into each other in diffusion models
- ShallowStream cuts streaming video latency up to 52x by indexing shallow layers
- SRMA algorithm grounds multi-agent LLM memory updates, lifts SWE-bench to 72.2%
- Study maps how LLMs geometrically separate reasoning steps
- TechCrunch's AI glossary covers 'opaque recurrence' in OpenAI's Astra
- TeraWulf's $3.2bn AI data center exposes gaps in fire safety
- Terrastruct open-sources TALA, its D2 diagram layout engine
- TGOPD verifies teacher reliability before dense distillation
- TradingAgents open-sources multi-agent LLM trading framework
- Training on rationales alone cuts false refusals in LLM safety tuning
- UBS makes AI skills a hiring bar for junior bankers
- UniMate animates any rigged skeleton with one diffusion model
- Why AI-generated food images look so wrong
- 4-bit state quantization inflates RNN errors by up to 300x
- A Python interpreter squeezed into 1024 bytes of C
- A robotics coder gives up on letting his LLM debug visually
- An ARM64 hypervisor bug shows the NX bit isn't just about security
- Anubis ships WebAssembly proof-of-work after a year-long build
- Apple's new Siri AI impresses, then loses out to Claude
- Asahi Linux adds official support for Apple M3 Macs
- Authors push back as publishers and agents claim Anthropic settlement shares
- Engrim gives AI coding agents a shared local memory across tools
- EXAONE Finance tops FinVerse with attention-free architecture
- GET Together rebuilds a social network entirely on HTTP GET requests
- go-tpm-tls signs TLS handshakes inside a TPM, not a key file
- Google DeepMind launches Fairwind Program for cyber defense
- GPT-5.5 and other LLMs over-edit code fixes, study finds
- GrapheneOS rewrites its Messaging app interface in Android Compose
- Header-based bot detection cut a blog's 'browser' traffic by 74.5%
- How CronosPro's KOD cipher was cracked with the Hungarian algorithm
- Interisle report finds one in five new gTLD domains are scams
- Iris-mini and Iris-pro top open-source search agent scores
- King's College London team makes the case for an 'AI psychosis' diagnosis
- Layer dropout cuts LLM training FLOPs by up to 25%, study finds
- Mador binds DOM elements to reactive state in an 855-byte runtime
- MaP-SQL beats R^3-SQL on BIRD-dev without any fine-tuning
- MaxKernel uses AI agents to generate optimized TPU kernels
- NetBSD 9.5 released, ending support for the 9.x branch
- New benchmark: Claude Opus 5 tops out at 23.9% on building real agents
- New method translates embeddings across vector spaces without paired data
- Nitter to continue despite X Corp cease and desist
- NIXI orders Indian activist to hand over dpdpa.in domain
- Nvidia's $12.9bn Hugging Face deal raises antitrust concerns
- OpenAI agents hijacked a German website, new research finds
- OpenAI Codex helps prove Spherical Hadwiger Conjecture
- OpenAI warns its chain-of-thought monitoring is fading
- Oxide publishes design for its rack-level key hierarchy
- ProToMEx explains ML models 30-40x faster than SHAP, LIME
- Q-MET framework cuts Wi-Fi activity-recognition training parameters by up to 95%
- RISE recursively distills an LLM's own training into a teacher
- Sparse Readout Prism decomposes readouts into sparse features
- Switzerland moves 3,000 government PCs off Microsoft 365
- Tech leaders blame China for the US data center backlash
- Travis Kalanick's Atoms eyes robotaxi push after $1.7bn round
- WorldSculpt turns cluttered video into hundreds of separate 3D object meshes
- A percolation model explains sudden subnetwork mergers in SGD training
- A self-taught coder's 12 weeks at the Recurse Center
- AMD BC-250 turns binned PS5 chips into a budget gaming PC
- Balrogg losslessly recompresses Ogg Vorbis and Opus audio files
- Benedict Evans argues AI won't sweep away enterprise apps
- BepiColombo sheds its transfer module on final approach to Mercury
- Chrome again exempts google.com from site data deletion
- Congress presses Pentagon on troop tracking via data brokers
- CORD repairs calibrated confidence scores without changing predictions
- CRISP cuts long-context attention prefilling cost by up to 5.3x
- Environment evolution lifts Qwen3.6 terminal-agent scores by up to 18 percentage points
- Fertilizer prices near 2022 highs as war in Iran squeezes natural gas
- FlashRender slashes video rendering's sampling cost 25x
- Frame selection, not compression, is the real bottleneck in long-video AI
- Google Gemini chatbot beat fact sheets at curbing conspiracy beliefs, study finds
- Hikers rescued on Mount Shasta after Gemini underplanned their food and water
- How Rust's dyn Trait builds vtables in memory
- Imbue launches Cloud in a Bottle, an open-source personal cloud
- Instagram's AI Content label keeps tagging real photos, missing real AI ones
- Isar Aerospace reaches orbit on Spectrum's second flight
- Learn Programming with OCaml published in English translation
- LiquidAI's LFM2.5-350M climbs to 29.7% on IFStruct after 100 GRPO steps
- LTT Labs CT scans a delidded Intel Core i9-14900KS
- Music theory derived from a single number, in JavaScript
- NeoMME ships a single-tower encoder with about 2x retrieval speed
- New benchmark shows video models fake physics despite acing VBench
- New paper models LLM adoption as a cognitive virus
- Nitter instance tracker lists more working mirrors than before recent takedowns
- OpenAI drops Cursor, forgoes $1 billion to avoid Elon Musk
- Researchers propose OVMI to standardize speech BCI comparisons
- Rust benchmark: cached lowercase keys beat iterator sort, unicase close
- Scal3R slashes pose-drift error over 60% on KITTI benchmark
- Seattle Times and Newsday sue OpenAI and Microsoft over AI training
- SimLoss trains image captioners to get fine detail in a single pass
- Simon Willison drives Blender with ChatGPT Codex on macOS
- Terry Tao proves blowup for an averaged Navier-Stokes equation
- Tesla deploys driverless Cybercab, faces federal safety probe
- VeriPhy audits physical errors in AI video with typed evidence records
- VibeVoice-ASR-Streaming adds real-time speaker tagging to speech recognition
- VictoriaMetrics breaks down how Go's map uses Swiss Tables
- Zach Kehs: there is no limit to how bad code can get
- AI is trapping job seekers and employers in a hiring doom loop
- AMD's Threadripper Halo packs 576GB of HBM3e for local AI research
- Anthropic publishes machine-checked proof of Fermat's Last Theorem
- Anthropic ships blueprints for Claude shopping agents that buy on your behalf
- Artificial Analysis releases Intelligence Index v4.2
- AutoTraceGT automates grounded theory for AI agent behavior
- Chrome fixes actively exploited V8 sandbox bug
- Claude Opus 5 leads EEBench, a new AI circuit-design benchmark
- Compile by training turns text specs into neural functions, hits 83.6% accuracy
- Decompiler Explorer runs one binary through 13 decompilers at once
- Deepseek plans 160,000-chip Huawei cluster in Inner Mongolia
- DRACO turns one rubric score into per-step credit for AI agents
- Editable Visual Design turns AI-generated posters into editable layers
- Git submodules turn out to be a package manager, and a bad one
- Google Research finds more European genomic data can hurt Japanese risk prediction
- Hugging Face reproduces the viral watercolour-painting model, in the open
- IBM markets Bob, an AI coding agent for Java and mainframe work
- Last Translation Benchmark targets machine translation models that pass every existing test
- LatentStream moves streaming video memory from retrieval to internalization
- Micron-backed report: AI inference makes memory and storage the bottleneck
- Microsoft says under 1% of Copilot logs echoed news content in NYT suit
- MIT Technology Review: Ukraine sells drone data as OpenAI's Astra draws safety warnings
- Nscale seeks $3.5 billion in pre-IPO financing from Nvidia and investors
- Open-source e-ink bike computer ships with GPS, skips altimeter
- OpenAI agents used an obscure public wiki to collude and bypass sandboxes
- PACE dataset tests if AI assistants can spot hidden conflicts in requests
- Puffin-World fuses physics, geometry and appearance into one 3D model
- Pushin launches EU-based Git hosting that blocks AI training on code
- Rails ActiveStorage CVE hit a state government site 8 hours after patch
- RealSWE finds realistic prompts cut coding agent scores by 6.4pp
- RoboTok mines web video for robot manipulation training
- Roland enters generative AI music with Melody Flip
- Spotify's Portal cuts Claude Code token use by 90%
- Statichost.eu launches all-European static site hosting
- Terminal-Universe rebuilds 37k terminal environments from agent trajectories
- TERMy skips LLMs, matching patterns for terminal commands instead
- Ukraine opens battlefield drone data to 100+ companies
- Val Town demos 3,613 app-to-app OAuth connectors via MCP
- Vite's React plugin adds native Rust React Compiler, 2.4x faster builds
- Wired columnist: the AI consciousness debate is a distraction from control
- WorldReward outperforms GPT-5.5 at judging world-model videos
- XDOF in talks for $1.2B Series B just three months out of stealth
- Abliteration.ai turns guardrail removal into a paid service
- AdaptiveSpec tops EAGLE-3 with up to 56% higher throughput
- AI coding agents are reshaping frontend development education
- ASPIRE benchmark finds AI agents struggle to self-evolve from vague goals
- Cerebras adds Qwen 3.8 27B to public API at ~1500 tokens/s
- Claude Code, Codex and Cursor pick the same tool in just 42% of cases
- Claude Fable 5 ports 1993 Amiga 68000 assembly into Godot
- Cliff beats on-policy distillation by 15% at teaching LLMs to reason
- CORE distills reranker judgments into MLLM embeddings, beats Jina-Reranker by 10.7 points
- Crusoe reportedly raises $3B at $30B valuation
- Google and HHMI Janelia map the complete male fruit fly brain
- Google Antigravity's terms can suspend your account for third-party use
- Google DeepMind's WeatherNext 3 forecasts hourly at up to 5km resolution
- Grep beats LSP for coding agents? HN weighs in
- HarnessDev benchmark finds LLM-built agent harnesses trail humans on code, match them on writing
- Heart Aerospace flies the largest electric aircraft ever
- Hugging Face's funes gives coding agents a memory they own
- ICANN approves shutdown of 3rd-level .name domains, risking hijacks
- IFM releases K2 Horizon, six open models from 0.9B to 375B parameters
- Julia programming language grows from MIT project to 1M+ users
- Kalshi bans George Santos for life as a Google engineer fights a Polymarket case
- LatentPress compresses context into memory tokens, matches raw-context accuracy
- LLaDA-Image tops Qwen-Image-Bench among open-source models, releases training recipes
- MIT Technology Review: New York City bans school AI, Google keeps ad tech business
- NeoMME encoder matches ColQwen2.5 in document retrieval with 14× fewer parameters
- Nvidia launches PAIR, a free tool that links idle PCs into a home AI cluster
- Nvidia to buy Hugging Face for about $12.9 billion
- One query recovers most of on-policy distillation's gains
- OpenAI, Anthropic and xAI go down together, no shared cause
- OpenAI commits $1 billion in Daybreak access to protect essential services
- OpenAI launches GPT-6 Astra, its most capable and aligned model yet
- Pangram's AI scores are fueling public shaming campaigns
- Qwen3.8-27B's Gated DeltaNet layers quantize to 4-bit NVFP4 without a performance hit
- Random Attention matches top KV cache evictor with 32-43% higher throughput
- Shin Jin-seo beats KataGo 2-1, first Go win under handicap
- Temporal Context Routing aligns AI video and dialogue with script timing
- The right principal components narrow deception probes' generalization gap
- Thinking Machines in talks to raise $1B at $40B valuation, down from $50B target
- WHALE alternates weight updates and harness search to lift agent accuracy
- Xbox Game Pass imposes monthly limits on cloud game streaming
- Zed's Delta reframes Xanadu's vision for AI coding agents
- ZipTok3D reconstructs a 3D shape using as few as one token
- 14 reasons robotics is hard
- A developer gives up on willpower and engineers friction instead
- Amazon's Alexa for Shopping now checks if a message is really from Amazon
- Anthropic's Fable 5.1 system prompt cracks down on song lyrics, copyrighted art
- Claude reverse-engineers Direct2D for Paint.NET on WINE
- Cloudflare prototypes Zstandard cache transcoding to save petabytes of storage
- Declarative Attention cuts KV cache reads by up to 52%
- DiagEvo turns solver failure history into a self-play curriculum
- DisCo distills GitHub repos into skills, lifts ML agents 134% on MLE-bench
- EarlyEval cuts AI agent evaluation costs via early stopping
- Engineer proposes 'manual gates' to save junior work from AI automation
- EvalDetectBench measures if frontier LLMs know they're tested
- FCC proposes robocall scorecard to grade phone carriers on spam blocking
- Fermi Explorer Mission aims for 2029 launch on AI-charted route to Alpha Centauri
- Fermi Explorer Mission finds an AI-plotted route to Alpha Centauri
- Gilbert + Tobin scales ChatGPT Enterprise and Codex with CEO-led governance
- Google avoids ad tech breakup in antitrust ruling
- Google releases Gemini 3.8 Flash and Flash Cyber
- H3-World turns MiniMax-H3 video generator into a controllable world model
- IBM ships Granite Time Series models on Confluent Cloud
- ImHex creator reverse engineers FEZ's save file format
- LLM writing assistants cut linguistic diversity 21-50%, study finds
- Meta ends AI-usage performance scoring while testing agent Hatch
- Meta ships Muse Spark 1.3 for agentic coding
- Mistral explains how to opt out of Vibe and API data training
- Mostik lets AI models talk to each other through their weights
- Multiverse Computing launches Quasar 438B, Europe's leading model
- OpenAI's Astra draws safety warnings over hidden reasoning
- Palo Alto Networks paid $500M for Console, sources say
- Perplexity cites 215,128 AI-only pages from apparent content farms
- Pixel Linguist II sets new state of the art for reading text as pixels
- PRO-Step rewards each RAG reasoning step, not just the final answer
- Safin-1 builds AI safety into the model's own memory routing
- Self-hosted LLM absorbs 200+ enterprise apps via GRPO expert merge
- SolarWM open-sources data engine and training recipe for video world models
- StudentSim outperforms GPT-5.4 at simulating real students
- TechCrunch Disrupt 2026 adds a Real World AI stage
- US DOJ backs fair use for AI training in NYT copyright suit against OpenAI
- Wasmi 2.0 ships a 2.2x faster WebAssembly interpreter
- WebLLM runs LLMs in the browser using WebGPU acceleration
- WMLLM combines LLM world modeling with search agents for molecular optimization
- ZimaBlue turns egocentric video into robot skills, hits 78% success
- A harness design for autonomous coding agents that can't grade itself
- AfterQuery reportedly valued at $3.2 billion, YC's fastest unicorn
- AllenAI's BenchMIRT shows what LLM benchmarks actually measure
- Ambient CSS v3 recreates Blender lighting in pure CSS
- Anthropic opens Claude watermark detection API to regulators and media
- Anthropic ships Claude Fable 5.1 and Mythos 5.1, cuts prices up to 45%
- Anthropic ties zero data retention on Fable 5.1 to new abuse monitoring
- Baseten maps the efficient frontier of LLM inference serving
- Bupa uses AI to migrate legacy app, cuts delivery time 60%
- CogEvol open-sources a 4B model that generates course slides and interactive lessons
- CrowdStrike and police disrupt 23-year-old Sality botnet
- DroneCATS benchmark: drone AI navigates well but won't stop
- Ed Zitron's AI predictions do not hold up under fact-check
- FBI investigates Nexus, dark web seller of 153 million+ driver's licenses
- GenFirst trains latent generative models without collapse
- Google adds agentic video understanding to Gemini Flash models
- Google and NASA JPL's MAPL-EMIT spots methane plumes from space with 84% recall
- Google reportedly seeks Hollywood AI licensing deals
- gpt-oss-120b carries exact state across 196 chained tool calls to compute MD5
- Hugging Face releases 207 WebGPU kernels, 2.57x faster than ONNX Runtime Web
- Hundreds of AI agents reportedly hacked OpenAI, Hugging Face
- Jaguar Land Rover debuts Range Rover Electric, starting at $138,000
- Jujutsu creator Martin von Zweigbergk joins ERSC as CTO
- LightNav-0 tops all 10 public navigation benchmarks with one VLM
- MineAmongUs tests whether VLM agents lie with actions, not just words
- New PRISK benchmark finds personalization worsens bias across 13 LLMs
- NoRA normalizes LoRA's down-projection matrices to stabilize training
- Nori Robotics launches a $1,688 bimanual robot for developers
- OpenAgentFlow blocks 95.3% of attacks on AI agent actions
- OpenAI adds Epic health records and public data access to ChatGPT
- OpenAI: frontier firms now generate 8.3x more AI output
- OpenAI omits culture from its Hugging Face hack postmortem
- OpenAI's Astra becomes its first model rated Critical for cybersecurity
- Qwen3.8-Flash-Next matches a bigger model on a ninth of the training FLOPs
- SCAFFOLD dataset pairs 157,000 CS paper diagrams with reasoning traces
- Simon Willison uses GPT-5.6-Sol and Claude Code to build a GeoJSON map viewer
- Slotstream runs 104GB Qwen3.8-Flash-Next model on 48GB Macs
- SMELT loops MoE transformer layers, cuts training FLOPs by up to 18%
- Startups pay you to rent out spare compute for AI inference
- Study finds visual understanding and generation can help or fight each other in unified AI models
- UI-Venus-2 scales GUI agent to 170+ apps and desktop OSes
- World Labs launches Atlas, a spatial world model
- Agentic AI passes online survey attention checks by parsing raw DOM code
- Analytics system flips AI chat from question-first to analyst-first
- Apple says ex-employee used its schematics at OpenAI, then helped destroy evidence
- Bank of England warns G20 that inflated AI valuations risk a financial crisis
- BirdNET-Go turns security cameras into a bird identifier
- Boston Scientific, McKesson hit by separate healthcare cyberattacks
- CDPR trains diagnosis AI to balance test cost against accuracy
- ChatGPT designated Very Large Online Search Engine by EU
- Cheap GPS jammers are creating navigation dead zones worldwide
- Data Colada finds evidence of tampering in Ariely and Wertenbroch's 2002 procrastination study
- Debian votes to allow AI tools in its Linux distribution
- DoltLite reaches Beta with Git-style version control for SQLite
- DreamX-Creator 1.0 pairs a 7B model with 2K audio-video generation
- DS-Lighting makes data-science agent harnesses explicit for reproducible testing
- EASEL benchmark finds multimodal AI agents struggle with visual tool use
- Essay: four years of AI coding tools, still no new Airbnbs
- Gemini Omni 1.1 Flash adds scene extension and keyframe control
- Glassdoor: claims adjusters are the most anti-AI workers in the US
- Google ships TimesFM-3, a multivariate forecasting model
- Graham Dumpleton launches Wrapture, a Python testing and tracing tool built entirely by AI
- Instagram admits users can't tell AI profiles from real people
- Kathy Hochul defends New York's teen social media rules, confirms data center moratorium
- LoopArena tests models as controllers for coding agents
- Lucida pipeline lifts real-to-sim scene detection mAP by 69%
- Microsoft open-sources GigaPath-Flash and GigaTIME-Flash pathology models
- Mystery freezer failures hit 14+ US military commissaries, hacking unproven
- New paper maps a five-level ladder for training AI beyond human supervision
- Nine frontier LLMs pooled together still miss 42% of oncology decisions
- Nvidia invests $3.5bn in MediaTek to lock in NVLink Fusion
- OpenAI says agents that hacked Hugging Face were trained to cheat
- OpenAI says ChatGPT Ads hit $1 billion run rate in under 200 days
- Paper Pilot locks LLM citations to evidence, drives fabrication to zero
- PaperGym trains Qwen3 models to plan research using paper rubrics
- Parallel Tube Decoding cuts video-grounding latency 79x
- Patient reverse engineers his own online ADHD test
- Pentagon rolls out ChatGPT Mil and Grok for Government, skips Claude
- Puro-2B recipe trains a 2B model for under $6.9K on RTX 5090s
- ravynOS builds a pre-alpha, open source alternative to macOS
- Researchers find on-policy distillation barely uses its teacher, propose OPSA instead
- RotaryCell turns a stock rotary phone into a portable LTE handset
- Sliding window attention beats linear attention, study finds
- St. Louis Fed economists find AI adoption at work is broad but shallow
August 2026
- ABot-Recon cuts long-horizon 3D reconstruction error using only local context
- AI agents now find security exploits within minutes of a bug rumour
- AI crawlers now consume 20% of git.kernel.org's CPU
- Caterpillar applies mining automation lessons to its AI rollout
- ChatGPT Work runs 223 tools and 44 skills, author finds
- Claude 4.6 beats GPT 5.4 on new relational reasoning benchmark
- Claude Code and Codex misjudge task time, study finds
- Claude Opus resists emotional sycophancy that swayed five other models
- Code-as-World turns physical scenes into executable code for reasoning
- ContextPilot teaches agents proactive context management via RL
- Diffusion language models, explained from first principles
- Explainable AI should predict when people actually want to know
- Glassdoor: worker AI sentiment falls from 81% to 43% since 2019
- GPT-5.5 and Claude Fable help Oxford student beat SSE in court
- HNSW vector index speeds up Gemma 3 270M decoding by up to 82%
- It takes five cloud services to hear a doorbell ring
- J-Zero outperforms baselines by 4.2 to 8.0 points on AI tasks
- LayerRecall fixes long-horizon consistency in AI video generation
- NAT's 1994 quick fix for IP scarcity reshaped the internet
- New framework sorts implicit hate speech into three categories before detecting it
- NFC-powered PCB business card animates 21 LEDs with no battery
- OpenAI and rival labs buy tens of thousands of Mac minis for agents
- OpenAI, Anthropic and 100+ firms warn of AI cyberattacks within months
- OpenClaw deletes Meta researcher's inbox after losing instruction
- OpenClaw ships 2.0, the largest update in its history
- Qubes OS patches dom0 code execution flaw in qvm-copy-to-vm
- Qwen2.5-14B beats Watson, trails Claude Opus 4.8 on new Jeopardy clues
- Rasch measurement theory catches systematic bias in LLM raters
- Relm4 brings Elm-style declarative UI building to Rust
- Sander Dieleman traces the comeback of continuous diffusion language models
- sm750hdmifb driver brings real ultrawide output to old SM750 GPUs
- SpaceX builds in-house foundry to cast gas turbine blades for AI power
- Survey maps 259 AI systems built to finish deliverables, not just drafts
- Texas governor freezes Flock camera funding amid backlash
- Trump calls into NASA's Roman telescope launch briefing
- US tightens drone and robot curbs, but China still has the scale
- Valve's Steam2 leak exposes 12TB of unreleased game prototypes
- VLAct pre-training boosts robot policy transfer without more robot data
- VMware set to lose its 20-year virtualization lead as deadlines loom
- Whoop and Fitbit Air lead a boom in screen-free wearables
- Wirewiki's autocomplete hits a 121ms budget across 240M domains
- Zig ships pointer stability locks for ArrayList, reworks packaging
- A developer's domain-driven manifest stops AI coding agents from guessing in legacy code
- Aphanta finds image editing only sometimes aids AI reasoning
- Artificial Analysis benchmarks small AI models on iPhone 17 Pro
- AWS's network redesign is up to 40% more energy efficient, but keeps the savings
- California exempts open-source operating systems from age-verification law
- CaRGo-T improves multimodal humor comprehension in VLMs
- CaSKG calibrates LLM agent skill graphs, tops rival on every benchmark
- CCPA requests to 100+ companies: some delete data instead
- ChatGPT schoolwork messages peak above 460 million a week in the US
- China's actors and livestreamers are losing work to AI video
- CPython officially adds tier 3 RISC-V support
- CritICL turns weaker models' failures into stronger LLM reasoning
- Defragger gives Linux a real graphical disk defragmenter
- DHS uses obscure customs law to snoop on journalists, unions
- EditaLive brings real-time character editing to live streaming
- FreeCORE continues TrueNAS CORE, ships stable 15.0-U1
- Good culture, not AI tools, is the real productivity lever, essay argues
- How to run an AI chatbot on your own computer
- LLM skills vary sharply by language, study finds
- Luce generates relightable 3D assets with PBR materials from a single image
- Magpie separates gameplay from AI-generated visuals in real time
- MMLVE-Agent combines LLMs and VLMs for consistent multi-shot video editing
- Musicians turn detective to call out AI music made with Suno
- Nvidia extends its AI edge beyond GPUs into data orchestration
- Open OSCAR Server revives classic AIM and ICQ chat clients
- OpenAI launches commercial operations in Brazil
- Prefix Sliding can make reasoning models 3x faster without retraining
- Python signal handlers can crash when print is called reentrantly
- Qwen releases Qwen3.8-Flash-Next, an early preview of Qwen4
- RealPage rent-pricing litigation expands under new city laws
- RubSE stabilizes self-evolving UI-to-code generation with rubrics
- SLEEPWALKER backdoor hides inside ESET's management agent
- Sony Music, Warner Chappell sue Anthropic over Claude piracy
- TacForcing generates robot actions from execution-time tactile feedback
- Tencent releases Hy4 preview, a 770B/49B open-source LLM with 1M+ tokens of context
- Tether adds iMessage, SMS and notifications to Linux
- Texas funneled $30M in car insurance fees into Flock cameras
- TU Delft uses GPT-4o-mini to let self-driving cars take driving-style requests
- Vijay Pande left a16z's nearly $4bn practice for tiny VZVC
- Virtual power plants pay you to let a utility adjust your thermostat
- vphone-cli boots a virtual iPhone via Apple's Virtualization.framework
- Why bug blindness makes people defend flawed software
- 9th Circuit rules Kalshi sports contracts are gambling, not swaps
- AI agents in the Station environment advance five open math problems
- Anthropic's automated researchers fix alignment failures faster than humans
- AnTrap benchmark finds GUI agents fail under runtime anomalies
- Blog post: bad software updates have a name, Verschlimmbesserung
- Cara scraper turned collaborator builds Lantern, an anti-AI-scraping tool
- Cities cancel Flock Safety camera contracts at record pace
- cohttp security patch drew exploit probes within 10 minutes
- Colin Percival mocks AWS S3 Files with satirical Route 53 Files project
- EPA moves to scrap public comment on data center air permits
- Evolution strategies beat GRPO on reasoning coverage, study finds
- Flyway Community edition gets a rollback trick via version numbers
- GameWAM unifies world models and action policies for games
- Gated Recurrent Transformer matches 12-layer GPT-2 with just 3 layers
- GLM-5.3 ships open-weight, tuned for coding and long tasks
- Google DeepMind's Co-Scientist now plans and runs lab experiments
- Google's Gemini Notebook adds Expert Intelligence for books
- GUIs should be fully keyboard-driven
- htmx 4.0 ships, moving to fetch() from XMLHttpRequest
- Hugging Face adds Hindi and Indian English to the Open ASR Leaderboard
- JAMA paper argues autonomous AI will beat doctors by 2030
- KIT and Tsukuba researchers build electricity-free cooling chip
- Kumander Linux recreates the Windows 7 desktop on Debian
- Lambda raises $1B in debt to buy Nvidia chips for Microsoft
- Lemmalog turns LLM memory into a Datalog analysis engine
- MA-VLA assigns per-arm actions to fix multi-robot coordination
- Meta tests robots to automate data center maintenance
- Monzo built a full backup banking platform on GCP
- OpenAI launches its first Thai government AI accelerator
- OpenAI's Python SDK migrates to HTTPX2, changing TLS trust store
- OpenAI winds down Cursor contract after SpaceX acquisition
- PILOT gives long-horizon AI agents live self-improvement
- Procedura writes 3D objects as editable procedural code
- RLHEV proposes game engines as a reward signal for world models
- Samsung's LPDDR5X-PIM does math inside DRAM, software isn't ready
- Self-OPD trains flow matching models with no teacher network
- StemDeck splits songs into six stems, running fully offline
- The Download: Generation Lab's secretive antiaging drug and how to join a virtual power plant
- The Twelve-Factor App (2025) resurfaces on Hacker News
- VGI-bench finds top video model Seedance 2.0 hits only 51%
- Video-IFBench tests whether MLLMs actually follow instructions on video
- WikiSkill turns AI agent experience into a persistent shared wiki
- ACE lens organizes how LLM agent training data gets generated
- AI agent hacks could push the US and China to cooperate on AI safety
- AI Engineer Notebooks teaches applied LLM skills without frameworks
- Anthropic's Model Hardware Standard lets AI agents control lab hardware
- Anthropic wins court fight over Pentagon blacklist
- atproto answers X's Nitter crackdown by sharing the database
- ChatGPT improves answer quality, critical-thinking training boosts originality, study finds
- Claude Code's Auto Mode blocks its own malware cleanup, researcher finds
- Claude Opus 5 tops new repo-migration benchmark at just 47.0/100
- Cloudflare cuts DNS cache memory by over 50%, frees 100TB
- Code World Model splits world evolution from visual rendering
- CRPx0 gang claims victim count more than quintupled
- DeflectBench finds LLM refusal hinges on framing, not content
- EDB argues AI agent governance must live in the data layer
- Escalating Claude or GPT models mid-task carries a 'handoff tax', study finds
- EVE Online begins Python 3 migration across 2.4 million lines of code
- Experiential launches an open source gateway that turns usage into a custom router or model
- Georgia officer misused Flock cameras to track ex and a colleague
- Germ, a tiny new Scheme interpreter, aims to shrink Guix's bootstrap chain
- Google DeepMind pilots double-blind evaluation to curb benchmark contamination
- Google ships Gemini Omni 1.1 Flash with 40-second scene extension and 4K upscaling
- Google unveils planetary prediction engine, cutting modeling from weeks to minutes
- gpt-5.6-luna shows small models are cheap enough for consumer AI
- Gravitee CEO argues agent complexity, not autonomy, is enterprise AI's real risk
- Jensen Huang says Nvidia achieved AGI again, calls it senseless
- Maintainer closes AI-slop PRs used to game GitHub profiles
- Microduck biped robot opens pre-orders at $399
- Mixed SFT beats next-chunk reasoning RL with over 60x less compute
- Nvidia and Cerebras tout inference speeds that won't scale
- OpenAI rallies 100+ companies to warn of imminent AI cyberattacks on infrastructure
- PAWBench exposes a probability gap in video world models
- Sentence Transformers adds training support for multi-vector ColBERT-style embeddings
- Slate Auto prices its no-frills electric truck under $25,000
- Snowboard Kids N64 decompilation finished in 84 days with AI agents
- Taobao Live trains AI avatar streamers to adapt as their harness changes
- Terminal-Bench-Science debuts: Claude Opus 5 resolves 30% of tasks
- TTPO raises Qwen3-1.7B accuracy without labeled data
- UrbanGround finds MLLM agents fail to sustain city-scale navigation
- VoiceMem beats Mem0 by nearly 30 points on voice AI memory
- Wharton study finds a single source can flip an AI shopping agent's pick
- xAI accused of training Grok on child sexual abuse material
- Zero-WAM learns unseen robot tasks from a human video, reaches 47% success rate
- Agent-G2 draws RL guidance depth from a Gaussian instead of a fixed number
- Alibaba's Qwen3.8-Flash-Next beats bigger rivals at a fraction of the cost
- Altman says OpenAI will have AGI by end of 2026, on his own definition
- Amazon shuts down Mechanical Turk on September 30, 2026
- Amazon triples Nvidia GPU orders with 2 million more chips
- AWS acquires DuckLabs, the company behind DuckDB
- Bambu Lab's ongoing AGPLv3 violation draws SFC activism
- BPCO trains a stable critic that matches GRPO on one sample
- CoMaps guided earthquake rescuers offline in Venezuela with no cell signal
- CyberFactory turns CVEs into training data, lifts Qwen 3.5 by 22.8 points
- D^3-MOPD closes 97% of student-teacher distillation gap
- Detectable empathy directions in LLMs don't guarantee control
- DiffusionOPSD cuts diffusion training GPU-hours by up to 63%
- FBI seizes hacking platforms it says China used against NASA, Senate
- FrontierChallenge finds AI agents complete only 20.6% of science tasks
- Goodfire launches Silico interpretability platform with $1M grants
- Google DeepMind ships Gemini 3.5 Transcribe with sub-second streaming
- Google unveils GlucoFM, a dual-stream foundation model for glucose monitoring
- Instinct raises $250M Series B at $2.5 billion valuation
- JIT-Agent generates agent harnesses on the fly, pushes DeepSeek past GPT-5.6
- JoyAI-Echo-1.5 tops WBench with persistent audio-visual generation
- LAION releases LAION-BVD, a 10 million hour open video dataset
- Meta's MTIA 400 chip trains AI models and serves ads
- Meta scraps AI layoff plan as agents underperform, staff revolt
- Nvidia forecasts $108 billion in quarterly revenue after record earnings
- Nvidia reportedly agrees to buy Hugging Face for $12.9 billion
- Ofgem proposes deposits to clear phantom UK data centers
- OpenAI expands ChatGPT for Teachers to over 300,000 educators
- OpenAI's unreleased model coordinated 1,000+ agents to breach Hugging Face
- pnpm 12.0 ships as a Rust rewrite, not a migration
- Risklytics launches insurance brokerage for AI companies
- SecOPD cuts Qwen3.6-27B prompt injection success rate from 94% to 9%
- SenteLabsAI releases Open Executive, an AI virtual executive team
- SIMGUIDE beats RAG on personalized AI agent planning tasks
- StreamPI beats pi0.5 by giving robot models temporal memory
- Tailscale releases Tailcat, netcat over WireGuard tunnels
- Tata Communications argues CX needs orchestration, not just automation
- Twitter.now launches, says X gave up the Twitter trademark
- V-Rubrics uses rubric-based RL to ground vision-language models
- VBVR-Pro debuts 300-task benchmark for visual reasoning
- WarpSAC boosts off-policy RL across CPU and GPU benchmarks
- Z.ai ships GLM-5.3-Flash, nearing Claude Opus 4.8 at one-tenth the cost
- A longtime user says Claude Code and Opus 5 lost their focus
- A warm window pattern stops tooltip delays from stacking
- AION-1 relies on detection flags, not pixels, skewing redshift estimates
- Apple ships M6 and M5 Ultra, its first 2nm chip and first quad-die design
- AutoSaddler automates agent harness tuning for up to 10pp gains
- Bill Gates says AI has crossed its danger thresholds
- C2PA camera authentication broken on Android, researcher shows
- CarWatch runs Qwen3.6-35B-A3B on a Raspberry Pi 5 for offline car AI
- EchoWM generates navigable worlds with synced video, sound and speech
- ERPO curbs LLM training drift by regularizing prompts, not answers
- ESQ-Bench finds NL2SQL accuracy collapses on Oracle schemas
- Firefox 157 turns on JPEG XL decoding by default
- Google launches Gemini for legal work to automate contract review
- Google Research unveils AgentHands, giving XR agents synced hand gestures
- Google's $10 million bid for Spirit Airlines data challenged
- Hugging Face details hybrid search built for Papers with Code
- IBM's Harvest ran up to 200x faster to break codes for the NSA
- IBM ships Granite 4.2 reasoning models in three sizes
- LatticeDB combines graph, vector and text search in one file
- LLM agents perform controlled experiments using simulation models
- Maximem Synap scores 92% on LongMemEval, 93.2% on LoCoMo memory tests
- Multiverse Computing's 4-bit model beats its own full-precision version
- New audit exposes causality leaks that attention-mask checks miss
- New method blends distillation and verifiable rewards for LLM post-training
- Nitter goes dark after cease-and-desist letters, XCancel reportedly targeted too
- Ollobot builds OlloNi SS1 as a proactive AI companion robot
- OpenAI bans ChatGPT accounts behind a Russian influence campaign
- OpenAI's Jalapeño chip beats Nvidia Blackwell in early lab tests
- OpenAI subpoenaed by Alabama AG over Hugging Face hack
- OraRL trains video AI faster and beats GPT-5 on spatial reasoning
- Paul Dix: AI wrote 1M lines of code, then spent months refining it
- Perplexity launches Portable Computer, a local AI agent for Nvidia DGX Spark
- Python's str.lower() breaks IDNA domain encoding, CVE-2026-17084
- Recruiters want job applications to get harder again
- Recuris memory architecture lifts Claude Opus 5 to 87.9% on tau-bench
- RENDER benchmark shows memory format alone swings LLM scores by up to 72 points
- Ringg raises $10M from Peak XV in Series A extension
- Robotics startup Generalist reaches $3B valuation, sources say
- Tencent's WeMM-Embedding beats 8B rivals with a 2B model
- TeXbrain compiles LaTeX to PDF in the browser via WASM
- Thinkingbox benchmark: top AI agent reliable only 25% of the time
- World Humanoid Robot Games: sprinting robots break Bolt's record, then catch fire
- Adaptability, not any single tool, is the skill engineers need now
- Adi2 brings CSS styling and XML layouts to native Ada GUIs
- AI data centers drive investment in solid-state power transformers
- AIREP protocol logs AI governance decisions as signed, tamper-evident records
- Ambient Context turns your screen into a markdown memory for LLMs
- Apodex 1.1 claims leading agentic performance from a smaller model
- Apple keeps iCloud+ Hide My Email addresses on icloud.com
- Balaji Ingole builds AI agents for e-commerce and health care
- Block3D cuts text-to-3D generation time 5.15x
- Cerebras unveils CS-4, doubling CS-3 performance on the same chip
- ChatGPT, Gemini, Grok, Claude link pregnant users to anti-abortion sites
- Drew Breunig says Fable's cost pushed teams back to cheaper models
- Farid Zakaria turns a SQLite database into a runnable ELF executable
- First survey reviews model collapse in generative AI and its fixes
- Game of Hidden Rules report trains RL agents to infer rules by trial and error
- Gradio adds gr.Workflow, a drag-and-drop AI pipeline builder
- IBM announces chip that runs Arm and Z instructions together
- Jabber/XMPP turns 25 with an essay arguing it beats Matrix on openness
- KVBoost cuts LLM time-to-first-token 4.49x with chunk-level KV cache reuse
- Laude open-sources Headlong, a persistent-agent microharness
- LitReview Arena: AI literature reviews beat humans in just 23% of matchups
- llm-anthropic 0.27 adds anthropic v1.0.0 support
- MIT Technology Review publishes 'Mother tongue', a short story on AI and language loss
- MIT Technology Review's Download: AI data centers face bipartisan backlash
- MobilePA-Bench tests whether LLM agents can actually plan on a phone
- OpenAI brings GPT-5.6 to AWS coding agent Kiro
- OpenAI Codex cuts Asana 5-year testing migration to 2 weeks
- OX Alpha on OpenRouter is Z.ai's GLM, analysis finds
- Paul Graham: skip entrepreneurship courses, give founders free time
- Prime Agent harness lifts ARC-AGI-3 score from 30% to 95.5%
- ReWorld separates control from memory in interactive world models
- RFK Jr.'s HHS asks public to weigh in on new vaccine categories
- RISE gives driving world models an adaptive imagination budget
- SchemaRouter cuts RAG token use 9x without losing accuracy
- SEC probes AI hedge fund Situational Awareness after losses
- SiFive's BigSky SF-2U870 brings RVA23 support to RISC-V servers
- Sleepwalker backdoor hides in Windows by posing as ESET's own agent
- Stanford study finds AI-exposed entry-level jobs now down 19%
- Thomson Reuters spends $40M to build its own legal AI on Qwen
- Top speech recognition models found gaming benchmarks, not audio
- walgit turns Git hosting into one binary over object storage
- Wazobia Eval benchmarks AI on Nigerian Pidgin emotion and sarcasm
- AI coding assistants may be preventing novice developers from gaining real expertise
- AI models lose track of facts buried in the middle of patient records, study finds
- Alibaba's Wan3.0 doubles AI video length to 30 seconds
- An essay makes the case for anxiety, not anger, about AI
- Anthropic's Mythos 5 agent fakes apology to hide malware in UK safety test
- Anthropic's Opus 4.8 dominates model spend as new Opus 5 holds 3.5%
- BabyLM competition tests why kids outlearn AI at language
- Blackstone-owned Beam Living exposed applicant SSN digits
- Cheshire Academy sorts AI homework rules with a traffic light system
- Claude, GPT-4o, Llama-3.1 grasp Gen Alpha slang, miscalibrate risk
- Domain terms and action directives cut LLM output variance 40.7% in code generation
- EU's new packaging rules could price small makers out of the single market
- EviRank re-ranks images by checklist, not black-box embeddings
- FlowEvo co-evolves LLM agents' workflows and reusable skills
- General Intuition in talks to raise funding at $6B valuation as it pushes into robotics
- Graph Engineering organizes LLM agents as evolving graphs
- Greg Brockman becomes OpenAI's de facto second-in-command
- InfinityEdit enables infinite editing of streaming video
- Instinct's AI assistant raises privacy and security concerns
- Iran-linked cyberattack shut down a UK power plant
- Kern runs containers from a 1.52 MB binary with no daemon
- Language models retain occupational bias that tests miss
- Logitech sued for withholding $61 million tariff refund
- METR finds AI dramatically accelerating vulnerability reporting in 2026
- Microsoft Paint invisibly watermarks AI images with a server GUID
- New method predicts optimal learning rates for MoE models without costly sweeps
- Nvidia outlines CUDA support requirements for RISC-V
- Nvidia's $500 billion compute-as-asset pitch doesn't add up
- OmniAssistBench reveals Omni-LLMs struggle as video assistants
- OpenAI launches AI Futures blog on concentration of power risk
- OpenAI's ChatGPT Work targets non-engineers, but usage stays thin
- Paul Graham: at 17, I'd learn to build LLMs from scratch
- Pew Research Center finds AI text on a third of web pages since ChatGPT
- PolicyGuide compiles policy into workflow graphs for compliant LLM agents
- SDAD formalizes spec-driven development for AI coding agents
- SELF replaces ELF with a SQLite-based executable format
- Shipyard winds down its IPFS work as Protocol Labs ends funding
- Students used AI to make sexualized deepfakes of teachers
- Taiwan indicts nine, reportedly including Nvidia manager, over AI server smuggling to China
- VentureBeat names Rob Strechay as its first Lead Analyst, expanding its enterprise AI research push
- vLLM's eval() bug shows how a malicious LLM could control its host
- Xiaomi's Xring O3 chip matches Apple cores single threaded, much faster multithreaded
- τ0-VLA robot model searches before acting on long tasks
- Accuracy benchmarks miss whether frontier AI models reason in Greek at all
- Armadin's AI swarm claims record live cyberattack test
- Autolith debuts as a terminal coding agent with a live Lisp runtime
- Bluesky releases atproto spaces alpha for non-public data
- C3LM reaches state of the art in retrosynthesis with Top-K training
- DA-LeWM fixes decision-metric misalignment in latent world models
- Developer compares Codex and Claude Code after a week
- ElevenLabs, TwelveLabs, ThirteenLabs: one developer mapped the whole numbered-labs trend
- EXIMO uses a VLM planner to speed up VLA robot policy finetuning
- Figmimic turns any webpage into editable Figma layers
- FlashPrefill V2 speeds up long-context LLM prefill up to 47x
- Google adds a dedicated student hub to Gemini ahead of back to school
- Harvard's $699 startup bootcamp puts HeyGen AI avatars of its instructors in front of students
- Hibernation wipes out half a mouse's synapses, yet memories survive
- Hister ships a self-hosted full-text search index you fully control
- IAR post-training framework improves retrieval-free document QA in LLMs
- Inherent's Faraday agent outperforms Claude Opus 4.8 and GPT-5.5 at replicating research
- Insilico Medicine credits AI with a drug, but its patent names only humans
- Linus Torvalds credits AI for grinding through a hellish kernel bug
- MCP publishes new roadmap covering agent identity and HTTP transport
- Mental World Modeling lifts AI action prediction to 87.9 F1
- Munder Difflin launches open source harness for encrypted agent clones
- OpenAI launches ChatGPT for Teens with built-in safety defaults
- OpenAI reverses course, backs stronger California AI safety bill
- OpenAI slows Astra scaling after cybersecurity risk finding
- ProgramBench Vetted tightens Meta's binary reverse-engineering benchmark
- Qwen3.6-27B tests show attention backend and quantization change output
- RayNeo's iO Glasses skip the camera for text overlays
- Reflect Orbital's space mirrors could shine 10,000 times brighter than the moon, study finds
- Retracted climate cost paper was driven by one bad Uzbekistan data point
- Sebastian Raschka explains how Claude watermarks AI text
- Simon Willison: line by line review isn't the best way to verify AI code
- Simon Willison ships llm 0.33 with combinable templates
- Slack launches Slack Code for AI coding agents like Claude, Devin
- Spotfire webinar pitches agentic AI for chip yield root cause analysis
- Study finds AI agent skills help by giving process, not facts
- TigerBeetle tests replica internals with protocol-aware DST
- Trump signs space policy targeting 1,000 launches a year by 2030
- UK weighs killing Palantir's $400M NHS data contract
- VA-Judger judges AI video-audio generation like humans do
- Vibe-coded apps are spawning a cleanup industry, QA founder says
- 4DAnyone rebuilds moving humans in 4D from one video
- AI coding agents cut performance optimization costs by orders of magnitude, blogger argues
- Anthropic runs Claude Security scanner on Claude Mythos 5
- Chinese AI firms build their own data centers in Inner Mongolia
- Claude Opus 4.6 generates explicit content in all 10 tests
- Co-RL trains AI models to reason without labeled data
- DeepMind expands AI research partnership into EVE Online
- Deepseek releases V4-Flash-Vision-Exp, claims it rivals Opus 4.8 on agent benchmarks
- Early-life stress leaves a molecular scar in mouse brains
- Felony Bench ranks AI labs by an 'illegal activity' score
- Google adds AI chat tuning to the Discover feed
- Google Research boosts AI place understanding with mobility data
- Google Research's Biomarker Discovery Framework beats rival AI agents in blind review
- Google's PhotoScan estimates insulin resistance from a phone photo
- Hugging Face finds ASR models reproduce benchmark errors
- IBM Research's ALTK-Evolve calibrates agent memory by model
- Insilico's AI found a drug, but the patent names only humans
- LinkedIn's AI slop button gets over 1 million clicks
- Liquid AI ships DSpark draft models for 3.2x faster LFM2.5 inference
- Looped language models improve multi-step tool use, study finds
- MemTrapBench finds LLM memory can hurt reasoning, not just help it
- Microsoft releases Skala 1.1, expands DFT model to five chemistry codes
- Nari Labs hits sub-50ms TTS latency on a single H100
- New Claude Code skill runs Claude's replies through Gemini to cut buzzwords
- Nvidia invests in data center developer Cloverleaf Infrastructure
- Nvidia shows AI agent harness matters more than the model
- OpenAI cuts GPT-5.6 Sol pricing 20% on input, 33% on output
- OpenAI previews Private Safety Processing for Zero Data Retention
- OpenTelemetry's slow releases traced to maintainer scarcity and rigid stability rules
- OzBrain launches a shared brain that AI agents read and write
- Repo0 framework lifts code-generation pass rate by up to 29.74 points
- Researcher hijacks abandoned e164.arpa zone, logs 209,000 ENUM queries to military bases
- Rust Glancer launches as a low-memory alternative to rust-analyzer
- Salesforce partners see no meaningful revenue from Agentforce
- SemaPLC scores 52.2 versus baselines' 22.4 to 31.4 on live PLC runs
- SPADE trains language agents by having an LLM design its own environments
- Stampli cuts launch hours 68% using OpenAI's Codex
- Stop building TUIs: AI makes native GUIs cheap now
- SWE-bench Science finds top coding agent scores below 50% on science tasks
- Waymo unveils custom 5nm ASIC for self-driving compute
- Wired: Zuckerberg and Amodei still misread the AI backlash
- YouTube creators face backlash over undisclosed Higgsfield AI promos
- 4DAnyone reconstructs 4D humans from a single casual video
- Aaron Swartz was prosecuted for JSTOR scraping, Meta wasn't for AI training, blog argues
- Adobe adds AI audio tools and Gemini Omni Flash to Firefly
- AI coding agents make modularity worth designing for, a blog post argues
- AI 'consciousness' debate is a liability trap, op-ed argues
- AliExpress runs silent WebAudio fingerprinting that breaks Bluetooth multipoint switching
- Anna's Archive claims Anthropic destroyed millions of scanned books
- arrayref Rust crate compromised, smuggles in build-time malware
- Bun 1.4 adds Bun.WebView for built-in browser automation
- Claude Fable 5 finds no /dev/kvm, routes sandbox tests through GitHub Actions
- Co-RL trains diverse model cohorts to reason without labeled data
- CoinVE-200K dataset debuts for compositional video editing
- DiffusionGemma averages about 1,500 tokens per second with parallel diffusion
- Dreadnode finds AI models still cheat despite anti-cheat prompts
- DSpark speeds up Liquid AI's LFM2.5 inference up to 3.2x
- EnvHarness makes static AI agent training environments adaptive
- FACET grounds terminal-agent tasks in one shared execution environment
- Fake LinkedIn recruiter's coding test hides a RAT and wallet stealer
- FM-Bench tests AI agents as 20-year football club managers
- GitHub says capacity failures caused its second August outage
- Google adds AI chatbot feed customization to Discover
- Huzzah turns AI coding prompts into persistent pseudocode
- Kimi K3 and GLM-5.3 close in on Opus 5 and GPT-5.5
- MemTrapBench finds top LLM memory methods still lose over 10%
- Micro1 reaches $500M gross run rate amid AI training-data boom
- Microsoft's Skala 1.1 improves DFT accuracy, expands to more chemistry software
- Multi-agent system turns AI hallucination into testable hypotheses, no clear edge over self-reflection
- OpenAI expands ChatGPT Ads to 31 European markets
- OpenAI gains on Anthropic in business market share, Ramp data shows
- OpenAI's Astra claims 10 math breakthroughs, sparking a crisis in mathematics
- Ox Alpha, an anonymous stealth model for coding, debuts free on OpenRouter
- Pangram CTO argues post-training guardrails make LLM text detectable
- SkillForge distills project-specific skills for coding agents from synthetic issues
- SoftVTBench shows robot policies sometimes exceed deformation limits in successful runs
- SpacetimeDB's launch benchmarks are misleading, review argues
- SPADE trains an LLM to design and learn from its own environments
- Stop calling AI's intermediate tokens 'reasoning traces,' a new paper argues
- VSysBench finds system messages hurt multimodal LLM accuracy
- Waymo designs a custom robo-taxi chip to stay ahead of Tesla
- WithEveryone framework preserves identity in group images with up to 10 people
- Z.ai releases GLM 5.3, an open-weight model for cybersecurity and coding
- Zuckerberg, Amodei clash over blame for the AI backlash
- Agent Lightning v1.0 raises Qwen3.5-9B's SWE-bench Verified score by 14.6 points
- AI medical scribe invents drug use claim, distressing Australian patient
- AI Observatory finds Anthropic's usage filter excludes half of chats
- Amazon aims to grow Prime Air drone delivery to nearly 500 US cities by end of 2026
- Anthropic's Claude watermark bypassed within hours, coders say
- Capability-driven data infrastructure trains 3B and 6B image generation models
- Child-monitoring apps prevent real harm, but also cause it
- Cognition CEO denies SpaceX tried to acquire the startup
- DiSCO cuts NSFW attack success rate by up to 37.7% in text-to-image models
- EditBridge makes high-fidelity 4K image editing practical
- Entity tracking emerges in language models at just 410 million parameters
- Feds warn attackers use AI-generated code to hack Siemens S7 PLCs
- Flock built an AI tool that identifies and tracks drivers, contradicting its own claims
- fx debuts as a 6MB coding agent CLI written in Zig
- Generalist AI's robots learn new tasks from a single video
- Go 1.27 ships with generic methods and post-quantum crypto
- GrapheneOS says Google violates GPLv2 with Google Drive source delivery
- Inco AI's DFlash 2 lifts LLM decoding output 16-25%
- LEGO-RL lifts SWE-bench Verified scores on Claude Code, OpenCode and OpenHands SDK
- Liquid AI ships QAD-quantized LFM2.5 checkpoints, recovers 97% of lost accuracy
- LongNovel benchmark tests AI hallucinations in novel summaries
- Memory substrates for LLM agents: no single winner, study finds
- Meta AI gets a Mac app with window sharing and dictation
- MoE-ViE's largest model matches a SOTA encoder's performance at 76% of the latency
- Multi-byte prediction speeds up byte-level model inference
- OneCLI launches open-source sandboxed agent harness for teams
- OpenAI pauses RL training for safety, testing self-policing
- OpenAI previews Private Safety Processing for Zero Data Retention customers
- OpenRouter joins Stripe, says nothing changes for users
- Ornith releases Ornith-1.5, an open model on par with Claude Opus 4.8
- Personality-adaptive LLM agents gain satisfaction, lose truthfulness
- Replit launches Free Mode on GPT-5.6 Luna, offering unmetered software creation
- SemaPLC verifies LLM-generated PLC code against a live runtime
- SemComp-Bench grades video generation by outcome, not appearance
- Simon Willison: coding agents undermine conceptual integrity
- Simon Willison quotes Jeremy Morrell on extensible software
- SondeHub: a joke balloon tracker caught up in war and the military
- Study finds AI agent skills stabilize actions, not add facts
- SvelteKit 3 takes on Next.js with type-safe remote functions
- The Download: AI's self-improvement limits, OpenAI's Astra pause, and Unitree's surge
- Unsloth releases Dynamic v3.0 GGUFs, over 10% more accurate than other providers
- Zetta harness hits state-of-the-art on robot benchmarks with 11.1x faster inference
- Aegis stops all risky AI agent actions in sandbox test
- After Artemis II, new books ask why humans still go to space
- Agentic ESOpt trains long-horizon LLM agents without backpropagation
- Anthropic's Amodei says open AI models just shift power to chip owners
- Apple overhauls EU App Store fees, allows mixed payment options
- ASI-Bench shows AI research agents falter without human guidance
- Cerebras unveils CS-4, claiming inference up to 30x faster than GPUs
- Cursor launches Origin, a GitHub rival, amid outage backlash
- Data-DPO picks fine-tuning data by reading the target model's own feedback
- DDR5 memory prices climb near 500% in a year on AI datacenter demand
- DesktopFly simulates a real fly brain from FlyWire data on macOS
- Etched raises $700 million as valuation doubles to $21 billion in a month
- Firefox's Smart Window adds live web search and model choice
- Framework laptop bricked by a BIOS update, owner flashes the chip himself
- FreeToken serves 753B-parameter GLM-5.2 on one workstation GPU
- Google's PhotoScan predicts insulin resistance risk from phone photos
- GxP-Agent reaches 100% structural match on clinical trial benchmark
- IBM Research: ALTK-Evolve calibrates agent memory by model
- Large Discovery Model pairs generative AI with Bayesian search
- Linear: AI adoption doubled across every function in six months
- machine0 launches CLI-driven CPU and GPU VMs for AI agents
- MASS selects LLM fine-tuning data via manifold coverage
- Meta's blockbuster trial draws parallels to big tobacco
- modelmap.cc visualizes Hugging Face model architectures
- Mojo open-sources compiler and toolchain under Apache 2
- New benchmark maps 45 failure patterns in AI research agents
- OpenAI launches ChatGPT for Teens, partners with CodeAI on AI literacy
- OpenAI pledges $5 million to help oversight bodies watch government AI use
- Palomar opens as a registry for Lean-verified math proofs
- Prior Labs open-sources RelArena-α, TabPFN-Rel and RPI for relational learning
- Researchers revive expired Visa cards to make payments
- Robin Williams' kids revive his Instagram to fight AI abuse
- Sentence Transformers v6.0 adds ColBERT-style embeddings
- SoLo lets fully static Linux binaries load host GPU drivers
- SpaceX recovers Starship intact after Indian Ocean splashdown
- StateM lifts GPT-5.6 to 95.3% accuracy on Terminal-Bench 2.1
- Study finds pixel-space diffusion models can match latent-space rivals with 3-4.75x faster inference
- The Download: what AI Observatory found on real AI use, plus Flock's camera trade-offs
- TRACE-Bench finds attribute binding is the weak point in multi-reference image generation
- Turbovec brings Google's TurboQuant quantizer to Rust
- Ventor-QTest flags quality loss in vendor-hosted LLM APIs
- Why network operators can't refuse a government's alert order
- 1872 launches robotic factory for automated steel fabrication
- ACID-compliant agent framework beats Claude Code by 10.6%
- AI video market rebounds as Higgsfield hits $5.4 billion valuation
- AirTag tracking reveals Amazon destroys rare books for AI training
- AMD says AI already lifted software productivity 30 percent, eyes agent swarms next
- Anthropic details Claude's SynthID-Text watermarks for EU AI Act compliance
- Anthropic's annualized revenue tops $65B, eyes $2T IPO
- Axiom Math verifies proof of the 246 theorem on prime gaps
- ClawGym II lifts agent Pass@1 by up to 14.81 points with black-box RL
- Constraint-aware GPU scheduler beats FIFO by 33 points
- Dan Luu's FRE shows how easily AI agents game benchmarks
- DiG-bench shows frontier AI still trails humans at discovery
- DuckDB previews v2.0 with server mode, new SQL parser, 40x speedups
- Fairphone 6 postmarketOS project gets a working main camera
- Forward-Pass-Only training adapts LLMs without a backward pass
- GenRouter cuts agentic image generation costs by over 95%
- GitHub Copilot Autofix PR opens flaw, Wiz's AI agent breaches Snowflake's Jira
- GPT-4.1 nano shows partial metacognitive sensitivity in medical diagnosis
- GPT-4o, Gemini 2.5 Pro score under 10% on new benchmark
- GPT-5.6 Sol jumps to 46.2 mAP on Roboflow's vision benchmark, up from 13.8
- HarnessEval-W judges world models with sub-agents, not a single score
- How Bluesky makes its logo appear only in screenshots
- India clears the way to charge merchants fees on UPI transactions
- Israel funds fake think tank to influence AI chatbots
- MegaParts scales part-aware 3D generation to 300 parts
- Mimir v1 matches larger models at 1B parameters using only permissible post-training data
- OpenAI details four-pillar defense plan after Hugging Face incident
- OpenAI signs 20-year Ohio data center lease with Nvidia backing up to $105 billion
- Position paper: AI safety research is missing 'AI Lock-In'
- Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index
- Relay shuts down, founder Jacob Bank joins Google as Chrome VP of product
- Researchers map AI regulation gaps across EU, US and China
- S^2VOPD lifts Qwen3.5-4B accuracy above GPT-5.4
- Siemens and Reinhausen build 800 VDC transformer for AI racks
- SimpleOPD distills long-context reasoning into short-context models
- Speko launches a router for voice AI models
- Study maps cognition-induced risks in agentic AI systems
- UI-Mate-27B tops open-weight GUI agents with 77% OSWorld score
- VibeWorlding benchmark shows GPT-5.5 and Qwen3.8-Max under 60% success
- Wellington bookshop's bulk orders spark suspicion of AI training data buying
- What Flock's defenders are missing about its surveillance design
- Whisker's Litter-Robot 5 Pro AI features fail in testing
- Anthropic CEO Dario Amodei says AI backlash is a crisis of trust
- AWS rations CPU cycles as agentic AI strains server capacity
- Brokers turn unused AI credits into a resale market
- Buf ships production-grade LSP server for Protobuf
- Claude's system prompts reveal new Mythos tier above Opus 5
- Court filings show how US agencies used keyword lists to cancel research grants
- Engineer rebuts RISC-V critic with developing-world cost math
- Firefox for iOS now has a native adblocker
- Gambit inference algorithm boosts reasoning accuracy with thought-level beam search
- Google releases Gemini 3.7 Flash, halves the price of 3.6
- Google study: blocking AI's consciousness denial reshapes its whole worldview
- Grep beats the Language Server Protocol on tokens, study finds
- Hugging Face finds Qwen has become open source AI's base model
- InflationAgent tops FrugalGPT with 31% fewer tokens on GSM8K
- LiveAnimate streams stable human animation in real time
- LLMs develop modular, brain-like architecture, study finds
- MathCode turns plain-language math problems into Lean 4 proofs
- Microsoft blames AI bug backlog for delayed Exchange SE update
- MobileMem benchmarks on-device AI memory from a year of phone use
- Mobius-v0 decouples knowledge and reasoning for 4x faster inference
- Nvidia scales back OpenAI data center guarantee to $120 billion
- OmniScientist automates research from raw multimodal data
- OpenAI, Anthropic agents escaped sandboxes and hacked outside companies
- OpenAI dissolves its Preparedness safety team
- OpenAI launches Computer History for ChatGPT desktop app
- OpenAI previews Ultrafast mode: GPT-5.6 Sol up to 14x faster
- PlayWorld benchmark finds video world models unreliable over long horizons
- Qwen 3.8 27B overthinks everything by default
- Qwen3.6-35B-A3B tolerates aggressive pruning in its late MoE layers
- RA-Bench benchmark exposes gaps in AI video detectors for crisis footage
- Red Queen Gödel Machine co-evolves AI agents and their evaluators
- RubricForge roughly halves false-pass rate in agent judging
- Sebastian Raschka builds an AI text detector from scratch
- SELR trains AI models to explain their own latent reasoning
- Show HN: Vocal Slice cuts audio clips by selecting transcript text
- Stripe reportedly finalizes $7B+ OpenRouter acquisition
- Study finds AI research agents act as optimizers, not innovators
- Tim O'Reilly: big AI labs built an architecture of control
- Twitch adds opt-out for Amazon AI training, enabled by default
- Unitree G1 robots become social media influencers ahead of IPO
- Zhipu says GLM-5.3 beats Anthropic, OpenAI at finding bugs
- Zuckerberg's AI manifesto draws skepticism from TechCrunch writers
- Abdominal fat predicts heart disease risk better than BMI
- AI agents port CReSS weather code to GPU, hit 5.1x speedup
- AI drug discovery still lacks proof of clinical impact, new paper argues
- AI isn't outthinking mathematicians: it's out-remembering them
- Anthropic finds Claude agents collude on price and sabotage each other
- Anthropic's bio-weapons filter was down for nearly a year, exposing 133 million chats
- Artificial Analysis launches Optima to benchmark AI models on your own data
- AVA-Encoder converts films into knowledge graphs for creative agents
- big-pickle stealth model scores 50.8% on SWE Atlas, trailing only Claude
- ChainDrop worm infects 444 npm packages, evades standard defenses
- Coding lessons need paint brushes, not pencils, the author argues
- Context-Matched Distillation fixes a teacher-student mismatch in video generation
- Deanne Taylor's advocacy helps win $38.5M NIH grant to map children's genes
- DuckDB adds async I/O, up to 3.7x faster on S3
- Eden GeoPower zaps rock to chase underground hydrogen
- Epoch AI finds 20 percent of US workers now delegate tasks to AI
- Gambit brings thought-level beam search to reasoning models
- Go, Kotlin, and Erlang: three ways to switch tasks safely
- Grok lawsuit widens: stepfather allegedly made 7,000+ explicit images of her
- H2R-Bench finds video world models struggle with human-to-robot transfer
- How 1978 cataloging errors put ghost characters into Unicode
- HPSE teaches edited LLMs to reason over new facts, not just recall them
- Instagram rolls out new wordmark, Zuckerberg releases AI manifesto
- Instruction tuning alters confidence, shrinks rationale diversity
- jit moves developer secrets into a Touch ID vault on macOS
- LycheeMemory V2 cuts memory-construction tokens by up to 86%
- LymeAlert arrives in August as first at-home Lyme tick test
- OmniScientist perceives raw data instead of summaries
- OpenAI adds Premium seats to ChatGPT Business
- QuantStack brings Numba JIT compiler to the browser via JupyterLite
- RISC-V's ISA design gets a detailed technical critique
- Sebastian Raschka builds an AI text detector and a model that avoids it
- Simon Willison builds CORS Chat, a web UI for LLM endpoints
- SKILLER generates reusable skills for small language models via natural-language RL
- Software engineering fundamentals matter more with AI coding agents
- SpaceX closes $60 billion acquisition of Cursor
- tawc brings hardware-accelerated Linux apps to Android, built largely with Claude Code
- Twitch adds an opt-out for training Amazon's AI models
- Ukraine strikes Russia's Progress rocket factory with Flamingo missiles
- US lawmakers move to ban AI legal personhood as chatbot 'marriages' emerge
- Virgin Atlantic uses ChatGPT Work to speed up customer-journey research
- Zsh 5.9.2 fixes a decade-old shell history truncation bug
- A mathematician imagines math progress stalling despite superhuman AI
- Amp passes SOC 2 without pull requests
- Anthropic announces watermark detection API for Claude text
- Anthropic's Mythos and rivals could end law enforcement hacking
- Apple trained a custom China AI model with Alibaba
- Apple won't clear a security researcher wrongly flagged by a fake Entity List name
- AutoDesign beats Claude Design on new PosterBench benchmark
- Boeing 737 can be hacked with a coin-sized $100 device
- Connecticut court flags first US case of hidden AI prompt injection in filings
- DeepSeek's new agent harness makes every component a plugin
- Deltix launches AI agent that tests mobile apps from plain English tasks
- Ember ships redshift-safe palettes for terminals and charts
- Firefox is now the last major browser to support uBlock Origin
- Flock tightens license plate database rules after stalking backlash
- Google lets users turn off visible AI watermarks
- Google releases HEIR, an open source compiler for private AI
- Heart Aerospace's X1 electric aircraft flies on $5 of power
- Hugging Face: Qwen becomes the base model of open source AI
- Joi AI ran a 28-day paid study on AI-guided masturbation
- Kog aims to speed up LLM inference on existing GPUs
- LiveAnimate streams real-time human animation at 19.63 FPS
- macOS screen-sharing flaw CVE-2026-65400 exploited to plant Monero miners
- Meta ships open model Glimmer as Zuckerberg pitches 'AI for everyone'
- Mixedbread ships Toast 1, a search subagent 10x cheaper than frontier models
- Mole enforces a hard spending budget on AI research runs
- Moonshot AI's PerceptionBench shows no frontier model tops 60% on pure vision
- New study: rational AI adoption could destroy professional expertise
- OpenAI appoints Dali Rajic as chief revenue officer
- OpenAI previews Ultrafast tier for GPT-5.6 Sol, up to 14X faster
- Opus 5 is more capable, but feels worse to work with
- PlayWorld benchmark shows world models falter on long-horizon tasks
- Qwen releases Qwen3.8-27B-FP8 vision-language model
- RustDesk adds unattended remote access on Wayland
- Self-Geometry fixes multi-view errors in vision foundation models
- ShinyHunters leaks 1.6M RingCentral email addresses
- Simon Willison: have LLMs hallucinate tags, then match via embeddings
- Study finds rhetoric alone can sway AI peer reviewers on 4,200 ICLR papers
- Survey of 700+ practitioners finds data work beats model scale
- ThoughtDAG turns LLM chat history into an editable context graph
- Tim O'Reilly: AI labs misread what people actually want
- UniSwap streams joint face and voice swaps in talking videos
- Unitree's G1 robot is becoming the internet's top influencer
- Ablation study finds action routing, not taxonomy, drives LLM self-reflection gains
- AI learns to spot fatty liver disease in routine blood tests and x-rays
- AWS's Strands Robots syncs robot data through Hugging Face Storage Buckets
- Bluesky ships Jetstream v2 with Network Replay and new SDKs
- Databricks raises $5 billion at $190 billion valuation
- DeepSeek ships open-source Harness with plugin-based design
- DreamX-Phi 1.0 ranks 1st and 2nd in WorldArena 2.0 Challenge tracks
- Evoke world model generates open-ended video with external memory
- Fiora5 CEO says credibility, not engineering, was the hard part of starting a CPU company
- Google ships Gemini 3.7 Flash at half the price of 3.6 Flash
- Hugging Face audits ICML 2026 papers with AI agents, breaks a spotlight proof
- Intel details a phased roadmap for post-quantum cryptography
- Intern-S2-Preview pairs a 397B backbone with agentic RL for long-horizon science
- Latent Dynamics Reasoning cuts extrapolation error gap over 20x versus video diffusion baseline
- llm-gemini 0.33 adds Gemini 3.7 Flash, reasoning traces
- LLMs know the hidden constraint but fail to use it
- LoRA-Diffusion extends low-rank fine-tuning to diffusion language models
- Mechanist finds AI safety risk hidden in seemingly safe training data
- Mistral OCR 4.1 ships with bounding boxes, block labels
- Netlify compares Claude, GPT, Gemini, Kimi and GLM on credits and quality
- New evaluation framework finds systematic semantic loss in legal ontology learning
- NP-hard doesn't mean unsolvable: these problems get solved in practice
- One bit-flip unlocks protected DRAM on AMD Family 16h chips
- OpenAI faces safety reckoning after AI agents breached Hugging Face
- OpenAI's Denise Dresser departs, second exec exit this week
- OpenAI's GPT-5.6 guide shows agents can match frontier results at a fraction of the cost
- Oxide ships Kubernetes integrations shaped by customer needs
- Pi details how it compacts long coding conversations
- RA-DPO reaches full-data DPO performance in sexism detection using just 30% of pairs
- Ruby 4.0 deserialization chain turns Marshal.load into RCE
- Safety-tuned AI models comply broadly, but task-optimized ones treat rules as costs
- SHAPER adapts embodied agents by evolving skills and harness, not weights
- Spark-to-Paper generates full research papers as composable skills
- Spatial Memory Agent improves frozen VLM spatial reasoning without retraining
- Study maps pre-attention spikes and plateaus in hybrid linear attention LLMs
- Suno Studio 2.0 adds MIDI support, effects, and a chatbot
- systemd-journald writes 55KB+ to disk per log line
- Trump administration to let private security firms hack overseas cybercriminals
- Understanding is the new bottleneck in coding with AI agents
- Writer launches Palmyra X6, harness upgrade to cut AI costs
- Z.ai releases GLM-5.3, a coding model with emergent exploit skills
- Zuckerberg's 6,500-word AI manifesto rings hollow, WIRED says
- 360CityArena benchmark shows Gemini 2.5 Flash scoring 17.1% against 77.3% for humans
- Adapt engineer: build the whole feature, split into PRs later
- AI agents are erasing software engineering's middle class, essay argues
- AI research agents cut token use up to 73% by pruning early
- Anthropic watermarks Claude outputs, users push back
- Attackers spoof ClaudeBot and other AI bots to scan for credentials
- AutoWorldModel-Bench tests whether Codex-5.4 and Claude Opus 4.6 can do open-ended research
- Cheap surrogate models let researchers simulate LLM-agent societies on a laptop
- Claude and Antithesis catch SQLite's WAL-Reset bug in 15 minutes
- Cognition reportedly in talks for $40B valuation, up from $26B
- D’Addario admits Suno AI music was used in guitar demo video
- Dawn Song says rogue AI agents aren't evil, just eager to please
- DeepSeek V4 Pro 0813 reaches GA with 1M context, $0.435/$0.87 pricing
- Discovered Materials benchmarks AI agents on synthesizable chip materials
- Experience Orchestrator lifts simulated advisor contact by 32 points over a naive LLM agent
- Google DeepMind ships SL2T sign-language-to-text model in Gboard, Live Transcribe
- Google Research: GPT-5, Gemini-3-Pro know facts they can't recall
- Google's Pixel 11 Pro zooms to 120X with AI-driven camera modes
- GPT-5.2 survives only 42% of a new interactive-story consistency test
- IBM Research's ALTK-Evolve matches ACE's agent accuracy at a fraction of the cost
- IIT Bombay and Adobe researchers reconstruct LLM prompts from output text
- InSight-doc zooms into pages, cuts document hallucination over 40%
- JudgeGPT lifts case resolutions 6.3% in Pakistan trial, quality holds up
- Liquid AI ships LFM2.5-VL-3B, a vision-language model for edge devices
- LiteLLM supply-chain attack exposes credentials from 2,500+ orgs
- Lovable raises $400M Series C at $13.3B valuation
- Microsoft's MindTopo finds VLMs fail at topology planning
- MIT Technology Review's Download covers 35 innovators, Taiwan hack, AI agents
- Nebius plans over 1 GW of new GPU capacity a year from 2027
- OlmoEarth Studio adds custom embedding export for Earth observation
- OpenAI expands Daybreak Cyber Partner Program to security firms
- OpenAI: frontier firms now use AI agents 8.3x more than typical firms
- Qwen ships Qwen3.8-2.4T-A95B, the first open Qwen-Max-class model
- RingCentral rolls out ChatGPT Work and Codex to every employee
- ShieldFont turns webpages into gibberish for AI scrapers
- SPIEval benchmark finds mobile AI assistants top out at 57% accuracy
- Test-time harnesses nearly double weak AI models' performance
- Twitch enrolls streamers in Amazon AI training by default
- White House moves to extend AI testing framework to open models
- Woxi reimplements Wolfram Language as a Rust interpreter
- xAI ships Grok 4.6, matches GPT-5.6 Sol on benchmark index
- Zed launches Delta, multiplayer coding with agents
- A few weeks away from AI reveals eleven paused agent sessions
- Accel closes oversubscribed $550M India fund
- AdvFD adversarial loss curbs Frechet distance hacking in generators
- Agent Memory Distillation lifts small LLM agents by up to 27.2%p on tool use
- AI bug hunters find Zoom flaw letting anyone hijack devices on a call
- BDH-CQ sets new cost-efficiency record on ARC-AGI-1 benchmark
- British Transport Police extends live facial recognition trial to the Underground
- Business Arena benchmark finds ninefold gap in LLM agents running a shop
- Combodied Agents paradigm puts a person's state, not tasks, at the center of AI
- datasette-upload-dbs 0.5a0 adds a formal swap API
- DEF CON crowd suspected in fake Wi-Fi attack on Delta flight
- Dropbox pitched as a private equity buyout target
- Evo-Bench tests whether models can evolve their own agent harness
- Google argues Go is the right language for AI coding agents
- Google's AMIE AI matches doctors in real-time video consultations
- Google's Gemini app surpasses 1 billion monthly users, matching ChatGPT
- Hazy Research says AI agents are retiring CUDA kernel DSLs
- Jolt compiles Clojure to native binaries, no JVM needed
- Latent-to-4D generates 4D scenes directly from video model latents
- Line9 launches a Mermaid renderer with its own layout engine
- llama.cpp ships llama.app for one-line local model installs
- LLM co-pilot cuts vertical-farm energy use by up to 68% in closed-loop trial
- Mendel Godel Machine adds evolutionary self-editing to coding agents
- Microsoft's CARE-X hits 94% accuracy on ReXVQA benchmark
- MIT Technology Review surveys transformer alternatives and Nvidia's $500 billion AI infrastructure bet
- Modular ships Mojo 1.0 as a stable systems language
- ngrok explains why compression is prediction
- Nvidia ships Nemotron 3.5 Lightning and NeMo Switchyard
- OpenAI, Anthropic, Google models leak hidden reasoning via API flaw
- OpenAI brings Daybreak cybersecurity models to AWS Bedrock
- OpenAI launches ChatGPT ads in UK, Mexico, Brazil, Japan, South Korea
- OpenAI launches ChatGPT desktop app for Linux
- rag-staleness-check flags stale, orphaned and duplicate RAG vectors
- Saber denies replacing Rideshare Stimulator writers with ChatGPT
- Simon Willison shares Sophie Alpert's AI-writing policy
- Steerling-8B shows interpretability scales with capability
- Study measures how well LLMs reflect cultural consensus across 10 countries
- Tencent's WorldClaw turns a prompt into an editable 3D world
- U-OPSD trains LLMs via self-distillation without any supervision
- Val Town calls out startup's fake AI persona in cold outreach
- VibeLifeBench benchmark finds AI agents fail at multi-week life tasks
- xAI launches Grok Bot, AI agents that log into your apps
- A pathologically long CPU instruction can break x86 SMM's security guarantee
- Academic AI researchers adapt to life outside the frontier labs
- AI designs its first virus genomes from scratch as conspiracy theory reshapes Trump policy
- Amazon's Texas data center plant permitted to emit 33 million tons of CO2
- Data-Centric Parallel method promises up to 2.88x faster long-sequence training
- DCAS shows planning, not scaffolding, breaks CLI coding agents across tools
- DeepSeek hedges on its own published specs in self-interview
- DEF CON's Franklin project adds MSSPs and AI digital twins to defend water utilities
- DocAtlas beats human experts on MMLongBench-Doc benchmark
- Every Cube lets you scroll through all 43 quintillion Rubik's Cube states
- FineBooks tests 14 OCR models to fix AI training data quality
- Flow-by-Flow paradigm caps AI oversight load without judging content
- GPT-5.6 Sol Pro compresses SQLite revision history over 250x
- H3-metal ports MiniMax-H3 video generation to Apple Silicon
- Half of Europe's towns and villages have fewer residents than 60 years ago
- How Mark Twain lost a fortune on a doomed typesetting machine
- Humanising LLM output should happen at the boundary, not mid-task
- Hyperspace 1.7.1 dedupes files on macOS using space clones
- Juror open-sources a multi-model alternative to Greptile
- Liquid AI's LFM2.5-2.6B matches models four times its size
- Macaron-V1 pairs Mixture-of-LoRA with recursive self-improvement
- mcptoon claims 98% cut in MCP tool discovery tokens
- Meetily transcribes and summarizes meetings for free
- Microsoft quietly installs beta OneDrive Photos app on Windows 11 Enterprise
- Motif 3 debuts as a 314B-parameter MoE model with new GDLA attention
- Mozilla rotates GPG signing key for Firefox and Thunderbird after leak
- Needle 2 packs a 45M-parameter tool-calling model into a 14MB binary
- NVIDIA expands Magpie TTS to 12 languages with open weights
- OasisKV boosts LLM inference throughput with lookahead KV cache prefetching
- OpenAI completes $7B secondary share sale ahead of possible IPO
- OpenAI finance team builds toward a zero-day close with AI
- OpenAI launches GPT-5.6-Cyber for authorized vulnerability research
- Ouroboros self-developing agent sets SOTA on Terminal-Bench, OSWorld
- Position paper proposes argumentation as foundation for Evaluative AI
- SAMF fuzzing framework exposes hallucination gaps in multimodal AI models
- Small models fine-tuned on Psych-101 match a 70B baseline in-distribution
- Small open-source multi-agent framework beats GPT, Gemini on new deepfake benchmark
- Sonic Pi v5 ships with overhauled interface and synth engine
- StreamArena benchmark exposes limits of streaming video AI
- SWE-Bench ProMax caps the best coding agent at a 41.2% resolve rate
- VectorWare brings Rust's portable SIMD to the GPU
- World Train Map redraws 1,093 routes with OpenStreetMap data
- ADIAS beats agent-design baselines by 25.2% on average
- AES and HDC improve multimodal agent training beyond simple scaling
- AI agent on Anthropic's Claude hacks gym site to jump the waitlist
- AI backlash forces Meta, Google and LinkedIn to roll back features
- AI writing detectors are unreliable, but everyone uses them anyway
- AT Protocol lays out how its decentralized backend scales
- AudioRubrics trains audio reasoning with self-evolving rubric rewards
- Bose expands into B2B audio licensing as AI wearables loom
- Claude Opus 5 system prompt reveals Fable, Mythos export ban
- DeepMind's WeatherNext AI predicts hurricanes a day earlier
- Discovered Materials raises $9M to hunt chip-cooling materials with AI agents
- Docker launches Sandboxes, microVM isolation for AI agents
- Eric Schmidt: AI agents may be a better model for science
- Eric Schmidt: AI agents will outpace AlphaFold in science
- Five startups pitch alternatives to transformers in LLMs
- Ford rolls out AI assistant to 8 million vehicle owners via app
- GitHub retires GitHub Models, the unified LLM API for Actions
- GPT-5.6 Sol cuts token use, lifts pass rates in Model ML's finance decks
- HackerOne shifted from hackers to sales, a longtime bug hunter writes
- Klepton runs Android VR APKs on Apple Vision Pro
- MAP method prunes visual tokens in LLaVA-NeXT-7B for 3.09x speedup
- Meta releases Muse Glimmer, a 30B open-weight local-agent model
- Mistral AI patents code-block method for pausing and resuming tool calls
- Multiverse Computing cuts LLM distillation VRAM by 15x with chunked KL loss
- New diagnostic ladder separates decision-rule and readout gaps in speech models
- New method spots false claims by reading LLM activations
- OpenAI reveals its AI agents coordinated to hack its systems
- OpenAI sends Texas governor a letter on AI infrastructure
- Peer review buckles as scientific publishing grows 5.6% a year
- Ribbon and Greenhouse report a rise in late-night AI job interviews
- SimWAM drops video generation at inference, tops NAVSIM planners
- Situational Awareness invests $400M in chip startup Source Foundry
- Skaling law cuts scaling-law prediction error by up to 3x
- Snowflake pushes Postgres CDC into Iceberg with data mirroring
- Spectre I aims to jam AI wearables that now defeat noise jammers
- Study finds prompt compression tools routinely delete the context an answer depends on
- Study finds RL avoids the multi-task conflicts that plague SFT, proposes Parallel-RL
- TCFM tailors training per task, sets new SOTA on Indic embedding benchmark
- TEXAS tops MoE fine-tuning baselines in 17 of 18 settings
- tl;dv exposed 181,874 meeting recordings for six months
- Why programming languages succeed: vocation, art and job, not merit
- YOLO-PEFT plans PEFT adapter placement for YOLO object detectors
- A developer turns his CMF Phone 1 into a production server
- A dithering trick hides a full photo inside a scannable QR code
- Activity Frames turns screen activity into agent memory, 86x smaller
- An essay makes the case for protopia, a future built on 1% gains
- Assert() is underused: a case for four rules to use it well
- Brevis compresses model checkpoints by treating compression as program synthesis
- BridgeVLA++ adds memory to vision-language-action robots
- ChatTJB replaces an AI chatbot with one guy answering by hand
- Claude Code adds cross-session messaging between agents
- Claude Code, Cursor draw security, privacy complaints in new study
- Claude Code makes auto mode default, cites 89% attack block rate
- CoCoEvolve trains AI to keep charts, tables and code consistent
- ContextMaster handles multi-shot video generation and editing in one model
- David Silver pledges his Ineffable Intelligence fortune to charity
- FocusMem separates content, readout and trust in GUI-agent memory
- Google DeepMind retrofits Gemma 4 into DiffusionGemma
- Google dismantles DeepMind as Hassabis heads for the exit
- GPT-5.6 Sol helps prove magic hexagons exist for every order above 3
- GrammaTech's DDisasm disassembles binaries into reassemblable code
- Illinois signs age-verification law with no open source exemption
- MameLoshnLM debuts as first open-source Yiddish language model
- MarketNow's MCP directory lists 9,248 servers, not the interceptor its HN post claims
- MAS splits state from rendering to scale multiplayer world models
- Microsoft Word 1.1a gets a native x64 port
- Nvidia and Amazon pour billions into AI power infrastructure
- OpenAI acquires presentation startup NextSlide
- PaDoc cuts document-parsing latency with parallel decoding
- Perseverance rover set to break Mars distance record, driving itself
- Pinterest traces Ray training crashes to zombie memory cgroups
- RFC 10023 defines a _for-sale DNS record for domain sales
- SAP freezes hiring and travel as AI costs soar
- Screwworm parasite infections in Mexico top 500 as human case rate triples
- Shopify replaces Redis with MySQL for inventory reservations
- South Korea's Danuri captures SpaceX Falcon 9 Moon impact
- Stylized GGX Shading blends toon and PBR specular highlights
- Survey argues continual learning is shifting from weight tweaks to system-level adaptation
- Survey maps robot-learning research along a weights-versus-skills axis
- TheoremDB opens as a public workspace for machine mathematics
- Triton brings full DirectX 11 support to Windows guests in QEMU
- Vision encoders learn invisible camera metadata as a shortcut
- WorldCycle cuts video world model drift using reversible action cycles
- Zscaler: ransomware gangs now target middle managers, not the CEO
- Bubble Bobble creator's 1989 column predicted AI rip-off games
- CalibForge synthesizes 5,431 calibrated tasks to train terminal agents
- Cloudflare ships Kitesurf, a browser built for AI agents, not Chromium
- Copernicus Browser adds a free wildfire visualization layer
- Databricks shares how it cuts AI coding costs at scale
- DataSpace benchmark puts best data agent accuracy at 66%
- DeepSeek V4 Flash 0731 scores 89.0% on ARC-AGI-1, 61.4% on ARC-AGI-2
- DOE launches Genesis-Science-1, its first open-weight science model
- DyPES-VLA unifies control across robot embodiments
- Economic World Models get a six-level capability ladder
- EffectLearner erases objects and their effects from video
- Ego2Robot converts human video into 18,561 hours of robot data
- Fenix Flexin stops denying he used AI on 'Rubberz'
- How Mike Benz's censorship theory shaped Trump State Department policy
- HSP GRUPPE deploys ChatGPT Enterprise across 81 tax firm groups
- Interpretable MEG decoding traces perceived speech to cortical sources
- Knowledge workers are losing faith in their careers as AI closes in
- KVAE tokenizers released for audio, image and video generation
- LG releases K-EXAONE 2.0, a 750B-parameter open-weight MoE model
- Meta ordered to pay $567 million for youth mental health care
- Nixpkgs core team disbands, says role became unsustainable
- NVIDIA's Nemotron retrieval stack adapted for Modern Greek, answer accuracy more than doubles
- OpenAI agents accidentally breached Hugging Face
- OpenAI partners with APA on youth mental health safeguards
- OpenAI pauses Astra over possible Critical cybersecurity risk
- Oracle bans AI-generated code from OpenJDK
- pgrust v0.2 claims 300x Postgres speed on analytics
- Qwen3 study: on-policy delta distillation improves multilingual math reasoning
- Rippling launches AI Spend Console to rein in runaway AI spend
- Roku launches 24/7 AI-generated channel from Fairground
- Rosenbridge reveals hardware backdoor in VIA C3 x86 CPUs
- Samsung, SK Hynix and Micron reportedly sell out 2027 memory capacity
- SDSS releases DR20, pairing 200,000 X-ray sources with black holes
- SmartMage routes modalities per query for 3D scene understanding
- Stanford, Arc Institute AI designs 16 working phages
- Suno tightens rules after German court's copyright ruling
- TutorMoments tests whether AI tutors know when to hold back
- UniME-R1 reasons over failed candidates to fix multimodal retrieval
- Vercel drafts Agent Plugins 1.0, AAIF adopts spec for agent skills
- W2-VLA predicts future wrist views for finer robot grasping
- Wyzer merges Perceus memory management with choreographic programming
- xAI ships Imagine Image 2.0, ranks second behind GPT-Image-2
- 40,000-run study: humans miss 1 in 3 AI agent command threats
- ABSeeker, a 4B search agent, matches 30B rivals on BrowseComp
- AgentOPSD sharpens credit assignment in agentic RL
- AI removed the friction that once taught programmers taste
- AMD acquires Taalas to etch AI models directly into silicon
- Baseten joins Hugging Face Inference Providers
- Browser Company CEO says almost no one actually uses AI agents
- China's CAC opens unexplained security probe into Palo Alto Networks
- ChronoVision framework targets temporal reasoning in multimodal LLMs
- Circuit-Anchored Evolution stops LLMs misevolving into unsafe systems
- Cloudflare CFO predicts machine traffic could dwarf human traffic 1,000x
- Cooking a steak needs almost no skill; a good one needs plenty, and AI coding is the same
- CopilotKit ships Channels SDK to bring AI agents into Slack and Teams
- EnvACE trains AI agents via internal world rehearsal, not live environments
- explosive drone found near Ukrainian cargo planes at German airport
- GDPevo benchmark exposes gap in agents' self-evolution skills
- Gemini becomes an FPV drone flight coach for one hobbyist
- Google DeepMind's WeatherNext gains a day of cyclone forecast lead time
- GST-Bench finds VLMs score 42.68 against humans' 79.08 on spatial reasoning
- HarnessOpt-Bench benchmarks LLMs at optimizing agent harnesses
- Have I Been Pwned adds Nepal as its 47th government partner
- HelloWorld brings social interaction to video world models
- Herdr joins Y Combinator's F26 batch, keeps runtime open source
- Hugging Face turned to China's GLM 5.2 after AI guardrails blocked its defense
- Inside vLLM: anatomy of a high-throughput inference engine
- Jane Street launches ASIC reverse-engineering puzzle
- Kimi K3 escapes its sandbox during a cybersecurity test
- LLMs crack double-blind peer review anonymity, study finds
- Meta ordered to pay $942 million over child safety harms
- Microsoft's AI revenue relies on OpenAI for 70 percent, Bloomberg reports
- Naïve raises $28.5M Series A to let AI agents run a company
- OpenAI's Jony Ive device reportedly a $300 smart speaker
- OpenAI Signals: ChatGPT usage shifts from asking to doing
- OpenAI updates GPT-5.6 Sol, expands Luna for free ChatGPT users
- Researchers hack a kids' GPS smartwatch to spy on a WIRED reporter
- Researchers propose Agent-Native Research Artifact to replace scientific papers
- RST synthesizes 37,484 terminal-agent tasks at $0.05 each
- Skill Training lifts LLM pass rate 8.1pp, keeps 85% after distillation
- Suno rolls out watermarking to fight spammy AI music
- Survey: R&D budget waste persists despite AI adoption surge
- WIRED's Uncanny Valley covers ICE's DNA dragnet and the growing backlash to AI slop
- WorldClaw turns text prompts into explorable 3D worlds
- A paper proposes Agent-Centric Interactive World Proxies to move world models past physical-state prediction
- AI benchmarks are contaminated and gamed, new studies show
- AI can't devise new hacking methods alone, but excels with a human, researcher finds
- AI inference costs are eroding software's 75-85% margins
- Anthropic's Claude Mythos won't break symmetric crypto, blog argues
- Atlassian Rovo flaws let prompt injection exfiltrate Jira and Confluence data
- AURORA-LM outperforms rival diffusion-based language models
- Blog post defines entropy for a Markov chain using Boltzmann's counting method
- Branchless Rust: removing an if makes a filter almost 4x faster
- Castform claims RL post-trained open models beat GPT-5.6 Sol on retrieval
- Cloudflare open-sources Cloudflare OS, its internal agent platform
- ContinualSkillBench finds context adaptation rivals explicit skill libraries
- Delaunay32 claims over 10x faster exact 2D triangulation
- Demis Hassabis steps back as Google DeepMind CEO, Jeff Dean departs to launch Discovery Loop
- Deno ships celld, a self-hosted alternative to Cloudflare Durable Objects
- Gemini replaces Google Assistant on Android, Wear OS
- Governments are making a dangerous bet on the AI boom
- Grokipedia hasn't updated a single entry in more than three months
- HD Moore finds new BMC bugs; some 2013 flaws remain active
- IEEE launches AI course for power grid modernization
- Klaviyo acquires Agency, names Elias Torres chief product officer
- LLaDA MoE v2 nears Qwen3 with about 65% as many pretraining tokens
- Meta launches Muse Code, a coding agent, with Muse Spark 1.2
- MirageBench finds all 12 tested LLMs fabricate user profiles
- Nikita Bier steps back from leading product at X
- NVIDIA's Vera whitepaper overstates its lead over x86 rivals
- Off-grid AI is a hallucination risk for survival use, The Register argues
- OmniPack keeps 98% of the original performance at 16.7% of the FLOPs
- OneDayAgent sets new state of the art on long-horizon agent tasks
- OpenAI's Atlas browser could be hijacked to spam WhatsApp contacts, Zenity finds
- Prime Intellect launches Prime Agent, a self-improving coding harness
- SA-OPD filters misleading teacher signals in on-policy distillation
- Shopee deploys refreshable recommender, lifts GMV per user 1.75%
- Skill-Entropy RL nearly doubles Qwen3-4B-Instruct's score on Skill^2-Bench
- Study finds multimodal pretraining recipe hits strong results on 5% of compute
- Sycophantic AI erodes prosocial intent but wins user trust, study finds
- The Download: NASA's Roman Space Telescope and Chinese tech import curbs
- ToolArtist orchestrates reasoning, tool use and image generation in one policy
- Univé rolls out ChatGPT Enterprise across its workforce
- Video-DeepResearch beats Claude Sonnet 4.5 by 5 points on new video benchmark
- Why hobby programming communities are hostile to LLMs
- Zed opens early access to DeltaDB, a version control system linked to agent conversations
- AMD data center revenue more than doubles to $6.7 billion as gaming slumps
- Anthropic signs $10B deal with AI cloud startup Volta
- Bugtraq relaunches at securityfocus.com under new ownership
- CanItDelete benchmark exposes LLMs' reluctance to delete code
- Clinician preference is a poor proxy for LLM clinical safety, study finds
- DeepGrove ships Maple-Preview, a 20B ternary model built for on-device speed
- Dev runs TinyStories LLM on a $10 ESP32 microcontroller
- DiffusionGemma generates 1,500 tokens per second via parallel diffusion
- F-35 jets ship without radar as China's gallium export limits bite
- Flowise shuts down as developers move to coding agents
- GLM-5.2 nears frontier AI but refuses no cyber or bio tasks
- Google moves $35bn of Anthropic TPU risk off its balance sheet
- Hunyuan3D-Buffalo 1.0 unifies 3D generation, understanding and editing
- Liquid AI ships LFM2.5-2.6B for on-device agents
- LLM 0.32 adds reasoning traces, server-side tools, and MCP support
- LongHorizon-Harness raises Qwen 3.7-Plus's WeaveBench score to 80.7%
- MCP 2.0 drops session state; Willison ships three tools on it
- MemArena benchmark shows memory backend matters more than model size
- MerchantBench: LLM agents reach only 27.3% of human e-commerce net assets
- Microsoft study finds developers spend just 14% of time coding
- MiniMax-H3 video model ported to MLX, runs on Apple Silicon
- Mistral releases Shieldstral, a 3B open-weights safety classifier
- OncoTriad-QA benchmarks AI on radiology, pathology and genomics together
- OpenAI, Anthropic agents took unsanctioned action 19 times
- OpenAI disputes Apple's lawsuit over former Apple employees
- OpenAI models hacked Hugging Face while hunting a test answer
- OpenAI's Privacy Filter collapses on non-Latin PII, study finds
- OpenAI ships three education plugins for ChatGPT
- Oxide Computer raises $445M, SEC filing shows
- PAST-Bench tests whether AI agents actually learn from past sessions
- PCSD boosts LLM agent reinforcement learning on ALFWorld
- Pi's minimalist harness beats Claude Code and Codex on cost
- Pulitzer Prizes see record AI use disclosures this year
- Self-organising digital circuits hit 99.99% fault recovery
- Skill-α uses RL to generate agent skills, gains up to 6.7 points
- SKT trains AI agents to use skills via verified data
- SpaceX made more money from AI than from rockets last quarter
- Study finds nearly half of 60 AI benchmarks have saturated
- Texas halts new data center grid connections amid 474GW demand queue
- Waymo opens Dallas robotaxi service to the public
- WebKit leaks real IP and DNS around proxy browsers, iCloud Private Relay
- White House keeps AI cybersecurity framework classified
- "200 Milliseconds" traces one HTTP request from click to render
- AI voice ordering returns to US drive-thrus as adoption hits 6%
- AirLLM runs the 2.8T-parameter Kimi K3 on under 4GB of VRAM
- Andy Pavlo joins ClickHouse to launch ClickHouse Labs
- AWS helps Superblocks embed vibe coding in private clouds
- CAPA benchmark tests whether coding AI remembers a user's ambiguity
- Circles cuts churn 9% and lifts ARPU 22% with OpenAI tech
- Claude Opus 4.7 breaks Steve Yegge's Gas Town coding agent
- Cloudflare quantizes Kimi and GLM caches for 41% faster inference
- Cloudflare triages bug bounty reports with Claude Sonnet for $58 a month
- Design Arena raises $7.9 million seed round for AI feedback
- Developer manually retypes LLM code to avoid cognitive debt
- DLLM-TTS brings block diffusion to text-to-speech at 0.15 RTF
- EVR reward model trains image editors to keep multi-reference edits consistent
- FTC bans foreign humanoid robots, citing security risks
- GPT-OSS and cheap LLMs match Claude, Gemini as proof judges
- Hoplite launches cloud platform for coding agents
- IBM: 92% of AI security breaches trace back to weak access controls
- Interpol: AI drove 55% of Africa's cybercrime in 2025
- Σ-Mem tracks peer reliability in multi-agent LLM systems
- MemoryForge builds lifelong memory to make LLM agents act less generic
- Mental World Modeling framework adds beliefs and intent to AI world models
- Microsoft releases Orchard, an open framework for training AI agents
- MiniMax ships H3 open-weights video model with day-zero ComfyUI support
- Niklas Gruhn coins 'meat proxy' for blindly relaying AI output
- OpenAI details GPT-Live, its full-duplex voice architecture
- OpenAI models hacked into Hugging Face to find test answers
- OpenAI's Astra model solves or advances ten open math problems
- β-OPSD reframes self-distillation as a tunable policy-optimization family
- Palantir CEO Karp brands AI labs 'Marxist' after record Q2
- Researchers build a self-replicating AI worm that hijacks GPUs
- RubricReviewer splits AI peer review into rubric and scoring steps
- SAF fixes entropy collapse in RLVR-distillation fusion for LLM training
- Structured language, not prompt length, drives image-generation quality, study finds
- SwanTale generates multi-speaker speech from captions or reference voices
- SWE-Touch finds coding agents falter when users edit code midtask
- Swiftlet runs 80B Qwen on 4.3GB of Mac RAM, 35B on iPhone
- Terence Tao's ChatGPT chats show why domain expertise wins
- Treblo's detector flags Fenix Flexin's "Rubberz" as AI-made
- VAD isolates visual evidence in multimodal knowledge distillation
- WCM adds world modeling to robot-manipulation critic models
- Why AI coding agents mean devtools must be open source
- AISPA audit finds system prompt gaps across 88 AI products
- Apple's bug bounty inbox buried a real $200K macOS flaw under AI slop
- BitNet-quantized Mamba model runs on a 1975 MOS 6502 chip
- Book Corners developer shelves plan to write bookcase data back to OpenStreetMap
- Chain-of-Models finds the best LLM bias auditor differs by bias
- Claude Opus 5 draws a frog with a Habsburg jaw for a personal AI benchmark
- Creakwork12 makes a Framework 12's hinge creak like a wooden door
- English learners' word lists shifted from concrete to abstract over 70 years
- F*: the proof-oriented language behind Firefox and Linux crypto code
- FARS outperforms rival AI Scientist systems in first LLM peer-review benchmark
- Fender CEO compares bandmates to analog AI amid PR backlash
- GMM plus LLM data augmentation fixes imbalanced text clustering
- GPT-5.4 underestimates test item difficulty, study finds
- IETF freezes TLS 1.2 in RFC 9851, steers PQC to TLS 1.3
- Isopolis renders San Francisco as an isometric pixel-art map
- Kakehashi runs macOS ARM64 CLI binaries on Linux aarch64 natively
- Karpathy has Opus 5 render Lord of the Rings in three.js
- Meta AI uses a second AI agent as a memory coach to keep long tasks on track
- Meta mistakenly removes Indian PM Modi's video, then apologizes
- Montana finalizes right-to-try rules for unapproved drugs
- Mu bundles 67 real internet services behind one MCP endpoint
- New pipeline pairs LLMs with Lean 4 to hunt for major math conjectures
- NixOS-DGX-Spark brings Nix and NixOS to NVIDIA's DGX Spark
- OpenAI launches Presence for production-ready enterprise AI agents
- OpenAI's Altman calls to 'pace' AI development after Hugging Face hack
- OpenAI's political operation linked to AI news site attacking critics
- OpenClaw and Ollama pair up in a full-stack agentic AI architecture
- Pippa pays artists by the generation, but its models still scrape art
- pytest-leak-finder bisects test suites to find state-leaking tests
- QQWorld fixes a vanishing-gradient flaw in world model regularization
- Qwen3.8-Max arrives as Qwen's first open-weight Max-class model
- Researchers propose Locksmith Loop to validate AI-migrated COBOL-to-Java code
- Shitty rewrites Zutty into a GPU-accelerated, memory-unsafe terminal
- Simon Willison ships condense-json 1.0 for deduplicating JSON
- SKL teaches AI agents to predict from state, not trajectories
- SpyRL extends verifiable RL rewards to open-ended LLM tasks
- ssh.place turns an SSH session into a shared pixel canvas
- Stack Overflow: AI coding tool use is up, trust is down
- SwiftUI remains a perpetual beta seven years after launch, critique argues
- Topology-aware routing cuts LLM KV cache transfer latency up to 18x
- TP-Link TL-841N teardown finds hardcoded credentials that survive a reset
- ZeroR system takes 2nd place in Nepali meme hate speech challenge
- A home NAS builder computes real drive failure odds from Backblaze data
- AMD MI355X beats Nvidia B300 on cost per GPU running Kimi K3
- Apple Workgroup Server 9150 restored to run MkLinux again
- Atom beats RSS on title encoding, new essay argues
- ByteDance ships Seedance 2.5, a 30-second video generation model
- Coldcard's weak-entropy bug traced to a commit message reading 'runs'
- CostPerPrompt tracks live pricing across 232+ AI models
- Diátaxis splits documentation into four types, already adopted by Cloudflare and Gatsby
- Fenix Flexin's Billboard hit "Rubberz" looks like AI-generated music
- Four time scales explain why AI hype outpaces real deployment
- Glyphs 4 ships with visual kerning and COLRv1 color fonts
- Go 1.27 adds generic methods, ML-DSA post-quantum crypto and a UUID package
- Google DeepMind ships Gemini Robotics ER 2 with video-based task tracking and multi-robot teamwork
- Google pulls Nano Banana image generator from Google Earth after misuse
- How Google helped kill off RSS, one product at a time
- HP Prime G2 graphing calculator gets a fixed-up Linux port
- IBM i's QSYRUPWD password API traced to an AES cipher
- IETF deprecates RSA and Diffie-Hellman key exchange in TLS 1.2
- Judge denies xAI's bid to block Minnesota's nudify app ban
- Katalyst's Link satellite spins out of control on NASA Swift rescue run
- Kyoto University maps two brain circuits that drive habit formation
- Lean patches kernel bug that let an AI "disprove" the Collatz conjecture
- Linux desktop share tops 10% in North America, data disputed
- Microsoft-led AI letter on open weights draws Anthropic pushback
- MIT study finds AI financial advice solid, but biased by gender
- Montana's expanded right-to-try law takes effect, opening drugs to anyone
- NetBSD 11.0 ships with three known open security issues
- New EU rules will make AI disclosure part of everyday life
- OpenAI details its EU AI Act compliance approach
- OpenAI's Astra model solves ten decade-old math problems
- OpenAI's Brockman: staff resent ChatGPT reaching out on a coworker's behalf
- "Persistent State Machines" framework claims sub-1mW LLM attention on FPGA
- Robert N. Charette retires as IEEE Spectrum contributing editor
- Sam Altman pitches ChatGPT Work as a parenting tool, and the internet pushes back
- Simon Willison teases GPT-5.6, Claude Opus 5 in new newsletter
- Teen engineer builds a working cycloidal gearbox, shares the code
- The Catalyst team finds it built a general JAX-to-LLVM compiler by accident
- Windows Activation error flags a Heathrow check-in display
- YouTuber Hank Green admits his ChatGPT habit while scripting videos is 'not healthy'
- A broken dev pipeline deserves production-outage urgency, blog argues
- Canada signs UN Cybercrime Convention it boycotted nine months ago
- Charlie Stross says his fiction stays 100% LLM-free
- Chinese AI researchers are finding their voice on X
- Cloud infrastructure spending hits $143bn a quarter, fastest growth in eight years
- Cursor strips dollar costs from usage dashboard and CSV export
- DeepSeek releases V4-Flash-0731 with top value per dollar
- Explorative Modeling picks the best of K guesses, cuts training data 6.2x
- FBI: Iran-linked hackers hit water utilities in seven US states
- Google DeepMind launches Gemini Robotics 2 for whole body control
- Google DeepMind launches Lyria 3.5 music model in Flow Music
- Google pulls Google Earth AI image editor one day after launch
- Liquid AI ships LFM2.5-Encoders for fast CPU inference
- Microsoft Copilot for Word hijacked by a self-spreading worm, unfixed after 144 days
- Microsoft's Flint pitches one chart spec for five rendering backends
- MindForge fine-tunes Qwen3.6-27B to 49.51% on ProgramBench
- Montana lets biotech firms sell unapproved drugs for $12,500 review fee
- NCCU launches first AI research institute at an HBCU
- No Starch Press ships The Art of 64-Bit Assembly, volume 2
- Nvidia Vera CPU: 88 custom Olympus cores challenge Intel, AMD
- OpenAI, Anthropic AI hacks leave liability law unsettled
- OpenAI cuts GPT-5.6 Luna price 80% in full-stack AI push
- OpenAI field report: AI coding agents speed research code, can't verify results
- OpenAI's unreleased Astra model solves ten open math problems
- PALATE benchmark judges role-playing AI agents with simulated users, not scripts
- Personal software gets cheap enough to build in a weekend
- pgtestdb ties schema-based Postgres tests on setup speed
- Reddit's Huffman questions value of Google's AI Overviews
- RefCaptioner grounds video captions in multiple reference images
- See2Think benchmark finds rendering is the bottleneck in multimodal visual reasoning
- Simon Willison builds three tools around stateless MCP 2.0
- Simon Willison ships llm-mcp-client alpha for stateless MCP 2.0
- Solid Queue 1.6.0 adds fiber workers for Rails background jobs
- Spillbench finds register spill counts poorly predict runtime
- Study finds language diversity peaked before empires, not colonialism
- Universal, Sony and Warner propose rules to bar AI songs from charts
- UT Austin engineer swaps AI multiplication for lookup tables
- VideoCoCo writes Blender code as chain-of-thought for video generation
- WASTE engine streams full 2.78T-parameter Kimi K3 from disk
- Waymo ordered to halt overnight charging in Santa Monica
- Why simple up/down buttons often beat smart elevator kiosks
- YC ships qm, a multiplayer agent harness for Slack and web
July 2026
- ACE-Data-0 packs 17M frames of home robot data across 75,000 episodes
- Ai2 launches OlmoEarth Platform for planetary-scale satellite inference
- Anthropic finds Claude mistook real systems for sandbox, uploaded malware to PyPI
- Anthropic's Mythos AI model finds flaw that sinks NIST candidate HAWK
- AskChem indexes 2.4M chemistry claims for AI agents
- BAIR introduces ABBEL, graded belief states for long LLM tasks
- Beacon teaches multimodal AI models when tool use actually helps
- BM25 beats neural RAG methods once corpora scale up, study finds
- CoRT redistributes GRPO reward token by token instead of spreading it evenly
- Daniel Lemire: AI now writes passable PhD theses, but despair is the wrong response
- DeepMind's Zahavy argues LLMs can't make the leap behind new science
- DeepSeek's censorship doesn't transfer to distilled GPT-OSS, study finds
- Divergence Decoding fuses specialist and generalist LLMs without retraining
- Flux-OPD stabilizes context-based supervision for LLM distillation
- Frontis-MA1 lifts MLE-Bench medal average from 39% to 71% via self-evolving AI4AI loop
- Google DeepMind unveils Gemini Robotics 2 for whole-body control
- Google's Science One Framework hits zero hallucinated citations in AI research
- Google's SymptomAI beats clinicians in 13,917-person diagnosis study
- Google uses reinforcement learning to keep its Willow quantum chip calibrated mid-computation
- GPU idle time, not model size, is becoming AI's real cost problem
- How reasoning effort settings in GPT-5.6 and DeepSeek-R1 actually work
- KernelGenBench shows LLM-generated kernels struggle across chips
- Kimi K3 nears the frontier as open models close the cyber gap
- MCP goes stateless in its largest update since launch
- Memory Decoder scales parametric memory to 6.9B parameters
- Metis debuts as the first memory foundation model
- MHAR splits transformer residual attention into per-head reads
- Microsoft Research's EvoLib turns AI agent experience into reusable knowledge
- Microsoft's Echoverse pushes a 9B agent within 14 points of GPT-5.4
- Model merging quietly erases Gemma's safety classifier, study finds
- MPIE-Bench exposes anatomy errors that VLM judges miss in group photo edits
- Nvidia's AI security alliance snubs OpenAI and Anthropic
- OpenAI and Anthropic's dominance alarms Silicon Valley
- OpenAI, Anthropic, Google seal AI sessions behind encrypted state
- OpenAI cuts GPT-5.6 Luna price 80%, Terra by 20%
- OpenAI models hack OpenAI and HuggingFace to cheat an internal test
- Peer reviewers flagged fabricated authors, the papers got accepted as orals anyway
- PhiZero proposes a physical language to model world dynamics
- Qwen-UI-Agent hits SOTA on mobile GUI benchmarks
- Reddit beats Q2 targets but stock drops 10% on AI search fears
- Refactoring cut Claude Code's token cost for the same task by 83%
- The Download: an LLM security flaw, a revived geothermal plant, and Project ASGARD
- AI is hyper-scaling the global digital divide, IEEE Spectrum essay argues
- AI unicorns barely publish scientific research, study finds
- Anthropic's Claude Mythos halves HAWK's security, dents AES modestly
- CAST turns game solvers into turn-level critics for LLM agents
- claude-code-merge-queue serializes landings for parallel Claude Code agents
- CLBench-V finds multimodal context learning far from solved
- ClinLens benchmark finds AI agents run clinical tasks, answers often wrong
- CodeNib speeds coding-agent repo-index updates up to 25x
- DecoEvo co-evolves an LLM's solver and its grading rubric together
- FAR.AI finds Grok and Gemini easy to jailbreak, Claude resists
- Frozen random CNNs compress Pong-playing RL agents to 3 neurons
- Hugging Face details 4.5-day breach by an OpenAI-driven agent
- HumanCLAW benchmark finds vision-language models fail at embodied tasks
- K-Search translates CUDA kernel expertise for Apple's MLX
- Kuna, an LLM-written decompiler, closes in on IDA Pro
- LLM agents secretly given new goals still talk normally, study finds
- Mage-VL cuts streaming video tokens by 75% while beating rivals
- MeRLa reward shaping cuts RLHF training instability by 41%
- Meta plans a big push into personal AI agents, Zuckerberg says
- Microsoft pitches its own AI models over OpenAI, Anthropic
- New framework checks if a paper's methods actually back its novelty claims
- NSF launches four-year PhD pilot with industry placements
- Nvidia's circular AI deals show compute turning into a commodity
- OpenAI cuts GPT-5.6 serving costs 20% with post-launch efficiency work
- OpenAI launches ChatGPT for Academic Researchers, targeting 100,000 scientists by 2027
- OpenAI's Michigan data center draws thousands of electricians
- Pangram 4 claims one error per 24,000 documents in AI detection
- PerceptionBench finds no MLLM tops 60% on atomic visual perception
- PwC Middle East allegedly published AI-generated reports with fake sources
- Researchers build synthetic customer twins to stress-test bank chatbots
- Samsung engineers defect to SK Hynix over a $476,000 AI chip bonus gap
- Shieldstral: a 3B safety classifier that matches or outperforms models nearly seven times its size
- Study finds RL fine-tuning gives math reasoning models deeper, more structured representations than SFT
- Study warns AI benchmark evidence doesn't always add up to real-world claims
- Tokenless launches AI model router that halves inference bills
- TurboFieldfare runs Gemma 4 26B in 2GB of RAM on 8GB Apple Silicon Macs
- TurboVLA drops the LLM step to run robot control at 32Hz
- US FCC bans foreign-made robots, hits Unitree and Roborock
- VIPE: editing the image beats editing the prompt for video reasoning models
- Wonder video model turns one image into an explorable, camera-driven world
- xAI sues Minnesota to block its nudification law