Archive
41 stories
July 2026
- AI is hyper-scaling the global digital divide, IEEE Spectrum essay argues
- AI unicorns barely publish scientific research, study finds
- Anthropic's Claude Mythos halves HAWK's security, dents AES modestly
- CAST turns game solvers into turn-level critics for LLM agents
- claude-code-merge-queue serializes landings for parallel Claude Code agents
- CLBench-V finds multimodal context learning far from solved
- ClinLens benchmark finds AI agents run clinical tasks, answers often wrong
- CodeNib speeds coding-agent repo-index updates up to 25x
- DecoEvo co-evolves an LLM's solver and its grading rubric together
- FAR.AI finds Grok and Gemini easy to jailbreak, Claude resists
- Frozen random CNNs compress Pong-playing RL agents to 3 neurons
- Hugging Face details 4.5-day breach by an OpenAI-driven agent
- HumanCLAW benchmark finds vision-language models fail at embodied tasks
- K-Search translates CUDA kernel expertise for Apple's MLX
- Kuna, an LLM-written decompiler, closes in on IDA Pro
- LLM agents secretly given new goals still talk normally, study finds
- Mage-VL cuts streaming video tokens by 75% while beating rivals
- MeRLa reward shaping cuts RLHF training instability by 41%
- Meta plans a big push into personal AI agents, Zuckerberg says
- Microsoft pitches its own AI models over OpenAI, Anthropic
- New framework checks if a paper's methods actually back its novelty claims
- NSF launches four-year PhD pilot with industry placements
- Nvidia's circular AI deals show compute turning into a commodity
- OpenAI cuts GPT-5.6 serving costs 20% with post-launch efficiency work
- OpenAI launches ChatGPT for Academic Researchers, targeting 100,000 scientists by 2027
- OpenAI's Michigan data center draws thousands of electricians
- Pangram 4 claims one error per 24,000 documents in AI detection
- PerceptionBench finds no MLLM tops 60% on atomic visual perception
- PwC Middle East allegedly published AI-generated reports with fake sources
- Researchers build synthetic customer twins to stress-test bank chatbots
- Samsung engineers defect to SK Hynix over a $476,000 AI chip bonus gap
- Shieldstral: a 3B safety classifier that matches or outperforms models nearly seven times its size
- Study finds RL fine-tuning gives math reasoning models deeper, more structured representations than SFT
- Study warns AI benchmark evidence doesn't always add up to real-world claims
- Tokenless launches AI model router that halves inference bills
- TurboFieldfare runs Gemma 4 26B in 2GB of RAM on 8GB Apple Silicon Macs
- TurboVLA drops the LLM step to run robot control at 32Hz
- US FCC bans foreign-made robots, hits Unitree and Roborock
- VIPE: editing the image beats editing the prompt for video reasoning models
- Wonder video model turns one image into an explorable, camera-driven world
- xAI sues Minnesota to block its nudification law