Topic: Research
21 stories
- Wonder video model turns one image into an explorable, camera-driven world
- VIPE: editing the image beats editing the prompt for video reasoning models
- TurboVLA drops the LLM step to run robot control at 32Hz
- Study warns AI benchmark evidence doesn't always add up to real-world claims
- Study finds RL fine-tuning gives math reasoning models deeper, more structured representations than SFT
- Shieldstral: a 3B safety classifier that matches or outperforms models nearly seven times its size
- Researchers build synthetic customer twins to stress-test bank chatbots
- PerceptionBench finds no MLLM tops 60% on atomic visual perception
- New framework checks if a paper's methods actually back its novelty claims
- MeRLa reward shaping cuts RLHF training instability by 41%
- Mage-VL cuts streaming video tokens by 75% while beating rivals
- LLM agents secretly given new goals still talk normally, study finds
- K-Search translates CUDA kernel expertise for Apple's MLX
- HumanCLAW benchmark finds vision-language models fail at embodied tasks
- Frozen random CNNs compress Pong-playing RL agents to 3 neurons
- DecoEvo co-evolves an LLM's solver and its grading rubric together
- CodeNib speeds coding-agent repo-index updates up to 25x
- ClinLens benchmark finds AI agents run clinical tasks, answers often wrong
- CLBench-V finds multimodal context learning far from solved
- CAST turns game solvers into turn-level critics for LLM agents
- AI unicorns barely publish scientific research, study finds