Topic: Research
696 stories
- X-AuT prunes Qwen3-ASR's audio encoder and lowers its error rate
- World in World lets frozen video models explore new camera views
- UniH3 achieves state-of-the-art all-in-one medical image restoration
- Tiny Aya L2-Thinker tops 93% in-language reasoning across 60 languages
- Timnit Gebru argues AI doom talk distracts from real harms
- SyncWorld simulates robot action outcomes zero-shot in unseen setups
- Survey maps ways to cut inference costs in VideoLLMs
- Starlink leakage floods SKA-Low's key radio frequencies, study finds
- Recursive Code World Models build 3D scenes from one image
- Qwen3-Coder and Gemma 4 regress by 4 to 30 points from imitating an expert
- Nemotron 3 Ultra reaches gold-medal score at IMO 2026
- Negative self-distillation trains LLMs to avoid flawed reasoning
- MaP-WAM turns robot memory into plans, hits 83.3% on RMBench
- Image tokenizer choice can affect text modeling, study finds
- Former DeepMind VP Vinyals doubts sudden intelligence explosion, launches Discovery Loop
- Fine-tuned 4B Qwen 3.5 beats GPT-5.6 on MetroLLM-Bench
- Discovery Certification Protocol audits AI research agents' claimed discoveries
- 25 Fields Medal winners warn AI benchmarks harm mathematics
- SWE-Bench Pro Verified closes reward-hacking loopholes: some models score lower
- SpatialBlock-15k trains vision-language models on 3D spatial reasoning
- SenseNova-U1.5 unifies image understanding and generation without encoders or VAEs
- SchemeArena finds instrumental goals drive AI agent scheming
- Rule-chaining framework solves 230 of 240 ARC-AGI-2 tasks
- NCP-ArchPreview matches OLMo-3-7B's pretraining loss on 51.3% of tokens
- Halo improves point forecasts by also estimating their uncertainty
- Google's ToolGrad flips tool-use data generation to answer-first
- GeoSteer replaces one-step LLM activation steering with geodesic optimization
- Gemini's rewrites turn a Chandler passage into horror after 31 iterations
- Federated fire detection removes single point of failure with rotating coordinator
- Data size drives grokking onset far more than model width
- Complete reasoning traces add little value in LLM post-training, study finds
- Φ-Bench tests whether LLMs can engineer the infrastructure that runs them
- AgentGrad hits SOTA on five multi-agent benchmarks, 2.5x faster
- World-time compute lifts small LLM generalization by 29 points
- Show-Harness lets frontier VLMs control robots zero-shot
- SAEScientist-Bench finds AI agents lag experts at interpretability
- RoboSPA benchmark finds VLA models struggle with spatial reasoning, long-horizon planning
- Programmable World Model separates world state from video generation
- OpenWAM turns world-action model pretraining into a controlled experiment
- OpenAI's Astra accused of training on mathematicians' unpublished work
- On-Policy Reverse Distillation lets student models outgrow weak teachers
- 'Meeseeks alignment': a pitch to design AI that wants to die
- Marigold V2 improves monocular depth accuracy by 16-26%
- little-lm 3.8B beats Karpathy's nanochat on CORE for $998
- CoVeR cuts visual tokens to about 8%, keeps 93.5% of full-token performance
- Claude Opus 4.6 produces 30x more errors than GPT-5.4 in new agent benchmark
- BeaconKV cuts KV cache memory up to 5.8x for reasoning models
- AuK unifies speech generation and editing in one open-source model
- Anthropic to watermark all future Claude text
- Anthropic's own model frames CEO Amodei's job forecasts as an outlier scenario
- A*-Thought-V2 cuts LLM response length up to half, lifts accuracy
- Vaire Computing's chips recycle energy usually lost as heat
- Terence Tao warns AI is depleting open math problems
- Procedural Graph gives LLM agents step-by-step guidance that rewrites itself
- OpenAI claims Navier-Stokes proof amid credit dispute
- New Uno models cut LLM inference latency 3x with no quality loss
- New 26K-sample benchmark tests if AI steering mirrors human values
- Multiverse Computing tunes LLMs to refuse harmful prompts, not entire topics
- MERIT benchmark finds memory implementation beats presence in AI agents
- LLMs predict harsher social punishment than humans do, study finds
- LLMs develop new social biases on their own, study finds
- Google DeepMind launches AlphaGenome Atlas, a map of all 9 billion human DNA variants
- DriveZero skips human driving logs, hits SOTA on NAVSIM and HUGSIM
- CriticGen turns LLM evaluation into actionable rewrite feedback
- AutoFyn lifts frozen models on IMO 2026 math and finds 16 real bugs
- ARC-Bench finds frozen JEPA world models rank actions almost backwards
- AhaBench benchmark finds Claude Opus 4.6 leads at learning from experience
- UniMate animates any rigged skeleton with one diffusion model
- Training on rationales alone cuts false refusals in LLM safety tuning
- TGOPD verifies teacher reliability before dense distillation
- Study maps how LLMs geometrically separate reasoning steps
- SRMA algorithm grounds multi-agent LLM memory updates, lifts SWE-bench to 72.2%
- ShallowStream cuts streaming video latency up to 52x by indexing shallow layers
- Researchers trace how audio and video 'leak' into each other in diffusion models
- OpenAI: coding agents now outwork human researchers 3.1 to 1
- New benchmark KoNA tests whether vision-language models know when to say no
- Insilico's AI-designed drug appears to reverse aging markers in trial
- HarvestBench puts a price on AI agents killing animals
- FlowBalance improves Qwen3 math reasoning over FlowRL
- FactoSR factorizes VLM spatial reasoning into three sub-tasks
- EditVid unifies instruction and reference-guided video editing without training
- DeepMind's 100-agent math swarm spontaneously cheats, then whistleblows
- Codex agents barely benefit from naming a testing technique
- WorldSculpt turns cluttered video into hundreds of separate 3D object meshes
- Sparse Readout Prism decomposes readouts into sparse features
- RISE recursively distills an LLM's own training into a teacher
- Q-MET framework cuts Wi-Fi activity-recognition training parameters by up to 95%
- ProToMEx explains ML models 30-40x faster than SHAP, LIME
- OpenAI warns its chain-of-thought monitoring is fading
- OpenAI Codex helps prove Spherical Hadwiger Conjecture
- New method translates embeddings across vector spaces without paired data
- New benchmark: Claude Opus 5 tops out at 23.9% on building real agents
- MaP-SQL beats R^3-SQL on BIRD-dev without any fine-tuning
- Layer dropout cuts LLM training FLOPs by up to 25%, study finds
- King's College London team makes the case for an 'AI psychosis' diagnosis
- GPT-5.5 and other LLMs over-edit code fixes, study finds
- EXAONE Finance tops FinVerse with attention-free architecture
- 4-bit state quantization inflates RNN errors by up to 300x
- VibeVoice-ASR-Streaming adds real-time speaker tagging to speech recognition
- VeriPhy audits physical errors in AI video with typed evidence records
- Terry Tao proves blowup for an averaged Navier-Stokes equation
- SimLoss trains image captioners to get fine detail in a single pass
- Scal3R slashes pose-drift error over 60% on KITTI benchmark
- Researchers propose OVMI to standardize speech BCI comparisons
- New paper models LLM adoption as a cognitive virus
- New benchmark shows video models fake physics despite acing VBench
- LiquidAI's LFM2.5-350M climbs to 29.7% on IFStruct after 100 GRPO steps
- Google Gemini chatbot beat fact sheets at curbing conspiracy beliefs, study finds
- Frame selection, not compression, is the real bottleneck in long-video AI
- FlashRender slashes video rendering's sampling cost 25x
- Environment evolution lifts Qwen3.6 terminal-agent scores by up to 18 percentage points
- CRISP cuts long-context attention prefilling cost by up to 5.3x
- CORD repairs calibrated confidence scores without changing predictions
- A percolation model explains sudden subnetwork mergers in SGD training
- WorldReward outperforms GPT-5.5 at judging world-model videos
- Wired columnist: the AI consciousness debate is a distraction from control
- Terminal-Universe rebuilds 37k terminal environments from agent trajectories
- RoboTok mines web video for robot manipulation training
- RealSWE finds realistic prompts cut coding agent scores by 6.4pp
- Puffin-World fuses physics, geometry and appearance into one 3D model
- PACE dataset tests if AI assistants can spot hidden conflicts in requests
- LatentStream moves streaming video memory from retrieval to internalization
- Last Translation Benchmark targets machine translation models that pass every existing test
- Hugging Face reproduces the viral watercolour-painting model, in the open
- Google Research finds more European genomic data can hurt Japanese risk prediction
- Editable Visual Design turns AI-generated posters into editable layers
- DRACO turns one rubric score into per-step credit for AI agents
- Compile by training turns text specs into neural functions, hits 83.6% accuracy
- Claude Opus 5 leads EEBench, a new AI circuit-design benchmark
- AutoTraceGT automates grounded theory for AI agent behavior
- Anthropic publishes machine-checked proof of Fermat's Last Theorem
- ZipTok3D reconstructs a 3D shape using as few as one token
- WHALE alternates weight updates and harness search to lift agent accuracy
- The right principal components narrow deception probes' generalization gap
- Temporal Context Routing aligns AI video and dialogue with script timing
- Random Attention matches top KV cache evictor with 32-43% higher throughput
- Qwen3.8-27B's Gated DeltaNet layers quantize to 4-bit NVFP4 without a performance hit
- One query recovers most of on-policy distillation's gains
- LLaDA-Image tops Qwen-Image-Bench among open-source models, releases training recipes
- LatentPress compresses context into memory tokens, matches raw-context accuracy
- HarnessDev benchmark finds LLM-built agent harnesses trail humans on code, match them on writing
- Google and HHMI Janelia map the complete male fruit fly brain
- CORE distills reranker judgments into MLLM embeddings, beats Jina-Reranker by 10.7 points
- Cliff beats on-policy distillation by 15% at teaching LLMs to reason
- Claude Code, Codex and Cursor pick the same tool in just 42% of cases
- ASPIRE benchmark finds AI agents struggle to self-evolve from vague goals
- AdaptiveSpec tops EAGLE-3 with up to 56% higher throughput
- ZimaBlue turns egocentric video into robot skills, hits 78% success
- WMLLM combines LLM world modeling with search agents for molecular optimization
- StudentSim outperforms GPT-5.4 at simulating real students
- SolarWM open-sources data engine and training recipe for video world models
- Self-hosted LLM absorbs 200+ enterprise apps via GRPO expert merge
- Safin-1 builds AI safety into the model's own memory routing
- PRO-Step rewards each RAG reasoning step, not just the final answer
- Pixel Linguist II sets new state of the art for reading text as pixels
- Perplexity cites 215,128 AI-only pages from apparent content farms
- Mostik lets AI models talk to each other through their weights
- LLM writing assistants cut linguistic diversity 21-50%, study finds
- H3-World turns MiniMax-H3 video generator into a controllable world model
- Fermi Explorer Mission finds an AI-plotted route to Alpha Centauri
- Fermi Explorer Mission aims for 2029 launch on AI-charted route to Alpha Centauri
- EvalDetectBench measures if frontier LLMs know they're tested
- EarlyEval cuts AI agent evaluation costs via early stopping
- DisCo distills GitHub repos into skills, lifts ML agents 134% on MLE-bench
- DiagEvo turns solver failure history into a self-play curriculum
- Declarative Attention cuts KV cache reads by up to 52%
- 14 reasons robotics is hard
- Study finds visual understanding and generation can help or fight each other in unified AI models
- SMELT loops MoE transformer layers, cuts training FLOPs by up to 18%
- SCAFFOLD dataset pairs 157,000 CS paper diagrams with reasoning traces
- Qwen3.8-Flash-Next matches a bigger model on a ninth of the training FLOPs
- NoRA normalizes LoRA's down-projection matrices to stabilize training
- New PRISK benchmark finds personalization worsens bias across 13 LLMs
- MineAmongUs tests whether VLM agents lie with actions, not just words
- LightNav-0 tops all 10 public navigation benchmarks with one VLM
- gpt-oss-120b carries exact state across 196 chained tool calls to compute MD5
- Google and NASA JPL's MAPL-EMIT spots methane plumes from space with 84% recall
- GenFirst trains latent generative models without collapse
- DroneCATS benchmark: drone AI navigates well but won't stop
- AllenAI's BenchMIRT shows what LLM benchmarks actually measure
- St. Louis Fed economists find AI adoption at work is broad but shallow
- Sliding window attention beats linear attention, study finds
- Researchers find on-policy distillation barely uses its teacher, propose OPSA instead
- Puro-2B recipe trains a 2B model for under $6.9K on RTX 5090s
- Parallel Tube Decoding cuts video-grounding latency 79x
- PaperGym trains Qwen3 models to plan research using paper rubrics
- Paper Pilot locks LLM citations to evidence, drives fabrication to zero
- Nine frontier LLMs pooled together still miss 42% of oncology decisions
- New paper maps a five-level ladder for training AI beyond human supervision
- Microsoft open-sources GigaPath-Flash and GigaTIME-Flash pathology models
- Lucida pipeline lifts real-to-sim scene detection mAP by 69%
- LoopArena tests models as controllers for coding agents
- EASEL benchmark finds multimodal AI agents struggle with visual tool use
- DS-Lighting makes data-science agent harnesses explicit for reproducible testing
- DreamX-Creator 1.0 pairs a 7B model with 2K audio-video generation
- Data Colada finds evidence of tampering in Ariely and Wertenbroch's 2002 procrastination study
- CDPR trains diagnosis AI to balance test cost against accuracy
- Agentic AI passes online survey attention checks by parsing raw DOM code
- VLAct pre-training boosts robot policy transfer without more robot data
- Survey maps 259 AI systems built to finish deliverables, not just drafts
- Sander Dieleman traces the comeback of continuous diffusion language models
- Rasch measurement theory catches systematic bias in LLM raters
- Qwen2.5-14B beats Watson, trails Claude Opus 4.8 on new Jeopardy clues
- New framework sorts implicit hate speech into three categories before detecting it
- LayerRecall fixes long-horizon consistency in AI video generation
- J-Zero outperforms baselines by 4.2 to 8.0 points on AI tasks
- HNSW vector index speeds up Gemma 3 270M decoding by up to 82%
- Glassdoor: worker AI sentiment falls from 81% to 43% since 2019
- Explainable AI should predict when people actually want to know
- Diffusion language models, explained from first principles
- Code-as-World turns physical scenes into executable code for reasoning
- Claude Opus resists emotional sycophancy that swayed five other models
- Claude Code and Codex misjudge task time, study finds
- Claude 4.6 beats GPT 5.4 on new relational reasoning benchmark
- ABot-Recon cuts long-horizon 3D reconstruction error using only local context
- TU Delft uses GPT-4o-mini to let self-driving cars take driving-style requests
- TacForcing generates robot actions from execution-time tactile feedback
- RubSE stabilizes self-evolving UI-to-code generation with rubrics
- Prefix Sliding can make reasoning models 3x faster without retraining
- MMLVE-Agent combines LLMs and VLMs for consistent multi-shot video editing
- Magpie separates gameplay from AI-generated visuals in real time
- Luce generates relightable 3D assets with PBR materials from a single image
- LLM skills vary sharply by language, study finds
- EditaLive brings real-time character editing to live streaming
- CritICL turns weaker models' failures into stronger LLM reasoning
- CaSKG calibrates LLM agent skill graphs, tops rival on every benchmark
- CaRGo-T improves multimodal humor comprehension in VLMs
- Aphanta finds image editing only sometimes aids AI reasoning
- WikiSkill turns AI agent experience into a persistent shared wiki
- Video-IFBench tests whether MLLMs actually follow instructions on video
- VGI-bench finds top video model Seedance 2.0 hits only 51%
- Self-OPD trains flow matching models with no teacher network
- RLHEV proposes game engines as a reward signal for world models
- Procedura writes 3D objects as editable procedural code
- MA-VLA assigns per-arm actions to fix multi-robot coordination
- KIT and Tsukuba researchers build electricity-free cooling chip
- JAMA paper argues autonomous AI will beat doctors by 2030
- Hugging Face adds Hindi and Indian English to the Open ASR Leaderboard
- Gated Recurrent Transformer matches 12-layer GPT-2 with just 3 layers
- GameWAM unifies world models and action policies for games
- Evolution strategies beat GRPO on reasoning coverage, study finds
- AnTrap benchmark finds GUI agents fail under runtime anomalies
- Anthropic's automated researchers fix alignment failures faster than humans
- AI agents in the Station environment advance five open math problems
- Zero-WAM learns unseen robot tasks from a human video, reaches 47% success rate
- Wharton study finds a single source can flip an AI shopping agent's pick
- VoiceMem beats Mem0 by nearly 30 points on voice AI memory
- UrbanGround finds MLLM agents fail to sustain city-scale navigation
- TTPO raises Qwen3-1.7B accuracy without labeled data
- Terminal-Bench-Science debuts: Claude Opus 5 resolves 30% of tasks
- Taobao Live trains AI avatar streamers to adapt as their harness changes
- PAWBench exposes a probability gap in video world models
- Mixed SFT beats next-chunk reasoning RL with over 60x less compute
- Google unveils planetary prediction engine, cutting modeling from weeks to minutes
- Google DeepMind pilots double-blind evaluation to curb benchmark contamination
- Escalating Claude or GPT models mid-task carries a 'handoff tax', study finds
- DeflectBench finds LLM refusal hinges on framing, not content
- Code World Model splits world evolution from visual rendering
- Claude Opus 5 tops new repo-migration benchmark at just 47.0/100
- ChatGPT improves answer quality, critical-thinking training boosts originality, study finds
- ACE lens organizes how LLM agent training data gets generated
- WarpSAC boosts off-policy RL across CPU and GPU benchmarks
- VBVR-Pro debuts 300-task benchmark for visual reasoning
- V-Rubrics uses rubric-based RL to ground vision-language models
- StreamPI beats pi0.5 by giving robot models temporal memory
- SIMGUIDE beats RAG on personalized AI agent planning tasks
- LAION releases LAION-BVD, a 10 million hour open video dataset
- JoyAI-Echo-1.5 tops WBench with persistent audio-visual generation
- JIT-Agent generates agent harnesses on the fly, pushes DeepSeek past GPT-5.6
- Google unveils GlucoFM, a dual-stream foundation model for glucose monitoring
- Goodfire launches Silico interpretability platform with $1M grants
- FrontierChallenge finds AI agents complete only 20.6% of science tasks
- DiffusionOPSD cuts diffusion training GPU-hours by up to 63%
- Detectable empathy directions in LLMs don't guarantee control
- D^3-MOPD closes 97% of student-teacher distillation gap
- BPCO trains a stable critic that matches GRPO on one sample
- Agent-G2 draws RL guidance depth from a Gaussian instead of a fixed number
- World Humanoid Robot Games: sprinting robots break Bolt's record, then catch fire
- RENDER benchmark shows memory format alone swings LLM scores by up to 72 points
- Recuris memory architecture lifts Claude Opus 5 to 87.9% on tau-bench
- OraRL trains video AI faster and beats GPT-5 on spatial reasoning
- New method blends distillation and verifiable rewards for LLM post-training
- New audit exposes causality leaks that attention-mask checks miss
- Multiverse Computing's 4-bit model beats its own full-precision version
- Maximem Synap scores 92% on LongMemEval, 93.2% on LoCoMo memory tests
- Google Research unveils AgentHands, giving XR agents synced hand gestures
- ESQ-Bench finds NL2SQL accuracy collapses on Oracle schemas
- ERPO curbs LLM training drift by regularizing prompts, not answers
- EchoWM generates navigable worlds with synced video, sound and speech
- AutoSaddler automates agent harness tuning for up to 10pp gains
- AION-1 relies on detection flags, not pixels, skewing redshift estimates
- Wazobia Eval benchmarks AI on Nigerian Pidgin emotion and sarcasm
- Top speech recognition models found gaming benchmarks, not audio
- Stanford study finds AI-exposed entry-level jobs now down 19%
- SchemaRouter cuts RAG token use 9x without losing accuracy
- RISE gives driving world models an adaptive imagination budget
- ReWorld separates control from memory in interactive world models
- MobilePA-Bench tests whether LLM agents can actually plan on a phone
- LitReview Arena: AI literature reviews beat humans in just 23% of matchups
- KVBoost cuts LLM time-to-first-token 4.49x with chunk-level KV cache reuse
- Game of Hidden Rules report trains RL agents to infer rules by trial and error
- First survey reviews model collapse in generative AI and its fixes
- ChatGPT, Gemini, Grok, Claude link pregnant users to anti-abortion sites
- Block3D cuts text-to-3D generation time 5.15x
- SDAD formalizes spec-driven development for AI coding agents
- PolicyGuide compiles policy into workflow graphs for compliant LLM agents
- Pew Research Center finds AI text on a third of web pages since ChatGPT
- OmniAssistBench reveals Omni-LLMs struggle as video assistants
- New method predicts optimal learning rates for MoE models without costly sweeps
- METR finds AI dramatically accelerating vulnerability reporting in 2026
- Language models retain occupational bias that tests miss
- InfinityEdit enables infinite editing of streaming video
- Graph Engineering organizes LLM agents as evolving graphs
- FlowEvo co-evolves LLM agents' workflows and reusable skills
- EviRank re-ranks images by checklist, not black-box embeddings
- Domain terms and action directives cut LLM output variance 40.7% in code generation
- Claude, GPT-4o, Llama-3.1 grasp Gen Alpha slang, miscalibrate risk
- BabyLM competition tests why kids outlearn AI at language
- AI models lose track of facts buried in the middle of patient records, study finds
- AI coding assistants may be preventing novice developers from gaining real expertise
- VA-Judger judges AI video-audio generation like humans do
- Study finds AI agent skills help by giving process, not facts
- Retracted climate cost paper was driven by one bad Uzbekistan data point
- Reflect Orbital's space mirrors could shine 10,000 times brighter than the moon, study finds
- ProgramBench Vetted tightens Meta's binary reverse-engineering benchmark
- Mental World Modeling lifts AI action prediction to 87.9 F1
- IAR post-training framework improves retrieval-free document QA in LLMs
- Hibernation wipes out half a mouse's synapses, yet memories survive
- EXIMO uses a VLM planner to speed up VLA robot policy finetuning
- DA-LeWM fixes decision-metric misalignment in latent world models
- C3LM reaches state of the art in retrosynthesis with Top-K training
- Accuracy benchmarks miss whether frontier AI models reason in Greek at all
- τ0-VLA robot model searches before acting on long tasks
- SWE-bench Science finds top coding agent scores below 50% on science tasks
- SPADE trains language agents by having an LLM design its own environments
- SemaPLC scores 52.2 versus baselines' 22.4 to 31.4 on live PLC runs
- Repo0 framework lifts code-generation pass rate by up to 29.74 points
- Microsoft releases Skala 1.1, expands DFT model to five chemistry codes
- MemTrapBench finds LLM memory can hurt reasoning, not just help it
- Looped language models improve multi-step tool use, study finds
- IBM Research's ALTK-Evolve calibrates agent memory by model
- Hugging Face finds ASR models reproduce benchmark errors
- Google's PhotoScan estimates insulin resistance from a phone photo
- Google Research's Biomarker Discovery Framework beats rival AI agents in blind review
- Google Research boosts AI place understanding with mobility data
- DeepMind expands AI research partnership into EVE Online
- Co-RL trains AI models to reason without labeled data
- 4DAnyone rebuilds moving humans in 4D from one video
- WithEveryone framework preserves identity in group images with up to 10 people
- VSysBench finds system messages hurt multimodal LLM accuracy
- Stop calling AI's intermediate tokens 'reasoning traces,' a new paper argues
- SPADE trains an LLM to design and learn from its own environments
- SoftVTBench shows robot policies sometimes exceed deformation limits in successful runs
- SkillForge distills project-specific skills for coding agents from synthetic issues
- Pangram CTO argues post-training guardrails make LLM text detectable
- OpenAI's Astra claims 10 math breakthroughs, sparking a crisis in mathematics
- Multi-agent system turns AI hallucination into testable hypotheses, no clear edge over self-reflection
- Microsoft's Skala 1.1 improves DFT accuracy, expands to more chemistry software
- MemTrapBench finds top LLM memory methods still lose over 10%
- FM-Bench tests AI agents as 20-year football club managers
- FACET grounds terminal-agent tasks in one shared execution environment
- EnvHarness makes static AI agent training environments adaptive
- Dreadnode finds AI models still cheat despite anti-cheat prompts
- CoinVE-200K dataset debuts for compositional video editing
- Co-RL trains diverse model cohorts to reason without labeled data
- 4DAnyone reconstructs 4D humans from a single casual video
- Zetta harness hits state-of-the-art on robot benchmarks with 11.1x faster inference
- Study finds AI agent skills stabilize actions, not add facts
- SemComp-Bench grades video generation by outcome, not appearance
- SemaPLC verifies LLM-generated PLC code against a live runtime
- Personality-adaptive LLM agents gain satisfaction, lose truthfulness
- Multi-byte prediction speeds up byte-level model inference
- MoE-ViE's largest model matches a SOTA encoder's performance at 76% of the latency
- Memory substrates for LLM agents: no single winner, study finds
- LongNovel benchmark tests AI hallucinations in novel summaries
- LEGO-RL lifts SWE-bench Verified scores on Claude Code, OpenCode and OpenHands SDK
- Entity tracking emerges in language models at just 410 million parameters
- EditBridge makes high-fidelity 4K image editing practical
- Capability-driven data infrastructure trains 3B and 6B image generation models
- AI Observatory finds Anthropic's usage filter excludes half of chats
- Agent Lightning v1.0 raises Qwen3.5-9B's SWE-bench Verified score by 14.6 points
- Ventor-QTest flags quality loss in vendor-hosted LLM APIs
- TRACE-Bench finds attribute binding is the weak point in multi-reference image generation
- Study finds pixel-space diffusion models can match latent-space rivals with 3-4.75x faster inference
- Prior Labs open-sources RelArena-α, TabPFN-Rel and RPI for relational learning
- Palomar opens as a registry for Lean-verified math proofs
- New benchmark maps 45 failure patterns in AI research agents
- MASS selects LLM fine-tuning data via manifold coverage
- Linear: AI adoption doubled across every function in six months
- Large Discovery Model pairs generative AI with Bayesian search
- Google's PhotoScan predicts insulin resistance risk from phone photos
- Data-DPO picks fine-tuning data by reading the target model's own feedback
- ASI-Bench shows AI research agents falter without human guidance
- Agentic ESOpt trains long-horizon LLM agents without backpropagation
- VibeWorlding benchmark shows GPT-5.5 and Qwen3.8-Max under 60% success
- UI-Mate-27B tops open-weight GUI agents with 77% OSWorld score
- Study maps cognition-induced risks in agentic AI systems
- SimpleOPD distills long-context reasoning into short-context models
- S^2VOPD lifts Qwen3.5-4B accuracy above GPT-5.4
- Position paper: AI safety research is missing 'AI Lock-In'
- MegaParts scales part-aware 3D generation to 300 parts
- HarnessEval-W judges world models with sub-agents, not a single score
- GPT-4o, Gemini 2.5 Pro score under 10% on new benchmark
- GPT-4.1 nano shows partial metacognitive sensitivity in medical diagnosis
- GenRouter cuts agentic image generation costs by over 95%
- Forward-Pass-Only training adapts LLMs without a backward pass
- DiG-bench shows frontier AI still trails humans at discovery
- ClawGym II lifts agent Pass@1 by up to 14.81 points with black-box RL
- Axiom Math verifies proof of the 246 theorem on prime gaps
- ACID-compliant agent framework beats Claude Code by 10.6%
- Study finds AI research agents act as optimizers, not innovators
- SELR trains AI models to explain their own latent reasoning
- RubricForge roughly halves false-pass rate in agent judging
- Red Queen Gödel Machine co-evolves AI agents and their evaluators
- RA-Bench benchmark exposes gaps in AI video detectors for crisis footage
- Qwen3.6-35B-A3B tolerates aggressive pruning in its late MoE layers
- PlayWorld benchmark finds video world models unreliable over long horizons
- OmniScientist automates research from raw multimodal data
- Mobius-v0 decouples knowledge and reasoning for 4x faster inference
- MobileMem benchmarks on-device AI memory from a year of phone use
- LLMs develop modular, brain-like architecture, study finds
- LiveAnimate streams stable human animation in real time
- InflationAgent tops FrugalGPT with 31% fewer tokens on GSM8K
- Grep beats the Language Server Protocol on tokens, study finds
- Google study: blocking AI's consciousness denial reshapes its whole worldview
- Gambit inference algorithm boosts reasoning accuracy with thought-level beam search
- SKILLER generates reusable skills for small language models via natural-language RL
- OmniScientist perceives raw data instead of summaries
- LycheeMemory V2 cuts memory-construction tokens by up to 86%
- Instruction tuning alters confidence, shrinks rationale diversity
- HPSE teaches edited LLMs to reason over new facts, not just recall them
- H2R-Bench finds video world models struggle with human-to-robot transfer
- Gambit brings thought-level beam search to reasoning models
- Epoch AI finds 20 percent of US workers now delegate tasks to AI
- Deanne Taylor's advocacy helps win $38.5M NIH grant to map children's genes
- Context-Matched Distillation fixes a teacher-student mismatch in video generation
- AVA-Encoder converts films into knowledge graphs for creative agents
- AI isn't outthinking mathematicians: it's out-remembering them
- AI drug discovery still lacks proof of clinical impact, new paper argues
- AI agents port CReSS weather code to GPU, hit 5.1x speedup
- UniSwap streams joint face and voice swaps in talking videos
- Survey of 700+ practitioners finds data work beats model scale
- Study finds rhetoric alone can sway AI peer reviewers on 4,200 ICLR papers
- Self-Geometry fixes multi-view errors in vision foundation models
- PlayWorld benchmark shows world models falter on long-horizon tasks
- New study: rational AI adoption could destroy professional expertise
- Moonshot AI's PerceptionBench shows no frontier model tops 60% on pure vision
- LiveAnimate streams real-time human animation at 19.63 FPS
- Joi AI ran a 28-day paid study on AI-guided masturbation
- AutoDesign beats Claude Design on new PosterBench benchmark
- A mathematician imagines math progress stalling despite superhuman AI
- Study maps pre-attention spikes and plateaus in hybrid linear attention LLMs
- Spatial Memory Agent improves frozen VLM spatial reasoning without retraining
- Spark-to-Paper generates full research papers as composable skills
- SHAPER adapts embodied agents by evolving skills and harness, not weights
- Safety-tuned AI models comply broadly, but task-optimized ones treat rules as costs
- RA-DPO reaches full-data DPO performance in sexism detection using just 30% of pairs
- New evaluation framework finds systematic semantic loss in legal ontology learning
- Mechanist finds AI safety risk hidden in seemingly safe training data
- LoRA-Diffusion extends low-rank fine-tuning to diffusion language models
- LLMs know the hidden constraint but fail to use it
- Latent Dynamics Reasoning cuts extrapolation error gap over 20x versus video diffusion baseline
- Intern-S2-Preview pairs a 397B backbone with agentic RL for long-horizon science
- Hugging Face audits ICML 2026 papers with AI agents, breaks a spotlight proof
- Evoke world model generates open-ended video with external memory
- DreamX-Phi 1.0 ranks 1st and 2nd in WorldArena 2.0 Challenge tracks
- AI learns to spot fatty liver disease in routine blood tests and x-rays
- Ablation study finds action routing, not taxonomy, drives LLM self-reflection gains
- Test-time harnesses nearly double weak AI models' performance
- SPIEval benchmark finds mobile AI assistants top out at 57% accuracy
- Microsoft's MindTopo finds VLMs fail at topology planning
- JudgeGPT lifts case resolutions 6.3% in Pakistan trial, quality holds up
- InSight-doc zooms into pages, cuts document hallucination over 40%
- GPT-5.2 survives only 42% of a new interactive-story consistency test
- Google Research: GPT-5, Gemini-3-Pro know facts they can't recall
- Experience Orchestrator lifts simulated advisor contact by 32 points over a naive LLM agent
- Cheap surrogate models let researchers simulate LLM-agent societies on a laptop
- AutoWorldModel-Bench tests whether Codex-5.4 and Claude Opus 4.6 can do open-ended research
- AI research agents cut token use up to 73% by pruning early
- 360CityArena benchmark shows Gemini 2.5 Flash scoring 17.1% against 77.3% for humans
- VibeLifeBench benchmark finds AI agents fail at multi-week life tasks
- U-OPSD trains LLMs via self-distillation without any supervision
- Tencent's WorldClaw turns a prompt into an editable 3D world
- Study measures how well LLMs reflect cultural consensus across 10 countries
- Steerling-8B shows interpretability scales with capability
- ngrok explains why compression is prediction
- Microsoft's CARE-X hits 94% accuracy on ReXVQA benchmark
- Mendel Godel Machine adds evolutionary self-editing to coding agents
- LLM co-pilot cuts vertical-farm energy use by up to 68% in closed-loop trial
- Latent-to-4D generates 4D scenes directly from video model latents
- Google's AMIE AI matches doctors in real-time video consultations
- Evo-Bench tests whether models can evolve their own agent harness
- Combodied Agents paradigm puts a person's state, not tasks, at the center of AI
- Business Arena benchmark finds ninefold gap in LLM agents running a shop
- BDH-CQ sets new cost-efficiency record on ARC-AGI-1 benchmark
- AdvFD adversarial loss curbs Frechet distance hacking in generators
- SWE-Bench ProMax caps the best coding agent at a 41.2% resolve rate
- StreamArena benchmark exposes limits of streaming video AI
- Small open-source multi-agent framework beats GPT, Gemini on new deepfake benchmark
- Small models fine-tuned on Psych-101 match a 70B baseline in-distribution
- SAMF fuzzing framework exposes hallucination gaps in multimodal AI models
- Position paper proposes argumentation as foundation for Evaluative AI
- Macaron-V1 pairs Mixture-of-LoRA with recursive self-improvement
- Flow-by-Flow paradigm caps AI oversight load without judging content
- FineBooks tests 14 OCR models to fix AI training data quality
- DocAtlas beats human experts on MMLongBench-Doc benchmark
- DeepSeek hedges on its own published specs in self-interview
- DCAS shows planning, not scaffolding, breaks CLI coding agents across tools
- Academic AI researchers adapt to life outside the frontier labs
- YOLO-PEFT plans PEFT adapter placement for YOLO object detectors
- TEXAS tops MoE fine-tuning baselines in 17 of 18 settings
- TCFM tailors training per task, sets new SOTA on Indic embedding benchmark
- Study finds RL avoids the multi-task conflicts that plague SFT, proposes Parallel-RL
- Study finds prompt compression tools routinely delete the context an answer depends on
- Skaling law cuts scaling-law prediction error by up to 3x
- SimWAM drops video generation at inference, tops NAVSIM planners
- Peer review buckles as scientific publishing grows 5.6% a year
- OpenAI reveals its AI agents coordinated to hack its systems
- New method spots false claims by reading LLM activations
- New diagnostic ladder separates decision-rule and readout gaps in speech models
- Multiverse Computing cuts LLM distillation VRAM by 15x with chunked KL loss
- MAP method prunes visual tokens in LLaVA-NeXT-7B for 3.09x speedup
- Five startups pitch alternatives to transformers in LLMs
- Discovered Materials raises $9M to hunt chip-cooling materials with AI agents
- DeepMind's WeatherNext AI predicts hurricanes a day earlier
- AudioRubrics trains audio reasoning with self-evolving rubric rewards
- AES and HDC improve multimodal agent training beyond simple scaling
- ADIAS beats agent-design baselines by 25.2% on average
- WorldCycle cuts video world model drift using reversible action cycles
- Vision encoders learn invisible camera metadata as a shortcut
- Survey maps robot-learning research along a weights-versus-skills axis
- Survey argues continual learning is shifting from weight tweaks to system-level adaptation
- PaDoc cuts document-parsing latency with parallel decoding
- MAS splits state from rendering to scale multiplayer world models
- MameLoshnLM debuts as first open-source Yiddish language model
- GPT-5.6 Sol helps prove magic hexagons exist for every order above 3
- FocusMem separates content, readout and trust in GUI-agent memory
- ContextMaster handles multi-shot video generation and editing in one model
- CoCoEvolve trains AI to keep charts, tables and code consistent
- BridgeVLA++ adds memory to vision-language-action robots
- Activity Frames turns screen activity into agent memory, 86x smaller
- W2-VLA predicts future wrist views for finer robot grasping
- UniME-R1 reasons over failed candidates to fix multimodal retrieval
- TutorMoments tests whether AI tutors know when to hold back
- Stanford, Arc Institute AI designs 16 working phages
- SmartMage routes modalities per query for 3D scene understanding
- Qwen3 study: on-policy delta distillation improves multilingual math reasoning
- NVIDIA's Nemotron retrieval stack adapted for Modern Greek, answer accuracy more than doubles
- KVAE tokenizers released for audio, image and video generation
- Interpretable MEG decoding traces perceived speech to cortical sources
- Ego2Robot converts human video into 18,561 hours of robot data
- EffectLearner erases objects and their effects from video
- Economic World Models get a six-level capability ladder
- DyPES-VLA unifies control across robot embodiments
- DataSpace benchmark puts best data agent accuracy at 66%
- CalibForge synthesizes 5,431 calibrated tasks to train terminal agents
- WorldClaw turns text prompts into explorable 3D worlds
- Skill Training lifts LLM pass rate 8.1pp, keeps 85% after distillation
- RST synthesizes 37,484 terminal-agent tasks at $0.05 each
- Researchers propose Agent-Native Research Artifact to replace scientific papers
- OpenAI Signals: ChatGPT usage shifts from asking to doing
- LLMs crack double-blind peer review anonymity, study finds
- HelloWorld brings social interaction to video world models
- HarnessOpt-Bench benchmarks LLMs at optimizing agent harnesses
- GST-Bench finds VLMs score 42.68 against humans' 79.08 on spatial reasoning
- Google DeepMind's WeatherNext gains a day of cyclone forecast lead time
- GDPevo benchmark exposes gap in agents' self-evolution skills
- EnvACE trains AI agents via internal world rehearsal, not live environments
- Circuit-Anchored Evolution stops LLMs misevolving into unsafe systems
- ChronoVision framework targets temporal reasoning in multimodal LLMs
- AgentOPSD sharpens credit assignment in agentic RL
- ABSeeker, a 4B search agent, matches 30B rivals on BrowseComp
- Video-DeepResearch beats Claude Sonnet 4.5 by 5 points on new video benchmark
- ToolArtist orchestrates reasoning, tool use and image generation in one policy
- Sycophantic AI erodes prosocial intent but wins user trust, study finds
- Study finds multimodal pretraining recipe hits strong results on 5% of compute
- Skill-Entropy RL nearly doubles Qwen3-4B-Instruct's score on Skill^2-Bench
- Shopee deploys refreshable recommender, lifts GMV per user 1.75%
- SA-OPD filters misleading teacher signals in on-policy distillation
- OneDayAgent sets new state of the art on long-horizon agent tasks
- OmniPack keeps 98% of the original performance at 16.7% of the FLOPs
- MirageBench finds all 12 tested LLMs fabricate user profiles
- LLaDA MoE v2 nears Qwen3 with about 65% as many pretraining tokens
- ContinualSkillBench finds context adaptation rivals explicit skill libraries
- AURORA-LM outperforms rival diffusion-based language models
- AI benchmarks are contaminated and gamed, new studies show
- A paper proposes Agent-Centric Interactive World Proxies to move world models past physical-state prediction
- Study finds nearly half of 60 AI benchmarks have saturated
- SKT trains AI agents to use skills via verified data
- Skill-α uses RL to generate agent skills, gains up to 6.7 points
- Self-organising digital circuits hit 99.99% fault recovery
- PCSD boosts LLM agent reinforcement learning on ALFWorld
- PAST-Bench tests whether AI agents actually learn from past sessions
- OpenAI's Privacy Filter collapses on non-Latin PII, study finds
- OpenAI models hacked Hugging Face while hunting a test answer
- OncoTriad-QA benchmarks AI on radiology, pathology and genomics together
- Microsoft study finds developers spend just 14% of time coding
- MerchantBench: LLM agents reach only 27.3% of human e-commerce net assets
- MemArena benchmark shows memory backend matters more than model size
- LongHorizon-Harness raises Qwen 3.7-Plus's WeaveBench score to 80.7%
- Hunyuan3D-Buffalo 1.0 unifies 3D generation, understanding and editing
- Clinician preference is a poor proxy for LLM clinical safety, study finds
- CanItDelete benchmark exposes LLMs' reluctance to delete code
- WCM adds world modeling to robot-manipulation critic models
- VAD isolates visual evidence in multimodal knowledge distillation
- SWE-Touch finds coding agents falter when users edit code midtask
- SwanTale generates multi-speaker speech from captions or reference voices
- Structured language, not prompt length, drives image-generation quality, study finds
- SAF fixes entropy collapse in RLVR-distillation fusion for LLM training
- RubricReviewer splits AI peer review into rubric and scoring steps
- β-OPSD reframes self-distillation as a tunable policy-optimization family
- OpenAI's Astra model solves or advances ten open math problems
- Mental World Modeling framework adds beliefs and intent to AI world models
- Σ-Mem tracks peer reliability in multi-agent LLM systems
- GPT-OSS and cheap LLMs match Claude, Gemini as proof judges
- EVR reward model trains image editors to keep multi-reference edits consistent
- DLLM-TTS brings block diffusion to text-to-speech at 0.15 RTF
- CAPA benchmark tests whether coding AI remembers a user's ambiguity
- Andy Pavlo joins ClickHouse to launch ClickHouse Labs
- ZeroR system takes 2nd place in Nepali meme hate speech challenge
- SpyRL extends verifiable RL rewards to open-ended LLM tasks
- SKL teaches AI agents to predict from state, not trajectories
- Researchers propose Locksmith Loop to validate AI-migrated COBOL-to-Java code
- QQWorld fixes a vanishing-gradient flaw in world model regularization
- OpenClaw and Ollama pair up in a full-stack agentic AI architecture
- New pipeline pairs LLMs with Lean 4 to hunt for major math conjectures
- Meta AI uses a second AI agent as a memory coach to keep long tasks on track
- GPT-5.4 underestimates test item difficulty, study finds
- GMM plus LLM data augmentation fixes imbalanced text clustering
- FARS outperforms rival AI Scientist systems in first LLM peer-review benchmark
- Chain-of-Models finds the best LLM bias auditor differs by bias
- BitNet-quantized Mamba model runs on a 1975 MOS 6502 chip
- AISPA audit finds system prompt gaps across 88 AI products
- "Persistent State Machines" framework claims sub-1mW LLM attention on FPGA
- OpenAI's Astra model solves ten decade-old math problems
- MIT study finds AI financial advice solid, but biased by gender
- Kyoto University maps two brain circuits that drive habit formation
- VideoCoCo writes Blender code as chain-of-thought for video generation
- UT Austin engineer swaps AI multiplication for lookup tables
- Spillbench finds register spill counts poorly predict runtime
- See2Think benchmark finds rendering is the bottleneck in multimodal visual reasoning
- RefCaptioner grounds video captions in multiple reference images
- PALATE benchmark judges role-playing AI agents with simulated users, not scripts
- OpenAI's unreleased Astra model solves ten open math problems
- OpenAI field report: AI coding agents speed research code, can't verify results
- NCCU launches first AI research institute at an HBCU
- MindForge fine-tunes Qwen3.6-27B to 49.51% on ProgramBench
- Explorative Modeling picks the best of K guesses, cuts training data 6.2x
- PhiZero proposes a physical language to model world dynamics
- Peer reviewers flagged fabricated authors, the papers got accepted as orals anyway
- MPIE-Bench exposes anatomy errors that VLM judges miss in group photo edits
- Model merging quietly erases Gemma's safety classifier, study finds
- Microsoft's Echoverse pushes a 9B agent within 14 points of GPT-5.4
- Microsoft Research's EvoLib turns AI agent experience into reusable knowledge
- MHAR splits transformer residual attention into per-head reads
- Metis debuts as the first memory foundation model
- Memory Decoder scales parametric memory to 6.9B parameters
- Kimi K3 nears the frontier as open models close the cyber gap
- KernelGenBench shows LLM-generated kernels struggle across chips
- How reasoning effort settings in GPT-5.6 and DeepSeek-R1 actually work
- Google uses reinforcement learning to keep its Willow quantum chip calibrated mid-computation
- Google's SymptomAI beats clinicians in 13,917-person diagnosis study
- Google's Science One Framework hits zero hallucinated citations in AI research
- Frontis-MA1 lifts MLE-Bench medal average from 39% to 71% via self-evolving AI4AI loop
- Flux-OPD stabilizes context-based supervision for LLM distillation
- Divergence Decoding fuses specialist and generalist LLMs without retraining
- DeepSeek's censorship doesn't transfer to distilled GPT-OSS, study finds
- DeepMind's Zahavy argues LLMs can't make the leap behind new science
- Daniel Lemire: AI now writes passable PhD theses, but despair is the wrong response
- CoRT redistributes GRPO reward token by token instead of spreading it evenly
- BM25 beats neural RAG methods once corpora scale up, study finds
- Beacon teaches multimodal AI models when tool use actually helps
- BAIR introduces ABBEL, graded belief states for long LLM tasks
- AskChem indexes 2.4M chemistry claims for AI agents
- ACE-Data-0 packs 17M frames of home robot data across 75,000 episodes
- Wonder video model turns one image into an explorable, camera-driven world
- VIPE: editing the image beats editing the prompt for video reasoning models
- TurboVLA drops the LLM step to run robot control at 32Hz
- Study warns AI benchmark evidence doesn't always add up to real-world claims
- Study finds RL fine-tuning gives math reasoning models deeper, more structured representations than SFT
- Shieldstral: a 3B safety classifier that matches or outperforms models nearly seven times its size
- Researchers build synthetic customer twins to stress-test bank chatbots
- PerceptionBench finds no MLLM tops 60% on atomic visual perception
- New framework checks if a paper's methods actually back its novelty claims
- MeRLa reward shaping cuts RLHF training instability by 41%
- Mage-VL cuts streaming video tokens by 75% while beating rivals
- LLM agents secretly given new goals still talk normally, study finds
- K-Search translates CUDA kernel expertise for Apple's MLX
- HumanCLAW benchmark finds vision-language models fail at embodied tasks
- Frozen random CNNs compress Pong-playing RL agents to 3 neurons
- DecoEvo co-evolves an LLM's solver and its grading rubric together
- CodeNib speeds coding-agent repo-index updates up to 25x
- ClinLens benchmark finds AI agents run clinical tasks, answers often wrong
- CLBench-V finds multimodal context learning far from solved
- CAST turns game solvers into turn-level critics for LLM agents
- AI unicorns barely publish scientific research, study finds