Topic: Models
106 stories
- YuE2 outperforms Suno v5 on WildSongBench music benchmark
- GPT-6 Astra: Sebastian Raschka digs into the looped-transformer rumor
- Cognition ships SWE-2, nearing GPT-6 Astra at a quarter of the cost
- Claude Fable 5.1 drops hedges and stock phrases, but answers grow longer
- Suno launches v6 music models built with Warner, BMG, and Believe
- IBM releases Granite Time Series PatchTST-FM-r2, tops GIFT-Eval among permissively licensed models
- GPT-6 Astra cuts unintended actions 89% versus GPT-5.6 Sol
- Desert Ant Labs launches 18 on-device AI models with one SDK
- DeepSeek ships V4.1-Flash with native visual understanding
- llm 0.35 adds support for OpenAI's GPT-6 Astra
- Inception Labs launches Mercury 2.5 diffusion LLM, 40% smarter than Mercury 2
- Gander releases an open real-time multimodal agent model
- Why AI-generated food images look so wrong
- Falcons AI's NSFW classifier ranks 6th on Hugging Face with 50.8M downloads
- NeoMME ships a single-tower encoder with about 2x retrieval speed
- Artificial Analysis releases Intelligence Index v4.2
- OpenAI launches GPT-6 Astra, its most capable and aligned model yet
- NeoMME encoder matches ColQwen2.5 in document retrieval with 14× fewer parameters
- IFM releases K2 Horizon, six open models from 0.9B to 375B parameters
- Google DeepMind's WeatherNext 3 forecasts hourly at up to 5km resolution
- Multiverse Computing launches Quasar 438B, Europe's leading model
- Meta ships Muse Spark 1.3 for agentic coding
- Google releases Gemini 3.8 Flash and Flash Cyber
- Anthropic's Fable 5.1 system prompt cracks down on song lyrics, copyrighted art
- World Labs launches Atlas, a spatial world model
- Google adds agentic video understanding to Gemini Flash models
- CogEvol open-sources a 4B model that generates course slides and interactive lessons
- Anthropic ships Claude Fable 5.1 and Mythos 5.1, cuts prices up to 45%
- Google ships TimesFM-3, a multivariate forecasting model
- Gemini Omni 1.1 Flash adds scene extension and keyframe control
- Tencent releases Hy4 preview, a 770B/49B open-source LLM with 1M+ tokens of context
- Qwen releases Qwen3.8-Flash-Next, an early preview of Qwen4
- GLM-5.3 ships open-weight, tuned for coding and long tasks
- gpt-5.6-luna shows small models are cheap enough for consumer AI
- Google ships Gemini Omni 1.1 Flash with 40-second scene extension and 4K upscaling
- Z.ai ships GLM-5.3-Flash, nearing Claude Opus 4.8 at one-tenth the cost
- Google DeepMind ships Gemini 3.5 Transcribe with sub-second streaming
- Altman says OpenAI will have AGI by end of 2026, on his own definition
- Alibaba's Qwen3.8-Flash-Next beats bigger rivals at a fraction of the cost
- Tencent's WeMM-Embedding beats 8B rivals with a 2B model
- IBM ships Granite 4.2 reasoning models in three sizes
- Thomson Reuters spends $40M to build its own legal AI on Qwen
- Drew Breunig says Fable's cost pushed teams back to cheaper models
- Alibaba's Wan3.0 doubles AI video length to 30 seconds
- OpenAI cuts GPT-5.6 Sol pricing 20% on input, 33% on output
- Deepseek releases V4-Flash-Vision-Exp, claims it rivals Opus 4.8 on agent benchmarks
- Ox Alpha, an anonymous stealth model for coding, debuts free on OpenRouter
- Kimi K3 and GLM-5.3 close in on Opus 5 and GPT-5.5
- DSpark speeds up Liquid AI's LFM2.5 inference up to 3.2x
- DiffusionGemma averages about 1,500 tokens per second with parallel diffusion
- Ornith releases Ornith-1.5, an open model on par with Claude Opus 4.8
- Liquid AI ships QAD-quantized LFM2.5 checkpoints, recovers 97% of lost accuracy
- Generalist AI's robots learn new tasks from a single video
- Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index
- Mimir v1 matches larger models at 1B parameters using only permissible post-training data
- GPT-5.6 Sol jumps to 46.2 mAP on Roboflow's vision benchmark, up from 13.8
- Tim O'Reilly: big AI labs built an architecture of control
- Qwen 3.8 27B overthinks everything by default
- Hugging Face finds Qwen has become open source AI's base model
- Google releases Gemini 3.7 Flash, halves the price of 3.6
- Claude's system prompts reveal new Mythos tier above Opus 5
- big-pickle stealth model scores 50.8% on SWE Atlas, trailing only Claude
- Tim O'Reilly: AI labs misread what people actually want
- Qwen releases Qwen3.8-27B-FP8 vision-language model
- Opus 5 is more capable, but feels worse to work with
- Meta ships open model Glimmer as Zuckerberg pitches 'AI for everyone'
- Hugging Face: Qwen becomes the base model of open source AI
- Apple trained a custom China AI model with Alibaba
- Writer launches Palmyra X6, harness upgrade to cut AI costs
- Netlify compares Claude, GPT, Gemini, Kimi and GLM on credits and quality
- Mistral OCR 4.1 ships with bounding boxes, block labels
- Google ships Gemini 3.7 Flash at half the price of 3.6 Flash
- xAI ships Grok 4.6, matches GPT-5.6 Sol on benchmark index
- Qwen ships Qwen3.8-2.4T-A95B, the first open Qwen-Max-class model
- Liquid AI ships LFM2.5-VL-3B, a vision-language model for edge devices
- DeepSeek V4 Pro 0813 reaches GA with 1M context, $0.435/$0.87 pricing
- Nvidia ships Nemotron 3.5 Lightning and NeMo Switchyard
- NVIDIA expands Magpie TTS to 12 languages with open weights
- Needle 2 packs a 45M-parameter tool-calling model into a 14MB binary
- Motif 3 debuts as a 314B-parameter MoE model with new GDLA attention
- Liquid AI's LFM2.5-2.6B matches models four times its size
- Meta releases Muse Glimmer, a 30B open-weight local-agent model
- Google DeepMind retrofits Gemma 4 into DiffusionGemma
- xAI ships Imagine Image 2.0, ranks second behind GPT-Image-2
- LG releases K-EXAONE 2.0, a 750B-parameter open-weight MoE model
- DOE launches Genesis-Science-1, its first open-weight science model
- DeepSeek V4 Flash 0731 scores 89.0% on ARC-AGI-1, 61.4% on ARC-AGI-2
- OpenAI updates GPT-5.6 Sol, expands Luna for free ChatGPT users
- MiniMax-H3 video model ported to MLX, runs on Apple Silicon
- Liquid AI ships LFM2.5-2.6B for on-device agents
- DiffusionGemma generates 1,500 tokens per second via parallel diffusion
- DeepGrove ships Maple-Preview, a 20B ternary model built for on-device speed
- MiniMax ships H3 open-weights video model with day-zero ComfyUI support
- Qwen3.8-Max arrives as Qwen's first open-weight Max-class model
- Karpathy has Opus 5 render Lord of the Rings in three.js
- Claude Opus 5 draws a frog with a Habsburg jaw for a personal AI benchmark
- Google DeepMind ships Gemini Robotics ER 2 with video-based task tracking and multi-robot teamwork
- ByteDance ships Seedance 2.5, a 30-second video generation model
- OpenAI cuts GPT-5.6 Luna price 80% in full-stack AI push
- Liquid AI ships LFM2.5-Encoders for fast CPU inference
- Google DeepMind launches Lyria 3.5 music model in Flow Music
- Google DeepMind launches Gemini Robotics 2 for whole body control
- DeepSeek releases V4-Flash-0731 with top value per dollar
- OpenAI cuts GPT-5.6 Luna price 80%, Terra by 20%
- Google DeepMind unveils Gemini Robotics 2 for whole-body control
- OpenAI cuts GPT-5.6 serving costs 20% with post-launch efficiency work