Strata runs 125B Qwen3.8-Flash-Next on a 12 GB gaming GPU

Strata is a free, open-source tool (MIT License) that runs Qwen3.8-Flash-Next, a 125-billion-parameter model from the Qwen team, on an ordinary gaming PC. The README says models of this size usually run on servers with hundreds of gigabytes of graphics memory, while a consumer card has 12-24 GB. Strata closes the gap by sharing the work across the whole machine. It needs an NVIDIA or AMD graphics card with 12 GB or more, on Windows or Linux. The model chats, writes code, reads pictures and works with apps and coding agents, and the README states that nothing leaves the PC.\n\nHow it fits. The README explains it with a kitchen: what you use all the time stays on the counter, the rest waits in the pantry. The model is a team of 24,576 small specialists ("experts"), and each word needs only 10 of them. The graphics card keeps the few thousand experts used most often. RAM holds all of them, and the processor works on the rest at the same time. The SSD holds a big lookup table. On top of that, a small helper model guesses the next few words and the big model checks them all at once; the README says this gives the same answer 1.6-1.8x sooner. Long prompts are read in pieces of up to 8,192 tokens at a time, at over 1,000 tokens per second.\n\nSpeed. The authors say they measured Strata on two ordinary gaming PCs, with full tables in a separate DETAILS.md file. The figure in the Hacker News title is not a measurement for an RTX 4090. The README's 100-140 tokens per second is a projection for an RTX 3090 (24 GB), worded as "should", and it adds that a card with more VRAM is faster. For scale, it notes that 60 tokens per second is faster than you can read. The experimental UD-Q4_K_XL size, the closest to the full model, writes only 7-8.5 tokens per second on a 64 GB PC because most of it is read from the SSD. The README's demo is a voxel pagoda garden built from a one-shot prompt on an RTX 5070 (IQ3_S, 128K context).\n\nSetup. You download and unzip Strata (or git clone it), then double-click START-HERE.bat on Windows or run ./setup.sh on Linux. The installer finds the card, sets up the right engine and asks which model and size, how much context, and whether to read pictures. Pressing Enter takes the recommended answers. It then downloads the model (about 70 GB), resumes if the download stops, and opens the Strata app at http://127.0.0.1:8080. While the model starts, the PC can be slow or unresponsive for 1-3 minutes (longest the first time), because Strata loads 35-55 GB into RAM and locks part of it for the graphics card. Users can also paste a prompt into an AI coding assistant (Claude Code, Cursor, Codex, GitHub Copilot) that follows docs/AI_SETUP.md, and AI tools can install, start and stop Strata through its MCP server.\n\nModel sizes. The same model comes in several compression levels; smaller is faster, larger is a bit smarter. The README lists: Coder, a coding version with half the experts removed that reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM, but is weaker outside code, including Chinese and other CJK text; Swift 1.5, a fine-tune that thinks for a much shorter time at about the same quality; Unsloth UD-IQ4_XS, a ~4-bit version with a 94 GB download that reads part of itself from the SSD with less than ~80 GB of RAM; the experimental UD-Q4_K_XL; and OrcaRouter's Uncensored IQ3_XXS, which is set up by hand and is not in the installer menu. The model was compressed by ISTA-DASLab, UkisAI (Swift 1.5) and Unsloth, and Strata uses parts of llama.cpp / ggml.\n\nUsing it. In the browser the app offers Chat, a live Monitor of the model and the GPU/CPU/RAM, and an About page with settings. Other apps connect through an "OpenAI-compatible" provider at http://127.0.0.1:8080/v1, through an Anthropic-style endpoint at /v1/messages (for Claude Code, ANTHROPIC_BASE_URL=http://127.0.0.1:8080), or through /v1/responses for Codex CLI. Thinking can be set to off, low, medium or high. By default Strata answers one request at a time; setting "parallel": 2 serves several, but on a 12 GB card each answer gets slower. The first message of a chat is read in full at about 1 minute per 30,000 tokens, and follow-ups start in seconds. Experimental, community-tested support covers older graphics cards, Intel Arc (built from source on Linux) and older CPUs without AVX2.
Key facts
- Strata is a free, MIT-licensed tool that runs the 125-billion-parameter Qwen3.8-Flash-Next on a Windows or Linux PC with an NVIDIA or AMD card of 12 GB or more.
- It spreads the model's 24,576 experts across GPU, RAM, CPU and SSD; each token needs only 10 experts, and speculative decoding gives the same answer 1.6-1.8x sooner.
- The 100-140 tokens per second figure is a projection ('should') for an RTX 3090 (24 GB), not a measured RTX 4090 result.
- The default install downloads about 70 GB and loads 35-55 GB into RAM, and the PC can stall for 1-3 minutes at startup.
- The Coder variant keeps 91% of the full model's SWE-bench Verified score (measured by its authors) but is weaker outside code, including CJK text.
Why it matters
A 125-billion-parameter model normally needs a server with hundreds of gigabytes of graphics memory. Strata's pitch is that a gaming PC with a 12-24 GB card can run it by sharing the work across GPU, RAM, CPU and SSD. The tricks it uses are splitting the model into 24,576 experts of which only 10 fire per word, keeping the busiest experts on the GPU, and guessing ahead with a small helper model. All of this runs locally, and the README says nothing leaves the PC. The model comes from the Qwen team; the compression comes from ISTA-DASLab, UkisAI (Swift 1.5) and Unsloth, with Strata building on llama.cpp / ggml.
Who it affects
People with a gaming PC who want a large model running locally: an NVIDIA or AMD card with 12 GB or more, Windows or Linux, and enough RAM and disk for a download of about 70 GB by default. It also reaches developers who want a local backend for coding agents, since Strata exposes OpenAI-compatible and Anthropic-style endpoints and an MCP server. The Coder size fits 32 GB of RAM. Users of older GPUs, Intel Arc or CPUs without AVX2 get experimental, community-tested paths.
How to use it
Strata is free and open source under the MIT License; a few parts and every model carry their own licenses. Download and unzip it (or git clone it), then run START-HERE.bat on Windows or ./setup.sh on Linux. Answer the installer's questions (model and size, context length, picture reading) or press Enter for the recommended choices. The app opens at http://127.0.0.1:8080. Point other apps at http://127.0.0.1:8080/v1 as an OpenAI-compatible provider; any API key and any model name work. To reach it from a phone or another PC, start it with --host 0.0.0.0 and always set an API key. Smaller sizes (Q2_0 or IQ2_XS) are the README's advice if RAM runs short.
How solid is it
The source is the project's own README, so the claims are the author's. The speed numbers shown come from the authors' tests on two ordinary gaming PCs, with full tables in DETAILS.md. The only 100-140 tokens per second figure is a projection ('should') for an RTX 3090. The README gives no quality comparison with the uncompressed model beyond the Coder variant's 91% SWE-bench Verified score, which the authors measured themselves. The text does not name the author or maintainer of Strata.
Risks and caveats
Startup is heavy: the PC can be slow or stop responding for 1-3 minutes while 35-55 GB loads into RAM. Too little free RAM leads to heavy disk activity or an engine crash. The README advises closing other programs or picking a smaller size. The largest, closest-to-full size (UD-Q4_K_XL, experimental) writes only 7-8.5 tokens per second on a 64 GB PC because it reads from the SSD. The Coder variant is weaker outside code, including Chinese and other CJK text. The first message of a long chat takes about 1 minute per 30,000 tokens to read. By default only one request is served at a time, and parallel mode slows each answer on a 12 GB card. AMD cards cannot yet read pictures on Windows. Exposing Strata beyond the local machine requires setting an API key.
“Think of a kitchen: the things you use all the time stay on the counter, and the rest waits in the pantry.”
— Strata README