Newsletter

Strata Says “Any Consumer Hardware.” The README Says 12GB of VRAM and 32GB of RAM.

Strata picked up 2,814 GitHub stars in a single day, the fastest real daily move we’ve measured. It installs QwStrata picked up 2,814 GitHub stars in a single day, the fastest real daily move we’ve measured. It installs Qwen3.8-Flash-Next, a 125 billion parameter model, on a Windows or Linux desktop and serves it at localhost on an OpenAI and Anthropic compatible API. Point Claude Code or Codex at it, then nothing leaves your machine. MIT licensed, free.

The repo’s one-line description says “any consumer hardware.” Its README says 12GB of VRAM, 32GB of RAM plus 80GB of free disk. Those aren’t the same sentence. On a 12GB RTX 5070 the developers measured 94 tokens a second, faster than you read, so it really does work. It works on an enthusiast gaming PC rather than on a laptop.

Every figure below comes from Strata’s own docs or Qwen’s model card. We haven’t installed it.

Best for anyone with a recent gaming desktop who wants a capable model running offline. Not ideal for laptop owners, 16GB machines, or anyone expecting a one-click install to be the whole story.

A 125 billion parameter model needs a server. Everybody knows that.

Strata’s pitch is that everybody’s wrong. It has the measurements to back that up.

It also has a one-line GitHub description reading “Qwen3.8-Flash-Next on any consumer hardware,” which is the part that’ll waste your afternoon. The README three clicks later wants a 12GB graphics card, 32GB of RAM and 80GB of free disk. Both statements come from the same repo. Only one of them is a spec.

So here’s the useful version. What it is, what it actually needs, how fast it really goes, plus the 30 second check that tells you whether to bother.


What Strata Actually Is

RepoNiko1221/Strata
Stars13,097, up 2,814 in one day
Forks1,103, a 8.4% fork ratio
CreatedSeptember 24, 2026, so 11 days old
LicenceMIT for Strata itself
LanguageC++
Model it runsQwen3.8-Flash-Next, licence qwen-community-1.0
Model size125B total, 6B active per token, plus a 51B n-gram embedding
Context262,144 tokens native, extensible to 1,000,000
PlatformsWindows 10 and 11, Linux
ServesOpenAI compatible at 127.0.0.1:8080/v1, Anthropic at /v1/messages, Codex Responses at /v1/responses

Strata isn’t a model. It’s the thing that makes somebody else’s model fit on hardware that shouldn’t hold it, built on parts of llama.cpp and ggml. The model is Qwen’s, released in August. Compressed versions come from ISTA-DASLab, UkisAI and Unsloth.

The star velocity is what pulled it into view. Our scanner measured +2,814 in 24 hours against the previous day’s count, which is the largest verified daily gain it has recorded since we started measuring day over day instead of dividing weekly figures. For context, the repo we ranked first in Fresh Commits 05 is doing +100 a day.


How It Fits, Which Is The Whole Trick

Qwen3.8-Flash-Next is a mixture of experts model. That phrase does a lot of work, so here’s what it means in practice.

The model holds 24,576 small specialist networks. Any single word you feed it wakes up 10 of them. That’s 0.041% of the model active at any moment, which is why the active parameter count is 6B against a 125B total, about 4.8%.

The Pantry, Not The Counter

Strata’s own docs use a kitchen metaphor and it’s a good one. The ingredients you reach for constantly stay on the counter. Everything else waits in the pantry.

Your graphics card holds the few thousand experts that get used most. Your RAM holds all of them. Whatever the card didn’t take, your processor works on at the same time. And your SSD holds a 29GB lookup table that never needs to move.

That’s the architecture. The model loads in two shards. Shard one goes into RAM and onto the card when Strata starts. Shard two, the lookup table, stays on disk and gets read from where it lies.

Which explains the requirements. The 12GB card isn’t holding the model, it’s holding the hot path. RAM is the real constraint, because RAM holds every expert. That’s backwards from how people usually think about running models locally, where VRAM is the figure people watch and system RAM is an afterthought. Here a bigger card makes it faster. It doesn’t make it fit.

Guess, Then Check

The second trick is speculative decoding. A small helper model guesses the next few words, then the big model checks all the guesses at once rather than generating them one at a time. Same output, arrived at 1.6x to 1.8x sooner by Strata’s measurement. The helper arrives as a separate 6GB download on first start.

Long prompts get read in chunks of up to 8,192 tokens, which is why the prompt reading numbers below are so much higher than the writing numbers.

There’s a paper in the repo if you want the longer version.


What You Actually Need

This is the section the GitHub one-liner skips.

ComponentRequirement
Graphics cardNVIDIA RTX 20, 30, 40 or 50 series, or one of a specific list of AMD Radeon cards. 12GB of VRAM minimum
RAM32GB minimum. 64GB runs every size
DiskAbout 80GB free. SSD strongly preferred
SystemWindows 10 or 11, or Linux, with a current driver
Download66GB to 76GB for the three smaller sizes, plus about 6GB for the draft layer

A 12GB card means an RTX 3060 12GB or better. 32GB of RAM is above what most prebuilt desktops ship with and well above most laptops. 80GB of free disk is a real ask on a 512GB drive that already has games on it.

None of that makes Strata bad. It makes “any consumer hardware” wrong. That’s the line people read first, so it’s the line that’ll have them downloading 70GB before they find out their machine can’t hold it.

There are experimental paths for older cards, Tesla P40 and V100 and GTX 10 series and some older Radeons, plus Intel Arc built from source on Linux, plus processors without AVX2 which the docs describe as working but slow. Community written and community tested, which is an honest label.


How Fast It Really Goes

Strata publishes measured numbers on two named machines, which is more than most projects in this space manage. These are the developers’ own figures, not ours.

NVIDIA RTX 5070 (12GB), Ryzen 5 7600, 64GB RAM:

SizeShort chatAt 128K contextReading your prompt
Q2_094 tokens/s76 tokens/s2,650 tokens/s
IQ2_XS79 tokens/s63 tokens/s2,090 tokens/s
IQ3_XXS62 tokens/s49 tokens/s1,750 tokens/s
IQ3_S53 tokens/s46 tokens/s1,620 tokens/s
Coder55 tokens/s43 tokens/s2,180 tokens/s

AMD RX 9070 XT (16GB), Ryzen 9 3900X, 47GB RAM, Linux:

SizeShort chatAt 128K contextReading your prompt
Q2_060 tokens/s48 tokens/s1,160 tokens/s
IQ2_XS52 tokens/s36 tokens/s1,110 tokens/s
Coder44 tokens/s33 tokens/s1,420 tokens/s

A token is about three quarters of a word, so 94 tokens a second is roughly 70 words a second. You can’t read that fast. The practical ceiling here isn’t speed.

Two things worth pulling out. Long context costs 19% on the NVIDIA box going from a short chat to 128K. It costs 31% on the AMD box, so the penalty isn’t fixed across hardware. And the best quality size, IQ3_S, runs 44% slower than the fastest, which is the tradeoff you’re actually choosing between.

The docs also say an RTX 3090 with 24GB should reach roughly 100 to 140 tokens a second, which is an estimate rather than a measurement and labelled as one.

There’s also a calibration pass worth knowing about. Running START-HERE.bat --calibrate measures a handful of engine settings on your specific machine and keeps whichever came out fastest. It takes five to ten minutes, it’s NVIDIA only for now, plus it made the Coder 7% quicker on the test PC above. Small gain, free, which means the published figures are a floor rather than a ceiling on tuned hardware.

Strata also ships a report template so people can submit their own measurements as pull requests, which is how the per card numbers for the Radeon AI PRO R9700, the RX 7800 XT, the RX 9060 XT plus the RX 6900 XT got collected. Community figures sit in their own file rather than being mixed into the headline tables, which is the correct way round.


The Sizes, And The RAM Maths That Decides For You

The same model comes compressed to different degrees. Smaller is faster. Bigger is smarter. Your RAM picks for you.

SizeRAM plus VRAM neededExperts in RAMSpeedQuality
Q2_037.6 GB34 GBfastestgood
IQ2_XS39.2 GB35.5 GBfastbetter, the recommended pick
IQ3_XXS47.0 GB43 GBslowergreat
IQ3_S54.8 GB50 GBslowestbest, matches the full model on published tests
Coder29.6 GB shard 123 GBfastcode only, see below

The rule the docs give is simple enough to do in your head. Your RAM needs to hold the experts plus about 10GB for Windows and everything else you’ve got open. So 34GB of experts plus 10 is 44, which is why Q2_0 wants a 48GB tier. The Coder’s 23GB plus 10 is 33, which is how it squeezes onto a 32GB machine.

There’s also a low RAM mode where the experts get mapped from disk instead of loaded, chosen automatically when your RAM can’t hold them beside the system. It works. With a small card it’s much slower, because most experts then come off the SSD. Setup tells you so rather than letting you find out.

The two Unsloth sizes are worth knowing about mainly as a warning. UD-IQ4_XS is a 94GB download. UD-Q4_K_XL is 111GB with 77GB of experts that don’t fit in RAM at all, so Strata streams most of it off the SSD and you get 7 to 8.5 tokens a second on a 64GB machine with a 12GB card. That’s 11x to 13x slower than Q2_0. The docs mark it experimental and publish the number, which is the right way to ship something like that.


The Coder Version Is Half A Model

ISTA-DASLab’s Coder build keeps 256 of each layer’s 512 experts, picked on code data, then drops the other half. That’s how it fits a 32GB PC when nothing else does.

Its authors report it reaching 91% of the full model’s SWE-bench Verified score and 99% of LiveCodeBench. Those are the authors’ own numbers and no independent benchmark exists, so treat them as a claim.

What the docs are refreshingly direct about is the cost. With half the experts gone it’s weaker outside code and weaker in languages other than English. Strata’s own issue #438 is cited by name: Chinese answers came out wrong or looping where English was fine. A project that links its own bug report in its model selection guide is doing something most don’t.

Verified vs Asserted
What We Checked And What We’re Repeating
13,097 stars, 1,103 forks, MIT, created Sep 24
GitHub API
+2,814 stars in 24h
measured, two scans
125B total, 6B active, 262K context, qwen-community-1.0
Qwen model card
12GB VRAM, 32GB RAM, 80GB disk
Strata README
94 tokens/s on an RTX 5070
developer measured
Coder at 91% of SWE-bench Verified
its authors, no third party
RTX 3090 at 100 to 140 tokens/s
estimate, labelled as one
“Nothing leaves your PC”
not independently audited
Orange rows came from a primary source we read. Grey rows are claims from the people who built the thing. We haven’t installed Strata, so no speed figure here is ours.

The Licence Split Nobody Will Notice

Strata is MIT. The model is not.

Qwen3.8-Flash-Next ships under qwen-community-1.0, which is Qwen’s own licence rather than an OSI approved one. Swift 1.5, the fine-tune that thinks for less time before answering, carries its own licence again. The README says so plainly, in the credits, in one sentence: Strata is MIT and a few parts and every model have their own licences.

That’s honest. It’s also three clicks from where anyone stops reading, while “free and open source” sits in bold at the top of the page.

The practical version: if you’re running this at home for yourself, none of it matters. If you’re putting it anywhere near a product, read the Qwen licence before you read anything else, because that’s the file that governs what you’re actually allowed to do with the output. We went through a version of this with Alibaba quietly changing what Qwen meant. Licence terms on open weight models keep turning out to be worth checking first rather than last.


The Rough Edges, All Of Which Are In The Docs

Credit where it’s due. Almost every annoyance here is documented by the project itself rather than discovered by users.

Your PC may freeze for one to three minutes on first start, because Strata loads 35GB to 55GB into RAM and locks part of it for the graphics card. The docs say wait and don’t close the window. If it’s still frozen at 10 minutes, something’s wrong.

It answers one request at a time by default. You can set parallel to 2. On a 12GB card that makes every answer slower.

The first message in a chat takes about a minute per 30,000 tokens to read. Follow-ups start in seconds. So a long first prompt feels broken and isn’t.

AMD cards can read images on Linux through the processor. On Windows they can’t yet.

The download is 70GB and if it stops you rerun the installer, which continues where it left off.

None of that’s hidden. All of it will surprise someone who read the one-liner and nothing else.


It Runs One Model, And That’s The Point

Every other local runner you’ve used is a model manager. You browse a library, pull what you want, swap between a dozen things.

Strata runs one model.

That’s not a limitation they’re apologising for, it’s the design. Its docs open the model page with exactly that: Strata runs one model, in several sizes and several versions. Everything in the engine is tuned for how Qwen3.8-Flash-Next specifically is shaped, which is how you get 94 tokens a second out of 125 billion parameters on a card that cost a few hundred dollars.

A general runner can’t do that. If it has to handle any architecture you throw at it, it can’t assume 24,576 experts with 10 firing per token, can’t decide in advance which few thousand live on the card, can’t ship a lookup table layout that suits one model’s routing. Generality costs throughput. Strata spent the generality and bought the speed.

Which also tells you when to use something else. If you want to compare four models this week, this is the wrong tool. If you want one capable model running locally and fast, forever, it’s arguably the right one. Same logic as picking a dedicated harness over a general one: narrower scope, better at the thing it narrowed to.

The versions give you some room inside that constraint. Four compression sizes of the original, the Coder with half its experts gone, Swift 1.5 which is a fine-tune that thinks for a shorter time before answering at roughly the same quality, plus two Unsloth builds at the heavy end. Sizes share files, so adding a second one later doesn’t re download what it already has.


You Can Make Your Agent Install It

VU readers will care most about an install path that skips the installer entirely.

Strata’s README gives you a line to paste into Claude Code, Cursor, Codex or Copilot:

Set up Strata on this PC for me: https://github.com/Niko1221/Strata and follow docs/AI_SETUP.md in that repository.

The agent reads a doc written for agents, checks your graphics card and RAM and disk, picks the size that fits, installs it, starts it and tells you how to connect your other apps. There’s an AI_SETUP.md in the repo written for exactly that, plus an MCP server so an agent can start and stop Strata afterwards.

This is a small thing that says something larger. Hardware detection, dependency choice and a sizing decision are precisely the fiddly work that used to make local model setup a weekend. Handing that to an agent and shipping the doc it reads beats another GUI. It also sidesteps the “any consumer hardware” problem, because the agent looks at your actual machine and tells you what fits.

If you’d rather do it yourself, it’s START-HERE.bat on Windows or ./setup.sh on Linux, then Enter through the questions for the recommended answers. Same result, more waiting.

There’s also a browser app at 127.0.0.1:8080 with a chat, the settings, plus a live monitor of the model and your GPU and RAM. Thinking effort is a dropdown with off, low, medium and high. Off is fastest and high is for the hard ones.


What You Should Actually Do

Thirty second version. Open Task Manager, Performance tab. Look at two numbers: total RAM, plus your GPU’s dedicated memory.

Under 32GB of RAM, stop here. Nothing in Strata’s lineup fits and the low RAM mode on a small card will be slow enough to annoy you. Under 12GB of VRAM, same answer.

At 32GB with a 12GB card, you get the Coder, which is good at code and noticeably weaker at everything else. Know that going in.

At 64GB with a 12GB card you get the real thing. Take IQ2_XS, which is what the installer recommends and the right default. Q2_0 if you want the extra speed more than the extra quality.

Then the part that makes it worth the 70GB: add it to your tools as an OpenAI compatible provider at http://127.0.0.1:8080/v1, any key, any model name. For Claude Code set ANTHROPIC_BASE_URL=http://127.0.0.1:8080. For Codex use the Responses endpoint. That’s the whole integration. It’s why this matters more than a local chat window would. Your existing agent setup keeps working and stops sending anything out.

If you’re weighing this against a subscription, the honest comparison is about what you’re buying rather than about quality. Our Claude Pro review covers where paying wins. Strata wins on exactly two things. Nothing leaves your machine, plus there’s no usage limit to run into. It loses on everything else, including the model being a compressed 2 bit or 3 bit version of something that was already smaller than a frontier model.


The Part Worth Keeping

Strata is a seriously impressive piece of engineering with a marketing problem it created in one line of text.

The engineering: 24,576 experts, the hot few thousand on your card, all of them in RAM, a 29GB lookup table left on the SSD, a small model guessing ahead so the big one can check in batches. 94 tokens a second on a mid range gaming card for a 125 billion parameter model. Measured, published, on named hardware, with the slow configurations published too.

The problem: “any consumer hardware” is going to send thousands of people into a 70GB download their 16GB laptop was never going to run. Some will conclude local models are a scam rather than that they read a compressed sentence.

2,814 stars in a day says the appetite is enormous. The gap between that appetite and a 12GB card plus 32GB of RAM is the actual story of local AI right now. No amount of clever expert routing closes it this year.

Check Task Manager first. Then download.


Charts and Blocks

Tokens Per Second By Size, RTX 5070 12GB

Developer Measured
Writing Speed By Size
RTX 5070 (12GB), Ryzen 5 7600, 64GB RAM. Short chat. Strata’s own figures, not ours. A token is about three quarters of a word.
Q2_0 (fastest)
94/s
IQ2_XS (default)
79/s
IQ3_XXS
62/s
Coder (32GB RAM)
55/s
IQ3_S (best quality)
53/s
UD-Q4_K_XL (SSD)
8/s
Bars are relative to the fastest. The bottom row is the near full quality build whose experts don’t fit in RAM, so most of it streams off the SSD. Same hardware, 11x slower.

Can Your PC Run It

Thirty Second Check
What You Get At Each RAM Tier
Under 32GB
Nothing fits. Stop here.
32GB
Coder only. Strong at code, weaker at everything else and at non English.
48GB
Q2_0 or IQ2_XS. Full expert set. The larger sizes won’t fit.
64GB
Every size fits. Take IQ2_XS. This is the tier the project is really built for.
96GB plus
IQ3_S or Unsloth’s UD-IQ4_XS with room to spare.
Separately, you need a 12GB graphics card or better and about 80GB of free disk at every tier. Figures from Strata’s README and MODELS.md, October 5, 2026.