AI NewsOpen sourceAnnouncement
Niko1221's Strata runs a 125-billion-parameter model on a gaming PC with one click
Strata installs Qwen3.8-Flash-Next on a Windows or Linux PC with one NVIDIA card and 64 GB of RAM, and the project collected 3,493 stars in its first week.

Why it mattersA frontier open model running at 60 to 95 tokens per second on a 2,000-dollar gaming PC changes which teams can afford to keep model weights in-house.
A frontier open model that normally needs a server now installs on a gaming PC by double-clicking a batch file. Strata, published by GitHub user Niko1221 on 24 September 2026, runs Qwen3.8-Flash-Next, a 125-billion-parameter model, on one NVIDIA card with 12 to 24 GB of memory plus 64 GB of RAM, under Windows or Linux. The repository collected 3,493 stars in its first seven days.
The project's own description says the model writes answers at 60 to 95 tokens per second (about three-quarters of a word per token), which Strata's author calls faster than the reader can read. Niko1221 reports these numbers on their own hardware, so they are the author's own measurements rather than an independent benchmark.
What the measurement actually says
Strata's README publishes a table measured on an RTX 5070 (12 GB), a Ryzen 5 7600 and 64 GB of RAM. The fastest size (Q2_0) writes short-chat answers at 93 tokens per second and ingests prompts at 2,170 tokens per second. The slowest and highest quality size (IQ3_S) writes at 53 tokens per second and ingests at 1,620 tokens per second. The author says the IQ3_S size matches the full-precision model on the model's published tests.
An RTX 3090 with 24 GB should run roughly 100 to 140 tokens per second, according to the author's estimate. Two or three NVIDIA cards can be used in a pipeline, and prompts on an RTX 5080 plus RTX 3090 were read 18 to 20 percent faster than on the 5080 alone. Every card must be RTX 20 series or newer with 8 GB or more.
The cost of adopting it
Install needs an 80 GB free disk, an SSD, and a current NVIDIA driver. Everything else, including Python, the engine and the model weights (about 70 GB), is downloaded by the setup script. First start loads 35 to 55 GB into RAM and locks part of it for the graphics card, during which the author warns the PC can be slow or stop responding for one to three minutes. The project is MIT-licensed, written in C++, and last pushed to on 1 October 2026.
For a team that wants to keep model weights off a vendor's servers, the practical question is whether the hardware already available in the office clears the bar. Strata's README answers that in plain terms: an NVIDIA RTX 20 or newer card with 12 GB, 64 GB of system RAM, 80 GB of free disk, Windows or Linux. The caveats are that these are the author's own numbers measured on their own hardware, that the model requires a batch script run on each start, and that only NVIDIA is officially supported (AMD Radeon is listed as experimental on Linux).
Source
Niko1221/Strata on GitHub, repository created 24 September 2026, star count read 1 October 2026.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.