Strata Runs a 125B AI Model on a Gaming PC
From the member
GBTI NetworkStrata is a free and open-source project designed to run Qwen3.8-Flash-Next, a 125-billion-parameter AI model, on ordinary gaming PCs. It supports Windows and Linux with compatible NVIDIA or AMD graphics cards starting at 12 GB of VRAM, at least 32 GB of system RAM, and roughly 80 GB of free storage.
Instead of requiring the entire model to fit on the graphics card, Strata distributes the workload across the computer. Frequently used experts remain on the GPU, the full expert set is held in system RAM, the CPU handles additional work, and the SSD stores supporting data. The project also uses speculative decoding, where a smaller model proposes upcoming tokens for the larger model to verify.
Strata exposes OpenAI-compatible, Anthropic-compatible, and Responses API endpoints, allowing it to work with coding tools and other applications that already support those interfaces. It also includes a browser interface for chat and system monitoring, configurable reasoning effort, optional image input, multi-GPU support, and local network access.
The project reports 53 to 94 output tokens per second across several quantizations on an RTX 5070 with 12 GB of VRAM, and 44 to 60 tokens per second on an RX 9070 XT with 16 GB. Strata is released under the MIT License, while the included models and some components retain their own licenses.
Footnote
- Niko1221, Strata, GitHub: https://github.com/Niko1221/Strata



0 Comments
No comments yet. Be the first. Members comment from the GBTI local client, where comments are submitted as pull requests and auto-published for paid members.
Become a memberComments are for members. Become a member to join the conversation, or if you already are.
You are signed in as a member. .