---
title: "177B Model on a Diet: GSQ-RCO Quantized Version Matches Original with 2.2M Downloads"
description: "This week's hottest thing on the local-model radar isn't a new model, it's Qwen3.8-Flash-Next-GSQ-RCO-GGUF, a 177B-class MoE squeezed down to IQ3 quantization, with 2.2 million downloads and 617 likes, the author even claims IQ3_S matches the original BF16 model's scores."
canonical_url: "https://blog.markkulab.net/en/local-ai-lab/local-ai-lab-2026-10-05"
author: "Mark Ku"
author_url: "https://blog.markkulab.net/en/author/mark-ku"
site: "Mark Ku's Tech Notes"
date_published: "2026-10-05T02:00:00.000Z"
category: "Local AI Lab"
tags: ["local-ai", "podcast", "地端模型", "lora", "開源模型", "QwenFlashNext", "GSQRCO", "FrogNano", "GGUF", "量化模型", "SWEbench"]
language: "en"
license: "CC BY 4.0"
license_url: "https://creativecommons.org/licenses/by/4.0/"
attribution: "when reusing or quoting, credit the author and link back to the original"
---

# 177B Model on a Diet: GSQ-RCO Quantized Version Matches Original with 2.2M Downloads

> **TL;DR** — Qwen3.8-Flash-Next-GSQ-RCO-GGUF, a 177B MoE model quantized to IQ3 via GSQ and RCO techniques, achieves 2.

# Opening

This week's hottest item on the local-model radar isn't a new model at all; it's a 177B-class MoE squeezed down to IQ3 quantization, Qwen3.8-Flash-Next-GSQ-RCO-GGUF, with 2.2 million downloads and 617 likes. The author even claims that IQ3_S matches the original BF16 model's scores. The other headliner is Microsoft's FrogNano-4B-2609, which hits 61.5% on SWE-bench at just 9.3 GB. But is it really install-and-go? Read on to find out.

## This week's picks

### 1. Qwen3.8-Flash-Next-GSQ-RCO-GGUF: a 177B MoE squeezed into IQ3, matching the original's scores

- **What it is**: A GGUF quantized release from IST Austria's DASLab, built on their own in-house method, with a 177B-class MoE underneath. GSQ (Gumbel-Softmax Quantization) and RCO (Riemannian Constrained Optimization) each have a dedicated arXiv paper (2604.18556, 2605.00649), so this isn't some casual community quant dump. It ships in four variants, Q2_0, IQ2_XS, IQ3_XXS, and IQ3_S, licensed under Apache-2.0, inherited from the base model.

- **Where it fits**:
  1. Teams whose proprietary code and client data can't leave the premises, who want to run a 177B-class model locally for inference and coding assistance: the README notes that with `-lm mmap --lazy-mode on`, the 28.8 GB n-gram lookup table can stay memory-mapped on disk instead of occupying resident memory.
  2. Engineers trying to balance budget against quality: the card lists a direct comparison table of recovery rates against the BF16 base (task average 93.12) for all four variants, IQ3_XXS at 92.57 and IQ3_S at 93.26, with the author stating that IQ3_S "matches or exceeds the base model on every task." You can pick a variant based on your own VRAM budget without having to rerun the benchmarks yourself.
  3. Anyone already comfortable with Ollama or LM Studio: the README provides `ollama run hf.co/<repo>`, while LM Studio users can just pick the `GSQ-RCO-*` build from the file list, no need to change your existing workflow.

- **Will it actually run**: The README's download table lists Shard 1 (weights) at Q2_0 37.6 GB, IQ2_XS 39.2 GB, IQ3_XXS 47.0 GB, IQ3_S 54.8 GB, while Shard 2 (the n-gram lookup table) at 28.8 GB can stay memory-mapped on disk. As for whether a consumer-grade single GPU can actually handle it, the official card makes no claims of its own: the llm-bench.io site states a peak of roughly 36.4 tok/s on consumer GPUs, and a third party mentions being able to start from as little as 8 GB VRAM via the Strata Engine. Both of these figures come from independent third-party testing sites and tools, not numbers verified by Qwen3.8-Flash-Next-GSQ-RCO-GGUF's own team; there's no public detail on testing methodology or whether the hardware configuration matches yours, so treat these as reference points only, not guarantees.Official Mac support status isn't listed either.

- **Buzz this week**: HF popularity score of 234 (2.2 million downloads, 617 likes), the highest download count among all candidates this round.

- **How it compares**: The author claims GSQ is "closing most of the gap between scalar and vector quantization at 2 to 3 bits," while still saving to the standard scalar GGUF format; RCO handles per-tensor quantization type assignment within an overall size budget. Worth distinguishing: a Hugging Face discussion thread has one user commenting "performance isn't as good as the official version, but excellent for an IQ3 quant." That's just one individual's personal impression, a sample size of one, and shouldn't be weighed on the same level as the recovery-rate numbers on the card, which are backed by a full benchmarking process.

- **How to get it**: `hf download <repo> --include "IQ3_XXS/*" --local-dir .`, then `llama-cli -m IQ3_XXS/Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001-of-00002.gguf -lm mmap --lazy-mode on -ngl 99 -p "..."`; also supports `ollama run hf.co/<repo>`, and LM Studio users can search the repo name directly.

- **Link**: [Hugging Face](https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF)

### 2. FrogNano-4B-2609: runs agent tasks at 4B, but it's not exactly plug-and-play

- **What it is**: A 4.66-billion-parameter model from Microsoft, 32 dense layers, combining Gated DeltaNet with gated-attention architecture, with roughly 131K context under the evaluation setup. Licensed under MIT, shipped in safetensors format. The technical report is on arXiv at 2609.07925, with the GitHub repo at microsoft/FrogNano.

- **Where it fits**:
  1. Individual developers who want an offline local coding agent: the model card describes it as being "given an authorized repository snapshot and an English natural-language issue," where the model produces text and structured Leaf tool calls, executed by a harness, forming a loop of repeatedly browsing files, searching, editing code, and running tests.
  2. Enterprise teams whose code can't leave the building, who want a local model to take a first pass before human review: the card lists applicable scopes as "bug diagnosis and repair, scoped feature implementation, regression fixing, test-driven code maintenance," explicitly positioned for human-supervised development.
  3. Anyone researching how small models get trained into agents: the technical report takes a pure-RL-plus-synthetic-tasks approach without distillation from frontier models; community posts summarize it as roughly 1,500 synthetic tasks over 5 iterations, a method you can reproduce by following along.

- **Will it actually run**: The official model card states that the BF16 checkpoint "requires about 9.3 GB for model weights alone, with additional memory needed for runtime state." However, the card also notes that exact minimum GPU model and VRAM configuration are "still to be validated before release," and the official team currently provides no quantized version. A community conversion already exists (roman220220/FrogNano-4B-2609-gptq-mlx-jang), but that's not an official artifact, and its stability and correctness haven't been verified by Microsoft.

- **Buzz this week**: HF popularity score of 105 (581 downloads, 107 likes).

- **How it compares**: Compared to other agent-oriented candidates this round that start at 27B and up (for example, BAAI/AREX-2 and autotrust/JEV-27B-VL also on this week's radar), FrogNano is just 4B. The official card's own comparison shows base model at 39.4% versus 61.5% after training, under the same pipeline. As for comparing this 4B model against much larger systems like GPT-5 mini, Grok 4, or Opus 4.1, that's a claim made by daily.dev roundups and posts from researchers on X (Rohan Paul, Minseon Kim), not an official statement from Microsoft.

- **How to get it**: On GitHub (microsoft/FrogNano, MIT), the commands are `uv venv --python 3.12` and `uv pip install -e .`, with evaluation via `frognano-eval run --config frognano/configs/eval/swebench-verified.yaml`. There's a gap worth knowing about upfront: the repo itself doesn't include vLLM or local-serving instructions; you're expected to set up an OpenAI-compatible endpoint yourself via `FROGNANO_MODEL_BASE_URL` and run tasks inside a Kubernetes sandbox. In other words, 9.3 GB is just the VRAM threshold for the model weights themselves; to actually turn it into an agent that "takes an issue and fixes it on its own," you'll still need to assemble your own harness, agent loop, and execution sandbox. For individual developers without a K8s environment who just want to double-click and run, be prepared for that gap, it's a bigger lift than the phrase "offline local agent" makes it sound.

- **Link**: [Hugging Face](https://huggingface.co/microsoft/FrogNano-4B-2609)

## Closing

These two represent two different trade-offs: Qwen3.8-Flash-Next-GSQ-RCO-GGUF uses quantization to bring a 177B-class model down to a scale you and I can actually run, with clear weight sizes and a clear path to getting it. FrogNano-4B-2609's VRAM threshold looks much lower, but turning it into an agent that can genuinely edit code on its own still requires you to build your own harness and sandbox, and those two things shouldn't be mistaken for equivalent. If you just want to get a feel for quantized models, the GGUF one is ready to run with Ollama right now. If you're more interested in the agent training methodology itself, FrogNano's technical report is worth reading in full before deciding whether to dive in. See you next Friday for whatever else turns up on the local-model radar.

---

## About this article and its author

Originally published on [Mark Ku's Tech Notes](https://blog.markkulab.net/en/local-ai-lab/local-ai-lab-2026-10-05)

License: [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) — when reusing or quoting, credit the author and link back to the original

### About the author

**[Mark Ku](https://blog.markkulab.net/en/author/mark-ku)** — Software engineer

- 10+ years as a software engineer
- Built North-American e-commerce and AI SaaS subscription billing

### Free tools built by the author

All of these are free to use:

- [Free PDF Sign Tool](https://blog.markkulab.net/en/tools/pdf-sign): Online PDF sign tool — draw, type, or upload a signature, then drag, resize, and download. Everything runs in your browser; nothing is uploaded.
- [VS Code Refactory](https://blog.markkulab.net/en/tools/refactory): Refactory is a VS Code refactoring extension: 34 actions plus a 37-rule code-smell inspection layer with a Code Health dashboard, across 18 languages, backed by 534 tests. It learns your repo's conventions: where interfaces live, where DI is registered, whether 'use client' belongs. It ranks files by git churn × complexity so you know what to fix first, and hands any smell to the Claude Code already on your machine. Free to use, and your source never leaves your computer.
- [DB-Kit Database Manager](https://blog.markkulab.net/en/tools/db-kit): DB-Kit is a lightweight, cross-platform database manager built with Tauri + Rust + React. Manage MySQL, MariaDB, PostgreSQL, SQL Server, Oracle, SQLite, MongoDB, Redis, Kafka, Elasticsearch and RabbitMQ from one consistent interface: passwords encrypted in the OS keychain, SSH tunnels, full CRUD, a visual query builder, stacked multi-statement result sets, cross-connection data transfer, schema & data compare with sync SQL and schema snapshots, Excel / CSV import & export, visualized execution plans, ER diagrams, scheduled backups, SQL stress testing with p50–p99 latency percentiles, a 15-rule SQL review engine, Review & Run (AI review plus per-statement backups and an auto-generated rollback script), Kafka message browsing with monitoring & alerts, an SSH terminal with SFTP / FTP, Docker / Kubernetes, remote desktop (RDP / VNC / RustDesk), file / folder compare, stored procedure integration tests, six UI languages, a built-in AI assistant (a local CLI or any Anthropic / OpenAI-compatible API; natural-language SQL, AI review and tuning advice) and the dbk CLI. Free and open source (MIT), with installers for Windows, macOS and Linux.
- [VS Code Super Mermaid](https://blog.markkulab.net/en/tools/super-mermaid): Super Mermaid is a VS Code extension for beautiful Mermaid diagrams out of the box: auto-colored live preview, mouse pan & zoom, high-res PNG / SVG export, 21 templates and multiple themes. Free and open source (MIT).
- [React Super Mermaid](https://blog.markkulab.net/en/tools/react-super-mermaid): react-super-mermaid is an open-source React component library: render beautiful Mermaid diagrams with a single <MermaidViewer>, with built-in colorful / sketch themes, pan & zoom, in-diagram search, and high-res SVG / PNG export. Lightweight, SSR-safe, fully typed. Free and open source (MIT).
- [Jira / Confluence Super Mermaid](https://blog.markkulab.net/en/tools/jira-super-mermaid): An Atlassian Forge app: write Mermaid syntax directly inside a Jira issue or a Confluence page and get flowcharts, sequence diagrams, state machines and Gantt charts. 11 diagram types, SVG / PNG export, light and dark themes, full CJK support. Runs on Atlassian: your diagrams live in your own site and the app calls no third-party service. Free, coming soon to the Atlassian Marketplace.
- [Mermaid Live Preview](https://blog.markkulab.net/en/tools/mermaid-preview): Write Mermaid in your browser, see it render instantly, and share the whole diagram as a single link. No sign-up, nothing uploaded to a server, and mermaid.live share links work as-is.
- [React Intl Phone Number](https://blog.markkulab.net/en/tools/react-intl-phone-number): react-intl-phone-number is an open-source React component: framework-agnostic and antd-free, with E.164 in/out, a searchable flag / country-code dropdown, configurable validation levels (strict / mobile-strict / loose), themeable CSS, and i18n — phone logic powered by google-libphonenumber. Lightweight and fully typed. Free and open source (MIT).
- [Uptime Kuma Cluster](https://blog.markkulab.net/en/tools/uptime-kuma-cluster): Turn single-node Uptime Kuma into a highly available cluster: OpenResty + Lua smart load balancing, shared MariaDB state, health checks and automatic failover, plus cluster-management REST APIs. One Docker Compose command to start. Free and open source (MIT).
- [AI Podcast Cut](https://blog.markkulab.net/en/tools/ai-podcast-cut): Drop in a recording and it removes fillers and stutters, levels loudness segment by segment, and sends a second agent to review every cut. Cut points snap to word boundaries and zero crossings, every splice gets a fade, and sentence-end breaths are preserved. Desktop app for Windows, macOS and Linux. MIT licensed; the Windows installer bundles ffmpeg.
- [AI Video Cut](https://blog.markkulab.net/en/tools/ai-video-cut): Pick something in your video: type it (face, license plate, phone screen, logo), click points or drag a box, or let Claude Code / Codex pick it. AI Video Cut tracks it frame by frame, then mosaics or blurs it, recolors it, pins stickers and text to it, swaps a screen or poster for your own image or video, or removes it using background that other frames actually captured. Plus sequence editing, local captions and an AI assistant. Free, open-source (MIT) Windows desktop app (macOS / Linux experimental) that runs locally, nothing uploaded.
- [open-pos restaurant POS](https://blog.markkulab.net/en/tools/open-pos): One computer and one receipt printer is enough to open the shop. Your data lives on your own disk, no subscription, no lock-in, MIT licensed. Money is integer New Taiwan dollars with tax split by the statutory formula, so sales + tax always equals the total. Printing goes straight over ESC/POS on TCP 9100, with no vendor driver. Tauri + Rust + SQLite desktop app, v1.0 in development.
- [.NET e-commerce API template](https://blog.markkulab.net/en/tools/dotnet-ecommerce-api): An e-commerce backend API template built on .NET 10: Controller → Service → Repository layering, Autofac convention-based registration, SqlSugar for MySQL / MariaDB and SQL Server, Redis cache and message queue, Hangfire jobs, JWT auth, and an extensible payment provider module (ECPay, Newebpay). Apache 2.0.

### Daily podcasts

- [Mark's Tech Insights — Daily AI News](https://blog.markkulab.net/en/category/tech-news): Daily curated AI and tech trends. Catch the latest developments via audio summaries — covering AI applications, software architecture, DevOps, and engineering practice. — RSS: https://blog.markkulab.net/feed.xml
- [AI股市蝦聊](https://blog.markkulab.net/en/category/ai-stock-chat): Every trading day, an AI-analyzed take on the Taiwan stock market, delivered as a two-host conversation covering the session and the next-day outlook. — RSS: https://blog.markkulab.net/ai-stock-chat/feed.xml
- [開源好物週報](https://blog.markkulab.net/en/category/open-source-weekly): A weekly two-host pick of free open-source tools surfaced from real Hacker News, GitHub, and Reddit buzz — what pain they solve and the fastest way to get started. — RSS: https://blog.markkulab.net/open-source-weekly/feed.xml
- [AI 運動週報](https://blog.markkulab.net/en/category/sports-weekly): Two hosts talk NBA, MLB and world sport three times a week — scores, records and the stories behind them, from a Taiwanese fan perspective. — RSS: https://blog.markkulab.net/sports-weekly/feed.xml
- [AI 國際新聞快報](https://blog.markkulab.net/en/category/world-news): A daily two-host briefing that makes sense of the past 24 hours in world news — geopolitics, the global economy, conflict and security, disasters and climate — from a Taiwanese perspective, neutral and fully sourced. — RSS: https://blog.markkulab.net/world-news/feed.xml
- [地端 AI 實驗室](https://blog.markkulab.net/en/category/local-ai-lab): Twice a week, a two-host look at open-weight models and LoRAs you can actually run on your own machine: what they are for, whether your GPU can handle them, and how they compare — sourced from official model cards. — RSS: https://blog.markkulab.net/local-ai-lab/feed.xml

### Deals

- [Saily eSIM](https://blog.markkulab.net/en/saily): Travel eSIM by the NordVPN team，Promo code：KUKU
- [NordVPN](https://blog.markkulab.net/en/nordvpn): The world's leading VPN, independently audited
- [PremLogin](https://blog.markkulab.net/en/premlogin): Subscription sharing for streaming and AI seats，Promo code：markku666

### Newsletter

[Subscribe to the newsletter](https://blog.markkulab.net/en/subscribe) — Be the first to know about new posts. No spam, unsubscribe anytime.
