Opening
This week's hottest item on the local-model radar isn't a new model at all; it's a 177B-class MoE squeezed down to IQ3 quantization, Qwen3.8-Flash-Next-GSQ-RCO-GGUF, with 2.2 million downloads and 617 likes. The author even claims that IQ3_S matches the original BF16 model's scores. The other headliner is Microsoft's FrogNano-4B-2609, which hits 61.5% on SWE-bench at just 9.3 GB. But is it really install-and-go? Read on to find out.
This week's picks
1. Qwen3.8-Flash-Next-GSQ-RCO-GGUF: a 177B MoE squeezed into IQ3, matching the original's scores
-
What it is: A GGUF quantized release from IST Austria's DASLab, built on their own in-house method, with a 177B-class MoE underneath. GSQ (Gumbel-Softmax Quantization) and RCO (Riemannian Constrained Optimization) each have a dedicated arXiv paper (2604.18556, 2605.00649), so this isn't some casual community quant dump. It ships in four variants, Q2_0, IQ2_XS, IQ3_XXS, and IQ3_S, licensed under Apache-2.0, inherited from the base model.
-
Where it fits:
- Teams whose proprietary code and client data can't leave the premises, who want to run a 177B-class model locally for inference and coding assistance: the README notes that with
-lm mmap --lazy-mode on, the 28.8 GB n-gram lookup table can stay memory-mapped on disk instead of occupying resident memory. - Engineers trying to balance budget against quality: the card lists a direct comparison table of recovery rates against the BF16 base (task average 93.12) for all four variants, IQ3_XXS at 92.57 and IQ3_S at 93.26, with the author stating that IQ3_S "matches or exceeds the base model on every task." You can pick a variant based on your own VRAM budget without having to rerun the benchmarks yourself.
- Anyone already comfortable with Ollama or LM Studio: the README provides
ollama run hf.co/<repo>, while LM Studio users can just pick theGSQ-RCO-*build from the file list, no need to change your existing workflow.
- Teams whose proprietary code and client data can't leave the premises, who want to run a 177B-class model locally for inference and coding assistance: the README notes that with
-
Will it actually run: The README's download table lists Shard 1 (weights) at Q2_0 37.6 GB, IQ2_XS 39.2 GB, IQ3_XXS 47.0 GB, IQ3_S 54.8 GB, while Shard 2 (the n-gram lookup table) at 28.8 GB can stay memory-mapped on disk. As for whether a consumer-grade single GPU can actually handle it, the official card makes no claims of its own: the llm-bench.io site states a peak of roughly 36.4 tok/s on consumer GPUs, and a third party mentions being able to start from as little as 8 GB VRAM via the Strata Engine. Both of these figures come from independent third-party testing sites and tools, not numbers verified by Qwen3.8-Flash-Next-GSQ-RCO-GGUF's own team; there's no public detail on testing methodology or whether the hardware configuration matches yours, so treat these as reference points only, not guarantees.Official Mac support status isn't listed either.
-
Buzz this week: HF popularity score of 234 (2.2 million downloads, 617 likes), the highest download count among all candidates this round.
-
How it compares: The author claims GSQ is "closing most of the gap between scalar and vector quantization at 2 to 3 bits," while still saving to the standard scalar GGUF format; RCO handles per-tensor quantization type assignment within an overall size budget. Worth distinguishing: a Hugging Face discussion thread has one user commenting "performance isn't as good as the official version, but excellent for an IQ3 quant." That's just one individual's personal impression, a sample size of one, and shouldn't be weighed on the same level as the recovery-rate numbers on the card, which are backed by a full benchmarking process.
-
How to get it:
hf download <repo> --include "IQ3_XXS/*" --local-dir ., thenllama-cli -m IQ3_XXS/Qwen3.8-Flash-Next-GSQ-RCO-IQ3_XXS-00001-of-00002.gguf -lm mmap --lazy-mode on -ngl 99 -p "..."; also supportsollama run hf.co/<repo>, and LM Studio users can search the repo name directly. -
Link: Hugging Face
2. FrogNano-4B-2609: runs agent tasks at 4B, but it's not exactly plug-and-play
-
What it is: A 4.66-billion-parameter model from Microsoft, 32 dense layers, combining Gated DeltaNet with gated-attention architecture, with roughly 131K context under the evaluation setup. Licensed under MIT, shipped in safetensors format. The technical report is on arXiv at 2609.07925, with the GitHub repo at microsoft/FrogNano.
-
Where it fits:
- Individual developers who want an offline local coding agent: the model card describes it as being "given an authorized repository snapshot and an English natural-language issue," where the model produces text and structured Leaf tool calls, executed by a harness, forming a loop of repeatedly browsing files, searching, editing code, and running tests.
- Enterprise teams whose code can't leave the building, who want a local model to take a first pass before human review: the card lists applicable scopes as "bug diagnosis and repair, scoped feature implementation, regression fixing, test-driven code maintenance," explicitly positioned for human-supervised development.
- Anyone researching how small models get trained into agents: the technical report takes a pure-RL-plus-synthetic-tasks approach without distillation from frontier models; community posts summarize it as roughly 1,500 synthetic tasks over 5 iterations, a method you can reproduce by following along.
-
Will it actually run: The official model card states that the BF16 checkpoint "requires about 9.3 GB for model weights alone, with additional memory needed for runtime state." However, the card also notes that exact minimum GPU model and VRAM configuration are "still to be validated before release," and the official team currently provides no quantized version. A community conversion already exists (roman220220/FrogNano-4B-2609-gptq-mlx-jang), but that's not an official artifact, and its stability and correctness haven't been verified by Microsoft.
-
Buzz this week: HF popularity score of 105 (581 downloads, 107 likes).
-
How it compares: Compared to other agent-oriented candidates this round that start at 27B and up (for example, BAAI/AREX-2 and autotrust/JEV-27B-VL also on this week's radar), FrogNano is just 4B. The official card's own comparison shows base model at 39.4% versus 61.5% after training, under the same pipeline. As for comparing this 4B model against much larger systems like GPT-5 mini, Grok 4, or Opus 4.1, that's a claim made by daily.dev roundups and posts from researchers on X (Rohan Paul, Minseon Kim), not an official statement from Microsoft.
-
How to get it: On GitHub (microsoft/FrogNano, MIT), the commands are
uv venv --python 3.12anduv pip install -e ., with evaluation viafrognano-eval run --config frognano/configs/eval/swebench-verified.yaml. There's a gap worth knowing about upfront: the repo itself doesn't include vLLM or local-serving instructions; you're expected to set up an OpenAI-compatible endpoint yourself viaFROGNANO_MODEL_BASE_URLand run tasks inside a Kubernetes sandbox. In other words, 9.3 GB is just the VRAM threshold for the model weights themselves; to actually turn it into an agent that "takes an issue and fixes it on its own," you'll still need to assemble your own harness, agent loop, and execution sandbox. For individual developers without a K8s environment who just want to double-click and run, be prepared for that gap, it's a bigger lift than the phrase "offline local agent" makes it sound. -
Link: Hugging Face
Closing
These two represent two different trade-offs: Qwen3.8-Flash-Next-GSQ-RCO-GGUF uses quantization to bring a 177B-class model down to a scale you and I can actually run, with clear weight sizes and a clear path to getting it. FrogNano-4B-2609's VRAM threshold looks much lower, but turning it into an agent that can genuinely edit code on its own still requires you to build your own harness and sandbox, and those two things shouldn't be mistaken for equivalent. If you just want to get a feel for quantized models, the GGUF one is ready to run with Ollama right now. If you're more interested in the agent training methodology itself, FrogNano's technical report is worth reading in full before deciding whether to dive in. See you next Friday for whatever else turns up on the local-model radar.



























Comments