---
title: "🧪 Xiaomi's Flagship Needs 8 GPUs, But This MiMo-V2.6-Distill-Qwen-9B Runs with a Single Command | Local AI Lab"
description: "This week Hugging Face's trending chart was completely dominated by Xiaomi's MiMo V2.6 family: the official flagship example command requires 8 GPUs minimum, but MiMo-V2.6-Distill-Qwen-9B within the family is the only version regular GPUs have a chance of running, under MIT license, with official GGUF from the llama.cpp team themselves."
canonical_url: "https://blog.markkulab.net/en/local-ai-lab/local-ai-lab-2026-09-28"
author: "Mark Ku"
author_url: "https://blog.markkulab.net/en/author/mark-ku"
site: "Mark Ku's Tech Notes"
date_published: "2026-09-28T02:00:00.000Z"
category: "Local AI Lab"
tags: ["local-ai", "podcast", "地端模型", "lora", "開源模型", "MiMo-V2.6-Distill-Qwen-9B", "TeleOCR", "GLiNER2.5-Decide", "llama.cpp", "GGUF", "OCR"]
language: "en"
license: "CC BY 4.0"
license_url: "https://creativecommons.org/licenses/by/4.0/"
attribution: "when reusing or quoting, credit the author and link back to the original"
---

# 🧪 Xiaomi's Flagship Needs 8 GPUs, But This MiMo-V2.6-Distill-Qwen-9B Runs with a Single Command | Local AI Lab

> **TL;DR** — Xiaomi's MiMo-V2.6-Distill-Qwen-9B is a 9B model requiring only a single GPU, MIT-licensed with official GGUF support from llama.cpp, outperforming its Qwen3.

## Opening

This week, Hugging Face's trending chart got completely swept by Xiaomi's MiMo V2.6 family. The official flagship's example command calls for at least 8 GPUs, but MiMo-V2.6-Distill-Qwen-9B, the smaller sibling in the family, is the only version a regular consumer GPU actually has a shot at running. It's MIT-licensed and even comes with a GGUF build put out personally by the llama.cpp team. This episode also picked two other equally deployable models: TeleOCR, a 1.2B model that turns scanned invoices into structured tables, and GLiNER2.5-Decide, which runs on pure CPU and lets you add new categories without retraining. Coming up, I'll walk through which workflow each of these three fits best.

## Featured This Episode

### 1. MiMo-V2.6-Distill-Qwen-9B: The Only Consumer-GPU-Friendly Model in Xiaomi's Flagship Takeover Week
- **What it is**: A 9B distilled language model from Xiaomi's MiMo V2.6 family, distilled from a Qwen3.5-9B base. MIT-licensed, with the official positioning centered on three application areas: coding, visual coding, and cybersecurity.
- **Where it fits**:
  - Solo developers who want a fully offline coding agent to work on their private repos without sending any code to the cloud. The official model card lists coding as the top of four training domains, with code data making up 29.9% of the training mix.
  - The GGUF build from ggml-org ships with an mmproj vision projector file, so you can feed it a screenshot and have it replicate the layout or spot UI issues directly. The official team calls this visual coding, and visual data accounts for 27.4% of the training set.
  - Security or ops teams looking to run semi-automated internal drills for terminal operations and vulnerability analysis. The official team lists cybersecurity as one of the four training domains (cyber data makes up 14.2%), with a self-reported Terminal Bench score of 37.1.
- **Can it actually run?**: The official model card doesn't specify VRAM requirements, only noting that it can be used with compatible tools like llama.cpp, Ollama, and LM Studio to browse quantized versions, of which 53 are listed. For actual numbers, you have to check third-party quantization pages: ggml-org's Q8_0 build is 9.53 GB (the page notes this includes the Q8_0 mmproj vision encoder), while bartowski's builds come in at Q4_K_M 5.84 GB, Q5_K_M 6.88 GB, Q6_K 7.79 GB, and Q8_0 9.55 GB, with the smallest, IQ2_M, at just 3.54 GB. The rule of thumb on bartowski's page is to pick a quantization whose file size is 1 to 2 GB smaller than your GPU's total VRAM.
- **This week's buzz**: 516 on HF's trending score, 8.8k downloads, 527 likes.
- **How it stacks up**: The official model card only benchmarks against its own Qwen3.5-9B base, claiming across-the-board wins with SWE Verified 61.1, SWE Pro 44.6, AutomationBench 30.3, and Terminal Bench 37.1, all self-reported figures. The real dividing line, though, is whether you can even self-host it. The flagship released alongside it, MiMo-V2.6-Pro-RL, is described in its own model card as a Sparse MoE architecture with 1.02T total parameters / 42B active, and its example command uses a tensor-parallel-size of 8 GPUs, meaning most people can only watch from the sidelines.
- **How to get it**: ggml-org's GGUF page gives a one-line command, `ollama run hf.co/ggml-org/MiMo-V2.6-Distill-Qwen-9B-GGUF:Q8_0`, or you can use llama.cpp's `llama serve -hf ggml-org/MiMo-V2.6-Distill-Qwen-9B-GGUF:Q8_0`. If you want to self-host an API, the official model card provides the command `sglang serve --model-path XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B --reasoning-parser mimo`, and it also supports vLLM and Docker.
- **Link**: [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B)

### 2. TeleOCR: A 1.2B Model That Turns Invoices, Contracts, and Lecture Notes into Structured Text
- **What it is**: A document parsing model from XingChen-AGI, roughly 1.2B parameters, BF16, apache-2.0 licensed. It handles text extraction, table parsing, and layout/reading-order restoration within a single unified framework.
- **Where it fits**:
  - Admin or accounting staff batch-converting phone-photographed invoices, receipts, and quotes into structured tables. The model card says tables are output in OTSL format, with an accompanying tool to convert them into HTML tables that can be imported directly into a spreadsheet or database.
  - Teams wanting to convert years of accumulated PDF specs and scanned contracts into markdown for a RAG knowledge base. The official team emphasizes that it's a unified framework handling text, tables, and layout restoration simultaneously, so there's no need to stitch together several separate tools.
  - Teachers or graduate students converting lecture notes and paper PDFs containing math formulas into editable text. The model card states that formulas are output as LaTeX wrapped in $$..$$, so there's no need to manually retype equations.
- **Can it actually run?**: The official model card doesn't specify VRAM or hardware requirements, only listing roughly 1.2B parameters and a BF16 tensor type. Quantization comes from the community: the card credits it directly, saying "Thanks to Nandraj for the GGUF conversion and llama.cpp support!" The official team also offers vLLM, SGLang, and Docker deployment options. At the 1.2B scale, a quantized version should reasonably fit on a consumer GPU, but that's just an inference from the parameter count; the official team doesn't provide actual figures.
- **This week's buzz**: 526 on HF's trending score, 27.8k downloads, 627 likes.
- **How it stacks up**: In the model card and the accompanying paper (arXiv 2608.12898), the authors claim it beats pipeline-based approaches like MinerU 2.5 Pro, PaddleOCR-VL 1.6, and GLM-OCR, as well as the end-to-end OvisOCR2, on OmniDocBench v1.6, with officially listed scores of 96.87 overall and 97.05 on Table TEDS, all self-reported numbers. Two things worth flagging honestly: the model card only tags Chinese and English as supported languages, without specifying how well it handles Traditional Chinese specifically; and the page also mentions "We have renamed NaviDC-OCR to TeleOCR," while a weights repo of the same name also exists under StarDoc-AI on Hugging Face, so it's hard to say definitively which one is the sole official source.
- **How to get it**: The model card provides `pip install transformers torch pillow`, then load it with `AutoProcessor.from_pretrained(..., trust_remote_code=True)` paired with `AutoModel.from_pretrained(..., trust_remote_code=True, torch_dtype=torch.bfloat16).cuda().eval()`. If you want a lightweight, fully local route, you can use the community GGUF build linked on the card together with llama.cpp, LM Studio, or Ollama.
- **Link**: [XingChen-AGI/TeleOCR](https://huggingface.co/XingChen-AGI/TeleOCR)

### 3. GLiNER2.5-Decide: A 340M Decision Classifier That Runs on Pure CPU
- **What it is**: A small classification model from fastino, built on a DeBERTa-v3-large encoder with 340M parameters, apache-2.0 licensed. It judges multiple classification fields in a single forward pass, with labels passed in at call time rather than baked into training, so adding a new category requires no retraining.
- **Where it fits**:
  - Judging intent, priority, and whether to escalate to a human agent all at once when a support ticket comes in. The model card's example is a hotel complaint scenario, returning intent, priority, needs_human, and a multi-label topics field all in one pass.
  - Serving as a front-end classifier for an LLM router, using the CPU to first decide which model or pipeline a given request should go to, reserving expensive large-model calls for cases that actually need them.
  - Multi-label tagging for content or logs. Since the label set is a dict passed in at call time, adding a new classification field on the fly is just a config change, no weight swap required.
- **Can it actually run?**: The model card states directly: "Runs on: CPU or GPU, through gliner2," explicitly supporting pure CPU execution. The encoder is DeBERTa-v3-large with 340M parameters (though the page's model size field lists 0.5B). The official team doesn't provide actual disk size, latency, or throughput figures; there are no speed numbers on the page at all.
- **This week's buzz**: 210 on HF's trending score, 19.8k downloads, 212 likes.
- **How it stacks up**: The comparison table on the model card pits it against fastino's own GLiNER2.5-Decide-1B (59.6%) and JevK5 (57.6%), with this 340M version coming out on top at 60.2% average accuracy, tested on fastino/fast-decisions, 17 domains with 300 questions each. The authors also draw a clear line themselves: "This release is not a general-purpose model. It does not reason, explain, or answer open questions." The most important limitation for teams here in Taiwan is that this version only handles English: the page states directly, "The suite is English. Use GLiNER2.5-multi-Decide when the input is multilingual," so for Chinese-language tickets you'd need the 287M multilingual version instead.
- **How to get it**: `pip install gliner2`, then `from gliner2 import AutoExtractor` and `model = AutoExtractor.from_pretrained('fastino/GLiNER2.5-Decide')`, then call `model.classify_text(text, label_dict)`. The label_dict can hold multiple decision fields at once, with multi_label and cls_threshold configurable per field.
- **Link**: [fastino/GLiNER2.5-Decide](https://huggingface.co/fastino/GLiNER2.5-Decide)

## Closing

The three models this episode happen to map to three different local deployment needs: MiMo-V2.6-Distill-Qwen-9B is the only one from Xiaomi's flagship lineage that fits on a consumer GPU, TeleOCR uses just 1.2B parameters to turn documents into structured data with scores that beat much larger models, and GLiNER2.5-Decide proves that classification tasks sometimes don't need a GPU at all. If you want to try one out first, GLiNER2.5-Decide has the lowest barrier to entry since it runs on CPU alone; if you're trying to save on GPU while still building a coding agent, MiMo-V2.6-Distill-Qwen-9B is worth putting on your list. See you next episode.

---

## About this article and its author

Originally published on [Mark Ku's Tech Notes](https://blog.markkulab.net/en/local-ai-lab/local-ai-lab-2026-09-28)

License: [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) — when reusing or quoting, credit the author and link back to the original

### About the author

**[Mark Ku](https://blog.markkulab.net/en/author/mark-ku)** — Software engineer

- 10+ years as a software engineer
- Built North-American e-commerce and AI SaaS subscription billing

### Free tools built by the author

All of these are free to use:

- [Free PDF Sign Tool](https://blog.markkulab.net/en/tools/pdf-sign): Online PDF sign tool — draw, type, or upload a signature, then drag, resize, and download. Everything runs in your browser; nothing is uploaded.
- [VS Code Refactory](https://blog.markkulab.net/en/tools/refactory): Refactory is a VS Code refactoring extension: 34 actions plus a 37-rule code-smell inspection layer with a Code Health dashboard, across 18 languages, backed by 534 tests. It learns your repo's conventions: where interfaces live, where DI is registered, whether 'use client' belongs. It ranks files by git churn × complexity so you know what to fix first, and hands any smell to the Claude Code already on your machine. Free to use, and your source never leaves your computer.
- [DB-Kit Database Manager](https://blog.markkulab.net/en/tools/db-kit): DB-Kit is a lightweight, cross-platform database manager built with Tauri + Rust + React. Manage MySQL, MariaDB, PostgreSQL, SQL Server, Oracle, SQLite, MongoDB, Redis, Kafka, Elasticsearch and RabbitMQ from one consistent interface: passwords encrypted in the OS keychain, SSH tunnels, full CRUD, a visual query builder, stacked multi-statement result sets, cross-connection data transfer, schema & data compare with sync SQL and schema snapshots, Excel / CSV import & export, visualized execution plans, ER diagrams, scheduled backups, SQL stress testing with p50–p99 latency percentiles, a 15-rule SQL review engine, Review & Run (AI review plus per-statement backups and an auto-generated rollback script), Kafka message browsing with monitoring & alerts, a bilingual UI (Traditional Chinese / English), a built-in AI assistant (a local CLI or any Anthropic / OpenAI-compatible API; natural-language SQL, AI review and tuning advice) and the dbk CLI. Free and open source (MIT), with installers for Windows, macOS and Linux.
- [VS Code Super Mermaid](https://blog.markkulab.net/en/tools/super-mermaid): Super Mermaid is a VS Code extension for beautiful Mermaid diagrams out of the box: auto-colored live preview, mouse pan & zoom, high-res PNG / SVG export, 21 templates and multiple themes. Free and open source (MIT).
- [React Super Mermaid](https://blog.markkulab.net/en/tools/react-super-mermaid): react-super-mermaid is an open-source React component library: render beautiful Mermaid diagrams with a single <MermaidViewer>, with built-in colorful / sketch themes, pan & zoom, in-diagram search, and high-res SVG / PNG export. Lightweight, SSR-safe, fully typed. Free and open source (MIT).
- [Jira / Confluence Super Mermaid](https://blog.markkulab.net/en/tools/jira-super-mermaid): An Atlassian Forge app: write Mermaid syntax directly inside a Jira issue or a Confluence page and get flowcharts, sequence diagrams, state machines and Gantt charts. 11 diagram types, SVG / PNG export, light and dark themes, full CJK support. Runs on Atlassian: your diagrams live in your own site and the app calls no third-party service. Free, coming soon to the Atlassian Marketplace.
- [Mermaid Live Preview](https://blog.markkulab.net/en/tools/mermaid-preview): Write Mermaid in your browser, see it render instantly, and share the whole diagram as a single link. No sign-up, nothing uploaded to a server, and mermaid.live share links work as-is.
- [React Intl Phone Number](https://blog.markkulab.net/en/tools/react-intl-phone-number): react-intl-phone-number is an open-source React component: framework-agnostic and antd-free, with E.164 in/out, a searchable flag / country-code dropdown, configurable validation levels (strict / mobile-strict / loose), themeable CSS, and i18n — phone logic powered by google-libphonenumber. Lightweight and fully typed. Free and open source (MIT).
- [Uptime Kuma Cluster](https://blog.markkulab.net/en/tools/uptime-kuma-cluster): Turn single-node Uptime Kuma into a highly available cluster: OpenResty + Lua smart load balancing, shared MariaDB state, health checks and automatic failover, plus cluster-management REST APIs. One Docker Compose command to start. Free and open source (MIT).
- [AI Podcast Cut](https://blog.markkulab.net/en/tools/ai-podcast-cut): Drop in a recording and it removes fillers and stutters, levels loudness segment by segment, and sends a second agent to review every cut. Cut points snap to word boundaries and zero crossings, every splice gets a fade, and sentence-end breaths are preserved. Desktop app for Windows, macOS and Linux. MIT licensed; the Windows installer bundles ffmpeg.
- [open-pos restaurant POS](https://blog.markkulab.net/en/tools/open-pos): One computer and one receipt printer is enough to open the shop. Your data lives on your own disk, no subscription, no lock-in, MIT licensed. Money is integer New Taiwan dollars with tax split by the statutory formula, so sales + tax always equals the total. Printing goes straight over ESC/POS on TCP 9100, with no vendor driver. Tauri + Rust + SQLite desktop app, v1.0 in development.

### Daily podcasts

- [Mark's Tech Insights — Daily AI News](https://blog.markkulab.net/en/category/tech-news): Daily curated AI and tech trends. Catch the latest developments via audio summaries — covering AI applications, software architecture, DevOps, and engineering practice. — RSS: https://blog.markkulab.net/feed.xml
- [AI股市蝦聊](https://blog.markkulab.net/en/category/ai-stock-chat): Every trading day, an AI-analyzed take on the Taiwan stock market, delivered as a two-host conversation covering the session and the next-day outlook. — RSS: https://blog.markkulab.net/ai-stock-chat/feed.xml
- [開源好物週報](https://blog.markkulab.net/en/category/open-source-weekly): A weekly two-host pick of free open-source tools surfaced from real Hacker News, GitHub, and Reddit buzz — what pain they solve and the fastest way to get started. — RSS: https://blog.markkulab.net/open-source-weekly/feed.xml
- [AI 運動週報](https://blog.markkulab.net/en/category/sports-weekly): Two hosts talk NBA, MLB and world sport three times a week — scores, records and the stories behind them, from a Taiwanese fan perspective. — RSS: https://blog.markkulab.net/sports-weekly/feed.xml
- [AI 國際新聞快報](https://blog.markkulab.net/en/category/world-news): A daily two-host briefing that makes sense of the past 24 hours in world news — geopolitics, the global economy, conflict and security, disasters and climate — from a Taiwanese perspective, neutral and fully sourced. — RSS: https://blog.markkulab.net/world-news/feed.xml
- [地端 AI 實驗室](https://blog.markkulab.net/en/category/local-ai-lab): Twice a week, a two-host look at open-weight models and LoRAs you can actually run on your own machine: what they are for, whether your GPU can handle them, and how they compare — sourced from official model cards. — RSS: https://blog.markkulab.net/local-ai-lab/feed.xml

### Deals

- [Saily eSIM](https://blog.markkulab.net/en/saily): Travel eSIM by the NordVPN team，Promo code：KUKU
- [NordVPN](https://blog.markkulab.net/en/nordvpn): The world's leading VPN, independently audited
- [PremLogin](https://blog.markkulab.net/en/premlogin): Subscription sharing for streaming and AI seats，Promo code：markku666

### Newsletter

[Subscribe to the newsletter](https://blog.markkulab.net/en/subscribe) — Be the first to know about new posts. No spam, unsubscribe anytime.
