Skip to main content
Mark Ku's Blog

Opening

This week, Hugging Face's trending chart got completely swept by Xiaomi's MiMo V2.6 family. The official flagship's example command calls for at least 8 GPUs, but MiMo-V2.6-Distill-Qwen-9B, the smaller sibling in the family, is the only version a regular consumer GPU actually has a shot at running. It's MIT-licensed and even comes with a GGUF build put out personally by the llama.cpp team. This episode also picked two other equally deployable models: TeleOCR, a 1.2B model that turns scanned invoices into structured tables, and GLiNER2.5-Decide, which runs on pure CPU and lets you add new categories without retraining. Coming up, I'll walk through which workflow each of these three fits best.

1. MiMo-V2.6-Distill-Qwen-9B: The Only Consumer-GPU-Friendly Model in Xiaomi's Flagship Takeover Week

  • What it is: A 9B distilled language model from Xiaomi's MiMo V2.6 family, distilled from a Qwen3.5-9B base. MIT-licensed, with the official positioning centered on three application areas: coding, visual coding, and cybersecurity.
  • Where it fits:
    • Solo developers who want a fully offline coding agent to work on their private repos without sending any code to the cloud. The official model card lists coding as the top of four training domains, with code data making up 29.9% of the training mix.
    • The GGUF build from ggml-org ships with an mmproj vision projector file, so you can feed it a screenshot and have it replicate the layout or spot UI issues directly. The official team calls this visual coding, and visual data accounts for 27.4% of the training set.
    • Security or ops teams looking to run semi-automated internal drills for terminal operations and vulnerability analysis. The official team lists cybersecurity as one of the four training domains (cyber data makes up 14.2%), with a self-reported Terminal Bench score of 37.1.
  • Can it actually run?: The official model card doesn't specify VRAM requirements, only noting that it can be used with compatible tools like llama.cpp, Ollama, and LM Studio to browse quantized versions, of which 53 are listed. For actual numbers, you have to check third-party quantization pages: ggml-org's Q8_0 build is 9.53 GB (the page notes this includes the Q8_0 mmproj vision encoder), while bartowski's builds come in at Q4_K_M 5.84 GB, Q5_K_M 6.88 GB, Q6_K 7.79 GB, and Q8_0 9.55 GB, with the smallest, IQ2_M, at just 3.54 GB. The rule of thumb on bartowski's page is to pick a quantization whose file size is 1 to 2 GB smaller than your GPU's total VRAM.
  • This week's buzz: 516 on HF's trending score, 8.8k downloads, 527 likes.
  • How it stacks up: The official model card only benchmarks against its own Qwen3.5-9B base, claiming across-the-board wins with SWE Verified 61.1, SWE Pro 44.6, AutomationBench 30.3, and Terminal Bench 37.1, all self-reported figures. The real dividing line, though, is whether you can even self-host it. The flagship released alongside it, MiMo-V2.6-Pro-RL, is described in its own model card as a Sparse MoE architecture with 1.02T total parameters / 42B active, and its example command uses a tensor-parallel-size of 8 GPUs, meaning most people can only watch from the sidelines.
  • How to get it: ggml-org's GGUF page gives a one-line command, ollama run hf.co/ggml-org/MiMo-V2.6-Distill-Qwen-9B-GGUF:Q8_0, or you can use llama.cpp's llama serve -hf ggml-org/MiMo-V2.6-Distill-Qwen-9B-GGUF:Q8_0. If you want to self-host an API, the official model card provides the command sglang serve --model-path XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B --reasoning-parser mimo, and it also supports vLLM and Docker.
  • Link: XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B

2. TeleOCR: A 1.2B Model That Turns Invoices, Contracts, and Lecture Notes into Structured Text

  • What it is: A document parsing model from XingChen-AGI, roughly 1.2B parameters, BF16, apache-2.0 licensed. It handles text extraction, table parsing, and layout/reading-order restoration within a single unified framework.
  • Where it fits:
    • Admin or accounting staff batch-converting phone-photographed invoices, receipts, and quotes into structured tables. The model card says tables are output in OTSL format, with an accompanying tool to convert them into HTML tables that can be imported directly into a spreadsheet or database.
    • Teams wanting to convert years of accumulated PDF specs and scanned contracts into markdown for a RAG knowledge base. The official team emphasizes that it's a unified framework handling text, tables, and layout restoration simultaneously, so there's no need to stitch together several separate tools.
    • Teachers or graduate students converting lecture notes and paper PDFs containing math formulas into editable text. The model card states that formulas are output as LaTeX wrapped in ...., so there's no need to manually retype equations.
  • Can it actually run?: The official model card doesn't specify VRAM or hardware requirements, only listing roughly 1.2B parameters and a BF16 tensor type. Quantization comes from the community: the card credits it directly, saying "Thanks to Nandraj for the GGUF conversion and llama.cpp support!" The official team also offers vLLM, SGLang, and Docker deployment options. At the 1.2B scale, a quantized version should reasonably fit on a consumer GPU, but that's just an inference from the parameter count; the official team doesn't provide actual figures.
  • This week's buzz: 526 on HF's trending score, 27.8k downloads, 627 likes.
  • How it stacks up: In the model card and the accompanying paper (arXiv 2608.12898), the authors claim it beats pipeline-based approaches like MinerU 2.5 Pro, PaddleOCR-VL 1.6, and GLM-OCR, as well as the end-to-end OvisOCR2, on OmniDocBench v1.6, with officially listed scores of 96.87 overall and 97.05 on Table TEDS, all self-reported numbers. Two things worth flagging honestly: the model card only tags Chinese and English as supported languages, without specifying how well it handles Traditional Chinese specifically; and the page also mentions "We have renamed NaviDC-OCR to TeleOCR," while a weights repo of the same name also exists under StarDoc-AI on Hugging Face, so it's hard to say definitively which one is the sole official source.
  • How to get it: The model card provides pip install transformers torch pillow, then load it with AutoProcessor.from_pretrained(..., trust_remote_code=True) paired with AutoModel.from_pretrained(..., trust_remote_code=True, torch_dtype=torch.bfloat16).cuda().eval(). If you want a lightweight, fully local route, you can use the community GGUF build linked on the card together with llama.cpp, LM Studio, or Ollama.
  • Link: XingChen-AGI/TeleOCR

3. GLiNER2.5-Decide: A 340M Decision Classifier That Runs on Pure CPU

  • What it is: A small classification model from fastino, built on a DeBERTa-v3-large encoder with 340M parameters, apache-2.0 licensed. It judges multiple classification fields in a single forward pass, with labels passed in at call time rather than baked into training, so adding a new category requires no retraining.
  • Where it fits:
    • Judging intent, priority, and whether to escalate to a human agent all at once when a support ticket comes in. The model card's example is a hotel complaint scenario, returning intent, priority, needs_human, and a multi-label topics field all in one pass.
    • Serving as a front-end classifier for an LLM router, using the CPU to first decide which model or pipeline a given request should go to, reserving expensive large-model calls for cases that actually need them.
    • Multi-label tagging for content or logs. Since the label set is a dict passed in at call time, adding a new classification field on the fly is just a config change, no weight swap required.
  • Can it actually run?: The model card states directly: "Runs on: CPU or GPU, through gliner2," explicitly supporting pure CPU execution. The encoder is DeBERTa-v3-large with 340M parameters (though the page's model size field lists 0.5B). The official team doesn't provide actual disk size, latency, or throughput figures; there are no speed numbers on the page at all.
  • This week's buzz: 210 on HF's trending score, 19.8k downloads, 212 likes.
  • How it stacks up: The comparison table on the model card pits it against fastino's own GLiNER2.5-Decide-1B (59.6%) and JevK5 (57.6%), with this 340M version coming out on top at 60.2% average accuracy, tested on fastino/fast-decisions, 17 domains with 300 questions each. The authors also draw a clear line themselves: "This release is not a general-purpose model. It does not reason, explain, or answer open questions." The most important limitation for teams here in Taiwan is that this version only handles English: the page states directly, "The suite is English. Use GLiNER2.5-multi-Decide when the input is multilingual," so for Chinese-language tickets you'd need the 287M multilingual version instead.
  • How to get it: pip install gliner2, then from gliner2 import AutoExtractor and model = AutoExtractor.from_pretrained('fastino/GLiNER2.5-Decide'), then call model.classify_text(text, label_dict). The label_dict can hold multiple decision fields at once, with multi_label and cls_threshold configurable per field.
  • Link: fastino/GLiNER2.5-Decide

Closing

The three models this episode happen to map to three different local deployment needs: MiMo-V2.6-Distill-Qwen-9B is the only one from Xiaomi's flagship lineage that fits on a consumer GPU, TeleOCR uses just 1.2B parameters to turn documents into structured data with scores that beat much larger models, and GLiNER2.5-Decide proves that classification tasks sometimes don't need a GPU at all. If you want to try one out first, GLiNER2.5-Decide has the lowest barrier to entry since it runs on CPU alone; if you're trying to save on GPU while still building a coding agent, MiMo-V2.6-Distill-Qwen-9B is worth putting on your list. See you next episode.

Author

Mark Ku

10 年以上的軟體工程師,做過北美電商與 AI SaaS 訂閱收費系統。現在經營貳陸資訊有限公司(www.226network.com),幫小公司做系統、網站、LINE BOT 與 AI 自動化,也在這裡分享開發筆記與開源工具。Read More

Found this useful?

The author's free tools, daily podcasts and newsletter are all here.

Mark Ku · This article is licensed under CC BY 4.0. Credit the author and link back to the original when reusing it.

Comments

Subscribe to Newsletter

Subscribe to get new posts delivered instantly — never miss a tech share.

By submitting, you agree to receive emails. You can anytime.

Popular Posts

View all
Mark Ku
··637

Oracle Cloud Always Free Tier: Linux Host and Static IP for a $0 Cloud Solution

Oracle Cloud Always Free Tier: Linux Host and Static IP for a $0 Cloud Solution
Mark Ku
··477

Say Goodbye to Postman's Fee Trap! A Hands-on Guide to Bruno, the Open-Source Git-Native API Testing Powerhouse.

Say Goodbye to Postman's Fee Trap! A Hands-on Guide to Bruno, the Open-Source Git-Native API Testing Powerhouse.
Mark Ku
··302

A Free, Open-Source, Notion-like Knowledge Base — A Complete Guide to Deploying and Backing Up Outline Wiki

A Free, Open-Source, Notion-like Knowledge Base — A Complete Guide to Deploying and Backing Up Outline Wiki
Mark Ku
··227

Building an Efficient API Management Platform: Deploying Kong Gateway from Scratch - Part 1

Building an Efficient API Management Platform: Deploying Kong Gateway from Scratch - Part 1
Mark Ku
··226

Setting Up Samba on Ubuntu to Share Folders with Windows 11

Setting Up Samba on Ubuntu to Share Folders with Windows 11
Mark Ku
··216

Training Your Own AI Voice: Hardware Requirements, Open-Source Model Comparison, and LoRA Fine-Tuning

Training Your Own AI Voice: Hardware Requirements, Open-Source Model Comparison, and LoRA Fine-Tuning