Mark Ku's Blog

The author's company

Vibe Coding architecture & coaching

Built internal tools with AI but afraid nobody can fix or change them safely? 226 Network helps you put the code in Git, keep secrets out of it, move schedules onto a stable host, and add handover docs and tests.

Opening

The most eye-popping download count this episode doesn't belong to a chat model, it belongs to a decision LoRA that's already blown past 900K downloads: GEV-26B-Decide, which lets an agent decide for itself whether it needs to "think a bit harder." We'll tell you exactly how much difference turning on reasoning makes. We've also got a speech recognition model that runs on CPU, moondream/parakeet-redux, and an ultra-lightweight decision model that only needs 1.6 GB of memory, Phocinae-Largha-150M-v1.

  • Dates: Oct 10 to Oct 18, 2026
  • Offer: buy 12 months of an eligible annual plan and get 1 month free (13 months total)
  • Eligible: Netflix, YouTube, Disney+, HBO Max, Amazon Prime, Spotify, Tidal, Canva Pro, Office 365, Duolingo, plus selected ChatGPT, Grok and Perplexity plans
  • Not eligible: standalone-account and top-up plans for ChatGPT, Grok and Perplexity, and Claude top-up plans
  • New-customer code: markku666 (5% off the first order; whether it stacks with the sale is shown at checkout)
  • Sale link: https://premlogin.com/?aff=hrQOPI0r (affiliate link)
  • Full review and risks: https://blog.markkulab.net/en/premlogin

PremLogin is a third-party subscription-sharing platform. It sells shared seats and is not an official authorized reseller, so read the usage rules before you buy.

This Episode's Picks

1. moondream/parakeet-redux: A 178 MB speech recognition model that runs without a GPU

  • What it is: A speech recognition (ASR) model. The official model card describes it as a version of parakeet-tdt-0.6b-v3 converted to 1.58-bit ternary weights (only three possible values: −1, 0, +1). The parameter count is listed as 0.1B, the weight file is only 178 MB, and it's licensed under CC-BY-4.0.
  • Where it fits:
    • If you've got a pile of English conference talks or online course videos and want to batch-generate subtitle files on a NAS or an old laptop with no discrete GPU, the model card says it supports segment- or word-level timestamps, which is exactly what you need to produce subtitle tracks.
    • Engineers building voice-memo or recording apps who want offline transcription as a built-in feature, without sending users' recordings to a cloud API. At 178 MB and running on the Photon runtime (the official docs say "AVX-512 VNNI on x86, NEON on ARM, Metal on Apple GPUs"), it can be bundled directly into a desktop or mobile app.
    • Processing multi-hour interviews or meeting recordings. The model card mentions built-in voice-activity detection for segmenting long audio, and lists a WER of 2.51 on long-form audio (TED-LIUM).
    • One thing worth flagging up front: Chinese isn't among the 25 languages the official docs list as supported, so this is a tool for English and European-language material, not a Chinese transcription solution.
  • Can your hardware run it: The example code in the official model card uses device="cpu", running on the Photon runtime. VRAM requirements aren't officially specified, but given the 178 MB of ternary weights and 0.1B parameter count, this is likely the lowest-barrier model of the episode.
  • This episode's buzz: HF trending score of 88, 14.1K downloads, 265 likes.
  • How it compares: Against its own original, parakeet-tdt-0.6b-v3, the model card says it outperforms the original on FLEURS and long-form audio, and for English it "stays within 0.3 WER of the original," though it does worse in noisy environments (9.04 vs. 6.72). No comparison against Whisper is provided in the card.
  • How to get it: The README's example is to first import moondream as md, then load it with with md.photon("moondream/parakeet-redux", device="cpu") as speech, then call speech.transcribe(audio="speech.wav") to get back result["text"].
  • Link: huggingface.co/moondream/parakeet-redux

2. autotrust/GEV-26B-Decide: A decision LoRA that lets an agent decide for itself whether to think harder

  • What it is: This isn't a repackage of someone else's weights, it's a LoRA adapter plus decision head that the authors trained themselves on top of Gemma-4-26B-A4B-it, purpose-built for structured decision tasks. It supports yes/no, 2-to-256-option, and 0-to-5 rating tasks. The stated license is Apache-2.0 (though this only covers the adapter, head, and calibration files; the base model is still bound by the Gemma 4 terms).
  • Where it fits:
    • Agent developers who want the system to decide for itself whether a given query needs extra thought. The model card describes "adaptive thinking," which only turns on reasoning when the model is uncertain, and lists results showing GPQA Diamond jumping from 42.9% to 78.6% and CRUXEval from 67.5% to 90.7% once reasoning is enabled, while a plain System 1 pass handles a decision in roughly 45 ms.
    • Anyone building computer-use or GUI automation who needs a model to decide "where do I click next." The comparison table in the model card shows it matching JEV-27B-VL on computer-use success rate (both 95%), but at 85 ms per click versus 260 ms, it's considerably faster.
    • Teams that need a single set of weights to handle both text and screenshot inputs for decision-making. The model card states: "One set of weights, one vLLM engine, for text and images."
  • Can your hardware run it: No explicit VRAM threshold is given officially; the model card only says testing was done on a single B200. What we can confirm is the base model architecture, "on Gemma-4-26B-A4B-it (26B parameters, ≈4B active per token)," an MoE setup, with a separate 24-slot fp32 decision head. To actually run it, the card notes that vLLM needs Gemma-4 support, and requires a LoRA patch for the tied lm_head.
  • This episode's buzz: HF trending score of 1777, 909.8K downloads, 2.1K likes, the highest buzz of this episode by far. The same family also has two quantized/repackaged variants.
  • How it compares: The model card itself names JEV-27B-VL as the comparison point. Both score 95% success on computer-use tasks, but this model is notably faster per click. That said, it flips the other way on robotic-arm pick-and-place tasks, where it trails at 40% versus 75%.
  • How to get it: The README gives two lines: hf download autotrust/GEV-26B-Decide --local-dir GEV-26B-Decide, then bash GEV-26B-Decide/serve.sh to spin up vLLM at :8000. The card also offers transformers plus peft System 1 inference code as an alternative path.
  • Link: huggingface.co/autotrust/GEV-26B-Decide

3. Phocinae/Phocinae-Largha-150M-v1: A 144M-parameter lightweight decision model light enough for CPU

  • What it is: A compact model with just 144.3M parameters and 288.6 MB of weights, licensed under Apache-2.0, built on an mmBERT-small encoder (MIT licensed). It's purpose-built for typed decisions (approval gates, routing, and similar tasks).
  • Where it fits:
    • AI agent developers who want an approval gate before every tool call. The model card says it does "one forward pass per decision," returning an answer with a confidence score. The official benchmark shows 21.0 ms per decision in fp16 on an RTX 5090, compared to 1.5+ seconds for the cloud API it's measured against.
    • Customer service or ticketing systems doing document triage and automated routing. The model card explicitly lists its use cases as approval gates, tool routing, escalation, and document triage, supporting yes/no, 2-to-10-option, and rating tasks. For Chinese specifically, the card lists a typed-decisions score of 0.848 (though it also notes the test questions were machine-translated, worth keeping in mind).
    • Budget-constrained teams running small VPS instances or CPU-only machines. The card lists a peak inference VRAM requirement of just 1.6 GB, and it even runs on a single CPU thread, albeit at 1.64 seconds per decision.
  • Can your hardware run it: Officially listed at 1.6 GB peak inference VRAM, 21.0 ms per decision in fp16 on a GPU (RTX 5090), or 1.64 seconds per decision on a single CPU thread. This is one of the two lowest-barrier models in this episode.
  • This episode's buzz: HF trending score of 107, 89 downloads, 117 likes.
  • How it compares: The card's own comparison lists zero-shot typed-decisions scores of 0.766 for Laya, 0.727 for JEV, and 0.768 for meraGPT, versus 0.906 for its own specialized version. That said, the card notes this figure was "fitted on the training split" rather than a true zero-shot result. Worth noting: the same card also openly discloses that it only scored 0.5455 (126/231) on JevBench public-231, below the 58.4% passing threshold, a fairly honest disclosure of a benchmark it didn't clear.
  • How to get it: Per the README, the flow is hf download Phocinae/Phocinae-Largha-150M-v1 --local-dir ./largha to grab the weights, then pip install phocinae-server, and spin up the service with PHOC_MODEL_DIR=./largha python -m phocinae.main. Alternatively, you can call it directly through the Python API with Engine("./largha", device="auto") invoking eng.run(state, questions).
  • Link: huggingface.co/Phocinae/Phocinae-Largha-150M-v1

Closing

These three models happen to map neatly onto different pain points in agent development: GEV-26B-Decide trades a larger parameter count for higher-precision judgment, Phocinae-Largha-150M-v1 trades an extremely low barrier for a cheap approval gate, and moondream/parakeet-redux lets offline speech transcription stop relying on cloud APIs altogether. If you don't have a GPU on hand and want to try one first, Phocinae-Largha-150M-v1 is probably the lowest-barrier starting point. Next episode, we'll dig up more gems from the local AI scene.

Author

Mark Ku

10 年以上的軟體工程師,做過北美電商與 AI SaaS 訂閱收費系統。Read More

Found this useful?

The author's free tools, daily podcasts and newsletter are all here.

Mark Ku · This article is licensed under CC BY 4.0. Credit the author and link back to the original when reusing it.

Comments

Subscribe to Newsletter

Subscribe to get new posts delivered instantly — never miss a tech share.

By submitting, you agree to receive emails. You can anytime.

Popular Posts

View all
Mark Ku
··637

Oracle Cloud Always Free Tier: Linux Host and Static IP for a $0 Cloud Solution

Oracle Cloud Always Free Tier: Linux Host and Static IP for a $0 Cloud Solution
Mark Ku
··472

Say Goodbye to Postman's Fee Trap! A Hands-on Guide to Bruno, the Open-Source Git-Native API Testing Powerhouse.

Say Goodbye to Postman's Fee Trap! A Hands-on Guide to Bruno, the Open-Source Git-Native API Testing Powerhouse.
Mark Ku
··268

A Free, Open-Source, Notion-like Knowledge Base — A Complete Guide to Deploying and Backing Up Outline Wiki

A Free, Open-Source, Notion-like Knowledge Base — A Complete Guide to Deploying and Backing Up Outline Wiki
Mark Ku
··215

Building an Efficient API Management Platform: Deploying Kong Gateway from Scratch - Part 1

Building an Efficient API Management Platform: Deploying Kong Gateway from Scratch - Part 1
Mark Ku
··210

Training Your Own AI Voice: Hardware Requirements, Open-Source Model Comparison, and LoRA Fine-Tuning

Training Your Own AI Voice: Hardware Requirements, Open-Source Model Comparison, and LoRA Fine-Tuning
Mark Ku
··209

Setting Up Samba on Ubuntu to Share Folders with Windows 11

Setting Up Samba on Ubuntu to Share Folders with Windows 11