---
title: "🧪 The most eye-popping download numbers this episode aren't a chat model, they're a decision LoRA: how GEV-26B-Decide lets the agent decide for itself whether to think more｜Local AI Lab"
description: "This episode's most astonishing download count isn't a chat model, it's a decision LoRA that's already surpassed 900,000 downloads: GEV-26B-Decide, which lets the agent judge for itself whether to think a bit longer, and I'll tell you in a moment how much the difference is once reasoning is turned on. Also featured: moondream/parakeet-redux, a speech recognition model that runs on CPU, and Phocinae-Largha-150M-v1, an ultra-lightweight decision model needing only 1.6 GB of memory."
canonical_url: "https://blog.markkulab.net/en/local-ai-lab/local-ai-lab-2026-10-09"
author: "Mark Ku"
author_url: "https://blog.markkulab.net/en/author/mark-ku"
site: "Mark Ku's Tech Notes"
date_published: "2026-10-09T02:00:00.000Z"
category: "Local AI Lab"
tags: ["local-ai", "podcast", "地端模型", "lora", "開源模型", "moondream", "parakeet-redux", "GEV-26B-Decide", "Phocinae-Largha-150M-v1", "LoRA", "語音辨識"]
language: "en"
license: "CC BY 4.0"
license_url: "https://creativecommons.org/licenses/by/4.0/"
attribution: "when reusing or quoting, credit the author and link back to the original"
---

# 🧪 The most eye-popping download numbers this episode aren't a chat model, they're a decision LoRA: how GEV-26B-Decide lets the agent decide for itself whether to think more｜Local AI Lab

> **TL;DR** — GEV-26B-Decide, a LoRA adapter on Gemma-4-26B with 909.8K downloads, lets agents decide whether to enable reasoning for harder problems, improving GPQA Diamond from 42.9% to 78.

## Opening

The most eye-popping download count this episode doesn't belong to a chat model, it belongs to a decision LoRA that's already blown past 900K downloads: GEV-26B-Decide, which lets an agent decide for itself whether it needs to "think a bit harder." We'll tell you exactly how much difference turning on reasoning makes. We've also got a speech recognition model that runs on CPU, moondream/parakeet-redux, and an ultra-lightweight decision model that only needs 1.6 GB of memory, Phocinae-Largha-150M-v1.

## Sponsor note: PremLogin Double Ten sale
- **Dates**: Oct 10 to Oct 18, 2026
- **Offer**: buy 12 months of an eligible annual plan and get 1 month free (13 months total)
- **Eligible**: Netflix, YouTube, Disney+, HBO Max, Amazon Prime, Spotify, Tidal, Canva Pro, Office 365, Duolingo, plus selected ChatGPT, Grok and Perplexity plans
- **Not eligible**: standalone-account and top-up plans for ChatGPT, Grok and Perplexity, and Claude top-up plans
- **New-customer code**: `markku666` (5% off the first order; whether it stacks with the sale is shown at checkout)
- **Sale link**: <https://premlogin.com/?aff=hrQOPI0r> (affiliate link)
- **Full review and risks**: <https://blog.markkulab.net/en/premlogin>

PremLogin is a third-party subscription-sharing platform. It sells shared seats and is not an official authorized reseller, so read the usage rules before you buy.

## This Episode's Picks

### 1. moondream/parakeet-redux: A 178 MB speech recognition model that runs without a GPU

- **What it is**: A speech recognition (ASR) model. The official model card describes it as a version of parakeet-tdt-0.6b-v3 converted to 1.58-bit ternary weights (only three possible values: −1, 0, +1). The parameter count is listed as 0.1B, the weight file is only 178 MB, and it's licensed under CC-BY-4.0.
- **Where it fits**:
  - If you've got a pile of English conference talks or online course videos and want to batch-generate subtitle files on a NAS or an old laptop with no discrete GPU, the model card says it supports segment- or word-level timestamps, which is exactly what you need to produce subtitle tracks.
  - Engineers building voice-memo or recording apps who want offline transcription as a built-in feature, without sending users' recordings to a cloud API. At 178 MB and running on the Photon runtime (the official docs say "AVX-512 VNNI on x86, NEON on ARM, Metal on Apple GPUs"), it can be bundled directly into a desktop or mobile app.
  - Processing multi-hour interviews or meeting recordings. The model card mentions built-in voice-activity detection for segmenting long audio, and lists a WER of 2.51 on long-form audio (TED-LIUM).
  - One thing worth flagging up front: Chinese isn't among the 25 languages the official docs list as supported, so this is a tool for English and European-language material, not a Chinese transcription solution.
- **Can your hardware run it**: The example code in the official model card uses `device="cpu"`, running on the Photon runtime. VRAM requirements aren't officially specified, but given the 178 MB of ternary weights and 0.1B parameter count, this is likely the lowest-barrier model of the episode.
- **This episode's buzz**: HF trending score of 88, 14.1K downloads, 265 likes.
- **How it compares**: Against its own original, parakeet-tdt-0.6b-v3, the model card says it outperforms the original on FLEURS and long-form audio, and for English it "stays within 0.3 WER of the original," though it does worse in noisy environments (9.04 vs. 6.72). No comparison against Whisper is provided in the card.
- **How to get it**: The README's example is to first `import moondream as md`, then load it with `with md.photon("moondream/parakeet-redux", device="cpu") as speech`, then call `speech.transcribe(audio="speech.wav")` to get back `result["text"]`.
- **Link**: [huggingface.co/moondream/parakeet-redux](https://huggingface.co/moondream/parakeet-redux)

### 2. autotrust/GEV-26B-Decide: A decision LoRA that lets an agent decide for itself whether to think harder

- **What it is**: This isn't a repackage of someone else's weights, it's a LoRA adapter plus decision head that the authors trained themselves on top of Gemma-4-26B-A4B-it, purpose-built for structured decision tasks. It supports yes/no, 2-to-256-option, and 0-to-5 rating tasks. The stated license is Apache-2.0 (though this only covers the adapter, head, and calibration files; the base model is still bound by the Gemma 4 terms).
- **Where it fits**:
  - Agent developers who want the system to decide for itself whether a given query needs extra thought. The model card describes "adaptive thinking," which only turns on reasoning when the model is uncertain, and lists results showing GPQA Diamond jumping from 42.9% to 78.6% and CRUXEval from 67.5% to 90.7% once reasoning is enabled, while a plain System 1 pass handles a decision in roughly 45 ms.
  - Anyone building computer-use or GUI automation who needs a model to decide "where do I click next." The comparison table in the model card shows it matching JEV-27B-VL on computer-use success rate (both 95%), but at 85 ms per click versus 260 ms, it's considerably faster.
  - Teams that need a single set of weights to handle both text and screenshot inputs for decision-making. The model card states: "One set of weights, one vLLM engine, for text and images."
- **Can your hardware run it**: No explicit VRAM threshold is given officially; the model card only says testing was done on a single B200. What we can confirm is the base model architecture, "on Gemma-4-26B-A4B-it (26B parameters, ≈4B active per token)," an MoE setup, with a separate 24-slot fp32 decision head. To actually run it, the card notes that vLLM needs Gemma-4 support, and requires a LoRA patch for the tied lm_head.
- **This episode's buzz**: HF trending score of 1777, 909.8K downloads, 2.1K likes, the highest buzz of this episode by far. The same family also has two quantized/repackaged variants.
- **How it compares**: The model card itself names JEV-27B-VL as the comparison point. Both score 95% success on computer-use tasks, but this model is notably faster per click. That said, it flips the other way on robotic-arm pick-and-place tasks, where it trails at 40% versus 75%.
- **How to get it**: The README gives two lines: `hf download autotrust/GEV-26B-Decide --local-dir GEV-26B-Decide`, then `bash GEV-26B-Decide/serve.sh` to spin up vLLM at `:8000`. The card also offers transformers plus peft System 1 inference code as an alternative path.
- **Link**: [huggingface.co/autotrust/GEV-26B-Decide](https://huggingface.co/autotrust/GEV-26B-Decide)

### 3. Phocinae/Phocinae-Largha-150M-v1: A 144M-parameter lightweight decision model light enough for CPU

- **What it is**: A compact model with just 144.3M parameters and 288.6 MB of weights, licensed under Apache-2.0, built on an mmBERT-small encoder (MIT licensed). It's purpose-built for typed decisions (approval gates, routing, and similar tasks).
- **Where it fits**:
  - AI agent developers who want an approval gate before every tool call. The model card says it does "one forward pass per decision," returning an answer with a confidence score. The official benchmark shows 21.0 ms per decision in fp16 on an RTX 5090, compared to 1.5+ seconds for the cloud API it's measured against.
  - Customer service or ticketing systems doing document triage and automated routing. The model card explicitly lists its use cases as approval gates, tool routing, escalation, and document triage, supporting yes/no, 2-to-10-option, and rating tasks. For Chinese specifically, the card lists a typed-decisions score of 0.848 (though it also notes the test questions were machine-translated, worth keeping in mind).
  - Budget-constrained teams running small VPS instances or CPU-only machines. The card lists a peak inference VRAM requirement of just 1.6 GB, and it even runs on a single CPU thread, albeit at 1.64 seconds per decision.
- **Can your hardware run it**: Officially listed at 1.6 GB peak inference VRAM, 21.0 ms per decision in fp16 on a GPU (RTX 5090), or 1.64 seconds per decision on a single CPU thread. This is one of the two lowest-barrier models in this episode.
- **This episode's buzz**: HF trending score of 107, 89 downloads, 117 likes.
- **How it compares**: The card's own comparison lists zero-shot typed-decisions scores of 0.766 for Laya, 0.727 for JEV, and 0.768 for meraGPT, versus 0.906 for its own specialized version. That said, the card notes this figure was "fitted on the training split" rather than a true zero-shot result. Worth noting: the same card also openly discloses that it only scored 0.5455 (126/231) on JevBench public-231, below the 58.4% passing threshold, a fairly honest disclosure of a benchmark it didn't clear.
- **How to get it**: Per the README, the flow is `hf download Phocinae/Phocinae-Largha-150M-v1 --local-dir ./largha` to grab the weights, then `pip install phocinae-server`, and spin up the service with `PHOC_MODEL_DIR=./largha python -m phocinae.main`. Alternatively, you can call it directly through the Python API with `Engine("./largha", device="auto")` invoking `eng.run(state, questions)`.
- **Link**: [huggingface.co/Phocinae/Phocinae-Largha-150M-v1](https://huggingface.co/Phocinae/Phocinae-Largha-150M-v1)

## Closing

These three models happen to map neatly onto different pain points in agent development: GEV-26B-Decide trades a larger parameter count for higher-precision judgment, Phocinae-Largha-150M-v1 trades an extremely low barrier for a cheap approval gate, and moondream/parakeet-redux lets offline speech transcription stop relying on cloud APIs altogether. If you don't have a GPU on hand and want to try one first, Phocinae-Largha-150M-v1 is probably the lowest-barrier starting point. Next episode, we'll dig up more gems from the local AI scene.

---

## About this article and its author

Originally published on [Mark Ku's Tech Notes](https://blog.markkulab.net/en/local-ai-lab/local-ai-lab-2026-10-09)

License: [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) — when reusing or quoting, credit the author and link back to the original

### About the author

**[Mark Ku](https://blog.markkulab.net/en/author/mark-ku)** — Software engineer

- 10+ years as a software engineer
- Built North-American e-commerce and AI SaaS subscription billing

### Free tools built by the author

All of these are free to use:

- [Free PDF Sign Tool](https://blog.markkulab.net/en/tools/pdf-sign): Online PDF sign tool — draw, type, or upload a signature, then drag, resize, and download. Everything runs in your browser; nothing is uploaded.
- [VS Code Refactory](https://blog.markkulab.net/en/tools/refactory): Refactory is a VS Code refactoring extension: 34 actions plus a 37-rule code-smell inspection layer with a Code Health dashboard, across 18 languages, backed by 534 tests. It learns your repo's conventions: where interfaces live, where DI is registered, whether 'use client' belongs. It ranks files by git churn × complexity so you know what to fix first, and hands any smell to the Claude Code already on your machine. Free to use, and your source never leaves your computer.
- [DB-Kit Database Manager](https://blog.markkulab.net/en/tools/db-kit): DB-Kit is a lightweight, cross-platform database manager built with Tauri + Rust + React. Manage MySQL, MariaDB, PostgreSQL, SQL Server, Oracle, SQLite, MongoDB, Redis, Kafka, Elasticsearch and RabbitMQ from one consistent interface: passwords encrypted in the OS keychain, SSH tunnels, full CRUD, a visual query builder, stacked multi-statement result sets, cross-connection data transfer, schema & data compare with sync SQL and schema snapshots, Excel / CSV import & export, visualized execution plans, ER diagrams, scheduled backups, SQL stress testing with p50–p99 latency percentiles, a 15-rule SQL review engine, Review & Run (AI review plus per-statement backups and an auto-generated rollback script), Kafka message browsing with monitoring & alerts, an SSH terminal with SFTP / FTP, Docker / Kubernetes, remote desktop (RDP / VNC / RustDesk), file / folder compare, stored procedure integration tests, six UI languages, a built-in AI assistant (a local CLI or any Anthropic / OpenAI-compatible API; natural-language SQL, AI review and tuning advice) and the dbk CLI. Free and open source (MIT), with installers for Windows, macOS and Linux.
- [VS Code Super Mermaid](https://blog.markkulab.net/en/tools/super-mermaid): Super Mermaid is a VS Code extension for beautiful Mermaid diagrams out of the box: auto-colored live preview, mouse pan & zoom, high-res PNG / SVG export, 21 templates and multiple themes. Free and open source (MIT).
- [React Super Mermaid](https://blog.markkulab.net/en/tools/react-super-mermaid): react-super-mermaid is an open-source React component library: render beautiful Mermaid diagrams with a single <MermaidViewer>, with built-in colorful / sketch themes, pan & zoom, in-diagram search, and high-res SVG / PNG export. Lightweight, SSR-safe, fully typed. Free and open source (MIT).
- [Jira / Confluence Super Mermaid](https://blog.markkulab.net/en/tools/jira-super-mermaid): An Atlassian Forge app: write Mermaid syntax directly inside a Jira issue or a Confluence page and get flowcharts, sequence diagrams, state machines and Gantt charts. 11 diagram types, SVG / PNG export, light and dark themes, full CJK support. Runs on Atlassian: your diagrams live in your own site and the app calls no third-party service. Free, coming soon to the Atlassian Marketplace.
- [Mermaid Live Preview](https://blog.markkulab.net/en/tools/mermaid-preview): Write Mermaid in your browser, see it render instantly, and share the whole diagram as a single link. No sign-up, nothing uploaded to a server, and mermaid.live share links work as-is.
- [React Intl Phone Number](https://blog.markkulab.net/en/tools/react-intl-phone-number): react-intl-phone-number is an open-source React component: framework-agnostic and antd-free, with E.164 in/out, a searchable flag / country-code dropdown, configurable validation levels (strict / mobile-strict / loose), themeable CSS, and i18n — phone logic powered by google-libphonenumber. Lightweight and fully typed. Free and open source (MIT).
- [Uptime Kuma Cluster](https://blog.markkulab.net/en/tools/uptime-kuma-cluster): Turn single-node Uptime Kuma into a highly available cluster: OpenResty + Lua smart load balancing, shared MariaDB state, health checks and automatic failover, plus cluster-management REST APIs. One Docker Compose command to start. Free and open source (MIT).
- [AI Podcast Cut](https://blog.markkulab.net/en/tools/ai-podcast-cut): Drop in a recording and it removes fillers and stutters, levels loudness segment by segment, and sends a second agent to review every cut. Cut points snap to word boundaries and zero crossings, every splice gets a fade, and sentence-end breaths are preserved. Desktop app for Windows, macOS and Linux. MIT licensed; the Windows installer bundles ffmpeg.
- [AI Video Cut](https://blog.markkulab.net/en/tools/ai-video-cut): Pick something in your video: type it (face, license plate, phone screen, logo), click points or drag a box, or let Claude Code / Codex pick it. AI Video Cut tracks it frame by frame, then mosaics or blurs it, recolors it, pins stickers and text to it, swaps a screen or poster for your own image or video, or removes it using background that other frames actually captured. Plus sequence editing, local captions and an AI assistant. Free, open-source (MIT) Windows desktop app (macOS / Linux experimental) that runs locally, nothing uploaded.
- [open-pos restaurant POS](https://blog.markkulab.net/en/tools/open-pos): One computer and one receipt printer is enough to open the shop. Your data lives on your own disk, no subscription, no lock-in, MIT licensed. Money is integer New Taiwan dollars with tax split by the statutory formula, so sales + tax always equals the total. Printing goes straight over ESC/POS on TCP 9100, with no vendor driver. Tauri + Rust + SQLite desktop app, v1.0 in development.
- [.NET e-commerce API template](https://blog.markkulab.net/en/tools/dotnet-ecommerce-api): An e-commerce backend API template built on .NET 10: Controller → Service → Repository layering, Autofac convention-based registration, SqlSugar for MySQL / MariaDB and SQL Server, Redis cache and message queue, Hangfire jobs, JWT auth, and an extensible payment provider module (ECPay, Newebpay). Apache 2.0.
- [Open Token Monitor team AI usage](https://blog.markkulab.net/en/tools/open-token-monitor): Collect token usage and equivalent cost from every employee's AI coding tools into a hub you host yourself: PostgreSQL, org roster, dashboards compared by company and department, and a reports API. Electron and Rust/Tauri clients upload every 30 minutes. MIT licensed, one Docker command.

### Daily podcasts

- [Mark's Tech Insights — Daily AI News](https://blog.markkulab.net/en/category/tech-news): Daily curated AI and tech trends. Catch the latest developments via audio summaries — covering AI applications, software architecture, DevOps, and engineering practice. — RSS: https://blog.markkulab.net/feed.xml
- [AI股市蝦聊](https://blog.markkulab.net/en/category/ai-stock-chat): Every trading day, an AI-analyzed take on the Taiwan stock market, delivered as a two-host conversation covering the session and the next-day outlook. — RSS: https://blog.markkulab.net/ai-stock-chat/feed.xml
- [開源好物週報](https://blog.markkulab.net/en/category/open-source-weekly): A weekly two-host pick of free open-source tools surfaced from real Hacker News, GitHub, and Reddit buzz — what pain they solve and the fastest way to get started. — RSS: https://blog.markkulab.net/open-source-weekly/feed.xml
- [AI 運動週報](https://blog.markkulab.net/en/category/sports-weekly): Two hosts talk NBA, MLB and world sport three times a week — scores, records and the stories behind them, from a Taiwanese fan perspective. — RSS: https://blog.markkulab.net/sports-weekly/feed.xml
- [AI 國際新聞快報](https://blog.markkulab.net/en/category/world-news): A daily two-host briefing that makes sense of the past 24 hours in world news — geopolitics, the global economy, conflict and security, disasters and climate — from a Taiwanese perspective, neutral and fully sourced. — RSS: https://blog.markkulab.net/world-news/feed.xml
- [地端 AI 實驗室](https://blog.markkulab.net/en/category/local-ai-lab): Twice a week, a two-host look at open-weight models and LoRAs you can actually run on your own machine: what they are for, whether your GPU can handle them, and how they compare — sourced from official model cards. — RSS: https://blog.markkulab.net/local-ai-lab/feed.xml

### Deals

- [Saily eSIM](https://blog.markkulab.net/en/saily): Travel eSIM by the NordVPN team，Promo code：KUKU
- [NordVPN](https://blog.markkulab.net/en/nordvpn): The world's leading VPN, independently audited
- [PremLogin](https://blog.markkulab.net/en/premlogin): Subscription sharing for streaming and AI seats，Promo code：markku666

### Newsletter

[Subscribe to the newsletter](https://blog.markkulab.net/en/subscribe) — Be the first to know about new posts. No spam, unsubscribe anytime.
