---
title: "🧪 164 MB, 174x Real-Time: Phonon-2 Makes Speech-to-Text Free From Cloud Taxes | Local AI Lab"
description: "Phonon-2 is officially claimed to be the most accurate English speech recognition model under 900 MB, with a download of only 164 MB, and officially rated at 174x real-time on an M5 MacBook Air. This episode also looks at Cloudflare's own post-trained 9B model clef-flash, and the LoRA-trainable image model Nanosaur2-670M, and we'll tell you which one is easiest to get started with later."
canonical_url: "https://blog.markkulab.net/en/local-ai-lab/local-ai-lab-2026-10-02"
author: "Mark Ku"
author_url: "https://blog.markkulab.net/en/author/mark-ku"
site: "Mark Ku's Tech Notes"
date_published: "2026-10-02T02:00:00.000Z"
category: "Local AI Lab"
tags: ["local-ai", "podcast", "地端模型", "lora", "開源模型", "Cloudflare", "clefflash", "Phonon2", "語音辨識", "Nanosaur2", "LoRA訓練"]
language: "en"
license: "CC BY 4.0"
license_url: "https://creativecommons.org/licenses/by/4.0/"
attribution: "when reusing or quoting, credit the author and link back to the original"
---

# 🧪 164 MB, 174x Real-Time: Phonon-2 Makes Speech-to-Text Free From Cloud Taxes | Local AI Lab

> **TL;DR** — Phonon-2, a 164 MB speech-to-text model, achieves 174x real-time speed on M5 MacBook Air with 5.21% WER across English datasets, eliminating cloud ASR costs.

## Opening

Phonon-2 claims to be the most accurate English speech recognition model under 900 MB, with a download size of just 164 MB. The official listing shows it running at 174x real-time on an M5 MacBook Air. This episode also looks at Cloudflare's own post-trained 9B model clef-flash, as well as Nanosaur2-670M, an image generation model you can train your own LoRA on. We'll tell you which one is easiest to get started with.

## This Week's Picks

### 1. Cloudflare/clef-flash: Answers with probabilities instead of free text, saving you the regex layer for parsing LLM responses
- **What it is**: A 9B model post-trained by Cloudflare themselves, built on Qwen/Qwen3.5-9B, released under the Apache-2.0 license. Its distinguishing feature isn't generating free text, but outputting probabilities for each option directly. The card lists ticket routing, invoice interpretation, and security incident classification as intended uses.
- **Where it can be used**:
  - Backend engineers can feed incoming support tickets as a set of typed questions, asking which product line they belong to, how severe they are, and who should handle them, getting back a probability for each option. No need to write regex to parse LLM responses or set up retry logic.
  - Teams building invoicing and expense systems can use it to read user-uploaded invoice images and determine the next action. The card explicitly lists invoice processing as an intended use, with input support for text, JSON, images, and video.
  - People running self-hosted monitoring can hand off security incident classification to it. The card lists security incident classification as a use case, and paired with a quantized version, it can run on Ollama within an internal network, keeping logs off the public internet.
- **Can it run on consumer hardware?**: The official model card only states the test environment was "a single H200" with PyTorch 2.11 and Transformers 5.10.2, without specifying a consumer-grade VRAM threshold. That said, the card also lists 20 quantized versions compatible with llama.cpp, Ollama, and LM Studio, making it the most likely of this episode's picks to run directly on your own machine.
- **This week's buzz**: HF trending score 234, 1.3k downloads, 235 likes.
- **Comparison with similar models**: Its sibling, Cloudflare/clef, is 27B, with a Qwen3.8-27B backbone. The clef-flash card explicitly states it was post-trained from Qwen/Qwen3.5-9B. Official numbers show median latency of 38.8ms versus clef's 209.3ms, but GSM8K score of 67.3% is lower than clef's 80.8%. The official positioning is that it's a trade-off version that sacrifices reasoning ability for speed.
- **How to get it**: The card's workflow is to first `snapshot_download("Cloudflare/clef-flash")`, then `sys.path.insert(0, path)` and `from joint_schema_model import load_release_model` to obtain the model and processor (`device="cuda"`); if you'd rather go through Ollama or LM Studio, use one of the quantized versions listed on the card instead.
- **Link**: [huggingface.co/Cloudflare/clef-flash](https://huggingface.co/Cloudflare/clef-flash)

### 2. FermionResearch/Phonon-2: A 164 MB speech-to-text model that claims to be the most accurate under 900 MB
- **What it is**: Built on NVIDIA's parakeet-tdt-0.6b-v3, the official team recompressed it into a 164 MB English speech recognition model, released under CC-BY-4.0. The card self-describes it as "the most accurate open speech recognition model for English under 900 MB."
- **Where it can be used**:
  - Content creators can transcribe English podcasts or meeting recordings on their own laptop. The official listing shows 174x real-time speed on an M5 MacBook Air, eliminating the need to pay for cloud ASR API fees.
  - Teams handling sensitive recordings, such as legal, medical, or HR interviews, can keep ASR entirely on an internal network. The card lists support for Apple silicon, Linux x86-64 and Arm, Windows CPU, and a GPU version via Docker, allowing the whole pipeline to stay closed-loop on your own machine.
  - Developers running self-hosted subtitle pipelines can turn it into a batch service to process entire video libraries. The official number given is 6,680x real-time on an H100 at batch size 128.
- **Can it run on consumer hardware?**: The card lists a 164 MB download, with the encoder "holding each weight at one of five learned levels in about 2.1 bits." Platform support includes Apple silicon, Linux (x86-64 and Arm), and Windows CPU, with GPU support via Docker. Specific RAM and VRAM figures are not provided officially.
- **This week's buzz**: HF trending score 144, 2.1k downloads, 147 likes.
- **Comparison with similar models**: Its base model is NVIDIA's parakeet-tdt-0.6b-v3. The author claims in the card an average WER of 5.21% across 7 English datasets, reaching 100.8% of the accuracy of a 2.5GB teacher model while being 15x smaller in size. One important caveat: the card's claims are limited to English only; there is no official claim about Chinese recognition accuracy.
- **How to get it**: According to the README, the steps are `pip install fermion-research`, then `pip install mlx mlx-lm mlx-audio soundfile scipy zstandard`, then `phonon transcribe recording.wav` (or `fermion transcribe phonon-2 recording.wav`); the card also includes CPU and GPU Docker versions.
- **Link**: [huggingface.co/FermionResearch/Phonon-2](https://huggingface.co/FermionResearch/Phonon-2)

### 3. well9472/Nanosaur2-670M: A 670M image generation model with train_lora.py included directly in the repo
- **What it is**: A 670M diffusion transformer image generation model, released under the MIT license. The author goes into fine detail on the architecture in the card, covering adaLN-single, 2D RoPE, SwiGLU, QK-norm, SPRINT sparse middle blocks, and x-prediction. The text encoder uses the second-to-last layer of a frozen Gemma-3-270M, and the VAE is a 129M semantic DINOv2 VAE.
- **Where it can be used**:
  - People who want to train a LoRA but don't have access to a big GPU can use it as a practice ground. The training command given in the card, `uv run python custom_nodes/nanosaur2_support/train_lora.py /path/to/images`, points directly at your own image folder.
  - Those already using ComfyUI can plug it in as a lightweight node in their existing workflow. The installation method described in the card is simply copying the custom nodes and model files into the corresponding ComfyUI directories.
  - Engineers who want to understand diffusion architecture can use it as study material. The official listing shows the base training took only 11 H100-days, and the architecture breakdown is far more detailed than a typical model card.
- **Can it run on consumer hardware?**: The official card doesn't specify RAM or VRAM thresholds for inference, only training-side numbers: the VAE took 12 hours on 1x H100, and base training took 11 H100-days. Here's an estimate: at 670M parameters, far smaller than image generation models with tens of billions of parameters, it should theoretically be friendlier to consumer GPUs, but this is inferred from parameter count, not an official hardware figure.
- **This week's buzz**: HF trending score 82, 0 downloads, 90 likes.
- **Comparison with similar models**: The author preemptively sets expectations in the card: "Do not expect this to compete with fully trained models like Anima in character knowledge or fine detail. The purpose to see what is possible with minimal compute." Compared with this episode's other open-source image generation pick, inclusionAI/Ming-Image-0.1-Design, whose card specifies a verification environment requiring an 80 GiB VRAM CUDA GPU, Nanosaur2's barrier to entry is clearly much lower.
- **How to get it**: Per the card's instructions, copy the custom nodes and model files to the corresponding ComfyUI directories to use it; to train a LoRA, use the `uv run python custom_nodes/nanosaur2_support/train_lora.py /path/to/images` script provided in the card.
- **Link**: [huggingface.co/well9472/Nanosaur2-670M](https://huggingface.co/well9472/Nanosaur2-670M)

## Closing

This episode's three models represent three distinct angles: clef-flash condenses LLM output into probabilities, cutting out the parsing layer entirely; Phonon-2 compresses speech recognition down to 164 MB, small enough to fit on a laptop; and Nanosaur2-670M is small enough to let you train your own LoRA hands-on. If you can only try one for now, clef-flash's 20 quantized versions probably make it the easiest starting point. See you next time.

---

## About this article and its author

Originally published on [Mark Ku's Tech Notes](https://blog.markkulab.net/en/local-ai-lab/local-ai-lab-2026-10-02)

License: [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) — when reusing or quoting, credit the author and link back to the original

### About the author

**[Mark Ku](https://blog.markkulab.net/en/author/mark-ku)** — Software engineer

- 10+ years as a software engineer
- Built North-American e-commerce and AI SaaS subscription billing

### Free tools built by the author

All of these are free to use:

- [Free PDF Sign Tool](https://blog.markkulab.net/en/tools/pdf-sign): Online PDF sign tool — draw, type, or upload a signature, then drag, resize, and download. Everything runs in your browser; nothing is uploaded.
- [VS Code Refactory](https://blog.markkulab.net/en/tools/refactory): Refactory is a VS Code refactoring extension: 34 actions plus a 37-rule code-smell inspection layer with a Code Health dashboard, across 18 languages, backed by 534 tests. It learns your repo's conventions: where interfaces live, where DI is registered, whether 'use client' belongs. It ranks files by git churn × complexity so you know what to fix first, and hands any smell to the Claude Code already on your machine. Free to use, and your source never leaves your computer.
- [DB-Kit Database Manager](https://blog.markkulab.net/en/tools/db-kit): DB-Kit is a lightweight, cross-platform database manager built with Tauri + Rust + React. Manage MySQL, MariaDB, PostgreSQL, SQL Server, Oracle, SQLite, MongoDB, Redis, Kafka, Elasticsearch and RabbitMQ from one consistent interface: passwords encrypted in the OS keychain, SSH tunnels, full CRUD, a visual query builder, stacked multi-statement result sets, cross-connection data transfer, schema & data compare with sync SQL and schema snapshots, Excel / CSV import & export, visualized execution plans, ER diagrams, scheduled backups, SQL stress testing with p50–p99 latency percentiles, a 15-rule SQL review engine, Review & Run (AI review plus per-statement backups and an auto-generated rollback script), Kafka message browsing with monitoring & alerts, a bilingual UI (Traditional Chinese / English), a built-in AI assistant (a local CLI or any Anthropic / OpenAI-compatible API; natural-language SQL, AI review and tuning advice) and the dbk CLI. Free and open source (MIT), with installers for Windows, macOS and Linux.
- [VS Code Super Mermaid](https://blog.markkulab.net/en/tools/super-mermaid): Super Mermaid is a VS Code extension for beautiful Mermaid diagrams out of the box: auto-colored live preview, mouse pan & zoom, high-res PNG / SVG export, 21 templates and multiple themes. Free and open source (MIT).
- [React Super Mermaid](https://blog.markkulab.net/en/tools/react-super-mermaid): react-super-mermaid is an open-source React component library: render beautiful Mermaid diagrams with a single <MermaidViewer>, with built-in colorful / sketch themes, pan & zoom, in-diagram search, and high-res SVG / PNG export. Lightweight, SSR-safe, fully typed. Free and open source (MIT).
- [Jira / Confluence Super Mermaid](https://blog.markkulab.net/en/tools/jira-super-mermaid): An Atlassian Forge app: write Mermaid syntax directly inside a Jira issue or a Confluence page and get flowcharts, sequence diagrams, state machines and Gantt charts. 11 diagram types, SVG / PNG export, light and dark themes, full CJK support. Runs on Atlassian: your diagrams live in your own site and the app calls no third-party service. Free, coming soon to the Atlassian Marketplace.
- [Mermaid Live Preview](https://blog.markkulab.net/en/tools/mermaid-preview): Write Mermaid in your browser, see it render instantly, and share the whole diagram as a single link. No sign-up, nothing uploaded to a server, and mermaid.live share links work as-is.
- [React Intl Phone Number](https://blog.markkulab.net/en/tools/react-intl-phone-number): react-intl-phone-number is an open-source React component: framework-agnostic and antd-free, with E.164 in/out, a searchable flag / country-code dropdown, configurable validation levels (strict / mobile-strict / loose), themeable CSS, and i18n — phone logic powered by google-libphonenumber. Lightweight and fully typed. Free and open source (MIT).
- [Uptime Kuma Cluster](https://blog.markkulab.net/en/tools/uptime-kuma-cluster): Turn single-node Uptime Kuma into a highly available cluster: OpenResty + Lua smart load balancing, shared MariaDB state, health checks and automatic failover, plus cluster-management REST APIs. One Docker Compose command to start. Free and open source (MIT).
- [AI Podcast Cut](https://blog.markkulab.net/en/tools/ai-podcast-cut): Drop in a recording and it removes fillers and stutters, levels loudness segment by segment, and sends a second agent to review every cut. Cut points snap to word boundaries and zero crossings, every splice gets a fade, and sentence-end breaths are preserved. Desktop app for Windows, macOS and Linux. MIT licensed; the Windows installer bundles ffmpeg.
- [open-pos restaurant POS](https://blog.markkulab.net/en/tools/open-pos): One computer and one receipt printer is enough to open the shop. Your data lives on your own disk, no subscription, no lock-in, MIT licensed. Money is integer New Taiwan dollars with tax split by the statutory formula, so sales + tax always equals the total. Printing goes straight over ESC/POS on TCP 9100, with no vendor driver. Tauri + Rust + SQLite desktop app, v1.0 in development.

### Daily podcasts

- [Mark's Tech Insights — Daily AI News](https://blog.markkulab.net/en/category/tech-news): Daily curated AI and tech trends. Catch the latest developments via audio summaries — covering AI applications, software architecture, DevOps, and engineering practice. — RSS: https://blog.markkulab.net/feed.xml
- [AI股市蝦聊](https://blog.markkulab.net/en/category/ai-stock-chat): Every trading day, an AI-analyzed take on the Taiwan stock market, delivered as a two-host conversation covering the session and the next-day outlook. — RSS: https://blog.markkulab.net/ai-stock-chat/feed.xml
- [開源好物週報](https://blog.markkulab.net/en/category/open-source-weekly): A weekly two-host pick of free open-source tools surfaced from real Hacker News, GitHub, and Reddit buzz — what pain they solve and the fastest way to get started. — RSS: https://blog.markkulab.net/open-source-weekly/feed.xml
- [AI 運動週報](https://blog.markkulab.net/en/category/sports-weekly): Two hosts talk NBA, MLB and world sport three times a week — scores, records and the stories behind them, from a Taiwanese fan perspective. — RSS: https://blog.markkulab.net/sports-weekly/feed.xml
- [AI 國際新聞快報](https://blog.markkulab.net/en/category/world-news): A daily two-host briefing that makes sense of the past 24 hours in world news — geopolitics, the global economy, conflict and security, disasters and climate — from a Taiwanese perspective, neutral and fully sourced. — RSS: https://blog.markkulab.net/world-news/feed.xml
- [地端 AI 實驗室](https://blog.markkulab.net/en/category/local-ai-lab): Twice a week, a two-host look at open-weight models and LoRAs you can actually run on your own machine: what they are for, whether your GPU can handle them, and how they compare — sourced from official model cards. — RSS: https://blog.markkulab.net/local-ai-lab/feed.xml

### Deals

- [Saily eSIM](https://blog.markkulab.net/en/saily): Travel eSIM by the NordVPN team，Promo code：KUKU
- [NordVPN](https://blog.markkulab.net/en/nordvpn): The world's leading VPN, independently audited
- [PremLogin](https://blog.markkulab.net/en/premlogin): Subscription sharing for streaming and AI seats，Promo code：markku666

### Newsletter

[Subscribe to the newsletter](https://blog.markkulab.net/en/subscribe) — Be the first to know about new posts. No spam, unsubscribe anytime.
