Opening
Phonon-2 claims to be the most accurate English speech recognition model under 900 MB, with a download size of just 164 MB. The official listing shows it running at 174x real-time on an M5 MacBook Air. This episode also looks at Cloudflare's own post-trained 9B model clef-flash, as well as Nanosaur2-670M, an image generation model you can train your own LoRA on. We'll tell you which one is easiest to get started with.
This Week's Picks
1. Cloudflare/clef-flash: Answers with probabilities instead of free text, saving you the regex layer for parsing LLM responses
- What it is: A 9B model post-trained by Cloudflare themselves, built on Qwen/Qwen3.5-9B, released under the Apache-2.0 license. Its distinguishing feature isn't generating free text, but outputting probabilities for each option directly. The card lists ticket routing, invoice interpretation, and security incident classification as intended uses.
- Where it can be used:
- Backend engineers can feed incoming support tickets as a set of typed questions, asking which product line they belong to, how severe they are, and who should handle them, getting back a probability for each option. No need to write regex to parse LLM responses or set up retry logic.
- Teams building invoicing and expense systems can use it to read user-uploaded invoice images and determine the next action. The card explicitly lists invoice processing as an intended use, with input support for text, JSON, images, and video.
- People running self-hosted monitoring can hand off security incident classification to it. The card lists security incident classification as a use case, and paired with a quantized version, it can run on Ollama within an internal network, keeping logs off the public internet.
- Can it run on consumer hardware?: The official model card only states the test environment was "a single H200" with PyTorch 2.11 and Transformers 5.10.2, without specifying a consumer-grade VRAM threshold. That said, the card also lists 20 quantized versions compatible with llama.cpp, Ollama, and LM Studio, making it the most likely of this episode's picks to run directly on your own machine.
- This week's buzz: HF trending score 234, 1.3k downloads, 235 likes.
- Comparison with similar models: Its sibling, Cloudflare/clef, is 27B, with a Qwen3.8-27B backbone. The clef-flash card explicitly states it was post-trained from Qwen/Qwen3.5-9B. Official numbers show median latency of 38.8ms versus clef's 209.3ms, but GSM8K score of 67.3% is lower than clef's 80.8%. The official positioning is that it's a trade-off version that sacrifices reasoning ability for speed.
- How to get it: The card's workflow is to first
snapshot_download("Cloudflare/clef-flash"), thensys.path.insert(0, path)andfrom joint_schema_model import load_release_modelto obtain the model and processor (device="cuda"); if you'd rather go through Ollama or LM Studio, use one of the quantized versions listed on the card instead. - Link: huggingface.co/Cloudflare/clef-flash
2. FermionResearch/Phonon-2: A 164 MB speech-to-text model that claims to be the most accurate under 900 MB
- What it is: Built on NVIDIA's parakeet-tdt-0.6b-v3, the official team recompressed it into a 164 MB English speech recognition model, released under CC-BY-4.0. The card self-describes it as "the most accurate open speech recognition model for English under 900 MB."
- Where it can be used:
- Content creators can transcribe English podcasts or meeting recordings on their own laptop. The official listing shows 174x real-time speed on an M5 MacBook Air, eliminating the need to pay for cloud ASR API fees.
- Teams handling sensitive recordings, such as legal, medical, or HR interviews, can keep ASR entirely on an internal network. The card lists support for Apple silicon, Linux x86-64 and Arm, Windows CPU, and a GPU version via Docker, allowing the whole pipeline to stay closed-loop on your own machine.
- Developers running self-hosted subtitle pipelines can turn it into a batch service to process entire video libraries. The official number given is 6,680x real-time on an H100 at batch size 128.
- Can it run on consumer hardware?: The card lists a 164 MB download, with the encoder "holding each weight at one of five learned levels in about 2.1 bits." Platform support includes Apple silicon, Linux (x86-64 and Arm), and Windows CPU, with GPU support via Docker. Specific RAM and VRAM figures are not provided officially.
- This week's buzz: HF trending score 144, 2.1k downloads, 147 likes.
- Comparison with similar models: Its base model is NVIDIA's parakeet-tdt-0.6b-v3. The author claims in the card an average WER of 5.21% across 7 English datasets, reaching 100.8% of the accuracy of a 2.5GB teacher model while being 15x smaller in size. One important caveat: the card's claims are limited to English only; there is no official claim about Chinese recognition accuracy.
- How to get it: According to the README, the steps are
pip install fermion-research, thenpip install mlx mlx-lm mlx-audio soundfile scipy zstandard, thenphonon transcribe recording.wav(orfermion transcribe phonon-2 recording.wav); the card also includes CPU and GPU Docker versions. - Link: huggingface.co/FermionResearch/Phonon-2
3. well9472/Nanosaur2-670M: A 670M image generation model with train_lora.py included directly in the repo
- What it is: A 670M diffusion transformer image generation model, released under the MIT license. The author goes into fine detail on the architecture in the card, covering adaLN-single, 2D RoPE, SwiGLU, QK-norm, SPRINT sparse middle blocks, and x-prediction. The text encoder uses the second-to-last layer of a frozen Gemma-3-270M, and the VAE is a 129M semantic DINOv2 VAE.
- Where it can be used:
- People who want to train a LoRA but don't have access to a big GPU can use it as a practice ground. The training command given in the card,
uv run python custom_nodes/nanosaur2_support/train_lora.py /path/to/images, points directly at your own image folder. - Those already using ComfyUI can plug it in as a lightweight node in their existing workflow. The installation method described in the card is simply copying the custom nodes and model files into the corresponding ComfyUI directories.
- Engineers who want to understand diffusion architecture can use it as study material. The official listing shows the base training took only 11 H100-days, and the architecture breakdown is far more detailed than a typical model card.
- People who want to train a LoRA but don't have access to a big GPU can use it as a practice ground. The training command given in the card,
- Can it run on consumer hardware?: The official card doesn't specify RAM or VRAM thresholds for inference, only training-side numbers: the VAE took 12 hours on 1x H100, and base training took 11 H100-days. Here's an estimate: at 670M parameters, far smaller than image generation models with tens of billions of parameters, it should theoretically be friendlier to consumer GPUs, but this is inferred from parameter count, not an official hardware figure.
- This week's buzz: HF trending score 82, 0 downloads, 90 likes.
- Comparison with similar models: The author preemptively sets expectations in the card: "Do not expect this to compete with fully trained models like Anima in character knowledge or fine detail. The purpose to see what is possible with minimal compute." Compared with this episode's other open-source image generation pick, inclusionAI/Ming-Image-0.1-Design, whose card specifies a verification environment requiring an 80 GiB VRAM CUDA GPU, Nanosaur2's barrier to entry is clearly much lower.
- How to get it: Per the card's instructions, copy the custom nodes and model files to the corresponding ComfyUI directories to use it; to train a LoRA, use the
uv run python custom_nodes/nanosaur2_support/train_lora.py /path/to/imagesscript provided in the card. - Link: huggingface.co/well9472/Nanosaur2-670M
Closing
This episode's three models represent three distinct angles: clef-flash condenses LLM output into probabilities, cutting out the parsing layer entirely; Phonon-2 compresses speech recognition down to 164 MB, small enough to fit on a laptop; and Nanosaur2-670M is small enough to let you train your own LoRA hands-on. If you can only try one for now, clef-flash's 20 quantized versions probably make it the easiest starting point. See you next time.



























Comments