About six months ago, I started working on a project to fine-tune Whisper locally on my M2 Ultra Mac Studio with a limited compute budget. I got into it. The problem I had at the time was I had 15,000 hours of audio data in Google Cloud Storage, and there was no way I could fit all the audio onto my local machine, so I built a system to stream data from my GCS to my machine during training.Gemma 3n came out, so I added that. Kinda went nuts, tbh.Then I put it on the shelf.When Gemma 4 came out a few days ago, I dusted it off, cleaned it up, broke out the Gemma part from the Whisper fine-tuning and added support for Gemma 4.I'm presenting it for you here today to play with, fork and improve upon.One thing I have learned so far: It's very easy to OOM when you fine-tune on longer sequences! My local Mac Studio has 64GB RAM, so I run out of memory constantly.Anywho, given how much interest there is in Gemma 4, and frankly, the fact that you can't really do audio fine-tuning with MLX, that's really the reason this exists (in addition to my personal interest). I would have preferred to use MLX and not have had to make this, but here we are. Welcome to my little side quest.And so I made this. I hope you have as much fun using it as I had fun making it.-Matt
FL score
out of 100
Verdict
high confidence
Competition
11
competitors found, emerging market, funded players
Trend
No signal yet
A local Gemma 4 multimodal audio fine-tuner for Apple Silicon, solving OOM issues and MLX limitations, built by a solo developer.
The pain
The gap
Build angle
Strengths
Questions about this idea?
FlyBot reads the scoring and gives you a second opinion on “Gemma 4 Multimodal Fine-Tuner for Apple Silicon”.
Risks
Next steps
Fly Labs Method
Is the pain real, is there a gap, is it the right time, can one person build it.
A strong technical solution to a specific pain point (Gemma 4 audio fine-tuning, OOM) on Apple Silicon, built by the developer. Monetization path for a solo builder is currently unclear but there is a clear market need for efficient local fine-tuning.
Value Equation
Dream outcome and how likely it feels, against the time and effort it costs.
This idea targets a high-pain, growing market with a clear value proposition, though monetizing it as a solo open-source project will be a key challenge.
One-Person Business
Curiosity pull, identity fit, and a path from free value to paid for a solo creator.
An excellent creator-problem fit with a clear, niche problem, but potential challenges in broad audience reach and defining a clear monetization path for a solo builder.
Viral Frameworks
Hook strength, shareability, and how cheaply it can be tested.
Strong target audience and clear value proposition for a validated technical problem. The main hurdle is transitioning from an open-source project to a micro-SaaS with a defined business model and distribution strategy.
Builder Lens
Evidence the problem exists, timing, defensibility, and a model that fits on a napkin.
Addresses a deeply felt pain for a specific, technical audience with a clear improvement over the status quo. The product exists as a narrow wedge, and future-fit is strong.
Why this verdict
Five lenses, one composite. How scoring works
The angle
This weekend
Who is already there, emerging market
An open-source library leveraging Apple's MLX framework for efficient local LLM fine-tuning on Apple Silicon, reducing reliance on cloud GPUs.
Pricing: Free (open-source)
A desktop application with a GUI for local LLM fine-tuning on Apple Silicon, built with Tauri + React and leveraging MLX.
Pricing: Free (open-source)
An open-source project by IBM and Red Hat that provides an easy way to tune and run models, including local fine-tuning on Apple Silicon.
Pricing: Free (open-source)
An MLX-powered library that brings the Unsloth fine-tuning experience to Apple Silicon, aiming for code portability between Mac and cloud GPUs.
Pricing: Free (open-source)
A framework that optimizes LLM fine-tuning, making it faster and reducing VRAM usage, with support for Gemma models.
Pricing: Offers free Colab notebooks for fine-tuning; likely commercial offerings for optimized use but specific pricing not found for self-hosted tooling.
A tool for running large language models locally, including Gemma 4, with an OpenAI-compatible API.
Pricing: Free (open-source)
A macOS speech-to-text app offering local-only models for privacy and fast input across applications.
Pricing: $249 lifetime plan or $84.99/year for Pro features; free version with limited access to small models.
A Mac dictation app optimized for Apple Silicon, offering offline privacy and using OpenAI's Whisper models for fast, accurate transcription.
Pricing: $99 lifetime plan or $44.10/year.
An open-source Mac-only voice-to-text app focused on speed, accuracy, and privacy, running fully offline with local AI models.
Pricing: Free by building from source or $39 one-time for pre-built convenience.
A voice-to-text tool built for fast writing on Mac, Windows, and iPhone, with cloud-based processing for consistent performance.
Pricing: Specific pricing not found, but noted as offering a 'premium experience'.
An open-source toolkit to fine-tune Gemma 4 and 3n with audio, images, and text on Apple Silicon, including streaming from cloud storage.
Pricing: Free (open-source)
What they charge
Recent news
Hacker News, April 07 2026
Agent Wars, April 06 2026
github.com/mattmireles, April 07 2026
Medium (by RK), April 06 2026
FlowHunt, April 04 2026
Market signals
The market for local LLM fine-tuning on Apple Silicon is a rapidly growing niche, driven by the desire for cost savings, privacy, and on-device AI capabilities. Recent developments, particularly with Gemma 4's release and the MLX framework, indicate a significant shift towards making advanced multimodal fine-tuning accessible on consumer hardware. The increasing number of open-source tools and community activity highlights strong interest and potential for expansion.
What frustrates people
Last summer we faced a conundrum at my company, Tiger Data, a Postgres cloud vendor whose main business is in timeseries data. We were trying to grow our business towards emerging AI-centric workloads and wanted to provide a state-of-the-art hybrid search stack in Postgres. We'd already built pgvectorscale in house with the goal of scaling semantic search beyond pgvector's main memory limitations. We just needed a scalable ranked keyword search solution too.The problem: core Postgres doesn't provide this; the leading Postgres BM25 extension, ParadeDB, is guarded behind AGPL; developing our own extension appeared daunting. We'd need a small team of sharp engineers and 6-12 months, I figured. And we'd probably still fall short of the performance of a mature system like Parade/Tantivy.Or would we? I'd be experimenting long enough with AI-boosted development at that point to realize that with the latest tools (Claude Code + Opus) and an experienced hand (I've been working in database systems internals for 25 years now), the old time estimates pretty much go out the window.I told our CTO I thought I could solo the project in one quarter. This raised some eyebrows.It did take a little more time than that (two quarters), and we got some real help from the community (amazing!) after open-sourcing the pre-release. But I'm thrilled/exhausted today to share that pg_textsearch v1.0 is freely available via open source (Postgres license), on Tiger Data cloud, and hopefully soon, a hyperscalar near you:https://github.com/timescale/pg_textsearchIn the blog post accompanying the release, I overview the architecture and present benchmark results using MS-MARCO. To my surprise, we were not only able to meet Parade/Tantivy's query performance, but exceed it substantially, measuring a 4.7x advantage on query throughput at scale:https://www.tigerdata.com/blog/pg-textsearch-bm25-fu
AI
Hi HN!I recently switched from a Fedora/GNOME laptop to a MacBook Air. My old setup served me well as a portable workstation, but I’ve started traveling more while working remotely and needed something with similar performance but better battery life. The main thing I missed was a simple taskbar that shows the windows in the current workspace instead of a Dock that mixes everything together.I built boringBar so I would not have to use the Dock. It shows only the windows in the current Space, lets you switch Spaces by scrolling on the bar, and adds a desktop switcher so you can jump directly to any Space. You can also hide the system Dock, pin apps, preview windows with thumbnails, and launch apps from a searchable menu (I keep Spotlight disabled because for some reason it uses a lot of system resources on my machine).I’ve been dogfooding it for a few months now, and it finally felt polished enough to share.It’s for people who like macOS but want window management to feel a bit more like GNOME, Windows, or a traditional taskbar. It’s also for people like me who wanted an easier transition to macOS, especially now that Windows feels increasingly user-hostile.I’d love feedback on the UX, bugs, and whether this solves the same Dock/Spaces pain for anyone else.P.S. It might also appeal to people who feel nostalgic for the GNOME 2 desktop of yore. I started my Linux journey with it, and boringBar brings back some of that feeling for me.
AI
### Describe the project you are working on Godot C# bindings ### Describe the problem or limitation you are having in your project For the past weeks, I've been discussing with several Unity users intending to move to Godot C# regarding dealing with the C# garbage collector. The most common complaint I hear from users is that, in Unity, allocations can trigger unexpected GC spikes into the game. In Godot, we target to make all of the high performance APIs (those that intended to be called every frame) not allocate any memory, so theoretically the GC should not be a problem. Additionally, Godot starting from 4.0, uses the Microsoft CoreCLR version of .net, which also supposedly has a better garbage collector than Unity. But in all, after several discussions with Unity users, neither is enough reassurance for them, and they would really feel safer if Godot exposed a zero allocation API. ### Describe the feature / enhancement and how it helps to overcome the problem or limitation The idea of this proposal is that Godot exposes zero allocation versions of many functions in the C# API, that users can use if they desire. Technically, this could be done from the binding generator itself, without breaking compatibility, and without doing any modification to Godot itself. ### Describe how your proposal will work, with code, pseudo-code, mock-ups, and/or diagrams **WARNING** I am not familiar with C#, so take this as pseudocode. Imagine you have two functions exposed as to C#: ```C# void MyClass.SetArray( Vector2[] array); Vector2[] MyClass.GetArray(); ``` This works and is pretty and intuitive. However, it has two problems: * GC is allocated on return * Memory is copied to Godot native formats every time there is a call. The idea is to add NoAlloc versions, which can be generated directly by the binder automatically when required: ```C# void MyClass.SetArrayNoAlloc( Godot.Collections.PackedVector2Array array); void MyCl
AI