Kitten TTS (https://github.com/KittenML/KittenTTS) is an open-source series of tiny and expressive text-to-speech models for on-device applications. We had a thread last year here: https://news.ycombinator.com/item?id=44807868.Today we're releasing three new models with 80M, 40M and 14M parameters.The largest model (80M) has the highest quality. The 14M variant reaches new SOTA in expressivity among similar sized models, despite being <25MB in size. This release is a major upgrade from the previous one and supports English text-to-speech applications in eight voices: four male and four female.Here's a short demo: https://www.youtube.com/watch?v=ge3u5qblqZA.Most models are quantized to int8 + fp16, and they use ONNX for runtime. Our models are designed to run anywhere eg. raspberry pi, low-end smartphones, wearables, browsers etc. No GPU required! This release aims to bridge the gap between on-device and cloud models for tts applications. Multi-lingual model release is coming soon.On-device AI is bottlenecked by one thing: a lack of tiny models that actually perform. Our goal is to open-source more models to run production-ready voice agents and apps entirely on-device.We would love your feedback!
FL score
out of 100
Verdict
high confidence
Competition
9
competitors found, emerging market, funded players
Trend
No signal yet
Build SOTA tiny, expressive, on-device TTS models (<25MB, no GPU) to enable private, low-latency voice applications, monetizing via commercial licenses and enterprise support.
The pain
The gap
Build angle
Strengths
Questions about this idea?
FlyBot reads the scoring and gives you a second opinion on “Three new Kitten TTS models – smallest less than 25MB”.
Risks
Next steps
Fly Labs Method
Is the pain real, is there a gap, is it the right time, can one person build it.
This idea targets a real, severe pain point in on-device AI with a clear gap for tiny, expressive, and easy-to-deploy models, and there are strong signals that developers would pay for a reliable solution. The technical complexity requires significant ML skill but leverages existing open-source work.
Value Equation
Dream outcome and how likely it feels, against the time and effort it costs.
Strong market demand for on-device TTS, excellent value proposition with tiny SOTA models, and a growing market. Differentiation is clear, but building and sustaining SOTA ML requires significant resources.
One-Person Business
Curiosity pull, identity fit, and a path from free value to paid for a solo creator.
A clear problem for a specific developer audience, but sustaining SOTA ML development and managing the open-source project for monetization will be complex for a solo builder.
Viral Frameworks
Hook strength, shareability, and how cheaply it can be tested.
Strong, measurable value prop for a specific developer audience, but needs more structured validation and careful business model execution around the open-source core.
Builder Lens
Evidence the problem exists, timing, defensibility, and a model that fits on a napkin.
Strong demand for solving real pain points with a specific target and narrow wedge, poised for increasing relevance in the future of AI.
Why this verdict
Five lenses, one composite. How scoring works
The angle
This weekend
Who is already there, emerging market
An open-source deep learning toolkit for Text-to-Speech synthesis with voice cloning capabilities and multilingual support.
Pricing: Free (open-source), but commercial use restricted without specific licensing. Cloud inference costs may apply for API usage.
The world's first on-device, super-realistic TTS model with instant voice cloning, built on a compact 0.5B-parameter LLM backbone.
Pricing: Not explicitly stated, but marketed for on-device deployment suggesting potential one-time licensing or enterprise solutions.
A high-performance, open-source TTS model family for low-latency, production-grade voice applications, using a streamlined 350M-parameter architecture.
Pricing: Pricing for Resemble AI's commercial offerings starts at $0.01/second for standard voices and $0.015/second for cloned voices.
A lightweight yet high-quality TTS model with just 82 million parameters, offering efficient running on modest hardware.
Pricing: Free (open-source, Apache 2.0 license).
A Llama-based TTS model available in multiple parameter versions (3B, 1B, 400M, 150M) optimized for natural, human-like speech.
Pricing: Not explicitly stated, likely enterprise-focused.
An ultra-fast, multilingual text-to-speech (TTS) model designed for real-time voice synthesis with sub-100ms latency and ultra-low VRAM usage.
Pricing: Likely API-based with free tier and enterprise options, specific numbers not publicly available for Lightning itself but Smallest.ai offers a TTS platform.
A fast, local neural text-to-speech system that emphasizes small model size for on-device applications.
Pricing: Free (open-source).
A compact open-source speech synthesizer that produces formant-synthesized speech.
Pricing: Free (open-source).
An open-source deep learning-based speech synthesis engine designed for high-quality and customizable speech output.
Pricing: Free (open-source).
What they charge
Recent news
On-Device AI shifts more processing onto devices across IoT systems
IoT News, March 18 2026
KAIST-created SoulMate to live on your device as personal AI
Korea JoongAng Daily, March 17 2026
Exploring open-source tools for integrating text to speech in conversational AI
ElevenLabs, March 14 2026
Kitten TTS: Open-Source Voice Synthesis for Edge Devices
Sesame Disk, March 19 2026
The best text-to-speech software in 2026 - Product Hunt
Product Hunt, March 21 2026
Market signals
The market for on-device and small footprint text-to-speech (TTS) models is a rapidly growing niche within the larger AI industry, driven by the increasing demand for privacy, reduced latency, and lower operational costs in edge AI and IoT applications. Recent signals from chipmakers and smartphone manufacturers indicate a strong trend towards shifting more AI processing onto devices. This includes significant advancements in Neural Processing Units (NPUs) embedded in consumer chips. This shift is also influenced by the high operational costs associated with cloud AI at scale.
What frustrates people
Last summer we faced a conundrum at my company, Tiger Data, a Postgres cloud vendor whose main business is in timeseries data. We were trying to grow our business towards emerging AI-centric workloads and wanted to provide a state-of-the-art hybrid search stack in Postgres. We'd already built pgvectorscale in house with the goal of scaling semantic search beyond pgvector's main memory limitations. We just needed a scalable ranked keyword search solution too.The problem: core Postgres doesn't provide this; the leading Postgres BM25 extension, ParadeDB, is guarded behind AGPL; developing our own extension appeared daunting. We'd need a small team of sharp engineers and 6-12 months, I figured. And we'd probably still fall short of the performance of a mature system like Parade/Tantivy.Or would we? I'd be experimenting long enough with AI-boosted development at that point to realize that with the latest tools (Claude Code + Opus) and an experienced hand (I've been working in database systems internals for 25 years now), the old time estimates pretty much go out the window.I told our CTO I thought I could solo the project in one quarter. This raised some eyebrows.It did take a little more time than that (two quarters), and we got some real help from the community (amazing!) after open-sourcing the pre-release. But I'm thrilled/exhausted today to share that pg_textsearch v1.0 is freely available via open source (Postgres license), on Tiger Data cloud, and hopefully soon, a hyperscalar near you:https://github.com/timescale/pg_textsearchIn the blog post accompanying the release, I overview the architecture and present benchmark results using MS-MARCO. To my surprise, we were not only able to meet Parade/Tantivy's query performance, but exceed it substantially, measuring a 4.7x advantage on query throughput at scale:https://www.tigerdata.com/blog/pg-textsearch-bm25-fu
AI
Hi HN!I recently switched from a Fedora/GNOME laptop to a MacBook Air. My old setup served me well as a portable workstation, but I’ve started traveling more while working remotely and needed something with similar performance but better battery life. The main thing I missed was a simple taskbar that shows the windows in the current workspace instead of a Dock that mixes everything together.I built boringBar so I would not have to use the Dock. It shows only the windows in the current Space, lets you switch Spaces by scrolling on the bar, and adds a desktop switcher so you can jump directly to any Space. You can also hide the system Dock, pin apps, preview windows with thumbnails, and launch apps from a searchable menu (I keep Spotlight disabled because for some reason it uses a lot of system resources on my machine).I’ve been dogfooding it for a few months now, and it finally felt polished enough to share.It’s for people who like macOS but want window management to feel a bit more like GNOME, Windows, or a traditional taskbar. It’s also for people like me who wanted an easier transition to macOS, especially now that Windows feels increasingly user-hostile.I’d love feedback on the UX, bugs, and whether this solves the same Dock/Spaces pain for anyone else.P.S. It might also appeal to people who feel nostalgic for the GNOME 2 desktop of yore. I started my Linux journey with it, and boringBar brings back some of that feeling for me.
AI
### Describe the project you are working on Godot C# bindings ### Describe the problem or limitation you are having in your project For the past weeks, I've been discussing with several Unity users intending to move to Godot C# regarding dealing with the C# garbage collector. The most common complaint I hear from users is that, in Unity, allocations can trigger unexpected GC spikes into the game. In Godot, we target to make all of the high performance APIs (those that intended to be called every frame) not allocate any memory, so theoretically the GC should not be a problem. Additionally, Godot starting from 4.0, uses the Microsoft CoreCLR version of .net, which also supposedly has a better garbage collector than Unity. But in all, after several discussions with Unity users, neither is enough reassurance for them, and they would really feel safer if Godot exposed a zero allocation API. ### Describe the feature / enhancement and how it helps to overcome the problem or limitation The idea of this proposal is that Godot exposes zero allocation versions of many functions in the C# API, that users can use if they desire. Technically, this could be done from the binding generator itself, without breaking compatibility, and without doing any modification to Godot itself. ### Describe how your proposal will work, with code, pseudo-code, mock-ups, and/or diagrams **WARNING** I am not familiar with C#, so take this as pseudocode. Imagine you have two functions exposed as to C#: ```C# void MyClass.SetArray( Vector2[] array); Vector2[] MyClass.GetArray(); ``` This works and is pretty and intuitive. However, it has two problems: * GC is allocated on return * Memory is copied to Godot native formats every time there is a call. The idea is to add NoAlloc versions, which can be generated directly by the binder automatically when required: ```C# void MyClass.SetArrayNoAlloc( Godot.Collections.PackedVector2Array array); void MyCl
AI