Hi HN, we're Sanchit and Shubham (YC W26). We built a fast inference engine for Apple Silicon. LLMs, speech-to-text, text-to-speech – MetalRT beats llama.cpp, Apple's MLX, Ollama, and sherpa-onnx on every modality we tested. Custom Metal shaders, no framework overhead.Also, we've open-sourced RCLI, the fastest end-to-end voice AI pipeline on Apple Silicon. Mic to spoken response, entirely on-device. No cloud, no API keys.To get started: brew tap RunanywhereAI/rcli https://github.com/RunanywhereAI/RCLI.git brew install rcli rcli setup # downloads ~1 GB of models rcli # interactive mode with push-to-talk Or: curl -fsSL https://raw.githubusercontent.com/RunanywhereAI/RCLI/main/install.sh | bash The numbers (M4 Max, 64 GB, reproducible via `rcli bench`):LLM decode – 1.67x faster than llama.cpp, 1.19x faster than Apple MLX (same model files): - Qwen3-0.6B: 658 tok/s (vs mlx-lm 552, llama.cpp 295) - Qwen3-4B: 186 tok/s (vs mlx-lm 170, llama.cpp 87) - LFM2.5-1.2B: 570 tok/s (vs mlx-lm 509, llama.cpp 372) - Time-to-first-token: 6.6 msSTT – 70 seconds of audio transcribed in *101 ms*. That's 714x real-time. 4.6x faster than mlx-whisper.TTS – 178 ms synthesis. 2.8x faster than mlx-audio and sherpa-onnx.We built this because demoing on-device AI is easy but shipping it is brutal. Voice is the hardest test: you're chaining STT, LLM, and TTS sequentially, and if any stage is slow, the user feels it. Most teams fall back to cloud APIs not because local models are bad, but because local inference infrastructure is.The thing that's hard to solve is latency compounding. In a voice pipeline, you're stacking three models in sequence. If each adds 200ms, you're at 600ms before the user hears a word, and that feels broken. You can't optimize one stage and call it done. Every stage needs to be fast, on one device, with no network round-trip to hide behind
FL score
out of 100
Verdict
high confidence
Competition
8
competitors found, emerging market
Trend
1 community mentions
A technically superior AI inference engine for Apple Silicon that delivers breakthrough performance for on-device, multi-modal AI, with clear market demand but complex monetization for a solo builder.
The pain
The gap
Build angle
Strengths
Questions about this idea?
FlyBot reads the scoring and gives you a second opinion on “RunAnywhere (YC W26) – Faster AI Inference on Apple Silicon”.
Risks
Next steps
Fly Labs Method
Is the pain real, is there a gap, is it the right time, can one person build it.
This idea targets a critical performance gap in on-device AI inference for Apple Silicon, driven by a real need for faster, more private local models. The technical innovation provides a strong solution, and market signals suggest a willingness to pay for superior performance, despite high build complexity.
Value Equation
Dream outcome and how likely it feels, against the time and effort it costs.
Excellent market fit and compelling value proposition for a growing trend, but monetization strategy and build complexity for a solo founder need careful consideration beyond open source.
One-Person Business
Curiosity pull, identity fit, and a path from free value to paid for a solo creator.
A highly technical idea with clear problem clarity and strong leverage, but the solo builder fit, build simplicity, and clear monetization are significant challenges.
Viral Frameworks
Hook strength, shareability, and how cheaply it can be tested.
Strong value proposition and specific target audience, but the open-source nature presents distribution and business model challenges for micro-SaaS, requiring further validation for commercial viability.
Builder Lens
Evidence the problem exists, timing, defensibility, and a model that fits on a napkin.
Strong evidence of a real, specific problem with clear user demand and a promising narrow wedge, with good future fit, but long-term competitive landscape needs careful navigation.
Why this verdict
Five lenses, one composite. How scoring works
The angle
This weekend
Who is already there, emerging market
An open-source platform that enables users to run large language models locally on their computers without requiring cloud connectivity.
Pricing: Free and open source. Offers a 'Pro' tier for cloud models and private model sharing.
A highly optimized inference engine designed to run LLMs efficiently on CPUs and GPUs, even on low-power or edge devices, with a focus on low-level performance and control.
Pricing: Free and open source.
An open-source array framework developed by Apple for machine learning on Apple Silicon, offering efficient, flexible, and highly tuned performance.
Pricing: Free and open source.
A user-friendly desktop application with a GUI for downloading, managing, and chatting with local LLMs, often using llama.cpp under the hood.
Pricing: Not explicitly stated, but generally positioned as a free tool for local use.
A fast and easy-to-use library for LLM inference and serving, designed for high-throughput, multi-user scenarios with optimizations like PagedAttention.
Pricing: Free and open source.
An open-source platform that operates locally and is available for free, intended to serve as a direct alternative to the OpenAI API, enabling on-premises inferencing for various AI functionalities.
Pricing: Free and open source.
An Ollama alternative that balances user-friendliness with a commitment to privacy, designed to run entirely on your local machine, completely offline, and optimized for consumer-grade hardware.
Pricing: Free.
A comprehensive web-based interface for running LLMs locally, built on top of frameworks like PyTorch and Hugging Face Transformers, offering chat, notebook-style prompting, and model management.
Pricing: Free and open source.
What they charge
What people say, 1 mentions
Just Launched: Document Processing SaaS With 99.9% OCR Accuracy + AI That Actually Understands What It's Reading
r/SaaS
Recent news
MetalRT Brings the First Unified AI Inference Engine to Apple Silicon
RunAnywhere Blog, March 18 2026
Tether's QVAC Launches World's First Cross-Platform BitNet LoRA Framework to Enable Billion-Parameter AI Training and Inference on Consumer GPUs and Smartphones
Tether.io, March 17 2026
Local LLMs Apple Silicon Mac 2026 | M1 M2 M3 Guide - SitePoint
SitePoint, March 13 2026
MLX is not faster. I benchmarked MLX vs llama.cpp on M1 Max across four real workloads. Effective tokens/s is quite an issue. What am I missing? Help me with benchmarks and M2 through M5 comparison. : r/LocalLLaMA
Reddit, March 12 2026
Sovereign LLM Inference on Apple Silicon - Medium
Medium, February 09 2026
Market signals
The market for AI inference on Apple Silicon is a growing and significant niche. Recent developments indicate a strong trend towards running powerful AI models locally on consumer devices, driven by desires for faster responses, greater privacy, and independence from cloud services. Apple's unified memory architecture is a key advantage, enabling Macs to run models that exceed typical consumer GPU VRAM limits. Several open-source projects are actively competing in this space, and recent funding rounds like those in the broader local AI / LLM space (e.g., for specialized tools or platforms integrating these inference engines) suggest continued investment and growth.
What frustrates people
Last summer we faced a conundrum at my company, Tiger Data, a Postgres cloud vendor whose main business is in timeseries data. We were trying to grow our business towards emerging AI-centric workloads and wanted to provide a state-of-the-art hybrid search stack in Postgres. We'd already built pgvectorscale in house with the goal of scaling semantic search beyond pgvector's main memory limitations. We just needed a scalable ranked keyword search solution too.The problem: core Postgres doesn't provide this; the leading Postgres BM25 extension, ParadeDB, is guarded behind AGPL; developing our own extension appeared daunting. We'd need a small team of sharp engineers and 6-12 months, I figured. And we'd probably still fall short of the performance of a mature system like Parade/Tantivy.Or would we? I'd be experimenting long enough with AI-boosted development at that point to realize that with the latest tools (Claude Code + Opus) and an experienced hand (I've been working in database systems internals for 25 years now), the old time estimates pretty much go out the window.I told our CTO I thought I could solo the project in one quarter. This raised some eyebrows.It did take a little more time than that (two quarters), and we got some real help from the community (amazing!) after open-sourcing the pre-release. But I'm thrilled/exhausted today to share that pg_textsearch v1.0 is freely available via open source (Postgres license), on Tiger Data cloud, and hopefully soon, a hyperscalar near you:https://github.com/timescale/pg_textsearchIn the blog post accompanying the release, I overview the architecture and present benchmark results using MS-MARCO. To my surprise, we were not only able to meet Parade/Tantivy's query performance, but exceed it substantially, measuring a 4.7x advantage on query throughput at scale:https://www.tigerdata.com/blog/pg-textsearch-bm25-fu
AI
Hi HN!I recently switched from a Fedora/GNOME laptop to a MacBook Air. My old setup served me well as a portable workstation, but I’ve started traveling more while working remotely and needed something with similar performance but better battery life. The main thing I missed was a simple taskbar that shows the windows in the current workspace instead of a Dock that mixes everything together.I built boringBar so I would not have to use the Dock. It shows only the windows in the current Space, lets you switch Spaces by scrolling on the bar, and adds a desktop switcher so you can jump directly to any Space. You can also hide the system Dock, pin apps, preview windows with thumbnails, and launch apps from a searchable menu (I keep Spotlight disabled because for some reason it uses a lot of system resources on my machine).I’ve been dogfooding it for a few months now, and it finally felt polished enough to share.It’s for people who like macOS but want window management to feel a bit more like GNOME, Windows, or a traditional taskbar. It’s also for people like me who wanted an easier transition to macOS, especially now that Windows feels increasingly user-hostile.I’d love feedback on the UX, bugs, and whether this solves the same Dock/Spaces pain for anyone else.P.S. It might also appeal to people who feel nostalgic for the GNOME 2 desktop of yore. I started my Linux journey with it, and boringBar brings back some of that feeling for me.
AI
### Describe the project you are working on Godot C# bindings ### Describe the problem or limitation you are having in your project For the past weeks, I've been discussing with several Unity users intending to move to Godot C# regarding dealing with the C# garbage collector. The most common complaint I hear from users is that, in Unity, allocations can trigger unexpected GC spikes into the game. In Godot, we target to make all of the high performance APIs (those that intended to be called every frame) not allocate any memory, so theoretically the GC should not be a problem. Additionally, Godot starting from 4.0, uses the Microsoft CoreCLR version of .net, which also supposedly has a better garbage collector than Unity. But in all, after several discussions with Unity users, neither is enough reassurance for them, and they would really feel safer if Godot exposed a zero allocation API. ### Describe the feature / enhancement and how it helps to overcome the problem or limitation The idea of this proposal is that Godot exposes zero allocation versions of many functions in the C# API, that users can use if they desire. Technically, this could be done from the binding generator itself, without breaking compatibility, and without doing any modification to Godot itself. ### Describe how your proposal will work, with code, pseudo-code, mock-ups, and/or diagrams **WARNING** I am not familiar with C#, so take this as pseudocode. Imagine you have two functions exposed as to C#: ```C# void MyClass.SetArray( Vector2[] array); Vector2[] MyClass.GetArray(); ``` This works and is pretty and intuitive. However, it has two problems: * GC is allocated on return * Memory is copied to Godot native formats every time there is a call. The idea is to add NoAlloc versions, which can be generated directly by the binder automatically when required: ```C# void MyClass.SetArrayNoAlloc( Godot.Collections.PackedVector2Array array); void MyCl
AI