Hey HN — I’m Veer and my cofounder is Suryaa. We're building Cumulus Labs (YC W26), and we're releasing our latest product IonRouter (https://ionrouter.io/), an inference API for open-source and fine tuned models. You swap in our base URL, keep your existing OpenAI client code, and get access to any model (open source or finetuned to you) running on our own inference engine.The problem we kept running into: every inference provider is either fast-but-expensive (Together, Fireworks — you pay for always-on GPUs) or cheap-but-DIY (Modal, RunPod — you configure vLLM yourself and deal with slow cold starts). Neither felt right for teams that just want to ship.Suryaa spent years building GPU orchestration infrastructure at TensorDock and production systems at Palantir. I led ML infrastructure and Linux kernel development for Space Force and NASA contracts where the stack had to actually work under pressure. When we started building AI products ourselves, we kept hitting the same wall: GPU infrastructure was either too expensive or too much work.So we built IonAttention — a C++ inference runtime designed specifically around the GH200's memory architecture. Most inference stacks treat GH200 as a compatibility target (make sure vLLM runs, use CPU memory as overflow). We took a different approach and built around what makes the hardware actually interesting: a 900 GB/s coherent CPU-GPU link, 452GB of LPDDR5X sitting right next to the accelerator, and 72 ARM cores you can actually use.Three things came out of that that we think are novel: (1) using hardware cache coherence to make CUDA graphs behave as if they have dynamic parameters at zero per-step cost — something that only works on GH200-class hardware; (2) eager KV block writeback driven by immutability rather than memory pressure, which drops eviction stalls from 10ms+ to under 0.25ms; (3) phantom-tile attention scheduling at small batch sizes that cuts attention time by over 60% in the
FL score
out of 100
Verdict
high confidence
Competition
19
competitors found, growing market, funded players
Trend
No signal yet
A deeply technical, hardware-optimized AI inference platform offering high-throughput and low-cost for open-source models, but entirely unsuitable for a solo builder.
The pain
The gap
Build angle
Strengths
Questions about this idea?
FlyBot reads the scoring and gives you a second opinion on “IonRouter (YC W26) – High-throughput, low-cost inference”.
Risks
Next steps
Fly Labs Method
Is the pain real, is there a gap, is it the right time, can one person build it.
The problem is real and painful, and there's a clear willingness to pay for a solution that bridges the cost-performance gap. However, this is a highly technical, hardware-level optimization project explicitly involving two co-founders with specialized backgrounds, making it completely unsuitable for a solo builder's weekend MVP.
Value Equation
Dream outcome and how likely it feels, against the time and effort it costs.
This idea targets a high-pain, growing market with a strong value proposition, but has very high build complexity and faces strong competition.
One-Person Business
Curiosity pull, identity fit, and a path from free value to paid for a solo creator.
High problem clarity and potential leverage, but extremely low creator fit for a solo builder due to immense complexity and a narrow technical niche.
Viral Frameworks
Hook strength, shareability, and how cheaply it can be tested.
Clear value proposition for a specific audience, but high execution risk in a competitive market and a challenging business model.
Builder Lens
Evidence the problem exists, timing, defensibility, and a model that fits on a napkin.
Strong demand and a painful status quo, but execution complexity and nascent observation data are challenges for this specific hardware-optimized wedge.
Why this verdict
Five lenses, one composite. How scoring works
The angle
This weekend
Who is already there, growing market
Fastest, reliable inference cloud for open models with high throughput and low latency.
Pricing: pay-per-token, competitive (e.g. low cost for frontier models)
World's fastest AI inference with wafer-scale chips, high TPS.
Pricing: 10c/M tokens Llama 8B, 60c/M 70B; pay-as-you-go
Low-cost hosting for various models including cheap TTS.
Pricing: unknown
Platform for inference, fine-tuning on open models.
Pricing: pay-per-token
Ultra-fast inference with custom chips.
Pricing: very low per token, fast
Serverless GPU inference platform.
Pricing: unknown
Large scale OpenAI-compatible inference, 85% cheaper, market pricing.
Pricing: market-based token pricing
Offers serverless GPU platforms that bill per-second of actual use, making it cost-effective for bursty AI agent workloads. It provides a flexible developer experience for custom Python-based agent pipelines.
Pricing: An A100 workload costs around $246 monthly on specialized providers like Modal at $1.10/hour, assuming 2 hours daily usage. For AI agents with variable load, serverless platforms like Modal can be 5x cheaper than SageMaker when GPU utilization is under 30%. IonRouter mentions Modal as 'cheap-but-DIY'.
Excels at cold start performance and GPU variety, best for batch processing and cost-sensitive production workloads. It provides on-demand GPU compute with a focus on containerized workflows and supports a wide range of consumer and data center GPUs.
Pricing: Offers competitive pricing, with RTX 4090 (24GB) at ~$0.35/hr and A100 SXM (80GB) at ~$0.35/hr. Reported RTX 3090 pricing is around $0.20/hr. RunPod achieves 48% of cold starts under 200ms. IonRouter mentions RunPod as 'cheap-but-DIY'.
Builds infrastructure for running AI models in production environments, focusing on inference. Their platform removes much of the complexity associated with operationalizing machine learning models.
Pricing: unknown
Founded by the creators of vLLM, focusing on optimizing how models use memory and compute for inference, particularly with its PagedAttention feature.
Pricing: unknown
Offers an LLM Inference Engine emphasizing high throughput and cost-effectiveness with serverless integration.
Pricing: Processes 130 tokens per second with Llama-2-70B-Chat and 180 tokens per second with Llama-2-13B-Chat at $0.20 per million tokens.
Gaps they leave open
Recent news
IonRouter (YC W26) Launches High-Throughput LLM Inference Platform with Proprietary IonAttention Engine
Agent Wars, March 15 2026
F5 and NVIDIA advance AI factory economics with new capabilities for accelerated AI inference
Business Wire, March 17 2026
NVIDIA Enters Production With Dynamo, the Broadly Adopted Inference Operating System for AI Factories
NVIDIA, March 17 2026
Nvidia targets inference as AI's next battleground with Groq 3 LPX
Network World, March 18 2026
NVIDIA, telecom leaders build AI grids to optimize inference on distributed networks
NVIDIA, March 18 2026
Market signals
The AI inference market is rapidly evolving, with a strong focus on high-throughput, low-latency, and cost-effective solutions, particularly for LLMs and agentic AI. Several well-funded startups and established tech giants are developing specialized hardware and software to address these demands, alongside a growing ecosystem of serverless GPU providers and platforms offering flexible pricing and deployment options.
Last summer we faced a conundrum at my company, Tiger Data, a Postgres cloud vendor whose main business is in timeseries data. We were trying to grow our business towards emerging AI-centric workloads and wanted to provide a state-of-the-art hybrid search stack in Postgres. We'd already built pgvectorscale in house with the goal of scaling semantic search beyond pgvector's main memory limitations. We just needed a scalable ranked keyword search solution too.The problem: core Postgres doesn't provide this; the leading Postgres BM25 extension, ParadeDB, is guarded behind AGPL; developing our own extension appeared daunting. We'd need a small team of sharp engineers and 6-12 months, I figured. And we'd probably still fall short of the performance of a mature system like Parade/Tantivy.Or would we? I'd be experimenting long enough with AI-boosted development at that point to realize that with the latest tools (Claude Code + Opus) and an experienced hand (I've been working in database systems internals for 25 years now), the old time estimates pretty much go out the window.I told our CTO I thought I could solo the project in one quarter. This raised some eyebrows.It did take a little more time than that (two quarters), and we got some real help from the community (amazing!) after open-sourcing the pre-release. But I'm thrilled/exhausted today to share that pg_textsearch v1.0 is freely available via open source (Postgres license), on Tiger Data cloud, and hopefully soon, a hyperscalar near you:https://github.com/timescale/pg_textsearchIn the blog post accompanying the release, I overview the architecture and present benchmark results using MS-MARCO. To my surprise, we were not only able to meet Parade/Tantivy's query performance, but exceed it substantially, measuring a 4.7x advantage on query throughput at scale:https://www.tigerdata.com/blog/pg-textsearch-bm25-fu
AI
Hi HN!I recently switched from a Fedora/GNOME laptop to a MacBook Air. My old setup served me well as a portable workstation, but I’ve started traveling more while working remotely and needed something with similar performance but better battery life. The main thing I missed was a simple taskbar that shows the windows in the current workspace instead of a Dock that mixes everything together.I built boringBar so I would not have to use the Dock. It shows only the windows in the current Space, lets you switch Spaces by scrolling on the bar, and adds a desktop switcher so you can jump directly to any Space. You can also hide the system Dock, pin apps, preview windows with thumbnails, and launch apps from a searchable menu (I keep Spotlight disabled because for some reason it uses a lot of system resources on my machine).I’ve been dogfooding it for a few months now, and it finally felt polished enough to share.It’s for people who like macOS but want window management to feel a bit more like GNOME, Windows, or a traditional taskbar. It’s also for people like me who wanted an easier transition to macOS, especially now that Windows feels increasingly user-hostile.I’d love feedback on the UX, bugs, and whether this solves the same Dock/Spaces pain for anyone else.P.S. It might also appeal to people who feel nostalgic for the GNOME 2 desktop of yore. I started my Linux journey with it, and boringBar brings back some of that feeling for me.
AI
### Describe the project you are working on Godot C# bindings ### Describe the problem or limitation you are having in your project For the past weeks, I've been discussing with several Unity users intending to move to Godot C# regarding dealing with the C# garbage collector. The most common complaint I hear from users is that, in Unity, allocations can trigger unexpected GC spikes into the game. In Godot, we target to make all of the high performance APIs (those that intended to be called every frame) not allocate any memory, so theoretically the GC should not be a problem. Additionally, Godot starting from 4.0, uses the Microsoft CoreCLR version of .net, which also supposedly has a better garbage collector than Unity. But in all, after several discussions with Unity users, neither is enough reassurance for them, and they would really feel safer if Godot exposed a zero allocation API. ### Describe the feature / enhancement and how it helps to overcome the problem or limitation The idea of this proposal is that Godot exposes zero allocation versions of many functions in the C# API, that users can use if they desire. Technically, this could be done from the binding generator itself, without breaking compatibility, and without doing any modification to Godot itself. ### Describe how your proposal will work, with code, pseudo-code, mock-ups, and/or diagrams **WARNING** I am not familiar with C#, so take this as pseudocode. Imagine you have two functions exposed as to C#: ```C# void MyClass.SetArray( Vector2[] array); Vector2[] MyClass.GetArray(); ``` This works and is pretty and intuitive. However, it has two problems: * GC is allocated on return * Memory is copied to Godot native formats every time there is a call. The idea is to add NoAlloc versions, which can be generated directly by the binder automatically when required: ```C# void MyClass.SetArrayNoAlloc( Godot.Collections.PackedVector2Array array); void MyCl
AI