Hi HN, we’re Lewis and Edgar, building Captain to simplify unstructured data search (https://runcaptain.com). Captain automates the building and maintenance of file-based RAG pipelines. It indexes cloud storage like S3 and GCS, plus SaaS sources like Google Drive. There’s a quick walkthrough at https://youtu.be/EIQkwAsIPmc.We also put up this demo site called “Ask PG’s Essays” which lets you ask/search the corpus of pg’s essays, to get a feel for how it works: https://pg.runcaptain.com. The RAG part of this took Captain about 3 minutes to set up.Here are some sample prompts to get a feel for the experience:“When do we do things that don't scale? When should we be more cautious?” https://pg.runcaptain.com/?q=When%20do%20we%20do%20things%20...“Give me some advice, I'm fundraising” https://pg.runcaptain.com/?q=Give%20me%20some%20advice%2C%20...“What are the biggest advantages of Lisp” https://pg.runcaptain.com/?q=what%20are%20the%20biggest%20ad...A good production RAG pipeline takes substantial effort to build, especially for file workloads. You have to handle ETL or text extraction, chunking, embedding, storage, search, re-ranking, inference, and often compliance and observability – all while optimizing for latency and reliability. It’s a lot to manage. grep works well in some cases, but for agents, semantic search provides significantly higher performance. Cursor uses both and reports 6.5%–23.5% accuracy gains from vector search over grep (https://cursor.com/blog/semsearch).We’ve spent the past four years scaling RAG pipelines for companies, and Edgar’s work at Purdue’s NLP lab directly informed our chunking techniques. In conversations with dozens of engineers, we repeatedly saw DIY pipelines produce inconsistent results, even after weeks of tuning. Many teams lacked clarity on which retrieval strategies best fit their data.We realized that a system t
FL score
out of 100
Verdict
high confidence
Competition
7
competitors found, emerging market, funded players
Trend
8 community mentions
Automated RAG for Files simplifies building and maintaining production-ready LLM pipelines from unstructured data, targeting engineers frustrated with complex frameworks and partial solutions.
The pain
The gap
Build angle
Strengths
Questions about this idea?
FlyBot reads the scoring and gives you a second opinion on “Captain (YC W26) – Automated RAG for Files”.
Risks
Next steps
Fly Labs Method
Is the pain real, is there a gap, is it the right time, can one person build it.
Captain tackles a significant pain point in building and maintaining RAG pipelines, with clear user frustrations about existing solutions. However, the market is highly competitive and the complexity of building a truly automated, robust solution for files is a very high bar for a solo builder.
Value Equation
Dream outcome and how likely it feels, against the time and effort it costs.
Captain has a strong value proposition in a growing market with clear pain, but building a defensible, comprehensive solution against well-funded incumbents is a significant challenge.
One-Person Business
Curiosity pull, identity fit, and a path from free value to paid for a solo creator.
Captain leverages strong founder expertise to address a clear problem, but the sheer complexity of the solution and market competition make it a challenging solo endeavor.
Viral Frameworks
Hook strength, shareability, and how cheaply it can be tested.
Captain has a clear value prop for a specific audience, but faces significant distribution challenges and high assumption risks in a competitive, complex market.
Builder Lens
Evidence the problem exists, timing, defensibility, and a model that fits on a napkin.
Captain addresses a high-demand, growing problem in the AI space with a clear target user, but the 'narrowest wedge' and initial validation against entrenched competition need refinement.
Why this verdict
Five lenses, one composite. How scoring works
The angle
This weekend
Who is already there, emerging market
A flexible framework for building LLM-powered applications, emphasizing orchestration of prompts, tools, and agents.
Pricing: Open-source and free to use; offers a managed service with potential costs.
A data framework focused on connecting custom data sources to LLMs for retrieval-augmented generation (RAG).
Pricing: Free plan available, Starter ($50/month), Pro ($500/month), Enterprise (contact for pricing).
An open-source framework for building production-ready LLM applications, particularly strong in document-centric RAG pipelines and semantic search.
Pricing: Free to use; contact deepset for Haystack Enterprise pricing.
A document parsing and extraction platform that transforms unstructured documents into structured data elements for AI applications.
Pricing: Offers open-source version, and API services with improved performance and enterprise features. Pricing for API covers extraction only; other costs are separate.
A managed AI platform that offers an agent builder and RAG features as part of the Google Cloud ecosystem.
Pricing: Part of Google Cloud pricing, usage-based.
An open-source, low-code platform for building LLM applications and agents with a visual canvas.
Pricing: Free plan available, Starter ($35/month), Pro ($65/month), Enterprise (contact for pricing).
A platform designed to transform unstructured data into optimized vector search indexes, facilitating retrieval-augmented generation pipelines.
Pricing: Not explicitly stated on public pages, but focuses on transforming unstructured data to optimized vector search indexes.
What they charge
What people say, 8 mentions
I built a mobile IV therapy company from $0 to $2M in 12 months, merged it into a competitor I ran as CEO and scaled from $2.4M to $10M, stepped down, and started completely over. 3 months in 2026 and we're doing $250K/month.
r/Entrepreneur
My YouTube videos get 100 views. They generate $12,000/month.
r/SaaS
I've seen hundreds of pitch decks this year and here is my learnings:
r/Entrepreneur
"Don't code. Just sell." : The rule that saved our SaaS
r/Entrepreneur
What are you working on?
r/SaaS
One huge client isn’t validation. $10,000 from 50 strangers is.
r/Entrepreneur
Built a file automation tool after getting tired of repetitive dev tasks — looking for honest feedback
r/SaaS
Launching MailSprinter: Automate Bulk File Emailing—with Smart PDF Compression for Large Documents 🚀
r/SaaS
Recent news
Launch HN: Captain (YC W26) – Automated RAG for Files
Hacker News, March 13 2026
We benchmarked Unstructured.io vs naive 500-token splits — both needed 1.4M+ tokens. We didn't expect them to tie. POMA AI needed 77% less. : r/Rag - Reddit
Reddit, March 20 2026
7 best LangChain alternatives I've tested in 2026 - Gumloop
Gumloop, March 20 2026
Best LlamaIndex Memory Alternatives for AI Agents (2026) - Vectorize
Vectorize, March 14 2026
Top 10 Haystack Alternatives & Competitors in 2026 - G2
G2, January 09 2026
Market signals
The market for automated RAG for files is a rapidly growing and significant niche within the broader AI industry. Recent funding rounds, such as Vectorize's $3.6M seed round in October 2024 and LangChain's $10M seed round in April 2023, indicate strong investor interest. The increasing complexity of LLM applications and the need for up-to-date, domain-specific information are driving demand for solutions that simplify RAG pipeline building and maintenance.
What frustrates people
Last summer we faced a conundrum at my company, Tiger Data, a Postgres cloud vendor whose main business is in timeseries data. We were trying to grow our business towards emerging AI-centric workloads and wanted to provide a state-of-the-art hybrid search stack in Postgres. We'd already built pgvectorscale in house with the goal of scaling semantic search beyond pgvector's main memory limitations. We just needed a scalable ranked keyword search solution too.The problem: core Postgres doesn't provide this; the leading Postgres BM25 extension, ParadeDB, is guarded behind AGPL; developing our own extension appeared daunting. We'd need a small team of sharp engineers and 6-12 months, I figured. And we'd probably still fall short of the performance of a mature system like Parade/Tantivy.Or would we? I'd be experimenting long enough with AI-boosted development at that point to realize that with the latest tools (Claude Code + Opus) and an experienced hand (I've been working in database systems internals for 25 years now), the old time estimates pretty much go out the window.I told our CTO I thought I could solo the project in one quarter. This raised some eyebrows.It did take a little more time than that (two quarters), and we got some real help from the community (amazing!) after open-sourcing the pre-release. But I'm thrilled/exhausted today to share that pg_textsearch v1.0 is freely available via open source (Postgres license), on Tiger Data cloud, and hopefully soon, a hyperscalar near you:https://github.com/timescale/pg_textsearchIn the blog post accompanying the release, I overview the architecture and present benchmark results using MS-MARCO. To my surprise, we were not only able to meet Parade/Tantivy's query performance, but exceed it substantially, measuring a 4.7x advantage on query throughput at scale:https://www.tigerdata.com/blog/pg-textsearch-bm25-fu
AI
Hi HN!I recently switched from a Fedora/GNOME laptop to a MacBook Air. My old setup served me well as a portable workstation, but I’ve started traveling more while working remotely and needed something with similar performance but better battery life. The main thing I missed was a simple taskbar that shows the windows in the current workspace instead of a Dock that mixes everything together.I built boringBar so I would not have to use the Dock. It shows only the windows in the current Space, lets you switch Spaces by scrolling on the bar, and adds a desktop switcher so you can jump directly to any Space. You can also hide the system Dock, pin apps, preview windows with thumbnails, and launch apps from a searchable menu (I keep Spotlight disabled because for some reason it uses a lot of system resources on my machine).I’ve been dogfooding it for a few months now, and it finally felt polished enough to share.It’s for people who like macOS but want window management to feel a bit more like GNOME, Windows, or a traditional taskbar. It’s also for people like me who wanted an easier transition to macOS, especially now that Windows feels increasingly user-hostile.I’d love feedback on the UX, bugs, and whether this solves the same Dock/Spaces pain for anyone else.P.S. It might also appeal to people who feel nostalgic for the GNOME 2 desktop of yore. I started my Linux journey with it, and boringBar brings back some of that feeling for me.
AI
### Describe the project you are working on Godot C# bindings ### Describe the problem or limitation you are having in your project For the past weeks, I've been discussing with several Unity users intending to move to Godot C# regarding dealing with the C# garbage collector. The most common complaint I hear from users is that, in Unity, allocations can trigger unexpected GC spikes into the game. In Godot, we target to make all of the high performance APIs (those that intended to be called every frame) not allocate any memory, so theoretically the GC should not be a problem. Additionally, Godot starting from 4.0, uses the Microsoft CoreCLR version of .net, which also supposedly has a better garbage collector than Unity. But in all, after several discussions with Unity users, neither is enough reassurance for them, and they would really feel safer if Godot exposed a zero allocation API. ### Describe the feature / enhancement and how it helps to overcome the problem or limitation The idea of this proposal is that Godot exposes zero allocation versions of many functions in the C# API, that users can use if they desire. Technically, this could be done from the binding generator itself, without breaking compatibility, and without doing any modification to Godot itself. ### Describe how your proposal will work, with code, pseudo-code, mock-ups, and/or diagrams **WARNING** I am not familiar with C#, so take this as pseudocode. Imagine you have two functions exposed as to C#: ```C# void MyClass.SetArray( Vector2[] array); Vector2[] MyClass.GetArray(); ``` This works and is pretty and intuitive. However, it has two problems: * GC is allocated on return * Memory is copied to Godot native formats every time there is a call. The idea is to add NoAlloc versions, which can be generated directly by the binder automatically when required: ```C# void MyClass.SetArrayNoAlloc( Godot.Collections.PackedVector2Array array); void MyCl
AI