Hi HN, I forked chromium and built agent-browser-protocol (ABP) after noticing that most browser-agent failures aren’t really about the model misunderstanding the page. Instead, the problem is that the model is reasoning from a stale state.ABP is designed to keep the acting agent synchronized with the browser at every step. After each action (click, type, etc), it freezes JavaScript execution and rendering, then captures the resulting state. It also compiles the notable events that occurred during that action loop, such as navigation, file pickers, permission prompts, alerts, and downloads, and sends that along with a screenshot of the frozen page state back to the agent.The result is that browser interaction starts to feel more like a multimodal chat loop. The agent takes an action, gets back a fresh visual state and a structured summary of what happened, then decides what to do next from there. That fits much better with how LLMs already work.A few common browser-use failures ABP helps eliminate: * A modal appears after the last Playwright screenshot and blocks the input the agent was about to use * Dynamic filters cause the page to reflow between steps * An autocomplete dropdown opens and covers the element the agent intended to click * alert() / confirm() interrupts the flow * Downloads are triggered, but the agent has no reliable way to know when they’ve completedAs proof, ABP with opus 4.6 as the driver scores 90.5% on the Online Mind2Web benchmark. I think modern LLMs already understand websites, they just need a better tool to interact with them. Happy to answer questions about the architecture, forking chrome or anything else in the comments below.Try it out: `claude mcp add browser -- npx -y agent-browser-protocol --mcp` (Codex/OpenCode instructions in the docs)Demo video: https://www.loom.com/share/387f6349196f417d8b4b16a5452c3369
FL score
out of 100
Verdict
high confidence
Competition
12
competitors found, emerging market, funded players
Trend
8 community mentions
An open-source browser protocol that solves AI agent reliability issues by synchronizing browser state, backed by strong benchmark performance, but faces a crowded market and high build complexity for a solo builder.
The pain
The gap
Build angle
Strengths
Questions about this idea?
FlyBot reads the scoring and gives you a second opinion on “Open-source browser for AI agents”.
Risks
Next steps
Fly Labs Method
Is the pain real, is there a gap, is it the right time, can one person build it.
The idea addresses a real, specific, and severe pain point for AI agent developers. While the technical solution is novel and highly desired, the market is crowded with funded competitors, posing a significant challenge for a solo builder.
Value Equation
Dream outcome and how likely it feels, against the time and effort it costs.
This idea targets a rapidly growing market with a strong value proposition, but faces significant technical complexity and competition.
One-Person Business
Curiosity pull, identity fit, and a path from free value to paid for a solo creator.
A highly technical, niche solution with strong problem clarity, but significant complexity for a solo builder to maintain and monetize effectively.
Viral Frameworks
Hook strength, shareability, and how cheaply it can be tested.
A well-defined value proposition for a specific, reachable audience, but the open-source nature requires a clear monetization strategy to be a viable micro-SaaS.
Builder Lens
Evidence the problem exists, timing, defensibility, and a model that fits on a napkin.
Strong demand and a clear solution for a critical pain point in a rapidly growing field, with a well-defined narrow wedge.
Why this verdict
Five lenses, one composite. How scoring works
The angle
This weekend
Who is already there, emerging market
An open-source framework that converts website interfaces into structured text, allowing AI agents to interact with the web more reliably and efficiently.
Pricing: Free tier available. Paid plans include: AI Agent Tasks ($0.01 init + per-step, with per-step costs varying by LLM model, e.g., Browser Use 2.0 at $0.006/step), Browser Sessions ($0.06/hour), Skill Creation ($2.00/skill), Skill Execution ($0.02/call), Proxy Data ($10/GB). Business/Scaleup plans offer discounts.
Provides serverless cloud browser infrastructure for AI agents and automation, supporting Playwright, Puppeteer, and Selenium at scale with features like stealth mode and session persistence.
Pricing: Not explicitly stated on the public page, but charges on a consumption basis.
An open-source CLI tool built in Rust that provides AI agents with direct browser control through command-line interface.
Pricing: Free and open-source.
An open-source headless browser API for AI agents focusing on providing infrastructure with transparency and self-hosted options.
Pricing: Free and open-source.
A platform that enables technical and non-technical users to build complex AI browser agents for reliable back-office task automation.
Pricing: Not explicitly stated on the public page.
A full Chromium-based browser with Perplexity's AI search engine built-in, offering autonomous navigation, form filling, and multi-step task completion.
Pricing: Free tier available. Agent mode requires a Plus subscription ($20/month).
OpenAI's AI-powered web browser that integrates browsing into the ChatGPT ecosystem, allowing agents to navigate and interact with websites.
Pricing: Requires a Plus subscription ($20/month) for Agent Mode.
A Chrome extension that automates browser tasks using natural language commands, enabling web scraping, form filling, and multi-tab support.
Pricing: $25/month with a 7-day free trial. Free with your own Gemini API key or Claude subscription.
Browser-based AI agents offering unlimited runs for a fixed cost, focusing on core automation, workflows, and AI power-ups.
Pricing: Free tier with 50 Cloud AI calls/month, 5 workflows, 10 nodes/workflow. Basic plan at $20/month with 5000 Cloud AI calls/month, 50 workflows, 50 nodes/workflow. Enterprise plan with custom pricing for unlimited features.
Transforms existing browsers into AI assistants for tasks like summarization, research, form filling, and content creation, emphasizing privacy and local operation.
Pricing: Not explicitly stated on the public page.
An open-source browser with built-in AI agents that emphasizes privacy and automation through natural language, acting as a privacy-first alternative.
Pricing: Free and open-source.
A web data API for AI agents and developers that handles web scraping, search, and full browser interaction, providing clean structured data.
Pricing: Free tier with 500 credits (includes 5 hours of free browser usage). Paid plans start from $16/month.
What they charge
What people say, 8 mentions
building saas with ai agent browser control, how do i not end up hosting illegal shit by accident?
r/SaaS
Market validation: AI-powered job application automation tool
r/Entrepreneur
How does one build Browser Agents?
r/SaaS
Lessons learned after launching my first SaaS after 3 months of nights and weekends
r/Entrepreneur
Non negotiable rule for browser agents: human approval before submit, send, or payment
r/SaaS
Which opensource AI agent do you use for your business?
r/Entrepreneur
Exploring new product category: Embeddable Web Agents
r/Entrepreneur
Building an AI agentic extention for your browser!!
r/SaaS
Recent news
MyNextBrowser: Make any browser agentic and automate workflows
Product Hunt, March 14 2026
Browser Use raises $17M to help steer AI agents through the internet
SiliconANGLE, March 23 2025
AI Browser: Automate anything online with a single prompt.
Product Hunt, 4 months ago (approx. November 2025)
Browserfly: AI agent that lives in your browser
Product Hunt, Not specified, but launched on Product Hunt.
Browser Use Raises $17 Million to Accelerate Development of Its Web Agent Competing with Operator
ActuIA, April 19 2025
Market signals
The AI browser market is experiencing significant growth, projected to reach $76.8 billion by 2034, from $4.5 billion in 2024, with a CAGR of 32.8%. North America dominates this market due to early adoption of AI and substantial investments in R&D. The market is shifting from passive browsing to active, intelligent agents capable of understanding user intent and automating complex tasks, driven by advancements in foundational AI models.
What frustrates people
Last summer we faced a conundrum at my company, Tiger Data, a Postgres cloud vendor whose main business is in timeseries data. We were trying to grow our business towards emerging AI-centric workloads and wanted to provide a state-of-the-art hybrid search stack in Postgres. We'd already built pgvectorscale in house with the goal of scaling semantic search beyond pgvector's main memory limitations. We just needed a scalable ranked keyword search solution too.The problem: core Postgres doesn't provide this; the leading Postgres BM25 extension, ParadeDB, is guarded behind AGPL; developing our own extension appeared daunting. We'd need a small team of sharp engineers and 6-12 months, I figured. And we'd probably still fall short of the performance of a mature system like Parade/Tantivy.Or would we? I'd be experimenting long enough with AI-boosted development at that point to realize that with the latest tools (Claude Code + Opus) and an experienced hand (I've been working in database systems internals for 25 years now), the old time estimates pretty much go out the window.I told our CTO I thought I could solo the project in one quarter. This raised some eyebrows.It did take a little more time than that (two quarters), and we got some real help from the community (amazing!) after open-sourcing the pre-release. But I'm thrilled/exhausted today to share that pg_textsearch v1.0 is freely available via open source (Postgres license), on Tiger Data cloud, and hopefully soon, a hyperscalar near you:https://github.com/timescale/pg_textsearchIn the blog post accompanying the release, I overview the architecture and present benchmark results using MS-MARCO. To my surprise, we were not only able to meet Parade/Tantivy's query performance, but exceed it substantially, measuring a 4.7x advantage on query throughput at scale:https://www.tigerdata.com/blog/pg-textsearch-bm25-fu
AI
Hi HN!I recently switched from a Fedora/GNOME laptop to a MacBook Air. My old setup served me well as a portable workstation, but I’ve started traveling more while working remotely and needed something with similar performance but better battery life. The main thing I missed was a simple taskbar that shows the windows in the current workspace instead of a Dock that mixes everything together.I built boringBar so I would not have to use the Dock. It shows only the windows in the current Space, lets you switch Spaces by scrolling on the bar, and adds a desktop switcher so you can jump directly to any Space. You can also hide the system Dock, pin apps, preview windows with thumbnails, and launch apps from a searchable menu (I keep Spotlight disabled because for some reason it uses a lot of system resources on my machine).I’ve been dogfooding it for a few months now, and it finally felt polished enough to share.It’s for people who like macOS but want window management to feel a bit more like GNOME, Windows, or a traditional taskbar. It’s also for people like me who wanted an easier transition to macOS, especially now that Windows feels increasingly user-hostile.I’d love feedback on the UX, bugs, and whether this solves the same Dock/Spaces pain for anyone else.P.S. It might also appeal to people who feel nostalgic for the GNOME 2 desktop of yore. I started my Linux journey with it, and boringBar brings back some of that feeling for me.
AI
### Describe the project you are working on Godot C# bindings ### Describe the problem or limitation you are having in your project For the past weeks, I've been discussing with several Unity users intending to move to Godot C# regarding dealing with the C# garbage collector. The most common complaint I hear from users is that, in Unity, allocations can trigger unexpected GC spikes into the game. In Godot, we target to make all of the high performance APIs (those that intended to be called every frame) not allocate any memory, so theoretically the GC should not be a problem. Additionally, Godot starting from 4.0, uses the Microsoft CoreCLR version of .net, which also supposedly has a better garbage collector than Unity. But in all, after several discussions with Unity users, neither is enough reassurance for them, and they would really feel safer if Godot exposed a zero allocation API. ### Describe the feature / enhancement and how it helps to overcome the problem or limitation The idea of this proposal is that Godot exposes zero allocation versions of many functions in the C# API, that users can use if they desire. Technically, this could be done from the binding generator itself, without breaking compatibility, and without doing any modification to Godot itself. ### Describe how your proposal will work, with code, pseudo-code, mock-ups, and/or diagrams **WARNING** I am not familiar with C#, so take this as pseudocode. Imagine you have two functions exposed as to C#: ```C# void MyClass.SetArray( Vector2[] array); Vector2[] MyClass.GetArray(); ``` This works and is pretty and intuitive. However, it has two problems: * GC is allocated on return * Memory is copied to Godot native formats every time there is a call. The idea is to add NoAlloc versions, which can be generated directly by the binder automatically when required: ```C# void MyClass.SetArrayNoAlloc( Godot.Collections.PackedVector2Array array); void MyCl
AI