I built Understudy because a lot of real work still spans native desktop apps, browser tabs, terminals, and chat tools. Most current agents live in only one of those surfaces.Understudy is a local-first desktop agent runtime that can operate GUI apps, browsers, shell tools, files, and messaging in one session. The part I'm most interested in feedback on is teach-by-demonstration: you do a task once, the agent records screen video + semantic events, extracts the intent rather than coordinates, and turns it into a reusable skill.Demo video: https://www.youtube.com/watch?v=3d5cRGnlb_0In the demo I teach it: Google Image search -> download a photo -> remove background in Pixelmator Pro -> export -> send via Telegram. Then I ask it to do the same for Elon Musk. The replay isn't a brittle macro: the published skill stores intent steps, route options, and GUI hints only as a fallback. In this example it can also prefer faster routes when they are available instead of repeating every GUI step.Current state: macOS only. Layers 1-2 are working today; Layers 3-4 are partial and still early. npm install -g @understudy-ai/understudy understudy wizard GitHub: https://github.com/understudy-ai/understudyHappy to answer questions about the architecture, teach-by-demonstration, or the limits of the current implementation.
FL score
out of 100
Verdict
high confidence
Competition
24
competitors found, emerging market, big tech present, funded players
Trend
1 community mentions
A local-first, cross-application desktop agent that learns complex tasks by observing a single demonstration and extracting intent, aiming to solve the unreliability of current AI agents.
The pain
The gap
Build angle
Strengths
Questions about this idea?
FlyBot reads the scoring and gives you a second opinion on “Understudy – Teach a desktop agent by demonstrating a task once”.
Risks
Next steps
Fly Labs Method
Is the pain real, is there a gap, is it the right time, can one person build it.
Strong problem validated by user frustration with current agents, a clear gap for reliable, cross-surface, intent-based automation, and demonstrated buildability, but with significant competition.
Value Equation
Dream outcome and how likely it feels, against the time and effort it costs.
High potential due to severe market pain and strong value proposition, but significant build complexity and competitive landscape.
One-Person Business
Curiosity pull, identity fit, and a path from free value to paid for a solo creator.
Strong idea with clear problem, capable builder, and intuitive user experience, but monetization strategy needs refinement.
Viral Frameworks
Hook strength, shareability, and how cheaply it can be tested.
Strong value proposition and reachable audience, but significant risks in reliability and competitive landscape require careful validation.
Builder Lens
Evidence the problem exists, timing, defensibility, and a model that fits on a napkin.
High potential with a clear, painful problem, an elegant solution, and initial validation through a working demo.
Why this verdict
Five lenses, one composite. How scoring works
The angle
This weekend
Who is already there, emerging market
Open-source local-first desktop AI agent with access to files, apps, and terminal for task automation.
Pricing: free (open source, BYOK)
Meta-backed local AI agent app for macOS/Windows controlling local files, apps, CLI, and desktop automation.
Pricing: unknown
Voice-first open-source macOS desktop agent using accessibility APIs for app control with local knowledge graph memory.
Pricing: free (open source MIT, BYOK)
Open-source desktop agent for executing tasks across apps, research, analysis, long-running automation.
Pricing: free (open source)
Browser extension to record interactions and teach AI agents by example/demonstration.
Pricing: free
Local browser/desktop agent controlling real browser, logged-in accounts, tasks like posting/research.
Pricing: unknown
Anthropic's real-time collaborative desktop agent/canvas for pair working.
Pricing: subscription (Claude Pro)
Chinese local AI workspace accessing files, email, calendar, agentic workflows.
Pricing: free
An AI agent for automating anything on any device (web, desktop, mobile). It clicks, types, and navigates based on a prompt.
Pricing: Not explicitly stated, but user reviews mention using it for repetitive tasks, implying a potential subscription or usage-based model.
A desktop robotic process automation system that automates tedious tasks like financial data entry with a text prompt, performing clicks and keystrokes.
Pricing: Not explicitly stated, but focuses on businesses and enterprises.
A smart agent that handles computer work from automation to web research to creating deliverables, using natural language input. Available on Web and Desktop (Mac Apple Silicon).
Pricing: Offers free credits during beta and a 30% off subscription code 'LAUNCH30' through August 2025.
A desktop app that brings Manus out of the cloud and onto your computer, allowing an AI agent to work directly with local files, tools, and apps on macOS and Windows, using text prompts to execute tasks across files, tools, and apps.
Pricing: Not explicitly stated on Product Hunt, but previous Manus launches exist.
Gaps they leave open
What people say, 1 mentions
Giving AI agents API keys is a nightmare, so I built a single gateway for them
r/SaaS
Recent news
Meta's Manus Launches Desktop App With AI Agent for Tasks Across Files, Apps
PCMag, March 18 2026
GTC Spotlights NVIDIA RTX PCs and DGX Sparks Running Latest Open Models and AI Agents Locally
NVIDIA, March 17 2026
The 10 Best AI Agents for Desktop Automation in 2026
Fazm Blog, March 14 2026
Understudy: Teach Desktop Agent Tasks with One Demonstration - AIToolly
AIToolly, March 12 2026
Show HN: Understudy – Teach a desktop agent by demonstrating a task once
James Routley (Hacker News), March 13 2026
Market signals
The market shows strong interest in AI desktop agents that learn from demonstration and natural language, with a growing number of YC-backed startups and recent product launches focusing on cross-application automation and intent-based learning rather than rigid scripting.
Last summer we faced a conundrum at my company, Tiger Data, a Postgres cloud vendor whose main business is in timeseries data. We were trying to grow our business towards emerging AI-centric workloads and wanted to provide a state-of-the-art hybrid search stack in Postgres. We'd already built pgvectorscale in house with the goal of scaling semantic search beyond pgvector's main memory limitations. We just needed a scalable ranked keyword search solution too.The problem: core Postgres doesn't provide this; the leading Postgres BM25 extension, ParadeDB, is guarded behind AGPL; developing our own extension appeared daunting. We'd need a small team of sharp engineers and 6-12 months, I figured. And we'd probably still fall short of the performance of a mature system like Parade/Tantivy.Or would we? I'd be experimenting long enough with AI-boosted development at that point to realize that with the latest tools (Claude Code + Opus) and an experienced hand (I've been working in database systems internals for 25 years now), the old time estimates pretty much go out the window.I told our CTO I thought I could solo the project in one quarter. This raised some eyebrows.It did take a little more time than that (two quarters), and we got some real help from the community (amazing!) after open-sourcing the pre-release. But I'm thrilled/exhausted today to share that pg_textsearch v1.0 is freely available via open source (Postgres license), on Tiger Data cloud, and hopefully soon, a hyperscalar near you:https://github.com/timescale/pg_textsearchIn the blog post accompanying the release, I overview the architecture and present benchmark results using MS-MARCO. To my surprise, we were not only able to meet Parade/Tantivy's query performance, but exceed it substantially, measuring a 4.7x advantage on query throughput at scale:https://www.tigerdata.com/blog/pg-textsearch-bm25-fu
AI
Hi HN!I recently switched from a Fedora/GNOME laptop to a MacBook Air. My old setup served me well as a portable workstation, but I’ve started traveling more while working remotely and needed something with similar performance but better battery life. The main thing I missed was a simple taskbar that shows the windows in the current workspace instead of a Dock that mixes everything together.I built boringBar so I would not have to use the Dock. It shows only the windows in the current Space, lets you switch Spaces by scrolling on the bar, and adds a desktop switcher so you can jump directly to any Space. You can also hide the system Dock, pin apps, preview windows with thumbnails, and launch apps from a searchable menu (I keep Spotlight disabled because for some reason it uses a lot of system resources on my machine).I’ve been dogfooding it for a few months now, and it finally felt polished enough to share.It’s for people who like macOS but want window management to feel a bit more like GNOME, Windows, or a traditional taskbar. It’s also for people like me who wanted an easier transition to macOS, especially now that Windows feels increasingly user-hostile.I’d love feedback on the UX, bugs, and whether this solves the same Dock/Spaces pain for anyone else.P.S. It might also appeal to people who feel nostalgic for the GNOME 2 desktop of yore. I started my Linux journey with it, and boringBar brings back some of that feeling for me.
AI
### Describe the project you are working on Godot C# bindings ### Describe the problem or limitation you are having in your project For the past weeks, I've been discussing with several Unity users intending to move to Godot C# regarding dealing with the C# garbage collector. The most common complaint I hear from users is that, in Unity, allocations can trigger unexpected GC spikes into the game. In Godot, we target to make all of the high performance APIs (those that intended to be called every frame) not allocate any memory, so theoretically the GC should not be a problem. Additionally, Godot starting from 4.0, uses the Microsoft CoreCLR version of .net, which also supposedly has a better garbage collector than Unity. But in all, after several discussions with Unity users, neither is enough reassurance for them, and they would really feel safer if Godot exposed a zero allocation API. ### Describe the feature / enhancement and how it helps to overcome the problem or limitation The idea of this proposal is that Godot exposes zero allocation versions of many functions in the C# API, that users can use if they desire. Technically, this could be done from the binding generator itself, without breaking compatibility, and without doing any modification to Godot itself. ### Describe how your proposal will work, with code, pseudo-code, mock-ups, and/or diagrams **WARNING** I am not familiar with C#, so take this as pseudocode. Imagine you have two functions exposed as to C#: ```C# void MyClass.SetArray( Vector2[] array); Vector2[] MyClass.GetArray(); ``` This works and is pretty and intuitive. However, it has two problems: * GC is allocated on return * Memory is copied to Godot native formats every time there is a call. The idea is to add NoAlloc versions, which can be generated directly by the binder automatically when required: ```C# void MyClass.SetArrayNoAlloc( Godot.Collections.PackedVector2Array array); void MyCl
AI