I built a browser-based tool for detecting objects in satellite imagery using vision-language models (VLMs). You draw a polygon on the map and enter a text prompt such as "swimming pools", "oil tanks", or "buses". The system scans the selected area tile-by-tile and returns detections projected back onto the map as GeoJSON.Pipeline: select area and zoom level, split the region into mercantile tiles, run each tile with the prompt through a VLM, convert predicted bounding boxes to geographic coordinates (WGS84), and render the results back on the map.It works reasonably well for distinct structures in a zero-shot setting. occluded objects are still better handled by specialized detectors like YOLO models.There is a public demo and no login required. I am mainly interested in feedback on detection quality, performance tradeoffs between VLMs and specialized detectors, and potential real-world use cases.
FL score
out of 100
Verdict
high confidence
Competition
6
competitors found, emerging market, funded players
Trend
1 community mentions
A browser-based tool for text-prompt object detection in satellite imagery, targeting the critical gap of VLM limitations with dense/occluded objects for enterprise users.
The pain
The gap
Build angle
Strengths
Questions about this idea?
FlyBot reads the scoring and gives you a second opinion on “Satellite imagery object detection using text prompts”.
Risks
Next steps
Fly Labs Method
Is the pain real, is there a gap, is it the right time, can one person build it.
Strong problem clarity with a clear gap in VLM accuracy for specific object types, but the market is crowded with well-funded competitors, and the buildability for a solo founder is challenging.
Value Equation
Dream outcome and how likely it feels, against the time and effort it costs.
A growing market with a clear pain, but challenges in competitive differentiation and solo execution.
One-Person Business
Curiosity pull, identity fit, and a path from free value to paid for a solo creator.
Clear problem with potential monetization, but significant challenges for a solo founder in execution, audience reach, and maintaining simplicity.
Viral Frameworks
Hook strength, shareability, and how cheaply it can be tested.
Clear value proposition for a specific pain, but significant risks in assumption validation, distribution, and solo execution in a competitive market.
Builder Lens
Evidence the problem exists, timing, defensibility, and a model that fits on a napkin.
Real demand and a clear specific problem, but needs a narrower focus and evidence of solving the hardest parts of the problem better than existing solutions.
Why this verdict
Five lenses, one composite. How scoring works
The angle
This weekend
Who is already there, emerging market
Esri integrates vision-language models into its ArcGIS platform for geospatial analysis, allowing users to extract features from imagery using natural language prompts for tasks like object detection and segmentation.
Pricing: Contact vendor for pricing (ArcGIS platform pricing varies, VLMs are integrated capabilities).
VisionAgent offers reasoning-driven object detection using text prompts without custom training, leveraging AI agents to understand object attributes, relationships, and dynamic states.
Pricing: Available via API, processing takes 20-30 seconds per image (working on speed improvements). Contact vendor for specific API pricing.
Element 84 utilizes remote sensing vision-language models like SkyCLIP and RemoteCLIP to enable 'queryable Earth' functionality, allowing retrieval of images and geolocations using text queries over large geographical areas.
Pricing: Information not readily available, likely custom solutions or enterprise-focused.
Picterra is a cloud-based geospatial intelligence platform that allows users to detect objects in satellite and aerial imagery, primarily through training custom detectors.
Pricing: Contact vendor for pricing; often offers different tiers based on usage and features.
Sinergise develops enterprise-level solutions for managing spatial data and processing satellite imagery, with Sentinel Hub being a key platform for accessing and processing Earth observation data.
Pricing: Offers various plans, including a free tier for basic usage, and paid plans based on data volume and services (e.g., €30/month for advanced features).
Hasty.ai provides an AI-powered annotation platform for computer vision teams to build, fine-tune, iterate, and manage AI models faster with high-quality training data.
Pricing: Contact vendor for pricing; offers a free trial.
What they charge
What people say, 1 mentions
I built "Google Lens for teams" – Search and ask anything about 1000s of your photos & videos – a ChatGPT for your photos and videos
r/SaaS
Recent news
Show HN: Satellite imagery object detection using text prompts
Hacker News, March 11 2026
Show HN: Detect any object in satellite imagery using a text prompt
Hacker News, March 12 2026
VisionAgent: Reasoning-Driven Agentic Object Detection
Product Hunt, February 12 2025
Talk to Your Imagery: Vision-Language Models for Geospatial Analysis
Esri, August 14 2025
Building a queryable Earth with vision-language foundation models
Element 84, March 19 2024
Market signals
The geospatial artificial intelligence market is a rapidly growing market, valued at $38 billion in 2023 and projected to reach $64.60 billion by 2029, with some reports indicating a potential increase to $126.58 billion by 2035. Key drivers include increased deployment of AI-driven geospatial solutions for real-time decision-making in urban planning, defense, agriculture, and disaster management. Recent trends highlight the adoption of cloud-based platforms and the integration of GeoAI with IoT and edge computing for real-time insights.
What frustrates people
Last summer we faced a conundrum at my company, Tiger Data, a Postgres cloud vendor whose main business is in timeseries data. We were trying to grow our business towards emerging AI-centric workloads and wanted to provide a state-of-the-art hybrid search stack in Postgres. We'd already built pgvectorscale in house with the goal of scaling semantic search beyond pgvector's main memory limitations. We just needed a scalable ranked keyword search solution too.The problem: core Postgres doesn't provide this; the leading Postgres BM25 extension, ParadeDB, is guarded behind AGPL; developing our own extension appeared daunting. We'd need a small team of sharp engineers and 6-12 months, I figured. And we'd probably still fall short of the performance of a mature system like Parade/Tantivy.Or would we? I'd be experimenting long enough with AI-boosted development at that point to realize that with the latest tools (Claude Code + Opus) and an experienced hand (I've been working in database systems internals for 25 years now), the old time estimates pretty much go out the window.I told our CTO I thought I could solo the project in one quarter. This raised some eyebrows.It did take a little more time than that (two quarters), and we got some real help from the community (amazing!) after open-sourcing the pre-release. But I'm thrilled/exhausted today to share that pg_textsearch v1.0 is freely available via open source (Postgres license), on Tiger Data cloud, and hopefully soon, a hyperscalar near you:https://github.com/timescale/pg_textsearchIn the blog post accompanying the release, I overview the architecture and present benchmark results using MS-MARCO. To my surprise, we were not only able to meet Parade/Tantivy's query performance, but exceed it substantially, measuring a 4.7x advantage on query throughput at scale:https://www.tigerdata.com/blog/pg-textsearch-bm25-fu
AI
Hi HN!I recently switched from a Fedora/GNOME laptop to a MacBook Air. My old setup served me well as a portable workstation, but I’ve started traveling more while working remotely and needed something with similar performance but better battery life. The main thing I missed was a simple taskbar that shows the windows in the current workspace instead of a Dock that mixes everything together.I built boringBar so I would not have to use the Dock. It shows only the windows in the current Space, lets you switch Spaces by scrolling on the bar, and adds a desktop switcher so you can jump directly to any Space. You can also hide the system Dock, pin apps, preview windows with thumbnails, and launch apps from a searchable menu (I keep Spotlight disabled because for some reason it uses a lot of system resources on my machine).I’ve been dogfooding it for a few months now, and it finally felt polished enough to share.It’s for people who like macOS but want window management to feel a bit more like GNOME, Windows, or a traditional taskbar. It’s also for people like me who wanted an easier transition to macOS, especially now that Windows feels increasingly user-hostile.I’d love feedback on the UX, bugs, and whether this solves the same Dock/Spaces pain for anyone else.P.S. It might also appeal to people who feel nostalgic for the GNOME 2 desktop of yore. I started my Linux journey with it, and boringBar brings back some of that feeling for me.
AI
### Describe the project you are working on Godot C# bindings ### Describe the problem or limitation you are having in your project For the past weeks, I've been discussing with several Unity users intending to move to Godot C# regarding dealing with the C# garbage collector. The most common complaint I hear from users is that, in Unity, allocations can trigger unexpected GC spikes into the game. In Godot, we target to make all of the high performance APIs (those that intended to be called every frame) not allocate any memory, so theoretically the GC should not be a problem. Additionally, Godot starting from 4.0, uses the Microsoft CoreCLR version of .net, which also supposedly has a better garbage collector than Unity. But in all, after several discussions with Unity users, neither is enough reassurance for them, and they would really feel safer if Godot exposed a zero allocation API. ### Describe the feature / enhancement and how it helps to overcome the problem or limitation The idea of this proposal is that Godot exposes zero allocation versions of many functions in the C# API, that users can use if they desire. Technically, this could be done from the binding generator itself, without breaking compatibility, and without doing any modification to Godot itself. ### Describe how your proposal will work, with code, pseudo-code, mock-ups, and/or diagrams **WARNING** I am not familiar with C#, so take this as pseudocode. Imagine you have two functions exposed as to C#: ```C# void MyClass.SetArray( Vector2[] array); Vector2[] MyClass.GetArray(); ``` This works and is pretty and intuitive. However, it has two problems: * GC is allocated on return * Memory is copied to Godot native formats every time there is a call. The idea is to add NoAlloc versions, which can be generated directly by the binder automatically when required: ```C# void MyClass.SetArrayNoAlloc( Godot.Collections.PackedVector2Array array); void MyCl
AI