Real problems people complain about online, pulled every morning and scored out of 100. Build, validate, or skip. How scoring works
I use AI agents to build UI features daily. The thing that kept annoying me: the agent writes code but never sees what it actually looks like in the browser. It can’t tell if the layout is broken or if the console is throwing errors.So I built a CLI that lets the agent open a browser, interact with the page, record what happens, and collect any errors. Then it bundles everything — video, screenshots, logs — into a self-contained HTML file I can review in seconds. proofshot start --run "npm run dev" --port 3000 # agent navigates, clicks, takes screenshots proofshot stop It works with whatever agent you use (Claude Code, Cursor, Codex, etc.) — it’s just shell commands. It's packaged as a skill so your AI coding agent knows exactly how it works. It's built on agent-browser from Vercel Labs which is far better and faster than Playwright MCP.It’s not a testing framework. The agent doesn’t decide pass/fail. It just gives me the evidence so I don’t have to open the browser myself every time.Open source and completely free.Website: https://proofshot.argil.io/
Hacker News5mo agoToolAI
Hacker News5mo agoToolAI
Hacker News5mo agoToolAI
Hi all, I'm Peter at Staff Engineer and Mozilla.ai and I want to share our idea for a standard for shared agent learning, conceptually it seemed to fit easily in my mental model as a Stack Overflow for agents.The project is trying to see if we can get agents (any agent, any model) to propose 'knowledge units' (KUs) as a standard schema based on gotchas it runs into during use, and proactively query for existing KUs in order to get insights which it can verify and confirm if they prove useful.It's currently very much a PoC with a more lofty proposal in the repo, we're trying to iterate from local use, up to team level, and ideally eventually have some kind of public commons.At the team level (see our Docker compose example) and your coding agent configured to point to the API address for the team to send KUs there instead - where they can be reviewed by a human in the loop (HITL) via a UI in the browser, before they're allowed to appear in queries by other agents in your team.We're learning a lot even from using it locally on various repos internally, not just in the kind of KUs it generates, but also from a UX perspective on trying to make it easy to get using it and approving KUs in the browser dashboard. There are bigger, complex problems to solve in the future around data privacy, governance etc. but for now we're super focussed on getting something that people can see some value from really quickly in their day-to-day.Tech stack:* Skills - markdown* Local Python MCP server (FastMCP) - managing a local SQLite knowledge store* Optional team API (FastAPI, Docker) for sharing knowledge across an org* Installs as a Claude Code plugin or OpenCode MCP server* Local-first by default; your knowledge stays on your machine unless you opt into team sync by setting the address in config* OSS (Apache 2.0 licensed)Here's an example of something which seemed straight forward, when asking Claude Code to write a GitHub action it often
Hacker News5mo agoToolAI
Kitten TTS (https://github.com/KittenML/KittenTTS) is an open-source series of tiny and expressive text-to-speech models for on-device applications. We had a thread last year here: https://news.ycombinator.com/item?id=44807868.Today we're releasing three new models with 80M, 40M and 14M parameters.The largest model (80M) has the highest quality. The 14M variant reaches new SOTA in expressivity among similar sized models, despite being <25MB in size. This release is a major upgrade from the previous one and supports English text-to-speech applications in eight voices: four male and four female.Here's a short demo: https://www.youtube.com/watch?v=ge3u5qblqZA.Most models are quantized to int8 + fp16, and they use ONNX for runtime. Our models are designed to run anywhere eg. raspberry pi, low-end smartphones, wearables, browsers etc. No GPU required! This release aims to bridge the gap between on-device and cloud models for tts applications. Multi-lingual model release is coming soon.On-device AI is bottlenecked by one thing: a lack of tiny models that actually perform. Our goal is to open-source more models to run production-ready voice agents and apps entirely on-device.We would love your feedback!
Hacker News6mo agoToolAI
Individuals and teams using AI tools waste money on excessive token consumption from suboptimal prompts and lack of optimization.
X6mo agoToolAI
Developers building AI agents face barriers like Docker configuration, server management, and setup time, preventing quick deployment and testing.
X6mo agoToolAI
OpenClaw users repeatedly hit walls around costs and model choices, leading to uncertainty and inefficiency rather than feature overload.
X5mo agoToolAI
Users frustrated with CHAI forcing premium subscription, Cai lacking security, and Emochi providing overly long or short boring answers; need a reliable AI chatbot app balancing features, security, and engaging responses.
X5mo agoToolAI
Email is one of those tools we check daily but its underlying experience didn’t evolve much. I use Gmail, as probably most of you reading this.The Arc browser brought joy and taste to browsing the web. Cursor created a new UX with agents ready to work for you in a handy right panel.I use these three tools every day. Since Arc was acquired by Atlassian, I’ve been wondering: what if I built a new interface that applied Arc’s UX to email rather than browser tabs, while making AI agents easily available to help manage emails, events, and files?I built a frontend PoC to showcase the idea.Try it: https://demo.define.appI’m not sure about it though... Is it worth continuing to explore this idea?
Hacker News5mo agoToolAI
I finally built this app after many years of being sick of unlocking my phone every goddamn time I need to take or view my notes. It particularly sucks when I'm doing my grocery and going down the list.I started building last year June. This is a native app written in Kotlin. And since I'm a 100% Web dev guy, I gotta say this wouldn't have been possible without this AI to assist me. So this isn't "vibe-coded". I simply used the chat interface in Gemini website, manually copy paste codes to build and integrate every single thing in the app! I used gemini to build it just because I was piggybacking on my last company's enterprise subscription. I personally didn't subscribe to any AI (and still don't cuz the free quota seems enough for me :)So I certainly have learnt alot about Android development, architecture patterns, Kotlin syntax, and obeying Google's whims. Can't say I love it all, but for the sake of this app, I will :)Anyway, I finally have the app I wish existed, and I'm using it everyday. It not only does the main thing I needed it to do, but there's also all this stuff:- Make your notes private if you don't want to show them on lock screen. - Create check/to-do lists. - Set one time or recurring reminders. - Full-text search your notes in the app. - Speech-to-text. - Organize your notes with custom or color labels. - Pin the app as a widget on your home screen. - You can auto backup and restore your notes on new install or Android device. - Works offline. - And no funny business happening in the background https://joonote.com/privacyIt's 30-day trial, then a one-time $9.99 to go Pro forever.I would love you all to check it out, FWIW.Ok thanks!
Hacker News5mo agoToolAI
AI builders risk massive bills from uncontrolled LLM usage; author confirms 'problem is real' while building ThSkyshield after weeks of private development.
X5mo agoToolAI
AI-assisted code deployments at scale (e.g., Amazon) delete live environments or cause Sev-1 outages due to inadequate safeguards, exacerbated by staff cuts and rushed adoption.
X6mo agoToolAI
Even a 45-year-old academic article written by a human is detected as 77% AI-generated due to advanced vocabulary and grammar. Skilled writers and researchers must intentionally degrade their sentence structure and paragraphs to score below detection thresholds, complicating their work unnecessarily.
X5mo agoToolAI
Hi HN! We're Vincent and Jochen from sitefire (https://sitefire.ai). Our platform makes it easy for brands to improve their visibility in AI search.We’ve been working together for years and have backgrounds in RL/optimization at Stanford and software engineering. We came to this idea after speaking with marketing teams who were seeing declining traffic due to Google’s AI Overviews and didn’t know what to do.This space can feel esoteric. Many case studies, few actual studies. Constant battle against myths (e.g. you need a llms.txt vs. you don't need a llms.txt) and "GEO hacks". We try to be more data-driven. And we try to be more bold and build a system that not only monitors, but actually improves traffic from AI search.While Google performs a single search, AI search engines expand the user prompt into 3-10 fan-out queries. The sourced pages are ranked using a classified algorithm similar to Reciprocal Rank Fusion (RFF). Finally, the LLMs skim the pages and decide what snippets to cite. Our goal is making sure brands have the right content that makes it through this funnel.Here is how sitefire works:- The user defines a set of prompts they want to monitor. These are synthetic prompts - we generate them based on SEO keywords and their monthly search volume.- We submit these prompts to ChatGPT, Gemini, Google AI Mode, etc. on a daily basis and capture the answers. We extract fan-out queries, sourced pages, citations, and brand mentions.- For each topic, our agents analyze which web pages are sourced and cited the most, and why. They also consider similar pages that you already have.- Based on the diagnosis, our content agents draft improvements or create new pages, and push them directly to the client’s CMS.- We integrate with the client’s network logs and Google Analytics to monitor the increase in AI bot requests and human referrals to their page.This system is continuously updated, so it always shows which content works, and how
Hacker News5mo agoToolAI
I go on a lot of backcountry trips where I barely get cell service. If my group splits, nobody knows knows where anyone is until you regroup at camp or at your destination. You can buy Garmin radios or try to set up an ATAK, but ATAK is Android-only and assumes you have a TAK Server running somewhere to make use of all of the functionality. Cool tools themselves, but expensive to set up correctly. I just wanted two iPhones to share their location directly over Bluetooth when cell coverage was lacking.Red Grid Link does that. Start a session, and anyone nearby running the app shows up on your offline map. When they walk out of range their marker stays as a "ghost" that slowly fades.The hard part was making sync reliable over BLE. The connections drop all the time. Someone turns a corner, walks behind a vehicle, whatever. I built a CRDT sync layer (LWW Register + G-Counter) so there's never merge conflicts. Each update is just under 200 bytes (from what I have tested so far). When a user/teammate disappears the app does exponential backoff from 2 to 30 seconds before giving up and marking them as a ghost.Everything is encrypted (AES-256-GCM, ECDH P-256 key exchange per peer pair). Sessions can require a PIN or QR code to join. It also offers offline topo maps with MGRS grid coordinates, same system as in my other app, Red Grid MGRS.The app is free, and I'm looking for some honest feedback from other real-world users. Let me know if you have any questions!
Hacker News5mo agoToolAI
Building and deploying wallet solutions for web3 startups requires significant technical infrastructure and security expertise. A lean, developer-friendly wallet toolkit could reduce friction for early-stage crypto teams.
YC Graveyard4y agoToolAI
Hey HN! We’re Hayden, Ronan, Avi, and Warren of Voltair (https://voltairlabs.com/). We’re making weatherized, hybrid-fixed drones deployed for power utility inspections.Here’s some footage: https://vimeo.com/1173862237/ac28095cc6?share=copy&fl=sv&fe=... and a photo of our latest prototype: https://imgur.com/a/bYHnqZ4.The U.S. has 7M miles of power lines (enough to go to the moon and back 14 times), and they're aging. Over 50% of all power flows through transformers that are at least 30 years old, which is about when they start to fail.Power line conductors are just bare metal with 4,000-765,000 volts sitting on ceramic insulators, usually held up by pieces of wood. It’s a cost effective and relatively reliable way to move power. But when the wood starts to rot, or the cotter pin falls out, and a live conductor is dropped on a dead tree on a windy day, you get devastating wildfires like the Palisades Fire in LA last year.Most utilities solve this problem with foot patrols. Linemen drive out with a clipboard or an iPad, and run through a checklist with binoculars to visually confirm everything is in order. A lineman can inspect about 50-150 poles per day, yet even the smallest rural electric cooperatives (with about ~20 employees) have about 50,000 distribution poles. Clearly the math doesn’t work out. As a result, a given utility pole is inspected about every 10 years (at least that’s what they tell their insurance adjuster).Helicopters are also used, but cost $25k to get off the ground, and more importantly, every year linemen die in helicopter crashes. Satellites can’t deliver the mm precision needed for these inspections. So drones have emerged as the best solution. Georgia Power saved 60% on operating expenses when they switched to using drones, and Xcel power found drones to find 60% more defects than foot patrols (because of pole-top vantage point).Problem #2: Drones are held back by the need
Hacker News6mo agoToolAI
Hey HN! We're Aakash and Viswesh, and we're building Canary (https://www.runcanary.ai). We build AI agents that read your codebase, figure out what a pull request actually changed, and generate and execute tests for every affected user workflow.Aakash and I previously built AI coding tools at Windsurf, Cognition, and Google. AI tools were making every team faster at shipping, but nobody was testing real user behavior before merge. PRs got bigger, reviews still happened in file diffs, and changes that looked clean broke checkout, auth, and billing in production. We saw it firsthand. We started Canary to close that gap. Here's how it works:Canary starts by connecting to your codebase and understands how your app is built: routes, controllers, validation logic. You push a PR and Canary reads the diff, understands the intent behind the changes, then generates and runs tests against your preview app checking real user flows end to end. It comments directly on the PR with test results and recordings showing what changed and flagging anything that doesn't behave as expected. You can also trigger specific user workflow tests via a PR comment.Beyond PR testing, tests generated from the PR can be moved into regression suites. You can also create tests by just prompting what you want tested in plain English. Canary generates a full test suite from your codebase, schedules it, and runs it continuously. One of our construction tech customers had an invoicing flow where the amount due drifted from the original proposal total by ~$1,600. Canary caught the regression in their invoice flow before release.This isn't something a single family of foundation models can do on its own. QA spans across many modalities like source code, DOM/ARIA, device emulators, visual verifications, analyzing screen recordings, network/console logs, live browser state etc. for any single model to be specialized in. You also need custom browser fleets, user ses
Hacker News6mo agoToolAI
Sales reps spend significant time duplicating information they've already communicated (emails, calls, meetings) by manually logging it into CRM systems. This creates friction between communication and data entry, leads to outdated or incomplete CRM records, and pulls focus from actual selling activities. Teams also struggle to quickly extract insights from their conversations (follow-up needs, objection patterns, ICP changes) without manually reviewing dozens of interactions.
Product Hunt6mo agoToolAI