Turnkey custom coding-agent benchmarks for small teams
Build a lightweight, hosted eval-generation service for small dev teams: given a GitHub repo, auto-generate PR-based regression benchmarks for coding agents and output a simple dashboard comparing model/agent performance on the team's own codebase, undercutting self-hosted or enterprise-grade tools with a turnkey SaaS for indie developers and small startups who want to know which coding agent actually works on their stack without setting up sandboxes themselves.
What to build
A hosted SaaS that connects to a small team's GitHub repo, auto-generates regression benchmarks from real past PRs and bugs in that codebase, then runs multiple coding agents (Claude, Copilot, Cursor, etc.) against them and shows a dashboard of which agent actually produces working, mergeable code on their specific stack.
Self-bench's Show HN proves there's appetite for benchmarking coding agents against real software instead of toy problems, but it's a self-hosted dev tool — a turnkey hosted version lowers the setup bar for small teams who just want a quick answer on which agent works on their repo.
Demand
Solo builders and small startups now juggle several coding agents and have no fast way to know which one actually performs on their own codebase and build environment, rather than on a generic public leaderboard.
- Hacker NewsNews
Show HN: Self-bench – benchmark coding agents on real-world software — 4 points, 0 comments, introducing an open-source tool for benchmarking agents against real repos instead of synthetic problems
Stack
- GitHub REST/Webhooks API
- Docker or E2B/Daytona sandboxes for isolated test runs
- Claude API, OpenAI API, GitHub Copilot API
- Postgres (via Supabase)
- Temporal or BullMQ for job orchestration
- Next.js dashboard
Solo + AI difficulty
The hard part is reliably sandboxing arbitrary repos (install steps, flaky or non-deterministic tests, varied build tools) well enough to score a PR as pass/fail — this is real infra work, not prompt engineering. The GitHub ingestion, agent API calls, and dashboard UI are comparatively easy with AI assistance. Realistic MVP for one language/framework (e.g. Node or Python) is 4-6 weeks solo.
- Entry threshold
- Core pieces (PR-to-eval generation, sandboxed harness execution, model API billing passthrough) are moderate to build solo with AI assistance over a few weeks, but reliable sandboxing infrastructure and cost-control for running many models adds real engineering overhead beyond a weekend project.
- Window
- 6-12 months
Where to find first users
- Show HN launch
- Indie Hackers community
- r/ExperiencedDevs and r/programming
- Product Hunt launch
Competitors
Counter-signals & risks
Established eval/observability platforms and model labs could add repo-specific regression benchmarking as a feature, eroding differentiation for a standalone SaaS.
Auto-generating meaningful PR-based regression benchmarks from arbitrary repos is technically hard (flaky tests, non-deterministic agent outputs, repo-specific build/test environments), limiting turnkey feasibility and raising support burden.
Price-sensitive indie developers and small startups may prefer free, self-hosted tools like the original self-bench project over a paid hosted service.
Original title: Show HN: Self-bench – benchmark coding agents on real-world software
Related signals
- Developer tools#2
Blacklist monitoring tool for cloud-hosted IPs
Build a quick ASN/IP reputation and blacklist-monitoring tool for indie hosts and SaaS founders on cloud providers (DigitalOcean, Linode, Vultr, Hetzner) that checks their IP/ASN against UCEPROTECT, Spamhaus and other RBLs, alerts them before mail or traffic gets silently filtered, and gives clear guidance on delisting or whitelisting options, including documenting predatory listing practices so users know what they are dealing with. Target developers and small businesses who get blocked without warning and only discover it by accident, as in this case.
Demand5/10measuredBuildability8/10est.Competition6/10est.via Hacker News33 competitors - Developer tools#1
Turning real site CSS into AI build prompts
Build a developer tool that extracts real CSS and design tokens from any live website and converts them into structured prompts or design specs for AI coding assistants like Claude, Cursor or v0, aimed at indie developers tired of generic 'AI slop' UI output. A paid version could add Figma export, Tailwind config generation, multi-page batch inspection, or team-shared design libraries, monetizing a problem the free pikspec extension only partially solves.
Demand5/10est.Buildability9/10est.Competition8/10est.via Reddit11 competitor - Developer tools#3
DNS setup wizard for custom domain email
Build a dead-simple custom domain email setup tool for solo founders and small businesses: a wizard that auto-generates and verifies DNS records (MX, SPF, DKIM, DMARC) for Google Workspace, Microsoft 365, or cheap forwarding providers like Cloudflare Email Routing, with plain-English error diagnostics when verification fails. Package it as a one-time paid tool or a thin SaaS layer on top of existing mailbox providers, targeting non-technical founders who get stuck on DNS setup when launching a new domain.
Demand3/10measuredBuildability8/10est.Competition6/10est.via Hacker News33 competitors