Field Manual / A Claude Code Skill for Idea Validation: The One That Killed Its Own Verdict
A Claude Code Skill for Idea Validation: The One That Killed Its Own Verdict
Field manual · last reviewed 2026-08-13
An idea-validation skill for Claude Code is a SKILL.md workflow that turns "should I build this?" into a repeatable evidence gate instead of a vibe check: mine the public corpus, take a position, and return one of four verdicts — BUILD, SLICE-OK, PIVOT or KILL. AgentSource's Demand Validation Engine is that skill, and its distinguishing move is a second gate most validations skip. After demand passes, the discovery-path gate asks you to name the one channel that will carry users and show proof it already carries them. Its captured field report is two real runs. In the first, a photo-cleanup app never needed a corpus mine at all — willingness to pay was already revealed by public app-intelligence estimates (the leading swipe-cleaner around $1M/month, the category leader around $5M/month), so Phase 0 short-circuited to BUILD and the app shipped. In the second, a bucket-list travel app passed demand cleanly — people pay $4.99 to $239.99/yr for list-keeping with no booking function — and then the discovery gate killed it anyway: the category's own influencer-backed entrant, promoting to millions of followers, had topped out at 679 App Store ratings. The verdict flipped from PIVOT to KILL and the reversal was recorded in the dossier rather than quietly dropped. One-time $9, or $29 in the Founder Kit with five more skills. Compatible with Claude Code, claude.ai & Codex.
Idea validation fails in two directions, and both of them feel like diligence at the time. The agent cheerleads — strong signal, real pain point, users would love this — and you spend a month building something nobody wanted. Or it applies a venture-capital bar that nothing survives: no moat, crowded category, market's too small, someone could clone it in a weekend. The second failure is the more expensive one, because it looks like rigor. An idea-validation skill for Claude Code is a fixed gate the agent applies instead of improvising a standard each time. This page is about the one AgentSource ships for the job — the Demand Validation Engine, $9 — and it carries a captured field report: two real validation runs that went opposite ways, readable before you pay.
What the skill actually encodes
A skill is a plain-markdown SKILL.md package — a folder of instructions the agent loads and follows. Nothing to compile, no SDK, no API key. This one holds the parts of a validation that are easy to skip and expensive to skip:
- The revealed-willingness-to-pay check, first. Before any forum gets opened, the skill asks whether incumbents are already visibly earning money. If they are, the market is proven and the mine is a waste of a day.
- A corpus-mining brief written to disconfirm. Sources in priority order, signals to capture, verbatim quotes with citations required — and the strongest counter-signal mandatory in the output, not optional.
- The forum-sampling-bias rule. Loud "nobody would pay for that" threads are the anti-buyer's voice. People who would happily pay are not in the thread complaining about paying. Revealed revenue outranks sentiment.
- Four verdicts and exactly four kill conditions. BUILD, SLICE-OK, PIVOT, KILL — and a closed list of what justifies a kill. Crowded category, free alternatives, clone-ability, and small TAM are explicitly not on it.
- The discovery-path gate. A second gate after demand passes, covered below, and the reason this skill exists in the form it does.
- Anti-sycophancy rules. The agent must take a position. Hedge-verdicts like "the signal is encouraging" are banned outright.
The gate that reverses its own verdict
Most validation stops when demand looks real. That's the point where this one runs a second check: name the one channel that will carry users, and show evidence it already carries demand for something like this.
It exists because of a run that passed demand and should not have shipped.
A bucket-list travel app — keep and share a life list, no booking — cleared the demand check cleanly. People demonstrably pay for pure list-keeping apps in that space, from $4.99 up to $239.99/yr, with no booking function attached. The read at that stage was PIVOT: thesis adjusted, wedge identified, worth building.
Then the discovery gate killed it. The category's own influencer-backed entrant — founded by a family-travel brand with millions of YouTube and Instagram followers, promoting their own app to their own audience — had topped out at 679 App Store ratings at the time of the check. That is the most favorable distribution conditions the category has ever seen, and it converted that thin. The honest reading isn't that a newcomer needs a better channel strategy; it's that the niche itself is small.
The verdict flipped from PIVOT to KILL, and the reversal was written into the dossier verbatim — gate result FAIL, verdict superseded, reasoning attached. That's the behavior worth paying for. An agent that quietly revises its earlier answer teaches you nothing; one that records the contradiction leaves you an audit trail.
The other run: when the mine is a waste of a day
The field report's first run goes the opposite way, and it's the cheaper lesson.
The idea was a photo-library cleanup app — swipe to keep or delete, reclaim storage. Phase 0 checked whether willingness to pay was already revealed, and it was, before a single forum was opened. Public app-intelligence revenue estimates at the time of the run put the leading swipe-cleaner around $1M/month, the category leader around $5M/month, and the combined annual revenue of the top-10 cleaner apps near $200M.
Money at that scale settles the demand question. The skill's call was BUILD, with the residual risk named honestly rather than waved off: funnel conversion and store-review differentiation in a crowded category — real problems, but build and go-to-market problems, not demand problems. The app was built and shipped.
What the skill prevented there is worth naming, because it's invisible: a week of forum mining that would have surfaced hundreds of "just do it manually" and "there are free ones" comments. Every one of those is the anti-buyer talking, already outranked by revenue that exists.
Those figures are public estimates from app-intelligence tooling as of the run, not audited financials, and the products in the report are anonymized. That's the honest description of the evidence, and it's the same standard the skill demands of you.
What ships in the package
SKILL.md— the phases, the four verdicts, the four kill conditions, the discovery-path gate, the honesty rulesreferences/demand-dossier-template.md— the full dossier format with a rubric for filling it honestlyQUICKSTART.md— install steps for Claude Code, claude.ai and Codex, plus a first-run promptCHANGELOG.mdand a license
Where it sits before the build
Validation is the first gate, not the only one. Two skills pick up where it stops:
- Positioning & Market Map ($9) runs on a BUILD or SLICE-OK verdict. It forces the ICP, the competitor matrix, and a UVP that isn't fuzz — one doc, no code. A validated idea with fuzzy positioning is still an idea.
- Reddit Growth System ($9) is the discovery path made concrete for the channel most builders name and most builders get banned from. It recons the rules before posting, earns standing first, and queues by hand — which matters, because naming a channel in the gate and then getting removed from it is the same zero.
Buy the set: the Founder Kit
Individually those three are $9 each. The Founder Kit issues all three for $29, along with the X growth system, the product SEO site builder, and the AI visibility audit — validation through positioning through the channels that carry the thing. If you're at the "what do I build next" stage rather than the "polish what I built" stage, that's the set shaped like the problem.
One limit, stated plainly: this decides from public evidence, and public evidence is not certainty. It cannot tell you that your execution will be good, and it will occasionally be wrong in both directions — that's what the recorded counter-signal is for. What it reliably prevents is the expensive pair: building the thing the forums told you to build while the revenue said otherwise, and shipping a validated idea into a niche where the best-distributed player in the category could only find 679 people. Read the full field report first, then decide.
QUESTIONS
How is this different from a free "validate my idea" prompt?›
Free validation prompts mostly run one shape: search around, weigh the pros and cons, return an encouraging verdict. The failure isn't that they're free — it's that they have no conditions under which they say no, so they've never killed anything they should have. This skill names exactly four kill conditions and forbids the rest, which means a crowded category, free alternatives, clone-ability, or a small market cannot kill an idea on their own. It also forbids hedge-verdicts like "the signal is encouraging." The agent has to take a position and show the strongest evidence against it, with a citation.
What is the discovery-path gate, and why does it exist?›
It's a second gate that runs after demand passes: name the one channel that will carry users, and show proof it already carries demand for something like this. It exists because of a run in the field report. A bucket-list travel app cleared the demand check — people demonstrably pay from $4.99 up to $239.99/yr for pure list-keeping apps — and the gate killed it anyway, because the category's own influencer-backed entrant, founded by a family-travel brand with millions of followers and promoting to its own audience, had topped out at 679 App Store ratings. A massive built-in audience converting that thin is evidence the niche is small, not that a newcomer needs better marketing. A good product nobody can find earns the same zero as a bad one.
Does it replace talking to customers?›
It replaces waiting on interviews to make the call. The verdict comes from public evidence — competitor revenue, real user voice in forums and reviews, and channel proof — which is available today rather than in three weeks. A Mom-Test interview kit ships with the skill but is strictly opt-in and never prescribed by default. If you want interviews, run them; the skill just refuses to treat their absence as a reason to stall.
Won't the agent just tell me my idea is great?›
That's the specific failure the file is written against. The research brief is written to disconfirm your thesis rather than support it, the dossier template requires the single strongest counter-signal with a citation attached, and the honesty rubric bans the hedge language agents reach for when they don't want to disappoint you. There's also a rule about where sentiment comes from: loud "nobody would pay for that" forum threads are the anti-buyer's voice, and revealed revenue outranks them. Both directions are constrained — it isn't allowed to cheerlead, and it isn't allowed to kill on snobbery either.
Is it iOS-specific, and does it work in Codex?›
Not iOS-specific. The worked examples lean consumer-app because that's where it was field-tested, but the phases, the four verdicts, and both gates apply to SaaS, B2B tools, and content businesses without modification. On runtime: it's plain SKILL.md markdown with no runtime lock-in — compatible with Claude Code, claude.ai & Codex. Only the install location differs, and the QUICKSTART carries the exact steps for each.
THE GEAR
field-tested · see it work before you payDemand Validation Engine
A straight verdict on your idea — build it, cut it down, change it, or drop it — before you write code.
v1.1.0 · updated 2026-08-25 · field report included
Positioning Market Map
One page that says who buys this, who else they'd pick, and why you win. No code.
v1.0.0 · updated 2026-07-17 · field report included
Reddit Growth System
Post about your app on Reddit without losing your account. Read every sub's rules first, submit by hand.
v1.1.0 · updated 2026-08-25 · field report included