AGENTSOURCE

Field Manual / A Claude Code Skill for AI Visibility: Measure Whether Engines Name You at All

A Claude Code Skill for AI Visibility: Measure Whether Engines Name You at All

Field manual · last reviewed 2026-08-27

An AI-visibility skill for Claude Code is a SKILL.md workflow that measures whether AI answer engines name and link your brand when a buyer asks them a purchase question — a different surface from classic SEO, which fights for a blue link on a results page rather than a sentence inside a synthesized answer. AgentSource's AI Visibility Audit is that skill, priced $9 one-time. It fires a frozen prompt set across four buyer-intent families (category/'best', problem-first, comparison, brand-defense), runs each prompt N times per engine because one AI answer is a single sample of a stochastic system, and scores mention rate separately from citation rate — being named is not the same as being linked. It covers OpenAI, Perplexity, Gemini and Claude automatically on your own API keys, plus an assisted-capture mode for the surfaces with no clean API (Google AI Overviews, AI Mode, consumer Copilot, consumer ChatGPT), and labels every engine by the mode that actually produced its data. Its anti-fabrication rule runs end to end: an engine you did not query is reported 'not captured this run', never scored zero. The captured field report publishes an honest pilot — 3 prompts, 2 of 4 intent families, N=1 — and says in its own header that every rate in it is a snapshot rather than a validated frequency. Compatible with Claude Code, claude.ai & Codex.

Someone who wants what you sell no longer types four words into Google and scans ten blue links. They ask an assistant a whole question — "best project management app for a five-person team" — and read one synthesized paragraph naming three or four brands. Either you are in that paragraph or you do not exist for that buyer, and your rank tracker cannot see the answer at all. An AI-visibility skill for Claude Code is a fixed, repeatable measurement of that surface. This page is about the one AgentSource ships for the job — the AI Visibility Audit, $9 one-time — and the captured run you can read before you buy it.

Being named and being cited are two different scores

The single most common mistake here is treating a mention as the outcome. The skill scores them apart, and the field report contains the case that shows why.

On the "Tasklet vs Asana" prompt, ChatGPT named the brand first and described it favourably: a lighter, cheaper tool aimed at small teams. It linked it zero times. A mention counter records a win. A citation counter records nothing. Both are correct, and only the second one tells you the engine is describing you from memory rather than reading your page — which is a completely different problem with a completely different fix.

The same run shows the reverse split across engines on identical prompts:

Engine Captured Mention rate Citation rate
ChatGPT 3/3 1/3 (33%) 0/3 (0%)
Perplexity 2/3 2/2 (100%) 2/2 (100%)
Google AI Overview 2/3 0/2 (0%) 0/2 (0%)
Gemini / Claude / Copilot NOT CAPTURED

Three engines, three unrelated verdicts on the same brand. Checking one of them and generalizing is how a brand concludes it is fine.

One query is one sample, not a measurement

Ask an engine the same question twice and you get different brands and different links. So a screenshot is not evidence, and the skill treats a single run as noise by construction: each prompt is fired N times per engine and mention is reported as a frequency with its spread visible.

The published field report is, deliberately, the weak version of its own method. It was run at N=1, on 3 prompts, covering 2 of the 4 required intent families, and its header says so before you reach a single number — every rate is labelled a snapshot rather than a validated frequency, and the reader is told not to treat the blended figure as a trend point. That is the honest shape of a pilot, published rather than rounded away.

The prompt set is the actual product

Most of the thinking is not in the scoring, it is in what you ask. The skill freezes a buyer-intent prompt set across four families and keeps them cohorted so the next run is comparable:

  • Category / "best" — the discovery question, where you are competing against everyone.
  • Problem-first — the buyer who describes a symptom and has no vendor in mind yet.
  • Comparison — "you versus a named rival", the only family where the model is guaranteed to have your name in front of it.
  • Brand-defense — what the engine says when someone asks about you, which is where a hallucinated or hostile claim shows up.

The field report ran only the first and third, and it draws the obvious conclusion out loud rather than letting the gap pass: sentiment came back positive with nothing negative anywhere, but the family built to catch bad claims was not run — so that clean sheet is absence of signal, not a clean bill of health.

It also surfaces the pattern those families exist to expose. Share of voice was 9% on the broad "best PM app" prompt, 0% on "affordable alternatives", and 50% on the head-to-head comparison. Read together, that is a brand the model only knows when you put its name in the question — which is a far more useful sentence than any single percentage.

An engine you didn't query is "not captured", never zero

This is the rule the whole thing rests on, and it is boring in exactly the right way. A zero looks like a finding. It drags a blended average down, it reads as "we checked and you were absent", and it is indistinguishable in a table from a real result. So an engine with no key and no paste is reported not captured this run, and the report says which ones.

The field report turns that rule on itself three separate times:

  • Perplexity is labelled ASSISTED (manual paste) rather than its usual automated bucket, because no API call actually happened that run. The label follows what happened, not what normally happens.
  • The page-citability grade is deferred, not run — the sample brand had no live URL, and grading a page nobody fetched would be inventing a score. No grade is given.
  • ChatGPT's and Google's 0% citation readings are flagged unconfirmed, because a missing link could equally be real product behaviour or a lossy paste, and the next capture is told to check for citation chips explicitly.

A report willing to write "we don't know" in three places is the only kind whose other numbers are worth anything.

Fix your page, or go where the engines are already reading

The action list splits in two, and the second half is the one people miss: repair your own page, or earn placement on the third-party page the engines keep quoting. The audit inventories the sources cited in the answers, so you can see whose page the model is actually building its paragraph from. Frequently that is a roundup or review site you do not own, and getting into it beats another round of on-page edits.

Notably, in the captured run every link Perplexity returned was brand-owned — the vendors' own domains, no third-party listicle or review site anywhere. That is itself a finding about which lever was available.

Where it sits next to the SEO skill

This measures a surface. Product SEO Site Builder ($9) builds one — a guide page per long-tail keyword, gated on the live SERP so you are not writing into a query you cannot win. They are complementary and they are not the same job: the audit tells you the engines describe you without linking you; the builder is one of the ways you give them something worth linking. Buying the builder first and the audit never is how you end up with pages and no idea whether anything reads them.

What $9 buys, and what it doesn't

The AI Visibility Audit is one-time, not a subscription, and the captured field report is free to read in full before you decide. What you get is the method — the frozen prompt families, the two-mode engine coverage with every mode labelled, mention scored apart from citation, share of voice against a competitor set you name, the cited-source inventory, the six-axis page grade scored only from what is actually on the page, and a re-run diff so the audit becomes a cadence instead of an artefact.

What it does not do is make an engine like you. It also cannot cover a surface you give it no way to reach — no key and no paste means not captured, and it will say so on the page rather than quietly averaging you down. If you would rather have a number that always fills every cell, this is the wrong tool. Read the field report — including the three places it admits what it could not measure — and decide from that.

QUESTIONS

Isn't this just an SEO audit with a new name?

No, and the skill says so plainly rather than implying a secret lever. Classic SEO competes for a ranked link on a results page. This measures whether you appear inside the answer the engine writes — named in the sentence, and separately, linked as a source. Those are different outcomes and the field report contains a clean example of the gap: on the 'Tasklet vs Asana' prompt, ChatGPT described the brand favourably and first, and linked it zero times. A rank tracker cannot see that response at all, and a mention counter that does not score citation separately would have called it a win.

Why run the same prompt more than once?

Because one answer is one sample. The same question returns different brands and different citations on consecutive calls, so a single screenshot is noise dressed up as a measurement. The skill fires each prompt N times per engine and reports mention as a rate with its variance shown. The published field report is deliberately the weak version of this — it was run at N=1 — and its own header states that every rate in it is a single snapshot, not a validated frequency, and should not be read as a trend point.

What do I need to run it?

At least one API key among OpenAI, Perplexity, Gemini and Anthropic drives the automated engines on your own account; more keys widen coverage. Perplexity is called out as the highest-value single key because it is the most representative of its own consumer product. Google AI Overviews, AI Mode and consumer Copilot have no clean API, so they run through assisted capture — a SERP API key such as SerpApi or DataForSEO, or a human paste. Anything you cannot capture is reported 'not captured', which is the point rather than a limitation being hidden.

Can the score be inflated?

Not without breaking the rule the skill is built around. Every figure traces back to a logged prompt, response and timestamp, and an engine that was not queried is reported 'not captured this run' rather than scored zero — a distinction that matters, because a zero silently drags a blended average down and reads as a real finding. The field report applies the same discipline against itself in three places: it labels Perplexity ASSISTED rather than its usual AUTOMATED because no API call actually happened that run, it refuses to grade page citability at all because the sample brand had no live URL to fetch, and it flags a zero-citation reading as unconfirmed rather than banking it.

What does it tell me to do with the result?

The action list splits into two piles, and the second is usually the higher-leverage one: fix your own page, or earn placement on the third-party page the engines keep quoting. That split exists because the audit also inventories the sources being cited, so you can see whose page the answer is actually built from. In the field report the top actions came straight out of the scorecard — go after the discovery-stage prompts where competitors swept 100 percent of mentions, build comparison pages for the rivals that had none to match the one that was working, and re-capture the engines whose zero-citation reading was unconfirmed.

Is a one-off audit worth anything?

Less than a repeated one, which is why the config is frozen and cohorted — the prompt set, the competitor list and the intent families are pinned so the next run is comparable to this one rather than a fresh unrelated sample. The skill ships a monitoring cadence and a re-run diff format for exactly that reason. If you change the prompts between runs you have two audits and no trend.