Skip to content
FlatRelay

How to measure AI supplier visibility without overclaiming

A repeatable protocol for manufacturing AI search tests: choose buyer prompts, log citations and errors, handle failed runs, and report limits.

Sep 14, 2026AI Visibility

By FlatRelay · Updated

Practical guidance from a consultancy selling SEO and AI visibility services. Sources, examples, and limitations are identified in the text. Editorial standards and corrections.

  • measurement
  • manufacturing
  • ai-citations

Measure AI supplier visibility with a fixed set of buyer questions, repeated runs, preserved answers, and separate counts for brand mentions, linked citations, and factual accuracy. Record engine, date, search mode, and test conditions. Treat the result as a sample of answers, not a market share estimate or a forecast of RFQs.

The protocol below is a suggested workflow for a manufacturing supplier. It has not been validated as a universal benchmark. Example prompts and calculations are illustrative; they do not represent a completed client study. For an actual account of our starting point, see FlatRelay’s own SEO experiment, which discloses the evidence it lacks.

Choose prompts from real procurement questions

Start with questions in sales emails, quotation requests, and qualification calls, after removing confidential details. Cover different stages: finding a process provider, verifying a capability, comparing alternatives, and understanding quote requirements. Keep branded questions separate because a prompt that names your company measures a different task from discovering it.

For an illustrative CNC supplier, a small starting set could include:

  • “Which suppliers offer low-volume CNC aluminum enclosures for industrial equipment?”
  • “What evidence should I request before selecting a CNC machining supplier?”
  • “How should I compare local machining with an overseas supplier for repeat orders?”
  • “What information does a machine shop need to quote an enclosure?”

Adapt these to the business’s actual capabilities and target market. Do not insert your own brand into a discovery prompt or keep only prompts that already return a favorable answer. State why each question was included and freeze the wording before collecting a baseline.

Define the test conditions before running it

Choose engines your buyers plausibly use and record the product surface, model if exposed, search setting, account state, language, and location setting. A location setting is not proof of the network’s physical location. Note that distinction when it matters.

Use new conversations for independent runs. A follow-up in the same conversation has additional context and should be labeled separately. A practical small pilot might use three repetitions per prompt, engine, and observation window. That is a proposed workload choice, not a statistically sufficient sample for every business. More repeats can reveal variation but still do not make a convenience sample representative of all buyers.

Capture the full answer and every citation URL. If the engine fails or returns a challenge, record the failure and retry policy. Never silently replace an unfavorable answer with a better one. Follow the provider’s usage rules and use supported access methods for automation.

Keep a reusable observation sheet

Field What to record
Run ID Unique reference connecting the row to the saved answer
Prompt Exact wording, intent, and branded/non-branded classification
Conditions UTC timestamp, engine, product, model if visible, search mode, language, location setting
Outcome Completed, failed, blocked, or no substantive answer
Brand mention Whether the supplier was named in the answer
Owned citation Whether an actual link resolves to the supplier’s domain
Third-party citation Source URL discussing the supplier, kept separate from owned citations
Accuracy Claim, source used to check it, correct/incorrect/unverifiable
Evidence Full text or permitted capture, reviewer, and notes

The reviewer should check the cited page rather than assuming that a visible link supports the adjacent statement. An engine can mention the right company and still misstate certification scope, geography, minimum order quantity, or service availability. Mark a claim unverifiable when evidence is missing; do not count uncertainty as correctness.

Calculate distinct measures

Mention rate: completed answers naming the supplier divided by all completed answers in the stated test group.

Owned citation rate: completed answers containing at least one verified supplier-domain citation divided by all completed answers in that group. Count each answer once, even if it contains several supplier links.

Reviewed claim accuracy: supported correct claims divided by all reviewed claims, with incorrect and unverifiable counts reported separately. Specify the denominator so readers can see whether a rate excludes uncertainty.

Illustrative calculation: 4 of 12 completed answers mention a supplier and 2 of 12 link to its domain. Mention rate is 33.3%; owned citation rate is 16.7%. If two additional attempts failed, report “12 completed, 2 failed” alongside the rates. Do not describe 33.3% as the proportion of real buyers who will discover the supplier.

Keep results by engine and prompt intent before considering any combined number. A combined result depends on how you weight engines and questions. If you create a score, publish the weights and missing-data policy and avoid presenting it as an industry standard.

Compare changes without claiming causation

Save a dated baseline before editing the site. Keep the next test’s questions, repetitions, and conditions as consistent as possible. Annotate page releases, new third-party mentions, provider changes, and other campaigns. Report absolute counts beside percentages: moving from one citation to two is a small sample even though the relative increase is large.

Run technical checks separately. A public page that loads in your browser does not prove that every crawler can retrieve it. Similarly, a page’s absence from an answer does not prove a crawler block. Google’s AI feature guidance describes eligibility; eligibility alone does not guarantee inclusion.

Connect referral visits to qualified RFQs where measurement is available, while respecting the site’s privacy arrangements. A mention without a click can still exist, but this protocol cannot establish that it influenced a purchase. Keep citation observations, traffic, and sales outcomes in separate columns rather than inferring one from another.

Common questions

How often should we repeat the test?

Use a schedule that matches how often you can act on findings and collect enough observations. A baseline and a repeat after a substantive release can be more useful than daily checks with no changes to investigate. Record the observation window and avoid interpreting a single day’s fluctuation as a durable trend.

Can a competitor citation tell us which page to create?

It can suggest a missing answer, but the cited page still needs inspection. Identify the buyer question it resolves, the evidence it supplies, and whether your business can provide an original, accurate answer. Copying a competitor’s wording or creating a page for a capability you do not have is not an appropriate response.

What should we fix first?

Prioritize verified errors and missing facts on commercially important pages. Confirm crawl access, then clarify capabilities and supporting evidence using the capability checklist and content actions guide. Re-test after publication. If results remain unchanged, keep that outcome in the report rather than inventing a benefit.

All articles