The methods column
How we measure, and what we[1][1] Refuse to claim: six specific non-guarantees, printed in full in chapter 06. Not a disclaimer buried in a footer link. refuse to claim
AI visibility is easy to fake and hard to measure. This page is the arithmetic, the sampling rules, and the explicit list of things CiteTrail does not guarantee. So you can judge our numbers instead of trusting them.
†: the central claim
A model never picks your score
Every number in CiteTrail is computed by deterministic code from responses we stored. Models do exactly three jobs: run the prompt, classify sentiment and extract entities from what came back, and write prose summaries of results already calculated. If a figure moved, a rule moved it, and the rule is on this page.
01: How prompts are selected
How prompts are selected
Three ways in: write them yourself, import a list, or start from prompts suggested for your brand and category.
Every prompt is then classified across 13 intent categories and 5 funnel stages and scored out of 100 for buyer value. High is 70 and above, medium is 45 and above. That score decides what gets run most often.
Region and language are separate measurements. A prompt belongs to one market, and results from different markets are never averaged into a single figure. Where a scoring signal has no connected data behind it, the factor falls back to neutral. We do not fill the gap with an estimate and then score it.
What the scoring rewards
Commercial intent, not comfort. A prompt is never weighted up because you perform well on it. The highest-value questions in most workspaces are the ones being lost.
What counts as one measurement
One prompt, on one engine, in one region, at one point in time. Everything above that (trends, share of answer, confidence) is built by aggregating those units. Nothing is aggregated across units that were not comparable in the first place.
02: Why one answer proves nothing
One answer is an anecdote
Language models are non-deterministic. Ask the same question twice and you can get two different answers naming two different sets of brands. This is not a flaw in the measurement. It is the thing being measured.
- The same question, twice
- Ask a model the same thing two minutes apart and you can get two answers naming two different sets of brands. That is how these systems work, not a fault in them.
- Prompts run on a schedule, not once
- Every tracked prompt runs repeatedly, across engines, over a window. A pattern is something that survives repetition. A single result is a story.
- Providers change things without telling you
- Models get swapped, system prompts get edited, grounding gets turned on. None of it is announced. Continuous sampling is the only way to see it happen.
Consequence: a finding based on fewer than two runs is labelled Low confidence. Always. No combination of agreement between those runs can override it.
03: How confidence is derived
How confidence is derived
Confidence is consistency multiplied by how much we sampled. Both halves are published.
Step one: how consistently did it repeat
Each sub-score is the share of runs that agree with the majority outcome, from 0 to 100. They are blended by how much each one matters to a business decision.
consistency = mention agreement × 0.40
+ competitor agreement × 0.25
+ citation agreement × 0.20
+ sentiment agreement × 0.15Step two: how much evidence backs it
Consistency alone would let one lucky run read as certainty. So it is gated by a sample-size factor that can only ever range from 0.3 to 1.0.
confidence = consistency × sample_factor(runs, engines)- Runs
- Contribute up to 0.7 of the factor, saturating at 8 runs.
- Engines
- Contribute up to 0.3 of the factor, saturating at 3 engines.
- Floor
- The factor never drops below 0.3, so a single consistent run keeps some confidence, but can never reach full confidence.
The caps are deliberate. Running a prompt 40 times does not buy more confidence than running it 8 times, and an engine beyond the 3 already tracked does not buy more. Beyond that point you are measuring the same thing again.
Step three: what you see
The score becomes a label, and the label always travels with the sample that produced it.
- High75 and above
- Medium50 and above
- Lowbelow 50
One override outranks the arithmetic: fewer than two runs is always low confidence, whatever the score says.
04: How citations are captured
How citations are captured
Sources are stored at run time, not reconstructed later. Each one keeps its URL, the engine that returned it, its position in the answer, and the time it was captured.
Domains are classified as editorial, review, community or owned, which is what makes a source-mix chart mean anything.
Engines that return no sources are recorded as returning none. That row stays empty. Guessing which pages a model probably read would fill the chart and break the page.
What a stored run keeps
- The response
- The answer text the assistant actually returned, kept so a finding can be checked rather than believed.
- The engine and the moment
- Which surface produced it, and when it was captured.
- The cited sources
- Every cited URL with its position, classified by domain and source type.
- What changed
- Mention flips, position deltas, and sources that appeared or dropped since the previous run.
We store the runs and responses we measured so findings stay auditable. See security for exactly what we hold and how it's handled.
05: Observed versus inferred
Four words, four columns, never one total
Observed
The referrer directly identifies an assistant. We saw the session arrive. This is the only tier we will ever describe as attribution.
Estimated
Modelled influence. A visit whose shape matches assistant-driven behaviour, without the referrer to prove it.
Inferred
Correlation between AI exposure and traffic movement. Suggestive. Not proof, and labelled as such.
Exposure
Your brand appeared in an answer. A fact about the answer, not about traffic. That is precisely why it's so often quietly counted as traffic elsewhere.
These four are never added together. A single blended “AI traffic” number would be the most quotable figure on this page and the least defensible.
06: What CiteTrail does not guarantee
What CiteTrail does not guarantee
This chapter is not a disclaimer we hid at the bottom. If a competitor promises you any of the following, they are describing something outside their control.
No guaranteed ranking, mention or traffic
Providers change models and grounding without notice, and outputs vary run to run. Anyone promising otherwise is selling something they cannot control.
Inferred influence is not revenue
It is a correlation with a label on it. Do not put it in a board deck as pipeline.
Provider APIs are not consumer apps
What we measure is the model behind the assistant. Your buyer's chat window may say something different, and we would rather tell you that than let you assume otherwise.
A visibility score is not a search ranking
There is no position one in an AI answer, and the two numbers are not comparable in either direction.
An action is a recommendation, not a causal claim
It is ranked by expected value. It becomes evidence only after a re-scan measures the change.
We do not report a number we did not measure
Empty stays empty. If a workspace has no data, it shows no data.
What we do promise is narrower and more useful: every number traces to a stored response, every finding shows the sample behind it, and anything we inferred says so.
07: Questions
Methodology, answered plainly
Does an AI model calculate my visibility score?
No. Scoring is deterministic code over stored responses. Models run prompts, classify sentiment, extract entities and write summaries. They never pick a number.
Why do my results change between runs?
Because language models are non-deterministic and providers change models and grounding without notice. That variance is the thing being measured, not noise being hidden.
What does a High confidence label actually mean?
That the finding is consistent across a decent sample. Consistency is a weighted agreement score across mention, competitors, citations and sentiment; it is multiplied by a sample factor that saturates at 8 runs and 3 engines. High is 75 and above. Fewer than 2 runs is always Low.
Where do the cited sources come from?
From the engine's own response, stored at the moment of the run with URL, engine, position and capture time. Engines that return no sources are recorded as returning none.
Can I compare AI visibility to my Google rankings?
No, and you should not try. There is no position one in an AI answer and no single answer everyone sees. They are different measurements of different systems.
What is the difference between observed and inferred traffic?
Observed traffic arrived with an identifiable assistant referrer. Inferred traffic is a correlation between AI exposure and a traffic movement. We keep them in separate columns and never add them together.
- “Refuse to claim” is not modesty. Every non-guarantee in chapter 06 is a claim a competitor in this category makes routinely. We would rather lose the sentence than sell you something we cannot control.
Judge the numbers for yourself.
Everything above is checkable inside a workspace. If the product disagrees with this page, that is a bug and we want to hear about it.