The methods column

How we measure, and what we[1] refuse to claim

AI visibility is easy to fake and hard to measure. This page is the arithmetic, the sampling rules, and the explicit list of things CiteTrail does not guarantee. So you can judge our numbers instead of trusting them.

: the central claim

A model never picks your score

Every number in CiteTrail is computed by deterministic code from responses we stored. Models do exactly three jobs: run the prompt, classify sentiment and extract entities from what came back, and write prose summaries of results already calculated. If a figure moved, a rule moved it, and the rule is on this page.

01: How prompts are selected

How prompts are selected

Three ways in: write them yourself, import a list, or start from prompts suggested for your brand and category.

Every prompt is then classified across 13 intent categories and 5 funnel stages and scored out of 100 for buyer value. High is 70 and above, medium is 45 and above. That score decides what gets run most often.

Region and language are separate measurements. A prompt belongs to one market, and results from different markets are never averaged into a single figure. Where a scoring signal has no connected data behind it, the factor falls back to neutral. We do not fill the gap with an estimate and then score it.

What the scoring rewards

Commercial intent, not comfort. A prompt is never weighted up because you perform well on it. The highest-value questions in most workspaces are the ones being lost.

What counts as one measurement

One prompt, on one engine, in one region, at one point in time. Everything above that (trends, share of answer, confidence) is built by aggregating those units. Nothing is aggregated across units that were not comparable in the first place.

02: Why one answer proves nothing

One answer is an anecdote

Language models are non-deterministic. Ask the same question twice and you can get two different answers naming two different sets of brands. This is not a flaw in the measurement. It is the thing being measured.

The same question, twice
Ask a model the same thing two minutes apart and you can get two answers naming two different sets of brands. That is how these systems work, not a fault in them.
Prompts run on a schedule, not once
Every tracked prompt runs repeatedly, across engines, over a window. A pattern is something that survives repetition. A single result is a story.
Providers change things without telling you
Models get swapped, system prompts get edited, grounding gets turned on. None of it is announced. Continuous sampling is the only way to see it happen.

Consequence: a finding based on fewer than two runs is labelled Low confidence. Always. No combination of agreement between those runs can override it.

03: How confidence is derived

How confidence is derived

Confidence is consistency multiplied by how much we sampled. Both halves are published.

Step one: how consistently did it repeat

Each sub-score is the share of runs that agree with the majority outcome, from 0 to 100. They are blended by how much each one matters to a business decision.

consistency = mention agreement     × 0.40
            + competitor agreement  × 0.25
            + citation agreement    × 0.20
            + sentiment agreement   × 0.15

Step two: how much evidence backs it

Consistency alone would let one lucky run read as certainty. So it is gated by a sample-size factor that can only ever range from 0.3 to 1.0.

confidence = consistency × sample_factor(runs, engines)
Runs
Contribute up to 0.7 of the factor, saturating at 8 runs.
Engines
Contribute up to 0.3 of the factor, saturating at 3 engines.
Floor
The factor never drops below 0.3, so a single consistent run keeps some confidence, but can never reach full confidence.

The caps are deliberate. Running a prompt 40 times does not buy more confidence than running it 8 times, and an engine beyond the 3 already tracked does not buy more. Beyond that point you are measuring the same thing again.

Step three: what you see

The score becomes a label, and the label always travels with the sample that produced it.

  • High75 and above
  • Medium50 and above
  • Lowbelow 50

One override outranks the arithmetic: fewer than two runs is always low confidence, whatever the score says.

04: How citations are captured

How citations are captured

Sources are stored at run time, not reconstructed later. Each one keeps its URL, the engine that returned it, its position in the answer, and the time it was captured.

Domains are classified as editorial, review, community or owned, which is what makes a source-mix chart mean anything.

Engines that return no sources are recorded as returning none. That row stays empty. Guessing which pages a model probably read would fill the chart and break the page.

Which engines return sources

What a stored run keeps

The response
The answer text the assistant actually returned, kept so a finding can be checked rather than believed.
The engine and the moment
Which surface produced it, and when it was captured.
The cited sources
Every cited URL with its position, classified by domain and source type.
What changed
Mention flips, position deltas, and sources that appeared or dropped since the previous run.

We store the runs and responses we measured so findings stay auditable. See security for exactly what we hold and how it's handled.

05: Observed versus inferred

Four words, four columns, never one total

Observed

The referrer directly identifies an assistant. We saw the session arrive. This is the only tier we will ever describe as attribution.

Estimated

Modelled influence. A visit whose shape matches assistant-driven behaviour, without the referrer to prove it.

Inferred

Correlation between AI exposure and traffic movement. Suggestive. Not proof, and labelled as such.

Exposure

Your brand appeared in an answer. A fact about the answer, not about traffic. That is precisely why it's so often quietly counted as traffic elsewhere.

These four are never added together. A single blended “AI traffic” number would be the most quotable figure on this page and the least defensible.

06: What CiteTrail does not guarantee

What CiteTrail does not guarantee

This chapter is not a disclaimer we hid at the bottom. If a competitor promises you any of the following, they are describing something outside their control.

No guaranteed ranking, mention or traffic

Providers change models and grounding without notice, and outputs vary run to run. Anyone promising otherwise is selling something they cannot control.

Inferred influence is not revenue

It is a correlation with a label on it. Do not put it in a board deck as pipeline.

Provider APIs are not consumer apps

What we measure is the model behind the assistant. Your buyer's chat window may say something different, and we would rather tell you that than let you assume otherwise.

A visibility score is not a search ranking

There is no position one in an AI answer, and the two numbers are not comparable in either direction.

An action is a recommendation, not a causal claim

It is ranked by expected value. It becomes evidence only after a re-scan measures the change.

We do not report a number we did not measure

Empty stays empty. If a workspace has no data, it shows no data.

What we do promise is narrower and more useful: every number traces to a stored response, every finding shows the sample behind it, and anything we inferred says so.

07: Questions

Methodology, answered plainly

Does an AI model calculate my visibility score?

No. Scoring is deterministic code over stored responses. Models run prompts, classify sentiment, extract entities and write summaries. They never pick a number.

Why do my results change between runs?

Because language models are non-deterministic and providers change models and grounding without notice. That variance is the thing being measured, not noise being hidden.

What does a High confidence label actually mean?

That the finding is consistent across a decent sample. Consistency is a weighted agreement score across mention, competitors, citations and sentiment; it is multiplied by a sample factor that saturates at 8 runs and 3 engines. High is 75 and above. Fewer than 2 runs is always Low.

Where do the cited sources come from?

From the engine's own response, stored at the moment of the run with URL, engine, position and capture time. Engines that return no sources are recorded as returning none.

Can I compare AI visibility to my Google rankings?

No, and you should not try. There is no position one in an AI answer and no single answer everyone sees. They are different measurements of different systems.

What is the difference between observed and inferred traffic?

Observed traffic arrived with an identifiable assistant referrer. Inferred traffic is a correlation between AI exposure and a traffic movement. We keep them in separate columns and never add them together.

  1. “Refuse to claim” is not modesty. Every non-guarantee in chapter 06 is a claim a competitor in this category makes routinely. We would rather lose the sentence than sell you something we cannot control.

Judge the numbers for yourself.

Everything above is checkable inside a workspace. If the product disagrees with this page, that is a bug and we want to hear about it.