How to Create an AI Visibility Baseline for a New Domain
A step-by-step framework for creating an AI visibility baseline for a new domain, including scope, exact surfaces, evidence, uncertainty, and the first report.
A new AI visibility program does not need a perfect historical dataset. It needs a baseline that is explicit enough to repeat. The first report should say what was measured, under which conditions, how much was observed, and what the result cannot prove yet.
Define the baseline boundary
Write down the domain, brand entities, products, markets, languages, and question clusters included in the first measurement. Decide whether the baseline is about awareness, recommendations, citations, or all three. A baseline that tries to answer every question can become too vague to guide action.
Name what is outside the boundary. If the program does not include local results, transactional questions, or a particular provider, say so. A clear limitation is more useful than an implied promise of complete market coverage.
A useful scope statement also names the decision the baseline should support. It may be a launch review, a market-entry question, a content investment, or a recurring monitoring program. Connecting the baseline to a decision keeps the first report from becoming an undirected audit with no owner and gives the team a reason to revisit the same questions.
Freeze the exact execution surface
Record the provider, model, model configuration hash, grounding mode, locale, location, probe version, campaign identity, and timestamp. These fields turn an answer into an observation with an execution context. Without them, two reports may appear comparable while measuring different surfaces.
Keep the same surface for the first comparison window. When the model or grounding mode changes, start a new labeled window or annotate the transition. Do not describe a surface change as a content win or loss without evidence that the content was the cause.
Freeze the fields in the report template as well as the provider settings. If a later run adds a new grounding mode or changes the location, the difference should be visible in the comparison. This matters when teams share dashboards across markets, because a regional result can otherwise be mistaken for a global trend.
- Question cluster and probe version.
- Provider, model, configuration, and grounding mode.
- Locale, location, timestamp, and campaign identity.
- Formula version, sample size, and comparison window.
Retain evidence without overstating it
The first report should point from an aggregate to the normalized observations that support it. Preserve the answer excerpt, detected mentions, citations, source ownership, and a hash reference for the artifact boundary your system retains. That gives a reviewer a path to inspect the result without claiming that every raw provider response is independently verified.
Separate missing evidence from negative evidence. A citation that was not retained is not the same as a response that contained no citation. Use explicit states for pending, unavailable, invalid, and complete observations so the baseline does not quietly convert collection gaps into zeros.
A reviewer should also be able to distinguish an observed absence from an unavailable record. Represent that distinction with explicit evidence states and a short reason. It protects the baseline from turning a collection failure into a false negative and gives operators a clear path for remediation before the next comparison window.
Publish the first report as a starting point
A baseline report should show the metric, the numerator and denominator, the uncertainty or confidence, and the top observations that explain the result. Add a short list of recommended next pages or questions, but keep the recommendations tied to evidence rather than to a generic content checklist.
The baseline earns its value through the second measurement. Schedule the remeasurement before making the first change, preserve the original snapshot, and compare like with like. The report is not a verdict on the domain; it is the reference point for a learning loop.
Use the first report to establish a recurring cadence rather than a one-time score. A monthly or campaign-based window may be enough at first; the right interval depends on answer variability and how quickly the team can make changes. State the cadence beside the metric so stakeholders understand what the result can and cannot tell them.
Before publishing, ask a second reviewer to challenge the boundary and the comparison assumptions. This lightweight review often catches a mixed market, an ambiguous brand entity, or a question cluster that cannot be answered consistently. Fixing those issues at the baseline stage is cheaper than explaining them after the first trend appears.