The science behind the score
How the Get Found Score works.
Get Found is your credit score for AI visibility. This page explains how the score is measured, what it can tell you, and what it cannot. We publish the method because a measurement is only worth trusting if you can see how it was built.
We are building and earning the standard for AI visibility the way FICO earned its seat: a published, transparent method, proven over time. This page is that method, described from the system running today.
What the score measures
One number, one question.
How consistently does AI recommend this business for the questions its real customers ask?
The Get Found Score is one number from 0 to 100. It is a recommendation rate. When a buyer asks an AI engine who to hire, does your company get named? We run the questions your customers actually ask, count how often you are recommended among the answers that name a real provider, and report the rate.
The score reports one thing: the share of real opportunities where AI recommended you. We do not average several unrelated numbers together to produce it.
The calculation
A valid opportunity is a question where the engine returned a real, provider-naming answer. A valid recommendation is an answer that named your business. Answers that name no provider, or that fail to return, are handled separately, so a technical failure never counts against you.
How questions are selected
A question set that flatters you is not a measurement.
The score is only as honest as the questions behind it. So every scored question has to earn its place.
Provenance
Where the question came from: a real search query, a query the engines expand into, a customer interview, a competitor citation, or trade research. “We thought it sounded good” is not a source.
Buyer intent
The question represents a real moment where a customer is deciding who to hire, not brand trivia and not a general information lookup.
Class
Local buyer-intent questions are what the headline counts. Informational and branded questions are recorded but kept out of the headline, because winning “what is an HVAC system” is not the same as being recommended to a buyer.
The full question-set governance policy, its inclusion rules, exclusion rules, weighting, review cadence, and how we decide two versions are comparable, is maintained as a versioned document. When the set changes, the version changes with it.
Planned: publishing that governance document publicly, so anyone can read why the selected questions represent a market. Provenance and class are necessary but not sufficient, and we hold ourselves to explaining the “why.”
How models are tested
Hold everything constant. Let only the answer move.
The same question is submitted to more than one AI engine during a controlled run. Today that is ChatGPT, Perplexity, and Gemini.
Within a run we hold constant the parts that would otherwise move the answer around: the question wording, the engines, the geographic context, the identity rules, and the question-set version. Only the AI’s answer is allowed to vary. That is what lets a change in the score reflect a change in AI behavior rather than a change in how we asked.
Each answer is classified by deterministic identity rules that record which company was named, which competitors were named, whether a citation was returned, and whether the answer failed or was too uncertain to score.
Planned: expanding and versioning the engine roster beyond today’s ChatGPT, Perplexity, and Gemini, and recording each model’s version when the engine exposes it, so a model update on their side becomes visible on ours.
How often measurement happens
One question on one day is not a conclusion.
Measurement is not a one-time snapshot. A single answer can swing on wording, timing, or which version of a model happened to respond. So the same questions are collected repeatedly over a series of runs. Repetition shows whether a result holds up, and it narrows the uncertainty around the observed rate as evidence accumulates.
Every score carries the date it was measured and how fresh the underlying data is. A number from two weeks ago is labeled as two weeks old, never shown as today’s.
Step 01
Approve the buyer questions
Each scored question asks an AI engine to recommend a provider for a real buyer need. Every question earns its place with a source, a buyer intent, and a class.
Step 02
Ask the assigned engines
The same question goes to more than one AI engine during a controlled run. Today that is ChatGPT, Perplexity, and Gemini.
Step 03
Classify the answer
Deterministic identity rules record which company was named, which competitors were named, whether a citation was returned, and whether the answer failed or was too uncertain to score.
Step 04
Repeat and compare
The same questions are collected again over a series of runs. Repetition shows whether a result holds up and narrows the uncertainty around the observed rate.
Planned: collection runs on a daily schedule today. Moving it to always-on infrastructure, so the cadence holds independent of any single machine, is planned.
How observed variation is shown
A point estimate can look more certain than the evidence allows.
So beside the score we show a range: the observed variation around the measured rate. A wide range means the system has not yet collected enough observations to be confident. As runs accumulate and the results stay consistent, the range usually narrows. Early on, when the sample is small, the range is honestly wide. That is the system telling the truth about how much it knows, not a flaw.
The range describes how much the observed rate moves under repeated sampling. It does not, by itself, prove that a change you made caused the score to move. Proving cause needs a controlled before-and-after comparison. A narrower range by itself does not establish it.
A note on the statistics, in plain terms
The range is computed from the observed recommendation rate and the number of valid observations. Because repeated answers can come from the same engines and questions, those observations are not fully independent, so we present the range as an estimate of variation, not as a formal statistical guarantee. We will only call a number a confidence interval once the calculation accounts for that structure in the data.
Evidence rules
A failed call never counts against you.
The system separates a business being absent from the answer from a problem in the collection itself. This keeps technical errors from silently lowering your score. The denominator is the count of valid opportunities, not every attempt. Failed and uncertain results are visible in the evidence, not buried in the number.
Valid recommendation
The answer named your company, or named another real provider. It counts toward the rate.
Uncertain identity
A possible name match without enough corroboration. It is excluded from the rate and flagged for review, never guessed.
Failed or unavailable call
A timeout, an engine error, or a missing answer. It is recorded as missing evidence, never as a mark against you.
Where SEO and content fit
The score and the improvement work have different jobs.
We keep them separate on purpose. Search Console and Google Analytics inform the action plan and measure what happens after a recommendation. They do not receive an invented weight inside the recommendation score. The measure of AI visibility stays clean of the things you do to improve it.
Inputs
Your published content, your search presence, your citations and authority sources. These are the levers you pull to improve.
The primary outcome
The AI recommendation rate. Does the company appear when a buyer asks who to hire?
Traffic
Search Console clicks, organic sessions, and AI-referral visits. These measure downstream discovery.
Business result
Calls, forms, booked meetings, and revenue, when those feeds are connected.
What the score can and cannot say
As clear about its limits as it is about its result.
The score can tell you
- How often your company was recommended among valid, provider-naming answers.
- Which buyer questions produced a recommendation, a competitor, or nothing.
- Which engines named you, named competitors, or failed to return evidence.
- How uncertain the current estimate is, given the sample collected.
- Whether the result changed after approved work and repeated measurement.
The score cannot, by itself, tell you
- Why an AI engine picked one provider over another.
- Whether one specific page change caused a score move, without a controlled comparison.
- How much revenue a recommendation produced, without connected conversion data.
- Whether a result holds for questions, markets, or engines outside the approved set.
- What permanent weight content, search, or authority should carry, without outcome calibration.
How business outcomes are reported
Outcomes sit beside the score, never inside it.
Calls, forms, appointments, qualified leads, and revenue are the point of visibility, so we report them. We do not fold outcomes into the visibility formula, and we do not claim that a recommendation caused a sale until a study design supports it. Keeping outcomes next to the score, rather than mixed into it, is what lets the visibility measure stay independent of the services we sell to improve it.
The conflict, named, and how we handle it
Get Found both measures your visibility and sells the work that improves it. So we freeze the scoring definition for each measurement period, we keep the raw observations and the historical question sets, we never rewrite a past score without labeling the change, and we keep the score formula independent of how much service you buy. The measurement has to be trustworthy first, or the improvement is not worth selling.
How methodology changes are disclosed
A change in the instrument never hides as a change in your performance.
The instrument itself is versioned. Every score is stamped with its formula version and its question-set version. When we change how the score is calculated, or which questions feed it, the version changes and the change is published. When a change means two scores can no longer be compared fairly, we show a visible break instead of pretending the new number continues the old line.
The operating promise
The score is worth what it can prove, and no more.
Get Found reports the score, the evidence behind it, the variation around it, the observations excluded from it, and the next action the data supports. Historical question sets and formulas stay versioned, so a change in how we measure can never hide as a change in how you are doing. We publish this method because that is how a measurement earns trust.