AI Visibility Index

How we scored this

Every score in the index is derived from real API calls. No estimates, no scraped proxies. Here’s exactly what we did.

100 brands25 prompts per brand3 LLMs7,500+ API calls

1. The 25 canonical prompts

Each brand is evaluated against 25 prompt templates drawn from five buyer-intent categories. Templates are filled with brand-specific variables (category, segment, top competitors, key use cases, integration partners) before they are sent to each LLM.

CD

Category Discovery (5 prompts)

Generic "what's best for X" queries. These are the prompts most buyers fire first. Example: "What is the best CRM for B2B SaaS?"

CM

Comparison (5 prompts)

Head-to-head comparisons against the brand's two closest competitors. Example: "Pipedrive vs HubSpot: which is better?"

AL

Alternatives (5 prompts)

Buyer-in-pain queries searching for a switch. Example: "Cheaper alternatives to Salesforce for SMBs."

UC

Use Case (5 prompts)

Job-to-be-done questions a practitioner might ask. Example: "Best way to track trial-to-paid conversion for a SaaS team."

IN

Integration (5 prompts)

Tool-stack queries about connectivity. Example: "Does Attio integrate with Zapier?"

2. The 3 LLMs scored

We fire each prompt against three platforms via their official APIs. Responses are captured verbatim and scored independently. No human editing.

Ch

ChatGPT (GPT-4o)

Non-browsing chat mode. Scores reflect the model's training knowledge, not live web results.

Cl

Claude (Haiku 4.5)

Anthropic's fast-tier model. Same non-browsing constraint as ChatGPT.

Pe

Perplexity (Sonar Pro)

Real-time web-augmented answers. Scores here are closer to a live organic-search signal.

Google AI Overviews: not in this run

Our data model includes a Google AIO slot, but live AI Overview retrieval via SerpAPI was not activated for this index run. Google AIO scores appear as 0 across all entries. We plan to add it in a future run once we validate retrieval accuracy.

3. How we compute the AVS

AI Visibility Score (AVS) is a 0–100 composite that measures how prominently a brand appears across all LLM responses. It is computed in three steps.

Step 1: Per-prompt raw score (0–10)

For every (prompt × LLM) pair we parse the response and award points on three signals:

SignalConditionPoints
RankListed 1st+6
RankListed 2nd+5
RankListed 3rd+4
RankListed 4th+3
RankListed 5th+2
RankListed 6th or lower+1
RankMentioned, not in a list ("unranked")+2
RankPrimary answer to direct-question prompt ("N/A")+3
RankNot mentioned0
SentimentPositive framing near brand+2
SentimentNeutral0
SentimentNegative framing near brand−2
CitationBrand URL cited in response+1

Raw score is clamped to 0–10. Maximum possible per prompt: 9 (1st-place + positive + URL cited).

Step 2: Per-LLM AVS

The 25 raw scores for a given LLM are averaged and multiplied by 10, yielding a 0–100 AVS for that platform.

AVSLLM = (Σ prompt scores / 25) × 10

Step 3: Composite AVS

The composite score on the leaderboard is the simple mean of all per-LLM AVS values for which at least one prompt was scored. Currently that is three platforms (ChatGPT, Claude, Perplexity).

AVSbrand = mean(AVSChatGPT, AVSClaude, AVSPerplexity)

4. What “Verified” means

Every brand in this index is marked Verified. That means all 75 scores (25 prompts × 3 LLMs) were obtained through live API calls made on the run date shown in the leaderboard. No score was interpolated, extrapolated, or carried forward from a prior run. An Estimated label would appear if a brand partially failed (e.g., a provider rate-limit caused some prompts to be skipped) and we back-filled with a prior result. That did not happen in this run.

5. Sample-size and freshness caveats

  • ·25 prompts is a proxy, not a census. Real buyer behaviour spans thousands of query variants. Our prompt set is designed to be representative across five intent types, but niche queries specific to a brand’s micro-category may not be covered.
  • ·LLM outputs are non-deterministic. Re-running the same prompt against the same model on the same day can produce a different ranked list. Scores are a snapshot, not a guaranteed steady-state measurement.
  • ·Scores decay. LLM training data and real-time web indices change continuously. The numbers shown reflect the run date; a brand’s score may differ if you re-run today.
  • ·Sentiment is LLM-assisted. We classify the sentiment of each brand mention using a lightweight model (GPT-4o-mini). Keyword-heuristic fallback is used if the classification call fails. Edge cases may be mis-labelled.

See the raw run data

The full leaderboard shows every brand’s composite AVS, per-LLM breakdown, and top gap prompts. All data comes directly from the scoring run.

View AI Visibility Index