AI Visibility Index
Every score in the index is derived from real API calls. No estimates, no scraped proxies. Here’s exactly what we did.
Each brand is evaluated against 25 prompt templates drawn from five buyer-intent categories. Templates are filled with brand-specific variables (category, segment, top competitors, key use cases, integration partners) before they are sent to each LLM.
Category Discovery (5 prompts)
Generic "what's best for X" queries. These are the prompts most buyers fire first. Example: "What is the best CRM for B2B SaaS?"
Comparison (5 prompts)
Head-to-head comparisons against the brand's two closest competitors. Example: "Pipedrive vs HubSpot: which is better?"
Alternatives (5 prompts)
Buyer-in-pain queries searching for a switch. Example: "Cheaper alternatives to Salesforce for SMBs."
Use Case (5 prompts)
Job-to-be-done questions a practitioner might ask. Example: "Best way to track trial-to-paid conversion for a SaaS team."
Integration (5 prompts)
Tool-stack queries about connectivity. Example: "Does Attio integrate with Zapier?"
We fire each prompt against three platforms via their official APIs. Responses are captured verbatim and scored independently. No human editing.
ChatGPT (GPT-4o)
Non-browsing chat mode. Scores reflect the model's training knowledge, not live web results.
Claude (Haiku 4.5)
Anthropic's fast-tier model. Same non-browsing constraint as ChatGPT.
Perplexity (Sonar Pro)
Real-time web-augmented answers. Scores here are closer to a live organic-search signal.
Google AI Overviews: not in this run
Our data model includes a Google AIO slot, but live AI Overview retrieval via SerpAPI was not activated for this index run. Google AIO scores appear as 0 across all entries. We plan to add it in a future run once we validate retrieval accuracy.
AI Visibility Score (AVS) is a 0–100 composite that measures how prominently a brand appears across all LLM responses. It is computed in three steps.
For every (prompt × LLM) pair we parse the response and award points on three signals:
| Signal | Condition | Points |
|---|---|---|
| Rank | Listed 1st | +6 |
| Rank | Listed 2nd | +5 |
| Rank | Listed 3rd | +4 |
| Rank | Listed 4th | +3 |
| Rank | Listed 5th | +2 |
| Rank | Listed 6th or lower | +1 |
| Rank | Mentioned, not in a list ("unranked") | +2 |
| Rank | Primary answer to direct-question prompt ("N/A") | +3 |
| Rank | Not mentioned | 0 |
| Sentiment | Positive framing near brand | +2 |
| Sentiment | Neutral | 0 |
| Sentiment | Negative framing near brand | −2 |
| Citation | Brand URL cited in response | +1 |
Raw score is clamped to 0–10. Maximum possible per prompt: 9 (1st-place + positive + URL cited).
The 25 raw scores for a given LLM are averaged and multiplied by 10, yielding a 0–100 AVS for that platform.
The composite score on the leaderboard is the simple mean of all per-LLM AVS values for which at least one prompt was scored. Currently that is three platforms (ChatGPT, Claude, Perplexity).
Every brand in this index is marked Verified. That means all 75 scores (25 prompts × 3 LLMs) were obtained through live API calls made on the run date shown in the leaderboard. No score was interpolated, extrapolated, or carried forward from a prior run. An Estimated label would appear if a brand partially failed (e.g., a provider rate-limit caused some prompts to be skipped) and we back-filled with a prior result. That did not happen in this run.
The full leaderboard shows every brand’s composite AVS, per-LLM breakdown, and top gap prompts. All data comes directly from the scoring run.