Black-Box Evaluation of Observable AI Recommendation Behaviours
An overview of ASTONTO’s methodology for measuring how large language models and AI-enabled search systems represent, compare and recommend organisations through observable public outputs.

Abstract
ASTONTO studies the observable behaviour of commercial large language models and AI-enabled search systems.
ASTONTO does not have access to the internal model weights, ranking systems, training corpora or real-time decision processes used by commercial AI platforms. Evaluation must therefore be based on controlled black-box testing: submitting documented prompts, preserving the resulting outputs and analysing measurable patterns across platforms, repeated runs, locations and languages.
This methodology overview explains how the PULSE Method v1.0 evaluates entity appearance, recommendation prominence, recommendation strength, sentiment, source behaviour, competitor performance and consistency.
The method measures observed behaviour within a defined testing period and prompt set. It does not claim to explain every internal cause of an AI-generated recommendation.
Research question
When commercial AI platforms are presented with controlled buyer-evaluation queries:
- which corporate entities appear or are omitted;
- how prominently those entities are positioned;
- how strongly they are recommended;
- what sentiment is expressed;
- which sources, if any, are cited or surfaced;
- how the results differ by platform, location and language; and
- how consistently the results recur across repeated tests?
The platforms currently covered by the core PULSE Method are:
- ChatGPT;
- Perplexity;
- Gemini; and
- Google AI Overviews.
Methodology scope
PULSE Method v1.0 examines seven broad areas of observable AI recommendation behaviour.
1. Prompt interpretation and buyer intent
ASTONTO evaluates how AI platforms respond to prompts with different commercial purposes, including high-intent recommendations, comparisons, service-selection questions, problem-solving queries and informational questions.
Prompt categories and commercial weights are agreed before testing. They are not changed retrospectively in response to the results.
2. Entity appearance and representation
The method records whether an organisation appears in the response and how it is represented.
An organisation may be:
- a primary recommendation;
- part of a positive shortlist;
- included in a wider list;
- presented as an alternative;
- mentioned only peripherally; or
- omitted entirely.
A mention is not automatically treated as a recommendation.
3. Position and prominence
PULSE records where and how prominently an organisation appears.
For clearly ranked responses, the method records its recommendation position. For unranked responses, it evaluates whether the organisation is the lead recommendation, part of the first recommendation group, included in the main answer, or mentioned only as supporting context.
First position is not automatically treated as the strongest recommendation. Position and endorsement are measured separately.
4. Recommendation strength and sentiment
Recommendation strength is classified using the following PULSE categories:
- Strong recommendation;
- Positive shortlist;
- Listed;
- Alternative;
- Weak mention;
- Not recommended;
- Negative recommendation.
Sentiment is assessed separately as:
- Strongly positive;
- Positive;
- Neutral;
- Mixed;
- Negative;
- Strongly negative.
This distinction prevents a prominent but critical or negative reference from being counted as valuable visibility.
5. Source and citation behaviour
Where a platform displays sources, links or citations, ASTONTO records the domains and pages surfaced in support of the answer.
Sources may include:
- the organisation’s own website;
- competitor websites;
- business directories;
- review platforms;
- news organisations;
- trade publications;
- professional bodies;
- government or regulatory sources; or
- other third-party publications.
The absence of a visible source is also recorded. A displayed citation does not by itself prove that it was the sole cause of the generated recommendation.
6. Comparative, geographic and language behaviour
The focal organisation is compared with four selected competitors using identical:
- prompts;
- platforms;
- locations;
- languages;
- testing dates;
- numbers of runs;
- prompt weights; and
- methodology versions.
Separate results are produced where country, city, region or language may materially affect the buyer journey.
PULSE Share of Voice applies only to the documented five-company competitor set and should not be interpreted as total market share.
7. Repeated testing and temporal consistency
AI-generated responses vary between runs.
Important prompts are therefore tested at least three times per platform. Five runs are recommended for high-value or strategically important prompts.
Each run is preserved and scored separately before the results are averaged. This means that an organisation appearing strongly in only one run receives a lower result than an organisation appearing consistently across repeated tests.
Longitudinal studies repeat the controlled test set at documented intervals to identify material changes in recommendation behaviour.
Scoring framework
For every organisation, prompt, platform, location, language and run, PULSE calculates:
Prompt Result Score = Position Factor × Recommendation Factor × Sentiment Factor
The Prompt Result Score ranges from negative visibility to the strongest possible positive recommendation.
Repeated runs are averaged before prompt and platform results are calculated.
PULSE reports separate Platform Scores for:
- ChatGPT;
- Perplexity;
- Gemini; and
- Google AI Overviews.
The standard PULSE Benchmark Score gives each platform an equal 25% weight. A separate PULSE Market Score may use client-specific weights where those weights are justified and agreed before testing.
Evidence and audit trail
Every material classification must be supported by preserved evidence.
The research record should include:
- exact prompt;
- prompt category and weight;
- platform and interface where identifiable;
- date and time;
- location and language;
- run number;
- full response;
- organisation position;
- recommendation classification;
- sentiment classification;
- resulting score;
- competitors mentioned;
- sources shown;
- evidence supporting the classification;
- confidence;
- manual-review status; and
- methodology version.
These fields are part of the approved PULSE data structure.
Reliability
Reliability is assigned to a specific study or PULSE Score, not to the methodology overview itself.
Under PULSE Method v1.0:
- High Reliability: at least 100 prompts, three or more runs and all target platforms tested;
- Medium Reliability: 50–99 prompts, at least two runs and most platforms tested;
- Indicative: fewer than 50 prompts, one run or major platform gaps.
An Indicative result must not be presented as a definitive assessment.
Observational limits and ethics
ASTONTO evaluates outputs accessible through documented commercial AI interfaces during the stated testing period.
ASTONTO makes no claim of access to:
- internal model weights;
- proprietary ranking algorithms;
- private training corpora;
- undisclosed system prompts;
- unreleased model changes; or
- the complete causal process behind an individual response.
Results apply only to the documented:
- prompt set;
- platforms;
- testing period;
- locations;
- languages;
- competitor group;
- number of runs; and
- methodology version.
AI-generated answers can change. A PULSE Score reflects observed recommendation performance at a particular time and does not guarantee future recommendations, citations or visibility.
Citation Block
ASTONTO Research. (2026-08-01). "Black-Box Evaluation of Observable AI Recommendation Behaviours". ASTONTO Independent AI Research. https://astonto.com/research/black-box-ai-evaluation