January–July 2026 evidence review

Six months of watching how AI systems interpret public websites.

SemanticRisk has now accumulated enough repeated observations to separate ordinary website change from changes in extracted meaning. This review summarizes what the evidence supports, what remains uncertain, and how the measurement process is being refined.

Dataset reviewed

Substantial longitudinal coverage

11,461scan results
16,494model runs
49,974claim observations
14,312scan comparisons

Coverage spans January 1 to July 28, 2026, with 1,035 domains present in the wider dataset. Longitudinal claim-drift comparisons currently cover a smaller actively monitored population and are described separately below.

Finding 1

Content change and interpretation change are different signals.

7,976ordinary content-change comparisons
3.84%classified as materially interpretive

Websites change frequently, but most changes do not materially alter the resulting AI interpretation. Monitoring only page edits would therefore produce substantial noise. The more useful signal is whether the claims, stance, confidence or evidence extracted from the page changed in a meaningful way.

Practical implication: organizations need semantic monitoring, not merely content-change alerts.
Finding 2

An unchanged website does not guarantee an unchanged AI interpretation.

798extraction-drift events
100%had identical normalized content hashes
87.72%were classified as material

SemanticRisk fingerprints the normalized content used for comparison. Every recorded extraction-drift event occurred while that fingerprint remained identical. This makes it possible to distinguish model-output movement from ordinary page edits.

The evidence supports a careful conclusion: repeated AI extraction can produce materially different finding sets from unchanged normalized content. It does not yet prove that every added or removed finding represents a genuinely different meaning; some variation may involve paraphrasing, sentence segmentation, evidence selection or claim splitting and merging.

Finding 3

Interpretation instability is not confined to a handful of websites.

70 of 87eligible longitudinal domains experienced extraction drift
80.46%of that monitored comparison population

Some domains were substantially more stable than others. In the observed sample, extraction-drift rates ranged from the low single digits for several heavily monitored sites to roughly 20–25% for a number of less stable domains.

This figure must not be read as 80.46% of all 1,035 domains. It applies only to the 87 domains with eligible longitudinal claim-drift comparisons. Monitoring duration and comparison counts also vary by domain.

Finding 4

Four major models agree often—but not consistently.

3,903jobs with all four models
72.92%average agreement
41.43%or lower for the bottom tenth

GPT, Claude, Gemini and Grok were evaluated together across approximately 3,900 jobs. Median agreement was 73.87%, while the highest-performing tenth reached complete agreement.

This is not evidence that one model is universally correct. It shows that interpretation is model-dependent enough to measure and that a single-model view can hide meaningful disagreement.

Finding 5

AI systems repeatedly surface both visible language and machine-readable page signals.

Visible business evidence

Repeated findings include direct descriptions of what a company is, what it sells, who it serves and what outcomes it claims. Concrete product, audience and capability language is easier to trace than broad aspirational language.

Page and metadata evidence

Findings also repeatedly reference title tags, meta descriptions, canonical URLs, Open Graph fields, robots directives, language declarations, navigation labels, accessibility signals and structured technical observations.

Observed pattern: the clearest company descriptions commonly follow a simple structure—company or product, category or capability, intended audience, and claimed outcome.
Finding 6

Raw claim churn can overstate true semantic change.

Across extraction-drift events, 4,517 claims were added and 4,342 removed. In mixed events, 14,496 were added and 14,434 removed. The near balance between additions and removals suggests that models often reorganize or reformulate findings rather than simply discovering more or less information.

Examples in the stored claim set include near-duplicate meta-description findings, equivalent canonical-URL observations with minor punctuation differences, and business descriptions that preserve the same underlying proposition while changing sentence structure.

Measurement refinement: SemanticRisk is separating true meaning drift from representational drift such as paraphrasing, claim splitting, claim merging and evidence changes.
What “claims” currently contain

Business meaning and technical findings need clearer separation.

The claim layer currently includes business descriptions, HTML and metadata observations, accessibility findings, and technical telemetry or scoring dimensions. These are useful signals, but they should not be treated as identical semantic units.

Business identity

Products, services, sectors, audiences, capabilities and claimed outcomes.

Page interpretation signals

Titles, descriptions, structured data, navigation, canonical URLs and accessibility.

Technical observations

Crawl behavior, telemetry, privacy, logging and related scoring evidence.

Fine-tuning the method

What SemanticRisk is improving next

Priority 1

Backfill multi-model claim history

Claims exist in structured outputs for roughly 4,000 runs from each provider, but historical claim-observation persistence is presently strongest for GPT. Backfilling Claude, Gemini and Grok will support valid model-by-model drift comparisons.

Priority 2

Semantic claim matching

Added and removed claims will be distinguished from paraphrases, split claims, merged claims, stance changes, confidence changes and evidence changes.

Priority 3

Claim-family separation

Business identity, product/service, audience, outcome, metadata, accessibility and technical findings will be reported as separate risk dimensions.

Priority 4

Richer structural evidence

Heading hierarchy, structured data, navigation text, visible main content and normalized model input will allow stronger testing of how HTML structure and wording relate to interpretation stability.

Method and limitations

What these results do—and do not—show

SemanticRisk measures outputs generated from public website evidence at specific points in time. Content hashes validate whether normalized input changed. Model agreement measures consistency across evaluated claim signals. Materiality is assigned by the current drift rules.

The results do not prove that a particular HTML element directly caused a model response, that one provider is universally more accurate, or that every output difference represents a different real-world belief. Technical failures are excluded from substantive drift conclusions, and model or prompt versions must be controlled when making longitudinal comparisons.

The current evidence is best described as an observed association between website evidence and model interpretation—not a universal specification of what every downstream AI system will do.

See current company-level evidence.

Explore the public benchmark, compare companies, and review the latest interpretation records.

Open live benchmark