January–July 2026 evidence review
Six months of watching how AI systems interpret public websites.
SemanticRisk has now accumulated enough repeated observations to separate ordinary website change from changes in extracted meaning. This review summarizes what the evidence supports, what remains uncertain, and how the measurement process is being refined.
Dataset reviewed
Substantial longitudinal coverage
11,461scan results
16,494model runs
49,974claim observations
14,312scan comparisons
Coverage spans January 1 to July 28, 2026, with 1,035 domains present in the wider dataset. Longitudinal claim-drift comparisons currently cover a smaller actively monitored population and are described separately below.
Finding 1
Content change and interpretation change are different signals.
7,976ordinary content-change comparisons
3.84%classified as materially interpretive
Websites change frequently, but most changes do not materially alter the resulting AI interpretation. Monitoring only page edits would therefore produce substantial noise. The more useful signal is whether the claims, stance, confidence or evidence extracted from the page changed in a meaningful way.
Practical implication: organizations need semantic monitoring, not merely content-change alerts.
Finding 2
An unchanged website does not guarantee an unchanged AI interpretation.
798extraction-drift events
100%had identical normalized content hashes
87.72%were classified as material
SemanticRisk fingerprints the normalized content used for comparison. Every recorded extraction-drift event occurred while that fingerprint remained identical. This makes it possible to distinguish model-output movement from ordinary page edits.
The evidence supports a careful conclusion: repeated AI extraction can produce materially different finding sets from unchanged normalized content. It does not yet prove that every added or removed finding represents a genuinely different meaning; some variation may involve paraphrasing, sentence segmentation, evidence selection or claim splitting and merging.
Finding 3
Interpretation instability is not confined to a handful of websites.
70 of 87eligible longitudinal domains experienced extraction drift
80.46%of that monitored comparison population
Some domains were substantially more stable than others. In the observed sample, extraction-drift rates ranged from the low single digits for several heavily monitored sites to roughly 20–25% for a number of less stable domains.
This figure must not be read as 80.46% of all 1,035 domains. It applies only to the 87 domains with eligible longitudinal claim-drift comparisons. Monitoring duration and comparison counts also vary by domain.
Finding 4
Four major models agree often—but not consistently.
3,903jobs with all four models
72.92%average agreement
41.43%or lower for the bottom tenth
GPT, Claude, Gemini and Grok were evaluated together across approximately 3,900 jobs. Median agreement was 73.87%, while the highest-performing tenth reached complete agreement.
This is not evidence that one model is universally correct. It shows that interpretation is model-dependent enough to measure and that a single-model view can hide meaningful disagreement.
Finding 5
AI systems repeatedly surface both visible language and machine-readable page signals.
Visible business evidence
Repeated findings include direct descriptions of what a company is, what it sells, who it serves and what outcomes it claims. Concrete product, audience and capability language is easier to trace than broad aspirational language.
Page and metadata evidence
Findings also repeatedly reference title tags, meta descriptions, canonical URLs, Open Graph fields, robots directives, language declarations, navigation labels, accessibility signals and structured technical observations.
Observed pattern: the clearest company descriptions commonly follow a simple structure—company or product, category or capability, intended audience, and claimed outcome.
Finding 6
Raw claim churn can overstate true semantic change.
Across extraction-drift events, 4,517 claims were added and 4,342 removed. In mixed events, 14,496 were added and 14,434 removed. The near balance between additions and removals suggests that models often reorganize or reformulate findings rather than simply discovering more or less information.
Examples in the stored claim set include near-duplicate meta-description findings, equivalent canonical-URL observations with minor punctuation differences, and business descriptions that preserve the same underlying proposition while changing sentence structure.
Measurement refinement: SemanticRisk is separating true meaning drift from representational drift such as paraphrasing, claim splitting, claim merging and evidence changes.
What “claims” currently contain
Business meaning and technical findings need clearer separation.
The claim layer currently includes business descriptions, HTML and metadata observations, accessibility findings, and technical telemetry or scoring dimensions. These are useful signals, but they should not be treated as identical semantic units.
Business identity
Products, services, sectors, audiences, capabilities and claimed outcomes.
Page interpretation signals
Titles, descriptions, structured data, navigation, canonical URLs and accessibility.
Technical observations
Crawl behavior, telemetry, privacy, logging and related scoring evidence.
Fine-tuning the method
What SemanticRisk is improving next
Priority 1Backfill multi-model claim history
Claims exist in structured outputs for roughly 4,000 runs from each provider, but historical claim-observation persistence is presently strongest for GPT. Backfilling Claude, Gemini and Grok will support valid model-by-model drift comparisons.
Priority 2Semantic claim matching
Added and removed claims will be distinguished from paraphrases, split claims, merged claims, stance changes, confidence changes and evidence changes.
Priority 3Claim-family separation
Business identity, product/service, audience, outcome, metadata, accessibility and technical findings will be reported as separate risk dimensions.
Priority 4Richer structural evidence
Heading hierarchy, structured data, navigation text, visible main content and normalized model input will allow stronger testing of how HTML structure and wording relate to interpretation stability.
Method and limitations
What these results do—and do not—show
SemanticRisk measures outputs generated from public website evidence at specific points in time. Content hashes validate whether normalized input changed. Model agreement measures consistency across evaluated claim signals. Materiality is assigned by the current drift rules.
The results do not prove that a particular HTML element directly caused a model response, that one provider is universally more accurate, or that every output difference represents a different real-world belief. Technical failures are excluded from substantive drift conclusions, and model or prompt versions must be controlled when making longitudinal comparisons.
The current evidence is best described as an observed association between website evidence and model interpretation—not a universal specification of what every downstream AI system will do.
See current company-level evidence.
Explore the public benchmark, compare companies, and review the latest interpretation records.
Open live benchmark