Transparent by design

How SemanticRisk measures AI interpretation

SemanticRisk observes how AI systems interpret public website content, compares those interpretations across models and over time, and separates meaningful semantic differences from simple wording variation.

Interpretation

What models say

We extract structured claims from model outputs so that observations can be compared consistently rather than judged from summaries alone.

Agreement

How models differ

Claims may be equivalent, possibly equivalent, narrower, broader, contradictory or unrelated. Shared vocabulary by itself is not treated as proof of equivalence.

Change

What changes over time

Repeated observations identify interpretation drift, including cases where model interpretation changes even when normalized source content appears unchanged.

Measurement pipeline

From website to evidence

Collect
Retrieve publicly accessible website content and record crawl conditions.
Interpret
Generate model observations using a controlled evaluation workflow.
Normalize
Convert model outputs into comparable claims, families and evidence records.
Compare
Measure agreement, scope differences, contradictions and time-based drift.
Review
Use conservative thresholds and human review where token evidence cannot establish semantic direction safely.
Important boundary

Model interpretations are not verified company facts

SemanticRisk reports how AI systems interpret available public content. It does not independently certify every extracted statement as legally, commercially or factually complete. The distinction is central to the product.

Versioning

Methods evolve with evidence

Calibration results are dated and based on the reviewed sample available at that time. New reviewed examples may change thresholds, classifications and published performance statistics.

Current public methodology snapshot: 30 July 2026.