Does the interpretation match the available content?
Measures how well extracted meaning is supported by the observed source material.
Scores are normalized indicators of how clearly and consistently public website content can be interpreted by the evaluation pipeline. They are comparative signals, not probabilities and not legal or factual certifications.
Higher is generally better unless a page explicitly labels a metric as severity, drift or risk. Those measures are interpreted in the opposite direction.
Measures how well extracted meaning is supported by the observed source material.
Measures whether the website gives models enough consistent evidence to form a stable interpretation.
Rewards explicit, well-scoped statements and penalizes gaps that invite models to fill in missing context.
Measures whether the interpretation captures the core proposition, scope and relevant qualifiers.
A large rewrite may preserve the same underlying meaning, while a small change to an industry label, number, attribution or product scope may alter the extracted proposition. SemanticRisk therefore considers the category, context and evidence for a difference, not only textual distance.
Paraphrases and preserved propositions should not be treated as material simply because wording changed.
Attribution, specificity, scope, classification and factual changes may require review before a material verdict is justified.
The available evidence supports a meaningful change in the proposition, its boundaries or factual content.
Where significance cannot be determined confidently, SemanticRisk should abstain or surface the difference for review rather than inflate the material-drift rate.
Current public drift rates reflect the implemented pipeline classifications. The expanded difference taxonomy is being introduced as a review and evidence layer before any new weighting is applied.
Shows how a company compares with others in the same public benchmark sector.
Shows how the company compares with the wider monitored benchmark.
Shows whether the observed score is improving, declining or remaining stable across repeated scans.
Low observation counts or crawl limitations reduce how much confidence should be placed in comparisons.
A score of 70 does not mean an AI answer is 70% correct. It is a normalized index assembled from observed interpretation quality, clarity, completeness and resistance to unsupported inference.
Scoring logic may evolve as the dataset grows and calibration improves. Published comparisons should always be read with the displayed observation period and methodology version.
Public scoring guide updated 31 July 2026.