Reading the benchmark

What the SemanticRisk scores mean

Scores are normalized indicators of how clearly and consistently public website content can be interpreted by the evaluation pipeline. They are comparative signals, not probabilities and not legal or factual certifications.

Scale

All public scores use a 0–100 scale

75–100Stronger clarity or lower observed semantic risk.
55–74Mixed signals or moderate interpretation risk.
0–54More ambiguity, missing evidence or unstable interpretation.

Higher is generally better unless a page explicitly labels a metric as severity, drift or risk. Those measures are interpreted in the opposite direction.

Accuracy

Does the interpretation match the available content?

Measures how well extracted meaning is supported by the observed source material.

Confidence

How unambiguous does the content appear?

Measures whether the website gives models enough consistent evidence to form a stable interpretation.

Hallucination resistance

How much room is there for unsupported inference?

Rewards explicit, well-scoped statements and penalizes gaps that invite models to fill in missing context.

Completeness

Are important details and boundaries present?

Measures whether the interpretation captures the core proposition, scope and relevant qualifiers.

Material claim drift

Semantic change is not measured by word count

A large rewrite may preserve the same underlying meaning, while a small change to an industry label, number, attribution or product scope may alter the extracted proposition. SemanticRisk therefore considers the category, context and evidence for a difference, not only textual distance.

Likely equivalent

Paraphrases and preserved propositions should not be treated as material simply because wording changed.

Potentially meaningful

Attribution, specificity, scope, classification and factual changes may require review before a material verdict is justified.

Material difference

The available evidence supports a meaningful change in the proposition, its boundaries or factual content.

Ambiguous

Where significance cannot be determined confidently, SemanticRisk should abstain or surface the difference for review rather than inflate the material-drift rate.

Current public drift rates reflect the implemented pipeline classifications. The expanded difference taxonomy is being introduced as a review and evidence layer before any new weighting is applied.

Comparisons

Use sector and trend context, not a number in isolation

Sector position

Shows how a company compares with others in the same public benchmark sector.

Overall position

Shows how the company compares with the wider monitored benchmark.

Trend

Shows whether the observed score is improving, declining or remaining stable across repeated scans.

Evidence volume

Low observation counts or crawl limitations reduce how much confidence should be placed in comparisons.

Do not over-read the score

A score is not a percentage chance of being correct

A score of 70 does not mean an AI answer is 70% correct. It is a normalized index assembled from observed interpretation quality, clarity, completeness and resistance to unsupported inference.

Versioning

Method changes are tracked

Scoring logic may evolve as the dataset grows and calibration improves. Published comparisons should always be read with the displayed observation period and methodology version.

Public scoring guide updated 31 July 2026.