Reading the benchmark

What the SemanticRisk scores mean

Scores are normalized indicators of how clearly and consistently public website content can be interpreted by the evaluation pipeline. They are comparative signals, not probabilities and not legal or factual certifications.

Scale

All public scores use a 0–100 scale

75–100Stronger clarity or lower observed semantic risk.
55–74Mixed signals or moderate interpretation risk.
0–54More ambiguity, missing evidence or unstable interpretation.

Higher is generally better unless a page explicitly labels a metric as severity, drift or risk. Those measures are interpreted in the opposite direction.

Accuracy

Does the interpretation match the available content?

Measures how well extracted meaning is supported by the observed source material.

Confidence

How unambiguous does the content appear?

Measures whether the website gives models enough consistent evidence to form a stable interpretation.

Hallucination resistance

How much room is there for unsupported inference?

Rewards explicit, well-scoped statements and penalizes gaps that invite models to fill in missing context.

Completeness

Are important details and boundaries present?

Measures whether the interpretation captures the core proposition, scope and relevant qualifiers.

Comparisons

Use sector and trend context, not a number in isolation

Sector position

Shows how a company compares with others in the same public benchmark sector.

Overall position

Shows how the company compares with the wider monitored benchmark.

Trend

Shows whether the observed score is improving, declining or remaining stable across repeated scans.

Evidence volume

Low observation counts or crawl limitations reduce how much confidence should be placed in comparisons.

Do not over-read the score

A score is not a percentage chance of being correct

A score of 70 does not mean an AI answer is 70% correct. It is a normalized index assembled from observed interpretation quality, clarity, completeness and resistance to unsupported inference.

Versioning

Method changes are tracked

Scoring logic may evolve as the dataset grows and calibration improves. Published comparisons should always be read with the displayed observation period and methodology version.

Public scoring guide updated 30 July 2026.