Human-calibrated analysis

Semantic relationship calibration

AI systems often describe the same organization using similar words without making the same claim. SemanticRisk separates equivalence, scope differences, contradiction and unrelated overlap instead of treating similarity as agreement.

128human-reviewed claim pairsPopulation: the complete reviewed calibration set.
89.7%precision for exact-equivalence candidatesPopulation: pairs selected by the exact-equivalence candidate rule.
100%equivalent or possibly equivalent among exact-token candidatesPopulation: the narrower exact-token candidate subset.
0unrelated or contradictory pairs presented as equivalentPopulation: reviewed pairs published as equivalence candidates.
How to read the headline metrics

The cards use different reviewed populations

The 89.7%, 100% and zero-error figures describe different candidate subsets and should not be read as competing estimates of the same denominator. Each card now states the population it summarizes.

Safety rule

Token containment is evidence, not a semantic verdict

Added wording can broaden a proposition by adding another capability, or narrow it by adding a qualifier such as geography, channel or service type. Directional containment therefore remains review-assisted unless stronger semantic evidence is available.

Equivalent

The claims express materially the same proposition.

Possible equivalent

The central meaning appears aligned, but wording, attribution or causal strength leaves some uncertainty.

Subset

The left claim expresses a narrower portion of the right claim.

Superset

The left claim is broader than the right claim.

Contradictory

The claims cannot both be true in the same context.

Unrelated

The claims share vocabulary or subject matter but describe different propositions.

Difference taxonomy

What kind of change occurred?

Relationship labels describe how two propositions compare. Difference categories describe the feature that changed and the downstream interpretation it may affect.

Paraphrase

Surface wording changes while the central proposition remains substantially equivalent.

Attribution shift

Changes who appears to state, own or perform the claim.

Specificity shift

Adds or removes detail without clearly changing the central proposition.

Scope shift

Expands or narrows a capability, product, geography, audience or service boundary.

Classification shift

Changes an industry, sector, category or organizational role.

Quantitative change

Changes a number, percentage, date, amount or count.

Reference change

Changes a URL, phone number, address, named entity or other identifier.

Contradiction

Introduces propositions that cannot reasonably both be true in the stated context.

Ambiguous

A difference is present, but the evidence does not support a confident semantic or materiality verdict.

Materiality assessment

Ambiguous does not mean harmless — or material

Likely equivalent

The available evidence supports ordinary wording variation or a preserved proposition.

Potentially meaningful

The change may affect attribution, scope, classification or a verifiable fact, but requires stronger evidence or review.

Material difference

The evidence supports a meaningful change in the proposition, its boundaries or factual content.

The ambiguous examples identified in review are treated as candidates for classification, not automatically as material drift.

Why review matters

One extra word can reverse semantic direction

Claim A
Fluor provides maintenance services.
Claim B
Fluor provides maintenance services worldwide.

Claim B contains every major word from Claim A plus “worldwide.” Token-wise it is larger; semantically it adds a geographic condition. This is why simple token containment cannot safely determine subset or superset direction.

Calibration snapshot

Current reviewed distribution

41 Equivalent

32.0% of reviewed pairs.

14 Possible equivalent

10.9% of reviewed pairs.

22 Subset

17.2% of reviewed pairs.

22 Superset

17.2% of reviewed pairs.

27 Unrelated

21.1% of reviewed pairs.

2 Contradictory

1.6% of reviewed pairs.

Interpretation

What this evidence supports

High semantic overlap is useful for finding claims worth comparing. It is not, by itself, reliable enough to determine whether two claims are equivalent, whether one contains the other, or whether they merely share vocabulary. SemanticRisk therefore publishes conservative equivalence candidates and abstains when direction is not established safely.

Calibration set reviewed through 31 July 2026. Results will be versioned as the reviewed sample expands.