Claude Opus 4.5 Negative AA-Omniscience Index but 51.3 FACTS
https://stateofseo.com/what-do-strategic-teams-lose-when-they-treat-ai-as-a-single-answer-tool/
Anthropic Conflicting Scores: Why Benchmark Disagreement Clouds Hallucination Rates Understanding Hallucination Metrics in Large Language Models As of April 2025, the AI landscape looks messier than ever when it comes to measuring