Skip to content
TopicTracker
From HackerNewsView original
TranslationTranslation

Anthropic Changed the Sonnet 5 Chart After It Made Sonnet Look Bad

Anthropic quietly updated its Claude Sonnet 5 benchmark chart after the original version showed Sonnet performing worse than competitors. The revision adjusted the chart's appearance, drawing scrutiny for allegedly making the model look better than the initial data indicated.

Background

- **Anthropic** is the AI company behind Claude, a family of large language models (LLMs) competing with OpenAI's GPT and Google's Gemini. "Sonnet" refers to Claude 3.5 Sonnet, one of Anthropic's mid-tier but widely used models. - **Benchmark charts** are standard in the AI industry: companies publish bar charts showing how their new model outperforms rivals on various tests (coding, math, reasoning, etc.). These charts are a major marketing tool and are scrutinized by the AI community. - The article claims Anthropic quietly altered a published benchmark comparison chart for "Sonnet 5" (likely a typo or internal codename for Claude 3.5 Sonnet or a newer variant) after the original chart made Sonnet look weaker than competitors — a practice that, if true, raises questions about transparency in how AI companies present performance data.