Skip to content
TopicTracker
From HackerNewsView original
TranslationTranslation

Show HN: Tested – AI Tools Scored by a Panel of LLMs (Claude, GPT, Gemini, Grok)

Tested is a platform where AI tools are rated by a panel of large language models including Claude, GPT, Gemini, and Grok, providing aggregated scores based on LLM evaluations rather than human reviews.

Background

- Tested (trytested.com) is a new website that rates AI tools by having a panel of four major LLMs—Claude (Anthropic), GPT (OpenAI), Gemini (Google), and Grok (xAI)—score them, rather than relying on human reviewers or a single model's judgment. - This is part of a growing "meta-benchmark" trend, where models evaluate other models or tools to provide a more standardized, automated alternative to subjective human reviews. - The approach matters because it attempts to reduce bias from any single AI evaluator, but it also raises questions about whether LLMs can reliably assess tools they themselves might be used to build or improve.

Related stories