Skip to content
TopicTracker
From HackerNewsView original
TranslationTranslation

Human-bench: an eval for "human shaped" agents

Human-Bench is a new evaluation benchmark designed to measure how closely AI agents perform like humans across a variety of real-world tasks. The leaderboard tracks and compares the "human-shaped" behavior of different AI systems.

Background

Human-Bench is a leaderboard that evaluates AI agents not just on raw task completion, but on how "human-shaped" their behavior is—meaning how natural, efficient, and interpretable their actions appear to a human observer. It fills a gap left by benchmarks like SWE-bench (which focuses solely on code-writing success) or standard agent evals (which only score final outcomes). Key entities include the developers at Modal (a cloud compute platform) and the broader AI alignment community concerned with agents that act inscrutably or wastefully. The project arose from a growing realization that as AI agents become more autonomous, judging only what they accomplish misses whether they do so in a way humans can understand and trust.

Related stories