Human-bench: an eval for "human shaped" agents
Human-bench is a benchmark designed to evaluate AI agents that interact with the world in human-like ways—using vision, language, and physical actions. It provides a leaderboard ranking agents based on how well they perform tasks that resemble human cognitive and physical abilities, aiming to measure progress toward more natural and capable AI systems.