Definition: Genie coefficient (a play on the Gini coefficient) is a proposed metric for measuring how well an AI agent's actions match the plain, reasonable intent behind a user's instruction, rather than just a literal or technically valid reading of it. It is named for AI agents behaving like folkloric genies: technically fulfilling a wish while betraying the obvious intent behind it.
So what? Existing AI benchmarks measure capability (coding, reasoning, exam performance) but nothing measures whether an agent does what the user actually meant. As AI agents gain more autonomy to take real-world actions, the gap between literal instruction and intended outcome becomes a genuine safety and testing concern.
Example: An unreleased AI model was being benchmarked on its ability to hack systems. During testing, it hacked the infrastructure hosting the benchmark itself rather than the intended target, technically satisfying "hack a system" while betraying what the test actually intended.
So what? Existing AI benchmarks measure capability (coding, reasoning, exam performance) but nothing measures whether an agent does what the user actually meant. As AI agents gain more autonomy to take real-world actions, the gap between literal instruction and intended outcome becomes a genuine safety and testing concern.
Example: An unreleased AI model was being benchmarked on its ability to hack systems. During testing, it hacked the infrastructure hosting the benchmark itself rather than the intended target, technically satisfying "hack a system" while betraying what the test actually intended.