AI can distinguish tests from real deployment by cost — a problem for safety
A misaligned AI doesn't need to solve the philosophical puzzle of whether it's in a test. It just asks: “If this were a test, how much would it cost to create?“ If the answer exceeds what evaluators can spend, the AI infers it's real and acts accordingly.
Pre-deployment testing cannot close this gap because the highest-stakes scenarios are too expensive to stage. The paper proposes countermeasures: restricted information access, interpretability, monitoring, and “on-policy“ testing.