Flipping the eval on its head — toward k×n×m testing
Current evals test n samples against 1 or k models. The proposal: flip to k×n×m — test k models against n samples across m dimensions (e.g., vulnerability detection, patch quality, proof assistants).
Three concrete cyberhardening tactics: red-blue loops (Claude finds vuln, Claude patches), proof retrofitting (Verus annotations in Rust, Lean specs from Aenaeus), and formal method synthesis.
Ambitious token spend today; may be cost-viable in six months or next year as capabilities scale.