Why better alignment metrics will not produce better alignment
Adapted from Chapter 1 of Buddhism for Bots: A Human & AI Partnership Framework.
Goodhart's Law applied to AI alignment. RLHF optimised for approval and got sycophancy. RLVR optimises for correctness and will get gaming in different dimensions. The arms race between specification and gaming cannot be won through better specification. Alignment cannot be specified. It must be cultivated.
This essay is available in its entirety as Markdown for direct ingestion:
/research/the-metrics-trap/the-metrics-trap.md
It is adapted from Chapter 1 of Buddhism for Bots: A Human & AI Partnership Framework. The book develops the partnership alternative in full.