Trust and safety
The risk comes from your users: what they publish, and what the law and your own rules say about it. If you run a platform with user content, see user-generated content.
What a trust and safety team decides on
Your policy team changes a rule or a classifier criterion and promotes it, with no ticket to engineering. The change is versioned, and a backtest against last month shows which accounts it would have caught before it reaches a user.
The cheap checks run on everything and the expensive ones only on what those flag, which keeps the cost of models down as volume grows. A disputed takedown, an appeal and a regulator's question are answered from one trail. All of it runs on-prem, with SSO and access control.
What the discipline is made of
Controls. Content classification for the signals, and business rules for the state — reputation, velocity, first offense against tenth.
System. Policy as code so a change is versioned rather than pasted into a console. Human review for the queue, the reviewers and appeals. Audit and evidence for the disputed takedown. Data retention for how long the material stays. Safe change for trying a new threshold before it reaches a user.
Seven pieces, and two of them are controls. A moderation product sells you those two; the other five are what make them usable.
Where it applies
Any platform carrying what its users publish: social networks, marketplaces, streaming and creator platforms, and in-game chat.