Your policy team changes a rule. Nobody opens a ticket.
Harm patterns change within hours, and a release takes weeks. Until it ships, your platform runs a policy you already know is out of date.
Building moderation tooling for platforms rather than running one? Yours is sell into the enterprise.
A rule your policy lead can read and edit
constants:
auto_ban_threshold: 3
rules:
block_repeat_offenders:
all:
- path: "event.type"
op: eq
value: "create_message"
- path: "state.user.counters.severe_violations"
op: gte
value: "{{ constants.auto_ban_threshold }}"
effects:
verdict: rejected
priority: 300
response:
reason: "This account is blocked after repeated violations." The engine keeps a count of severe violations for each account, so the rule can treat the third violation differently from the first. Change the threshold from 3 to 2 and promote it without waiting for an engineering release. Test it in shadow on live traffic first, or backtest it against last month and see which accounts it would have blocked.
Bring your own detectors
A detector can be a classifier that returns a score, a check on a criterion you write as one sentence, or a judge model that reads your whole rulebook. It can be a free open model, the vendor you already pay, or one you host yourself. Cheap checks run on everything, and the expensive ones only on what the cheap ones flag: how that works.
Regulators, by name
The EU Digital Services Act is the one that changes moderation work in Europe, and Article 24(5) sends every statement of reasons you issue to the Commission's public database, where anyone can read it.
| Obligation | What the engine gives you toward it |
|---|---|
| DSA Article 21: out-of-court dispute settlement | a certified body reopens a decision you already made, and reads the rule, the version and the reviewer from the same trail |
| DSA Article 24: transparency reporting | the counts come out of the same records that made the decisions |
| DSA Article 17: statement of reasons | every action already carries the rule version and the reason that produced it |
| DSA Article 20: internal complaint handling | an appeal is a second queue you declare, decided by someone other than the original reviewer, on the same trail |
| UK Online Safety Act risk assessment | what you enforce, since when, and what it caught |
| Child safety reporting | a flagged item goes to a child-safety queue with a priority, and the reviewer's decision and reason land on the same tamper-evident audit trail as the automated decisions: the record a CyberTipline report needs. Your categories and filing path are declared |
Your team files the report, with the decision, the reason and its tamper-evidence in one place.
Built for the people who review
Blurring an image by default, or holding a field back until a reviewer asks for it, is declared in the screen itself, so it happens every time.
By platform type
Social networks — posts, replies, coordinated behavior across accounts. Marketplaces — listings, seller messages, fraud patterns that only show up as a sequence. Streaming and creator platforms — uploads, live chat, monetization eligibility. In-game chat — text at very high volume, and grooming patterns you see in behavior rather than in single words.
One rule pack covers all four, because the regulator, the buyer and the engine are the same.