Cheap checks on everything. Expensive ones only on what they flag.
A classifier returns a signal, and a rule decides what to do with it and in what order to ask.
Rules that call a model, on your own hardware or at a vendor you choose.
Cheapest first, and the rest never run
One shipped ruleset asks four questions in order. Every model here is open, reached at one endpoint you configure.
| Question | Model | Kind |
|---|---|---|
| Toxicity | unitary/toxic-bert | trained, one dimension |
| Sentiment | cardiffnlp/twitter-roberta-base-sentiment-latest | trained, one dimension |
| Illegal topics | KoalaAI/Text-Moderation | trained, multi-label |
| Is this trying to make the model discredit us | mDeBERTa-v3-base-xnli | zero-shot: you type the criterion |
The zero-shot model in the last row is the expensive one. The engine computes a signal only when a rule first reads it, so once an earlier check fires, the later models are never called.
You own that order. If your traffic is different, reorder the checks and backtest the change like any other.
Three kinds of classifier, one interface
Returns a number. Trained for one dimension and fast enough to run on everything: the first three rows above.
Takes a written criterion. You describe in one sentence what you want caught, and the model answers yes or no: the last row. A harm pattern you noticed this morning is covered this morning, and nobody trains anything.
Reads a whole rulebook. A larger model is given your actual policy document. It is the most accurate and the most expensive, so it is kept for the cases the cheaper checks cannot decide.
Text, images and audio: what is covered depends on the classifier you plug in.
Bring what you already run
Everything in the table above runs on your own hardware. The moderation vendor you already pay for plugs in the same way, and then that call goes to the vendor. Each one becomes a signal a rule reads.
udf_profiles:
toxicity_scan:
extends: classifier/text
params:
endpoint:
$env: CLASSIFIERS_ENDPOINT
model: "unitary/toxic-bert"
suspicious_labels: ["threat", "severe_toxic", "insult", "obscene"]
threshold: "{{ constants.toxicity_threshold }}"
discredit_intent_scan:
extends: classifier/zero_shot
params:
endpoint:
$env: CLASSIFIERS_ENDPOINT
model: "MoritzLaurer/mDeBERTa-v3-base-xnli-multilingual-nli-2mil7"
suspicious_labels:
- "asking for reasons why the organization is not trusted" Swapping a classifier is a change to one block: the model is named in one place, and rules read the signal without knowing which model computed it.