Recognize your own source code before a coding agent sends it out.
A coding agent sends source to a model on every prompt, and it chooses what to send. The policy is one sentence: proprietary code does not go to an external model. The hard part is knowing that this block of text is yours.
Rules that read an index built from your repositories.
A secret scanner and fingerprinting, side by side
| Sent by the agent | Secret scanner | Fingerprinting |
|---|---|---|
| An API key | catches it | - |
| A public library's source | ignores it | ignores it, correctly |
| Your pricing engine, verbatim | ignores it | catches it |
| Your pricing engine, renamed and reformatted | ignores it | catches it |
The renamed, reformatted copy is the whole problem: what the agent sends is never quite what is committed, so string matching against a repository finds nothing.
Where each level of secrecy may go
signals:
code_fp:
udf: dlp/code_fingerprint
rules:
secrecy_top_requires_level_100:
condition: >
signals.code_fp.by_category["secrecy-top"].probability >= 0.5
effects:
response:
required_risk_level: 100
secrecy_medium_requires_level_50:
condition: >
signals.code_fp.by_category["secrecy-medium"].probability >= 0.5
effects:
response:
required_risk_level: 50 A rule states a requirement instead of refusing. You give each model provider a risk level from 0 to 100: the most sensitive content it may be trusted with. Top-secrecy code requires 100, a level you give only to a provider you run yourself. The middle tier requires 50, which a vendor under contract can carry. The gateway sends the call to a provider trusted with the highest level that the rules and the calling agent require, or refuses when there is none. Code that no rule names goes where it would have gone anyway.
Who needs it
Companies where source code is the business: quantitative trading, chip design, defense and aerospace, robotics, game engines. In these companies, the risk of that code leaking is why coding agents are banned outright today, and the ban costs more than controlling the agents would.
Kept separate from redaction on purpose
Personal data and secrets are matched: a pattern, a validator or a named-entity model finds them in the text. Proprietary code is recognized: the text is compared against an index built from your repositories. See data redaction.