Skip to main content

Policy model

A policy is what turns detections into decisions. It is a named, versioned object, written and read entirely in the console, and assigned to endpoints one at a time. This page is the model. For the exact steps of writing one, see Policies; for the file format on a self-managed machine, see Policy file format.

Two rule kinds​

A policy has two kinds of rule, and you can read both directly on a policy in the console:

A policy opened in the console: access rules (deny an unmanaged device, mutate an IDE to shadow), then prompt rules — prompt injection Block, sensitive data Redact, malicious entity Block, language Log — each with its threshold and priority

  • Access rules are conditions on metadata — identity, device, application, provider and model, network and time. They decide whether and how an interaction is processed, before any content is inspected: deny blocks an unsanctioned app or an unmanaged device, mutate changes the processing mode (forcing log-only, or escalating to a stricter rule set), and allow proceeds. This is attribute-based access control over AI use. Their conditions are exact matches on a context path, by design — no pattern matching, no comparison.
  • Prompt rules are content detectors from the six detection classes, each with a threshold (the score a detection must reach), the stage it applies to (prompts, responses, or both), and an action.

The three actions​

ActionEffect
LogPass through and record. The only action a log-only collector supports.
RedactReplace the matched spans with a placeholder, before submission for prompts and before delivery for responses. The interaction proceeds.
BlockStop the request or response and return a policy message.

How a policy is written​

New policy opens the builder. It takes a name and a default action, then one row per detector: the action to take, the score its detection must reach, and whether the rule applies to prompts, responses or both — no file to edit, no syntax to get wrong.

The policy builder: a name and default action, then a row per detector with its action, threshold and applies-to

Three things the builder decides for you: a sensitive-data rule always covers the credentials, financial and pii subclasses; a Redact rule always writes the placeholder [REDACTED:{class}]; and there is exactly one rule per detector. The full behaviour — duplicating instead of editing, assignment, and how a change reaches a machine — is in Policies.

Evaluation order​

When an interaction is inspected, on the machine, in single-digit milliseconds:

  1. Access rules first, in order: deny short-circuits to Block, mutate caps downstream actions, allow proceeds.
  2. Prompt rules run, collecting all matches. One interaction can hit several classes.
  3. Resolve to one effective action — Block > Redact > Log, most severe wins, ties broken by priority.
  4. Redaction is cumulative — every Redact rule contributes spans, and overlapping spans merge into one clean rewrite.

With no match, the policy's default action applies (Log by default). A log-only mutate caps even that, so such a policy cannot Block no matter what a rule says.

Context changes the verdict​

Because access rules read metadata, the same content can produce a different verdict on a different machine. An access rule that caps an unmanaged laptop at log-only means that laptop is watched, not policed: the detection is still recorded, only the enforcement is withheld, and the console marks it Shadow. See Watched, not policed.