Policy file format
This page is the file format. It is what a policy looks like on disk, for
someone editing policy.yaml on a machine or reading the file the console
assigned. For the model itself — what a policy decides and how it is built — read
Policy model, which shows it in the console with no
YAML at all.
An enrolled endpoint pulls its policy from the console and overwrites the local
file whenever the assigned version changes, so editing it by hand is replaced at
the next check-in. Editing policy.yaml directly is how you change a machine
that was never enrolled. See Configuration.
Two rule kinds
- Access rules are attribute conditions on metadata such as identity,
device, application, provider and model, network and time. They decide
whether and how a request is processed, before any content is inspected:
allowproceeds,denyblocks an unsanctioned app or an unmanaged device, andmutatechanges the processing mode, either forcing log-only or escalating to a stricter rule set. This is attribute-based access control over AI use. - Prompt rules are content detectors from the six detection classes, each with a threshold and an action.
Actions
| Action | Effect |
|---|---|
| Log | Pass through and record. The only action a log-only collector supports. |
| Redact | Replace matched spans with a placeholder, before submission for prompts and before delivery for responses. The interaction proceeds. |
| Block | Stop the request or response. Return a policy message. |
Evaluation order
- Access rules first, in author order:
denyshort-circuits to Block,mutatecaps downstream actions,allowproceeds. - Prompt rules run locally and inline, collecting all matches. One interaction can hit several classes.
- Resolve to one effective action: Block > Redact > Log, most
severe wins, ties broken by
priority. - Redaction is cumulative. Every Redact rule contributes spans. Overlapping spans are merged into one clean rewrite.
With no match, the policy's default_action applies, and that is log by
default. A log-only mutate caps even that. Such a policy cannot Block no
matter what a rule says.
A complete example
policy:
id: pol_workforce_std
name: "Workforce - standard"
version: 7
default_action: log
access_rules:
- { id: ar_block_unsanctioned_mcp, when: { application.id: "mcp:pastebin" }, effect: deny }
- { id: ar_monitor_lab, when: { device.hostname: build-lab-01 }, effect: mutate, mode: log_only }
prompt_rules:
- { id: pr_injection, detector: prompt_injection, applies_to: [prompt],
threshold: 0.80, action: block, priority: 100 }
- { id: pr_secrets, detector: sensitive_data, subclasses: [credentials, financial, pii],
applies_to: [prompt, response], action: redact,
redaction: { placeholder: "[REDACTED:{class}]" } }
- { id: pr_malicious_url, detector: malicious_entity, applies_to: [response], action: block }
- { id: pr_topic, detector: topic, subclasses: [legal_advice, medical_advice],
threshold: 0.75, action: log }
- { id: pr_language, detector: language, allow: [en, es], action: block }
custom_definitions:
- { id: cd_caseid, kind: regex, pattern: "\\bWZ-\\d{6}\\b",
class: confidential.case_id, action: redact }
The lab-machine access rule is worth reading twice: that machine gets watched, not policed. Detections are still recorded, enforcement is withheld. Identical content produces a different verdict because policy reads the context. See Watched, not policed.
An access rule's conditions are exact matches on a context path, by design. There is no pattern matching in them and no comparison.
A condition can only match a key some collector actually sends. The keys the collectors fill today:
| Key | Filled by | Example values |
|---|---|---|
device.hostname | every collector | build-lab-01 |
user.name | the hook, the MCP wrapper, the browser's native-messaging host, and the proxy when it can name the connection's owner | alice |
application.id | the proxy (from the client's user agent), the hook, the MCP wrapper | claude-code, cursor, mcp:<server name> |
provider.name | the proxy; the hook only when Claude Code sends the model, which it rarely does | anthropic, openai |
provider.model | the proxy, from the request; the hook under the same condition | claude-sonnet-4-5 |
provider.id | the browser extension | openai-chatgpt, anthropic-claude, perplexity, google-gemini |
A key no collector sends, such as device.managed or application.category,
never matches, so a rule on it never fires. policy validate cannot warn about
that, because context keys are open-ended.
For a topic rule, subclasses: picks the categories it acts on. The grammar
also accepts a categories: list, but the agent does not read it: a topic rule
with only categories: acts on every category the classifier scores.
Keys are checked
A key the policy grammar does not define is an error, at every level. A
misspelled thresold:, or deny_languages: where the key is deny:, is refused
instead of loading a rule that silently ignores the line:
lumen-agent policy validatefails and names the key. Run it before you ship a policy file.- When the file changes under a running agent, or the console pushes a policy, the agent refuses the new version and keeps running the previous one.
- The console refuses the same keys before it saves a policy.
When the agent starts, it is tolerant, because a policy file an earlier
agent accepted must never keep the service from running. A file that only these
checks refuse loads the way agents up to v2.5 loaded it: unknown keys are
ignored and a language rule with neither list is dropped. The agent logs a
warning naming what it refused, and policy show prints the same warning. Fix
the file, or reassign the policy from the console, to clear it.
A language rule needs allow:, deny: or both. With neither it could never
match, so it is refused. When the agent is confident which language a prompt is
in, the rule acts on that. When it is not, which is common for prompts of a few
words, the rule decides among the languages it lists and a set of widely used
languages in the same script, so a short prompt in a listed language is still
recognised. Short Russian and Korean prompts are caught this way.
This has limits on prompts of a few words, and they matter most for rules that block:
- A language outside that set can be read as its nearest neighbour inside
it. A few words of Bulgarian, Macedonian or Serbian are often read as
Russian by a
deny: [ru]rule. The same happens in Latin script: short Swedish or Danish can be read as Dutch bydeny: [nl], short Romanian as French bydeny: [fr], and short Indonesian or Czech as Italian bydeny: [it]. The block's reason then names a language the prompt is not in. - Short Russian is often read as Ukrainian, sometimes with full
confidence. It then escapes a
deny: [ru]rule and can trip adeny: [uk]rule. - A short English question can be read as another Latin-script language
by an
allow: [en]rule.
List the neighbouring language yourself when it matters to you
(deny: [ru, bg], deny: [nl, sv, da]), which lets it be recognised as
itself, and keep a language rule at Log until you have seen what it
matches.
Redaction spans
A span targets one field and carries a zero-based [start, end) character
offset over the normalized text, the detection class, and the placeholder
substituted. Spans are stored on the finding, so the console shows what
class was redacted without ever storing the raw value.
Validating and inspecting
lumen-agent policy validate --policy /etc/lumen/policy.yaml
lumen-agent policy effective --config /etc/lumen/agent.yaml
policy effective resolves the policy exactly the way the daemon does, so it
cannot report something the daemon does not run. See
Configuration for why a hand-edited agent.yaml that
drops the policy: key silently runs the log-only policy compiled into the
binary instead, so that edits to this file never take effect.