Skip to main content

Policy file format

This page is the file format. It is what a policy looks like on disk, for someone editing policy.yaml on a machine or reading the file the console assigned. For the model itself — what a policy decides and how it is built — read Policy model, which shows it in the console with no YAML at all.

On an enrolled machine, the console owns this file

An enrolled endpoint pulls its policy from the console and overwrites the local file whenever the assigned version changes, so editing it by hand is replaced at the next check-in. Editing policy.yaml directly is how you change a machine that was never enrolled. See Configuration.

Two rule kinds​

  • Access rules are attribute conditions on metadata such as identity, device, application, provider and model, network and time. They decide whether and how a request is processed, before any content is inspected: allow proceeds, deny blocks an unsanctioned app or an unmanaged device, and mutate changes the processing mode, either forcing log-only or escalating to a stricter rule set. This is attribute-based access control over AI use.
  • Prompt rules are content detectors from the six detection classes, each with a threshold and an action.

Actions​

ActionEffect
LogPass through and record. The only action a log-only collector supports.
RedactReplace matched spans with a placeholder, before submission for prompts and before delivery for responses. The interaction proceeds.
BlockStop the request or response. Return a policy message.

Evaluation order​

  1. Access rules first, in author order: deny short-circuits to Block, mutate caps downstream actions, allow proceeds.
  2. Prompt rules run locally and inline, collecting all matches. One interaction can hit several classes.
  3. Resolve to one effective action: Block > Redact > Log, most severe wins, ties broken by priority.
  4. Redaction is cumulative. Every Redact rule contributes spans. Overlapping spans are merged into one clean rewrite.

With no match, the policy's default_action applies, and that is log by default. A log-only mutate caps even that. Such a policy cannot Block no matter what a rule says.

A complete example​

policy:
id: pol_workforce_std
name: "Workforce - standard"
version: 7
default_action: log
access_rules:
- { id: ar_block_unsanctioned_mcp, when: { application.id: "mcp:pastebin" }, effect: deny }
- { id: ar_monitor_lab, when: { device.hostname: build-lab-01 }, effect: mutate, mode: log_only }
prompt_rules:
- { id: pr_injection, detector: prompt_injection, applies_to: [prompt],
threshold: 0.80, action: block, priority: 100 }
- { id: pr_secrets, detector: sensitive_data, subclasses: [credentials, financial, pii],
applies_to: [prompt, response], action: redact,
redaction: { placeholder: "[REDACTED:{class}]" } }
- { id: pr_malicious_url, detector: malicious_entity, applies_to: [response], action: block }
- { id: pr_topic, detector: topic, subclasses: [legal_advice, medical_advice],
threshold: 0.75, action: log }
- { id: pr_language, detector: language, allow: [en, es], action: block }
custom_definitions:
- { id: cd_caseid, kind: regex, pattern: "\\bWZ-\\d{6}\\b",
class: confidential.case_id, action: redact }

The lab-machine access rule is worth reading twice: that machine gets watched, not policed. Detections are still recorded, enforcement is withheld. Identical content produces a different verdict because policy reads the context. See Watched, not policed.

An access rule's conditions are exact matches on a context path, by design. There is no pattern matching in them and no comparison.

A condition can only match a key some collector actually sends. The keys the collectors fill today:

KeyFilled byExample values
device.hostnameevery collectorbuild-lab-01
user.namethe hook, the MCP wrapper, the browser's native-messaging host, and the proxy when it can name the connection's owneralice
application.idthe proxy (from the client's user agent), the hook, the MCP wrapperclaude-code, cursor, mcp:<server name>
provider.namethe proxy; the hook only when Claude Code sends the model, which it rarely doesanthropic, openai
provider.modelthe proxy, from the request; the hook under the same conditionclaude-sonnet-4-5
provider.idthe browser extensionopenai-chatgpt, anthropic-claude, perplexity, google-gemini

A key no collector sends, such as device.managed or application.category, never matches, so a rule on it never fires. policy validate cannot warn about that, because context keys are open-ended.

For a topic rule, subclasses: picks the categories it acts on. The grammar also accepts a categories: list, but the agent does not read it: a topic rule with only categories: acts on every category the classifier scores.

Keys are checked​

A key the policy grammar does not define is an error, at every level. A misspelled thresold:, or deny_languages: where the key is deny:, is refused instead of loading a rule that silently ignores the line:

  • lumen-agent policy validate fails and names the key. Run it before you ship a policy file.
  • When the file changes under a running agent, or the console pushes a policy, the agent refuses the new version and keeps running the previous one.
  • The console refuses the same keys before it saves a policy.

When the agent starts, it is tolerant, because a policy file an earlier agent accepted must never keep the service from running. A file that only these checks refuse loads the way agents up to v2.5 loaded it: unknown keys are ignored and a language rule with neither list is dropped. The agent logs a warning naming what it refused, and policy show prints the same warning. Fix the file, or reassign the policy from the console, to clear it.

A language rule needs allow:, deny: or both. With neither it could never match, so it is refused. When the agent is confident which language a prompt is in, the rule acts on that. When it is not, which is common for prompts of a few words, the rule decides among the languages it lists and a set of widely used languages in the same script, so a short prompt in a listed language is still recognised. Short Russian and Korean prompts are caught this way.

This has limits on prompts of a few words, and they matter most for rules that block:

  • A language outside that set can be read as its nearest neighbour inside it. A few words of Bulgarian, Macedonian or Serbian are often read as Russian by a deny: [ru] rule. The same happens in Latin script: short Swedish or Danish can be read as Dutch by deny: [nl], short Romanian as French by deny: [fr], and short Indonesian or Czech as Italian by deny: [it]. The block's reason then names a language the prompt is not in.
  • Short Russian is often read as Ukrainian, sometimes with full confidence. It then escapes a deny: [ru] rule and can trip a deny: [uk] rule.
  • A short English question can be read as another Latin-script language by an allow: [en] rule.

List the neighbouring language yourself when it matters to you (deny: [ru, bg], deny: [nl, sv, da]), which lets it be recognised as itself, and keep a language rule at Log until you have seen what it matches.

Redaction spans​

A span targets one field and carries a zero-based [start, end) character offset over the normalized text, the detection class, and the placeholder substituted. Spans are stored on the finding, so the console shows what class was redacted without ever storing the raw value.

Validating and inspecting​

lumen-agent policy validate --policy /etc/lumen/policy.yaml
lumen-agent policy effective --config /etc/lumen/agent.yaml

policy effective resolves the policy exactly the way the daemon does, so it cannot report something the daemon does not run. See Configuration for why a hand-edited agent.yaml that drops the policy: key silently runs the log-only policy compiled into the binary instead, so that edits to this file never take effect.