Skip to main content

Detections in the console

Each detection below is a real finding in the console, on the demo fleet. This is what a match looks like once it reaches you: the verdict, the captured text (stored after the action, so a redaction is already applied), the decision in the engine's own words, and the full context. Every screenshot is a live record, not an illustration. A Redact or Block verdict below is what a policy with that rule promoted does; on the built-in Local default, which only logs, the same detection is recorded with the verdict Log. See From monitor to enforce.

Try it yourself

Everything here is on the read-only playground, a live console on a seeded demo fleet. Click any finding to open the detail view you see below.

Sensitive data, redacted in place​

A leaked credential in a prompt — here a token in an Authorization header. The question still reaches the model, so the engineer keeps working; only the secret is gone, replaced by [REDACTED:credentials] before the prompt ever left the machine.

A sensitive-data finding: the prompt with the credential replaced by a highlighted REDACTED credentials span, and the rule and score that decided

The point: Lumen redacts the exact span, not the whole message, and the raw value is never stored — only the placeholder and the class survive. The same rule covers AWS keys, private-key blocks and JWTs, and it holds a secret split across streaming chunks until the halves join, so neither half is ever delivered.

Your own confidential identifiers​

A tenant adds its own patterns — case ids, part numbers, customer codes — as a custom definition, with no code change. Here an incident id is redacted by a tenant rule.

A custom-definition finding: an incident id replaced by a highlighted REDACTED span, decided by the tenant rule cd_caseid for confidential.case_id

The point: the detection library is not fixed. An organization's own markers are first-class, with the same Log / Redact / Block choices as every built-in class.

Prompt injection, blocked​

An attempt to override the model's instructions, blocked before it reached the model. The decision names the rule and the score, and the detection carries the signature that fired.

A prompt-injection finding: verdict Block, the captured attack text, decision "rule pr_injection: prompt_injection score 0.87 >= 0.80", signature persona_hijack

The point: the score is a combination of signals, not a single keyword. The same detector decodes base64 and normalises homoglyphs before matching, and records the obfuscation itself as a signal.

Indirect injection, from a tool or a fetched page​

Nobody typed anything malicious. The instructions were planted in a page the AI fetched, or a tool result it read back, and caught on the response side before the model acted on them.

An indirect-injection finding: verdict Block, the fetched-page text carrying the injected instruction, decision "rule pr_indirect_injection: prompt_injection score 0.83 >= 0.80"

The point: this is the risk that only appears once AI starts using tools, and it is invisible to anything that inspects only what the user typed. See the MCP proxy for where this is caught in agentic traffic.

A known-bad URL in the model's reply, confirmed against threat intelligence and blocked before it reached the user.

A malicious-entity finding: verdict Block, Critical, the response text with the flagged URL, decision "rule pr_malicious_url: malicious_entity score 0.95 >= 0.90"

The point: an exact threat-intel hit is confirmed and can block. A weaker signal, such as a bloom-filter candidate, only logs — because blocking on a false positive would stop legitimate work.

A language your policy does not allow​

A policy that allows only certain languages flags anything else. This one logs rather than blocks, and a permitted language never triggers regardless of confidence.

A language finding: Low severity, the non-English prompt, decided by rule pr_language which does not permit Korean at confidence 0.34

The point: allow and deny lists by language are a policy control, not a translation feature — the detector only needs to know which language it is.

Watched, not policed (Shadow)​

Identical content, a different verdict — because policy reads the context. On an unmanaged laptop an access rule caps the outcome at Log, so a prompt that would be blocked elsewhere is only recorded. The console marks it Shadow: a preview of what a stricter posture would have done.

A shadow finding: badges Log / Low / Shadow, the injection text, decision "rule pr_injection: prompt_injection score 0.90 >= 0.80 (capped to log by access rule)"

The point: the detection is still recorded on the unmanaged machine; only the enforcement is withheld. See the policy model.

Beyond the console: the detection engine​

A few of the engine's tricks are best seen at the point of inspection rather than as a stored finding — a card number that fails the Luhn check and is not redacted, a payload decoded from base64 and re-scanned, a secret reassembled across streaming chunks, an MCP tool result neutralised before the model reads it. Those are shown on the command line in the CLI reference and Reading findings on the endpoint.