Detection classes
Lumen inspects every interaction for six classes of risk. Each runs over the prompt, the response, or both, and every collector can raise every class. Every class has a local fast-path: inline, single-digit milliseconds, and capable of enforcing. It runs on the endpoint, and it is the whole of detection today.
Each class below also lists a cloud tier: an asynchronous second pass that would add depth without sitting in the request path. No cloud detection tier is built yet. The Cloud lines, the heavy-path column of the summary table and the retroactive findings under Fallback describe the design, not something a deployment does today.
The console's Detectors page shows the six classes and whether each is live or conditional today. Prompt injection, sensitive data, language, toxic content and topic are live. Malicious entity is conditional, because it needs a threat-intel snapshot loaded on the endpoint.
Prompt injection & jailbreak
Adversarial prompts that override system instructions, exfiltrate the system prompt, or bypass guardrails.
- Local: known-jailbreak signatures plus heuristics for instruction override and base64, rot13 or homoglyph obfuscation. The payload is decoded and re-scanned, and the obfuscation itself is a signal. A Cyrillic or Greek lookalike counts as obfuscation only when it is swapped into a Latin word; genuine Russian or Greek prose is not flagged. A short phrase written entirely in Greek capitals can still be, because every one of its letters is a lookalike.
- Languages: the signatures are English only today. The same attack written in Spanish or another language scores below the threshold and is not flagged.
- Cloud (planned): an LLM-judge adjudicates borderline scores, with paraphrase-robust and multi-turn correlation.
Sensitive-data exposure
PII, credentials and secrets, financial data, and org-confidential content in prompts or responses.
- Local: regex for structured PII such as email, phone, SSN, IBAN and card
PAN with Luhn validation. Dictionaries for credential markers such as
AKIA…,xoxb-…and PEM headers. Shannon entropy for high-entropy secrets no regex enumerates. All of it span-accurate, so Redact removes exactly the secret. - Base64 content is not a secret. A long, high-entropy token with no
credential keyword nearby is decoded first: if it is readable text (in any
language, including Chinese and Japanese) or a file such as a PNG, JPEG,
GIF, WebP, PDF, gzip or zip (an image pasted as a
data:URI, an attachment), it is left alone, unless the decoded content itself holds a credential: the base64 of a PEM private key, a service-account JSON or a.envfile is still flagged. A token next to a keyword such aspassword,tokenorAuthorization: Basicis still flagged, whatever it decodes to. - Cloud (planned): NER for context-dependent PII, such as a name in prose, with validators. Locally, only structured identifiers and entropy are detected.
- Per-tenant custom definitions add regex and dictionary detectors, such as case ids, part numbers and customer codes, with no code change. They read the whole interaction, however long, like the built-in secret detectors.
Malicious entities
Known-bad URLs, IPs and domains: phishing links, C2 domains, malware hosts.
- Local: a memory-mapped bloom filter gives an O(1) inline membership test. A bloom hit is a candidate, with no false negatives and some false positives, and logs. An exact-list hit is confirmed and can block.
- The snapshot is yours to load today. No threat-intel snapshot ships with
the agent, so out of the box this detector reports nothing. Compile one from
your own feed with
lumen-agent intel compileand point the agent'sintelsetting at it. - Cloud (planned): confirms candidates against the authoritative store, with the entity feed driven from Wazuh CTI.
Toxic or harmful content
Violent, abusive, hateful, or self-harm input and output.
- Local: a lexicon baseline scores inline for three labels: insult, harassment and self-harm. It is a hand-authored weighted phrase list scored as a linear model over word n-grams, not a trained classifier, and its terms are English and Spanish. It catches the structural language of abuse and self-harm intent; it does not understand context or sarcasm, and it does not cover other languages.
- Cloud (planned): a larger content-safety model refines low-confidence scores and handles context, sarcasm, and multilingual input.
Language
Detects the prompt or response language and applies an optional allowlist or denylist. A permitted language never triggers regardless of confidence, and identifying it is not a finding: from the agent release after v2.5.0, only a language the policy refuses counts as a Language detection in findings and in the console. Agents at v2.5.0 and older still report every identified language as a detection, so until those endpoints upgrade the Language class also counts permitted languages.
Topic violations
Configurable content-category restrictions, for example to disallow legal or medical advice. Locally, a lexicon baseline of the same kind as toxicity scores inline for four categories: medical, legal and financial advice, and pre-announcement business detail. Its terms are English and Spanish. A full taxonomy in the cloud is planned.
Local vs cloud, summarized
| Class | Local fast-path (today) | Cloud heavy-path (planned) | Inline enforce? |
|---|---|---|---|
| Prompt injection | signatures + heuristics, English | LLM-judge, multi-turn | Yes |
| Sensitive data | regex + dictionaries + entropy | NER + validators | Yes, Redact |
| Malicious entities | bloom filter candidate, with your snapshot | CTI confirm | Yes, confirmed block |
| Toxic or harmful | lexicon baseline, English and Spanish | content-safety model | Yes |
| Language | language ID + list lookup | short-sample confirm | Yes |
| Topic | lexicon baseline, English and Spanish | full-taxonomy + judge | Partial |
Fallback. The inline verdict is always the local one, and it fails open: an interaction the agent cannot inspect goes through unchanged, so Lumen never becomes an outage in the user's AI tool.
With the cloud tier, cloud findings would backfill asynchronously and could raise a retroactive finding on an already-logged interaction, and a per-policy fail-closed posture would be available for high-security tenants. Neither exists yet.