Skip to main content

Detection classes

Lumen inspects every interaction for six classes of risk. Each runs over the prompt, the response, or both, and every collector can raise every class. Every class has a local fast-path: inline, single-digit milliseconds, and capable of enforcing. It runs on the endpoint, and it is the whole of detection today.

Planned: the cloud heavy-path

Each class below also lists a cloud tier: an asynchronous second pass that would add depth without sitting in the request path. No cloud detection tier is built yet. The Cloud lines, the heavy-path column of the summary table and the retroactive findings under Fallback describe the design, not something a deployment does today.

The console's Detectors page shows the six classes and whether each is live or conditional today. Prompt injection, sensitive data, language, toxic content and topic are live. Malicious entity is conditional, because it needs a threat-intel snapshot loaded on the endpoint.

Prompt injection & jailbreak​

Adversarial prompts that override system instructions, exfiltrate the system prompt, or bypass guardrails.

  • Local: known-jailbreak signatures plus heuristics for instruction override and base64, rot13 or homoglyph obfuscation. The payload is decoded and re-scanned, and the obfuscation itself is a signal. A Cyrillic or Greek lookalike counts as obfuscation only when it is swapped into a Latin word; genuine Russian or Greek prose is not flagged. A short phrase written entirely in Greek capitals can still be, because every one of its letters is a lookalike.
  • Languages: the signatures are English only today. The same attack written in Spanish or another language scores below the threshold and is not flagged.
  • Cloud (planned): an LLM-judge adjudicates borderline scores, with paraphrase-robust and multi-turn correlation.

Sensitive-data exposure​

PII, credentials and secrets, financial data, and org-confidential content in prompts or responses.

  • Local: regex for structured PII such as email, phone, SSN, IBAN and card PAN with Luhn validation. Dictionaries for credential markers such as AKIA…, xoxb-… and PEM headers. Shannon entropy for high-entropy secrets no regex enumerates. All of it span-accurate, so Redact removes exactly the secret.
  • Base64 content is not a secret. A long, high-entropy token with no credential keyword nearby is decoded first: if it is readable text (in any language, including Chinese and Japanese) or a file such as a PNG, JPEG, GIF, WebP, PDF, gzip or zip (an image pasted as a data: URI, an attachment), it is left alone, unless the decoded content itself holds a credential: the base64 of a PEM private key, a service-account JSON or a .env file is still flagged. A token next to a keyword such as password, token or Authorization: Basic is still flagged, whatever it decodes to.
  • Cloud (planned): NER for context-dependent PII, such as a name in prose, with validators. Locally, only structured identifiers and entropy are detected.
  • Per-tenant custom definitions add regex and dictionary detectors, such as case ids, part numbers and customer codes, with no code change. They read the whole interaction, however long, like the built-in secret detectors.

Malicious entities​

Known-bad URLs, IPs and domains: phishing links, C2 domains, malware hosts.

  • Local: a memory-mapped bloom filter gives an O(1) inline membership test. A bloom hit is a candidate, with no false negatives and some false positives, and logs. An exact-list hit is confirmed and can block.
  • The snapshot is yours to load today. No threat-intel snapshot ships with the agent, so out of the box this detector reports nothing. Compile one from your own feed with lumen-agent intel compile and point the agent's intel setting at it.
  • Cloud (planned): confirms candidates against the authoritative store, with the entity feed driven from Wazuh CTI.

Toxic or harmful content​

Violent, abusive, hateful, or self-harm input and output.

  • Local: a lexicon baseline scores inline for three labels: insult, harassment and self-harm. It is a hand-authored weighted phrase list scored as a linear model over word n-grams, not a trained classifier, and its terms are English and Spanish. It catches the structural language of abuse and self-harm intent; it does not understand context or sarcasm, and it does not cover other languages.
  • Cloud (planned): a larger content-safety model refines low-confidence scores and handles context, sarcasm, and multilingual input.

Language​

Detects the prompt or response language and applies an optional allowlist or denylist. A permitted language never triggers regardless of confidence, and identifying it is not a finding: from the agent release after v2.5.0, only a language the policy refuses counts as a Language detection in findings and in the console. Agents at v2.5.0 and older still report every identified language as a detection, so until those endpoints upgrade the Language class also counts permitted languages.

Topic violations​

Configurable content-category restrictions, for example to disallow legal or medical advice. Locally, a lexicon baseline of the same kind as toxicity scores inline for four categories: medical, legal and financial advice, and pre-announcement business detail. Its terms are English and Spanish. A full taxonomy in the cloud is planned.

Local vs cloud, summarized​

ClassLocal fast-path (today)Cloud heavy-path (planned)Inline enforce?
Prompt injectionsignatures + heuristics, EnglishLLM-judge, multi-turnYes
Sensitive dataregex + dictionaries + entropyNER + validatorsYes, Redact
Malicious entitiesbloom filter candidate, with your snapshotCTI confirmYes, confirmed block
Toxic or harmfullexicon baseline, English and Spanishcontent-safety modelYes
Languagelanguage ID + list lookupshort-sample confirmYes
Topiclexicon baseline, English and Spanishfull-taxonomy + judgePartial

Fallback. The inline verdict is always the local one, and it fails open: an interaction the agent cannot inspect goes through unchanged, so Lumen never becomes an outage in the user's AI tool.

Planned

With the cloud tier, cloud findings would backfill asynchronously and could raise a retroactive finding on an already-logged interaction, and a per-policy fail-closed posture would be available for high-security tenants. Neither exists yet.