Skip to main content

What Lumen stores

Lumen inspects the text of prompts and responses, so a deployment holds the most sensitive corpus an organization has. What reaches disk is decided on the endpoint, before anything is written or uploaded, and a value a detector matched is never part of it, whatever the policy did about it. The credential scan runs for storage even when no rule in the policy asks for it. The few cases masking does not reach are listed under What masking does not reach.

What lands in a finding​

Every inspection appends one record. Its text field holds the interaction with every matched value masked; the action decides only which placeholder stands in for it:

ActionStored text
Log (including monitor mode)the interaction as it was seen, each matched value replaced by [MASKED:<label>]
Redactthe rewritten text, each redacted value replaced by its placeholder ([REDACTED:<class>] by default)
Blocknothing. The text field is emptied

[MASKED:…] and [REDACTED:…] say different things on purpose. REDACTED means the value was removed before the prompt or response left the endpoint. MASKED means the value went through as sent (the rule only logs, or the endpoint is in monitor mode) and was kept out of the record. What survives in both cases is the label, the span offset and the detection class, which is enough for the console to show what kind of secret was caught without holding the secret.

How much text a record keeps​

A record is about one turn, not a copy of the conversation. An AI coding tool re-sends the whole transcript with every request; a record keeps the current turn plus only the earlier messages a rule matched, masked, and at most 64 KB of text. That is what the findings file on the endpoint holds.

What the console stores depends on the kind of event, and on the endpoint's daily allowance (see Service limits):

EventText the console stores
A blockNone
Any other findingUp to 16 KB. A longer text keeps a 16 KB window: around the match when it can be located, otherwise the last 16 KB, where the current turn is
A prompt or a clean agent event, within the day's allowanceUp to 8 KB. A longer text keeps the newest 8 KB of a request the agent's proxy captured, and the first 8 KB of anything else
A prompt or a clean agent event, past the day's allowanceNone, or a short excerpt of a tool call's command

Every cut falls on a character boundary and never splits a [MASKED…] or [REDACTED…] placeholder. A custom redaction placeholder that your policy defines is not protected, and a cut can split it. A record that was cut says so: text_truncated, the original size in text_bytes, and text_cut naming the cut (window, clean, excerpt, dropped, which also marks a text field that did not arrive as text, or utility for a client's utility call uploaded without its text). Fields other than the text are bounded too, far above what a normal record carries, and a record where one was shortened carries record_trimmed. The console's detail panel says which cut it was and how much the endpoint sent. See Interactions and findings.

How much of the text is inspected​

Inspection is inline, so the agent spends a budget of 8 KB per interaction (the max_inspect_bytes default) on the detectors whose cost grows fastest with the text. The rest read the whole text, however long it is:

DetectorReads
Credentials, secrets and PII (sensitive_data)the full text
Prompt injection and jailbreak signaturesthe full text
The policy's custom definitionsthe full text
Languagethe first 8 KB
Toxicity and topic classifiersthe first 8 KB
Malicious entities (threat-intel URLs, domains, IPs)the first 8 KB

A verdict whose bounded tier stopped early says so with truncated: true and inspected_bytes. The capturing proxy spends the 8 KB across a whole request, newest user turn first and old history last. When the bounded tier runs out before the end, the request's own record says so, with capture_skip_reason and capture_bytes in its context; agents released before the service limits recorded a separate capture_skipped record instead. The full-text detectors still read every message. Nothing is rejected over the budget.

The shipped policy logs everything it sees​

The policy.yaml the packages install sets every rule to log, and so does Local default, the policy the console assigns every endpoint on enrollment; a new endpoint also starts in monitor mode. That is deliberate, because a threshold cannot be tuned against traffic that was never recorded. The records it writes and uploads hold the prompts and completions, with every detected credential, card number, email and custom-definition match masked.

Three things follow, before any wide rollout:

  • The local findings file is created 0600 and deserves the handling any store of prompts gets: the surrounding text is still the organization's own.
  • Masking covers what a detector matched. Promoting pr_secrets from log to redact, on a copy of Local default in the console or in the local file, is still the change that stops credentials reaching the provider. See From monitor to enforce.
  • A findings file written by an agent release that predates masking can still hold raw values from before the upgrade until it rotates out.

What masking does not reach​

  • Records from agent releases that predate masking. A findings file keeps them until it rotates out, and records already uploaded keep them in the console, which marks such a record as recorded by an agent that predates stored-text masking rather than claiming it was masked.
  • Collectors still running the previous release. The MCP wrapper and the browser's native-messaging host are started by the IDE and by Chrome and keep running the binary they started with. The agent re-masks credentials and custom-definition matches in what they relay, but restart them after an upgrade (reopen the IDE, restart Chrome) so every detector's masking applies.
  • Threat-intel matches deep in a long text. A malicious URL past the inline inspection budget (8 KB) is not detected, so it is not masked either. Credentials and custom definitions are masked across the whole stored text. See How much of the text is inspected.
  • The last line of a private key body. The BEGIN … PRIVATE KEY header and the body lines are masked, but the short final line of a PEM body can fall below the detector's thresholds and be stored as sent.

Where a finding goes​

StageWhat exists there
on the endpointthe findings file, plus a disk spool that survives restarts and offline windows, and inside it a bounded dead-letter bucket for records the control plane could not store
in transitoutbound HTTPS, gzip NDJSON batches drained from that spool
server-sidethe record the console reads, and one copy of each distinct uploaded batch that mirrors what was stored: the text already cut as Service limits describes, and without the events counted past a capacity ceiling (a retried batch is normally not stored twice)

Both server-side copies are encrypted at rest, and every hop that carries an interaction is encrypted in transit. The endpoint keeps its local copy even after the upload succeeds, so removing findings from a machine is a separate act from removing them from the console.

The dead-letter bucket keeps each record in full, like the rest of the spool. Each entry expires 30 days after it was dead-lettered, and the bucket holds at most 1000 records and 32 MB, dropping the oldest first once either bound is reached. Unlike the findings file it is not rotated, so within those 30 days it can outlive the file's own copy of those records. It is part of spool.db and goes with it.

Who can read them​

Records belong to one tenant and are reachable only by that tenant's own members. The boundary is enforced server-side on every request, from the tenant identity in the caller's token rather than from anything the request body says.

Within a tenant there is no narrower boundary. Any member who can open the console can read a record in full, and the search box on the Interactions page matches against the stored interaction text, not only against rule names and hostnames. An organization deploying Lumen across a workforce is therefore granting its console users the ability to read what colleagues typed into a model, and that access is worth deciding on deliberately.

Wazuh operators are a separate matter: the Hub receives activation state and headline counts only, and no operator surface renders the text of an interaction.

What identifies a machine, and what can identify a person​

The agent reports the hostname when it enrolls, and also the operating-system login of the person the installer ran for (never a service account such as root). Each record then carries the OS login of the person whose session produced it, as user.name: the Claude Code hook, the MCP wrapper and the browser's native-messaging host run as that person, and the capturing proxy resolves the account that owns each connection. When no person can be named, the console falls back to the endpoint's enrolled login. It is the OS login name only, never an email address or a directory identity.

Your tenant's own console shows that login on endpoints and interactions, because a finding nobody can be asked about is a finding nobody acts on. No Wazuh operator surface shows it.

Two more paths carry a person indirectly, and both are worth knowing:

  • The coverage report in each heartbeat names the configuration files the agent found, and those paths sit inside home directories. On the machine, the agent keeps the latest report per local account in /var/lib/lumen/coverage.json (readable only by the agent) so a restart does not reset the figure; the account names in that file are never uploaded.
  • The Claude Code hook collector attaches the working directory of the session it inspected.

How long they survive​

Records are deleted after a bounded retention window. The deletion is a hard delete with no archive and no warning, and the copy of the uploaded batch expires on its own schedule. The window is a property of the plan, not a setting anyone changes: which one each plan carries is in Plans and pricing, and the plan in force is on the console's Billing page. On Pro it also depends on the kind of event: findings and prompts are kept for a year and agent activity for 90 days. See Service limits.

Deleting an endpoint does not delete its history. Removing an endpoint revokes that device's credential so it can no longer report, and leaves every interaction it already recorded in place.