What Lumen stores
Lumen inspects the text of prompts and responses, so a deployment holds the most sensitive corpus an organization has. What reaches disk is decided on the endpoint, before anything is written or uploaded, and a value a detector matched is never part of it, whatever the policy did about it. The credential scan runs for storage even when no rule in the policy asks for it. The few cases masking does not reach are listed under What masking does not reach.
What lands in a finding
Every inspection appends one record. Its text field holds the interaction with every matched value masked; the action decides only which placeholder stands in for it:
| Action | Stored text |
|---|---|
| Log (including monitor mode) | the interaction as it was seen, each matched value replaced by [MASKED:<label>] |
| Redact | the rewritten text, each redacted value replaced by its placeholder ([REDACTED:<class>] by default) |
| Block | nothing. The text field is emptied |
[MASKED:…] and [REDACTED:…] say different things on purpose. REDACTED means
the value was removed before the prompt or response left the endpoint. MASKED
means the value went through as sent (the rule only logs, or the endpoint is in
monitor mode) and was kept out of the record. What survives in both cases is the
label, the span offset and the detection class, which is enough for the console
to show what kind of secret was caught without holding the secret.
How much text a record keeps
A record is about one turn, not a copy of the conversation. An AI coding tool re-sends the whole transcript with every request; a record keeps the current turn plus only the earlier messages a rule matched, masked, and at most 64 KB of text. That is what the findings file on the endpoint holds.
What the console stores depends on the kind of event, and on the endpoint's daily allowance (see Service limits):
| Event | Text the console stores |
|---|---|
| A block | None |
| Any other finding | Up to 16 KB. A longer text keeps a 16 KB window: around the match when it can be located, otherwise the last 16 KB, where the current turn is |
| A prompt or a clean agent event, within the day's allowance | Up to 8 KB. A longer text keeps the newest 8 KB of a request the agent's proxy captured, and the first 8 KB of anything else |
| A prompt or a clean agent event, past the day's allowance | None, or a short excerpt of a tool call's command |
Every cut falls on a character boundary and never splits a [MASKED…] or
[REDACTED…] placeholder. A custom redaction placeholder that your policy
defines is not protected, and a cut can split it. A record that was cut says so: text_truncated, the
original size in text_bytes, and text_cut naming the cut (window, clean,
excerpt, dropped, which also marks a text field that did not arrive as text,
or utility for a client's utility call uploaded without its text). Fields other
than the text are bounded too, far above what a normal record carries, and a
record where one was shortened carries record_trimmed. The console's detail panel says which cut it was and how much the
endpoint sent. See Interactions and findings.
How much of the text is inspected
Inspection is inline, so the agent spends a budget of 8 KB per interaction
(the max_inspect_bytes default) on the detectors whose cost grows fastest
with the text. The rest read the whole text, however long it is:
| Detector | Reads |
|---|---|
Credentials, secrets and PII (sensitive_data) | the full text |
| Prompt injection and jailbreak signatures | the full text |
| The policy's custom definitions | the full text |
| Language | the first 8 KB |
| Toxicity and topic classifiers | the first 8 KB |
| Malicious entities (threat-intel URLs, domains, IPs) | the first 8 KB |
A verdict whose bounded tier stopped early says so with truncated: true and
inspected_bytes. The capturing proxy spends the 8 KB across a whole request,
newest user turn first and old history last. When the bounded tier runs out
before the end, the request's own record says so, with capture_skip_reason
and capture_bytes in its context; agents released before the
service limits recorded a separate capture_skipped record instead.
The full-text detectors still read every message. Nothing is rejected over the
budget.
The shipped policy logs everything it sees
The policy.yaml the packages install sets every rule to log, and so does
Local default, the policy the console assigns every endpoint on enrollment; a
new endpoint also starts in monitor mode. That is deliberate, because a
threshold cannot be tuned against traffic that was never recorded. The records it writes and
uploads hold the prompts and completions, with every detected credential,
card number, email and custom-definition match masked.
Three things follow, before any wide rollout:
- The local findings file is created
0600and deserves the handling any store of prompts gets: the surrounding text is still the organization's own. - Masking covers what a detector matched. Promoting
pr_secretsfromlogtoredact, on a copy of Local default in the console or in the local file, is still the change that stops credentials reaching the provider. See From monitor to enforce. - A findings file written by an agent release that predates masking can still hold raw values from before the upgrade until it rotates out.
What masking does not reach
- Records from agent releases that predate masking. A findings file keeps them until it rotates out, and records already uploaded keep them in the console, which marks such a record as recorded by an agent that predates stored-text masking rather than claiming it was masked.
- Collectors still running the previous release. The MCP wrapper and the browser's native-messaging host are started by the IDE and by Chrome and keep running the binary they started with. The agent re-masks credentials and custom-definition matches in what they relay, but restart them after an upgrade (reopen the IDE, restart Chrome) so every detector's masking applies.
- Threat-intel matches deep in a long text. A malicious URL past the inline inspection budget (8 KB) is not detected, so it is not masked either. Credentials and custom definitions are masked across the whole stored text. See How much of the text is inspected.
- The last line of a private key body. The
BEGIN … PRIVATE KEYheader and the body lines are masked, but the short final line of a PEM body can fall below the detector's thresholds and be stored as sent.
Where a finding goes
| Stage | What exists there |
|---|---|
| on the endpoint | the findings file, plus a disk spool that survives restarts and offline windows, and inside it a bounded dead-letter bucket for records the control plane could not store |
| in transit | outbound HTTPS, gzip NDJSON batches drained from that spool |
| server-side | the record the console reads, and one copy of each distinct uploaded batch that mirrors what was stored: the text already cut as Service limits describes, and without the events counted past a capacity ceiling (a retried batch is normally not stored twice) |
Both server-side copies are encrypted at rest, and every hop that carries an interaction is encrypted in transit. The endpoint keeps its local copy even after the upload succeeds, so removing findings from a machine is a separate act from removing them from the console.
The dead-letter bucket keeps each record in full, like the rest of the spool.
Each entry expires 30 days after it was dead-lettered, and the bucket holds at
most 1000 records and 32 MB, dropping the oldest first once either bound is
reached. Unlike the findings file it is not rotated, so within those 30 days it
can outlive the file's own copy of those records. It is part of spool.db and
goes with it.
Who can read them
Records belong to one tenant and are reachable only by that tenant's own members. The boundary is enforced server-side on every request, from the tenant identity in the caller's token rather than from anything the request body says.
Within a tenant there is no narrower boundary. Any member who can open the console can read a record in full, and the search box on the Interactions page matches against the stored interaction text, not only against rule names and hostnames. An organization deploying Lumen across a workforce is therefore granting its console users the ability to read what colleagues typed into a model, and that access is worth deciding on deliberately.
Wazuh operators are a separate matter: the Hub receives activation state and headline counts only, and no operator surface renders the text of an interaction.
What identifies a machine, and what can identify a person
The agent reports the hostname when it enrolls, and also the
operating-system login of the person the installer ran for (never a service
account such as root). Each record then carries the OS login of the person
whose session produced it, as user.name: the Claude Code hook, the MCP
wrapper and the browser's native-messaging host run as that person, and the
capturing proxy resolves the account that owns each connection. When no person
can be named, the console falls back to the endpoint's enrolled login. It is
the OS login name only, never an email address or a directory identity.
Your tenant's own console shows that login on endpoints and interactions, because a finding nobody can be asked about is a finding nobody acts on. No Wazuh operator surface shows it.
Two more paths carry a person indirectly, and both are worth knowing:
- The coverage report in each heartbeat names the configuration files the agent
found, and those paths sit inside home directories. On the machine, the agent
keeps the latest report per local account in
/var/lib/lumen/coverage.json(readable only by the agent) so a restart does not reset the figure; the account names in that file are never uploaded. - The Claude Code hook collector attaches the working directory of the session it inspected.
How long they survive
Records are deleted after a bounded retention window. The deletion is a hard delete with no archive and no warning, and the copy of the uploaded batch expires on its own schedule. The window is a property of the plan, not a setting anyone changes: which one each plan carries is in Plans and pricing, and the plan in force is on the console's Billing page. On Pro it also depends on the kind of event: findings and prompts are kept for a year and agent activity for 90 days. See Service limits.
Deleting an endpoint does not delete its history. Removing an endpoint revokes that device's credential so it can no longer report, and leaves every interaction it already recorded in place.