Local API
The daemon serves a loopback HTTP API on 127.0.0.1:7645 by default, or on
a Unix socket. It is the seam every co-located collector uses: SDKs, the hook
guard, wrapped MCP servers, your own scripts.
Health and introspection
curl -s 127.0.0.1:7645/v1/health | jq # liveness, what installers poll
curl -s 127.0.0.1:7645/v1/policy | jq # which policy id and version is loaded
curl -s 127.0.0.1:7645/v1/status | jq # what `lumen-agent status` reads:
# pid, supervisor, proxy upstreams,
# spool depth and dead-letter count,
# control state, heartbeat
One-shot inspection
curl -s 127.0.0.1:7645/v1/inspect \
-d '{"stage":"input","text":"the key is AKIAIOSFODNN7EXAMPLE"}'
{
"action": "redact",
"text": "the key is [REDACTED:credentials]",
"spans": [[11, 31, "credentials.aws_access_key"]],
"truncated": false,
"inspected_bytes": 31
}
stage is input for prompt rules or output for response rules. The response
carries the effective action, the post-action text, and the span list.
The examples on this page run a policy with pr_secrets promoted to redact.
Every policy Lumen ships, the built-in Local default included, only logs, and
on it the same request answers "action": "log" with the text unchanged and an
empty span list; the detection is still recorded. Enforcing is a deliberate
step, covered in From monitor to enforce.
On a connected endpoint that record reaches the console like any other
finding. To see what the detectors flag without leaving one, run
lumen-agent inspect --text '<text>'. It is a local check, not this daemon's
verdict: it reads no agent.yaml, so it checks the text against the policy
built into the binary, with no console enforcement ceiling and no threat-intel
snapshot, and it records and uploads nothing. With the policy this page uses,
the request above answers redact and inspect answers log.
Optional fields let a caller name itself: collector_id labels the source, and
collector_type files the finding under a collector surface (for example a
co-located gateway sends "llm_gateway") so the console can tell it from host
capture. Both are optional; collector_type defaults to endpoint_agent. The
same two fields are accepted on /v1/stream/start.
truncated and inspected_bytes qualify the verdict for the bounded
detector tier only: language, the toxicity and topic classifiers, and
threat-intel entity extraction, which read the first 8 KB of the text (the
agent's inline byte budget). Credentials and PII, prompt injection, and the
policy's custom definitions always read the full text, so a truncated: true
verdict is still complete for them; it means the bounded tier stopped at
inspected_bytes. The split is in
What Lumen stores.
Streaming: hold and release
A secret split across streaming chunks matches nothing on its own. The stream API accumulates deltas, forwards only the provably clean prefix, and holds the tail until it can be judged:
SID=$(curl -s 127.0.0.1:7645/v1/stream/start -d '{"stage":"output"}' | jq -r .stream_id)
curl -s 127.0.0.1:7645/v1/stream/chunk -d "{\"stream_id\":\"$SID\",\"text\":\"The backup key is AKIAIOSFO\"}"
curl -s 127.0.0.1:7645/v1/stream/chunk -d "{\"stream_id\":\"$SID\",\"text\":\"DNN7EXAMPLE, keep it safe.\"}"
curl -s 127.0.0.1:7645/v1/stream/end -d "{\"stream_id\":\"$SID\"}"
Each chunk call returns what may be delivered now and how much is
held. end flushes the remainder post-action: with pr_secrets at
redact the user never sees a character of the secret, and at log (the
built-in) what is delivered is the text as sent. See
Sensitive data, redacted in place
for what a caught secret looks like in the console.
The capturing proxy uses this same hold-and-release window internally for SSE responses, so clients keep their streaming UX while the output is still enforced.