Skip to main content

Local API

The daemon serves a loopback HTTP API on 127.0.0.1:7645 by default, or on a Unix socket. It is the seam every co-located collector uses: SDKs, the hook guard, wrapped MCP servers, your own scripts.

Health and introspection​

curl -s 127.0.0.1:7645/v1/health | jq # liveness, what installers poll
curl -s 127.0.0.1:7645/v1/policy | jq # which policy id and version is loaded
curl -s 127.0.0.1:7645/v1/status | jq # what `lumen-agent status` reads:
# pid, supervisor, proxy upstreams,
# spool depth and dead-letter count,
# control state, heartbeat

One-shot inspection​

curl -s 127.0.0.1:7645/v1/inspect \
-d '{"stage":"input","text":"the key is AKIAIOSFODNN7EXAMPLE"}'
{
"action": "redact",
"text": "the key is [REDACTED:credentials]",
"spans": [[11, 31, "credentials.aws_access_key"]],
"truncated": false,
"inspected_bytes": 31
}

stage is input for prompt rules or output for response rules. The response carries the effective action, the post-action text, and the span list.

The examples on this page run a policy with pr_secrets promoted to redact. Every policy Lumen ships, the built-in Local default included, only logs, and on it the same request answers "action": "log" with the text unchanged and an empty span list; the detection is still recorded. Enforcing is a deliberate step, covered in From monitor to enforce.

On a connected endpoint that record reaches the console like any other finding. To see what the detectors flag without leaving one, run lumen-agent inspect --text '<text>'. It is a local check, not this daemon's verdict: it reads no agent.yaml, so it checks the text against the policy built into the binary, with no console enforcement ceiling and no threat-intel snapshot, and it records and uploads nothing. With the policy this page uses, the request above answers redact and inspect answers log.

Optional fields let a caller name itself: collector_id labels the source, and collector_type files the finding under a collector surface (for example a co-located gateway sends "llm_gateway") so the console can tell it from host capture. Both are optional; collector_type defaults to endpoint_agent. The same two fields are accepted on /v1/stream/start.

truncated and inspected_bytes qualify the verdict for the bounded detector tier only: language, the toxicity and topic classifiers, and threat-intel entity extraction, which read the first 8 KB of the text (the agent's inline byte budget). Credentials and PII, prompt injection, and the policy's custom definitions always read the full text, so a truncated: true verdict is still complete for them; it means the bounded tier stopped at inspected_bytes. The split is in What Lumen stores.

Streaming: hold and release​

A secret split across streaming chunks matches nothing on its own. The stream API accumulates deltas, forwards only the provably clean prefix, and holds the tail until it can be judged:

SID=$(curl -s 127.0.0.1:7645/v1/stream/start -d '{"stage":"output"}' | jq -r .stream_id)
curl -s 127.0.0.1:7645/v1/stream/chunk -d "{\"stream_id\":\"$SID\",\"text\":\"The backup key is AKIAIOSFO\"}"
curl -s 127.0.0.1:7645/v1/stream/chunk -d "{\"stream_id\":\"$SID\",\"text\":\"DNN7EXAMPLE, keep it safe.\"}"
curl -s 127.0.0.1:7645/v1/stream/end -d "{\"stream_id\":\"$SID\"}"

Each chunk call returns what may be delivered now and how much is held. end flushes the remainder post-action: with pr_secrets at redact the user never sees a character of the secret, and at log (the built-in) what is delivered is the text as sent. See Sensitive data, redacted in place for what a caught secret looks like in the console.

note

The capturing proxy uses this same hold-and-release window internally for SSE responses, so clients keep their streaming UX while the output is still enforced.