Gateways
Organizations that already route LLM traffic through an API gateway get Lumen as a plugin at that choke point, with no per-machine deployment and full inline enforcement for every application behind the gateway.
| Gateway | Plugin form | Status |
|---|---|---|
| LiteLLM | Python CustomLogger callbacks | Available (co-located) |
| Kong | Lua plugin | Planned |
| Portkey | Guardrail webhook | Planned |
Each adapter is a thin plugin in the gateway's native extension language that calls Lumen's inline inspection API on the request and response and enforces the verdict in the request path — Log, Redact or Block. All detection and policy live in Lumen; the plugin only serializes, asks, and acts.
LiteLLM
The LiteLLM plugin (the lumen_litellm Python package) inspects every model
call routed through your LiteLLM proxy. Because
LiteLLM normalizes every provider it fronts — OpenAI, Anthropic, Bedrock, Azure,
and 100+ others — to one shape before its hooks run, a single plugin covers all
of them.
It runs against a co-located Lumen agent: the endpoint agent
(install it) on the same host, serving its local
inspection API on 127.0.0.1:7645. That path needs no gateway API key, adds only
single-digit-millisecond latency, and keeps working if the network to the control
plane is down. The plugin reuses the agent's engine, findings and uplink, so
gateway traffic shows up in the console like any other collector's.
What it inspects
- Requests — the new user input (and any tool results being fed back to the
model) are inspected before the call leaves. Block returns a
403in the provider's own error shape so client SDKs fail cleanly; Redact rewrites the prompt in place so a secret never reaches the provider. Earlier conversation history is not re-judged, so an agentic loop never wedges on its own transcript. - Responses, streaming and non-streaming both, are inspected before delivery.
A streamed completion runs through the same hold-and-release window as every
other inline collector: the last stretch of text is held back until Lumen can
prove it clear, so a secret spanning two chunks is caught before the user sees
either half. Block ends the stream with a
content_filterfinish; Redact masks the completion.
Install
The plugin is not published as a package or on the download site yet, so its
lumen_litellm directory comes with your access for now, like the console
address. Make it importable by your LiteLLM proxy and register it as a callback:
pip install httpx # the plugin's only dependency
# put the directory that holds lumen_litellm/ on the proxy's PYTHONPATH
export PYTHONPATH="/path/to/lumen-litellm:$PYTHONPATH"
# litellm config.yaml
litellm_settings:
callbacks: ["lumen_litellm.LumenGuard"]
The co-located agent's OS service serves the inspection API once the agent is installed. Confirm it is up before starting the proxy:
lumen-agent status # the service answers on 127.0.0.1:7645 by default
Configure
The plugin reads these environment variables once at startup:
| Variable | Default | Meaning |
|---|---|---|
LUMEN_INSPECT_URL | http://127.0.0.1:7645 | the co-located agent's local API |
LUMEN_COLLECTOR_ID | gw-litellm | how this gateway's findings are labelled in the console |
LUMEN_TIMEOUT_S | 0.15 | hard per-call budget for an inspection |
LUMEN_FAIL_OPEN | true | true: a slow or unreachable agent never breaks a call (it passes through as Log). false: fail closed — block when a verdict is unavailable |
LUMEN_ENFORCE | true | false: observe only — still records findings, never blocks or redacts (see monitor to enforce) |
fail_open defaults to true: the security layer must never be the reason a tool
breaks. Set it false on a route where control matters more than availability.
Current limits
- Co-located only. The plugin talks to a local agent; a remote inline API for gateways with no agent on the host is a later addition.
- The fresh turn is enforced. A secret pasted several turns earlier (and already sent) is not re-scanned on every later request; full-transcript scanning of long histories is a planned refinement.
Kong and Portkey
Planned, with the same contract: a Lua plugin (access + body_filter) for Kong,
and a guardrail webhook for Portkey. Both call the same inline inspection API and
map the verdict onto the gateway's native block/redact.
When to choose a gateway collector
- You already operate a supported gateway in front of your model traffic.
- You want one enforcement point for many applications, rather than instrumenting each with the SDK.
- The traffic you care about is server-side. Developer laptops and browsers still need the host-local collectors.