Skip to main content

Gateways

Organizations that already route LLM traffic through an API gateway get Lumen as a plugin at that choke point, with no per-machine deployment and full inline enforcement for every application behind the gateway.

GatewayPlugin formStatus
LiteLLMPython CustomLogger callbacksAvailable (co-located)
KongLua pluginPlanned
PortkeyGuardrail webhookPlanned

Each adapter is a thin plugin in the gateway's native extension language that calls Lumen's inline inspection API on the request and response and enforces the verdict in the request path — Log, Redact or Block. All detection and policy live in Lumen; the plugin only serializes, asks, and acts.

LiteLLM​

The LiteLLM plugin (the lumen_litellm Python package) inspects every model call routed through your LiteLLM proxy. Because LiteLLM normalizes every provider it fronts — OpenAI, Anthropic, Bedrock, Azure, and 100+ others — to one shape before its hooks run, a single plugin covers all of them.

It runs against a co-located Lumen agent: the endpoint agent (install it) on the same host, serving its local inspection API on 127.0.0.1:7645. That path needs no gateway API key, adds only single-digit-millisecond latency, and keeps working if the network to the control plane is down. The plugin reuses the agent's engine, findings and uplink, so gateway traffic shows up in the console like any other collector's.

What it inspects​

  • Requests — the new user input (and any tool results being fed back to the model) are inspected before the call leaves. Block returns a 403 in the provider's own error shape so client SDKs fail cleanly; Redact rewrites the prompt in place so a secret never reaches the provider. Earlier conversation history is not re-judged, so an agentic loop never wedges on its own transcript.
  • Responses, streaming and non-streaming both, are inspected before delivery. A streamed completion runs through the same hold-and-release window as every other inline collector: the last stretch of text is held back until Lumen can prove it clear, so a secret spanning two chunks is caught before the user sees either half. Block ends the stream with a content_filter finish; Redact masks the completion.

Install​

The plugin is not published as a package or on the download site yet, so its lumen_litellm directory comes with your access for now, like the console address. Make it importable by your LiteLLM proxy and register it as a callback:

pip install httpx # the plugin's only dependency
# put the directory that holds lumen_litellm/ on the proxy's PYTHONPATH
export PYTHONPATH="/path/to/lumen-litellm:$PYTHONPATH"
# litellm config.yaml
litellm_settings:
callbacks: ["lumen_litellm.LumenGuard"]

The co-located agent's OS service serves the inspection API once the agent is installed. Confirm it is up before starting the proxy:

lumen-agent status # the service answers on 127.0.0.1:7645 by default

Configure​

The plugin reads these environment variables once at startup:

VariableDefaultMeaning
LUMEN_INSPECT_URLhttp://127.0.0.1:7645the co-located agent's local API
LUMEN_COLLECTOR_IDgw-litellmhow this gateway's findings are labelled in the console
LUMEN_TIMEOUT_S0.15hard per-call budget for an inspection
LUMEN_FAIL_OPENtruetrue: a slow or unreachable agent never breaks a call (it passes through as Log). false: fail closed — block when a verdict is unavailable
LUMEN_ENFORCEtruefalse: observe only — still records findings, never blocks or redacts (see monitor to enforce)

fail_open defaults to true: the security layer must never be the reason a tool breaks. Set it false on a route where control matters more than availability.

Current limits​

  • Co-located only. The plugin talks to a local agent; a remote inline API for gateways with no agent on the host is a later addition.
  • The fresh turn is enforced. A secret pasted several turns earlier (and already sent) is not re-scanned on every later request; full-transcript scanning of long histories is a planned refinement.

Kong and Portkey​

Planned, with the same contract: a Lua plugin (access + body_filter) for Kong, and a guardrail webhook for Portkey. Both call the same inline inspection API and map the verdict onto the gateway's native block/redact.

When to choose a gateway collector​

  • You already operate a supported gateway in front of your model traffic.
  • You want one enforcement point for many applications, rather than instrumenting each with the SDK.
  • The traffic you care about is server-side. Developer laptops and browsers still need the host-local collectors.