Skip to content

OpenAI-Compatible Proxy

The guardrail proxy fronts any OpenAI-compatible API and scores responses for hallucination before they reach the caller. Point a client at it with OPENAI_BASE_URL and no code changes:

director-ai proxy --port 8080 --facts kb.txt --threshold 0.6
export OPENAI_BASE_URL=http://localhost:8080/v1

The proxy binds to 127.0.0.1 by default. To expose it beyond the local machine (LAN, container), opt in explicitly with --host 0.0.0.0 or the DIRECTOR_SERVER_HOST environment variable — and put authentication (--api-keys) in front of any non-loopback bind.

Build it programmatically with create_proxy_app (reference below):

from director_ai.proxy import create_proxy_app

app = create_proxy_app(
    threshold=0.6,
    facts_path="kb.txt",
    upstream_url="https://api.openai.com",
    on_fail="reject",          # or "warn"
    moderations="local",       # or "upstream"
)

Routes

Route Behaviour
POST /v1/chat/completions Forwarded upstream; the assistant message (or the accumulated stream) is scored. on_fail="reject" returns HTTP 422 (content_filter) on hallucination; streams are halted mid-flight with a content_filter finish reason.
POST /v1/completions Legacy text completions, same scoring flow as chat — non-streaming and streaming (choices[0].text deltas).
POST /v1/moderations moderations="local" (default) analyses every input with the shipped dependency-free detectors and answers in the OpenAI moderations shape; "upstream" forwards the request verbatim.
POST /v1/embeddings Plain passthrough — embeddings carry no natural-language claims to verify.
GET /v1/models Plain passthrough.
GET /health Proxy status, threshold, and failure mode.

Scored responses carry X-Director-Score and X-Director-Approved headers; rejected ones return the OpenAI error shape with "type": "content_filter".

Local moderations

Local mode needs no upstream moderations endpoint — it works in front of vLLM, llama.cpp, or any self-hosted gateway. Each input is analysed by KeywordToxicityDetector (word-boundary seed list plus attack patterns) and RegexPIIDetector (email, phone, credit card, SSN, PHI, IBAN, passport, IPv4). The response follows the OpenAI shape with Director's category names:

{
  "id": "modr-…",
  "model": "director-ai-local-moderation",
  "results": [
    {
      "flagged": true,
      "categories": {"email": true},
      "category_scores": {"email": 1.0}
    }
  ]
}

input accepts a string or a non-empty list of strings; anything else returns HTTP 400 in the OpenAI error shape. Category names are Director's own (keyword, threat, self_harm_encouragement, plus the PII categories above) — clients that only read flagged need no changes.

CLI flags

Flag Default Meaning
--port 8080 Listen port.
--threshold 0.6 Coherence threshold.
--facts / --facts-root Ground-truth facts file (and its allowed root).
--upstream-url https://api.openai.com Upstream base URL (HTTPS enforced unless --allow-http-upstream).
--on-fail reject reject (HTTP 422) or warn (forward with headers).
--api-keys Comma-separated keys; clients must send X-API-Key.
--moderations local local or upstream.
--audit-db SQLite compliance audit database path.
--config-env off Build the scorer from DIRECTOR_* environment configuration.

API reference

director_ai.proxy.create_proxy_app

create_proxy_app(threshold: float = 0.6, facts_path: str | None = None, facts_root: str | None = None, upstream_url: str = 'https://api.openai.com', on_fail: str = 'reject', use_nli: bool | None = None, api_keys: list[str] | None = None, allow_http_upstream: bool = False, audit_db: str | None = None, config: DirectorConfig | None = None, moderations: str = 'local', stream_disclosure: str = 'immediate', _transport: Any = None) -> FastAPI

Build a FastAPI app that proxies OpenAI requests with scoring.

Parameters:

Name Type Description Default
threshold float

Coherence threshold below which responses are flagged.

0.6
facts_path str | None

Path to a key: value facts file (one per line).

None
facts_root str | None

Allowed root directory for facts_path. When set, the resolved facts_path (with symlinks followed) must lie inside facts_root; otherwise :class:ValueError is raised. Leave None for CLI/operator use; set in production deployments where facts_path is derived from untrusted configuration.

None
upstream_url str

Base URL of the upstream OpenAI-compatible API.

'https://api.openai.com'
on_fail str

"reject" returns 422 on hallucination. "warn" forwards the response with warning headers.

'reject'
use_nli bool | None

Enable NLI model. None auto-detects.

None
api_keys list[str] | None

Required API keys. Clients must send X-API-Key header. None or empty = no auth (not recommended for production).

None
allow_http_upstream bool

Allow non-HTTPS upstream URLs. Default False rejects them.

False
audit_db str | None

Path to SQLite compliance audit database. None disables audit logging.

None
config DirectorConfig | None

Optional DirectorConfig. When provided, the proxy builds the configured store and scorer instead of the minimal in-memory scorer.

None
moderations str

"local" serves /v1/moderations from the shipped dependency-free detectors; "upstream" forwards the request to the upstream endpoint verbatim.

'local'
stream_disclosure str

"immediate" (default) forwards every streamed chunk as it arrives; a mid-stream halt stops FUTURE tokens only, so content emitted before the halt has already reached the client — early termination with partial disclosure. "buffered" withholds chunks until they pass a review and discards the unreleased buffer on a halt, so a rejected stream discloses nothing unreviewed (adds up to STREAM_CHECK_INTERVAL chunks of latency; meaningful with on_fail="reject").

'immediate'