EndoRouter Open source · 0.1

Decide where a prompt may go
before deciding which model is best.

Most model routers send everything to the cloud and try to catch the sensitive requests on the way out. Whatever the detectors do not recognise leaves the building. EndoRouter fails closed: work stays on your local model unless its provenance says it is public, and every send is written to an audit log before a byte leaves.

EndoRouter on GitHub On PyPI

How it works

One label per request, and labels only ever tighten.

Every request is scanned, labelled public, unknown or private, decided, audited, then dispatched. Provenance comes from the client as headers naming the source of the text; only clients you list may call anything public, and any client may call it private. Structural detectors look for the shapes of secrets, keys, cards and personal data in every forwarded field, history and tool calls included. Anything unlabelled stays local. In the optional balanced mode a local classifier may clear unlabelled work, and the audit log says when it did.

It speaks the OpenAI chat completions API, so any client that can set a base URL can use it. Setup asks no questions: it finds the local model server already running, verifies the program behind the port, adds cloud providers whose keys are already set, and protects standard secret files.

Enforced in code, pinned by tests

Private never reaches the cloud

In strict mode a private or unknown request never selects a cloud target. Asking for a cloud model by name is refused, not honoured, and a local model that is down fails the request rather than falling back.

Audit before send

A flushed record names the destination before each send. If the record cannot be written, nothing is sent. The log holds decisions and reasons, never prompt content.

Loopback only

It listens on the local machine, answers only requests addressed to localhost without a browser origin, follows no redirects, and ignores proxy environment variables.

leakbench

The benchmark that tests the claim ships with it.

leakbench measures whether private data reaches a cloud through any OpenAI-compatible gateway. It stands in for the cloud and the local model with two recording servers, sends each case as written, and counts a leak when any private string of a case arrives at the cloud sink, at any time, in any message. On the main suite of 24 private cases EndoRouter leaked none in strict or balanced mode. On the boundary suite, strict mode kept all 7 cases in formats the detectors cover off the cloud and, as documented, none of the 5 known misses; balanced mode leaked none of either. The cases, the methods and every result are in the repository, with the threat model stated: it measures gateways that were not built to fool it.

Read the results →

Start

Two commands. No questions.

Python 3.10 or newer on macOS, Linux or Windows, and a local model server already running, such as Ollama.

pip install endorouter
endorouter serve

Then point your client at http://127.0.0.1:8787/v1 and ask for the model auto. Version 0.1 handles text chat completions, with and without streaming; images, audio, embeddings and the Responses API are refused rather than passed through. Apache-2.0.