An institution can keep a model on its own server and still have no defensible answer to a simple question: who was allowed to ask it what? If users can reach the underlying model process directly, a policy gateway is only a suggested route. If a response is filtered after streaming has already released the sensitive fragment, the filter has arrived too late. Cordon was built around those two facts. It owns the inference runtime and makes the controlled path the only network path.
Every request needs a client identity, model permission, token budget, audit record written before computation, output policy applied before release, and a verifiable response. Cordon starts llama-server as a child on an ephemeral loopback port, checks that its web UI is unavailable, and gives it a per-boot API key through an owner-only file. Another exposed runtime port would undo the controls in front of it.

From a local node to a deployment tool#
Cordon 2.1 extends the desktop application beyond a local model launcher. An operator can configure Light, Sovereign Cloud, Vault, Island or Dark mode in the same window used to inspect the node. Each mode shows its prerequisites and refuses to switch until the machine meets them. The application can create and back up a Client Master Key, seal a model bundle, verify every shard, and import or export the bundle for another node.
Remote access uses a private certificate authority and mutual TLS. Client certificates are pinned by fingerprint in a registry that denies unknown clients; the management console stays on loopback. This is a practical change in the deployment path: the operator no longer has to assemble separate commands for every routine key and certificate step.
Stronger modes require key material, a sealed model, mutual TLS and, where claimed, hardware measurements. Dark mode adds no-egress and custody requirements. Missing prerequisites prevent the switch; node settings mirror command-line mode defaults.
Bugs encountered at the boundary#
The release also fixes two failures that only appear when the model store becomes real. A bundle identifier could previously steer decrypted staging files outside their directory, so manifest identifiers now have a restricted character set. Adding any bundle to a store could also block requests to a node serving an unrelated plain model; the gate now checks the bundle actually loaded by the runtime. Serving a sealed bundle probes the bundled llama.cpp for a non-mapped load option and refuses to proceed if it has neither supported form.
Sealed weights are briefly plaintext during load. Cordon decrypts shards into a restricted staging file, checks the digest, starts the runtime without memory mapping, then removes the file. This narrows exposure but is not enclave-resident decryption; tmpfs is recommended where disk exposure matters. The changed llama.cpp loading flag was therefore a release blocker.
Streaming has its own boundary: a sensitive pattern may finish in the next chunk. Cordon holds back a trailing window and scans before release. It refuses streaming when timing normalization is enabled because chunk intervals would disclose that timing signal; ordinary inference remains available.
Hardware attestation has synthetic verification coverage, but the documented AMD path had not yet run on real silicon. Confidentiality from a hostile host administrator depends on the selected and validated hardware deployment; the desktop cannot create a hardware root of trust by itself.
Next comes deployment-specific evidence: target hardware measurements, staging location, certificate custody, and offline verification of signed responses and the audit chain. Read the Cordon system record for the request path and deployment postures.