Skip to content

Platform

A private inference endpoint

A dedicated model deployment on isolated GPUs, tuned per workload, with a signed data processing agreement, zero retention and a named US engineer.

The deployment

Each customer gets a dedicated model deployment. It runs on isolated GPUs, on hardware we own, in a US datacenter. No other customer’s workload runs on those GPUs.

We tune the serving stack to your workload and maintain it against that workload. You call the endpoint. We run the hardware and the serving stack underneath it.

The agreement

Every deployment comes with a signed data processing agreement. Where HIPAA applies, we also sign a business associate agreement.

Retention is zero, and the agreement says so. Prompts are not logged. Customer data is not used for training. We use no foreign subprocessors.

The engineer

A named US engineer is assigned to your deployment. They answer the phone.

Model versions

Model versions are pinned. A pinned version does not change underneath the workload. When an examiner asks which model produced a decision in March, you can name the version.

Capacity

Capacity is reserved rather than requested. The GPUs are held for your deployment, not allocated on demand. Reserved GPUs and pinned versions keep costs predictable.

Who it is for

We sell to hospitals, banks, law firms, defense contractors and federal agencies.

Talk to us