The deployment
Each customer gets a dedicated model deployment. It runs on isolated GPUs, on hardware we own, in a US datacenter. No other customer’s workload runs on those GPUs.
We tune the serving stack to your workload and maintain it against that workload. You call the endpoint. We run the hardware and the serving stack underneath it.
The agreement
Every deployment comes with a signed data processing agreement. Where HIPAA applies, we also sign a business associate agreement.
Retention is zero, and the agreement says so. Prompts are not logged. Customer data is not used for training. We use no foreign subprocessors.
The engineer
A named US engineer is assigned to your deployment. They answer the phone.
Model versions
Model versions are pinned. A pinned version does not change underneath the workload. When an examiner asks which model produced a decision in March, you can name the version.
Capacity
Capacity is reserved rather than requested. The GPUs are held for your deployment, not allocated on demand. Reserved GPUs and pinned versions keep costs predictable.
Who it is for
We sell to hospitals, banks, law firms, defense contractors and federal agencies.