Inference
Route workloads across DMS model tiers while keeping an OpenAI-compatible integration.
/v1/chat/completionsA familiar API surface backed by routing, private context, tool-aware models, and deployment controls.
Route workloads across DMS model tiers while keeping an OpenAI-compatible integration.
/v1/chat/completionsGround model responses with organization context and deployment boundaries designed around your requirements.
context.source = privateUse structured output and tool calls to move from an answer to the next operation.
finish_reason = tool_callsScope routing, capacity, data flow, and service terms for the production environment.
deployment = managed | privateOne OpenAI-compatible endpoint routes each workload to the model tier built for it.
Model selection and streamingYour knowledge layer grounds the request without turning private data into training material.
Retrieval and context assemblyTool-aware models choose the next operation, return structured output, and continue the workflow.
Tool use and orchestrationPolicies, deployment boundaries, and audit signals keep production workloads observable.
Policy and audit controls