Skip to content

Three inference tiers. One endpoint. Zero config.

Three purpose-built tiers. One unified OpenAI-compatible endpoint. Set dms-auto-router as your model and the platform selects the strongest available model for your workload — your endpoint improves without a single code change.

Automatic model routingdms-auto-router

One model ID. The strongest available model for each workload. No integration changes as the platform improves.

lite-64k

Real-Time Fast

64,000token context window

Low-latency inference for customer-facing chat, high-volume classification, and real-time AI agents.

Use cases

  • Customer-facing chat and support agents
  • High-volume classification and routing
  • Interactive real-time AI assistants
  • Embedding preprocessing and lightweight tasks

Fastest model meeting quality threshold, prioritized by latency.

pro-256kMost popular

Code & Precision

256,000token context window

Structured generation accuracy, multi-file coherence, and instruction fidelity for precision-sensitive workloads.

Use cases

  • Multi-file code generation and refactoring
  • Complex enterprise logic with high accuracy requirements
  • Precision-sensitive content generation
  • Instruction-heavy structured outputs

Top-performing code model, selected by benchmark performance.

max-1mFlagship

Flagship Reasoning

1,000,000token context window

Long-context reasoning, multi-document synthesis, and multi-step agent orchestration at scale.

Use cases

  • Full-codebase analysis in a single pass
  • Legal document review across thousands of pages
  • Multi-step agent workflows with extended history
  • Deep research and synthesis

Best-in-class reasoning model, continuously evaluated and deployed.