Skip to content

Three inference tiers. One path to production.

Lite, Pro, and Max map directly to the workload profiles on the Models page. Weekly prices stay synced with the existing DMS API; enterprise capacity and private deployment are scoped separately.

Lite

lite-64k · Real-Time Fast

Low-latency inference for customer-facing chat and real-time agents.

$2.50/wk64K
  • Customer-facing chat and support
  • High-volume classification and routing
  • Latency-prioritized model selection
Choose plan

Pro

pro-256k · Code & Precision

Structured generation accuracy and multi-file coherence for precision-sensitive work.

$6/wk256K
  • Multi-file generation and refactoring
  • Instruction-heavy structured outputs
  • Benchmark-led model selection
Choose plan

Max

max-1m · Flagship Reasoning

Long-context reasoning and multi-step agent orchestration at scale.

$25/wk1M
  • Full-codebase and document analysis
  • Extended-history agent workflows
  • Availability shown in dashboard
Choose plan