Lite
lite-64k · Real-Time FastLow-latency inference for customer-facing chat and real-time agents.
$2.50/wk64K- Customer-facing chat and support
- High-volume classification and routing
- Latency-prioritized model selection
Lite, Pro, and Max map directly to the workload profiles on the Models page. Weekly prices stay synced with the existing DMS API; enterprise capacity and private deployment are scoped separately.
Low-latency inference for customer-facing chat and real-time agents.
$2.50/wk64KStructured generation accuracy and multi-file coherence for precision-sensitive work.
$6/wk256KLong-context reasoning and multi-step agent orchestration at scale.
$25/wk1M