lite-64kReal-Time Fast
Low-latency inference for customer-facing chat, high-volume classification, and real-time AI agents.
Use cases
- Customer-facing chat and support agents
- High-volume classification and routing
- Interactive real-time AI assistants
- Embedding preprocessing and lightweight tasks
Fastest model meeting quality threshold, prioritized by latency.
