
Preview
Fast, lightweight model for text and vision tasks.

Duce
High-performance multimodal with SOTA music generation and Mixture of Models (MoM).

Magma
Maximum capability with 1M context, SOTA video, and Mixture of Models (MoM).
Preview
The fastest model layer, optimized for text and vision tasks.
Capabilities:
- Function calling
- Structured output
- Reasoning
Duce
High-performance multimodal model with state-of-the-art music generation and broad media capabilities.
Capabilities:
- SOTA music generation
- Function calling
- Reasoning
- Mixture of Models (MoM)
- Text-to-Speech (TTS)
- Image-to-Image
- Text/Image-to-Video
- Video-to-Video
- Text-to-Music / Music-to-Music
Magma
The most capable model layer with the largest context window and state-of-the-art video generation.
Capabilities:
- SOTA video generation
- Function calling
- Reasoning
- Mixture of Models (MoM)
- Text-to-Speech (TTS)
- Image-to-Image
- Text/Image-to-Video
- Video-to-Video
- Text-to-Music / Music-to-Music