preview

Preview

Fast, lightweight model for text and vision tasks.
duce

Duce

High-performance multimodal with SOTA music generation and Mixture of Models (MoM).
magma

Magma

Maximum capability with 1M context, SOTA video, and Mixture of Models (MoM).

Preview

The fastest model layer, optimized for text and vision tasks. Capabilities:
  • Function calling
  • Structured output
  • Reasoning
Best for: Quick text completions, image analysis, structured data extraction, and function calling where speed matters most.

Duce

High-performance multimodal model with state-of-the-art music generation and broad media capabilities. Capabilities:
  • SOTA music generation
  • Function calling
  • Reasoning
  • Mixture of Models (MoM)
  • Text-to-Speech (TTS)
  • Image-to-Image
  • Text/Image-to-Video
  • Video-to-Video
  • Text-to-Music / Music-to-Music
Best for: Music generation, multimodal content creation, and complex reasoning tasks requiring media output.

Magma

The most capable model layer with the largest context window and state-of-the-art video generation. Capabilities:
  • SOTA video generation
  • Function calling
  • Reasoning
  • Mixture of Models (MoM)
  • Text-to-Speech (TTS)
  • Image-to-Image
  • Text/Image-to-Video
  • Video-to-Video
  • Text-to-Music / Music-to-Music
Best for: Long-context tasks, video generation, complex agentic workflows, and maximum quality output across all modalities.

Comparison


Mixture of Models (MoM)

Duce and Magma support Mixture of Models (MoM) — a technique that routes requests across multiple specialized models within the layer to produce the best output. MoM is automatically applied when using these model layers; no additional configuration is required.