Skip to main content

1. Model Introduction

NVIDIA Cosmos3 is an omnimodal world-model family spanning text/image/video generation, optional synchronized sound, and robot action prediction. Its main advantage is breadth: the same native SGLang pipeline can serve media-generation checkpoints and the DROID policy checkpoint without routing through an LLM sampler. Choose Nano for the broadest modality coverage and lower deployment cost, Super for the larger 64B image/video model, and a specialized checkpoint when only T2I or I2V is needed. Sound and action are checkpoint-specific heads, so they are not available from every Cosmos3 repository. Sound and action generation require the corresponding checkpoint heads. The pipeline reads the transformer and scheduler configs at startup, so Edge and distilled checkpoints do not require architecture-specific server flags. Non-distilled checkpoints use the flow-native FlowUniPCMultistepScheduler; distilled checkpoints use the fixed sigma schedule stored in the checkpoint. The default flow_shift is 3.0 for T2I, 10.0 for non-Edge video and all action modes, and 3.0 for Edge video modes. Distilled checkpoints bake the schedule into their sigmas and do not use a request-level flow_shift.

2. Installation

Install SGLang with the diffusion dependencies:
Command
Cosmos3 guardrails are enabled by default when the package is available:
Command
cosmos-guardrail downloads gated NVIDIA guardrail weights, so pass a Hugging Face token if your environment needs one. If the package is not installed, SGLang skips Cosmos3 guardrails and logs a warning. To disable Cosmos3 guardrails for local experiments, set SGLANG_DISABLE_COSMOS3_GUARDRAILS=1 before starting the server. There may be problems loading the Cosmos-1.0-Guardrail weights on Ascend NPU. If the _pickle.UnpicklingError error occurs during startup, you should change weight_only=True to weights_only=False parameter in cosmos_guardrail/cosmos_utils.py:

3. Serve Cosmos3

Serve Cosmos3-Nano directly from the Hugging Face model ID:
Command
With --performance-mode auto, a Cosmos3 Nano checkpoint keeps its DiT and VAE resident when every selected GPU has at least 120 GiB available at startup. Below that threshold, auto mode retains the conservative DiT component-offload policy. This high-memory override is limited to Nano; Cosmos3 Super checkpoints keep their existing multi-GPU placement defaults. For Cosmos3-Super, split the model across multiple GPUs:
Command
The server also accepts the specialized nvidia/Cosmos3-Super-Text2Image and nvidia/Cosmos3-Super-Image2Video checkpoint IDs.

Edge checkpoints

Cosmos3-Edge is a 4B dense model and can be served on one GPU:
Command
Edge is trained for 256p and 480p generation. Its default video configuration is 832x480 with guidance_scale=5.0; its default image configuration is 640x640 with guidance_scale=7.0. Supported sizes are 832x480, 480x832, 640x480, 480x640, 480x480, 640x640, 448x256, 256x448, and 256x256. Serve the Edge DROID policy checkpoint with the same single-GPU configuration, replacing the model path with nvidia/Cosmos3-Edge-Policy-DROID.

Distilled checkpoints

The distilled Super checkpoints are 64B models. Use multiple GPUs unless the complete model and request workload fit on one GPU:
Command
For distilled I2V, replace the model path with nvidia/Cosmos3-Super-Image2Video-4Step. SGLang detects both checkpoints from scheduler/scheduler_config.json, uses the checkpoint’s fixed four-step sigma schedule, and forces guidance_scale=1.0. Do not tune num_inference_steps or flow_shift for these checkpoints.

4. OpenAI-Compatible Requests

Text to image

Cosmos3 text-to-image uses /v1/images/generations. The default Cosmos3 image response is b64_json, matching vLLM-Omni’s examples.
Command
With a server running nvidia/Cosmos3-Super-Text2Image-4Step, omit the scheduler controls and use guidance_scale=1.0:
Command

Text to video with sound

Use /v1/videos to create an asynchronous job, then poll the job and download the completed MP4. Set generate_sound=true to generate and mux a stereo 48 kHz audio track; omit it for a silent video.
Command

Image to video

This mirrors the official nvidia/Cosmos3-Nano Hugging Face image-to-video example:
Python
For the distilled I2V checkpoint, use the same API with a server running nvidia/Cosmos3-Super-Image2Video-4Step. The recommended request is 480p and does not specify scheduler controls:
Command
Poll and download this job with the same status and content endpoints used by the T2V example.

Video to video

Upload a source video with video_reference. Cosmos3 keeps latent frames [0, 1] by default and generates the remaining frames. Use condition_frame_indexes to select different latent frames, and condition_video_keep to take conditioning frames from the start or end of the source.
Command
Poll and download this job with the same status and content endpoints used by the T2V example.

Action generation

For DROID policy generation, start a single-GPU server with either the Nano or Edge policy checkpoint. Cosmos3 action generation does not currently support CFG or sequence parallelism.
Command
Use nvidia/Cosmos3-Edge-Policy-DROID in the same command to serve the smaller 4B policy checkpoint. policy and inverse_dynamics return actions, so their canonical API is the synchronous /v1/actions/generations endpoint. The following request predicts a 16-step action chunk from one observation image. action_horizon=16 maps to the model’s num_frames=17 convention.
Python
Use GET /v1/actions/metadata to inspect the action modes, default horizon, padded action dimension, and accepted observation modalities. Msgpack requests and the /v1/actions/realtime websocket use the same action envelope. inverse_dynamics also uses /v1/actions/generations; set action_mode="inverse_dynamics" and pass an observation video URL or server-local path as input.observation.video. Select the embodiment head with domain_name or domain_id; set raw_action_dim explicitly when it cannot be inferred from the domain name. forward_dynamics is intentionally different: it consumes an action array and predicts video, so it remains on /v1/videos. Action-producing modes submitted to /v1/videos return HTTP 400 with the canonical action endpoint in the error message.

5. Cosmos3 Parameters

Cosmos3 supports the standard SGLang video and image fields such as size, num_frames, fps, num_inference_steps, guidance_scale, negative_prompt, and seed. For distilled checkpoints, SGLang replaces num_inference_steps with the checkpoint’s fixed four-step schedule and forces guidance_scale=1.0; negative-prompt CFG and request-level flow_shift do not apply. Top-level Cosmos3 request fields:
  • max_sequence_length: maximum text token length used by the Cosmos3 tokenizer.
  • flow_shift: per-request scheduler shift for non-distilled checkpoints. If omitted, SGLang uses --flow-shift, then the mode default (3.0 for T2I, 10.0 for non-Edge video and all action modes, or 3.0 for Edge video).
  • guidance_interval: optional [start, end] noise interval for CFG. Non-distilled T2I defaults to [400, 1000]; video modes guide at every step.
Cosmos3 omnimodal fields are accepted as extra JSON fields or multipart form fields:
  • generate_sound: generate a sound track whose duration follows num_frames / fps.
  • sound_duration: explicit sound duration in seconds; takes precedence over the derived duration.
  • condition_frame_indexes: V2V latent-frame indexes to keep from the source video; defaults to [0, 1].
  • condition_video_keep: use the first or last source frames for V2V conditioning.
  • action_mode: policy, forward_dynamics, or inverse_dynamics.
  • domain_name / domain_id: select the action embodiment head.
  • raw_action_dim: number of active action dimensions; inferred for known domain names.
  • action: action array with shape [T, D], required by forward_dynamics.
  • action_fps: action-token frame rate for temporal mRoPE; defaults to the video FPS.
  • action_view_point: viewpoint used in the structured action caption.
  • action_normalization: dataset normalization mode, such as quantile, meanstd, or minmax.
Put model-specific compatibility knobs in extra_params for video requests, or extra_args for image requests:
  • use_duration_template: whether to append SGLang’s generated duration suffix to video prompts.
  • use_resolution_template: accepted for vLLM-Omni request compatibility.
  • use_system_prompt: whether to add the Cosmos3 system prompt to the chat template.
  • guardrails or use_guardrails: per-request guardrail toggle when the server started with guardrails enabled.