Interface VideoProjectParams

Video-specific parameters for video workflows (t2v, i2v, s2v, ia2v, a2v, animate). Only applicable when using video models like wan_v2.2-14b-fp8_t2v or ltx25-22b-int8_t2v_distilled. Includes frame count, fps, shift, and reference assets (image, audio, video).

  • Always generate video at 16fps internally
  • The fps parameter (16 or 32) only controls post-render frame interpolation
  • fps=32 doubles the frames via interpolation after generation
  • Frame count is always calculated as: duration * 16 + 1
  • Example: 5 seconds at 32fps = 81 frames generated, then interpolated to 161 output frames
  • Generate video at the actual specified FPS (1-60 fps range)
  • No post-render interpolation - fps directly affects generation
  • Frame count is calculated as: duration * fps + 1
  • Frame count must follow the pattern: 1 + n*8 (i.e., 1, 9, 17, 25, 33, ...)
  • Example: 5 seconds at 24fps = 121 frames (since 121 = 1 + 15*8)
  • External API-backed video models for text-to-video, image-to-video, multimodal reference generation, image+audio-to-video, and video-to-video
  • Generate at fixed 24fps
  • Full tier supports up to 4K output; Fast caps at 720p
  • Direct SDK project duration range is 4 to 15 seconds
  • Frame count is calculated as: duration * 24 + 1
  • Vendor reference limits are 9 images, 3 videos, 3 audios, and 12 asset files total
  • External API-backed (Alibaba) video models with native audio
  • happyhorse-1.1-t2v (text-to-video), happyhorse-1.1-i2v (single first-frame image-to-video), and happyhorse-1.1-r2v (1-9 reference images-to-video)
  • Generate at fixed 24fps; output up to 720P/1080P
  • Direct SDK project duration range is 3 to 15 seconds
  • Frame count is calculated as: duration * 24 + 1
  • Image-only reference context: no reference video or reference audio assets
  • Text-to-video, endpoint-conditioned image-to-video (first frame, last frame, or both), first-and-last-frame video, and multi-reference video. Two checkpoints ship: FL2VA (minimax-h3-fl2va-fp8_t2v / _i2v / _flf2v) and Ref2VA (minimax-h3-ref2va-fp8_r2v).
  • Turbo adds _turbo to the FL2VA t2v/i2v/flf2v IDs and to Ref2VA r2v. Ref2VA Turbo is minimax-h3-ref2va-fp8_r2v_turbo; it uses its dedicated LightX2V v0.1 four-step LoRA with Euler/simple, not the FL2VA Turbo LoRA.
  • Video and 32kHz stereo audio are generated jointly. Audio is included by default; set generateAudio: false to return a video without an audio track.
  • Generation is fixed at 24fps and guidance 1, with no separate negative-prompt input. Standard H3 uses 20 steps and res_multistep/simple; Turbo uses its fixed 4-step sampling path.
  • Frames follow 124 + n*17 from 124 through 362. Dimensions use a 32px grid, with a 1344px per-axis limit and a 1032192-pixel canvas limit.
  • The i2v model accepts referenceImage, referenceImageEnd, or both, and requires at least one of them. The flf2v model requires both.
  • minimax-h3-ref2va-fp8_r2v (standard) and minimax-h3-ref2va-fp8_r2v_turbo (four-step Turbo) condition on labelled reference material rather than on frame anchors. The checkpoint accepts up to 9 reference images, 3 reference videos (24fps, 2-15 seconds each), and 3 reference audio clips, with at most 12 reference files in total.

  • All of those ceilings apply, and at least one visual reference (image or video) is required. Audio-only reference sets are rejected.

  • r2v is the only reference workflow that runs on Sogni's own workers rather than at a vendor, so every reference is uploaded to S3 before the request is sent. Images use referenceImage plus contextImages; videos use referenceVideo plus referenceVideos; audio uses referenceAudio plus referenceAudios.

  • Upload order is preserved, so a prompt ordinal refers to a predictable file.

  • References are presented to the model in a fixed order - images, then videos (each video's own soundtrack immediately before it), then standalone audio - and are numbered from 1 per type. The H3 text encoder splices a literal label in front of each one before your prompt text (comfy/text_encoders/minimax.py emits "<Picture %d>: ", "<Video %d>: " and "<Audio %d>: "), so write the SAME form in the prompt - <Picture 1>, <Video 1>, <Audio 1>, angle brackets included - and the reference and the sentence about it share one token sequence. Prose aliases like "Image 1" or "the second photo" do not.

  • Give every reference an explicit job, or the model averages them. Separate identity from style, motion from appearance, and voice character from spoken words, and state the priority when two references conflict: "Use <Picture 1> for the character's face and hairstyle. Use <Picture 2> only for environment and lighting. When <Picture 1> and <Picture 2> disagree, <Picture 1> wins."

  • Reference resolution has a real cost/quality tradeoff. The workflow's ref_image_size is match by default, which scales references down to the generation pixel area. max uses a 2048px short edge for the best identity fidelity, but its reference tokens ride through every sampling step, making the render several times slower. match is what Sogni ships; max is not exposed as an SDK parameter.

  • r2v has no frame anchors, so referenceImageEnd is rejected. There is no closing frame to pin.

  • See the repository's authoring examples: https://github.com/Sogni-AI/sogni-client/blob/alpha/examples/workflow_minimax_h3_video.mjs

interface VideoProjectParams {
    appSource?: string;
    attribution?: WorkloadAttributionInput;
    audioDuration?: number;
    audioIdentityStrength?: number;
    audioStart?: number;
    billingMode?: BillingMode;
    contextImages?: InputMedia[];
    controlNet?: VideoControlNetParams;
    detailerStrength?: number;
    disableNSFWFilter?: boolean;
    duration?: number;
    firstFrameStrength?: number;
    fps?: number;
    frames?: number;
    generateAudio?: boolean;
    guidance?: number;
    height?: number;
    lastFrameStrength?: number;
    loras?: string[];
    loraStrengths?: number[];
    modelId: string;
    negativePrompt?: string;
    network?: SupernetType;
    numberOfMedia: number;
    outpaintPosition?: "center" | "top" | "bottom" | "left" | "right";
    outputFormat?: "mp4";
    positivePrompt: string;
    referenceAudio?: InputMedia;
    referenceAudioIdentity?: InputMedia;
    referenceAudios?: InputMedia[];
    referenceAudioUrls?: string[];
    referenceImage?: InputMedia;
    referenceImageEnd?: InputMedia;
    referenceImageUrls?: string[];
    referenceMask?: InputMedia;
    referenceVideo?: InputMedia;
    referenceVideos?: InputMedia[];
    referenceVideoUrls?: string[];
    sam2Coordinates?: { x: number; y: number }[];
    sampler?: string;
    scheduler?: string;
    seed?: number;
    shift?: number;
    steps?: number;
    stylePrompt?: string;
    teacacheThreshold?: number;
    tokenType?: TokenType;
    trimEndFrame?: boolean;
    type: "video";
    videoStart?: number;
    width?: number;
}

Hierarchy (View Summary)

Properties

appSource?: string

Optional client app/source label to attach to the project request for server-side attribution.

Optional workload attribution for this project. Fields override the immutable defaults configured on SogniClient.

audioDuration?: number

Audio duration in seconds for audio-driven workflows (s2v, ia2v, a2v). Specifies how many seconds of audio to use. If not provided, defaults to 30 seconds on the server.

audioIdentityStrength?: number

Controls how strongly the speaker's vocal identity is applied. Uses an extra forward pass per denoising step to amplify identity features. Range: 0-10. Default: 3.0. Set to 0 to disable (skips extra forward pass). Only used when referenceAudioIdentity is provided.

audioStart?: number

Audio start position in seconds for audio-driven workflows (s2v, ia2v, a2v). Specifies where to begin reading from the audio file. Default: 0

billingMode?: BillingMode

Select how eligible jobs should be billed.

  • auto: use Unlimited subscription coverage when available, otherwise use tokens.
  • subscription: require Unlimited subscription coverage; fail if unavailable.
  • tokens: opt out of Unlimited coverage and use Spark/SOGNI tokens.
contextImages?: InputMedia[]

Uploaded reference images for the MiniMax H3 r2v multi-reference workflow (minimax-h3-ref2va-fp8_r2v), in the order the model is shown them.

This is the video counterpart of ImageProjectParams.contextImages, and it uses the same contextImage1..contextImage9 upload slots used by Qwen Image Edit and GPT Image. It exists because H3 is Comfy-native: the worker builds the ComfyUI graph locally from Sogni-hosted assets.

The uploaded reference set is [referenceImage, ...contextImages]. Both fields count against the same 9-image ceiling. Images are optional when at least one reference video is supplied. Prompt ordinals follow that order: with referenceImage set, contextImages[0] is <Picture 2>; without it, contextImages[0] is <Picture 1>. Entries must not be empty - a hole would renumber every reference after it.

Any other video model rejects this field.

Control parameters for LTX 2.5 or LTX 2.3 v2v workflows. Specifies which control signal to extract from the reference video.

detailerStrength?: number

Detailer LoRA strength for LTX 2.5 or LTX 2.3 v2v IC-Control workflows. The detailer LoRA is always loaded alongside the control LoRA (canny/pose/depth). Range: 0.0-1.0, default 0.6.

disableNSFWFilter?: boolean

Requested content-filter policy. The server remains authoritative.

duration?: number

Duration of the video in seconds. Supported range 1 to 10 (WAN), 2 to 20 (LTX 2.5), 4 to 20 (LTX 2.3), 4 to 15 (Seedance direct SDK projects), 3 to 15 (HappyHorse direct SDK projects), or 124/24 to 362/24 seconds (MiniMax H3).

The SDK automatically calculates the correct frame count based on the model:

  • WAN 2.2: duration * 16 + 1 (always 16fps generation)
  • LTX 2.x: duration * fps + 1, snapped to frame step constraint
  • Seedance: duration * 24 + 1
  • HappyHorse: duration * 24 + 1
  • MiniMax H3: duration * 24 snapped to the 124 + n*17 grid and clamped to 124-362 frames (always 24fps generation, and no +1 term)
firstFrameStrength?: number

First frame strength for LTX-2.3 keyframe interpolation (when referenceImageEnd is provided). Controls how strictly the first frame is matched. Range: 0.0-1.0, default 0.6. Set to 0 to disable first frame (last-frame-only mode).

fps?: number

Frames per second for output video.

WAN 2.2 Models: Only 16 or 32 fps allowed. The 32fps option is post-render frame interpolation that doubles the output frames. Internal generation is always 16fps.

LTX 2.x Models: Any value from 1-60 fps. This directly controls the generation frame rate - there is no post-render interpolation.

Seedance Models: Fixed 24fps external API generation.

HappyHorse Models: Fixed 24fps external API generation.

MiniMax H3 Models: Fixed 24fps. Omit this field or pass 24.

frames?: number

Number of frames to generate.

Use duration instead. When using duration, the SDK automatically calculates the correct frame count based on the model type.

generateAudio?: boolean

Include the model's generated/native audio track when supported. Audio is enabled by default; set to false to return a video without an audio track.

guidance?: number

Guidance scale. For most Stable Diffusion models, optimal value is 7.5. For video models: Regular models range 0.7-8.0, LoRA version (lightx2v) range 0.7-1.6, step 0.01. This maps to guidanceScale in the keyFrame for both image and video models.

height?: number

Output video height. Only used if sizePreset is "custom"

lastFrameStrength?: number

Last frame strength for LTX-2.3 keyframe interpolation (when referenceImageEnd is provided). Controls how strictly the last frame is matched. Range: 0.0-1.0, default 0.6.

loras?: string[]

LoRA IDs to apply, in the order they should be chained.

Which LoRAs are available depends on the model; the Krea 2 family carries the largest set. Workers download a LoRA on first use, so the first render with an uncached one takes longer to start.

Order is significant. The LoRAs are applied in sequence and the same set in a different order produces a measurably different image, because these models run fp8-quantized and the patches do not commute.

Up to 8 per render. IDs are resolved to filenames by the worker. Example: ['krea2-detail-enhancer', 'krea2-amateur']

loraStrengths?: number[]

Strength for each entry in loras, positionally matched. Defaults to 1.0.

Not restricted to positive values. Most Krea 2 LoRAs are bipolar sliders where a negative strength applies the inverse effect and 0 does nothing - Warm Light warms at 2 and cools at -2. Each LoRA has its own valid range and its author's recommended band; values outside the valid range are clamped server-side, and pushing past the recommended band usually costs detail rather than adding effect.

Example: [3, -2]

modelId: string

ID of the model to use, available models are available in the availableModels property of the ProjectsApi instance.

negativePrompt?: string

Prompt for what to be avoided. LTX 2.5, LTX 2.3, and WAN video workflows accept this field; provider workflows such as MiniMax H3 and Seedance do not. If not provided, the server or workflow default is used.

network?: SupernetType

Override current network type. Default value can be read from sogni.account.currentAccount.network

numberOfMedia: number

Number of media files to generate. Depending on project type, this can be number of images or number of videos.

outpaintPosition?: "center" | "top" | "bottom" | "left" | "right"

Outpaint canvas anchor for distilled LTX 2.5 or LTX 2.3 v2v outpaint workflows. Determines where the original frame is placed within the expanded canvas. Default: 'center'.

outputFormat?: "mp4"

Output video format. For now only 'mp4' is supported, defaults to 'mp4'.

positivePrompt: string

Prompt for what to be created

referenceAudio?: InputMedia

Reference audio for audio-driven video workflows (s2v, ia2v, a2v).

On the MiniMax H3 r2v workflow this is standalone reference audio 1 - a voice or soundtrack the prompt assigns a job to, not a track the video is driven by. It is the first item in [referenceAudio, ...referenceAudios].

referenceAudioIdentity?: InputMedia

Reference audio for ID-LoRA speaker identity transfer (LTX-2.3 only). Provide a ~5 second audio clip of the target speaker's voice. The model uses this to transfer vocal identity into the generated video. Available on t2v, i2v, and v2v LTX-2.3 workflows. Not compatible with audio-driven workflows (s2v, ia2v, a2v).

referenceAudios?: InputMedia[]

Additional uploaded standalone audio references for MiniMax H3 r2v. Together with referenceAudio, at most three clips are accepted. Entries are uploaded to distinct S3 objects and retain array order.

referenceAudioUrls?: string[]

Audio context references for Seedance. These must be publicly accessible HTTPS URLs. Seedance does not support text+audio-only requests; include at least one image or video reference when using audio URL references. MiniMax H3 r2v uses uploaded referenceAudios instead.

referenceImage?: InputMedia

Reference image for video workflows. Maps to: startImage (i2v), characterImage (animate), referenceImage (s2v, ia2v)

MiniMax H3 i2v accepts this first-frame anchor by itself, together with referenceImageEnd, or can omit it when referenceImageEnd is supplied.

On the MiniMax H3 r2v workflow (minimax-h3-ref2va-fp8_r2v) this is reference image 1 (<Picture 1>) rather than a frame anchor, and it is optional: the same slot can be filled from contextImages instead.

referenceImageEnd?: InputMedia

Optional end image for i2v workflows. It can be provided alone for last-frame-only generation, or with referenceImage to interpolate between two images.

MiniMax H3 i2v accepts either endpoint independently or both together, with at least one required.

Required, together with referenceImage, for the MiniMax H3 flf2v workflow (minimax-h3-fl2va-fp8_flf2v), which always interpolates between two anchor frames.

Rejected by the MiniMax H3 r2v workflow, which has no closing frame to pin. Its second reference image is the next entry in contextImages.

referenceImageUrls?: string[]

Loose image context references for Seedance and HappyHorse. These must be publicly accessible HTTPS URLs and are handed to the external vendor. Use referenceImage / referenceImageEnd when the image should lock the first or last frame. HappyHorse r2v accepts 1-9 reference images here.

MiniMax H3 r2v does not accept URL references; use referenceImage and contextImages, which the SDK uploads through Sogni's asset path.

referenceMask?: InputMedia

Inpaint mask IMAGE for distilled LTX 2.5 or LTX 2.3 v2v inpaint workflows. White pixels mark the region to regenerate. Maps to jobKey 'referenceMask'. Used by the 'inpaint' control type.

referenceVideo?: InputMedia

Reference video for animate and v2v (ControlNet) workflows. Maps to: drivingVideo (animate-move), sourceVideo (animate-replace), referenceVideo (v2v)

On the MiniMax H3 r2v workflow this is reference video 1 (<Video 1>, read as 24fps) that the prompt assigns a job to - camera movement, blocking, or subject motion - rather than a source clip to transform. Its own soundtrack is also presented to the model, numbered before any standalone reference audio. It is the first item in [referenceVideo, ...referenceVideos].

referenceVideos?: InputMedia[]

Additional uploaded video references for MiniMax H3 r2v. Together with referenceVideo, at most three clips are accepted. Entries are uploaded to distinct S3 objects and retain array order.

referenceVideoUrls?: string[]

Video context references for Seedance. These must be publicly accessible HTTPS URLs, and map to Seedance reference_video assets. MiniMax H3 r2v uses uploaded referenceVideos instead.

sam2Coordinates?: { x: number; y: number }[]

SAM2 click coordinates for subject detection in animate-replace workflows. Array of {x, y} coordinate objects indicating where the subject is located in the reference image.

Coordinates can be normalized (0.0-1.0) or absolute pixel values. Normalized coordinates are automatically converted to pixel values by the server. If not provided, the server defaults to the center of the frame.

Example: [{ x: 0.5, y: 0.5 }] for center of frame

sampler?: string

Sampler, available options depend on the model. Use sogni.projects.getModelOptions(modelId) to get the list of available samplers.

scheduler?: string

Scheduler, available options depend on the model. Use sogni.projects.getModelOptions(modelId) to get the list of available schedulers.

seed?: number

Seed for one of images in project. Other will get random seed. Must be Uint32

shift?: number

Shift parameter for video diffusion models. Controls motion intensity. Range: 1.0-8.0, step 0.1. Default: 8.0 for regular models, 5.0 for speed lora (lightx2v) except s2v and animate which use 8.0

steps?: number

Number of steps. For most Stable Diffusion models, optimal value is 20.

stylePrompt?: string

Image style prompt. If not provided, server default is used.

teacacheThreshold?: number

TeaCache optimization threshold for T2V and I2V models. Range: 0.0-1.0. 0.0 = disabled. Recommended: 0.15 for T2V (~1.5x speedup), 0.2 for I2V (conservative quality-focused)

tokenType?: TokenType

Select which tokens to use for the project. If not specified, the Sogni token will be used.

trimEndFrame?: boolean

Trim the last frame from the generated video. Used for seamless stitching of transition videos where the last frame duplicates the end reference image. Default: false

type: "video"
videoStart?: number

Video start position in seconds for animate workflows (animate-move, animate-replace). Specifies where to begin reading from the reference video file. Default: 0

width?: number

Output video width. Only used if sizePreset is "custom"