Interface AudioProjectParams

interface AudioProjectParams {
    appSource?: string;
    attribution?: WorkloadAttributionInput;
    billingMode?: BillingMode;
    bpm?: number;
    composerMode?: boolean;
    creativity?: number;
    disableNSFWFilter?: boolean;
    duration?: number;
    guidance?: number;
    keyscale?: string;
    language?: string;
    loras?: string[];
    loraStrengths?: number[];
    lyrics?: string;
    modelId: string;
    negativePrompt?: string;
    network?: SupernetType;
    numberOfMedia: number;
    outputFormat?: AudioOutputFormat;
    positivePrompt: string;
    promptStrength?: number;
    sampler?: string;
    scheduler?: string;
    seed?: number;
    shift?: number;
    steps?: number;
    stylePrompt?: string;
    timesignature?: string;
    tokenType?: TokenType;
    type: "audio";
}

Hierarchy (View Summary)

Properties

appSource?: string

Optional client app/source label to attach to the project request for server-side attribution.

Optional workload attribution for this project. Fields override the immutable defaults configured on SogniClient.

billingMode?: BillingMode

Select how eligible jobs should be billed.

  • auto: use Unlimited subscription coverage when available, otherwise use tokens.
  • subscription: require Unlimited subscription coverage; fail if unavailable.
  • tokens: opt out of Unlimited coverage and use Spark/SOGNI tokens.
bpm?: number

Beats per minute (30-300, default: 120)

composerMode?: boolean

Enable AI composer mode for higher quality music generation (default: true). Disable for faster generation or when using reference audio. Maps to generate_audio_codes in the ComfyUI workflow.

creativity?: number

Composition variation / temperature (0-2, default: 0.85). Higher = more creative, lower = more predictable. Maps to temperature in the ComfyUI workflow.

disableNSFWFilter?: boolean

Requested content-filter policy. The server remains authoritative.

duration?: number

Duration of the audio in seconds (10-600, default: 30)

guidance?: number

Guidance scale. For most Stable Diffusion models, optimal value is 7.5. For video models: Regular models range 0.7-8.0, LoRA version (lightx2v) range 0.7-1.6, step 0.01. This maps to guidanceScale in the keyFrame for both image and video models.

keyscale?: string

Key/scale setting (e.g., "C major", "A minor"). Omitted to use server default.

language?: string

Lyrics language code (default: en)

loras?: string[]

LoRA IDs to apply, in the order they should be chained.

Which LoRAs are available depends on the model; the Krea 2 family carries the largest set. Workers download a LoRA on first use, so the first render with an uncached one takes longer to start.

Order is significant. The LoRAs are applied in sequence and the same set in a different order produces a measurably different image, because these models run fp8-quantized and the patches do not commute.

Up to 8 per render. IDs are resolved to filenames by the worker. Example: ['krea2-detail-enhancer', 'krea2-amateur']

loraStrengths?: number[]

Strength for each entry in loras, positionally matched. Defaults to 1.0.

Not restricted to positive values. Most Krea 2 LoRAs are bipolar sliders where a negative strength applies the inverse effect and 0 does nothing - Warm Light warms at 2 and cools at -2. Each LoRA has its own valid range and its author's recommended band; values outside the valid range are clamped server-side, and pushing past the recommended band usually costs detail rather than adding effect.

Example: [3, -2]

lyrics?: string

Song lyrics. Omit for instrumental generation.

modelId: string

ID of the model to use, available models are available in the availableModels property of the ProjectsApi instance.

negativePrompt?: string

Prompt for what to be avoided. LTX 2.5, LTX 2.3, and WAN video workflows accept this field; provider workflows such as MiniMax H3 and Seedance do not. If not provided, the server or workflow default is used.

network?: SupernetType

Override current network type. Default value can be read from sogni.account.currentAccount.network

numberOfMedia: number

Number of media files to generate. Depending on project type, this can be number of images or number of videos.

outputFormat?: AudioOutputFormat

Output audio format. Can be 'mp3', 'flac', or 'wav'. Defaults to 'mp3'.

positivePrompt: string

Prompt for what to be created

promptStrength?: number

How closely the AI composer follows your prompt (0-10, default: 2.0). Higher values = stricter prompt adherence. Maps to cfg_scale in the ComfyUI workflow.

sampler?: string

Sampler, available options depend on the model.

scheduler?: string

Scheduler, available options depend on the model.

seed?: number

Seed for one of images in project. Other will get random seed. Must be Uint32

shift?: number

Shift parameter for ModelSamplingAuraFlow (1-6, default: 3 for turbo). Controls how denoising effort is distributed across generation steps. Higher values front-load structure/composition, producing more coherent arrangements. Lower values distribute effort evenly, focusing more on detail/texture. Official ComfyUI template uses shift=3 for ACE-Step 1.5 Turbo.

steps?: number

Number of steps. For most Stable Diffusion models, optimal value is 20.

stylePrompt?: string

Image style prompt. If not provided, server default is used.

timesignature?: string

Time signature (2, 3, 4, or 6 - default: 4)

tokenType?: TokenType

Select which tokens to use for the project. If not specified, the Sogni token will be used.

type: "audio"