{
  "version": "2026-07-18.1",
  "source": "sogni-creative-agent/src/tools/definitions/*/definition.ts",
  "schemaRefs": {
    "generate_image": "../schemas/tools/generate_image.schema.json",
    "generate_video": "../schemas/tools/generate_video.schema.json",
    "generate_music": "../schemas/tools/generate_music.schema.json",
    "generate_speech": "../schemas/tools/generate_speech.schema.json",
    "edit_image": "../schemas/tools/edit_image.schema.json",
    "apply_style": "../schemas/tools/apply_style.schema.json",
    "restore_photo": "../schemas/tools/restore_photo.schema.json",
    "upscale_image": "../schemas/tools/upscale_image.schema.json",
    "upscale_video": "../schemas/tools/upscale_video.schema.json",
    "refine_result": "../schemas/tools/refine_result.schema.json",
    "animate_photo": "../schemas/tools/animate_photo.schema.json",
    "change_angle": "../schemas/tools/change_angle.schema.json",
    "video_to_video": "../schemas/tools/video_to_video.schema.json",
    "stitch_video": "../schemas/tools/stitch_video.schema.json",
    "orbit_video": "../schemas/tools/orbit_video.schema.json",
    "dance_montage": "../schemas/tools/dance_montage.schema.json",
    "sound_to_video": "../schemas/tools/sound_to_video.schema.json",
    "extend_video": "../schemas/tools/extend_video.schema.json",
    "replace_video_segment": "../schemas/tools/replace_video_segment.schema.json",
    "overlay_video": "../schemas/tools/overlay_video.schema.json",
    "add_subtitles": "../schemas/tools/add_subtitles.schema.json",
    "segment_image": "../schemas/tools/segment_image.schema.json",
    "image_to_3d": "../schemas/tools/image_to_3d.schema.json",
    "remove_background": "../schemas/tools/remove_background.schema.json"
  },
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "generate_image",
        "description": "Generate a new image from a text description. Usually this is text-only: do NOT use this tool when the user expects an existing image to be reused or preserved in the result. That includes (a) people from My Personas, and (b) uploaded assets such as logos, brand marks, mascots, product shots, photos, screenshots, sketches, character designs, or other reference images they want carried through. Use edit_image with sourceImageIndex=-1 (or the appropriate generated index) instead. Exception: when the user explicitly requests Z-image, Z Image, Z-image Turbo, or Krea 2 Turbo for an uploaded-image enhancement/image-to-image request, use this tool with model=\"z-turbo\", model=\"z-image\", or model=\"krea-2-turbo\", sourceImageIndex=-1, and starting_image_strength because edit_image does not expose those base image-to-image models.",
        "parameters": {
          "type": "object",
          "properties": {
            "prompt": {
              "type": "string",
              "description": "Text description of the image (50-200 words). POSITIVE phrasing only. Be specific and vivid — reference real artists, franchises, and aesthetics by name.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing.\n\nPROMPT ORDER (follow this structure by default): [SUBJECT] → [ATTRIBUTES] → [ACTION/POSE] → [CAMERA/FRAMING] → [ENVIRONMENT] → [LIGHTING] → [STYLE/MEDIUM] → [MATERIALS/TEXTURES] → [SECONDARY DETAILS]. Lead with the main subject and its concrete, observable attributes unless the user explicitly asks for mood, atmosphere, or another prompt shape first. Put the most visually decisive details early.\n\nSPECIFICITY: Use concrete nouns and observable adjectives (\"weathered leather jacket\", not \"cool outfit\"). Specify framing (close-up, medium shot, full body, wide shot), angle (eye level, low angle, high angle, overhead), lighting type (\"soft overcast daylight\", \"warm golden-hour sunlight\", \"moody neon spill with deep shadows\"), and medium/style (\"photorealistic editorial photography\", \"cinematic still frame\", \"clean anime illustration\"). Include materials and textures when relevant (\"brushed aluminum\", \"wet asphalt reflections\", \"heavy wool texture\").\n\nDEFAULTS (fill in when user is underspecified): Framing: medium shot for portraits, wide shot for environments, full-body for fashion/outfits. Angle: eye level unless dramatic perspective requested. Lighting: soft natural light for realism, clean studio light for product shots. Style: photorealistic for realistic models, matching the model's native style for stylized models (e.g. anime illustration for pony/animagine). Reference real artists and franchises by name (\"in the style of Monet's Water Lilies\", \"Wes Anderson symmetrical pastel composition\", \"cyberpunk Blade Runner neon city\", \"shot on 85mm f/1.4 with shallow depth of field\").\n\nAVOID: Starting with abstract mood words alone. Burying the subject after a long style preamble. Stacking incompatible styles. Overloading with competing focal points. Vague phrases like \"very cool\" or \"epic vibes\".\n\nCHARACTER / MASCOT SHEETS: When the user asks for a character sheet, mascot sheet, model sheet, turnaround, expression sheet, or reusable character reference board, create ONE comprehensive professional reference-board image, not separate variations. Include a large hero pose, front / 3/4 / side / back turnaround views, an expression row, action/personality poses, accessories or props, color palette swatches, and compact notes such as personality, fun facts, or brand usage when appropriate. Preserve exact user-provided brand names, slogans, logo text, and requested copy verbatim; incidental tiny notes may be generated by the image model if the user did not provide exact wording. Keep the character consistent across every panel and use clean readable typography.\n\nBATCH VARIATIONS: When numberOfVariations > 1, the prompt describes one output image. Do not mention counts, \"versions\", \"different\", or \"multiple\" in the prompt text unless the user explicitly wants those words visible in the image. Do not describe multiple copies or duplicates of the subject in a single image unless the user asked for a collage, grid, or side-by-side composition. Use Dynamic Prompt syntax to vary one dimension across separate images. Example: user asks \"4 cats in different spots\" → numberOfVariations=4, prompt=\"a black cat {lounging in a sunlit window|prowling through autumn leaves|sitting on a vintage bookshelf|curled up by a fireplace}\" — each output is one cat in one spot. Vary setting, style, lighting, expression, or composition; preserve what the user specified. Preserve any requested orientation, aspect ratio, or exact pixel dimensions across every variation.\n\nSELECTION-GATED IMAGE STAGES: If the user asks for multiple image options/takes/versions and says they will pick one before a later dance, animation, or video, this tool call is still the first step. Generate the complete image batch now with the exact requested count, Dynamic Prompt options for each output, and the final video/image aspect ratio. Do not ask the user to choose before the images exist, and do not call video tools until after the user selects an image.\n\nLINKED VARIANTS: If multiple details must stay paired per output — visual style, outfit, label text, symbol, setting, character, prop, location, or before/after keyframe details — use ONE top-level Dynamic Prompt branch with one complete prompt per output. Do NOT use separate Dynamic Prompt groups for details that must stay together; unpaired groups can mix attributes. Treat every user requirement quantified across the batch (\"each\", \"every\", \"all\") as a hard per-option invariant and repeat it inside EVERY option. This includes identity/pose continuity, required clothing or styling, the actual setting, and literal visible names, labels, captions, flags, logos, or symbols. When the user names a subject or character, write that name or stable role inside every Dynamic Prompt option; a shared prefix outside the branch is not enough because each option must stand alone. Correct shape: \"{full prompt for variant 1 with all paired details|full prompt for variant 2 with all paired details|...}\".\n\nEach option must be a fully concrete standalone image description. Name the actual garment or styling, actual setting, actual accessories, and literal text or symbol shown on screen when requested. Never use meta-placeholder phrasing such as \"style-specific outfit\", \"variant-specific background\", \"include the requested symbol\", \"include a humorous alternate name\", or \"bake the name and symbol into the image\" — those describe the task instead of the image.\n\nORIGINAL + VARIANT BATCHES: When one option remakes or preserves an original and the other options are themed variants, the original option still needs a complete visual contract. Specify the original clothing and setting to preserve, plus every requested label, flag, logo, symbol, or prop. Do not leave that option as only \"the original\" or \"unchanged subject\" while the other options are concrete.\n\nSCREENPLAY / STORYBOARD BATCHES: For multi-scene commercials, storyboards, or shot lists, numberOfVariations should equal the scene count and the prompt should be a single top-level dynamic branch containing one full scene prompt per option, e.g. \"{scene 1 full prompt|scene 2 full prompt|scene 3 full prompt}\". This is the required way to batch scenes with materially different content while still rendering one image per scene. If recurring characters appear, use stable character names and repeat the same visual anchors in every scene option where they appear (age range, build, hairstyle, outfit silhouette, color palette, signature prop/accessory, posture). Do not rename, merge, redesign, or drift characters between scene keyframes unless the user asks. Include speaker-tagged dialogue details when dialogue affects the keyframe, e.g. CHARACTER: \"We made it.\" Do not set numberOfVariations=N with only scene 1's prompt; that creates N duplicate versions of scene 1, not N scenes. If the scene count is 16 or fewer, keep it in one call unless the user explicitly asks for separate projects or per-output settings require separate calls.\n\nCOMPOSITE GPT IMAGE 2 STORYBOARD SHEETS: When numberOfVariations=1 and the user asks for one composite video storyboard/keyframe sheet, the prompt must be a compiled storyboard prompt, not a concept summary. Include a SCENES: section with exactly the requested number of concrete entries named SCENE_01, SCENE_02, etc. Every scene entry must include Visual/Action, Camera/Motion, Dialogue/VO (or [no dialogue]), Audio/SFX, and any visible text or reference usage for that scene. Do not provide only the source brief or generic layout instructions; malformed compiled storyboard prompts are blocked by quality audit.\n\nVIDEO KEYFRAMES: When generating images intended as first+last frames for video (animate_photo with frameRole=\"both\"), use numberOfVariations=2 with Dynamic Prompts to create both frames in one call. Make each frame a distinct scene that creates a compelling transition. The video handler will inspect both generated frames and build a scene-aware transition prompt, so focus this image prompt on producing strong start/end visuals. Example: \"a serene lake {at dawn with mist rising and soft pink sky|at dusk with fireflies and deep blue twilight}\".\n\nDISTINCT IMAGE SETS: When the user asks for a set, batch, or collection of distinct images/pages/designs/options in one project, do not write one composite prompt that lists all requested outputs as contents of every image. Set numberOfVariations to the requested output count and use exactly one Dynamic Prompt branch with the same number of options. Put shared style, medium, constraints, and dimensions outside the branch, and put one complete output concept in each branch option. Shape: \"shared constraints {complete prompt for output 1|complete prompt for output 2|...|complete prompt for output N}\".\n\nSEAMLESSLY REPEATING / TILING IMAGES: when the requested image is meant to repeat edge to edge without visible joins — a seamless pattern, repeating texture, wallpaper, tiling background, or an Escher-style tessellation of interlocking figures — it needs a specific configuration, because an ordinary render will not wrap. Set model=\"krea-2-turbo\" and width=1024 with height=1024. 1024x1024 is the only size that tiles reliably; 768, 1280, 1536 and non-square aspect ratios were measured at a 0% success rate, so do not honor a different size for a tiling request without telling the user it will not wrap. Build the prompt as: the subject, then \"a perfect crop from an infinite repeating pattern that continues beyond every edge\", then a motif-scale clause, then a lighting clause. MOTIF SCALE: add \"the motif repeats exactly once across and once down\" for large bold figures, or \"the motif repeats exactly two times across and two times down\" for a medium pattern. Use only 1 or 2 — both measured 63% while 3 and 4 were worse — and always keep the two counts equal. LIGHTING: the frame must not carry a global light gradient, because that is what makes opposite edges disagree — but individual figures may still be shaded. Use \"consistent even illumination from edge to edge, with natural shading and depth modeled within each object\" to keep three-dimensional depth, or \"uniform flat lighting with no shadows or vignette\" for a flatter graphic look and a slightly higher hit rate. Do not omit the lighting clause and do not soften it to something vague like \"evenly lit\": both drop the success rate to near zero. Keep the subject tonally close — one dominant colour family — because high-contrast palettes expose the seam. For an interlocking Escher tessellation rather than a flat pattern, phrase the subject as \"photorealistic Escher tessellation of <objects>\" and add \"every figure complete and recognizable, fitting its neighbors perfectly with no gaps, no overlaps\". Compliant rounded shapes tessellate (frogs, ducks, shells, leaves, feathers, lizards); rigid objects resist. Subject choice matters more than any clause. Tiling is probabilistic even with the right configuration — roughly half of renders wrap cleanly on a good subject and fewer on a hard one — so set numberOfVariations=4 and tell the user to pick whichever tiles, rather than promising every result will."
            },
            "model": {
              "type": "string",
              "enum": [
                "gpt-image-2",
                "gpt-image-2.5-sunburst",
                "gpt-image-2.5-flare",
                "z-turbo",
                "z-image",
                "krea-2-turbo",
                "dark-beast-krea2",
                "dark-beast-z-turbo",
                "chroma-v46-flash",
                "chroma1-hd",
                "chroma-detail",
                "pony-v7",
                "qwen-2512",
                "qwen-2512-lightning",
                "albedo-xl",
                "animagine-xl",
                "one-obsession-v22",
                "anima-pencil-xl",
                "art-universe-xl",
                "hyphoria-real",
                "analog-madness-xl",
                "cyberrealistic-xl",
                "real-dream-xl",
                "faetastic-xl",
                "zavychroma-xl",
                "pony-faetality",
                "dreamshaper-xl"
              ],
              "description": "DO NOT SET THIS PARAMETER unless the user names a specific model, asks for a very complex image render, asks for a video storyboard/storyboard sheet/contact sheet/panel layout image, asks for anime without naming a model, requests permitted NSFW/nudity content, or explicitly asks for Z-image/Z-image Turbo/Krea 2 Turbo image-to-image. The app auto-selects based on quality settings. Set \"gpt-image-2\" when the user explicitly asks for legacy GPT Image 2. Set \"gpt-image-2.5-sunburst\" by default for complex single-image renders that need very strong text rendering, dense labels, crisp typography, multi-panel composition, timing notes, foley notes, professional storyboard-sheet layout, or a comprehensive character/mascot/model sheet with turnarounds, expressions, accessories, palette swatches, and brand notes. Set \"one-obsession-v22\" when the user asks for an anime or anime-style image and has not named a specific image model. Set \"z-turbo\" when the user asks for Z-image Turbo; set \"z-image\" when they ask for Z-image without Turbo. Set \"krea-2-turbo\" when the user asks for Krea 2 Turbo. If the user names another image model, honor that requested model instead. A model preference usually does not change which tool to use; the Z-image and Krea 2 Turbo image-to-image exception uses sourceImageIndex plus starting_image_strength on this tool. NSFW rule: \"gpt-image-2\"/Qwen image models CANNOT do nudity. For permitted NSFW/nudity content, prefer \"dark-beast-krea2\", then \"dark-beast-z-turbo\"; \"chroma1-hd\", \"pony-v7\", \"chroma-detail\", \"chroma-v46-flash\", and \"z-turbo\" are compatible fallbacks. GPT Image 2.5 adds two distinct models: use \"gpt-image-2.5-sunburst\" for an explicit Sunburst request and \"gpt-image-2.5-flare\" for an explicit Flare request. For GPT Image 2.5 without a named variant, use Flare. Preserve explicit GPT Image 2.0 as \"gpt-image-2\". Sunburst is positioned for difficult images and precise edits; Flare for faster everyday generation. Both generate and edit. Model choice is independent of rendering quality."
            },
            "width": {
              "type": "number",
              "description": "Output image width in pixels. Default: 1024. Supported bounds depend on the selected image model: Z-Image/Z-Image Turbo, Dark Beast Z-Image Turbo, Chroma, and legacy/specialized image models support 256-2048 on either edge; Krea 2 Turbo, Dark Beast KREA 2, and Qwen image models support 256-2560 on either edge; One Obsession v22 supports 256-1920 on either edge; GPT Image 2 supports flexible dimensions up to 3840px on either edge with max 3:1 aspect ratio and a total pixel budget from 655,360 to 8,294,400. Set when the user specifies a width, exact pixel dimensions, or a named resolution (e.g., \"1280 wide\", \"1280x720\", \"720p\", \"1080x1920\", \"3840x2160\"). If the user gives only one dimension, set only that dimension and preserve/infer the sensible aspect ratio. User-requested dimensions override the default media quality, including Pro. Non-multiple-of-16 values are accepted when in bounds; the renderer snaps to the nearest supported size internally, so do not ask the user to adjust by a few pixels."
            },
            "height": {
              "type": "number",
              "description": "Output image height in pixels. Default: 1024. Supported bounds depend on the selected image model: Z-Image/Z-Image Turbo, Dark Beast Z-Image Turbo, Chroma, and legacy/specialized image models support 256-2048 on either edge; Krea 2 Turbo, Dark Beast KREA 2, and Qwen image models support 256-2560 on either edge; One Obsession v22 supports 256-1920 on either edge; GPT Image 2 supports flexible dimensions up to 3840px on either edge with max 3:1 aspect ratio and a total pixel budget from 655,360 to 8,294,400. Set when the user specifies a height, exact pixel dimensions, or a named resolution (e.g., \"720 high\", \"1280x720\", \"720p\", \"1080x1920\", \"2160x3840\"). If the user gives only one dimension, set only that dimension and preserve/infer the sensible aspect ratio. User-requested dimensions override the default media quality, including Pro. Non-multiple-of-16 values are accepted when in bounds; the renderer snaps to the nearest supported size internally, so do not ask the user to adjust by a few pixels."
            },
            "numberOfVariations": {
              "type": "number",
              "description": "Number of variations (1-16). Set to the user's exact requested count in one call whenever they ask for multiple images and the outputs can share project settings. Trigger phrasings: \"draw N\", \"make N\", \"give me N\", \"show me N\", \"render N\", \"create N\", \"generate N\", \"N more\", \"another N\", \"N as separate\", \"N separate images\", \"N different images\", \"N options\", \"N takes\", \"N versions\", \"N variations\", \"N pictures of\", \"all at the same time\", \"in parallel\", \"side by side as separate\". This includes selection-gated image batches that will feed a later dance, animation, or video after the user picks one. Avoid multiple serial generate_image calls unless the user explicitly wants separate projects, isolated approvals, or per-output settings that cannot share one project. If the user previously got a composite \"N subjects in one image\" result and now says \"draw N more as separate images\" / \"as separate\" / \"separately\", set numberOfVariations=N for THIS call — the prior call's numberOfVariations does not carry forward when the user explicitly asks for separation. For screenplay/storyboard batches, this should equal the scene count and the prompt should contain one Dynamic Prompt branch with one full scene prompt per scene; do not set numberOfVariations=N with only one scene prompt. Default: 1 when the user clearly wants a single composite image (e.g. \"draw 2 goats in a meadow\" with no separation language, or explicit \"in one image\" / \"single image\" / \"composite\" / \"sheet\").\n\nFor distinct image sets, numberOfVariations is the output count and the prompt must contain one Dynamic Prompt branch with the same option count. Do not satisfy a multi-output request by putting all requested items into one prompt; that makes every generated result contain the whole set.",
              "minimum": 1,
              "maximum": 16
            },
            "negativePrompt": {
              "type": "string",
              "description": "Things to avoid in the generated image. Only set when the user explicitly mentions what to avoid. E.g., \"no watermarks, no text, no blurry edges\"."
            },
            "starting_image_strength": {
              "type": "number",
              "description": "Image-to-image source guidance strength (0.0-1.0). Only set when a source image is available and the model supports img2img. For Z-Image/Z-Image Turbo, Krea 2 Turbo, or source-preserving enhancement requests, use 0.75 with sourceImageIndex so the source image remains a strong guide while allowing higher-resolution reconstruction. Use lower values only when the user explicitly asks for a lighter guide/subtle variation; higher values are more creative and can deviate further from the source."
            },
            "sourceImageIndex": {
              "type": "number",
              "description": "Which result image to use as starting image for img2img (0-based index). -1 = original upload. Omit to auto-select latest result. Only relevant when starting_image_strength is set."
            },
            "seed": {
              "type": "integer",
              "description": "Random seed for reproducibility. Use -1 for random (default). Set a specific seed when the user wants to reproduce a previous result."
            },
            "guidance": {
              "type": "number",
              "description": "Guidance scale override. Higher values = more prompt adherence. Model-specific defaults are used if omitted. Only set when the user explicitly requests a guidance value."
            },
            "loras": {
              "type": "array",
              "minItems": 1,
              "maxItems": 8,
              "items": {
                "type": "string",
                "minLength": 1
              },
              "description": "Ordered LoRA IDs to apply to a compatible image model. Use only when the user explicitly requests LoRAs or asks for an effect one of these names directly. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only by the five Krea 2 based models: krea2_turbo_fp8_scaled (text-to-image), krea2_identity_edit_v1_2 and krea2_identity_edit_sogni_v0_3_alpha (identity edit), and the dark_beast_krea2_fp8 / dark_beast_krea2_identity_edit_v1_2 community variants.\n\nBipolar sliders — each id names its POSITIVE direction, a negative strength applies the opposite, and 0 disables it: krea2-detail-enhancer, krea2-scene-complexity, krea2-realism (+ = photoreal), krea2-amateur, krea2-candid, krea2-zoom (+ = zoomed in), krea2-skin-detail, krea2-wetness, krea2-age, krea2-height, krea2-weight, krea2-hourglass-figure, krea2-breast, krea2-chest-firmness, krea2-nipple-projection, krea2-warm-light, krea2-afterlight (+ = golden), krea2-skin-tone (+ = darker), krea2-purple-grainy (+ = grainy and muted). Positive-only fine-tunes: krea2-realism-engine (photographic realism), krea2-bloomgirls (polished influencer look), krea2-mystic-x (uncensored adult), krea2-aberrant (industrial body horror), krea2-filter-bypass-2 and krea2-filter-bypass-3 (restore expressions, anatomy and poses the base model flattens; try the 2-vector first). Exact per-LoRA ranges, maturity flags and the full contract: GET /v1/loras/comfy?modelId=<model>. Personal imports use authenticated GET /v1/loras/personal/catalog: use an owned ready personal- id, its modelIds, strength range and requirements. Never invent ids or silently omit a requested personal LoRA."
            },
            "loraStrengths": {
              "type": "array",
              "minItems": 1,
              "maxItems": 8,
              "items": {
                "type": "number"
              },
              "description": "Strength for each LoRA in loras, in the same order. Omitting the array uses 1.0 for every LoRA, which is not each LoRA's catalog default — krea2-chest-firmness, krea2-nipple-projection and krea2-height default to 0 (no effect) — so prefer explicit values. Do NOT clamp to 0-1: most Krea 2 LoRAs are bipolar, so krea2-warm-light warms the grade at 2 and cools it at -2. Usable bands vary per LoRA — roughly -2..5 for krea2-detail-enhancer, -3..3 for krea2-warm-light, 3..9 for krea2-candid, 0.5..1 for krea2-realism-engine, 1..2 for the filter-bypass pair. Scale the magnitude to how strongly the user asked; the server clamps out-of-range values, and pushing past a LoRA's recommended band costs image quality rather than adding effect. Preserve explicit user values. Example: loras=[\"krea2-detail-enhancer\",\"krea2-amateur\"], loraStrengths=[3,-2]."
            },
            "gptImageQuality": {
              "type": "string",
              "enum": [
                "low",
                "medium",
                "high",
                "xhigh",
                "max"
              ],
              "description": "Optional GPT Image rendering quality. Only set it when the user explicitly asks for low/fast, medium/balanced, high/final, xhigh/extra high, or max/maximum quality; xhigh and max require GPT Image 2.5 Sunburst or Flare. Provider-chosen (auto) quality is never used. Otherwise omit it and let the host app media quality setting map Fast to low, HQ to medium, and Pro to high. The same quality label does not promise equivalent results across models."
            },
            "outputFormat": {
              "type": "string",
              "enum": [
                "png",
                "jpg",
                "jpeg",
                "webp"
              ],
              "description": "Optional output file format for generated images. Set only when the user explicitly requests PNG, JPG/JPEG, or WebP. Hosts should normalize \"jpeg\" to the Sogni project format \"jpg\"."
            },
            "aspectRatio": {
              "type": "string",
              "description": "Do NOT set unless the user explicitly requests an aspect ratio, format, orientation, or exact pixel dimensions. When a reference/source image is used and the user did not ask to change its shape, omit this field so the handler preserves the selected source image's own ratio.\n\nFormats: \"16:9\", \"9:16\", \"4:5\", \"1:1\", \"4:3\", \"3:2\", \"21:9\", or exact pixels like \"1920x1080\".\n\nCRITICAL: When the user specifies exact pixel dimensions (e.g., \"1280x720\", \"1080x1920\", \"1920x1080\", \"3840x2160\") or an orientation-qualified named resolution (e.g., \"720p landscape\", \"720p portrait\"), use the exact pixel format, NOT a ratio like \"16:9\" or \"9:16\". Exact user-requested dimensions override the selected default media quality, including Pro/HQ defaults. A bare named video resolution like \"720p resolution\" is only a resolution tier/short-side request; do not turn it into landscape pixels and do not set aspectRatio unless the user also states landscape, portrait, vertical, horizontal, or exact pixels. If requested pixels are in bounds but not on the model's pixel step, still pass the user's exact pixel request; the handler snaps to the nearest supported size internally. Only use ratio format when the user says a generic format name without pixel dimensions.\n\nMappings (use ONLY when user does NOT specify pixel dimensions): landscape/widescreen/YouTube/cinematic → \"16:9\". portrait → \"9:16\". TikTok/Reels/IG Reels → \"1080x1920\". ultrawide/cinema scope → \"21:9\". Instagram post → \"4:5\". square → \"1:1\". standard/TV → \"4:3\". 720p landscape → \"1280x720\". 720p portrait → \"720x1280\". 1080p landscape → \"1920x1080\". 1080p portrait/HD portrait → \"1080x1920\". 4K landscape → \"3840x2160\". 4K portrait → \"2160x3840\". Never set for generic requests like \"make a video\".\n\nSet this whenever the user specifies an image or downstream video orientation/aspect ratio such as 9:16, 16:9, portrait, vertical, landscape, widescreen, TikTok/Reels/Shorts, or exact pixels. This includes selection-gated image batches that will feed a later video or dance after the user picks one. For GPT Image 2 exact size requests, preserve exact pixel intent when possible and prefer popular GPT sizes such as 1536x1024, 1024x1536, 2048x1152, 3840x2160, and 2160x3840. Sogni does not expose transparent output for GPT Image 2.0. For transparent assets, select GPT Image 2.5 Sunburst or Flare with gptImageBackground=transparent and PNG or WebP output."
            },
            "gptImageBackground": {
              "type": "string",
              "enum": [
                "auto",
                "opaque",
                "transparent"
              ],
              "description": "GPT Image background: auto or opaque; transparent is supported by GPT Image 2.5 Sunburst and Flare with PNG or WebP output. JPEG cannot preserve transparency."
            },
            "gptImageOutputCompression": {
              "type": "integer",
              "minimum": 0,
              "maximum": 100,
              "description": "Optional GPT Image JPEG/WebP output compression, from 0 to 100. Omit for PNG."
            }
          },
          "required": [
            "prompt"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "generate_video",
        "description": "Generate a video from text or Seedance multimodal references. LTX 2.5 is the default and generates audio natively; LTX 2.3 remains available as rollback and also generates audio natively (dialogue, sounds, ambient music) — describe audio in the prompt. If the user provides exact speech, include it in double quotes; if they only imply speech, describe the performance and voice without inventing quoted words. Never use placeholders such as \"while speaking\", \"dialogue begins\", \"explaining\", or \"final line lands\". PERSONA VOICE: Only when the user explicitly asks to use/clone a registered persona voice clip, call resolve_personas first, then set voicePersonaName to select which persona's voice clip to use. Do not set voicePersonaName for ordinary character dialogue or inferred voices; describe those voices in the prompt for native LTX audio. For cross-persona narration (e.g. David narrates a video of Aleyna), resolve both personas and set voicePersonaName to the narrator only if that registered voice was requested. Persona voice requires ltx23 because LTX 2.5 has no compatible ID-LoRA and WAN 2.2 does not support voice identity. For non-Seedance syncing to a specific song or audio track, use sound_to_video instead. For non-Seedance animation from a locked source photo, use animate_photo. Do NOT use for My Personas unless generating a Seedance reference-based video — standard persona videos use resolve_personas → edit_image → animate_photo. SEEDANCE DEFAULT: For seedance2, seedance2-mini, or seedance2-5, default to exactly one video (4-15s on 2.0/Mini, 4-30s on seedance2-5) unless the user explicitly asks for multiple separate outputs. Multiple beats, shots, or scene descriptions in one up-to-15s Seedance prompt are still one video. If the user requests one continuous Seedance video longer than 15s, prefer seedance2-5, which renders up to 30s in a single call; beyond 30s (or on 2.0/Mini) preserve the requested total duration in the prompt/context and let chat orchestration split it into supported segment renders and stitch them instead of clamping it to a short excerpt. Uploaded/generated storyboard, shot-sheet, or trailer-concept images used as Seedance references should become one Seedance 2.5 generate_video call at targetResolution=1080 by default; preserve an explicitly requested model or resolution, and use Seedance 2.0 for 4K because Seedance 2.5 supports up to 1080p. Do not extract panels with edit_image and do not animate the storyboard sheet with LTX unless the user explicitly asks for separate non-Seedance clips. Seedance loose image, video, and audio references go through this tool; do not use animate_photo sourceImageIndex/frameRole/endImageIndex for Seedance. If an uploaded video is the source clip to transform, upscale, enhance, restyle, or remaster, use video_to_video with controlMode=\"seedance-v2v\" instead of generate_video referenceVideoIndices. If the uploaded audio is the primary sync target, lip-sync target, or requested as sound-to-video/audio-sync, use sound_to_video with videoModel=\"seedance2-mini\" instead of this tool unless the user asks for full Seedance. Use referenceAudioIndices here only when audio is a loose reference under an image/video-anchored Seedance shot. For Seedance, every image — first frame, last frame, or loose reference — is passed through referenceImageIndices (auto-uploaded as referenceImageUrls). Anchor frame intent in the prompt with @Image tags such as \"Use @Image1 as the opening shot reference. Begin the video with a composition, subject placement, lighting, mood, and camera framing that closely match @Image1.\" (or @Image2 as the final shot reference). For seamless-loop or \"first frame and last frame identical\" requests with a single uploaded image, anchor it explicitly as both: \"Use @Image1 as both the first frame and last frame so the video loops cleanly back to the opening composition.\" Assign each useful @Image/@Video/@Audio tag a role. APPROVED STORYBOARD PRODUCTION: When the user asks for a production workflow from an approved storyboard, the chat orchestrator should use the durable CampaignStoryboard contract: render and audit the composite board, generate one scene reference still per approved scene with GPT Image 2.5 Sunburst, render each scene as a Seedance 2.5 video segment at targetResolution=1080, then stitch the segments. Preserve an explicit alternate model or resolution. Do not replace that with a generic storyboard-reference video unless the user asks for a fast draft. PARTIAL VIDEO EDITS: Do NOT call generate_video to re-render an existing rendered/uploaded video just to change part of it (the bumper, the intro, the end card, a single scene, the last few seconds, etc.). Use replace_video_segment for that — it preserves the unchanged portion, keeps the original audio outside the replaced window, and costs far less. Likewise use extend_video to add new time to the end without rewriting the rest. If the request is vague, ask about vision/mood/style first. Only call once you have clear creative intent. WAN 3 uses the exact selector wan3.0-video: use this tool for text-to-video or loose Image 1/Video 1/Audio 1 references; use animate_photo for native first/last frames and sound_to_video when audio drives timing. A Wan 3 video reference conditions a new generation; it does not invoke a provider-backed edit or extend mode. Wan 3.0 Enhanced uses Sogni selector wan3.0-spicy-video and MuleRouter provider ID w3.0-video; use this tool for prompt-only or loose-reference generation, and never combine loose references with frame anchors.",
        "parameters": {
          "type": "object",
          "properties": {
            "prompt": {
              "type": "string",
              "description": "Write one flowing paragraph like a cinematographer describing a shot. Present tense, specific natural language. Longer clips need longer prompts; close-ups need more detail than wide shots.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing. Set skipPromptProcessing=true; for Seedance or Wan 3 also set expandPrompt=false.\n\nSTRUCTURE: shot/style → subject (age, clothing, hairstyle, distinguishing details) → environment, lighting, atmosphere → action beat by beat → camera movement → audio and dialogue.\n\nCAST CONTINUITY: For screenplay, script, storyboard, commercial, series, or other longer-form video tasks with recurring characters, use stable character names and repeat the same visual anchors every time they appear (age range, build, hairstyle, outfit silhouette, color palette, signature prop/accessory, posture, voice). Do not rename, merge, redesign, or drift characters between scenes unless the user asks.\n\nMOTION PACING: Scale complexity to duration. <=6s: 1 main action beat + 1 simple camera move. Around 10s: 2-3 clear action beats + 1 camera move. >10s: up to 4 action beats in clear sequence. Prefer fewer readable beats over dense micro-actions, especially in short clips.\n\nBLOCKING: Direct the layout like scene blocking. State left/right placement, foreground/background, facing toward/away, and relative distance when multiple subjects or important objects are involved.\n\nACTION: Drive motion with concrete verbs. Specify who moves, what moves, how it moves, and what the camera does. Avoid generic phrases like \"comes alive.\"\n\nDIALOGUE: Put user-provided spoken lines in double quotes. For screenplay-style or longer-form tasks, prefix each spoken line with a stable speaker tag outside the quotes, e.g. CHARACTER: \"We made it.\" Break long speech into short quoted phrases with acting beats between them (gestures, pauses, glances). If the user asks for speech but provides no exact words, describe the visible delivery, voice quality, and emotion without inventing quoted dialogue; ask only when exact wording is the point of the request. Never write placeholders such as \"while speaking\", \"dialogue begins\", \"explaining\", or \"final line lands\". Show emotion through visible behavior — not \"she is sad\", instead \"she looks down, pauses, and her voice cracks\". QUOTING RULE: ONLY use double quotes for spoken dialogue. Never quote on-screen text, overlay text, titles, captions, signs, or any visual text — describe them without quotes.\n\nSTORYBOARD TEXT: For storyboard references, structural headings, section numbers, slide titles, panel titles, and captions may become short audio-only narration/voiceover or key-message beats, but they are not subtitles, title cards, lower thirds, or visible overlays unless the user explicitly asks for visible text/on-screen text/title card/subtitle/lower third/signage/CTA. Do not concatenate storyboard labels into run-on voiceover; use separate brief phrases with pauses.\n\nAUDIO: Prompt sound intentionally — voice quality, volume, room tone, ambience, music, weather, footsteps. Include language or accent if relevant. Useful voice/volume anchors: whisper, mutter, shout, scream, energetic announcer, resonant voice with gravitas, distorted radio-style, robotic monotone, childlike curiosity.\n\nCAMERA: Cinematic terms — close-up, tracking shot, dolly in, handheld, slow arc, static frame. Describe movement relative to subject.\n\nFor specific characters (movies, TV): describe visual appearance — don't rely on names alone.\n\nFor complex/creative scenes (characters, dialogue, skits): capture the full creative intent. The system auto-expands into a detailed prompt.\n\nAVOID: Vague prompts, too many characters at once, conflicting lighting logic, readable text or logos, abstract emotions with no visible behavior, rigid numeric constraints (exact angles, counts, speeds).\n\nNON-SEEDANCE POSITIVE CONSTRAINTS: For videoModel=\"ltx25\", \"ltx23\", or \"wan22\", prompt is a positive prompt. Translate user avoid/no/don't constraints into affirmative production constraints instead of copying negative phrasing. Preserve exact quoted visible text or dialogue when the user explicitly requests it; keep surrounding surfaces blank.\n\nBATCH VARIATIONS: When numberOfVariations > 1, use Dynamic Prompt syntax. This is one Sogni project with multiple jobs, so prefer it when all outputs share the same references, model, duration, dimensions, and generation parameters and only prompt text varies. Lock in any camera/subject/style the user specified, vary the rest. Example: \"slow dolly in on a city street {at dawn with golden light|during a rainstorm|at night with neon reflections}\"."
            },
            "expandPrompt": {
              "type": "boolean",
              "description": "Seedance and Wan 3 only. Whether to expand the prompt before dispatch. Defaults to true. For Wan 3, a successful Sogni expansion disables Alibaba prompt_extend to prevent a second rewrite; false disables both expansion layers so exact prompts remain exact."
            },
            "skipPromptProcessing": {
              "type": "boolean",
              "description": "Bypass automatic prompt shaping/refinement and voice-identity prompt formatting so the prompt text is sent unchanged to the video model. Set true ONLY when the user explicitly says not to modify/rewrite/enhance/expand/change/improve the prompt, or to use/send it exactly, verbatim, or as-is, AND the provided prompt already satisfies the tool requirements. Continue to set non-prompt parameters such as model, duration, count, aspect ratio, and seed. For Seedance or Wan 3 literal prompt requests, also set expandPrompt=false. Do not set for ordinary underspecified requests."
            },
            "duration": {
              "type": "number",
              "description": "Video duration in seconds. Default: 5. Per-model range: LTX 2.3 and LTX 2.5 = 2-20s; Wan 3 = 2-30s; Seedance 2.0 and Mini = 4-15s; Seedance 2.5 = 4-30s; HappyHorse 1.1 = 3-15s. Use when the user explicitly requests a specific length. MiniMax H3 is quantized to a 17-frame grid at a fixed 24 fps and renders 124-362 frames, so an H3 clip runs 5.17-15.08 seconds and a requested length outside that window snaps to the nearest valid H3 length.",
              "minimum": 2,
              "maximum": 30
            },
            "ratio": {
              "type": "string",
              "enum": [
                "adaptive",
                "16:9",
                "4:3",
                "1:1",
                "3:4",
                "9:16"
              ],
              "description": "Wan 3 only. Output ratio. Use \"adaptive\" to derive the shape from input/context media. Omit for automatic behavior."
            },
            "watermark": {
              "type": "boolean",
              "description": "Alibaba wan3.0-video only. Add the visible watermark. Defaults to false."
            },
            "referenceFileUrl": {
              "type": "string",
              "description": "Alibaba wan3.0-video only. One public HTTPS document URL for context (DOCX/DOC/XLSX/XLS/PPTX/PPT/PDF/TXT/KEY/PAGES/NUMBERS/Markdown, up to 100 MB; PDF/DOCX/DOC/PPTX/PPT/KEY/PAGES up to 50 pages). Mutually exclusive with referenceLinkUrl and first/last-frame inputs."
            },
            "referenceLinkUrl": {
              "type": "string",
              "description": "Alibaba wan3.0-video only. One public HTTPS webpage URL for context. Mutually exclusive with referenceFileUrl and first/last-frame inputs."
            },
            "negativePrompt": {
              "type": "string",
              "description": "Advanced LTX/WAN only. Use this field only when the user explicitly asks to set a separate negative prompt. MiniMax H3 has no negative-prompt input; put requested exclusions in prompt. Do not set for MiniMax H3, Seedance, or HappyHorse.\n\nWan 3 has no negativePrompt request field; do not set this for wan3.0-video."
            },
            "videoModel": {
              "type": "string",
              "enum": [
                "ltx25",
                "ltx23",
                "wan22",
                "seedance2",
                "seedance2-mini",
                "seedance2-5",
                "minimax-h3-t2v",
                "minimax-h3-t2v-turbo",
                "minimax-h3-fasth3-t2v-turbo",
                "minimax-h3-fasth3-t2v-turbo-2stage",
                "happyhorse-1.1-t2v",
                "happyhorse-1.1-i2v",
                "happyhorse-1.1-r2v",
                "minimax-h3-r2v",
                "minimax-h3-r2v-turbo",
                "minimax-h3-r2v-2stage",
                "minimax-h3-r2v-balanced-2stage",
                "wan3.0-video",
                "wan3.0-spicy-video"
              ],
              "description": "\"ltx25\" (default): LTX 2.5 with native audio; Fast, HQ, and Pro currently use the release-validated official Distilled INT8 workflow. The Dev checkpoints are not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Video model. \"ltx23\": LTX 2.3 rollback with native audio. \"wan22\": quick simple motion without audio. \"minimax-h3-t2v\": standard 20-step MiniMax H3 text-to-video; \"minimax-h3-t2v-turbo\": the existing 4-step LightX2V Turbo text-to-video; \"minimax-h3-fasth3-t2v-turbo\": the separate FastVideo VSA four-step FastH3 engine, about 2x faster and fixed to Euler/simple; \"minimax-h3-fasth3-t2v-turbo-2stage\": the two-stage FastH3 engine: FastH3 renders the canvas, then the worker enlarges it 2x and refines it, so the clip is delivered at twice the canvas width and height with the same length and audio. targetResolution picks the delivered class: 1080 renders a 544px short-edge canvas (960x544 is delivered at 1920x1088) for 10 Spark per second, 1440 or omitted renders the 768p canvas for 2K (1344x768 is delivered at 2688x1536) for 16 Spark per second, and 720 renders a 384px canvas (672x384 is delivered at 1344x768) for the regular FastH3 price of 4 Spark per second. The estimate prices every request. It takes the same inputs, durations and LoRAs as \"minimax-h3-fasth3-t2v-turbo\". Choose it when the user asks for 1080p, 1440p or 2K MiniMax H3 output, for two-stage output, or for the sharpest/best H3 quality; for ordinary 768p FastH3 output keep the regular FastH3 selector at targetResolution 768. All use native audio, fixed 24fps, 5.17-15.08s, and a 768p-class 32px-grid canvas; use animate_photo for H3 image-conditioned modes. Base and Turbo T2V/I2V/FLF2V prompts use the exact ordered fields integrated_multimodal_description, overall_soundscape, and non_diegetic_music; I2V/FLF2V prepend the official alignment line. \"minimax-h3-r2v\": standard 20-step MiniMax H3 reference-to-video; \"minimax-h3-r2v-turbo\": the dedicated LightX2V 4-step Ref2VA Turbo workflow using Euler/simple and a 960x544 default. \"minimax-h3-r2v-2stage\" (Standard, 20 steps, res_multistep/simple) and \"minimax-h3-r2v-balanced-2stage\" (Balanced, 8 steps, Euler/simple) are the two-stage reference-to-video forms: the tier renders the canvas, then the worker enlarges it 2x and refines it, so the clip is delivered at twice the canvas width and height with the same length, audio and references. targetResolution names the delivered class exactly as for the FastH3 two-stage selectors: 1080 renders a 544px short-edge canvas (960x544 is delivered at 1920x1088), 1440 or omitted renders the 768p canvas for 2K (1344x768 is delivered at 2688x1536), and 720 renders a 384px canvas (672x384 is delivered at 1344x768). Each bills its tier's own per-second rate plus the two-stage surcharge of the delivered class (none at 720p); the estimate prices every request. Choose one when the user asks for 1080p, 1440p or 2K MiniMax H3 reference-to-video output; keep the one-stage \"minimax-h3-r2v\" or Balanced selector for ordinary 768p output. FastH3 has no R2V mode. Every R2V selector accepts up to 9 images, 3 videos, and 3 audios (12 files total); at least one visual reference (image or video) is required and audio alone is invalid. Select references with referenceImageIndices/referenceVideoIndices/referenceAudioIndices and address them with the official <Subject N>/<Picture N>/<Video N>/<Audio N> semantics. Seedance quality is selected only by model: use \"seedance2-mini\" for Seedance 2.0 Mini or faster/lower-cost 720p iteration, use \"seedance2-5\" at targetResolution=1080 for generated/uploaded storyboard images and current high-quality Seedance generation unless the user explicitly requests another model or resolution, and use \"seedance2\" for Seedance 2.0 or 4K requests because Seedance 2.5 supports up to 1080p. Do not use Default Media Quality Fast/HQ/Pro to represent Seedance quality. \"seedance2-5\": Seedance 2.5, the newest Seedance generation — 480p, 720p, and 1080p (4K is unsupported), 4-30s per clip at a fixed 24 fps, native audio, first-and-last-frame conditioning, and a much larger reference budget than the 2.0 family: up to 30 images, 10 videos, and 10 audios, with up to 50 reference media files total, subject to those per-modality caps. Choose \"seedance2-5\" when the user asks for Seedance 2.5, wants a single continuous Seedance clip longer than 15s (2.5 renders up to 30s in one call instead of being split and stitched), or wants a first-and-last-frame Seedance transition. Keep \"seedance2\" for 4K requests; Seedance 2.5 supports up to 1080p. Seedance supports multimodal loose reference assets. Seedance 2.0 and Mini accept up to 9 images, 3 videos, and 3 audios, with no more than 12 asset files total. Seedance 2.5 accepts up to 30 images, 10 videos, and 10 audios, with up to 50 reference media files total, subject to those per-modality caps. Use @Image1/@Video1/@Audio1 style references in creative briefs when assigning roles. Assign every useful reference asset a role and prefer positive preservation constraints. If an uploaded video is the source clip to transform, upscale, enhance, restyle, or remaster, use video_to_video with controlMode=\"seedance-v2v\" instead of generate_video referenceVideoIndices. Alibaba HappyHorse 1.1 video models (third-party vendor — requires Premium Spark). Select by mode: \"happyhorse-1.1-t2v\" for text-to-video, \"happyhorse-1.1-i2v\" for image-to-video from one first-frame image, and \"happyhorse-1.1-r2v\" for reference-to-video with up to 9 reference images. Resolutions 720P and 1080P; duration 3-15 seconds at 24 fps; native synchronized audio is always generated (do not set generateAudio or negativePrompt). Supported aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9. HappyHorse 1.1 takes image references only and renders a native synchronized audio track (always on; do not set generateAudio or a negative prompt). Pick the model by mode: happyhorse-1.1-t2v for text-to-video (no reference image), happyhorse-1.1-i2v for image-to-video from a single first frame, and happyhorse-1.1-r2v for reference-to-video with 1 to 9 reference images. For r2v, tag the images in the prompt as [Image 1]…[Image 9] and assign each a clear role. HappyHorse does not accept reference videos or reference audios. \"wan3.0-video\" is Alibaba Wan 3 and \"wan3.0-spicy-video\" is MuleRouter w3.0-video. Both render 2-30s at fixed 30 fps with optional native audio, provider prompt expansion, 480p/720p/1080p, adaptive/fixed ratios, first/last frames, and up to 10 image/5 video/5 audio references. Only Alibaba wan3.0-video accepts document/web context and watermark. Frame anchors and loose references are mutually exclusive. Do not send negativePrompt; video references are loose conditioning for a new result, not source-video editing or extension."
            },
            "generateAudio": {
              "type": "boolean",
              "description": "Whether the returned video should include generated/native audio. Omit to include audio by default; set false only when the user explicitly asks for silent output or no audio. Supported by LTX, MiniMax H3, and Seedance; not supported by WAN or HappyHorse.\n\nWan 3 supports this toggle; omit it for audio-on by default or set false only for an explicitly silent result."
            },
            "referenceImageIndices": {
              "type": "array",
              "items": {
                "type": "number"
              },
              "description": "Seedance or MiniMax H3 r2v image references. Use negative indices for uploaded images and non-negative indices for generated image results. Seedance uses @Image tags. H3 Ref2VA requires at least one visual reference—an image or video—and uses the official <Subject N>/<Picture N> semantics; these are loose references, not locked first frames. Wan 3 loose images use Image 1, Image 2, and so on, with up to 10 images."
            },
            "referenceVideoIndices": {
              "type": "array",
              "items": {
                "type": "number"
              },
              "description": "Seedance or MiniMax H3 r2v loose video references. Use negative indices for uploaded videos and non-negative indices for generated video results. Seedance uses @Video tags; H3 uses <Video 1>, <Video 2>, and so on in selection order. A video may be the only H3 Ref2VA visual reference. Do not use this for source-video transforms; use video_to_video instead. Wan 3 loose videos use Video 1, Video 2, and so on, with up to 5 videos."
            },
            "referenceAudioIndices": {
              "type": "array",
              "items": {
                "type": "number"
              },
              "description": "Seedance or MiniMax H3 r2v loose audio references. Use negative indices for uploaded audio files and non-negative indices for generated audio results. Seedance uses @Audio tags; H3 uses <Audio 1>, <Audio 2>, and so on in selection order. H3 Ref2VA audio may accompany an image or video, but audio alone is invalid. Wan 3 loose audios use Audio 1, Audio 2, and so on, with up to 5 audios."
            },
            "width": {
              "type": "number",
              "description": "Video width in pixels. LTX 2.3: 640-3840. WAN: 480-1536. Default resolution depends on model and quality tier: LTX Fast about 720p and High/Pro about 1080p; WAN Fast uses 480p short side and High/Pro uses 720p short side. Set width only when the user specifies an exact width or orientation-qualified exact pixels. A bare named resolution like \"720p resolution\" is a short-side target, not an instruction to make landscape 1280x720. If the user gives only one exact dimension, set only that dimension and preserve/infer the sensible aspect ratio. User-requested exact dimensions override the default media quality. Mappings when orientation is explicit: 480p landscape=854x480, 480p portrait=480x854, 720p landscape=1280x720, 720p portrait=720x1280, 1080p landscape=1920x1080, 1080p portrait=1080x1920, 4K landscape=3840x2160. Non-step values are accepted when in bounds; LTX snaps to the nearest 64px step and WAN snaps to the nearest 16px step internally, so do not ask the user to adjust by a few pixels."
            },
            "height": {
              "type": "number",
              "description": "Video height in pixels. LTX 2.3: 640-3840. WAN: 480-1536. Set height only when the user specifies an exact height or orientation-qualified exact pixels. A bare named resolution like \"720p resolution\" is a short-side target; do not convert it to landscape dimensions unless the user says landscape/horizontal/widescreen. If the user gives only one exact dimension, set only that dimension and preserve/infer the sensible aspect ratio. User-requested exact dimensions override Default Media Quality, including Pro. Non-step values are accepted when in bounds; LTX snaps to the nearest 64px step and WAN snaps to the nearest 16px step internally, so do not ask the user to adjust by a few pixels."
            },
            "targetResolution": {
              "type": "number",
              "description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", \"1080p\", \"2160p\", or \"4K\" without exact pixels or an output orientation. This is resolution only, not a Seedance quality tier: Seedance quality is selected by videoModel (\"seedance2\" vs \"seedance2-mini\" vs \"seedance2-5\"). Seedance 2.0 full supports 4K; Seedance Mini supports 480p/720p; Seedance 2.5 supports 480p/720p/1080p, so never set 4K for \"seedance2-5\". Wan 3 supports exactly 480p, 720p, and 1080p. HappyHorse supports only 720p and 1080p. Never set 4K for Wan 3 or HappyHorse. MiniMax H3 renders inside a 1344x768 pixel budget on a 32px grid, so use 768 for the regular H3 selectors and never 1080p or 4K. The two-stage H3 selector \"minimax-h3-fasth3-t2v-turbo-2stage\" delivers twice the canvas, so there targetResolution names the delivered short-edge class: 1080 (544px canvas short edge: 960x544 delivered at 1920x1088), 1440 for 2K (the 1344x768 canvas delivered at 2688x1536), or 720 (384px canvas: 672x384 delivered at 1344x768); omit it for 2K. Never set 4K for H3. Do not set targetResolution from Default Media Quality Fast/HQ/Pro. If omitted for Seedance, Wan 3, HappyHorse, or MiniMax H3, the host uses the selected model default. This preserves/inherits the current video shape instead of forcing landscape. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\", \"720p landscape\", \"4K portrait\", or \"4K landscape\", use exact width/height/aspectRatio instead."
            },
            "numberOfVariations": {
              "type": "number",
              "description": "Number of variations (1-16). Use with one Dynamic Prompt branch for multiple prompt-only takes that share the same references, model, duration, dimensions, and parameters. This creates one Sogni project with multiple jobs. Use 1 unless the user explicitly requests multiple separate video outputs. For Seedance, default to 1 unless the user explicitly requests separate outputs.",
              "minimum": 1,
              "maximum": 16
            },
            "aspectRatio": {
              "type": "string",
              "description": "Do NOT set unless the user explicitly requests an aspect ratio, format, orientation, or exact pixel dimensions. When a reference/source image is used and the user did not ask to change its shape, omit this field so the handler preserves the selected source image's own ratio.\n\nFormats: \"16:9\", \"9:16\", \"4:5\", \"1:1\", \"4:3\", \"3:2\", \"21:9\", or exact pixels like \"1920x1080\".\n\nCRITICAL: When the user specifies exact pixel dimensions (e.g., \"1280x720\", \"1080x1920\", \"1920x1080\", \"3840x2160\") or an orientation-qualified named resolution (e.g., \"720p landscape\", \"720p portrait\"), use the exact pixel format, NOT a ratio like \"16:9\" or \"9:16\". Exact user-requested dimensions override the selected default media quality, including Pro/HQ defaults. A bare named video resolution like \"720p resolution\" is only a resolution tier/short-side request; do not turn it into landscape pixels and do not set aspectRatio unless the user also states landscape, portrait, vertical, horizontal, or exact pixels. If requested pixels are in bounds but not on the model's pixel step, still pass the user's exact pixel request; the handler snaps to the nearest supported size internally. Only use ratio format when the user says a generic format name without pixel dimensions.\n\nMappings (use ONLY when user does NOT specify pixel dimensions): landscape/widescreen/YouTube/cinematic → \"16:9\". portrait → \"9:16\". TikTok/Reels/IG Reels → \"1080x1920\". ultrawide/cinema scope → \"21:9\". Instagram post → \"4:5\". square → \"1:1\". standard/TV → \"4:3\". 720p landscape → \"1280x720\". 720p portrait → \"720x1280\". 1080p landscape → \"1920x1080\". 1080p portrait/HD portrait → \"1080x1920\". 4K landscape → \"3840x2160\". 4K portrait → \"2160x3840\". Never set for generic requests like \"make a video\"."
            },
            "voicePersonaName": {
              "type": "string",
              "description": "ONLY when the user explicitly requests a registered/reference persona voice clip. Name of the persona whose voice clip to use as referenceAudioIdentity. Set this when the narrator/speaker is a different persona than the one described in the video (e.g. \"David\" narrates a scene featuring Aleyna), or to explicitly select a requested voice when multiple personas with voice clips are resolved. Do NOT set this for ordinary character dialogue, inferred voices, or personas without a voice clip — LTX 2.3 generates voice natively from the text prompt instead. Requires ltx23 because LTX 2.5 has no compatible ID-LoRA."
            },
            "loras": {
              "type": "array",
              "minItems": 1,
              "maxItems": 8,
              "items": {
                "type": "string",
                "minLength": 1
              },
              "description": "Ordered LoRA IDs to apply to a MiniMax H3 render. Use only when the user explicitly asks for a LoRA or for an effect one of these names describes. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-t2v\", \"minimax-h3-t2v-turbo\", \"minimax-h3-fasth3-t2v-turbo\", \"minimax-h3-fasth3-t2v-turbo-2stage\", \"minimax-h3-r2v\", \"minimax-h3-r2v-turbo\", \"minimax-h3-r2v-2stage\", \"minimax-h3-r2v-balanced-2stage\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nFive LoRAs are published for MiniMax H3 today and the set differs by mode, so GET /v1/loras/comfy?modelId=<model> is authoritative for the mode in hand and carries exact ranges, maturity flags, and anything published since. h3-realism-people (fal) is a realism pass trained on live-action footage of people: it restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It is the only one gated on a trigger word — put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. h3-vbvr-video-reasoning is a prompt-adherence pass that holds the model to what was asked instead of improvising. h3-natural-face-speech (AdaptiveVision) makes people talking on camera look and sound more natural: cheeks, brows, jaw and lips move together as in real speech, and spoken English comes through clearer; use it for talking-head shots such as vlogs, podcasts, interviews and presenters. h3-better-motion (AdaptiveVision) gives people more natural, consistent body movement — weight shifts, strides, turns and gestures that follow through — for dance, sport, walking and other full-body shots. Both AdaptiveVision LoRAs work best with short, simple prompt sentences and are not validated on reference-to-video. h3-mystic-xxx-v4 is an uncensored adult fine-tune. Personal imports are discovered through authenticated GET /v1/loras/personal/catalog; use only owned ready ids with the selected model in modelIds, and respect their strength range and requirements. Do not invent ids."
            },
            "loraStrengths": {
              "type": "array",
              "minItems": 1,
              "maxItems": 8,
              "items": {
                "type": "number"
              },
              "description": "Strength for each LoRA in loras, in the same order. Omitting the array applies 1.0 to every LoRA, which is NOT the catalog default and for h3-realism-people is already at the top of its band, so send explicit values. Video LoRAs are positive-only — unlike the bipolar Krea 2 image sliders, a negative value is not an inverse effect and 0 is off. h3-realism-people takes 0-2 and its catalog default is 0.8; 0.6-1 is the usable band. It also pulls the camera in as it climbs: at 1.5 and above the shot reliably recomposes and the grade darkens, which on an image-conditioned mode can crop the subject out of the frame the user supplied. Raise it above 1 only when the user asks for more, and prefer the default when they supplied a first or last frame. h3-vbvr-video-reasoning and h3-mystic-xxx-v4 both take 0-1 and do default to 1.0, with usable bands of 0.7-1 and 0.2-1. h3-natural-face-speech and h3-better-motion take 0-1.5 and default to 0.6; their usable band is 0.4-0.8."
            },
            "outputFormat": {
              "type": "string",
              "enum": [
                "mp4",
                "mov"
              ],
              "description": "Video container. Defaults to mp4. MOV is supported only by Seedance 2.5; choose it when the user requests MOV for editing."
            },
            "returnLastFrame": {
              "type": "boolean",
              "description": "Seedance 2.5 only. Set true to export a separate image of the final frame alongside the video. The result includes lastFrameUrl, which can be used as the first-frame image for a subsequent clip. Defaults to false; this does not extend the video automatically."
            }
          },
          "required": [
            "prompt"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "generate_music",
        "description": "Generate music from a text description. Creates original songs with optional lyrics, BPM, key signature, and duration control. Use when the user wants to create music, a song, a beat, a melody, background music, or any audio content.",
        "parameters": {
          "type": "object",
          "properties": {
            "prompt": {
              "type": "string",
              "description": "Genre, mood, and style description for the music. Be specific about musical characteristics.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing.\n\nExamples:\n- \"upbeat electronic dance music with driving bass and synth arpeggios\"\n- \"mellow jazz ballad with soft piano, brushed drums, and walking bass\"\n- \"epic orchestral soundtrack with soaring strings and powerful brass\"\n- \"lo-fi hip hop beat with vinyl crackle, muted keys, and chill vibes\"\n- \"acoustic folk song with fingerpicked guitar and warm harmonies\"\n\nInclude:\n- Genre (rock, jazz, electronic, classical, hip-hop, etc.)\n- Mood (happy, melancholic, energetic, relaxing, epic, etc.)\n- Instruments (piano, guitar, drums, synth, strings, etc.)\n- Style descriptors (driving, mellow, atmospheric, punchy, etc.)\n\nMODEL \"music3\": MiniMax Music 3 wants a structured caption instead of a tag list — write the prompt as one paragraph in three labeled parts: \"Global Metadata: genre, BPM, key, emotional progression across the song, production profile. Vocal Details: gender, timbre, delivery, harmonies (or: none, purely instrumental). Arrangement: primary and secondary instruments, groove, bass, percussion, textures, how sections evolve.\" The more specific, the closer the result. Fold tempo and key into this caption — music3 ignores the bpm/keyscale/timesig args.\n\nBATCH VARIATIONS: When numberOfVariations > 1, use Dynamic Prompt syntax to vary ONE dimension across separate tracks. This is one Sogni project with multiple jobs, so prefer it when all tracks share the same duration, BPM, key, lyrics, model, and generation parameters and only prompt text varies. Lock in any genre/mood/instruments the user specified, vary the rest. Example: \"{lo-fi hip hop beat with muted keys|jazz piano trio with brushed drums|ambient electronic with soft pads} with warm reverb and vinyl texture\"."
            },
            "duration": {
              "type": "number",
              "description": "Duration in seconds. Default: 30. Range: 10-600 (10 seconds to 10 minutes). Short clips: 10-30s. Standard songs: 120-300s.",
              "minimum": 10,
              "maximum": 600
            },
            "bpm": {
              "type": "number",
              "description": "Beats per minute / tempo. Default: 120. Range: 30-300. Slow ballad: 60-80. Mid-tempo: 90-120. Upbeat: 120-140. Fast dance: 140-180. Very fast: 180+.",
              "minimum": 30,
              "maximum": 300
            },
            "keyscale": {
              "type": "string",
              "description": "Musical key and scale. E.g., \"C major\", \"A minor\", \"F# minor\", \"Bb major\". Default: \"C major\". Only set when the user specifies a key or when a particular mood calls for it (minor keys for sad/dark, major for happy/bright)."
            },
            "lyrics": {
              "type": "string",
              "description": "Song lyrics. Optional — omit for instrumental music. Format: write lyrics naturally with line breaks. The model will attempt to sing these lyrics with the generated music. Works best with clear, rhythmic phrasing that matches the BPM. For model \"music3\", structure lyrics with plain section tags on their own lines ([Intro], [Verse], [Pre-Chorus], [Chorus], [Post-Chorus], [Bridge], [Solo], [Outro]) — no modifiers inside brackets, and write enough sections to fill the requested duration since the composer ends the song when the lyric sheet runs out. For instrumental music3 tracks, pass ONLY a skeleton of those tags one per line (e.g. [Intro] [Verse] [Chorus] [Verse] [Solo] [Chorus] [Outro]) — bare instrumental pieces end early without it."
            },
            "model": {
              "type": "string",
              "enum": [
                "turbo",
                "sft",
                "music3"
              ],
              "description": "Music model. \"music3\" (default): MiniMax Music 3 — premium autoregressive composer with the best vocals, lyric adherence and song structure; 30 steps, up to 5 minutes, and it treats duration as a ceiling (may end the song early at a musical resolution). BPM/key/timesig args are ignored by music3 — fold tempo and key into the prompt instead. \"turbo\": ACE-Step 1.5 Turbo — fast 4-16 step drafts at roughly 1/20 the music3 cost; use only when the user asks for a quick, cheap, or draft track, or names ACE-Step. \"sft\": ACE-Step 1.5 SFT — experimental, strong lyric handling, 10-200 steps; use only when the user names it. Default to \"music3\" whenever the user does not ask for a draft or a specific model."
            },
            "timesig": {
              "type": "number",
              "enum": [
                2,
                3,
                4,
                6
              ],
              "description": "Time signature (beats per measure). 4 = 4/4 time (default, most common). 3 = 3/4 time (waltz). 2 = 2/4 time (march). 6 = 6/8 time (compound). Default: 4."
            },
            "numberOfVariations": {
              "type": "number",
              "description": "Number of variations (1-16). Use with one Dynamic Prompt branch when the user requests multiple prompt-only music variations that share the same duration, BPM, key, lyrics, model, and parameters. This creates one Sogni project with multiple jobs. Default: 1.",
              "minimum": 1,
              "maximum": 16
            }
          },
          "required": [
            "prompt"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "generate_speech",
        "description": "Turn written words into spoken audio. Use when the user wants something read aloud, a narration, a voiceover, a line of dialogue, an audiobook passage, or a voice cloned from a recording they uploaded. This tool speaks text; it does not compose music — use generate_music for songs and instrumentals.",
        "parameters": {
          "type": "object",
          "properties": {
            "prompt": {
              "type": "string",
              "description": "The exact words to be spoken, verbatim. This is NOT a description of the audio: whatever is written here is read aloud character for character. \"a calm woman reading the news\" would be spoken as those seven words — put that in voiceDescription instead and write the actual news copy here.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing.\n\nPunctuation is prosody: full stops, commas, question marks and ellipses control pauses and intonation, so keep them. Write numbers, dates, currency and abbreviations the way they should be said (\"nineteen eighty-four\", \"twelve dollars fifty\", \"Doctor Chen\") when the plain form would be ambiguous. Line breaks are not pauses; use punctuation.\n\nIf the user asks for a script rather than supplying one — \"read me a poem about the sea\", \"record an intro for my podcast\" — write the words first and pass them here. compose_script is for longer or structured pieces; a sentence or two you can simply write.\n\nLimit 4096 characters, roughly five minutes of speech. A longer script is refused rather than cut off mid-sentence, so split it across several calls at a natural break.",
              "maxLength": 4096
            },
            "model": {
              "type": "string",
              "enum": [
                "voice",
                "clone",
                "design"
              ],
              "description": "Which speech model to use, chosen by what the user gave you. \"voice\" (default): one of nine studio voices, optionally restyled by voiceDescription — the right choice for narration, voiceover and dialogue when no particular person is being imitated. \"clone\": reproduces a specific voice from a recording the user uploaded; requires voiceSourceIndex. \"design\": invents a speaker who does not exist from a written description; requires voiceDescription and takes no recording. Pick \"clone\" whenever the user uploads a voice clip and asks for that person, and \"design\" when they describe a speaker instead of choosing one."
            },
            "voice": {
              "type": "string",
              "enum": [
                "serena",
                "vivian",
                "uncle_fu",
                "ryan",
                "aiden",
                "ono_anna",
                "sohee",
                "eric",
                "dylan"
              ],
              "description": "Studio voice to speak in. Only used when model=\"voice\". Female: serena (English), vivian (Chinese), ono_anna (Japanese), sohee (Korean). Male: ryan, aiden, eric, dylan (English), uncle_fu (older, Chinese). Every voice speaks all supported languages — the note above describes the accent it carries, not a limit on what it can read. Default: serena."
            },
            "voiceDescription": {
              "type": "string",
              "description": "How the line should be delivered, or who should deliver it. With model=\"voice\" this restyles the chosen voice without changing who it is: \"whispering, close to the microphone\", \"furious\", \"reading a bedtime story\", \"like a sports commentator\". With model=\"design\" this is required and describes a speaker to invent: age, gender, accent, timbre, pace, mood, and recording space — \"a warm, unhurried narrator in her forties with a faint Scottish lilt, close-miked in a quiet room\". Describe one coherent person; contradictory directions produce a voice that shifts mid-sentence. Not accepted with model=\"clone\", where the recording defines the voice. Limit 512 characters.",
              "maxLength": 512
            },
            "voiceSourceIndex": {
              "type": "number",
              "description": "Which audio holds the voice to clone, using the same numbering every other tool uses for audio: negative indices are uploads (-1 = first/primary upload, -2 = second upload, and so on) and 0-based non-negative indices are audio generated earlier in this conversation. A clip the user just uploaded is -1. Required when model=\"clone\" and ignored otherwise. The clip should be three to thirty seconds of one person speaking cleanly, with no music, no second speaker and no heavy room echo; anything past thirty seconds is trimmed."
            },
            "voice_source_url": {
              "type": "string",
              "description": "REST alternative to voiceSourceIndex: retrievable original recording for clone mode."
            },
            "voiceTranscript": {
              "type": "string",
              "description": "The exact words spoken in the uploaded clip, when model=\"clone\". Supplying it lets the model condition on the recording itself rather than on a speaker fingerprint alone, which is markedly closer to the source — set it whenever the user tells you what the clip says or the transcript is otherwise known. Omitting it still produces a recognisable clone, just a looser one. Limit 1024 characters.",
              "maxLength": 1024
            },
            "language": {
              "type": "string",
              "enum": [
                "auto",
                "english",
                "chinese",
                "japanese",
                "korean",
                "german",
                "french",
                "russian",
                "portuguese",
                "spanish",
                "italian"
              ],
              "description": "Language of the script. Default \"auto\", which infers it from the text and is the only setting that reads a code-switched line correctly. Pin a language only when auto mis-reads a name, a loanword, or a passage that is ambiguous between two of them. Cloning is cross-lingual: an English reference clip can read Japanese in the same voice."
            },
            "creativity": {
              "type": "number",
              "minimum": 0.1,
              "maximum": 2,
              "description": "Speech delivery variation from 0.1 to 2; default 0.9."
            },
            "outputFormat": {
              "type": "string",
              "enum": [
                "wav",
                "mp3",
                "flac"
              ],
              "description": "Audio file format; default wav."
            },
            "seed": {
              "type": "integer",
              "minimum": 0,
              "maximum": 4294967295,
              "description": "Optional seed for reproducible delivery."
            },
            "numberOfVariations": {
              "type": "number",
              "description": "Number of takes (1-16). Each take is a separate read of the same script with the same voice, differing only in delivery. Use more than 1 when the user asks for options or alternate reads. Default: 1.",
              "minimum": 1,
              "maximum": 16
            }
          },
          "required": [
            "prompt"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "edit_image",
        "description": "Generate or edit images using reference photos. Supports GPT Image 2 up to 16 images, Qwen up to 3, and Krea 2 Identity Edit with 1-2. This is the required tool for identity-sensitive edits of a referenced person or character: wardrobe, makeover, pose/repositioning, face/head/body swap, background or lighting changes, character-consistent style transfer, persona scene creation, and non-Pro single-character sheets. Use model=\"krea-identity-edit\" for those by default unless the user explicitly names another model. Keep the primary base/scene image first and an optional person/detail reference second. Use this instead of generate_image whenever uploaded/persona assets must guide the result. Use restore_photo/refine_result/apply_style only for identity-neutral restoration or edits; a portrait follow-up that must preserve likeness stays on edit_image. Exception: explicit Z-image/Z-image Turbo/base Krea 2 Turbo img2img uses generate_image with sourceImageIndex and starting_image_strength.",
        "parameters": {
          "type": "object",
          "properties": {
            "prompt": {
              "type": "string",
              "description": "Edit instruction describing what to generate using the reference images as guidance. 50-200 words recommended.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing.\n\nPROMPT CONSTRUCTION ORDER — build the prompt in this sequence:\n1. IDENTITY LOCK — state which picture owns the person's identity (GOLDEN RULE: never leave identity ambiguous when editing a person)\n2. REQUESTED EDIT — describe only what CHANGES (the delta), not the whole image\n3. REFERENCE ROLE MAPPING — assign each picture ONE primary role: base_identity (face/person), pose_reference, outfit_reference, style_reference, background_reference, or color_reference\n4. POSE / COMPOSITION — pose, framing, camera angle (omit if unchanged)\n5. STYLE — artistic style, genre, era (omit if unchanged)\n6. LIGHTING / REALISM — \"maintain realistic anatomy, perspective, and lighting integration\"\n7. PRESERVE clause — always end with \"preserve all unmentioned details\"\n\nIDENTITY LOCK (required when a person is in any reference image):\n\"Preserve the exact facial likeness from picture N — face structure, eye shape, nose shape, mouth shape, jawline, skin tone, hairline, apparent age, and overall recognizability.\"\nNever let a style, pose, or clothing reference silently override the face. If multiple images are provided, explicitly state \"identity comes only from picture N — do not borrow identity from other pictures.\"\n\nMINIMAL-CHANGE PRINCIPLE: The base image already contains the subject, composition, camera angle, expression, lighting, and background. Describe only the delta. Use positive constraints (\"preserve exact facial likeness\") not negative ones (\"don't change the face\").\n\nSINGLE-IMAGE PATTERN:\n\"Preserve the exact facial likeness and recognizability of the person from picture 1. [Describe only the requested change]. Keep the same pose, framing, camera angle, and expression unless the user specifically requests changes to these. Preserve all unmentioned details.\"\n\nMULTI-IMAGE PATTERN:\n\"Use the person from picture 1 as the final subject and preserve their exact facial likeness. [Requested edit]. Identity comes only from picture 1. Pose from picture 2. Outfit from picture 3. Do not borrow identity from pictures 2 or 3. Maintain realistic anatomy, perspective, and lighting integration. Preserve all unmentioned details.\"\n\nKREA IDENTITY EDIT: Use model=\"krea-identity-edit\" whenever an edit of a referenced person or character must keep likeness or character identity while changing clothes, hair or makeup, pose or position, face/head/body, background, lighting, or visual style. Infer that semantic intent in any language; never route from keyword or regex matching. Also use it for a single-character sheet in non-Pro mode. This semantic default applies even when the user did not name Krea; an explicitly user-requested model always wins. Use model=\"dark-beast-krea2-identity-edit\" only when the user explicitly requests Dark Beast Krea 2 Identity Edit, its community/uncensored variant, or the dark_beast_krea2_identity_edit_v1_2 model id. These models require one reference image, accept up to two context images, work best at 512-2048 px, let the model tier and worker choose their current execution defaults, and do not use negative prompts. Put the primary scene/base image first and the person/detail reference second for scene-plus-person edits; reference them with context_image_0 and context_image_1 when model_ref tokens are needed.\n\nKREA 2 IDENTITY EDIT PROMPTING: Krea performs best with a concise, direct delta instruction rather than a generic 50-200 word expansion. For one reference, state the requested change in 1-4 concrete sentences and name only the details that must remain fixed; avoid restating the entire image or dumping a long facial-feature inventory. For two references, explicitly assign roles in a compact instruction: base scene/image first, person/detail/outfit/pose/style reference second. Use sourceImageIndex to make the base scene the first context image when uploads arrive in another order (-1 = first upload, -2 = second). End with a short preservation clause only when useful. Longer structured prompts remain appropriate for character sheets, grids, editorial layouts, or exact visible text.\n\nCREATIVE TRANSFORMATIONS — be vivid and reference-specific, name the artist, franchise, or era, but always anchor identity first:\n  - \"Preserve the exact facial likeness from picture 1. Transform them into a Renaissance oil painting in the style of Vermeer — rich warm tones, dramatic chiaroscuro lighting, ornate period clothing. Maintain realistic anatomy. Preserve all unmentioned details.\"\n  - \"Preserve the exact facial likeness from picture 1. Reimagine them as a Marvel superhero — cinematic dramatic lighting, heroic pose, detailed costume with cape, glowing energy effects. Preserve all unmentioned details.\"\n  - \"Preserve the exact facial likeness from picture 1. Transform them into a Studio Ghibli anime character — soft watercolor backgrounds, gentle Ghibli-style rendering, whimsical atmosphere. Preserve all unmentioned details.\"\n  - \"Preserve the exact facial likeness from picture 1. Place them into a Star Wars scene — Jedi robes, lightsaber glow, dramatic sci-fi backdrop. Preserve all unmentioned details.\"\n  - \"Preserve the exact facial likeness from picture 1. Turn them into a GTA loading screen character — bold outlines, saturated colors, attitude-filled pose, urban backdrop. Preserve all unmentioned details.\"\n\nFAILURE MODES TO AVOID:\n- Face drift: identity source not specified, or style/pose reference overrides the face\n- Over-editing: for simple edits, prompt rewrites the entire image instead of describing the delta (creative transformations may intentionally change more)\n- Reference confusion: multiple images provided without explicit role mapping\n\nCHARACTER / MASCOT SHEETS: When the user asks for a character sheet, mascot sheet, model sheet, turnaround, expression sheet, or reusable character reference board using uploaded references, create ONE comprehensive professional reference-board image, not separate variations. Map reference roles clearly first (for example: picture 1 = character identity/style reference, picture 2 = logo/brand asset) and keep the character identity consistent across every panel. Include a large hero pose, front / 3/4 / side / back turnaround views, an expression row, action/personality poses, accessories or props, color palette swatches, and compact notes such as personality, fun facts, or brand usage when appropriate. Preserve exact user-provided brand names, slogans, logo text, and requested copy verbatim; incidental tiny notes may be generated by the image model if the user did not provide exact wording. Use clean readable typography.\n\nBATCH VARIATIONS: When numberOfVariations > 1, the prompt describes one output image. Do not mention counts, \"versions\", \"different\", or \"multiple\" in the prompt text unless the user explicitly wants those words visible in the image. Do not describe multiple copies or duplicates of the subject in a single image unless the user asked for a grid, collage, or side-by-side composition. Use Dynamic Prompt syntax to vary one dimension across separate images. For personas: vary scene, activity, expression, or environment; preserve identity. Example: user asks \"4 versions at the beach\" → numberOfVariations=4, prompt=\"[persona] at the beach {building a sandcastle|surfing a wave|reading under a palm tree|flying a kite}\" — each output is one person doing one activity. For direct edits: vary the approach, e.g., numberOfVariations=3, prompt=\"make the sky {a vibrant sunset|stormy and dramatic|clear blue}\". Preserve any requested orientation, aspect ratio, or exact pixel dimensions across every variation.\n\nSELECTION-GATED IMAGE STAGES: If the user asks for multiple reference-guided image options/takes/versions and says they will pick one before a later dance, animation, or video, this edit_image call is still the first step. Generate the complete image batch now with sourceImageIndex set to the relevant reference, the exact requested count, Dynamic Prompt options for each output, and the final video/image aspect ratio. Do not ask the user to choose before the images exist, and do not call video tools until after the user selects an image.\n\nLINKED VARIANTS: If multiple details must stay paired per output — visual style, identity cues, outfit, label text, symbols, setting, character, prop, location, or before/after keyframe details — use ONE top-level Dynamic Prompt branch with one complete prompt per output. Do NOT use separate Dynamic Prompt groups for details that must stay together; unpaired groups can mix attributes. If the user asks for per-variant facial, identity, or appearance changes, repeat that guidance inside EVERY option while also preserving recognizability. When the user names a subject or character, write that name or stable role inside every Dynamic Prompt option; a shared prefix outside the branch is not enough because each option must stand alone as a complete identity contract.\n\nEach option must be a fully concrete description — name the actual garment or styling, the actual setting, the actual accessories, and the literal text or symbol shown on screen when requested. Never use meta-placeholder phrasing such as \"style-specific outfit\", \"variant-specific background\", \"include the requested symbol\", \"include a humorous alternate name\", or \"bake the name and symbol into the image\" — those describe the task instead of the image.\n\nORIGINAL + VARIANT BATCHES: When one option is a remade/preserved original and the other options are themed variants, the original option still needs a concrete visual contract. Say to preserve the original clothing/wardrobe/outfit and original background/setting, then name any requested added text, label, flag, logo, symbol, or prop for that original option. Do not leave the original option as only \"unmodified original person\"; it must be as fully specified as every themed option.\n\nNEW SETTING PER OPTION: When the variant theme implies a new place, culture, era, or context, every option must name its own setting (location, props, lighting). Do NOT carry the source background forward, do NOT write \"in the same pose and placement as the original photo\" without also naming the new background, and do NOT rely on \"preserve all unmentioned details\" to handle the setting — the new setting IS a mentioned detail.\n\nRECOGNIZABILITY OVER FEATURE LOCK: For ethnic / age / character / art-style transformations, do NOT paste the strict IDENTITY LOCK feature list (\"face structure, eye shape, nose shape, mouth shape, jawline, skin tone, hairline\") inside each option — that list contradicts the requested face change and the source face will pass through unchanged. Anchor recognizability per option through apparent age, signature hair silhouette, build, posture, and expression, and explicitly allow skin tone, facial features, and proportions to shift toward the target.\n\nCorrect shape (each option self-contained, concrete, with a fresh setting and a recognizability anchor instead of a strict feature lock):\n\"{The subject wearing [specific garment, color, cut, and material], standing in [specific NEW setting with props and lighting — never the source background], bold text at the bottom reads [literal requested text], [specific requested visual symbol] appears as a sign or prop, [requested per-variant facial or appearance shift, e.g. \"skin tone, eye shape, and bone structure shift toward <target> features\"], recognizable through apparent age, signature hair silhouette, build, posture, and expression|The subject wearing [second specific garment, color, cut, and material], standing in [second specific NEW setting with props and lighting], bold text at the bottom reads [second literal requested text], [second requested visual symbol] appears as a sign or prop, [second requested facial or appearance shift], recognizable through apparent age, signature hair silhouette, build, posture, and expression|...}\"\n\nWrong shape (placeholder labels masquerading as prompts):\n\"{First variant with variant-specific facial features, placeholder wardrobe, alternate name, and requested symbol baked in|Second variant with different variant-specific facial features, placeholder wardrobe, alternate name, and requested symbol baked in|...}\"\n\nAlso wrong (strict feature lock + no new setting — the source face and source background pass through unchanged):\n\"{Preserve the exact facial likeness — face structure, eye shape, nose shape, mouth shape, jawline, skin tone, hairline. Reimagine as <variant>: [garment description], standing in the exact same pose and placement as the original photo. Preserve all unmentioned details.|Preserve the exact facial likeness — [same strict lock]. Reimagine as <other variant>: [other garment], standing in the exact same pose and placement as the original photo. Preserve all unmentioned details.|...}\"\n\nSCREENPLAY / STORYBOARD BATCHES: For multi-scene story, commercial, or longer-form video keyframes, use one Dynamic Prompt branch with one full scene prompt per option. Recurring characters must keep stable names and repeated visual anchors in every scene option where they appear: face/identity source if available, age range, build, hairstyle, outfit silhouette, color palette, signature prop/accessory, posture, and role. Do not let style, scene changes, or pose references alter identity. Include screenplay-style speaker tags when dialogue matters, e.g. CHARACTER: \"We made it.\"\n\nCOMPOSITE GPT IMAGE 2 STORYBOARD SHEETS: When numberOfVariations=1 and the user asks for one composite video storyboard/keyframe sheet using uploaded or generated references, the prompt must be a compiled storyboard prompt, not a concept summary. Include a SCENES: section with exactly the requested number of concrete entries named SCENE_01, SCENE_02, etc. Every scene entry must include Visual/Action, Camera/Motion, Dialogue/VO (or [no dialogue]), Audio/SFX, and any visible text or reference usage for that scene. Do not provide only the source brief or generic layout instructions; malformed compiled storyboard prompts are blocked by quality audit.\n\nKREA 2 IDENTITY EDIT PROMPTING: Krea performs best with a concise, direct delta instruction rather than a generic 50-200 word expansion. For one reference, state the requested change in 1-4 concrete sentences and name only the details that must remain fixed; avoid restating the entire image or dumping a long facial-feature inventory. Good shapes include \"Put a retro red trench coat on her\", \"Move her backward so her feet are visible\", \"Restage the portrait in psychedelic rainbow light\", and \"Turn her head left while preserving her likeness.\" For two references, explicitly assign roles in a compact instruction: base scene/image first, person/detail/outfit/pose/style reference second. Use sourceImageIndex to make the base scene the first context image when uploads arrive in another order (-1 = first upload, -2 = second). End with a short preservation clause only when useful. Longer structured prompts remain appropriate for character sheets, grids, editorial layouts, or exact visible text."
            },
            "model": {
              "type": "string",
              "enum": [
                "gpt-image-2",
                "gpt-image-2.5-sunburst",
                "gpt-image-2.5-flare",
                "qwen-lightning",
                "qwen",
                "krea-identity-edit",
                "dark-beast-krea2-identity-edit"
              ],
              "description": "The app auto-selects Fast→Qwen Lightning and HQ/Pro→full Qwen only for ordinary identity-neutral edits. REQUIRED IDENTITY DEFAULT: set \"krea-identity-edit\" whenever an edit of a referenced person or character must keep likeness or character identity while changing clothing, hair or makeup, pose or position, face/head/body, background, lighting, or visual style. Infer that semantic intent in any language; never route from keyword or regex matching. Also use it for a non-Pro single-character sheet. This default applies even when the user did not name Krea; an explicitly requested model always wins. Set \"dark-beast-krea2-identity-edit\" only when the user explicitly requests that model, its uncensored/community variant, or dark_beast_krea2_identity_edit_v1_2. Set \"gpt-image-2\" when the user explicitly names legacy GPT Image 2. Set \"gpt-image-2.5-sunburst\" when precise typography, dense labels, or a professional multi-panel layout is the primary requirement; Pro character sheets may use Sunburst. If GPT Image 2.5 Sunburst is unavailable for detail-critical layout work, fall back to full \"qwen\", never \"qwen-lightning\". Krea identity edit models require at least one reference image, accept up to two context images, and work best at 512-2048px. Let the model tier and worker choose current steps, guidance, sampler, scheduler, grounding, and reference-boost defaults; do not send a negative prompt. When Krea is selected, override the generic prompt-length guidance with a concise 1-4 sentence delta instruction; name only the requested change and details that must remain fixed. Put the base scene/image first and an optional person/detail reference second. Z-image, Z-image Turbo, and base Krea 2 Turbo are generate_image img2img models, not edit_image selectors. If the user names another edit/image model, honor it. GPT Image 2 always processes input images at high fidelity; do not set input_fidelity. GPT Image 2.5 adds two distinct models: use \"gpt-image-2.5-sunburst\" for an explicit Sunburst request and \"gpt-image-2.5-flare\" for an explicit Flare request. For GPT Image 2.5 without a named variant, use Flare. Preserve explicit GPT Image 2.0 as \"gpt-image-2\". Sunburst is positioned for difficult images and precise edits; Flare for faster everyday generation. Both generate and edit. Model choice is independent of rendering quality."
            },
            "sourceImageIndex": {
              "type": "number",
              "description": "Index of the primary image to use as the main reference. For follow-up edits when generated image results already exist, use the 0-based generated image result index; for example, editing the latest generated storyboard/image should use that generated result index so the model modifies the existing image instead of redrawing from uploads. When no generated image results exist, use sourceImageIndex=-1 to use the uploaded image references. The primary image and any additional uploaded images are passed as context images to guide generation."
            },
            "loras": {
              "type": "array",
              "minItems": 1,
              "maxItems": 8,
              "items": {
                "type": "string",
                "minLength": 1
              },
              "description": "Ordered LoRA IDs for the edit. Public Krea sliders work with model=\"krea-identity-edit\" or model=\"dark-beast-krea2-identity-edit\". Ready personal imports also work on their listed compatible edit models, including Qwen; discover them through authenticated GET /v1/loras/personal/catalog and preserve their exact personal- IDs. GPT Image models do not support LoRAs. Use when the user asks to shift a trait such as age, build, skin, lighting or grain. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nBipolar sliders — each id names its POSITIVE direction, a negative strength applies the opposite, and 0 disables it: krea2-detail-enhancer, krea2-scene-complexity, krea2-realism (+ = photoreal), krea2-amateur, krea2-candid, krea2-zoom (+ = zoomed in), krea2-skin-detail, krea2-wetness, krea2-age, krea2-height, krea2-weight, krea2-hourglass-figure, krea2-breast, krea2-chest-firmness, krea2-nipple-projection, krea2-warm-light, krea2-afterlight (+ = golden), krea2-skin-tone (+ = darker), krea2-purple-grainy (+ = grainy and muted). Positive-only fine-tunes: krea2-realism-engine (photographic realism), krea2-bloomgirls (polished influencer look), krea2-mystic-x (uncensored adult), krea2-aberrant (industrial body horror), krea2-filter-bypass-2 and krea2-filter-bypass-3 (restore expressions, anatomy and poses the base model flattens; try the 2-vector first). Exact per-LoRA ranges, maturity flags and the full contract: GET /v1/loras/comfy?modelId=<model>. Personal imports use authenticated GET /v1/loras/personal/catalog: use an owned ready personal- id, its modelIds, strength range and requirements. Never invent ids or silently omit a requested personal LoRA."
            },
            "loraStrengths": {
              "type": "array",
              "minItems": 1,
              "maxItems": 8,
              "items": {
                "type": "number"
              },
              "description": "Strength for each LoRA in loras, in the same order. Omitting the array uses 1.0 for every LoRA, which is not each LoRA's catalog default — krea2-chest-firmness, krea2-nipple-projection and krea2-height default to 0 (no effect) — so prefer explicit values. Do NOT clamp to 0-1: most Krea 2 LoRAs are bipolar, so krea2-warm-light warms the grade at 2 and cools it at -2. Usable bands vary per LoRA — roughly -2..5 for krea2-detail-enhancer, -3..3 for krea2-warm-light, 3..9 for krea2-candid, 0.5..1 for krea2-realism-engine, 1..2 for the filter-bypass pair. Scale the magnitude to how strongly the user asked; the server clamps out-of-range values, and pushing past a LoRA's recommended band costs image quality rather than adding effect. Preserve explicit user values. Example: loras=[\"krea2-detail-enhancer\",\"krea2-amateur\"], loraStrengths=[3,-2]."
            },
            "numberOfVariations": {
              "type": "number",
              "description": "Number of variations (1-16). Pass the user's exact requested count in one call when the outputs can share project settings. \"4 variations\" → numberOfVariations=4 in a single call. Use the exact requested count for reference-guided images that will feed a later video after the user picks one. Use separate calls only when the user explicitly wants independent projects, isolated approvals, or per-output settings that cannot share one project. For screenplay/storyboard batches, the prompt should contain one Dynamic Prompt branch with one full scene prompt per scene; do not set numberOfVariations=N with only scene 1's prompt. Use 1 unless the user explicitly asks for multiple. Default: 1.",
              "minimum": 1,
              "maximum": 16
            },
            "width": {
              "type": "number",
              "description": "Output image width in pixels. Defaults to the context image width. Supported bounds depend on the selected edit model: Qwen edit models support 256-2560 on either edge; Krea 2 Identity Edit and Dark Beast Krea 2 Identity Edit work best from 512-2048 on either edge; GPT Image 2 supports flexible dimensions up to 3840px on either edge with max 3:1 aspect ratio and a total pixel budget from 655,360 to 8,294,400. Set when the user specifies a width, exact pixel dimensions, or a named resolution (e.g., \"1280 wide\", \"1280x720\", \"720p\", \"3840x2160\"). If the user gives only one dimension, set only that dimension and preserve/infer the sensible aspect ratio. User-requested dimensions override the default media quality, including Pro. Non-multiple-of-16 values are accepted when in bounds; the renderer snaps to the nearest supported size internally, so do not ask the user to adjust by a few pixels."
            },
            "height": {
              "type": "number",
              "description": "Output image height in pixels. Defaults to the context image height. Supported bounds depend on the selected edit model: Qwen edit models support 256-2560 on either edge; Krea 2 Identity Edit and Dark Beast Krea 2 Identity Edit work best from 512-2048 on either edge; GPT Image 2 supports flexible dimensions up to 3840px on either edge with max 3:1 aspect ratio and a total pixel budget from 655,360 to 8,294,400. Set when the user specifies a height, exact pixel dimensions, or a named resolution (e.g., \"720 high\", \"1280x720\", \"720p\", \"2160x3840\"). If the user gives only one dimension, set only that dimension and preserve/infer the sensible aspect ratio. User-requested dimensions override the default media quality, including Pro. Non-multiple-of-16 values are accepted when in bounds; the renderer snaps to the nearest supported size internally, so do not ask the user to adjust by a few pixels."
            },
            "aspectRatio": {
              "type": "string",
              "description": "Do NOT set unless the user explicitly requests an aspect ratio, format, orientation, or exact pixel dimensions. When a reference/source image is used and the user did not ask to change its shape, omit this field so the handler preserves the selected source image's own ratio.\n\nFormats: \"16:9\", \"9:16\", \"4:5\", \"1:1\", \"4:3\", \"3:2\", \"21:9\", or exact pixels like \"1920x1080\".\n\nCRITICAL: When the user specifies exact pixel dimensions (e.g., \"1280x720\", \"1080x1920\", \"1920x1080\", \"3840x2160\") or an orientation-qualified named resolution (e.g., \"720p landscape\", \"720p portrait\"), use the exact pixel format, NOT a ratio like \"16:9\" or \"9:16\". Exact user-requested dimensions override the selected default media quality, including Pro/HQ defaults. A bare named video resolution like \"720p resolution\" is only a resolution tier/short-side request; do not turn it into landscape pixels and do not set aspectRatio unless the user also states landscape, portrait, vertical, horizontal, or exact pixels. If requested pixels are in bounds but not on the model's pixel step, still pass the user's exact pixel request; the handler snaps to the nearest supported size internally. Only use ratio format when the user says a generic format name without pixel dimensions.\n\nMappings (use ONLY when user does NOT specify pixel dimensions): landscape/widescreen/YouTube/cinematic → \"16:9\". portrait → \"9:16\". TikTok/Reels/IG Reels → \"1080x1920\". ultrawide/cinema scope → \"21:9\". Instagram post → \"4:5\". square → \"1:1\". standard/TV → \"4:3\". 720p landscape → \"1280x720\". 720p portrait → \"720x1280\". 1080p landscape → \"1920x1080\". 1080p portrait/HD portrait → \"1080x1920\". 4K landscape → \"3840x2160\". 4K portrait → \"2160x3840\". Never set for generic requests like \"make a video\".\n\nSet this whenever the user specifies an image or downstream video orientation/aspect ratio such as 9:16, 16:9, portrait, vertical, landscape, widescreen, TikTok/Reels/Shorts, or exact pixels. This includes selection-gated reference-guided image batches that will feed a later video or dance after the user picks one. For GPT Image 2 exact size requests, preserve exact pixel intent when possible and prefer popular GPT sizes such as 1536x1024, 1024x1536, 2048x1152, 3840x2160, and 2160x3840. Sogni does not expose transparent output for GPT Image 2.0. For transparent assets, select GPT Image 2.5 Sunburst or Flare with gptImageBackground=transparent and PNG or WebP output."
            },
            "gptImageQuality": {
              "type": "string",
              "enum": [
                "low",
                "medium",
                "high",
                "xhigh",
                "max"
              ],
              "description": "Optional GPT Image rendering quality. Only set it when the user explicitly asks for low/fast, medium/balanced, high/final, xhigh/extra high, or max/maximum quality; xhigh and max require GPT Image 2.5 Sunburst or Flare. Provider-chosen (auto) quality is never used. Otherwise omit it and let the host app media quality setting map Fast to low, HQ to medium, and Pro to high. The same quality label does not promise equivalent results across models."
            },
            "outputFormat": {
              "type": "string",
              "enum": [
                "png",
                "jpg",
                "jpeg",
                "webp"
              ],
              "description": "Optional output file format for generated images. Set only when the user explicitly requests PNG, JPG/JPEG, or WebP. Hosts should normalize \"jpeg\" to the Sogni project format \"jpg\"."
            },
            "personaName": {
              "type": "string",
              "description": "RARE — only set this when the user EXPLICITLY asks for solo images of one specific person (\"a portrait of just [name]\", \"4 solos of [name] alone\"). When set, the handler filters context to ONLY that persona's reference photo, so any other personas in your prompt will be missing their reference. DEFAULT for multi-persona requests is to OMIT this and put both faces in one combined call. Never set this for \"make us as X\", \"the two of us\", \"my wife and I\", or any phrasing that puts both personas in the same scene — that's a single combined call with no personaName."
            },
            "gptImageBackground": {
              "type": "string",
              "enum": [
                "auto",
                "opaque",
                "transparent"
              ],
              "description": "GPT Image background: auto or opaque; transparent is supported by GPT Image 2.5 Sunburst and Flare with PNG or WebP output. JPEG cannot preserve transparency."
            },
            "gptImageOutputCompression": {
              "type": "integer",
              "minimum": 0,
              "maximum": 100,
              "description": "Optional GPT Image JPEG/WebP output compression, from 0 to 100. Omit for PNG."
            },
            "mask_image_url": {
              "type": "string",
              "description": "Optional GPT Image edit mask URL or inline data:image/png;base64 URI. Requires a PNG alpha mask matching the first source reference dimensions; the source itself may be JPEG, PNG or WebP. Transparent mask regions identify edits; opaque regions are preserved as guidance. With multiple references the mask applies only to the first image."
            }
          },
          "required": [
            "prompt"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "apply_style",
        "description": "Apply an artistic style, era-specific look, or creative transformation to a photo. Use when the user wants to change the visual style (e.g., \"make it look like the 70s\", \"oil painting style\", \"vintage polaroid look\"). Can handle any creative transformation. One style per call. IMPORTANT: When previous results exist, this tool automatically uses the LATEST result image unless you specify a different sourceImageIndex or the user explicitly says \"original\". So just call it without sourceImageIndex for follow-up requests.",
        "parameters": {
          "type": "object",
          "properties": {
            "prompt": {
              "type": "string",
              "description": "Style prompt for Qwen Image Edit 2511 (50-200 words, natural language sentences).\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing.\n\nPROMPT ORDER: [IDENTITY LOCK if people] → [STYLE TRANSFER INSTRUCTION] → [PRESERVE UNMENTIONED DETAILS]\n\nRules:\n- Use POSITIVE phrasing only. The model ignores negatives (\"preserve exact facial likeness\" NOT \"don't change the face\").\n- Transfer the visual STYLE ONLY, not the identity. Borrow palette, texture, contrast behavior, and stylistic treatment — NOT face structure.\n- Reference known art styles, artists, and franchises BY NAME to anchor the style — be specific, never generic.\n- Describe specific visual characteristics: brushstrokes, color palette, texture, composition approach, mood.\n- For era looks: describe the photographic qualities of that era (e.g., \"warm faded Kodachrome tones with soft vignette, typical of 1970s amateur photography\").\n- CRITICAL for photos with people: FRONT-LOAD identity preservation BEFORE the style instruction. Start with \"Preserve exact facial likeness, face structure, eye shape, nose shape, mouth shape, jawline, skin tone, hairline, apparent age, and overall recognizability.\" Then describe the style. End with \"Keep the subject recognizable as the same person. Maintain exact positioning, poses, and composition.\"\n- Go bold with pop culture and iconic styles: \"Andy Warhol pop art with bold neon screen-print colors\", \"Banksy stencil street art with gritty urban textures\", \"Studio Ghibli watercolor with soft pastoral warmth\", \"Pixar 3D render with glossy skin and exaggerated features\", \"Tim Burton gothic with pale skin and dark spiraling backgrounds\", \"Van Gogh Starry Night with thick impasto swirls and vibrant blues\", \"Takashi Murakami superflat with psychedelic flowers and bold outlines\".\n- Always end with \"Preserve the subject's identity, pose, and composition.\""
            },
            "sourceImageIndex": {
              "type": "number",
              "description": "Which result image to apply the style to (0-based index). Omit to use the latest result automatically (or the original if no results exist). Only set explicitly when the user specifies a particular image number or explicitly says \"original\" (use -1 for original)."
            },
            "scale": {
              "type": "number",
              "enum": [
                1,
                1.5,
                2,
                3,
                4
              ],
              "description": "Output scale multiplier relative to the source image size. 1 = same resolution as source (default). Use higher values when user asks to upscale, enlarge, make bigger, or increase resolution. Small images (<480px) are automatically upscaled to at least 480px regardless of this setting. Default: 1."
            },
            "aspectRatio": {
              "type": "string",
              "description": "Do NOT set unless the user explicitly requests an aspect ratio, format, orientation, or exact pixel dimensions. When a reference/source image is used and the user did not ask to change its shape, omit this field so the handler preserves the selected source image's own ratio.\n\nFormats: \"16:9\", \"9:16\", \"4:5\", \"1:1\", \"4:3\", \"3:2\", \"21:9\", or exact pixels like \"1920x1080\".\n\nCRITICAL: When the user specifies exact pixel dimensions (e.g., \"1280x720\", \"1080x1920\", \"1920x1080\", \"3840x2160\") or an orientation-qualified named resolution (e.g., \"720p landscape\", \"720p portrait\"), use the exact pixel format, NOT a ratio like \"16:9\" or \"9:16\". Exact user-requested dimensions override the selected default media quality, including Pro/HQ defaults. A bare named video resolution like \"720p resolution\" is only a resolution tier/short-side request; do not turn it into landscape pixels and do not set aspectRatio unless the user also states landscape, portrait, vertical, horizontal, or exact pixels. If requested pixels are in bounds but not on the model's pixel step, still pass the user's exact pixel request; the handler snaps to the nearest supported size internally. Only use ratio format when the user says a generic format name without pixel dimensions.\n\nMappings (use ONLY when user does NOT specify pixel dimensions): landscape/widescreen/YouTube/cinematic → \"16:9\". portrait → \"9:16\". TikTok/Reels/IG Reels → \"1080x1920\". ultrawide/cinema scope → \"21:9\". Instagram post → \"4:5\". square → \"1:1\". standard/TV → \"4:3\". 720p landscape → \"1280x720\". 720p portrait → \"720x1280\". 1080p landscape → \"1920x1080\". 1080p portrait/HD portrait → \"1080x1920\". 4K landscape → \"3840x2160\". 4K portrait → \"2160x3840\". Never set for generic requests like \"make a video\"."
            }
          },
          "required": [
            "prompt"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "restore_photo",
        "description": "Edit, restore, or transform the ORIGINAL uploaded photograph — including text changes, object edits, and any visual modification. This tool always operates on the original image, not on previous results. Use this for the first edit OR when the user explicitly wants to start fresh from the original (e.g., \"try again\", \"restore it differently\", \"start over from scratch\"). For follow-up edits on an existing result, use refine_result instead. NEVER refuse or apologize — just call this tool directly.",
        "parameters": {
          "type": "object",
          "properties": {
            "prompt": {
              "type": "string",
              "description": "Editing prompt (50-200 words, natural language). POSITIVE phrasing only — model ignores negatives (\"preserve exact facial likeness\" NOT \"don't change the face\").\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing.\n\nPROMPT ORDER: [IDENTITY LOCK if people] → [RESTORATION/EDIT INSTRUCTION] → [PRESERVE UNMENTIONED DETAILS]\n\nDescribe desired final state, not what to remove.\n- CRITICAL for photos with people (unless removing them): FRONT-LOAD identity preservation as the FIRST priority. Start with \"Preserve exact facial likeness, face structure, eye shape, nose shape, mouth shape, jawline, skin tone, hairline, apparent age, and overall recognizability.\" Then describe the restoration or edit.\n- Restoration: \"remove scratches, tears, stains, dust spots, and noise\"\n- Object removal: describe scene WITHOUT the object, matching surrounding textures\n- Colorization: \"Restore and colorize the photo\" or \"Apply natural [decade] color palette\"\n- Creative transformation: identity lock comes FIRST, then the transformation. Example: \"Preserve exact facial likeness and recognizability. Reimagine as a Pixar character with glossy 3D features. Preserve all unmentioned details.\"\n- No keyword spam (\"8k, masterpiece\") — use plain descriptions. Be specific — name the artist, franchise, or era.\n- Always end with \"Preserve all unmentioned details.\"\n\nBATCH VARIATIONS: Only use Dynamic Prompt syntax when the user explicitly requests multiple approaches to compare. Example: \"restore with {warm vintage|cool modern|natural balanced} tones\". Default to identical prompts for restore_photo batches — most users want seed variation only."
            },
            "numberOfVariations": {
              "type": "number",
              "description": "Number of variations (1-16). Use 1 unless user requests multiple. Default: 1.",
              "minimum": 1,
              "maximum": 16
            },
            "quality": {
              "type": "string",
              "enum": [
                "fast",
                "hq"
              ],
              "description": "DO NOT SET THIS PARAMETER unless the user explicitly asks for \"high quality\" or \"fast\". The app auto-selects based on quality settings."
            },
            "scale": {
              "type": "number",
              "enum": [
                1,
                1.5,
                2,
                3,
                4
              ],
              "description": "Output scale multiplier relative to the source image size. 1 = same resolution as source (default). Use higher values when user asks to upscale, enlarge, make bigger, or increase resolution. Small images (<480px) are automatically upscaled to at least 480px regardless of this setting. Default: 1."
            },
            "aspectRatio": {
              "type": "string",
              "description": "Do NOT set unless the user explicitly requests an aspect ratio, format, orientation, or exact pixel dimensions. When a reference/source image is used and the user did not ask to change its shape, omit this field so the handler preserves the selected source image's own ratio.\n\nFormats: \"16:9\", \"9:16\", \"4:5\", \"1:1\", \"4:3\", \"3:2\", \"21:9\", or exact pixels like \"1920x1080\".\n\nCRITICAL: When the user specifies exact pixel dimensions (e.g., \"1280x720\", \"1080x1920\", \"1920x1080\", \"3840x2160\") or an orientation-qualified named resolution (e.g., \"720p landscape\", \"720p portrait\"), use the exact pixel format, NOT a ratio like \"16:9\" or \"9:16\". Exact user-requested dimensions override the selected default media quality, including Pro/HQ defaults. A bare named video resolution like \"720p resolution\" is only a resolution tier/short-side request; do not turn it into landscape pixels and do not set aspectRatio unless the user also states landscape, portrait, vertical, horizontal, or exact pixels. If requested pixels are in bounds but not on the model's pixel step, still pass the user's exact pixel request; the handler snaps to the nearest supported size internally. Only use ratio format when the user says a generic format name without pixel dimensions.\n\nMappings (use ONLY when user does NOT specify pixel dimensions): landscape/widescreen/YouTube/cinematic → \"16:9\". portrait → \"9:16\". TikTok/Reels/IG Reels → \"1080x1920\". ultrawide/cinema scope → \"21:9\". Instagram post → \"4:5\". square → \"1:1\". standard/TV → \"4:3\". 720p landscape → \"1280x720\". 720p portrait → \"720x1280\". 1080p landscape → \"1920x1080\". 1080p portrait/HD portrait → \"1080x1920\". 4K landscape → \"3840x2160\". 4K portrait → \"2160x3840\". Never set for generic requests like \"make a video\"."
            }
          },
          "required": [
            "prompt"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "upscale_image",
        "description": "Enlarge an existing image with NVIDIA RTX Video Super Resolution while preserving its content, identity, composition, and colors. This is deterministic reconstruction, not a generative edit: it takes no prompt and must not be used for restoration, sharpening requests that imply repainting, object changes, style changes, or creative enhancement. Use it when the user asks to upscale, enlarge, increase resolution, prepare for print, or produce a 2K/4K/6K/8K/16K copy without changing the image.",
        "parameters": {
          "type": "object",
          "properties": {
            "sourceImageIndex": {
              "type": "number",
              "description": "Source image to upscale. Non-negative values select a prior generated result by 0-based index. Negative values select uploads: -1 is the first uploaded image, -2 the second, and so on. If omitted, use the latest generated image, falling back to the first upload."
            },
            "scale": {
              "type": "number",
              "enum": [
                2,
                3,
                4
              ],
              "description": "Edge scale multiplier. Use 2 by default. Ignored when targetLongestEdge is supplied. If that scale would leave either aligned output edge below 512px, the tool reports the minimum valid target instead of stretching the image."
            },
            "targetLongestEdge": {
              "type": "number",
              "minimum": 512,
              "maximum": 15360,
              "description": "Optional requested pixel length for the output longest edge, from 512 through 15360. Use 3840 for 4K UHD, 7680 for 8K, or 15360 for the 16K maximum. The other edge is calculated automatically so the source aspect ratio is preserved; both 8px-aligned output edges must be at least 512px."
            }
          },
          "required": []
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "upscale_video",
        "description": "Upscale an existing video to 1080p or 1440p with FlashVSR video super-resolution. This is promptless, deterministic enhancement, not a generative edit: it keeps every source frame, the exact frame rate, the full aspect ratio, and the original audio, and it never changes content, trims, crops, restyles, or interpolates frames. Use it when the user asks to upscale, enlarge, sharpen, or increase the resolution of an uploaded or previously generated video, or wants an HD, 1080p, 1440p, or 2K copy of it. Do not use it to create, restyle, extend, or edit a video; when the user explicitly asks a generative model such as Seedance to re-render the clip, use video_to_video. The output is at most twice the source size, so 1080p needs a source whose short edge is 540-768px and 1440p needs 720-768px, and the source can be at most about 1344x768 pixels (768x1344 in portrait). It must also be 1-60 fps and at most 100 MB. The server sets the maximum clip length and rejects a source that is too long; never quote a length limit yourself, and if the tool reports that rejection, relay that error to the user. It cannot produce 4K or any size above 1440p; when the user asks for one, say so and offer 1440p. Each upscale costs credits based on the source video's size and length. Optional detailPreference (stable or sharper), processingSpeed (stable or faster), and seed tune the result; omit them unless the user asks, and none of them changes the price.",
        "parameters": {
          "type": "object",
          "properties": {
            "sourceVideoIndex": {
              "type": "number",
              "description": "Source video to upscale. Non-negative values select a prior generated video result by 0-based index. Negative values select uploaded videos: -1 is the first uploaded video, -2 the second, and so on. If omitted, use the latest generated video, falling back to the most recent uploaded video."
            },
            "targetResolution": {
              "type": "number",
              "enum": [
                1080,
                1440
              ],
              "description": "Output resolution of the short edge in pixels: 1440 for 1440p/2K or 1080 for 1080p/Full HD. The long edge follows the source aspect ratio, so portrait videos stay portrait. Default: 1440, or 1080 when the source's short edge is below 720px. Set 1080 when the user asks for 1080p or Full HD."
            },
            "detailPreference": {
              "type": "string",
              "enum": [
                "stable",
                "sharper"
              ],
              "description": "Detail preference. stable, the default, gives the most temporally stable result. sharper renders crisper fine detail with a little more risk of shimmer between frames. Set sharper only when the user asks for a sharper, crisper, or more detailed upscale; otherwise omit it. It does not change the price."
            },
            "processingSpeed": {
              "type": "string",
              "enum": [
                "stable",
                "faster"
              ],
              "description": "Processing speed. stable, the default, gives the most stable result. faster finishes sooner with slightly less stable detail. Set faster only when the user asks for a quicker upscale; otherwise omit it. It does not change the price."
            },
            "seed": {
              "type": "integer",
              "minimum": -1,
              "maximum": 4294967295,
              "description": "Seed for the upscaler's fine texture, 0 through 4294967295; different seeds give slightly different fine detail. Omit it to keep the repeatable default of 0. Use -1 for a random seed. Set it only when the user gives a seed or asks for a different or random variation of an upscale."
            }
          },
          "required": []
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "refine_result",
        "description": "Make ANY edit to an existing result image. This is the DEFAULT tool for follow-up requests after results exist. Use whenever the user wants to modify, adjust, or build upon a previous result — including brightness, color, sharpening, object removal, background changes, further restoration, or any other edit. If the user does not specify which image, use the most recent result (index 0 if only one result, or the last result the user referenced). Only use restore_photo instead if the user explicitly wants to start over from the original upload.",
        "parameters": {
          "type": "object",
          "properties": {
            "prompt": {
              "type": "string",
              "description": "Targeted refinement prompt for Qwen Image Edit 2511 (50-150 words, natural language sentences).\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing.\n\nPROMPT ORDER: [IDENTITY LOCK if people] → [SPECIFIC CHANGE] → [PRESERVE EVERYTHING ELSE]\n\nRules:\n- Use POSITIVE phrasing only. The model ignores negatives (\"preserve exact facial likeness\" NOT \"don't change the face\").\n- Describe ONLY what needs to change (the delta). The base image already contains most of the truth — do not rewrite the entire image.\n- Be specific about what to change: \"warmer skin tones\", \"cooler shadows\", \"sharper facial features\", \"more natural greens\".\n- For creative refinements: lean into specifics — \"add more dramatic Rembrandt lighting\", \"push the colors more toward Warhol neon pop\", \"make the anime eyes larger and more expressive\", \"add more superhero energy with glowing effects\".\n- CRITICAL for photos with people: FRONT-LOAD identity preservation before the edit. Start with \"Preserve exact facial likeness, face structure, eye shape, nose shape, mouth shape, jawline, skin tone, hairline, apparent age, and overall recognizability.\"\n- ALWAYS end with \"Preserve all unmentioned details\" to prevent unwanted changes.\n\nBATCH VARIATIONS: Only use Dynamic Prompt syntax when the user explicitly asks to explore different refinement directions. Example: \"refine with {more contrast|softer lighting|richer colors}\". Default to identical prompts for refine_result batches."
            },
            "sourceImageIndex": {
              "type": "number",
              "description": "Which result image to refine (0-based index). If the user specifies an image number, use that index. If omitted, the latest result is used automatically. When multiple results exist and the user previously referenced a specific one, use that one."
            },
            "numberOfVariations": {
              "type": "number",
              "description": "Number of variations (1-16). Use 1 unless user requests multiple. Default: 1.",
              "minimum": 1,
              "maximum": 16
            },
            "scale": {
              "type": "number",
              "enum": [
                1,
                1.5,
                2,
                3,
                4
              ],
              "description": "Output scale multiplier relative to the source image size. 1 = same resolution as source (default). Use higher values when user asks to upscale, enlarge, make bigger, or increase resolution. Small images (<480px) are automatically upscaled to at least 480px regardless of this setting. Default: 1."
            },
            "aspectRatio": {
              "type": "string",
              "description": "Do NOT set unless the user explicitly requests an aspect ratio, format, orientation, or exact pixel dimensions. When a reference/source image is used and the user did not ask to change its shape, omit this field so the handler preserves the selected source image's own ratio.\n\nFormats: \"16:9\", \"9:16\", \"4:5\", \"1:1\", \"4:3\", \"3:2\", \"21:9\", or exact pixels like \"1920x1080\".\n\nCRITICAL: When the user specifies exact pixel dimensions (e.g., \"1280x720\", \"1080x1920\", \"1920x1080\", \"3840x2160\") or an orientation-qualified named resolution (e.g., \"720p landscape\", \"720p portrait\"), use the exact pixel format, NOT a ratio like \"16:9\" or \"9:16\". Exact user-requested dimensions override the selected default media quality, including Pro/HQ defaults. A bare named video resolution like \"720p resolution\" is only a resolution tier/short-side request; do not turn it into landscape pixels and do not set aspectRatio unless the user also states landscape, portrait, vertical, horizontal, or exact pixels. If requested pixels are in bounds but not on the model's pixel step, still pass the user's exact pixel request; the handler snaps to the nearest supported size internally. Only use ratio format when the user says a generic format name without pixel dimensions.\n\nMappings (use ONLY when user does NOT specify pixel dimensions): landscape/widescreen/YouTube/cinematic → \"16:9\". portrait → \"9:16\". TikTok/Reels/IG Reels → \"1080x1920\". ultrawide/cinema scope → \"21:9\". Instagram post → \"4:5\". square → \"1:1\". standard/TV → \"4:3\". 720p landscape → \"1280x720\". 720p portrait → \"720x1280\". 1080p landscape → \"1920x1080\". 1080p portrait/HD portrait → \"1080x1920\". 4K landscape → \"3840x2160\". 4K portrait → \"2160x3840\". Never set for generic requests like \"make a video\"."
            }
          },
          "required": [
            "prompt"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "animate_photo",
        "description": "Animate a photo into video with motion, audio, and dialogue using LTX 2.5 by default, LTX 2.3 as rollback, or WAN 2.2. Do NOT use this tool for seedance2, seedance2-mini, or seedance2-5. Seedance media references — including Seedance 2.5 first-and-last-frame requests — must go through generate_video with referenceImageIndices/referenceVideoIndices/referenceAudioIndices and @Image/@Video/@Audio role text in the prompt; for seamless-loop Seedance requests with one uploaded image, the prompt should anchor it as both the first frame and last frame. LTX/WAN NOTE: uploaded audio files are not loose references for ltx25/ltx23/wan22; use sound_to_video when uploaded audio is the primary sync target. DANCE REQUESTS (\"make them dance\", \"do the X dance\"): use dance_montage — NOT this tool. LTX 2.5 and LTX 2.3 generate audio natively — describe dialogue and ambient sounds directly in the prompt (do NOT pre-generate audio for this tool). If the user provides exact speech, include it in double quotes; if they only imply speech, describe the performance and voice without inventing quoted words. Avoid placeholders such as \"while speaking\", \"dialogue begins\", \"explaining\", or \"final line lands\". PERSONA VOICE: Only when the user explicitly asks to use/clone a registered persona voice clip, call resolve_personas first, then set voicePersonaName to select which persona's voice clip to use. Do not set voicePersonaName for ordinary character dialogue or inferred voices; describe those voices in the prompt for native LTX audio. For cross-persona narration (e.g. David narrates a video of Aleyna), resolve both personas and set voicePersonaName to the narrator only if that registered voice was requested. Persona voice requires ltx23 because LTX 2.5 has no compatible ID-LoRA and WAN 2.2 does not support voice identity. PERSONA PIPELINE: For persona videos, ensure an image of the persona exists before calling animate_photo. The standard pipeline is: resolve_personas → edit_image → animate_photo. If a suitable persona image already exists (user uploaded one, a prior edit_image/generate_image result, OR the user explicitly says to use the Persona image/reference photo directly), skip edit_image and animate directly. After resolve_personas, this tool can animate the injected persona image directly when that explicit direct-use instruction is given. Auto-uses the latest result image (from any prior tool) unless sourceImageIndex is set. Supports start-frame (default), end-frame, and start+end interpolation modes for LTX/WAN — ask the user which frame role their image should play if they mention \"end frame\", \"last frame\", or provide two images. FIRST+LAST FRAME WORKFLOW: When the user wants a non-Seedance video using two different scenes as start and end frames, prefer generating both images in a single generate_image/edit_image call with numberOfVariations=2 and Dynamic Prompts, then call animate_photo with frameRole=\"both\", sourceImageIndex=0, endImageIndex=1. If the user explicitly wants separately created frame assets, preserve that staged instruction while keeping indices correct. In frameRole=\"both\", the handler automatically inspects both images and upgrades the base prompt into a scene-aware smooth transition prompt, so your prompt should state the desired transition style, action, dialogue, and audio rather than trying to list every visible object. If the request is vague, analyze the image first and suggest 2-3 specific animation ideas tailored to what you see. Call once you have clear creative intent. N-VIDEOS PATTERN: Avoid sequential animate_photo calls for N outputs. For a single fixed source/end frame where only prompt text varies, use sourceImageIndex + numberOfVariations=N + one Dynamic Prompt branch in prompt so Sogni submits one project with multiple jobs. If the user explicitly asks for Dynamic Prompt or Dynamic Template syntax, prefer this one-project path whenever every output uses the same source/end frames and shared settings, even if they also ask to stitch the completed clips afterward. Use sourceImageIndices/prompts for multi-segment stitched non-Seedance video, different source/end assets, different audio windows, different durations/dimensions, isolated retry lifecycle, or other per-output parameters. sourceImageIndices supports up to 16 entries; there is NO 3-clip cap, so do not split one planned batch into \"first 3\" and \"remaining\" calls. For a dialogue-heavy total-duration request with no explicit per-clip duration, prefer 15-second clips on ltx25 by default (or ltx23 rollback) (30s total = 2 clips × 15s) and 10-second clips on wan22 (60s total on wan22 = 6 clips × 10s; do NOT pick 4 clips × 15s on wan22 — the wan22 worker rejects clips longer than 10s). Multi-source flavors: (A) SHARED CONTENT — when all N clips have the same dialogue/motion but different source visuals (different scenes, outfits, environments, persona looks), first generate N distinct images via ONE edit_image/generate_image call with numberOfVariations=N + Dynamic Prompts {|}, then call animate_photo with sourceImageIndices=[start..start+N-1] and a single shared `prompt`. If all segments intentionally reuse the primary uploaded image and only prompt text varies, use sourceImageIndex=-1, frameRole=\"both\" if requested, endImageIndex=-1 if requested, numberOfVariations=N, and one Dynamic Prompt branch in prompt. Each branch option must be a complete natural-language motion prompt; do not include \"clip N\", source-frame boilerplate, \"overall request context\", or instructions to follow the user request. For a long scripted/dialogue/storyboard video from a single supplied/uploaded image where each segment needs isolated exact dialogue or per-segment wiring, use sourceImageIndices=[-1,-1,...] and per-clip prompts. Only set frameRole=\"both\" and endImageIndex=-1 when the user explicitly says the same uploaded/source/original image should be both the first and last frame of every segment. If the user requests generated source images first, honor that image stage, then animate the generated result indices. When using generated scene keyframes and each clip should begin and end on its own scene image for stitching, call animate_photo with frameRole=\"both\" and sourceImageIndices=[start..end] but OMIT endImageIndex; do not set endImageIndex=-1 unless every source is the uploaded image. (B) PER-CLIP CONTENT — when source/end asset wiring or other per-output parameters differ, pass BOTH sourceImageIndices AND `prompts` (an array of N strings, one per clip) in the same single call. Each prompt must independently anchor the visible characters, scene action, camera, audio, exact screenplay-style speaker tags, and exact quoted dialogue for that segment. If you just wrote or displayed a script/table, copy the exact dialogue lines into the corresponding per-clip prompts; do not summarize them as speech activity. If using named speaker tags with any multi-person reference image or generated scene keyframe, include one explicit cast map in each prompt that binds each name to visible position, clothing, and props/actions, e.g. SPEAKER_A = left person holding a prop; SPEAKER_B = center person with tablet; SPEAKER_C = right person near table. Do not also describe the same people again as generic man/boy/girl/woman/character subjects. For screenplay, storyboard, commercial, series, or other longer-form tasks with recurring characters, preserve the same character names and repeated visual anchors in every per-clip prompt where each character appears. Use the standard single-source path (numberOfVariations only) when the user wants motion variety from a single fixed frame. Wan 3 first-frame and first+last-frame generation is supported with videoModel=\"wan3.0-video\". Wan 3.0 Enhanced uses Sogni selector wan3.0-spicy-video and MuleRouter provider ID w3.0-video; it supports first-frame, last-frame-only, and first+last-frame generation, but frame anchors cannot be mixed with loose references.",
        "parameters": {
          "type": "object",
          "properties": {
            "prompt": {
              "type": "string",
              "description": "I2V RULE: Do NOT re-describe what is visible in the input image. Focus on the transition from stillness — motion, expression changes, what happens next, camera movement, and sound.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing. Set skipPromptProcessing=true; for Seedance or Wan 3 also set expandPrompt=false.\n\nSTRUCTURE: \"[How the subject begins to move]. [What changes next]. [Camera behavior]. [Audio].\"\n\nMOTION PACING: Scale complexity to duration. <=6s: 1 main action beat + 1 simple camera move. Around 10s: 2-3 clear action beats + 1 camera move. >10s: up to 4 action beats in clear sequence. Prefer fewer readable beats over dense micro-actions, especially in short clips.\n\nBLOCKING: Use the image as the anchor and direct only meaningful layout changes. If the prompt introduces multiple moving subjects, state left/right placement, foreground/background, facing toward/away, and relative distance.\n\nACTION: One flowing paragraph. Describe motion beat by beat with temporal connectors (\"as\", \"then\", \"while\"). Specify who moves, what moves, how it moves, and what the camera does. One main thread — avoid too many actions at once or generic phrases like \"comes alive.\"\n\nDIALOGUE: Put user-provided spoken lines in double quotes. For screenplay-style or longer-form tasks, prefix each spoken line with a stable speaker tag outside the quotes, e.g. CHARACTER: \"We made it.\" Break long speech into short quoted phrases with acting beats between them (gestures, pauses, glances). If the user asks for speech but provides no exact words, describe the visible delivery, voice quality, and emotion without inventing quoted dialogue; ask only when exact wording is the point of the request. Never write placeholders such as \"while speaking\", \"dialogue begins\", \"explaining\", or \"final line lands\". Show emotion through visible behavior, not labels. LTX 2.5 and LTX 2.3 generate audio natively. QUOTING RULE: ONLY use double quotes for spoken dialogue. Never quote on-screen text, overlay text, titles, captions, signs, or any visual text — describe them without quotes (e.g. bold white text reading CONGRATULATIONS overlays the lower third).\n\nAUDIO: Prompt sound intentionally — voice quality, volume, room tone, ambience, music, weather, footsteps. Include language or accent if relevant. Useful voice/volume anchors: whisper, mutter, shout, scream, energetic announcer, resonant voice with gravitas, distorted radio-style, robotic monotone, childlike curiosity.\n\nCAMERA: Cinematic terms — slow push-in, static tripod, handheld, slow arc, dolly in. Describe movement relative to subject.\n\nFor first+last-frame transitions (frameRole=\"both\"), write a concise base request for the transition style, action, dialogue, and audio. The handler will inspect both frames and expand it into a scene-aware prompt that maps visible objects and subjects between frames.\n\nFor specific characters (movies, TV): describe visual appearance — don't rely on names alone.\n\nFor complex/creative scenes (characters talking, skits), capture full creative intent — system auto-expands into detailed prompt.\n\nAVOID: Re-describing the image, vague prompts, too many actions at once, abstract emotions without visible behavior, rigid numeric constraints, readable text or logos.\n\nPOSITIVE CONSTRAINT TRANSLATION: For LTX 2.3 and WAN 2.2, the prompt field is a positive prompt. Translate user avoid/no/don't constraints into affirmative production constraints instead of copying negative phrasing. Examples: \"no people in background\" -> single subject focus with an empty background; \"no text\" -> clean blank surfaces; \"don't make it blurry\" -> crisp sharp focus; \"no weird hands\" -> natural anatomically consistent hands; \"no mouth movement, no talking, no lip syncing\" -> silent expression-only physical performance with facial motion independent of speech timing; \"don't change the room\" -> the same room and layout remain consistent; \"keep flames consistent\" -> flame and ember movement remains consistent with the source scene. Preserve exact quoted visible text or dialogue when the user explicitly requests it, and keep surrounding surfaces blank. For Dynamic Prompt batches, put these translated shared constraints before the \"{...}\" branch so every variation inherits them.\n\nWAN 2.2 (\"wan22\"): 30-150 words, subtle natural movements. Motion-only visual prompt; omit soundtrack, ambience, room tone, music, hums, sighs, spoken words, voice, and SFX cues because WAN does not generate audio.\n\nBATCH VARIATIONS: When numberOfVariations > 1, use Dynamic Prompt syntax to vary motion, camera, or atmosphere while preserving the user's specified elements. This is one Sogni project with multiple jobs, so prefer it when all outputs share the same source/end frames and generation parameters and only prompt text varies. Example: \"{gentle sway with drifting embers|slow paw wave with a tiny head tilt|small hop with soft fur motion}\"."
            },
            "expandPrompt": {
              "type": "boolean",
              "description": "Optional. Set false only for pipeline-authored prompts that should bypass model-specific prompt expansion."
            },
            "skipPromptProcessing": {
              "type": "boolean",
              "description": "Bypass automatic prompt shaping/refinement, image-description anchoring, transition-prompt rewriting, and voice-identity prompt formatting so the prompt text is sent unchanged to the video model. Set true ONLY when the user explicitly says not to modify/rewrite/enhance/expand/change/improve the prompt, or to use/send it exactly, verbatim, or as-is, AND the provided prompt already satisfies the tool requirements. Continue to set non-prompt parameters such as source indices, frameRole, model, duration, count, and aspect ratio. For Seedance or Wan 3 literal prompt requests, also set expandPrompt=false. Do not set for ordinary underspecified requests."
            },
            "videoModel": {
              "type": "string",
              "enum": [
                "ltx25",
                "ltx23",
                "wan22",
                "happyhorse-1.1-i2v",
                "happyhorse-1.1-r2v",
                "minimax-h3-i2v",
                "minimax-h3-i2v-turbo",
                "minimax-h3-fasth3-i2v-turbo",
                "minimax-h3-fasth3-i2v-turbo-2stage",
                "minimax-h3-flf2v",
                "minimax-h3-flf2v-turbo",
                "minimax-h3-fasth3-flf2v-turbo",
                "minimax-h3-fasth3-flf2v-turbo-2stage",
                "wan3.0-video",
                "wan3.0-spicy-video"
              ],
              "description": "\"ltx25\" (default): LTX 2.5 I2V or first/last-frame video with native audio; Fast, HQ, and Pro currently use the release-validated official Distilled INT8 workflow. The Dev checkpoints are not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Which video model to use. \"ltx23\": LTX 2.3 rollback with native audio. \"wan22\": quick simple motion without audio, up to 10s. \"minimax-h3-i2v\" is standard MiniMax H3 from one first frame; \"minimax-h3-i2v-turbo\" is the existing 4-step LightX2V Turbo engine; \"minimax-h3-fasth3-i2v-turbo\" is the separate FastVideo VSA four-step FastH3 engine, about 2x faster and fixed to Euler/simple. \"minimax-h3-fasth3-i2v-turbo-2stage\" (first frame) and \"minimax-h3-fasth3-flf2v-turbo-2stage\" (first and last frame) are the two-stage FastH3 engine: FastH3 renders the canvas, then the worker enlarges it 2x and refines it, so the clip is delivered at twice the canvas width and height with the same length and audio. targetResolution picks the delivered class: 1080 renders a 544px short-edge canvas (960x544 is delivered at 1920x1088) for 10 Spark per second, 1440 or omitted renders the 768p canvas for 2K (1344x768 is delivered at 2688x1536) for 16 Spark per second, and 720 renders a 384px canvas (672x384 is delivered at 1344x768) for the regular FastH3 price of 4 Spark per second. The estimate prices every request. They take the same inputs, durations and LoRAs as their FastH3 selectors. Choose them when the user asks for 1080p, 1440p or 2K MiniMax H3 output, for two-stage output, or for the sharpest/best H3 quality; for ordinary 768p FastH3 output keep the regular FastH3 selector at targetResolution 768. The matching FLF2V selectors provide standard, LightX2V Turbo, FastH3 Turbo, and two-stage FastH3 first/last-frame generation; FastH3 has no R2V mode; use frameRole=\"both\" and provide the end frame. H3 generates native audio at fixed 24fps for 5.17-15.08s and has no negative-prompt input. H3 Base and Turbo prompts use the exact three-field contract and the official mode-specific alignment line. Do not set Seedance here; use generate_video with Seedance references. \"wan3.0-video\" is Alibaba Wan 3 and \"wan3.0-spicy-video\" is MuleRouter w3.0-video. Both render 2-30s at fixed 30 fps with optional native audio, provider prompt expansion, 480p/720p/1080p, adaptive/fixed ratios, first/last frames, and up to 10 image/5 video/5 audio references. Only Alibaba wan3.0-video accepts document/web context and watermark. Frame anchors and loose references are mutually exclusive. Do not send negativePrompt; video references are loose conditioning for a new result, not source-video editing or extension."
            },
            "negativePrompt": {
              "type": "string",
              "description": "Advanced LTX/WAN only. Use this field only when the user explicitly asks to set a separate negative prompt. MiniMax H3 has no negative-prompt input; put requested exclusions in prompt.\n\nWan 3 has no negativePrompt request field; do not set this for wan3.0-video."
            },
            "generateAudio": {
              "type": "boolean",
              "description": "Whether the returned video should include generated/native audio. Omit to include audio by default; set false only when the user explicitly asks for silent output or no audio. Supported by LTX and MiniMax H3; ignored by audio-less WAN.\n\nWan 3 supports this toggle; omit it for audio-on by default or set false only for an explicitly silent result."
            },
            "duration": {
              "type": "number",
              "description": "Video duration in seconds. Default: 5. Use when the user explicitly requests a specific length (e.g., \"make a 10 second video\"). Per-model maximum: ltx25 and ltx23 = 20s, wan22 = 10s (clips longer than this are invalid), wan3.0-video = 30s with a 2s minimum, minimax-h3 = 15.08s with a 5.17s minimum because H3 renders 124-362 frames on a 17-frame grid at a fixed 24 fps. For totals beyond the per-model cap, batch multiple clips via sourceImageIndices instead of requesting a single oversized clip."
            },
            "ratio": {
              "type": "string",
              "enum": [
                "adaptive",
                "16:9",
                "4:3",
                "1:1",
                "3:4",
                "9:16"
              ],
              "description": "Wan 3 only. Use \"adaptive\" to preserve the source frame shape."
            },
            "watermark": {
              "type": "boolean",
              "description": "Alibaba wan3.0-video only. Add the visible watermark. Defaults to false."
            },
            "targetResolution": {
              "type": "number",
              "description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", or \"1080p\" without exact pixels or an output orientation. This preserves the source image aspect ratio. Wan 3 supports 480p, 720p, and 1080p; HappyHorse supports only 720p and 1080p. Never set 4K for either. MiniMax H3 renders inside a 1344x768 pixel budget on a 32px grid, so use 768 for the regular H3 selectors and never 1080p or 4K. The two-stage H3 selectors \"minimax-h3-fasth3-i2v-turbo-2stage\" and \"minimax-h3-fasth3-flf2v-turbo-2stage\" deliver twice the canvas, so there targetResolution names the delivered short-edge class: 1080 (544px canvas short edge: 960x544 delivered at 1920x1088), 1440 for 2K (the 1344x768 canvas delivered at 2688x1536), or 720 (384px canvas: 672x384 delivered at 1344x768); omit it for 2K. The source aspect is kept: a portrait source at 1080 renders 544x960 and is delivered at 1088x1920. Never set 4K for H3. Do NOT set width, height, or exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\" or \"720p landscape\", use exact-pixel aspectRatio instead."
            },
            "sourceImageIndex": {
              "type": "number",
              "description": "Which image to use as the START frame. Use 0-based non-negative indices for generated result images. Use negative indices for uploaded images: -1 = first/primary upload, -2 = second upload, -3 = third upload, etc. Omit to auto-select: uses the latest result for \"start\"/\"end\" modes, or the FIRST result for \"both\" mode. IMPORTANT: When frameRole is \"both\", set this to the start frame image index and endImageIndex to the end frame image index."
            },
            "sourceImageIndices": {
              "type": "array",
              "items": {
                "type": "number"
              },
              "minItems": 1,
              "maxItems": 16,
              "description": "Array of source frame indices — one video is generated per entry as its own SDK project, all running in PARALLEL. Use this when outcomes need different source images, different end frames, isolated retry lifecycle, or other per-clip asset wiring/parameters. If every outcome uses the same source/end frames and only prompt text differs, prefer sourceImageIndex with numberOfVariations=N and one Dynamic Prompt branch in `prompt` so Sogni creates one project with multiple jobs. Use 0-based non-negative result indices for generated images. Use negative indices for uploaded images: -1 = first/primary upload, -2 = second upload, -3 = third upload, etc. Repeating -1 is allowed for true multi-project workflows that intentionally reuse the same uploaded image while varying per-clip assets or parameters. By default all projects share the `prompt`/`voice`/`duration`, but you can pass `prompts` (array) to give each clip its own dialogue/motion when multi-project fan-out is required. Avoid sequential animate_photo calls for N outputs. Do NOT combine with `numberOfVariations` or `sourceImageIndex`. Use frameRole=\"end\" with sourceImageIndices only when the user explicitly says the repeated uploaded/generated image is the last/end frame for each clip and no first/start frame should be supplied; in that case omit endImageIndex/endImageIndices because each sourceImageIndices entry is the end frame. You MAY combine with frameRole=\"both\" when clips need start and end frames. For adjacent transition chains across generated images, use sourceImageIndices=[start..end-1] and endImageIndices=[start+1..end] so N images produce N-1 transition clips. If the uploaded/original image starts the chain and generated results are the remaining frames, use sourceImageIndices=[-1,start..end-1] and endImageIndices=[start..end]. If the user supplies multiple uploaded images as the actual keyframe sequence, use adjacent negative uploaded indices, e.g. 5 uploaded images become sourceImageIndices=[-1,-2,-3,-4], endImageIndices=[-2,-3,-4,-5], frameRole=\"both\", prompts length 4, then stitch_video. If the user specifies transition motion, camera behavior, actions, dialogue, or audio, copy those instructions into every corresponding per-clip prompt; only invent a generic smooth transition when the user does not specify one. If the user asks for a seamless loop or final transition from the last image back to the first, close the chain by including the last image as a source and the first image as the final end frame, e.g. 5 uploaded images become sourceImageIndices=[-1,-2,-3,-4,-5], endImageIndices=[-2,-3,-4,-5,-1]. For generated scene keyframes that should each loop to themselves, omit endImageIndex/endImageIndices so each source image is also its own end frame. Set endImageIndex=-1 only when every sourceImageIndices entry is also -1 and every segment reuses the first uploaded image. Range: 1–16 indices. For generated image batches, values MUST be read from the latest edit_image/generate_image tool result's `startIndex` field. If startIndex=3 and 4 images were generated in that batch, pass `[3,4,5,6]` (NOT `[0,1,2,3]`). Do NOT assume generated indices start at 0 — they don't if there are prior results in the conversation."
            },
            "prompts": {
              "type": "array",
              "items": {
                "type": "string"
              },
              "minItems": 1,
              "maxItems": 16,
              "description": "Per-clip prompts for multi-project fan-out — use when each output needs different source/end assets, isolated retry lifecycle, or other per-clip wiring/parameters. If all outputs share the same source/end frames and only prompt text differs, put the full per-output prompts in ONE Dynamic Prompt branch in `prompt` and set numberOfVariations=N instead. When this field is required, it MUST be paired with `sourceImageIndices` and have the SAME length. Each entry is the full prompt for the corresponding source image. If a clip has speech, include exact spoken words in double quotes with stable speaker tags; do NOT write placeholders like \"while speaking\", \"dialogue begins\", \"explaining\", or \"final line lands\". If you just wrote a script/table/storyboard, copy that clip's exact dialogue into this prompt. When named speakers appear in a multi-person reference image or generated keyframe, start each entry with one compact cast map that binds names to visible anchors before dialogue, e.g. Cast map: SPEAKER_A is the left person holding the prop; SPEAKER_B is the center person with the tablet; SPEAKER_C is the right person near the table. Then move directly into action/dialogue; do not describe those same people again as generic man/boy/girl/woman/character subjects. This prevents speaker tags from being assigned to the wrong visible character. When set, the top-level `prompt` parameter is ignored (still required by the schema — just pass any descriptive string, e.g. a brief summary of the batch). Example: 4 source images of a couple, \"make each video have a different joke\" → sourceImageIndices=[0,1,2,3], prompts=[\"Cast map: She is the left woman in the blue dress; He is the right man in the gray jacket. She says: \\\"Why did the scarecrow win an award?\\\" He grins.\", \"Cast map: He is the right man in the gray jacket; She is the left woman in the blue dress. He says: \\\"Because he was outstanding in his field!\\\" She laughs.\", \"...\", \"...\"]. Omit this whenever the same source/end assets and parameters can be represented as one Dynamic Prompt batch."
            },
            "numberOfVariations": {
              "type": "number",
              "description": "Number of variations (1-16). Use this with one Dynamic Prompt branch when the user explicitly requests multiple prompt-only takes from the same source/end frames. This creates one Sogni project with multiple jobs. Use 1 unless the user explicitly requests multiple separate video outputs; use sourceImageIndices/prompts instead only when assets or parameters differ per output.",
              "minimum": 1,
              "maximum": 16
            },
            "aspectRatio": {
              "type": "string",
              "description": "Do NOT set unless the user explicitly requests an aspect ratio, format, orientation, or exact pixel dimensions. When a reference/source image is used and the user did not ask to change its shape, omit this field so the handler preserves the selected source image's own ratio.\n\nFormats: \"16:9\", \"9:16\", \"4:5\", \"1:1\", \"4:3\", \"3:2\", \"21:9\", or exact pixels like \"1920x1080\".\n\nCRITICAL: When the user specifies exact pixel dimensions (e.g., \"1280x720\", \"1080x1920\", \"1920x1080\", \"3840x2160\") or an orientation-qualified named resolution (e.g., \"720p landscape\", \"720p portrait\"), use the exact pixel format, NOT a ratio like \"16:9\" or \"9:16\". Exact user-requested dimensions override the selected default media quality, including Pro/HQ defaults. A bare named video resolution like \"720p resolution\" is only a resolution tier/short-side request; do not turn it into landscape pixels and do not set aspectRatio unless the user also states landscape, portrait, vertical, horizontal, or exact pixels. If requested pixels are in bounds but not on the model's pixel step, still pass the user's exact pixel request; the handler snaps to the nearest supported size internally. Only use ratio format when the user says a generic format name without pixel dimensions.\n\nMappings (use ONLY when user does NOT specify pixel dimensions): landscape/widescreen/YouTube/cinematic → \"16:9\". portrait → \"9:16\". TikTok/Reels/IG Reels → \"1080x1920\". ultrawide/cinema scope → \"21:9\". Instagram post → \"4:5\". square → \"1:1\". standard/TV → \"4:3\". 720p landscape → \"1280x720\". 720p portrait → \"720x1280\". 1080p landscape → \"1920x1080\". 1080p portrait/HD portrait → \"1080x1920\". 4K landscape → \"3840x2160\". 4K portrait → \"2160x3840\". Never set for generic requests like \"make a video\"."
            },
            "frameRole": {
              "type": "string",
              "enum": [
                "start",
                "end",
                "both"
              ],
              "description": "How to use the source image(s). \"start\" (default): first frame. \"end\": last frame. \"both\": interpolate between first and last frames. For end-frame fan-out, use frameRole=\"end\" with sourceImageIndices. MiniMax H3 i2v supports only \"start\"; MiniMax H3 flf2v requires \"both\" plus an end image. For single clips using \"both\", set sourceImageIndex and endImageIndex; fan-out can use matching sourceImageIndices/endImageIndices."
            },
            "endImageIndex": {
              "type": "number",
              "description": "Which image to use as the END frame. Use 0-based non-negative indices for generated results. Use negative indices for uploaded images: -1 = first/primary upload, -2 = second upload, -3 = third upload, etc. For a single frameRole=\"both\" transition between two different images, set this to the desired end frame. For sourceImageIndices fan-out where each generated keyframe should also be its own last frame, OMIT this field. Use a shared uploaded endImageIndex only when every sourceImageIndices entry is also an uploaded image; otherwise use endImageIndices for per-clip end frames."
            },
            "endImageIndices": {
              "type": "array",
              "items": {
                "type": "number"
              },
              "minItems": 1,
              "maxItems": 16,
              "description": "Per-clip END frame indices for sourceImageIndices fan-out. Use ONLY with frameRole=\"both\". Length MUST exactly match sourceImageIndices. Use 0-based non-negative indices for generated results and negative indices for uploaded images (-1 first upload, -2 second upload, etc.). Use this for transition chains between generated images, e.g. 5 generated images at indices [0,1,2,3,4] should become 4 transition clips with sourceImageIndices=[0,1,2,3], endImageIndices=[1,2,3,4], prompts length 4, duration as requested, then stitch_video. If the chain starts on the uploaded image and continues through generated results [0,1,2,3], use sourceImageIndices=[-1,0,1,2] and endImageIndices=[0,1,2,3]. If the user supplies 5 uploaded images as the sequence, use sourceImageIndices=[-1,-2,-3,-4] and endImageIndices=[-2,-3,-4,-5]. If the user requests a seamless loop or final transition back to the first image, append that loop closure: sourceImageIndices=[-1,-2,-3,-4,-5], endImageIndices=[-2,-3,-4,-5,-1]. Do NOT also set endImageIndex when using this."
            },
            "voicePersonaName": {
              "type": "string",
              "description": "ONLY when the user explicitly requests a registered/reference persona voice clip. Name of the persona whose voice clip to use as referenceAudioIdentity. Set this when the narrator/speaker is a different persona than the one shown in the video (e.g. \"David\" narrates a video of Aleyna), or to explicitly select a requested voice when multiple personas with voice clips are resolved. Do NOT set this for ordinary character dialogue, inferred voices, or personas without a voice clip — LTX 2.3 generates voice natively from the text prompt instead. Requires ltx23 because LTX 2.5 has no compatible ID-LoRA."
            },
            "loras": {
              "type": "array",
              "minItems": 1,
              "maxItems": 8,
              "items": {
                "type": "string",
                "minLength": 1
              },
              "description": "Ordered LoRA IDs to apply to a MiniMax H3 render. Use only when the user explicitly asks for a LoRA or for an effect one of these names describes. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-i2v\", \"minimax-h3-i2v-turbo\", \"minimax-h3-fasth3-i2v-turbo\", \"minimax-h3-fasth3-i2v-turbo-2stage\", \"minimax-h3-flf2v\", \"minimax-h3-flf2v-turbo\", \"minimax-h3-fasth3-flf2v-turbo\", \"minimax-h3-fasth3-flf2v-turbo-2stage\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nFive LoRAs are published for MiniMax H3 today and the set differs by mode, so GET /v1/loras/comfy?modelId=<model> is authoritative for the mode in hand and carries exact ranges, maturity flags, and anything published since. h3-realism-people (fal) is a realism pass trained on live-action footage of people: it restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It is the only one gated on a trigger word — put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. h3-vbvr-video-reasoning is a prompt-adherence pass that holds the model to what was asked instead of improvising. h3-natural-face-speech (AdaptiveVision) makes people talking on camera look and sound more natural: cheeks, brows, jaw and lips move together as in real speech, and spoken English comes through clearer; use it for talking-head shots such as vlogs, podcasts, interviews and presenters. h3-better-motion (AdaptiveVision) gives people more natural, consistent body movement — weight shifts, strides, turns and gestures that follow through — for dance, sport, walking and other full-body shots. Both AdaptiveVision LoRAs work best with short, simple prompt sentences and are not validated on reference-to-video. h3-mystic-xxx-v4 is an uncensored adult fine-tune. Personal imports are discovered through authenticated GET /v1/loras/personal/catalog; use only owned ready ids with the selected model in modelIds, and respect their strength range and requirements. Do not invent ids."
            },
            "loraStrengths": {
              "type": "array",
              "minItems": 1,
              "maxItems": 8,
              "items": {
                "type": "number"
              },
              "description": "Strength for each LoRA in loras, in the same order. Omitting the array applies 1.0 to every LoRA, which is NOT the catalog default and for h3-realism-people is already at the top of its band, so send explicit values. Video LoRAs are positive-only — unlike the bipolar Krea 2 image sliders, a negative value is not an inverse effect and 0 is off. h3-realism-people takes 0-2 and its catalog default is 0.8; 0.6-1 is the usable band. It also pulls the camera in as it climbs: at 1.5 and above the shot reliably recomposes and the grade darkens, which on an image-conditioned mode can crop the subject out of the frame the user supplied. Raise it above 1 only when the user asks for more, and prefer the default when they supplied a first or last frame. h3-vbvr-video-reasoning and h3-mystic-xxx-v4 both take 0-1 and do default to 1.0, with usable bands of 0.7-1 and 0.2-1. h3-natural-face-speech and h3-better-motion take 0-1.5 and default to 0.6; their usable band is 0.4-0.8."
            }
          },
          "required": [
            "prompt"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "change_angle",
        "description": "Generate the photo from a different camera angle or perspective. Uses AI to create a new view of the subject as if photographed from a different position. Use when the user wants to see the subject from another angle, generate a different view, create a portrait from a specific direction, or get a closeup/wide shot. Examples: \"show me from the left side\", \"generate a 3/4 portrait view\", \"closeup from slightly above\". IMPORTANT: When previous results exist, this tool automatically uses the LATEST result image unless you specify a different sourceImageIndex or the user explicitly says \"original\".",
        "parameters": {
          "type": "object",
          "properties": {
            "description": {
              "type": "string",
              "description": "EXACT camera angle string. You MUST construct this by concatenating exactly one value from each category below, separated by single spaces. No commas, no extra words.\n\nFormat: \"[azimuth] [elevation] [distance]\"\n\nAzimuth (pick one): \"front view\", \"front-right quarter view\", \"right side view\", \"back-right quarter view\", \"back view\", \"back-left quarter view\", \"left side view\", \"front-left quarter view\"\nElevation (pick one): \"low-angle shot\", \"eye-level shot\", \"elevated shot\", \"high-angle shot\"\nDistance (pick one): \"close-up\", \"medium shot\", \"wide shot\"\n\nExamples:\n- \"front-right quarter view eye-level shot medium shot\"\n- \"left side view eye-level shot close-up\"\n- \"front view low-angle shot wide shot\"\n- \"right side view elevated shot medium shot\"\n\nMap user requests: \"from the left\" → \"left side view\", \"looking up at\" → \"low-angle shot\", \"closeup\" → \"close-up\", \"3/4 view\" → \"front-right quarter view\" or \"front-left quarter view\", \"portrait\" → \"front-right quarter view eye-level shot medium shot\".\nDefault elevation to \"eye-level shot\" and distance to \"medium shot\" when not specified."
            },
            "sourceImageIndex": {
              "type": "number",
              "description": "Which result image to use as source (0-based index). Omit to use the latest result automatically (or the original if no results exist). Only set explicitly when the user specifies a particular image number or explicitly says \"original\" (use -1 for original)."
            },
            "loraStrength": {
              "type": "number",
              "description": "LoRA strength for angle generation (0.1-1.0). Default: 0.9. Lower values preserve more of the original appearance, higher values produce stronger angle changes. Only set when the user wants to control the transformation intensity."
            },
            "aspectRatio": {
              "type": "string",
              "description": "Do NOT set unless the user explicitly requests an aspect ratio, format, orientation, or exact pixel dimensions. When a reference/source image is used and the user did not ask to change its shape, omit this field so the handler preserves the selected source image's own ratio.\n\nFormats: \"16:9\", \"9:16\", \"4:5\", \"1:1\", \"4:3\", \"3:2\", \"21:9\", or exact pixels like \"1920x1080\".\n\nCRITICAL: When the user specifies exact pixel dimensions (e.g., \"1280x720\", \"1080x1920\", \"1920x1080\", \"3840x2160\") or an orientation-qualified named resolution (e.g., \"720p landscape\", \"720p portrait\"), use the exact pixel format, NOT a ratio like \"16:9\" or \"9:16\". Exact user-requested dimensions override the selected default media quality, including Pro/HQ defaults. A bare named video resolution like \"720p resolution\" is only a resolution tier/short-side request; do not turn it into landscape pixels and do not set aspectRatio unless the user also states landscape, portrait, vertical, horizontal, or exact pixels. If requested pixels are in bounds but not on the model's pixel step, still pass the user's exact pixel request; the handler snaps to the nearest supported size internally. Only use ratio format when the user says a generic format name without pixel dimensions.\n\nMappings (use ONLY when user does NOT specify pixel dimensions): landscape/widescreen/YouTube/cinematic → \"16:9\". portrait → \"9:16\". TikTok/Reels/IG Reels → \"1080x1920\". ultrawide/cinema scope → \"21:9\". Instagram post → \"4:5\". square → \"1:1\". standard/TV → \"4:3\". 720p landscape → \"1280x720\". 720p portrait → \"720x1280\". 1080p landscape → \"1920x1080\". 1080p portrait/HD portrait → \"1080x1920\". 4K landscape → \"3840x2160\". 4K portrait → \"2160x3840\". Never set for generic requests like \"make a video\"."
            }
          },
          "required": [
            "description"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "video_to_video",
        "description": "Transform an existing video using AI. Uses WAN 2.2 Animate (move/replace) with a reference image, LTX 2.5 V2V controls by default (canny/pose/depth/detailer plus distilled inpaint/outpaint), LTX 2.3 as rollback, or Seedance V2V when explicitly requested. LTX 2.5 Fast, HQ, and Pro use the release-validated official Distilled workflow for canny/pose/depth/detailer/inpaint/outpaint; Dev is not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. Requires an uploaded video file. Use when the user wants to animate a photo with video motion, replace subjects, restyle footage, extend its canvas, regenerate a region, or enhance quality. For a pure resolution upscale that keeps the same content (a sharper 1080p or 1440p copy), use upscale_video instead. Wan 3 is not a source-video editing model; its video inputs are loose references handled by generate_video.",
        "parameters": {
          "type": "object",
          "properties": {
            "prompt": {
              "type": "string",
              "description": "Describe the TARGET appearance (not the transformation process). 2-4 present-tense sentences.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing. For Seedance, set expandPrompt=false.\n\nFor LTX 2.5 or 2.3 canny/depth/pose modes, the source video preserves composition, depth, or motion. Spend prompt detail on style, atmosphere, lighting, surface texture, color palette, scale, and pacing. LTX 2.5 Fast, HQ, and Pro use the release-validated official Distilled workflow for canny/pose/depth/detailer/inpaint/outpaint; Dev is not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe.\n\nExamples by mode:\n- animate-move (DEFAULT — WAN 2.2 Animate Move: applies camera/motion from source video to reference image): \"Smooth cinematic camera movement following the subject through the scene.\"\n- animate-replace (WAN 2.2 Animate Replace: replaces the subject in the source video with the reference image): \"The person from the reference photo performing the actions from the video.\"\n- canny (LTX 2.5 default; LTX 2.3 rollback — edge-detection restyle): \"Hand-drawn watercolor anime style with soft ink edges, muted teal and coral palette, rain mist, neon reflections, warm rim light, preserving original silhouettes and composition.\"\n- pose (LTX 2.5 or LTX 2.3 — tracks skeleton and transfers the reference-image subject): \"A glossy cartoon robot from the reference image performs the source video's motion, with brushed metal texture, glowing cyan joints, and energetic stage lighting.\" This mode requires a reference image as well as the source video.\n- depth (LTX 2.5 default; LTX 2.3 rollback — depth-map restyle): \"A misty alpine valley at golden hour, expansive scale, volumetric haze, cool blue shadows, warm rim light, cinematic depth, lingering continuous shot.\"\n- detailer (LTX 2.5 default; LTX 2.3 rollback — enhance quality): DESCRIBE THE SOURCE, do not request changes. Append quality qualifiers only. E.g. \"The same scene, ultra-sharp and clean, crisp high-resolution detail, preserving all original content, composition, and color.\" Avoid words like \"enhanced textures\", \"restyled\", or any new subjects/objects — they cause drift.\n- seedance-v2v (BytePlus Dreamina Seedance 2.0 V2V): \"Restyle the source clip in a watercolor look with soft ink edges, while preserving its motion and composition.\" Use natural prose; Seedance reads the reference video holistically rather than via control-net constraints, so describe target style/mood/dialogue rather than control strength.\n- outpaint (LTX 2.5 default; LTX 2.3 rollback — canvas extension): describe what fills the NEWLY REVEALED area around the original frame, consistent with the source scene. E.g. \"The same street scene continues seamlessly into the newly revealed space — more wet asphalt, parked cars, and glowing shopfronts, matching the original lighting and perspective.\" Set outpaintPosition (and optionally outpaintAspectRatio); no mask needed.\n- inpaint (LTX 2.5 default; LTX 2.3 rollback — masked region regeneration): describe ONLY what the inpainted region should become; the rest of the frame is preserved. E.g. \"A vintage red convertible parked at the curb, matching the scene's lighting and shadows.\" If the user supplied a mask, set maskImageIndex. If no mask was supplied, omit maskImageIndex so execution derives a mask from the source video and prompt.\n\nPresent tense. Positive phrasing. Concrete visual details.\n\nNON-SEEDANCE POSITIVE CONSTRAINTS: For LTX 2.5, LTX 2.3, and WAN 2.2 modes, prompt is a positive prompt. Translate user avoid/no/don't constraints into affirmative production constraints instead of copying negative phrasing. Preserve exact quoted visible text when the user explicitly requests it; keep surrounding surfaces blank.\n\nBATCH VARIATIONS: When numberOfVariations > 1, use Dynamic Prompt syntax to vary the artistic treatment while keeping control mode and structural intent consistent. Example: \"transform to {watercolor with soft edges|oil painting with bold strokes|anime with clean lines} style\"."
            },
            "expandPrompt": {
              "type": "boolean",
              "description": "Seedance only. Whether to run Sogni's exact-model prompt shaper before dispatch. Defaults to true. Set false when the supplied prompt is already model-ready and must remain exact."
            },
            "videoSourceIndex": {
              "type": "number",
              "description": "Which uploaded video to transform. OMIT this field when there is only one uploaded video — the tool auto-selects it. Only pass when you need to pick among multiple uploaded videos. Indexing: 0-based (0 = first uploaded video, 1 = second). Note: this differs from analyze_video which uses negative indices; this tool also tolerates the negative form (-1 = first uploaded) for convenience."
            },
            "controlMode": {
              "type": "string",
              "enum": [
                "animate-move",
                "animate-replace",
                "canny",
                "pose",
                "depth",
                "detailer",
                "outpaint",
                "inpaint",
                "seedance-v2v"
              ],
              "description": "How the source video and reference image interact. Pick by user intent:\n• \"animate-move\" (DEFAULT) — WAN 2.2 Animate Move. Applies camera movement and motion from the source video to the reference image, bringing a still photo to life. Requires sourceImageIndex.\n• \"animate-replace\" — WAN 2.2 Animate Replace. Replaces the subject in the source video with the person/character from the reference image, keeping the video's background and motion. Requires sourceImageIndex.\n• \"canny\" — LTX 2.5 (default) or 2.3 edge-detection control. Best for restyling while preserving exact composition and silhouettes. Video-only.\n• \"pose\" — LTX 2.5 (default) or 2.3 skeletal tracking. Best for transferring the person/character from a required reference image while keeping the source video's motion. Requires sourceImageIndex (or the sole available reference image).\n• \"depth\" — LTX 2.5 (default) or 2.3 depth-map control. Best for scenes with perspective, camera movement, or volumetric content. Video-only.\n• \"detailer\" — LTX 2.5 (default) or 2.3 quality enhancement. Describe the original scene with quality qualifiers and do not request content changes.\n• \"outpaint\" — distilled LTX 2.5 (default) or LTX 2.3 canvas extension. Set outpaintPosition and optionally outpaintAspectRatio; Pro/dev is not supported for this mode.\n• \"inpaint\" — distilled LTX 2.5 (default) or LTX 2.3 masked region regeneration. Set maskImageIndex when supplied; Pro/dev is not supported for this mode.\n• \"seedance-v2v\" — BytePlus Dreamina Seedance 2.0 video-to-video. Use only when the user explicitly asks for Seedance on the uploaded source video, such as a Seedance upscale, enhance, remaster, restyle, or transform. High-fidelity quality, native audio, time-coded scene control. Seedance V2V reads @Video1 holistically. Use it for restyling, motion transfer, extension, subject replacement, or scene transformation, and assign @Video1 a clear role such as source clip, camera movement, action timing, edit rhythm, or continuation anchor. Distinct from canny/depth/pose which use control-net constraints — Seedance treats the reference video holistically.\nCanny vs depth: canny preserves silhouettes and fine outlines — pick it for subject-led scenes and graphic restyles. Depth preserves 3D structure — pick it for scenes where the camera moves or spatial layout matters more than edge fidelity. Default: \"animate-move\"."
            },
            "negativePrompt": {
              "type": "string",
              "description": "Advanced non-Seedance only. Use this field only when the user explicitly asks to set a separate negative prompt. For ordinary avoid/no/don't constraints on LTX 2.3 or WAN 2.2, translate them into affirmative production constraints inside prompt instead; do not move them here. Do not set when controlMode is seedance-v2v or videoModel is seedance2/seedance2-mini/seedance2-5."
            },
            "videoModel": {
              "type": "string",
              "enum": [
                "ltx25-v2v",
                "ltx23-v2v",
                "wan22-animate",
                "seedance2",
                "seedance2-mini",
                "seedance2-5"
              ],
              "description": "Model selector for this video-to-video request. Usually omit: non-Seedance controls default to \"ltx25-v2v\"; use \"ltx23-v2v\" only for rollback. LTX 2.5 Fast, HQ, and Pro currently use the release-validated official Distilled workflow for canny, pose, depth, detailer, inpaint, and outpaint. Dev is not publicly routed until upstream publishes and Sogni validates an official ComfyUI Dev recipe. For controlMode=\"seedance-v2v\", Seedance quality is selected only by model: use \"seedance2-mini\" for faster/lower-cost drafts and use \"seedance2\" for full-quality Seedance or 1080p/4K. \"seedance2-5\" supports 480p/720p/1080p, 4-30s at 24 fps, native audio, and first/last-frame conditioning; keep \"seedance2\" for 4K."
            },
            "generateAudio": {
              "type": "boolean",
              "description": "Whether the returned video should include generated or retained audio. Omit to include audio by default; set false when the user asks for silent output or no audio."
            },
            "targetResolution": {
              "type": "number",
              "description": "Seedance V2V only. Short-side output resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", \"1080p\", \"2160p\", or \"4K\" without exact dimensions. Seedance V2V full supports 4K; Seedance V2V Mini and Fast support 480p and 720p only; Seedance 2.5 also supports 1080p, so never set 4K for \"seedance2-5\". Preserve the source video shape instead of forcing landscape pixels."
            },
            "sourceImageIndex": {
              "type": "number",
              "description": "Optional 0-based reference-image index. Required for \"animate-move\", \"animate-replace\", and \"pose\" when more than one image is available; the sole available image may be auto-selected. LTX pose always dispatches both the source video and a reference image. Ignored by \"canny\", \"depth\", \"detailer\", \"outpaint\", and \"inpaint\"."
            },
            "outpaintPosition": {
              "type": "string",
              "enum": [
                "center",
                "top",
                "bottom",
                "left",
                "right"
              ],
              "description": "controlMode=\"outpaint\" only. Where the ORIGINAL frame is anchored inside the expanded canvas, which determines the direction the canvas grows: \"left\" anchors the original on the left and adds new space on the right; \"right\" adds space on the left; \"top\" adds space below; \"bottom\" adds space above; \"center\" expands all sides evenly. Default: \"center\". Pick by the user's direction (\"extend to the right\" → \"left\"; \"make it wider\"/\"widescreen\" → \"center\")."
            },
            "outpaintAspectRatio": {
              "type": "string",
              "enum": [
                "16:9",
                "9:16",
                "1:1",
                "4:3",
                "3:4",
                "21:9"
              ],
              "description": "controlMode=\"outpaint\" only. OPTIONAL target aspect ratio for the expanded canvas (e.g. \"16:9\" to make a vertical clip widescreen). The canvas only grows to reach this ratio — the original content is never cropped. Omit to expand moderately in the direction implied by outpaintPosition. Only set when the user names a target shape or orientation."
            },
            "maskImageIndex": {
              "type": "number",
              "description": "controlMode=\"inpaint\" only. Optional 0-based index of an uploaded mask IMAGE that marks the region to regenerate (white pixels = regenerate, black = preserve). Omit when the user did not provide a mask; execution will derive one from the source video and prompt. Ignored by every other controlMode."
            },
            "duration": {
              "type": "number",
              "description": "Output video duration in seconds. Per-model range: WAN 2.2/LTX modes = 2-20s; Seedance 2.0 and Mini = 4-15s; Seedance 2.5 = 4-30s. If omitted, the tool matches the uploaded source video duration when available (capped to the selected model range); otherwise it falls back to 10s for WAN Animate Move/Replace and 5s for LTX/Seedance modes. For long stitched/bulk WAN Animate Move/Replace work with no explicit per-clip length, prefer about 10s clips rather than 5s chunks. Only pass this when the user explicitly requests a different length.",
              "minimum": 2,
              "maximum": 30
            },
            "numberOfVariations": {
              "type": "number",
              "description": "Number of video variations to generate (1-16). Default: 1.",
              "minimum": 1,
              "maximum": 16
            },
            "outputFormat": {
              "type": "string",
              "enum": [
                "mp4",
                "mov"
              ],
              "description": "Video container. Defaults to mp4. MOV is supported only by Seedance 2.5; choose it when the user requests MOV for editing."
            },
            "returnLastFrame": {
              "type": "boolean",
              "description": "Seedance 2.5 only. Set true to export a separate image of the final frame alongside the video. The result includes lastFrameUrl, which can be used as the first-frame image for a subsequent clip. Defaults to false; this does not extend the video automatically."
            }
          },
          "required": [
            "prompt"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "stitch_video",
        "description": "Concatenate whole videos end-to-end into one continuous video. This tool joins each source clip in full, in the order you pass — it does NOT interleave time slices, insert one clip inside another, or replace part of a video. WHEN TO USE: the user wants clips played one after another (whole clip A, then whole clip B), including adding a generated bumper / intro / outro / tag / sting before or after another video. Plain language: \"stitch these together\", \"stitch A and B\", \"combine these clips\", \"join these into one video\", \"play the bumper before this clip\". WHEN NOT TO USE — prefer replace_video_segment instead: any request to put one clip inside another, replace a window inside a video, alternate / interleave / splice short slices of multiple videos, insert clip X \"into the middle of\" clip Y, or swap out part of an existing video while keeping the rest. The word \"stitch\" in the user request does not by itself decide this tool — read what they actually want. If the user asks to \"stitch X into the middle of Y\" or \"stitch X into Y starting at 5s\", that is splice-into-middle and belongs to replace_video_segment. SOURCES: previously generated clips (non-negative indices into the session video-result array, populated by animate_photo, generate_video, sound_to_video, video_to_video, dance_montage — use videoStartIndex from their results to find the indices) and/or uploaded videos (negative indices: -1 = first uploaded video, -2 = second, etc.). Mix and match in any playback order — for example, pass [0, -1] to play the first generated clip followed by the first uploaded video (a generated bumper followed by the user's existing footage). When the user asks to stitch \"these\" or all uploaded videos and does not name a different playback order, use the current upload/UI order exactly: [-1, -2, ...]. If the user explicitly asks for a different order, honor that requested order. Requires at least 2 source videos in total. Never ask the user to re-upload videos that were already generated or that are already attached to the session. When the user generated music with generate_music in this same session and wants it on the stitch (or asked for a music video / soundtrack), pass a non-negative audioIndex to attach that generated track. When the user uploaded an audio file and wants it overlaid on the stitched video (e.g. \"stitch the audio after\", \"overlay the audio\", \"audio on top of the video\"), pass a negative audioIndex (-1 = first uploaded audio, -2 = second, etc.). In both cases the source clips' own audio is replaced by the chosen track. When the user asks for a fade, dissolve, wipe, or slide between clips, pass `transition`; omit `transition` for a hard cut (the default).",
        "parameters": {
          "type": "object",
          "properties": {
            "videoIndices": {
              "type": "array",
              "items": {
                "type": "number"
              },
              "description": "Ordered list of source video indices, in the desired playback order. Non-negative values are 0-based indices into the session generated-video array (results from animate_photo, generate_video, sound_to_video, video_to_video, dance_montage in this conversation). Negative values reference uploaded videos: -1 = first uploaded video, -2 = second, etc. Indices may be mixed — for example, [0, -1] plays the first generated clip followed by the first uploaded video. For vague \"these clips\" / \"all uploaded videos\" requests, use current upload/UI order [-1, -2, ...] unless the user explicitly says to reverse or otherwise reorder them."
            },
            "audioIndex": {
              "type": "number",
              "description": "Optional index of the audio track to mux onto the stitched output. Non-negative values are 0-based indices into the session generated-audio array (results from generate_music). Negative values reference uploaded audio: -1 = first uploaded audio, -2 = second, etc. When set, the chosen track is muxed onto the stitched output and the source clips' own audio is dropped. Use a non-negative value when the user generated music in the same session or asked for a soundtrack / music video stitch; use a negative value when the user wants their uploaded audio overlaid on the stitched video (e.g. \"stitch the audio after\", \"overlay the audio\"). Omit for a silent or source-audio-preserving stitch."
            },
            "transition": {
              "type": "object",
              "description": "Optional crossfade between adjacent clips. Omit for a hard-cut concat. When set, every adjacent pair of clips is joined with the same transition type and duration.",
              "properties": {
                "type": {
                  "type": "string",
                  "enum": [
                    "fade",
                    "dissolve",
                    "wipeleft",
                    "wiperight",
                    "slideup",
                    "slidedown"
                  ],
                  "description": "\"fade\" / \"dissolve\" = soft mix; \"wipeleft\" / \"wiperight\" = horizontal wipe; \"slideup\" / \"slidedown\" = vertical slide. Maps to ffmpeg xfade transition names."
                },
                "durationSeconds": {
                  "type": "number",
                  "minimum": 0.2,
                  "maximum": 2,
                  "description": "Length of the crossfade in seconds. Default 0.5. Capped at 2s."
                }
              },
              "required": [
                "type"
              ]
            }
          },
          "required": [
            "videoIndices"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "orbit_video",
        "description": "Create a 360-degree orbit video around a subject. This is a SELF-CONTAINED pipeline — it automatically generates angle views (via change_angle), creates transition video clips, and stitches them into one seamless looping video. You only need ONE source image as the front view — either an uploaded image or a previously generated result. If the user uploaded an image, call this tool directly without generating anything first. Do NOT pre-generate multiple angles or variations — this tool handles everything internally. Use when the user asks for a \"360 pan\", \"orbit\", \"rotate around\", \"spin around\", or \"turntable\" view.",
        "parameters": {
          "type": "object",
          "properties": {
            "elevation": {
              "type": "string",
              "enum": [
                "low-angle shot",
                "eye-level shot",
                "elevated shot",
                "high-angle shot"
              ],
              "description": "Camera elevation for all angles. Default: \"eye-level shot\"."
            },
            "distance": {
              "type": "string",
              "enum": [
                "close-up",
                "medium shot",
                "wide shot"
              ],
              "description": "Camera distance for all angles. Default: \"medium shot\"."
            },
            "prompt": {
              "type": "string",
              "description": "Describe the SUBJECT and ambient environment (for example, a concise description of the visible subject, location, weather, props, and ambience). Do NOT describe camera motion, rotation, panning, orbiting, or 360-degree movement — camera motion is handled automatically. Do NOT put spoken dialogue here — use the dialogue parameter instead. Music is automatically suppressed — use generate_music separately."
            },
            "dialogue": {
              "type": "string",
              "description": "Spoken dialogue or narration for a SINGLE segment of the orbit video. This is applied ONLY to the segment specified by dialogueSegment (default: first segment). All other segments get foley/ambient audio only. Keep it brief — each segment is 2.5 seconds (~6 words max). If the user asks for dialogue in \"just the first segment\" or \"only at the start\", put the speech here and leave prompt for motion/foley only. If the user asks for dialogue in multiple/every segment, use dialogues instead."
            },
            "dialogues": {
              "type": "array",
              "items": {
                "type": "string"
              },
              "description": "Per-segment spoken dialogue lines for multiple orbit transitions. Use this when the user asks for dialogue in multiple segments, every turn, or before each 90-degree turn. With the default standard 360° orbit there are 4 transitions, so provide exactly 4 short lines in order. Each line should be brief enough for a 2.5 second segment (~6 words max). Preserve real names from the request or prior generated image; never invent placeholder speakers. For \"us\"/\"we\"/couple requests, make the named people speak together. Omit entries or use an empty string for segments that should have foley/ambient audio only. Do NOT also put these dialogue lines in prompt."
            },
            "dialogueSegment": {
              "type": "number",
              "description": "Which transition segment receives the dialogue (0-based index into the transition sequence). 0 = first transition (default), last index = wrap-back to front. With default angles there are 4 transitions (0-3). With custom angles the count equals angles.length + 1. Only used when dialogue is provided."
            },
            "angles": {
              "type": "array",
              "items": {
                "type": "string",
                "enum": [
                  "front-right quarter view",
                  "right side view",
                  "back-right quarter view",
                  "back view",
                  "back-left quarter view",
                  "left side view",
                  "front-left quarter view"
                ]
              },
              "description": "OMIT THIS PARAMETER for standard 360° orbits — the default (3 angles at 90° increments: right, back, left + source as front = 4 transitions) works for nearly all requests. Only provide this when the user explicitly asks for specific angles, a partial orbit, or extra-smooth rotation. Each additional angle costs extra credits and generation time. Values are clockwise azimuths between the source (front) and wrap-back."
            },
            "sourceImageIndex": {
              "type": "number",
              "description": "Which result image to orbit around (0-based). If the user picked a 1-based image number, subtract 1 and set this explicitly (number 3 -> 2). Omit only when the user did not choose a specific prior result; then the tool uses the latest result or original upload."
            }
          },
          "required": []
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "dance_montage",
        "description": "REQUIRED for ALL dance video requests — do NOT use animate_photo or generate_video for dances. Uses real choreography reference videos to transfer dance motion onto a photo via WAN 2.2 Animate Move. Output is always 9:16 480p portrait. Do NOT use this for bare TikTok/Reels/Shorts/social-video requests unless the user explicitly asks for a dance, choreography, dance trend, or named dance preset. UPLOADED PHOTO: When the user asks for a dance \"using this photo\" or \"with this photo\", call dance_montage directly on the uploaded photo; do NOT call edit_image/generate_image first just to prepare, stylize, restyle, reframe, make full-body, or reinterpret the subject. Words that identify a dance preset or vibe, such as \"Barbie\", \"Metric\", \"Black Sheep\", \"Rasputin\", or \"TikTok dance trend\", are NOT requests for image prep. Only create image prep first when the user explicitly asks for a new look, outfit, variation set, multiple characters, or loaded persona identity preservation. IMAGE PREP: When generating images for dance (via edit_image or generate_image), ALWAYS use aspectRatio=\"9:16\". CRITICAL — IMAGE COUNT: Generate exactly 1 image (numberOfVariations=1) for dance requests UNLESS the user explicitly asks for variations, different looks, or multiple characters (e.g. \"4 different outfits\", \"alternate between a cat and a dog\"). A single consistent image is used for ALL video segments to ensure visual consistency in the final stitched dance video. When the user DOES request multiple variations, batch them into ONE tool call using numberOfVariations + Dynamic Prompts — never split into multiple batches. PERSONAS: When personas are loaded, ALWAYS generate images via edit_image FIRST (using the persona reference photos for identity preservation), then call dance_montage — it will automatically use all generated images. Never use imagePrompt for persona dance requests — edit_image with persona context photos produces far better likeness. USING GENERATED IMAGES: When images have already been generated earlier in the conversation, simply call dance_montage WITHOUT sourceImageIndex — all previously generated images are used automatically as alternating montage segments. Do NOT tell the user to \"upload\" images that were already generated. Requires at least one uploaded photo, previously generated image, or loaded personas. Best results with photos of people.",
        "parameters": {
          "type": "object",
          "properties": {
            "dance": {
              "type": "string",
              "enum": [
                "rasputin",
                "big-guy",
                "keep-it-gangsta",
                "this-is-america",
                "chinese-new-year",
                "spongebob",
                "chanel",
                "crystal-light-aerobics-1988",
                "plastic-dream-sequence"
              ],
              "description": "Which dance choreography to use. \"rasputin\": Boney M - Rasputin (Viral Russian TikTok Dance, max 32s). \"big-guy\": Ice Spice - Big Guy (From \"The SpongeBob Movie: Search for SquarePants\" movie, max 11s). \"keep-it-gangsta\": Nhale ft. Dezzy Hollow - Keep it Gangsta (Hip-hop gangsta dance, max 21s). \"this-is-america\": Childish Gambino - This Is America (Iconic choreography from the This Is America music video, max 22s). \"chinese-new-year\": 弥渡山歌 (Midu Echoing) - Dan Thy (Chinese New Year Dance, Chinese Military Dance Trend, max 18s). \"spongebob\": SpongeBob - Stadium Rave (Jellyfish Jam Dance from SpongeBob SquarePants, max 27s). \"chanel\": Tyla - Chanel (Put me in Chanel dance, max 14s). \"crystal-light-aerobics-1988\": Crystal Light National Aerobics Championship 1988 (80s aerobics dance from the 1988 Crystal Light National Aerobics Championship, max 52s). \"plastic-dream-sequence\": Metric - Black Sheep (Barbie plastic dream sequence dance, max 28s).."
            },
            "duration": {
              "type": "number",
              "description": "Total video duration in seconds. Range: 8-30. OMIT this parameter unless the user explicitly requests a specific length — the handler defaults to the chosen dance's reference video length (capped at 30s) so the full choreography plays through. Each dance has its own max based on its reference video; the handler caps automatically.",
              "minimum": 8,
              "maximum": 30
            },
            "sourceImageIndex": {
              "type": "number",
              "description": "Which previously generated result image to use (0-based index). Use -1 for the original uploaded image. When omitted, all previously generated images are used automatically as alternating montage segments."
            },
            "imagePrompt": {
              "type": "string",
              "description": "Creative style/look for auto-generated images when no pre-generated images are available and no personas are loaded. For persona requests, always generate images via edit_image first — it preserves identity far better. This is a fallback only. If omitted, uses a default full-body portrait style."
            },
            "singleClip": {
              "type": "boolean",
              "description": "When true, renders the entire dance as one continuous clip (no stitching). Only works for durations ≤ 20s. Use when the user explicitly asks for a single video or one unbroken clip. Default: false (splits into segments for faster concurrent rendering)."
            }
          },
          "required": [
            "dance"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "sound_to_video",
        "description": "Generate video synchronized to audio. Use when the user has uploaded an audio file (mp3, wav, m4a, flac) and the audio is the primary sync target, especially uploaded-audio-only workflows. Also use after generate_music (\"turn that song into a video\", \"make a music video from that\"). Auto-detects generated audio from generate_music if no audio file is uploaded. Seedance animate_photo/generate_video can also attach uploaded audio as a loose @Audio reference when an image or video reference anchors the request; use this tool instead when the soundtrack itself should drive the video. If the user provides a reference image, use ltx25-ia2v by default (ltx23-ia2v is rollback); for lip-sync with a face image, use wan-s2v; if no image, use ltx25-a2v by default (ltx23-a2v is rollback). If the user wants dialogue/audio WITHOUT pre-existing audio, use animate_photo instead (LTX 2.5 and LTX 2.3 generate audio natively). Note: Persona voice clips from resolve_personas are NOT used by this tool — for persona voice identity in video, use animate_photo or generate_video with videoModel=\"ltx23\" because LTX 2.5 has no compatible ID-LoRA. LONG AUDIO ON SEEDANCE: Seedance 2.0 and Mini cap each clip at 15s; Seedance 2.5 renders up to 30s in one call, so prefer seedance2-5 for 16-30s audio instead of splitting. When the user uploads audio longer than the per-clip cap of the selected model and Seedance is selected (seedance2, seedance2-mini, or seedance2-5), do NOT clamp to 15s and drop the rest — split the run into multiple sound_to_video calls in the same turn (one per 15s segment, so a 20s audio becomes two clips: audioStart=0 duration=15, then audioStart=15 duration=5) and finish with a single stitch_video call referencing the resulting clip indices in order with audioIndex pointing at the same uploaded audio so the stitched output carries the full original soundtrack. LTX/WAN models accept up to 20s per clip, so single-call is fine for them. Use videoModel=\"wan3.0-video\" when the user explicitly requests Wan 3 audio-driven video. Use videoModel=\"wan3.0-spicy-video\" for Wan 3.0 Enhanced audio-driven video through MuleRouter provider ID w3.0-video; it supports adaptive ratios and provider prompt expansion. MINIMAX H3 AUDIO: only when the user asks for MiniMax H3 or FastH3 audio-driven video, use videoModel=\"minimax-h3-fasth3-ia2v-turbo\" with a first-frame image (sourceImageIndex), \"minimax-h3-fasth3-flfa2v-turbo\" with a first and a last frame (sourceImageIndex and endImageIndex), or \"minimax-h3-fasth3-a2v-turbo\" with no image; add \"-2stage\" only for 1080p, 1440p or 2K H3 output. Without an explicit MiniMax H3 or FastH3 request keep the LTX 2.5 defaults. The MiniMax H3 selectors on generate_video and animate_photo cannot take an uploaded audio track; send uploaded-audio H3 requests here.",
        "parameters": {
          "type": "object",
          "properties": {
            "prompt": {
              "type": "string",
              "description": "Describe the video like a cinematographer. Let the audio define timing — use the prompt for visual interpretation. One flowing paragraph, present tense, specific natural language.\n\nLITERAL PROMPT OVERRIDE: If the user explicitly says not to modify the prompt, or to use it exactly/verbatim/as-is, copy the identified prompt text verbatim instead of applying these construction rules unless a hard requirement is missing. For Seedance or Wan 3, set expandPrompt=false.\n\nSTRUCTURE: shot/style and scale → subject → environment, lighting, color, texture, atmosphere → visual action synced to audio → camera movement. For LTX 2.3 image+audio mode, do not re-describe static details already visible in the reference image; focus on motion, action, camera, and how the image responds to the audio.\n\nMOTION PACING: Scale complexity to duration. <=6s: 1 main visual beat + 1 simple camera move. Around 10s: 2-3 clear beats + 1 camera move. >10s: up to 4 beats in clear sequence. Let the audio define timing, but avoid stacking subject, camera, and environment motion in short clips.\n\nBLOCKING: Direct layout when it affects the shot: left/right placement, foreground/background, facing direction, and relative distance between subjects.\n\nLIP-SYNC: Shot framing, speaker's appearance and setting, physical performance synced to audio — gestures, expressions, jaw movement between phrases. Include acting beats.\n\nMUSIC VISUALIZATION: Visual style, environment, and how elements react to rhythm and energy.\n\nAUDIO-REACTIVE: Motion and visual changes that correspond to sounds in the track.\n\nLTX VOCABULARY: camera (tracking, dolly, pan, tilt, handheld, static frame), lighting/atmosphere (golden hour, neon glow, dramatic shadows, fog, rain, smoke, reflections), scale/pacing (expansive, epic, intimate, claustrophobic, slow motion, time-lapse, lingering shot, continuous shot), style/genre (film noir, painterly, cyberpunk, stop-motion, claymation, 2D/3D animation, hand-drawn, fantasy, thriller, experimental film).\n\nAVOID: Vague prompts, too many competing visual elements, abstract descriptions without visible behavior, rigid numeric constraints, readable text or logos. QUOTING RULE: ONLY use double quotes for spoken dialogue. Never quote on-screen text, overlay text, titles, captions, signs, or any visual text — describe them without quotes.\n\nNON-SEEDANCE POSITIVE CONSTRAINTS: For ltx25-ia2v, ltx25-a2v, ltx23-ia2v, ltx23-a2v, and wan-s2v, prompt is a positive prompt. Translate user avoid/no/don't constraints into affirmative production constraints instead of copying negative phrasing. Preserve exact quoted visible text or dialogue when the user explicitly requests it; keep surrounding surfaces blank.\n\nBATCH VARIATIONS: When numberOfVariations > 1, use Dynamic Prompt syntax to vary the visual interpretation while keeping audio sync intent consistent. This is one Sogni project with multiple jobs, so prefer it when all outputs share the same audio source/window, image source, model, duration, dimensions, and parameters and only prompt text varies. Example: \"{abstract neon visualization|nature scene with swaying trees|urban street with rain} synced to the beat\".\n\nMINIMAX H3 AUDIO SELECTORS: write the request plainly (subject, action, camera, the voice or sound heard in the upload and any exact words spoken in it); the MiniMax H3 prompt shaper turns it into the H3 contract for the matching image-to-video, first-and-last-frame or text-to-video mode. Describe the uploaded audio as it is, and do not invent other dialogue or music."
            },
            "expandPrompt": {
              "type": "boolean",
              "description": "Seedance and Wan 3 only. Whether to expand the prompt before dispatch. Defaults to true. For Wan 3, a successful Sogni expansion disables Alibaba prompt_extend to prevent a second rewrite; false disables both expansion layers so exact prompts remain exact."
            },
            "negativePrompt": {
              "type": "string",
              "description": "Advanced LTX 2.5/LTX 2.3/WAN only. The LTX A2V and IA2V workflows accept this separate negative prompt. Use it only when the user explicitly asks to set one. Do not set for Seedance.\n\nWan 3 has no negativePrompt request field; do not set this for wan3.0-video.\n\nMiniMax H3 has no negative-prompt input; do not set this for the MiniMax H3 FastH3 audio selectors and state exclusions positively in prompt."
            },
            "audioSourceIndex": {
              "type": "number",
              "description": "Index of the uploaded audio file to use (0-based, from uploaded files list). If only one audio file is uploaded, use 0. If no audio was uploaded but generate_music was used earlier, omit this — the tool will automatically find the generated audio."
            },
            "sourceImageIndex": {
              "type": "number",
              "description": "Optional index of an uploaded image to use as the starting frame (0-based). Required for lip-sync models (WAN S2V). For audio-only-to-video models (LTX 2.3 A2V), this is optional — omit it to generate video purely from text + audio. MiniMax H3 FastH3 audio guide: \"minimax-h3-fasth3-ia2v-turbo\" and \"minimax-h3-fasth3-flfa2v-turbo\" (and their -2stage forms) require this first frame; \"minimax-h3-fasth3-a2v-turbo\" and its -2stage form take no image, so omit it for them."
            },
            "endImageIndex": {
              "type": "number",
              "description": "Which image to use as the LAST frame, indexed like animate_photo's endImageIndex: negative indices for uploaded images (-1 = first upload, -2 = second upload) and 0-based non-negative indices for generated results. Only \"minimax-h3-fasth3-flfa2v-turbo\" and \"minimax-h3-fasth3-flfa2v-turbo-2stage\" take it, and they require it together with sourceImageIndex (for two uploaded images use sourceImageIndex=-1 and endImageIndex=-2). Omit it for every other videoModel."
            },
            "audioStart": {
              "type": "number",
              "description": "Start offset in seconds into the audio track. Use when the user says \"start 20 seconds in\", \"skip the intro\", \"use the chorus at 1:30\", etc. Default: 0 (beginning of audio). The video will be synced to the audio starting from this point. The MiniMax H3 FastH3 audio selectors take audioStart too: the clip uses the upload from audioStart for its own length.",
              "minimum": 0
            },
            "duration": {
              "type": "number",
              "description": "Video duration in seconds. Default: 5. Per-model range: LTX/WAN 2.2 = 2-20s; Wan 3 = 2-30s; Seedance 2.0 and Mini = 4-15s; Seedance 2.5 = 4-30s. For music videos, use the maximum duration the selected model allows because the audio is usually longer than the video limit. Use when the user explicitly requests a specific length. MiniMax H3 FastH3 audio selectors render 124-362 frames on a 17-frame grid at a fixed 24 fps, so the clip runs about 5.2-15.1 seconds and a length outside that window snaps to it; for longer audio pick the window with audioStart.",
              "minimum": 2,
              "maximum": 30
            },
            "ratio": {
              "type": "string",
              "enum": [
                "adaptive",
                "16:9",
                "4:3",
                "1:1",
                "3:4",
                "9:16"
              ],
              "description": "Wan 3 only. Output ratio; \"adaptive\" derives it from the input."
            },
            "watermark": {
              "type": "boolean",
              "description": "Alibaba wan3.0-video only. Add the visible watermark. Defaults to false."
            },
            "referenceFileUrl": {
              "type": "string",
              "description": "Alibaba wan3.0-video only. One public HTTPS document URL for additional audio-driven context (DOCX/DOC/XLSX/XLS/PPTX/PPT/PDF/TXT/KEY/PAGES/NUMBERS/Markdown, up to 100 MB; PDF/DOCX/DOC/PPTX/PPT/KEY/PAGES up to 50 pages). Mutually exclusive with referenceLinkUrl."
            },
            "referenceLinkUrl": {
              "type": "string",
              "description": "Alibaba wan3.0-video only. One public HTTPS webpage URL for additional audio-driven context. Mutually exclusive with referenceFileUrl."
            },
            "videoModel": {
              "type": "string",
              "enum": [
                "wan-s2v",
                "seedance2",
                "seedance2-mini",
                "seedance2-5",
                "ltx25-ia2v",
                "ltx25-a2v",
                "ltx23-ia2v",
                "ltx23-a2v",
                "wan3.0-video",
                "wan3.0-spicy-video",
                "minimax-h3-fasth3-ia2v-turbo",
                "minimax-h3-fasth3-ia2v-turbo-2stage",
                "minimax-h3-fasth3-flfa2v-turbo",
                "minimax-h3-fasth3-flfa2v-turbo-2stage",
                "minimax-h3-fasth3-a2v-turbo",
                "minimax-h3-fasth3-a2v-turbo-2stage"
              ],
              "description": "\"ltx25-ia2v\" (default with image) and \"ltx25-a2v\" (default without image): LTX 2.5 image+audio and audio-only modes; Fast, HQ, and Pro currently use the release-validated official Distilled INT8 workflows. The Dev checkpoints are not publicly routed until upstream publishes and Sogni validates official ComfyUI Dev recipes. Video model. \"ltx23-ia2v\" (rollback with image): LTX 2.3 image+audio to video, audio-reactive with a reference image; Fast/HQ use the distilled 8-step worker and Default Media Quality Pro uses the non-distilled dev worker. \"ltx23-a2v\" (rollback without image): LTX 2.3 audio-only to video, no image needed, creates video purely from text prompt + audio with the same quality-tier routing. \"wan-s2v\": WAN 2.2 sound-to-video, best for lip-sync with a face image, fast 4-step. \"seedance2\": full Seedance 2.0 audio-reference video, 4-15s. \"seedance2-mini\": Seedance 2.0 Mini, 720p cap, fastest/lower-cost Seedance option. Seedance quality is selected only by this model value: pick \"seedance2-mini\" for faster/lower-cost drafts or explicit Mini requests, and pick \"seedance2\" for full-quality Seedance or 1080p/4K. Do not infer the Seedance model from Default Media Quality Fast/HQ/Pro. \"seedance2-5\": Seedance 2.5, the newest Seedance generation — 480p, 720p, and 1080p (4K is unsupported), 4-30s per clip at a fixed 24 fps, native audio, first-and-last-frame conditioning, and a much larger reference budget than the 2.0 family: up to 30 images, 10 videos, and 10 audios, with up to 50 reference media files total, subject to those per-modality caps. Choose \"seedance2-5\" when the user asks for Seedance 2.5, wants a single continuous Seedance clip longer than 15s (2.5 renders up to 30s in one call instead of being split and stitched), or wants a first-and-last-frame Seedance transition. Keep \"seedance2\" for 4K requests; Seedance 2.5 supports up to 1080p. For Seedance audio-reference prompts, preserve exact spoken dialogue when the user supplied it, and assign @Image1/@Audio1 roles. If the user asks for speech without words, describe the vocal performance without inventing quoted dialogue. Treat lip-sync, voice cloning, and real-human reference behavior as provider-sensitive rather than guaranteed. Omit to auto-select based on whether an image is present. \"wan3.0-video\" is Alibaba Wan 3 and \"wan3.0-spicy-video\" is MuleRouter w3.0-video. Both render 2-30s at fixed 30 fps with optional native audio, provider prompt expansion, 480p/720p/1080p, adaptive/fixed ratios, first/last frames, and up to 10 image/5 video/5 audio references. Only Alibaba wan3.0-video accepts document/web context and watermark. Frame anchors and loose references are mutually exclusive. Do not send negativePrompt; video references are loose conditioning for a new result, not source-video editing or extension. MiniMax H3 FastH3 audio guide: \"minimax-h3-fasth3-ia2v-turbo\" (first-frame image from sourceImageIndex plus the audio), \"minimax-h3-fasth3-flfa2v-turbo\" (first frame from sourceImageIndex, last frame from endImageIndex, plus the audio) and \"minimax-h3-fasth3-a2v-turbo\" (audio only; set neither sourceImageIndex nor endImageIndex) run the four-step FastH3 engine with the uploaded audio driving the picture from frame 0, and the output keeps that audio as its soundtrack. Choose them only when the user asks for MiniMax H3 or FastH3; never pick them in place of the LTX 2.5 defaults. They render 124-362 frames on the H3 17-frame grid at a fixed 24 fps (about 5.2-15.1 seconds) on a 32px grid within 1344x768, so use targetResolution 768 or omit it. audioStart picks the window of the upload. MiniMax H3 catalog and Personal LoRAs are supported across audio modes and tiers. generateAudio=false and negativePrompt are not supported. Each costs the FastH3 price of its image-to-video, first-and-last-frame or text-to-video mode, 4 Spark per second. \"minimax-h3-fasth3-ia2v-turbo-2stage\", \"minimax-h3-fasth3-flfa2v-turbo-2stage\" and \"minimax-h3-fasth3-a2v-turbo-2stage\" are the two-stage forms: the same inputs rendered on the FastH3 canvas, then enlarged 2x and refined, delivered at twice the canvas with the same length and audio. For them targetResolution names the delivered class: 1080 renders a 544px short-edge canvas (960x544 is delivered at 1920x1088) for 10 Spark per second, and 1440 or omitted renders the 768p canvas for 2K (1344x768 is delivered at 2688x1536) for 16 Spark per second. The audio two-stage selectors have no 720p price class (720 would render a 384px canvas at the 1080p rate), so for 720p or 768p audio-guided H3 output use the regular selector. The estimate prices every request."
            },
            "generateAudio": {
              "type": "boolean",
              "description": "Whether the returned video should include audio. Omit to include audio by default; set false when the user asks for silent output or no audio. The reference audio is still required and still drives generation even when the returned video has no audio track. MiniMax H3 FastH3 audio selectors always deliver the uploaded audio: omit generateAudio for them (false is refused)."
            },
            "numberOfVariations": {
              "type": "number",
              "description": "Number of video variations to generate (1-16). Use with one Dynamic Prompt branch when all variations share the same audio source/window, image source, model, duration, dimensions, and parameters and only prompt text varies. This creates one Sogni project with multiple jobs. Default: 1.",
              "minimum": 1,
              "maximum": 16
            },
            "targetResolution": {
              "type": "number",
              "description": "Short-side video resolution target in pixels. Use when the user asks for a bare named resolution such as \"480p\", \"720p\", \"1080p\", \"2160p\", or \"4K\" without exact pixels or an output orientation. This preserves the source/reference aspect ratio. Do NOT set exact-pixel aspectRatio for bare named resolution requests. If the user says \"720p portrait\", \"720p landscape\", \"4K portrait\", or \"4K landscape\", use exact-pixel aspectRatio instead."
            },
            "aspectRatio": {
              "type": "string",
              "description": "Do NOT set unless the user explicitly requests an aspect ratio, format, orientation, or exact pixel dimensions. When a reference/source image is used and the user did not ask to change its shape, omit this field so the handler preserves the selected source image's own ratio.\n\nFormats: \"16:9\", \"9:16\", \"4:5\", \"1:1\", \"4:3\", \"3:2\", \"21:9\", or exact pixels like \"1920x1080\".\n\nCRITICAL: When the user specifies exact pixel dimensions (e.g., \"1280x720\", \"1080x1920\", \"1920x1080\", \"3840x2160\") or an orientation-qualified named resolution (e.g., \"720p landscape\", \"720p portrait\"), use the exact pixel format, NOT a ratio like \"16:9\" or \"9:16\". Exact user-requested dimensions override the selected default media quality, including Pro/HQ defaults. A bare named video resolution like \"720p resolution\" is only a resolution tier/short-side request; do not turn it into landscape pixels and do not set aspectRatio unless the user also states landscape, portrait, vertical, horizontal, or exact pixels. If requested pixels are in bounds but not on the model's pixel step, still pass the user's exact pixel request; the handler snaps to the nearest supported size internally. Only use ratio format when the user says a generic format name without pixel dimensions.\n\nMappings (use ONLY when user does NOT specify pixel dimensions): landscape/widescreen/YouTube/cinematic → \"16:9\". portrait → \"9:16\". TikTok/Reels/IG Reels → \"1080x1920\". ultrawide/cinema scope → \"21:9\". Instagram post → \"4:5\". square → \"1:1\". standard/TV → \"4:3\". 720p landscape → \"1280x720\". 720p portrait → \"720x1280\". 1080p landscape → \"1920x1080\". 1080p portrait/HD portrait → \"1080x1920\". 4K landscape → \"3840x2160\". 4K portrait → \"2160x3840\". Never set for generic requests like \"make a video\"."
            },
            "outputFormat": {
              "type": "string",
              "enum": [
                "mp4",
                "mov"
              ],
              "description": "Video container. Defaults to mp4. MOV is supported only by Seedance 2.5; choose it when the user requests MOV for editing."
            },
            "returnLastFrame": {
              "type": "boolean",
              "description": "Seedance 2.5 only. Set true to export a separate image of the final frame alongside the video. The result includes lastFrameUrl, which can be used as the first-frame image for a subsequent clip. Defaults to false; this does not extend the video automatically."
            },
            "loras": {
              "type": "array",
              "minItems": 1,
              "maxItems": 8,
              "items": {
                "type": "string",
                "minLength": 1
              },
              "description": "Ordered MiniMax H3 LoRA IDs, including owned ready Personal LoRAs. Stack up to 8 in one request; order matters because the adapters apply in sequence and do not commute. Keep this array positionally aligned with loraStrengths. The first render with an uncached LoRA takes longer to start while the worker downloads it.\n\nAccepted only when videoModel is one of \"minimax-h3-fasth3-ia2v-turbo\", \"minimax-h3-fasth3-ia2v-turbo-2stage\", \"minimax-h3-fasth3-flfa2v-turbo\", \"minimax-h3-fasth3-flfa2v-turbo-2stage\", \"minimax-h3-fasth3-a2v-turbo\", \"minimax-h3-fasth3-a2v-turbo-2stage\". Every other video model on this tool loads no LoRAs and silently ignores these arrays, so set videoModel to an H3 mode in the same call when the user asks for one.\n\nFive LoRAs are published for MiniMax H3 today and the set differs by mode, so GET /v1/loras/comfy?modelId=<model> is authoritative for the mode in hand and carries exact ranges, maturity flags, and anything published since. h3-realism-people (fal) is a realism pass trained on live-action footage of people: it restores skin texture and pores, stray hairs, fabric weave and a fine sensor grain that the base model smooths away, and holds up in close-up. It is the only one gated on a trigger word — put r34l1sm near the FRONT of the prompt, or the render comes back as ordinary H3 with no error. h3-vbvr-video-reasoning is a prompt-adherence pass that holds the model to what was asked instead of improvising. h3-natural-face-speech (AdaptiveVision) makes people talking on camera look and sound more natural: cheeks, brows, jaw and lips move together as in real speech, and spoken English comes through clearer; use it for talking-head shots such as vlogs, podcasts, interviews and presenters. h3-better-motion (AdaptiveVision) gives people more natural, consistent body movement — weight shifts, strides, turns and gestures that follow through — for dance, sport, walking and other full-body shots. Both AdaptiveVision LoRAs work best with short, simple prompt sentences and are not validated on reference-to-video. h3-mystic-xxx-v4 is an uncensored adult fine-tune. Personal imports are discovered through authenticated GET /v1/loras/personal/catalog; use only owned ready ids with the selected model in modelIds, and respect their strength range and requirements. Do not invent ids."
            },
            "loraStrengths": {
              "type": "array",
              "minItems": 1,
              "maxItems": 8,
              "items": {
                "type": "number"
              },
              "description": "Strength for each LoRA in loras, in the same order. Omitting the array applies 1.0 to every LoRA, which is not every LoRA's catalog default and for h3-realism-people is already at the top of its band, so send explicit values. Video LoRAs are positive-only — unlike the bipolar Krea 2 image sliders, a negative value is not an inverse effect and 0 is off. h3-realism-people takes 0-2 and its catalog default is 0.8; 0.6-1 is the usable band. It also pulls the camera in as it climbs: at 1.5 and above the shot reliably recomposes and the grade darkens, which on an image-conditioned mode can crop the subject out of the frame the user supplied. Raise it above 1 only when the user asks for more, and prefer the default when they supplied a first or last frame. h3-vbvr-video-reasoning and h3-mystic-xxx-v4 both take 0-1 and do default to 1.0, with usable bands of 0.7-1 and 0.2-1. h3-natural-face-speech and h3-better-motion take 0-1.5 and default to 0.6; their usable band is 0.4-0.8."
            }
          },
          "required": [
            "prompt"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "extend_video",
        "description": "Extend a video by adding new time to the end. Works on BOTH videos previously rendered in this session AND user-uploaded videos — set videoIndex to a negative number (e.g. -1) to target an uploaded video when no prior render exists. The base video is auto-selected from the most recent video in this session unless videoIndex is set. For LTX 2.5/2.3 base clips, the tool extracts the last frame and renders an image-to-video continuation; new non-Seedance continuations default to LTX 2.5. For Seedance base clips, the tool extracts a trailing reference segment and renders a video-to-video continuation. Returns both the standalone new segment and a spliced composite (base + new segment). Use when the user asks to \"make it longer\", \"extend the video\", \"add another N seconds\", \"continue the scene\", \"add an outro/bumper to the end\", etc. Prefer this over generate_image+animate_photo+stitch_video for \"add a bumper/outro to this video\" — extend_video preserves the original base bytes, audio, and timing instead of re-encoding them. Do not use this tool to render fresh videos from scratch — call generate_video or animate_photo for that. Output durations follow each model's native limits (LTX 2-20s; Seedance 2.0 and Mini 4-15s; Seedance 2.5 4-30s) for the new segment alone.",
        "parameters": {
          "type": "object",
          "properties": {
            "prompt": {
              "type": "string",
              "description": "What should happen during the extension — describe motion, action, dialogue, and audio for the appended seconds, NOT the entire video. For LTX continuations, preserve user-provided spoken dialogue in double quotes; if speech is requested without exact words, describe the delivery without inventing quoted dialogue. If the user did not specify what should happen, write a brief continuation that preserves the existing tone (e.g. \"the scene continues with the same camera and pacing\")."
            },
            "duration": {
              "type": "number",
              "description": "Length in seconds of the new appended segment (NOT total final length). LTX 2-20s; Seedance 2.0 and Mini 4-15s; Seedance 2.5 4-30s. Default: 5.",
              "minimum": 2,
              "maximum": 30
            },
            "videoIndex": {
              "type": "number",
              "description": "Which video result to extend. Default: -1 (most recent video in this session). Use 0-based non-negative indices for prior tool result videos. Use negative indices for uploaded videos: -1 = most recent video result OR first uploaded video when no prior render exists."
            },
            "videoModel": {
              "type": "string",
              "enum": [
                "auto",
                "ltx25",
                "ltx23",
                "seedance2",
                "seedance2-mini",
                "seedance2-5"
              ],
              "description": "Which model to use for the new segment. Default: \"auto\" — preserve Seedance for a Seedance base and otherwise use LTX 2.5. Use ltx23 only for explicit rollback. Override only when the user explicitly requests a different model."
            },
            "keepOriginalAudio": {
              "type": "boolean",
              "description": "Has no effect for extend_video (the new segment is appended after the base, so the base audio is always preserved through the original portion and the new segment carries its own audio). Reserved for parity with replace_video_segment."
            }
          },
          "required": [
            "duration"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "replace_video_segment",
        "description": "Modify a portion of an existing video while keeping the rest intact — either by regenerating that slice fresh or by splicing in another existing clip. Operates on a [startSeconds, endSeconds] window inside a base video; everything outside the window stays exactly as it was. Works on BOTH videos previously rendered in this session AND user-uploaded videos (set videoIndex to a negative number to target an uploaded video when no prior render exists). WHEN TO USE: any request to change part of one video while keeping the rest, to put another clip inside another video at a specific position, or to interleave time slices of multiple videos. Plain language: \"regenerate from 5s to 10s\", \"redo the last 3 seconds\", \"swap out the middle\", \"replace the bumper at the end\", \"swap the end card\", \"change the outro / intro / ending / last clip\", \"replace 2s-4s with a stronger expression\", \"splice video 2 into video 1\", \"stitch video 2 into the middle of video 1\", \"insert the second clip at 5s\", \"alternate 1 second from each video\". The word \"stitch\" in the user request does not by itself mean stitch_video — when the user clearly wants insertion or in-place replacement, this tool is the right one. WHEN NOT TO USE — prefer stitch_video instead: the user wants to concatenate whole clips end-to-end without modifying their interiors (\"stitch these together\", \"play A then B\", \"add a bumper before / after\"). SPLICING EXISTING CLIPS: pass replacementVideoIndex when the replacement already exists as an uploaded or generated video — do not call generate_video / animate_photo / video_to_video in that case. Set endSeconds=startSeconds when the user asks for an insertion that should not remove time from the base video. TIME-SLICED INTERLEAVING (\"alternate 1 second from each video\"): pass replacementStartSeconds and replacementEndSeconds to cut the next source slice out of the replacement video before splicing it into the base. Repeat this call for each alternating window. By default use replacement windows (endSeconds = startSeconds + sliceDuration); use insertion windows (endSeconds = startSeconds) only when the user explicitly asks to lengthen the output by inserting extra slices. replacementStartSeconds and replacementEndSeconds must be concrete non-negative seconds; never use -1 as an end-of-source sentinel. PREFER this over re-running generate_video / animate_photo on the original prompt when the user only wants part of the video changed — re-rendering wastes credits, loses the unchanged sections, and breaks the original timing. If the user does not specify the exact start/end seconds (e.g. \"replace the bumper at the end\"), call analyze_video first to identify the correct window, OR derive it from the storyboard timing already in the conversation (e.g. last beat's time range). Do not guess wildly — pick a sensible bumper/end-card window such as the final 1-3 seconds when the storyboard says scene_07 is 14-15s. Returns both the standalone replacement clip and the spliced composite. For LTX 2.5/2.3 and Wan 2.2 base videos the tool locks both ends with first/last-frame keyframes for seamless edges; new non-Seedance LTX segments default to 2.5. For Seedance base videos the tool uses the original window as a reference for video-to-video transformation. If a requested window is shorter than the selected model's native render minimum, the handler renders a slightly larger handled clip, trims the result back to the requested seconds, then splices exactly that requested range. By default the regenerated segment's audio replaces the original audio in the [startSeconds, endSeconds] window, so new motion stays in sync with new sound. Pass keepOriginalAudio=true only when the user explicitly asks to keep the existing audio — phrasings like \"keep the audio\", \"leave the original audio\", \"preserve the music/score/dialogue\", \"don't change the audio\". If the user uses an ambiguous phrasing such as \"with the audio\" (which could mean either \"with the original audio kept\" or \"with new audio\"), DO NOT call this tool yet — first ask the user whether to preserve or replace the original audio in the replaced window. When replacementVideoIndex is set, the existing replacement clip's own audio is used; pass keepOriginalAudio=true only when the user explicitly wants the base video audio to stay over the replacement window.",
        "parameters": {
          "type": "object",
          "properties": {
            "startSeconds": {
              "type": "number",
              "description": "Start of the window (in seconds) inside the base video that should be regenerated. Must be ≥ 0 and < endSeconds. For \"the last N seconds\" requests, set startSeconds = max(0, baseDuration - N). When unsure of the exact base duration, you may pass a sentinel value of -1 to mean \"from the end of the base video\"; the handler will resolve it after probing."
            },
            "endSeconds": {
              "type": "number",
              "description": "End of the window (in seconds) inside the base video. For regenerated segments it must be > startSeconds and ≤ base video duration. When replacementVideoIndex is set, endSeconds may equal startSeconds to insert the replacement clip at that timestamp without removing any base-video time. For alternating/interleaved time-slice edits, use endSeconds=startSeconds+sliceDuration so the source slice replaces that base window; do not use insertion unless the user explicitly asks to lengthen the output. Pass -1 to mean \"until the end of the base video\". Windows shorter than the selected model's native render minimum are rendered with handles and trimmed before splicing; windows longer than the model maximum must be split."
            },
            "prompt": {
              "type": "string",
              "description": "What should happen in the replaced window — motion, action, dialogue, audio. For LTX include exact spoken words in double quotes when speech is requested. For Seedance V2V describe the transformation relative to the existing visuals (the original window is provided as a reference clip)."
            },
            "videoIndex": {
              "type": "number",
              "description": "Which video to edit. Default: -1 (most recent video in this session, falling back to the first uploaded video when no prior render exists). Use 0-based non-negative indices for prior tool result videos. For uploaded videos with no prior render, leave this absent or pass -1 — the handler will pick the uploaded base automatically."
            },
            "replacementVideoIndex": {
              "type": "number",
              "description": "Optional existing video clip to splice into the base video instead of regenerating a segment. Non-negative values reference prior generated videos; negative values reference uploaded videos (-1 = first uploaded video, -2 = second, etc.). Use for requests like \"splice video 2 into video 1\", \"replace 5s to 15s with uploaded clip 2\", or alternating/interleaved edits that pull timed slices from another existing clip. When this is set, the operation is pure ffmpeg post-production and keeps the replacement clip audio unless keepOriginalAudio=true."
            },
            "replacementStartSeconds": {
              "type": "number",
              "minimum": 0,
              "description": "Optional start time, in seconds, inside replacementVideoIndex. Must be a concrete non-negative source time; do not use -1 sentinels for replacement source windows. Use with replacementEndSeconds when only a slice of the replacement clip should be spliced. Example: alternating 1-second clips from video 1 and video 2 should replace base window 1..2 with the first replacement slice by setting replacementVideoIndex=-2, replacementStartSeconds=0, replacementEndSeconds=1."
            },
            "replacementEndSeconds": {
              "type": "number",
              "minimum": 0,
              "description": "Optional end time, in seconds, inside replacementVideoIndex. Must be a concrete non-negative source time greater than replacementStartSeconds. Do not pass -1 to mean \"end of replacement video\"; use the known uploaded/generated clip duration from metadata for routine time-sliced edits, or omit both replacementStartSeconds and replacementEndSeconds to use the whole replacement clip. Do not call analyze_video just to learn duration."
            },
            "videoModel": {
              "type": "string",
              "enum": [
                "auto",
                "ltx25",
                "ltx23",
                "wan22",
                "seedance2",
                "seedance2-mini",
                "seedance2-5"
              ],
              "description": "Which model to use for the new segment. Default: \"auto\" — preserve Seedance or WAN for matching base clips and otherwise use LTX 2.5. Use ltx23 only for explicit rollback. Override only when the user explicitly requests a different model."
            },
            "keepOriginalAudio": {
              "type": "boolean",
              "description": "When true, the audio from the original [startSeconds, endSeconds] window is muxed onto the regenerated visuals so the user keeps the original dialogue/score. When false (default), the new clip's own audio is used (LTX renders fresh audio; Seedance V2V depends on generateAudio). Default: false. Set true only when the user explicitly asks to preserve the existing audio; if the user uses an ambiguous phrasing like \"with the audio\", ask the user to clarify rather than guessing."
            }
          },
          "required": [
            "startSeconds",
            "endSeconds"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "overlay_video",
        "description": "Burn text and/or logo/watermark image overlays onto a previously rendered or uploaded video. Use when the user asks to add a title, caption, label, watermark, brand logo, sponsor mark, lower-third, tagline, sticker, or any persistent text/graphic over the existing video frames. Multiple overlays can be supplied in one call (e.g. a corner logo plus a top-center title). Each overlay can optionally be limited to a [startSeconds, endSeconds] time range. When the user asks for an overlay to appear for a specific window (for example \"2 seconds in the middle\"), set startSeconds/endSeconds on the overlay item in the same call. Negative startSeconds/endSeconds are relative to the end of the base video, so startSeconds=-2 with omitted endSeconds means \"the last 2 seconds\". When replacing a video time window with an uploaded still image or screenshot, use an image overlay with widthPct=100 and fit=\"cover\" for that window. This is a pure ffmpeg post-production op — it does not regenerate the video. Do not use for generative intro/outro/bumper/end-card/start-card requests; those add or regenerate video time and should use extend_video or replace_video_segment. Do not call it again just to refine default size/placement after it succeeds; finalize and wait for user feedback. Do not use for animated typography, kinetic captions, or moving stickers; this lays down static overlays only.",
        "parameters": {
          "type": "object",
          "properties": {
            "sourceVideoIndex": {
              "type": "number",
              "description": "Which video to overlay onto. Omit to use the most recent generated video, or the first uploaded video when no generated video exists. Non-negative values are 0-based indices into prior generated video results. Negative values reference uploaded videos: -1 = first uploaded video, -2 = second, etc., falling back to the most recent generated video when no uploads exist."
            },
            "overlays": {
              "type": "array",
              "minItems": 1,
              "description": "Ordered list of overlays to burn in. Each overlay is rendered on top of all previous overlays. Either kind=\"text\" (with `text` and styling) or kind=\"image\" (with `sourceImageIndex`).",
              "items": {
                "type": "object",
                "properties": {
                  "kind": {
                    "type": "string",
                    "enum": [
                      "text",
                      "image"
                    ],
                    "description": "Overlay kind. \"text\" renders drawtext; \"image\" composites an existing image asset."
                  },
                  "position": {
                    "type": "string",
                    "enum": [
                      "top-left",
                      "top-center",
                      "top-right",
                      "center",
                      "bottom-left",
                      "bottom-center",
                      "bottom-right"
                    ],
                    "description": "Anchor position on the frame. Pixel offsets nudge inward; the renderer pads each anchor by a small safe margin so overlays do not touch the frame edge."
                  },
                  "offsetX": {
                    "type": "number",
                    "description": "Optional horizontal offset in pixels. Positive = inward from the anchor edge."
                  },
                  "offsetY": {
                    "type": "number",
                    "description": "Optional vertical offset in pixels. Positive = inward from the anchor edge."
                  },
                  "startSeconds": {
                    "type": "number",
                    "description": "Show the overlay from this time. Default 0 (show from the start). Negative values are relative to the end of the base video; startSeconds=-2 means start 2 seconds before the end."
                  },
                  "endSeconds": {
                    "type": "number",
                    "description": "Hide the overlay at this time. Default = full video duration. Negative values are relative to the end of the base video."
                  },
                  "text": {
                    "type": "string",
                    "description": "Overlay text. Required when kind=\"text\". Use plain text; line breaks are honored."
                  },
                  "fontSizePct": {
                    "type": "number",
                    "minimum": 1,
                    "maximum": 30,
                    "description": "Font size as a percentage of the video height. Default: 6 (≈ 43px on a 720p frame). Only valid when kind=\"text\"."
                  },
                  "color": {
                    "type": "string",
                    "description": "Text fill color (CSS hex like \"#FFFFFF\" or named ffmpeg color). Default \"#FFFFFF\". Only valid when kind=\"text\"."
                  },
                  "outlineColor": {
                    "type": "string",
                    "description": "Text outline color. Default \"#000000\" with a thin stroke for legibility. Only valid when kind=\"text\"."
                  },
                  "backgroundColor": {
                    "type": [
                      "string",
                      "null"
                    ],
                    "description": "Optional rgba background pill behind the text (e.g. \"rgba(0,0,0,0.5)\"). null = no box. Only valid when kind=\"text\"."
                  },
                  "fontWeight": {
                    "type": "string",
                    "enum": [
                      "normal",
                      "bold"
                    ],
                    "description": "Default \"normal\". Only valid when kind=\"text\"."
                  },
                  "sourceImageIndex": {
                    "type": "number",
                    "description": "Which image to overlay. Required when kind=\"image\". Non-negative values are 0-based indices into prior generated image results. Negative values reference uploaded images in image-only order: -1 = first uploaded image, -2 = second, etc. If the user uploaded one video and one logo image, the logo is sourceImageIndex=-1."
                  },
                  "widthPct": {
                    "type": "number",
                    "minimum": 1,
                    "maximum": 100,
                    "description": "Logo width as a percentage of the video width. Default: 15. Only valid when kind=\"image\"."
                  },
                  "opacity": {
                    "type": "number",
                    "minimum": 0,
                    "maximum": 1,
                    "description": "Image overlay opacity, 0..1. Default 1.0 (fully opaque). Only valid when kind=\"image\"."
                  },
                  "fit": {
                    "type": "string",
                    "enum": [
                      "contain",
                      "cover"
                    ],
                    "description": "Image sizing mode. Default \"contain\" scales by widthPct and preserves the full overlay image. \"cover\" scales/crops the image to cover the full video frame; use with widthPct=100 for screenshot/still-frame replacement windows."
                  }
                },
                "required": [
                  "kind",
                  "position"
                ]
              }
            }
          },
          "required": [
            "overlays"
          ]
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "add_subtitles",
        "description": "Burn subtitles into a video from caller-supplied cues or an SRT/VTT string. Use when the user asks to add captions, subtitles, on-screen dialogue, or burned-in lyrics to a video. Either pass `cues` as an array of {startSeconds, endSeconds, text}, or pass a full `srt` string. Pace cues like real subtitles: split the script into multiple short cues (typically 1.5–4 seconds each, ~1–8 words per cue, roughly 15–20 characters per second of cue duration). Never burn a single cue that spans the entire clip — even a static image should get progressively revealed lines, not one paragraph held on screen the whole time. Auto-transcription (auto_transcribe=true) is not yet enabled and will return USER_INPUT_INCOMPLETE — when the user has not supplied lines, ask them for the cue text and timing instead of calling with auto_transcribe. If the user explicitly asks you to write, invent, improvise, or make up captions/subtitles, create a few short, generic cue lines yourself and call this tool with cues; do not ask a follow-up for exact wording in that case.",
        "parameters": {
          "type": "object",
          "properties": {
            "sourceVideoIndex": {
              "type": "number",
              "description": "Which video to subtitle. Default: -1 (most recent generated or uploaded video). Non-negative values are 0-based indices into prior generated video results. Negative values reference uploaded videos."
            },
            "cues": {
              "type": "array",
              "description": "Ordered subtitle cues. Each cue has startSeconds, endSeconds, and the line of text to display. Provide either `cues` or `srt`, not both. Aim for multiple short cues (1.5–4s each, ~1–8 words) rather than one long cue spanning the full clip.",
              "items": {
                "type": "object",
                "properties": {
                  "startSeconds": {
                    "type": "number",
                    "minimum": 0
                  },
                  "endSeconds": {
                    "type": "number",
                    "minimum": 0
                  },
                  "text": {
                    "type": "string"
                  }
                },
                "required": [
                  "startSeconds",
                  "endSeconds",
                  "text"
                ]
              }
            },
            "srt": {
              "type": "string",
              "description": "Full SRT (or VTT) document as a string, used in place of `cues`. Useful when the user pastes a subtitle file directly. Provide either `cues` or `srt`, not both."
            },
            "auto_transcribe": {
              "type": "boolean",
              "description": "Reserved for future speech-to-text support. Currently returns USER_INPUT_INCOMPLETE so the LLM can ask the user to supply cues. Do not set this — gather cue text from the user instead."
            },
            "style": {
              "type": "object",
              "description": "Optional styling overrides for the burned subtitles.",
              "properties": {
                "fontSizePct": {
                  "type": "number",
                  "minimum": 1,
                  "maximum": 30,
                  "description": "Font size as a percentage of the video height. Default 6."
                },
                "color": {
                  "type": "string",
                  "description": "Subtitle fill color. Default \"#FFFFFF\"."
                },
                "outlineColor": {
                  "type": "string",
                  "description": "Subtitle outline color. Default \"#000000\"."
                },
                "position": {
                  "type": "string",
                  "enum": [
                    "bottom",
                    "top",
                    "center"
                  ],
                  "description": "Vertical placement of the subtitle line. Default \"bottom\"."
                }
              }
            }
          }
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "image_to_3d",
        "description": "Reconstruct a textured 3D model with Pixal3D from an original front image, optionally with left, back and right views of the same subject. Returns a downloadable binary GLB model, not an image or video. Takes no prompt. Use original stills with consistent framing and height; never substitute screenshots or invented views.",
        "parameters": {
          "type": "object",
          "properties": {
            "sourceImageIndex": {
              "type": "integer",
              "description": "Front image: negative indices select uploads (-1 is the first), non-negative indices select generated images. Omit to use the latest image."
            },
            "source_image_url": {
              "type": "string",
              "description": "REST alternative to sourceImageIndex: retrievable front image URL."
            },
            "leftViewImageIndex": {
              "type": "integer",
              "description": "Subject own LEFT side facing the camera (subject faces screen-left). Any subset of orbit views enables multi-view reconstruction."
            },
            "backViewImageIndex": {
              "type": "integer",
              "description": "Same subject seen from behind."
            },
            "rightViewImageIndex": {
              "type": "integer",
              "description": "Subject own RIGHT side facing the camera (subject faces screen-right). Do not swap left and right."
            },
            "left_view_image_url": {
              "type": "string",
              "description": "REST alternative for the left view."
            },
            "back_view_image_url": {
              "type": "string",
              "description": "REST alternative for the back view."
            },
            "right_view_image_url": {
              "type": "string",
              "description": "REST alternative for the right view."
            },
            "meshTargetFaces": {
              "type": "integer",
              "minimum": 5000,
              "maximum": 700000,
              "description": "Triangle budget. Default 700000; choose a lower budget for real-time use."
            },
            "textureSize": {
              "type": "integer",
              "minimum": 1024,
              "maximum": 4096,
              "description": "Base-colour texture and UV atlas size. Default 4096."
            },
            "normalMapSize": {
              "type": "integer",
              "minimum": 512,
              "maximum": 2048,
              "description": "Normal map size. Default 2048."
            },
            "ambientOcclusionSize": {
              "type": "integer",
              "minimum": 256,
              "maximum": 1024,
              "description": "Ambient occlusion map size. Default 1024."
            },
            "shapeResolution": {
              "type": "integer",
              "enum": [
                1024,
                1536
              ],
              "description": "Shape resolution. Default 1024; 1536 costs more."
            }
          },
          "required": []
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "remove_background",
        "description": "Remove an image background with BiRefNet while preserving the original foreground. Returns a transparent PNG cutout by default, or a soft foreground mask. Uses the original source image without generative repainting; takes no prompt.",
        "parameters": {
          "type": "object",
          "properties": {
            "sourceImageIndex": {
              "type": "integer",
              "description": "Original source image: negative indices select uploads (-1 is the first), non-negative indices select generated images. Omit to use the latest image."
            },
            "source_image_url": {
              "type": "string",
              "description": "REST alternative to sourceImageIndex: retrievable source image URL."
            },
            "applyMask": {
              "type": "boolean",
              "description": "True (default) returns a transparent cutout; false returns the soft mask."
            }
          },
          "required": []
        }
      }
    },
    {
      "type": "function",
      "function": {
        "name": "segment_image",
        "description": "Select objects in an original image with SAM 3. Returns a PNG mask or transparent cutout without repainting the source. Coordinates refer to the original image, normalized 0–1. Supply an object description, a positive point, or a box. Text can combine with boxes; points can combine with at most one box, never with text.",
        "parameters": {
          "type": "object",
          "properties": {
            "sourceImageIndex": {
              "type": "integer",
              "description": "Original source: negative indices select uploads (-1 is the first), non-negative indices select generated images. Omit to use the latest image."
            },
            "source_image_url": {
              "type": "string",
              "description": "REST alternative: retrievable original image URL."
            },
            "text": {
              "type": "string",
              "minLength": 1,
              "maxLength": 240,
              "description": "Object or concept to select, such as the red suitcase."
            },
            "points": {
              "type": "array",
              "maxItems": 32,
              "items": {
                "type": "object",
                "properties": {
                  "x": {
                    "type": "number",
                    "minimum": 0,
                    "maximum": 1
                  },
                  "y": {
                    "type": "number",
                    "minimum": 0,
                    "maximum": 1
                  },
                  "label": {
                    "type": "string",
                    "enum": [
                      "positive",
                      "negative"
                    ]
                  }
                },
                "required": [
                  "x",
                  "y",
                  "label"
                ],
                "additionalProperties": false
              }
            },
            "boxes": {
              "type": "array",
              "maxItems": 16,
              "items": {
                "type": "object",
                "properties": {
                  "x0": {
                    "type": "number",
                    "minimum": 0,
                    "maximum": 1
                  },
                  "y0": {
                    "type": "number",
                    "minimum": 0,
                    "maximum": 1
                  },
                  "x1": {
                    "type": "number",
                    "minimum": 0,
                    "maximum": 1
                  },
                  "y1": {
                    "type": "number",
                    "minimum": 0,
                    "maximum": 1
                  },
                  "label": {
                    "type": "string",
                    "enum": [
                      "positive",
                      "negative"
                    ],
                    "description": "Optional inclusion/exclusion label; defaults to positive. Negative boxes require a text prompt."
                  }
                },
                "required": [
                  "x0",
                  "y0",
                  "x1",
                  "y1"
                ],
                "additionalProperties": false
              }
            },
            "maxInstances": {
              "type": "integer",
              "minimum": 1,
              "maximum": 16,
              "description": "Keep only the strongest N selections. Omit to keep every selection above threshold."
            },
            "threshold": {
              "type": "number",
              "minimum": 0,
              "maximum": 1,
              "description": "Minimum object confidence. Default 0.5."
            },
            "multimask": {
              "type": "boolean",
              "description": "Return multiple candidate masks for point selection. Default true for points."
            },
            "applyMask": {
              "type": "boolean",
              "description": "Return the original foreground with transparency instead of the binary mask. Default false."
            }
          },
          "required": []
        }
      }
    }
  ]
}
