---
name: gsk-video-generation
version: 1.0.0
description: Generate videos using AI models. Supports text-to-video and image-to-video
  generation.
metadata:
  category: general
  requires:
    bins:
    - gsk
  cliHelp: gsk video --help
---

# gsk-video-generation

**PREREQUISITE:** Read `../gsk-shared/SKILL.md` for auth, global flags, and security rules.

Generate videos using AI models. Supports text-to-video and image-to-video generation.

## Usage

```bash
gsk video [options]
```

**Aliases:** `video`

## Flags

| Flag | Required | Description |
|------|----------|-------------|
| `<query>` (positional) | No | Detailed, self-contained description of the video to generate, written in English. This text is sent to the video model verbatim as the generation prompt, so put all the visual detail here — but NEVER output settings: text like '16:9' or '4 seconds' in the prompt may be rendered literally into the scene. Pass aspect ratio via the `aspect_ratio` parameter and length via `duration`. duration is recommended to be 5 ~ 10 seconds. For image-to-image transitions (user provides start and end frames) e.g., 'A timelapse of a flower blooming in a garden'. Omit when passing the prompt by `query_file` instead. (string) |
| `--query_file` | No | Optional. Repo-relative path to a UTF-8 text file containing the prompt, used verbatim in place of `query` (lets you pass a long prompt by path instead of inlining it). Requires `repo_id`. When set, `query` may be omitted. (string) |
| `--file_name` | No | The name of the video to generate.  (string, null) |
| `-m`, `--model` | Yes | The model to use for video generation. kling/v3: Latest Kling V3. Supports native audio, but generates a SILENT video unless audio_enable=true. image_urls are frames, not references: 1 image = first frame (i2v); 2 images = first + last frame (the 2nd is the END frame). No reference mode (use kling/o3 with reference_mode for reference-driven generation). Pro/Standard/Turbo tiers (set via the tier parameter; turbo is the fastest arm but has no audio and no first/last-frame control). Supported ratios: 16:9, 9:16, 1:1. Duration: 3-15s. kling/o3: Kling O3. REQUIRES at least one input image or a video_url (no text-to-video). image_urls are frames (1 = first, 2 = first + last) unless reference_mode=true: style/character references, max 4 total including elements. video_url switches to video-to-video (video_mode 'reference' = motion guide, 'edit' = restyle in footage). Native audio via audio_enable. Pro/Standard tiers. Supported ratios: 16:9, 9:16, 1:1. Duration: 3-15s. gemini/veo3.1: Gemini Veo 3.1 - Latest version with enhanced quality and features. Supports text-to-video and image-to-video generation with improved quality. image_urls are frames unless reference_mode: 1 image = first frame, exactly 2 = first/last frame transition, 3 (or reference_mode=true) = style/character references (max 3, 8s only). video_url (MP4 source) = video EXTENSION mode: continues from the last frame, 7s, up to 1080p. Duration: 4s/6s/8s. Supported ratios: 16:9, 9:16. Supports tier (standard \| fast). Resolution: 720p, 1080p, 4k. minimax/h3: MiniMax H3 2K video model. Supports text-to-video, image-to-video with first/last frame control, and reference-to-video (up to 9 reference images, 3 reference videos, 3 reference audio clips — refer to them in the prompt as 'Image 1', 'Video 1', 'Audio 1', not @-style). Keyframes and references are mutually exclusive. Supported ratios: adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16. Duration: 5-15s. Resolution: 768p or 2K (default). minimax/h3-max: MiniMax H3 Max, a distinct post-trained model with faster generation, stronger prompt adherence, and improved aesthetics. Supports text, first/last frames, and image/video/audio references. tier=turbo selects the faster, half-price text/image arm; reference inputs use standard. Duration: 5-15s. Resolution: 480p or 768p. wan/v2.7: Wan v2.7 with enhanced motion smoothness and scene fidelity. Supports text-to-video, image-to-video, reference-to-video, and edit-video. Ratios: 16:9, 9:16, 1:1. Duration: 5s. Resolution: 480p, 720p vidu/q3: Vidu Q3 model with enhanced quality and audio generation. Supports text-to-video, image-to-video, and reference-to-video (reference_mode with 1-4 image_urls keeps subjects/scenes consistent; no turbo tier in reference mode). tier=turbo runs the faster turbo arm at half the price. Supported ratios: 16:9, 9:16, 4:3, 3:4, 1:1. Duration: 1-16s. Resolution: 720p, 1080prunway/gen4_turbo: A model for generate video with high quality, fast.Supported ratios: 5:3, 3:5. Only support i2v. Duration: 5s, 10spixverse/v6: PixVerse V6 latest model with lifelike motion, richer skin detail, real emotions. Full cinematic control including choreography and camera. Supports text-to-video, image-to-video, transition, and video extend generation. VFX, time-lapse, transformation scenes, product demos, 360° views, multi-shot storytelling. Extend: pass video_url to extend an existing video. Supported ratios: 16:9, 9:16, 4:3, 1:1, 3:4. Duration: 5s, 8s.  pixverse/c1: PixVerse C1 with native audio (audio_enable=true), text-to-video, image-to-video, transitions (2 images = first/last frame) and reference-to-video (reference_mode with 1-7 image_urls; address them in the prompt as @image1..@image7). No video extend. Supported ratios: 16:9, 9:16, 4:3, 3:4, 1:1, 3:2, 2:3, 21:9. Duration: 1-15s. Resolution: 360p, 540p, 720p, 1080p.  fal-ai/bytedance/seedance-2.0: Bytedance Seedance 2.0 model for highest quality video generation with native audio and lip-sync. Supports text-to-video and image-to-video with first/last frame control. Supported ratios: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16. Duration: 4-15s. Resolution: 480p, 720p, or 1080p on the standard tier (~2.25x cost; fast/mini cap at 720p). tier picks the variant: standard, fast, or mini (cheapest, ~half the standard price). Also supports reference-to-video with up to 9 images, 3 videos, or 3 audio refs via settings.fal-ai/bytedance-upscaler/upscale/video: ByteDance Video Upscaler for enhancing video quality. Upscales videos to higher resolutions (2k) Requires video_url parameter. Does NOT generate new videos. Use when user wants to improve quality of existing videos.xai/grok-imagine-video: xAI Grok Imagine Video v1.5 for high-quality video generation. Supports text-to-video and image-to-video generation. 2+ image_urls (or reference_mode=true) = style/character references, up to 7 (reference mode caps at 720p). video_url (MP4, 2-15s source) = video EXTENSION mode: continues from the last frame, 2-10s extension, 720p. 480p/720p/1080p output. Supported ratios: 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, 9:16. Duration: 1-15s. alibaba/happy-horse: Alibaba Happy Horse video model. High-quality generation from text, or image-to-video from a single first-frame image (one image_url without reference_mode). Supported ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21, 5:4, 4:5. Duration: 3-15s. Resolution: 720p, 1080p.alibaba/happy-horse/reference-to-video: Alibaba Happy Horse reference-to-video. Generate video from 1-9 reference images; address subjects in the prompt as character1..character9 (order matches image_urls). Requires image_urls. Supported ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21, 5:4, 4:5. Duration: 3-15s. Resolution: 720p, 1080p.alibaba/happy-horse/video-edit: Alibaba Happy Horse video-edit. Edit a source video using a text prompt; optional 1-5 reference images addressed as @Image1..@Image5. Requires video_url (MP4/MOV, 3-60s). Output preserves source aspect ratio, capped at 15s. Resolution: 720p, 1080p.gemini/omni-flash: Google's multimodal video model with native audio. Supports text-to-video, image-to-video (1 image as first frame), reference-to-video (set reference_mode with image_urls to use images as subject/style references — even a single image; 2+ images are references by default), and video editing (pass video_url to edit an existing clip — restyle, restage, add objects). Bind image roles in the prompt with <FIRST_FRAME> / <IMAGE_REF_0>. Ratios: 16:9, 9:16. Duration: 3-10s. Resolution: 720p. bfl/flux-3-preview-high: BFL FLUX 3 video model with native audio. Supports text-to-video and 1-10 ordered keyframes, or continuation from one source video selected with input_mode=start_video. Ratios: auto, 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, 9:16. Duration: 5-20s. Resolution: 720p, 1080p. Only one input family may be used per request.fal-ai/bytedance/seedance-2.5: Bytedance Seedance 2.5 for native 30-second single-shot video with synchronized audio. Supports text-to-video and image-to-video with first/last frame control. No tier variants. Supported ratios: auto, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16. Duration: 4-30s, or 0 = auto (needs video refs). Resolution: 480p, 720p, or 1080p (~2.46x cost). Also supports reference-to-video (up to 30 images, 10 videos, 10 audio refs; no image required), video EDITING and EXTENSION — source in video_urls, the prompt describes the change or continuation.wan/v3.0: Alibaba Wan 3.0 — next-gen open-weight Wan with native audio and single-pass clips up to 30s. Supports text-to-video, image-to-video with first/last frame control, and reference-to-video (reference_mode with up to 10 reference images, 5 videos and 5 audio clips, each ref modality <=15s combined — address them positionally in the prompt as 'Image 1', 'Video 1', 'Audio 1'). tier=fast runs the accelerated Prime arm (~1.4x cost). Supported ratios: adaptive, 16:9, 4:3, 1:1, 3:4, 9:16. Duration: 2-30s. Resolution: 480p, 720p, 1080p. Duration ranges above are hard caps; out-of-range values are clamped to the nearest supported value. (string, one of: kling/v3, kling/o3, gemini/veo3.1, minimax/h3, minimax/h3-max, wan/v2.7, vidu/q3, runway/gen4_turbo, pixverse/v6, pixverse/c1, fal-ai/bytedance/seedance-2.0, fal-ai/bytedance-upscaler/upscale/video, xai/grok-imagine-video, alibaba/happy-horse, alibaba/happy-horse/reference-to-video, alibaba/happy-horse/video-edit, gemini/omni-flash, bfl/flux-3-preview-high, fal-ai/bytedance/seedance-2.5, wan/v3.0) |
| `-i`, `--image_urls` | No | The URLs of the images to use as reference key frames for the video generation. For single image models, provide 1 image. For multi-image models (like kling/v1.6/pro/elements), provide 1-4 images in sequence order. For start-end models (like vidu, pixverse), provide exactly 2 images (start and end frames). (default is [], if the task is based on one or more reference images, it is required) IMPORTANT: without reference_mode, the first image becomes the video's literal opening frame — identity markers visible in it (flags, jerseys, text, logos, faces) persist in the output and the text prompt CANNOT override them. When generating a batch of subject-specific clips (per-country / per-person / per-product), each clip needs an image matching its own subject — never reuse another subject's image. (array) |
| `-r`, `--aspect_ratio` | No | The aspect ratio of the video to generate. For image-to-video (i2v) generation, this should match the aspect ratio of your input image(s). Ratios unsupported by the chosen model fall back to its default (5:4 and 4:5 are only supported by alibaba/happy-horse; 3:2 and 2:3 by pixverse/c1 and grok-imagine-video). Use 'auto' (or omit this parameter) for FLUX 3. Omit it for MiniMax H3 native adaptive framing. (string, one of: auto, 16:9, 9:16, 1:1, 4:3, 3:4, 2:1, 3:2, 2:3, 21:9, 9:21, 5:4, 4:5) |
| `-d`, `--duration` | No | The duration in seconds. 0 = auto (Seedance 2.5 with video_urls; required for edit prompts — output keeps the source length; extensions: pass the new segment's length). (number, default: `5`) |
| `-a`, `--audio_url` | No | Audio URL for audio integration in video generation. - Optional for Wan v2.5 model (custom audio: music or speech, WAV/MP3, 3-10s, up to 15MB) - Required for OmniHuman model (audio-driven animation) Leave empty if no audio is needed. (string, default: ``) |
| `--video_url` | No | Source video URL for video extension or video-to-video. On Grok Imagine Video and Gemini Veo 3.1 it switches to EXTENSION mode; on Kling O3 it switches to video-to-video (see video_mode). Supports both web URLs and AI Drive paths (starting with '/' or 'aidrive://') (string, default: ``) |
| `--video_mode` | No | How video_url is used. Kling O3: 'reference' (default — motion/camera guide) or 'edit' (restyle inside the source footage); no extend. Gemini Veo 3.1 / Grok Imagine Video: 'extend' only. Ignored without video_url. (string, one of: reference, edit, extend) |
| `--keep_audio` | No | Kling O3 video-to-video only: keep the source video's audio track in the output. (boolean) |
| `--input_mode` | No | FLUX 3 Video only: set start_video to continue video_url from its final frames. Omit for all other models and for text/image input. (string, one of: start_video) |
| `--video_urls` | No | Reference video URLs for reference-to-video models. For Wan v2.7 and Seedance 2.0 reference-to-video: provide 1-3 reference videos and refer to them in the prompt as @Video1, @Video2, @Video3. Seedance 2.0 also needs reference_mode plus at least one image_url; combined video duration 2-15s. For Seedance 2.5: up to 10 videos (≤30s combined), no image required; also the edit/extension source (edit: duration=0, output keeps the source length; extension: duration = the new segment's length). For MiniMax H3: up to 3 motion reference videos (each 2-15s, ≤15s combined), referred to in the prompt as 'Video 1' etc. (array, default: `[]`) |
| `--audio_urls` | No | Reference audio URLs for reference-to-video models. For Seedance 2.0: provide up to 3 audio clips (MP3/WAV) and refer to them in the prompt as @Audio1, @Audio2, @Audio3 (e.g. to drive lip-sync or match a voice). Combined duration must not exceed 15s. Requires reference_mode plus at least one image_url (image_urls is what activates the seedance-2.0/ref route). For Seedance 2.5: up to 10 clips (≤30s combined), no image required. For MiniMax H3: up to 3 audio references (each 2-15s, ≤15s combined; 'Audio 1' prompt style), must be paired with at least one reference image or video. Distinct from audio_url, the single-clip Wan v2.5 / OmniHuman field. (array, default: `[]`) |
| `--repo_id` | No | Optional. Second Brain repo id (provided in the agent's system prompt). Required only when image_urls / audio_url / video_url / video_urls contain relative paths from a Second Brain project (e.g. 'assets/clip.mp4'). Absolute http(s) URLs and data: URLs do not need this parameter. (string) |
| `--audio_enable` | No | Set true when the user wants audio, voiceover, or narration. kling/v3 outputs a SILENT video unless true (audio costs more per second). Omit to keep each model's default (kling: off; seedance/veo: on). (boolean) |
| `--tier` | No | Model tier (options are model-specific — check get_model_info tier_options). Seedance 2.0: standard \| fast (quicker/cheaper, 720p cap) \| mini (cheapest, ~half the standard price, 720p cap). Veo 3 / Veo 3.1 family: standard \| fast (quicker, lower cost). Kling V3: pro \| standard \| turbo (fastest, no audio/frame control). Kling O3/Motion Control: pro \| standard. Vidu Q3: standard \| turbo (half price). Wan 3.0: standard \| fast (the accelerated Prime arm — quicker but ~1.4x the price). MiniMax H3 Max: standard \| turbo (faster, half price; references use standard). MiniMax H3 ignores tier. Models without tiers ignore it. (string, one of: standard, fast, mini, pro, turbo, max) |
| `--video_size` | No | Output resolution. 'auto' uses the model default (usually 720p). MiniMax H3 defaults to 2K with a cheaper 768p tier; MiniMax H3 Max uses 480p or 768p. Values a model does not publish fall back to its default. Higher tiers bill more per second — see each model's Resolution line. Default: auto. (string, one of: auto, 480p, 720p, 768p, 1080p, 2K, 4k, default: `auto`) |
| `--reference_mode` | No | Enable reference-to-video mode: provided images guide character/style rather than being used as the first frame. Requires image_urls. For Seedance 2.0: routes to seedance-2.0/ref (combinable with the fast/mini tiers, supports up to 9 reference images plus optional reference video/audio). For Veo 3.1 base model: upgrades to veo3.1/reference-to-video. For MiniMax H3: treats images as subject/style references instead of keyframes (any video_urls/audio_urls also force reference mode). For Vidu Q3: routes to the reference-to-video mix endpoint (1-4 reference images). For PixVerse C1: 1-7 subject references (@image1..@image7 in the prompt). For Grok Imagine Video: 1-7 references (2+ images switch automatically; caps output at 720p). For Kling O3: 1-4 references incl. elements (3+ images switch automatically). Ignored on other models. Default: false. (boolean, default: `False`) |

## Local File Support

Parameters that accept URLs (`--image_urls`, `--audio_url`, `--video_url`, `--video_urls`, `--audio_urls`) also accept local file paths. The CLI automatically uploads local files before sending to the API.

## Output File

Use `-o <path>` / `--output-file <path>` to download the generated result directly to a local file.

## See Also

- [gsk-shared](../gsk-shared/SKILL.md) — Authentication and global flags
