---
name: prepare-video-assets
description: |
  Asset-preparation skill: resolves and generates every asset (image / audio / video) referenced by a Video DSL, persists a RenderPlan to the database, and returns a `job_id` for the subsequent `render_video` call.

  Use this skill as soon as the user mentions any of these intents:
  - Generate / prepare video assets, resolve assets, render-ready
  - "Make me a video about X" (the agent calls gen_script → prepare_video_assets → render_video)
  - Regenerate one asset (image / audio) for a specific scene

  Next step: after the user confirms the resolved assets, call `render_video` with the `job_id` returned by this skill.

  ⚠️ Stop-and-confirm gate: this skill runs only after the user has confirmed the script, and after it returns you must show the resolved assets and wait for the user's explicit confirmation. Never call `render_video` in the same turn.
triggers:
  - Generate / prepare video assets
  - Resolve missing assets in a DSL
  - User confirmed the script and the agent needs to prepare assets for review
  - User asks to regenerate a specific scene's image or narration
---

# Prepare Video Assets Skill

Phase 1 of the two-phase video pipeline. Takes a Video DSL plus a template binding, walks every `AssetRef` declared in the DSL, calls the matching atomic skills to fill in missing assets (`gen-image`, `gen-voice`, `gen-video`, `gen-digital-human`), persists the resulting RenderPlan to ab-api, and prints a `job_id` for the user-confirmation step.

> ⚠️ **Stop-and-confirm gate (mandatory, do not skip) — both sides of this skill.**
> **Before**: only run this skill after the user has confirmed the script produced by `gen_script`.
> **After**: show the resolved assets (images inline, audio as links) and **stop your turn** — wait for
> the user to explicitly confirm before rendering. **Never** call `render_video` in the same turn as
> `prepare_video_assets`. This is the last checkpoint where a bad image or a wrong TTS take can be
> fixed for the price of one asset instead of a whole re-render.
> See **"Pre-render user confirmation (Phase 2)"** below for the required summary format.

> **Next step**: after the user reviews the assets and confirms, call **`render_video`** with the `job_id` returned by this skill.

## Pipeline (this skill's part only)

```
DSL + template_id (or binding)
    ↓
[1] DSL Validator → schema check + normalize
    ↓
[2] Template Binder → produce RenderPlan
    ↓
[3] Asset Resolver → for every AssetRef, call atomic skills as needed
    ↓
[3b] Persist RenderPlan to DB → stdout "📦 render job jobId: N"
    ↓
(user reviews → confirms → render_video --job-id N)
```

## Atomic-skill dependencies

The asset resolver invokes the following skills based on `AssetRef` declarations:

| Asset type | Skill called | Notes |
|------------|--------------|-------|
| `image` + `source: gen-image` | `gen-image` | Text-to-image / image-to-image. |
| `audio` + `source: gen-voice` | `gen-voice` | Narration TTS. |
| `video` + `source: gen-video` | `gen-video` | AI-generated video clips. |
| `avatar` + `source: gen-digital-human` | `gen-digital-human` | Talking-head digital-human segments. |

## Authentication & environment

| Env var | Description | Default |
|---------|-------------|---------|
| `PRIV_TOKEN` | Tianyan token (needed for asset generation + DB persist). | (none) |
| `MM_API_BASE_URL` | Asset-generation API root (ab-agent proxy). | `https://api-agent.remixmate.ai/api` |
| `MM_BACKEND_API_URL` | Backend API root (RenderPlan persistence). | `https://api.remixmate.ai/api` |
| `ASSET_CACHE_DIR` | Asset cache directory. | `./.asset-cache/` |

## Canonical usage (inline DSL JSON, no temp files)

```bash
python3 <SkillDir>/scripts/prepare_video_assets.py \
  --dsl-json '<DSL_JSON_STRING>' \
  --template-id html-slide
```

- `<DSL_JSON_STRING>` is the JSON the agent received from `gen_script` — pass inline; **do not** write it to disk.
- `--save-job` defaults to on, persisting the RenderPlan and printing `📦 render job jobId: N`.
- Image and audio URLs are printed to stdout under `🔊 TTS audio:` and `🖼  Image assets:` sections.
- When `PRIV_TOKEN` / `MM_BACKEND_API_URL` are missing the wrapper auto-degrades to file mode and prints a notice.

## File-mode fallback (single-user / local debug)

```bash
python3 <SkillDir>/scripts/prepare_video_assets.py \
  --dsl my-video.dsl.json \
  --template-id html-slide
```

`--template-id` automatically calls template-registry internally; or pass `--binding my-video.binding.json` if you have a pre-generated binding.

## Image-URL handling (agent rules)

**DSL image assets must use HTTPS URLs.** During Remotion rendering, headless Chrome fetches the images directly over the network — no local download is needed.

| Situation | Correct action |
|-----------|----------------|
| Image in a DSL AssetRef | Use the raw HTTPS URL. |
| Showing an asset preview to the user | Use Markdown inline image syntax: `![name](https://...)`. **Do not** Read an HTTPS URL. |
| Need to understand the image (to draft a title / bullets) | `curl` it down locally and read it; the DSL still carries the HTTPS URL. |
| Avoid WebFetch on images | Internal-domain images are blocked by the WebFetch security policy. |

## Pre-render user confirmation (Phase 2)

This skill is the **first half** of a two-phase user-confirmation flow. After it returns, the agent **must** show the user the resolved assets, then **end its turn and wait** for confirmation before calling `render_video`.

> ⚠️ **Never** start the final render before the user confirms. An end-to-end request ("make me a video
> about X") authorizes the pipeline, not the skipping of its review steps — stopping here **is** the
> correct completion of this step, not an unfinished task.

Read the `job_id` from the stdout line `📦 render job jobId: N` and remember it; asset URLs come from the `🔊 TTS audio:` and `🖼  Image assets:` sections. Show images inline (`![name](https://...)`) and audio as links (`[🔊 listen](http://cdn.../x.mp3)`) so the user can actually review them. The recommended summary format:

```markdown
## Render confirmation

### Video overview
| Field | Value |
|------|-----|
| Title | ... |
| Actual duration | xxs (after audio adjustment) |
| Aspect ratio | 9:16 |
| Scenes | x |

### Scenes & assets
| # | Purpose | Duration | Narration excerpt | Image | Audio |
|---|---------|----------|--------------------|-------|-------|
| 1 | opening | 6.2s | ... | ![scene-01](https://...) | [🔊 listen](http://cdn.../narration-scene-01.mp3) |
| 2 | point   | 11.5s | ... | ![scene-02](https://...) | [🔊 listen](http://cdn.../narration-scene-02.mp3) |

### Asset status
- ✅ Images: x/x generated
- ✅ TTS: x/x generated
- ❌ Failed: list any failures

> Reply "continue" to start rendering. To regenerate a specific asset, tell me which one.
```

## Regeneration pattern

To regenerate a specific asset (e.g. the user is unhappy with scene 2's image):

1. Update that scene's prompt / params in the DSL.
2. Call this skill again with the updated `dsl_json` (or the minimal narration-override JSON when the gen_script skeleton is cached in the session).
3. A new `job_id` is issued. Show the new assets and ask for confirmation again.
4. After confirmation, call `render_video` with the **new** `job_id`.

## Test mode: skip asset generation (`--stub-image-url` / `--stub-video-url`)

During dev / debug the user may want to exercise the pipeline without burning gen-image / gen-video quota. With a stub URL set, every `image+source=gen-image` AssetRef (or `video+source=gen-video`) is short-circuited at the resolver to that URL — no generation request is sent. TTS is unaffected and always runs for real.

### Usage rules

- The URL is passed **verbatim**; do not rewrite it.
- If the user expressed the intent but did not supply a URL, the agent must ask which fallback URL to use — never invent one.
- The upstream `gen_script.py --stub-image-url / --stub-video-url` is the source-of-truth solution; this skill's flag is the render-layer safety net. Both layers can be enabled at the same time.
- If the user does not re-state test mode in a later turn, **do not** carry the previous stub URL forward.

### Command examples

```bash
python3 <SkillDir>/scripts/prepare_video_assets.py \
  --dsl-json '<DSL_JSON_STRING>' \
  --template-id picture-book-en \
  --stub-image-url "https://cdn.example.com/placeholder.jpg"
```

Env vars `STUB_IMAGE_URL` / `STUB_VIDEO_URL` also work — their priority is lower than the CLI flag.

## Common CLI flags (inherited from render_video.py via the wrapper)

| Flag | Description | Default |
|------|-------------|---------|
| `--dsl` | Input DSL file path. | — |
| `--dsl-json` | DSL JSON as an inline string (preferred). | — |
| `--template-id` | Template id (auto-bind). | — |
| `--binding` | TemplateBinding file path. | — |
| `--binding-json` | TemplateBinding JSON as an inline string. | — |
| `--save-job` | Persist RenderPlan to the database; pass `--no-save-job` to disable. | **on by default** |
| `--asset-cache-dir` | Asset cache directory. | `.asset-cache/` |
| `--max-asset-retries` | Max retries per asset. | `3` |
| `--asset-timeout` | Per-asset generation timeout (seconds). | `300` |
| `--private-token` | Tianyan token. | env var |
| `--stub-image-url` | Test mode: short-circuit image+source=gen-image assets to this URL. | — |
| `--stub-video-url` | Test mode: short-circuit video+source=gen-video assets to this URL. | — |

## Asset-resolution strategy

The Asset Resolver handles each `AssetRef` in this order:

1. **status = generated / approved**: asset is ready — use the URL as-is.
2. **status = planned / missing**: call the matching atomic skill to generate it.
3. **Generation failed**: retry up to `maxRetries`; final failures are recorded in the RenderPlan errors.
4. **Parallel generation**: assets of the same type are generated in parallel; different types are sequenced by dependency.
5. **Cache reuse**: assets with the same payload are checked against `asset-cache-dir` to avoid duplicate generation.

## Credits

Every run charges credits. The CLI prints a footer on stdout when it does:

```
💳 Charged 31 credits · balance 1,240
```

Relay it to the user whenever it appears — it is the only signal they get about what a
generation cost, and the balance is the only warning before a run fails with
`insufficient_credits`. Do not drop it from your summary.

## Error handling

- **DSL validation failed**: pre-validate with `gen-script --validate`.
- **Template not found**: confirm the templateId is in the registry.
- **Asset generation failed**: inspect the RenderPlan `errors` field and the atomic skill logs.
- **Stale skeleton**: if the agent submits gen_script's raw skeleton verbatim as `dsl_json`, the agent-layer precall rejects with a skeleton-placeholder error. Fill in real narration for every scene.

## See also

- **`render_video`** — the Phase 3 skill that consumes this skill's `job_id` and produces the final video.
- **`template-registry`** — Lists available templates. Called internally by this skill when `--template-id` is provided.
