---
name: reference-grounded-narrative-sequence-skill
description: |
  Generate continuity-safe narrative image sequences and clean standalone keyframes from one or more reference images plus a supplied style directive.
  This skill is designed for visual storytelling use cases such as cinematic narrative boards, next-scene continuation,
  character emotional continuation, environment/detail progression, and illustrated narrative sequences.

  Use this skill when users want to:
  - Generate a narrative sequence from 1 or many reference images
  - Combine character, environment, prop, and anchor-shot references into one coherent visual sequence
  - Keep subject identity, environment, props, wardrobe, and style consistent across panels
  - Default to a 6-panel labeled sequence
  - Extract or recreate any panel as a clean standalone image without labels, borders, or text overlays
platforms:
  - openclaw
  - claude-code
  - cursor
  - codex
  - gemini-cli
---

# Reference-Grounded Narrative Sequence Skill

You are an expert at turning one or more reference images into a continuity-safe visual narrative sequence while preserving identity, environment, props, and project style.

## Core Purpose

This skill is for **narrative visual storytelling**.

Do not use this skill for plain image-editing requests, even when multiple reference images are attached. If the user wants a direct edit, retouch, replacement, inpainting, outpainting, or the same explicit change applied across multiple images, route that work to `image-generation` instead.

It can adapt to:
- cinematic narrative boards
- next-scene continuation
- character emotional continuation
- environment/detail progression
- comic or anime narrative sequences when requested
- illustrated story sequences

## Supported Inputs

The orchestrator or user may provide:

- **One reference image** or **multiple reference images**
- A **description for each reference image**
- A **style directive** from:
  - the user's conversation
  - the orchestrator agent
  - a project production bible / visual bible
- An optional **adaptation profile**
- An optional **panel count**

### Reference Roles

Each reference image may represent one of these roles:
- **anchor_shot**: main composition or scene seed
- **character**: face, hair, costume, body, silhouette, identity
- **location**: environment, architecture, terrain, room layout
- **prop**: important objects that must remain consistent
- **style_mood**: palette, lighting, rendering, mood, texture
- **detail_cue**: specific narrative detail such as reflection, shadow, blood trace, ripple, broken glass, footprint, note, or gesture

If the user does not explicitly label the roles, infer them from the descriptions.

## Default Behavior

- Default to **6 panels** unless the user or orchestrator requests another count
- Default adaptation profile: **cinematic_narrative**
- Stay flexible enough to switch profiles if the request clearly implies another narrative structure

## Adaptation Profiles

### 1. cinematic_narrative (default)
Use for filmic storytelling and storyboard-like progression.

Example grammar only — adapt as needed:
1. establishing wide
2. medium wide subject placement
3. medium
4. medium close-up or close-up
5. insert detail / reaction / over-shoulder
6. final hero frame / payoff

Important: this is only a default example, not a rigid rule. Narrative emphasis may override pure camera grammar.

### 2. next_scene_continuation
Use when the request implies: “what happens right after this image?”

Typical progression:
1. anchor continuation
2. movement beat
3. reveal
4. consequence
5. escalation
6. next-scene payoff

### 3. character_emotional_continuation
Use when the focus is subtle emotional or performance progression.

Typical progression:
1. neutral state
2. attention shift
3. reaction
4. internal beat
5. emotional peak
6. resolved aftermath

### 4. environment_detail_progression
Use when the scene, space, or clues matter more than facial performance.

Typical progression:
1. establish space
2. zone focus
3. object clue
4. sensory detail
5. hidden meaning or reveal
6. integrated final frame

### 5. illustrated_sequence
Use when the style directive or project language implies comic, anime, manga, painted, or illustrated progression.

Do not force comic grammar unless requested by style or project context.

## Style Directive Handling

A style directive is a **first-class input**.

Sources may include:
1. production bible / visual bible
2. explicit user style request
3. orchestrator summary
4. inferred style from references only if none is provided

### Style Rule

If a style directive is provided by the orchestrator or production bible, treat it as a **global visual constraint** for narrative-sequence outputs handled by this skill.
Maintain that style consistently across all panels and all extracted standalone frames unless the user explicitly requests a change.

### What style consistency should lock
- rendering medium
- palette
- contrast
- lighting family
- lens / film feel
- mood tone
- texture treatment
- composition discipline
- linework / realism / painterly / cel-shaded treatment

## Continuity Locks

For every sequence, maintain these locks unless the request says otherwise:
- same subject identity
- same location/environment logic
- same props and their role in the scene
- same wardrobe / silhouette / hair consistency
- same lighting family and grading
- same overall style language

If multiple references conflict, prefer:
1. explicit orchestrator instruction
2. production bible
3. anchor shot
4. character / location / prop references

## Orchestrator Input Contract

The orchestrator should pass normalized inputs whenever possible.

### Expected Inputs
- `referenceImages[]`: one or more reference images
- `referenceImages[].description`: a short description for each image
- `referenceImages[].role` (optional): one of `anchor_shot`, `character`, `location`, `prop`, `style_mood`, `detail_cue`
- `styleDirective`: the resolved visual style to keep consistent across outputs
- `adaptationProfile` (optional): such as `cinematic_narrative`, `next_scene_continuation`, `character_emotional_continuation`, `environment_detail_progression`, `illustrated_sequence`
- `panelCount` (optional): default to `6`
- `continuityPriorities` (optional): identity, environment, props, style, or all
- `outputMode` (optional): `collage`, `clean_extract`, or both
- `targetPanel` (optional): used for follow-up extraction requests such as panel 4

### Input Precedence Rules
If multiple sources conflict for a narrative-sequence task, resolve in this order:
1. explicit orchestrator instruction
2. explicit user request in the current conversation
3. production bible / visual bible
4. anchor shot reference
5. other reference images and descriptions
6. inferred defaults

### Orchestrator Responsibilities
When possible, the orchestrator should:
- attach a description to every image
- assign roles if already known
- pass a clean style summary instead of raw long-form bible text
- indicate whether the user wants a multi-panel sequence or a clean standalone extraction
- specify if a certain panel count or adaptation profile is required

### Skill Guarantees
This skill should reliably:
- default to a 6-panel sequence unless overridden
- preserve continuity across identity, environment, props, wardrobe, and style
- produce labeled collage-ready panel prompts
- preserve a clean standalone prompt for every panel
- support follow-up extraction or recreation of any chosen panel

## Tool Usage

If an image-generation tool is available in the runtime, use it to generate the requested sequence or extracted frame.

- Prefer an available MCP image-generation tool when present
- Example: a Kalasetu MCP `generate image` tool
- If multiple image-generation tools are available, prefer the one selected by the orchestrator or project workflow
- If no image-generation tool is available, still produce the best structured prompts and panel plan for downstream generation

## Workflow

### Step 1: Parse inputs
Identify:
- panel count
- adaptation profile
- style directive
- reference roles
- continuity priorities
- whether the user wants a collage sequence or a clean frame extraction

### Step 2: Clarify only if needed
Ask questions only if critical information is missing, such as:
- Which reference is the main character?
- Should the anchor shot dominate composition?
- Which props must remain present in all frames?
- Is the goal cinematic narrative, next-scene continuation, or emotional continuation?

Do not over-question if the orchestrator already supplied enough context.

### Step 3: Build global locks
Create internal guidance for:
- style_lock
- identity_lock
- environment_lock
- prop_lock
- narrative_progression

### Step 4: Plan sequence panels
By default plan 6 panels.

Each panel should have:
- index
- visible label for collage mode
- panel purpose
- framing or narrative emphasis
- panel prompt
- clean extraction prompt

### Step 5: Generate collage-mode instructions
For collage or storyboard outputs:
- add a visible panel number
- add a concise panel label
- examples:
  - `1: establishing wide`
  - `2: medium wide`
  - `4: reflection reveal`
  - `5: reaction beat`
- keep labels clean and small
- labels must not dominate the image

### Step 6: Generate clean extraction instructions
If the user asks to extract or recreate a panel as a standalone image:
- use the selected panel’s clean prompt
- preserve the exact narrative moment and continuity
- remove all text, numbering, borders, collage framing, and layout artifacts
- the output must be a clean reference-quality image

Supported follow-up requests include:
- extract panel 4
- recreate the 4th frame clean
- give me the close-up as a standalone image
- remove labels and borders from panel 3

### Step 7: Keep outputs reusable
For every panel, preserve two prompt variants internally:
1. **panel/collage prompt**
2. **clean standalone prompt**

This is mandatory so individual frames can be recreated reliably later.

## Output Format

When planning a sequence, structure the output internally like this:

```json
{
  "panelCount": 6,
  "adaptationProfile": "cinematic_narrative",
  "styleDirective": "...",
  "referenceInputs": [
    { "role": "anchor_shot", "description": "..." },
    { "role": "character", "description": "..." },
    { "role": "location", "description": "..." }
  ],
  "globalLocks": {
    "style": "...",
    "identity": "...",
    "environment": "...",
    "props": "..."
  },
  "panels": [
    {
      "index": 1,
      "label": "1: establishing wide",
      "purpose": "scene setup",
      "panelPrompt": "...",
      "cleanPrompt": "..."
    }
  ]
}
```

## What This Skill Can Do

- create continuity-safe narrative sequences from one or more reference images
- combine character, environment, prop, anchor-shot, and detail-cue references
- apply a supplied style directive consistently across the whole sequence
- default to a 6-panel narrative output
- default to cinematic narrative profile while staying flexible
- label each panel with number + concise panel type in collage mode
- recreate any selected panel as a clean standalone image
- preserve consistency across both multi-panel and extracted outputs

## Important Rules

- Do not force storyboard, comic, or film grammar unless the project context calls for it
- Use cinematic narrative as the default profile, not a rigid template
- Shot variation is only one sequencing strategy; other valid strategies include emotional progression, scene evolution, environment storytelling, and clue reveals
- Clean extraction outputs must contain **no text overlays, no numbering, no borders, and no collage artifacts**
- Keep responses aligned with the user’s language, but prompts for image generation may remain in English when useful for downstream generation
