---
name: image-generation
description: Generate or edit images with iterative QC refinement loop (max 3 iterations)
allowed-tools: [kalaasetu]
---

# Iterative Image Generation & QC Skill

## Purpose
This skill generates new keyframes and performs direct image edits with an iterative QC loop. Classify the task correctly first, then load the right reference guide before prompt construction.

## Reference Sub-Guides
This skill is split into separate reference files so mode-specific policy stays explicit.

**You MUST read the relevant reference file(s) before building the prompt:**

- `references/task-classification-and-routing.md`
- `references/style-precedence-policy.md`
- `references/scene-and-continuity-generation.md` for `scene_generation` or `continuity_generation`
- `references/image-editing-single-reference.md` for single-image direct edits
- `references/image-editing-multi-reference.md` for multi-image direct edits
- `references/qc-and-retry-playbook.md` before retrying failed outputs

## Context & Inputs
**CRITICAL**: You must EXTRACT and RETAIN the following from the user's request (or Director):
1.  `prompt` (string): The user's requested image instruction.
2.  `reference_images` (array): The list of reference or source image objects.
3.  `global_style` (string): The visual style defined in the production bible. Use this for `scene_generation` or `continuity_generation`. For `image_editing`, do not inject it unless the user explicitly asks to restyle or match a bible-defined look.
4.  `output_path` (string): The absolute path where the final image MUST be saved.
5.  `task_mode` (derived): Classify the request as `scene_generation`, `continuity_generation`, or `image_editing` before building the prompt.
6.  `style_policy` (derived): Resolve as `enforce_bible_style`, `preserve_source_style`, or `user_requested_restyle` before prompt construction.
7.  `edit_scope` (derived, optional): The exact region, object, or user-marked area that must change for editing tasks.

## Routing Summary
- Use `image_generation` for new keyframes and for direct edits to existing image files.
- Multiple references alone do **not** make the task a narrative-sequence request.
- Route to `reference-grounded-narrative-sequence-skill` only when the requested output is a multi-panel narrative sequence, storyboard-style collage, next-scene continuation sequence, or clean extraction from such a sequence.

## Instructions
**Procedure for Image Generation / Editing:**

1.  **Direct Execution**:
    *   **Do NOT** create a sub-agent or delegate this task.
    *   Execute the tool calls yourself in the current context.

2.  **Execution Loop**:
    *   **Step 0: Classify and Load Guides**
        *   Determine whether the task is `scene_generation`, `continuity_generation`, or `image_editing`.
        *   Resolve `style_policy` before building the prompt.
        *   Read `references/task-classification-and-routing.md` and `references/style-precedence-policy.md`.
        *   Read the mode-specific guide for the current task.

    *   **Step 1: Analyze Inputs**
        *   First, use the `read` tool to view the actual files in `reference_images`. Analyze them to understand composition, subject details, spatial arrangement, continuity needs, and any marked edit regions.
        *   If the task is `scene_generation` or `continuity_generation`, read `project/production_bible.json` to extract `global_style` and any required aspect ratio guidance.
        *   If the task is `image_editing`, infer the style primarily from the provided source image(s) and the user's explicit instruction. Do **not** append production-bible style by default.
        *   If the user marked a red box, highlight, mask, or bounding area, treat it as the authoritative `edit_scope`.

    *   **Step 2: Build the Prompt**
        *   For `scene_generation` or `continuity_generation`, combine the user's request with the required global styling and only the continuity details needed to preserve approved character, location, prop, wardrobe, and lighting consistency.
        *   For `image_editing`, keep the prompt minimal and surgical. Describe only the requested change, the target scope, and the preservation constraints needed to stop unintended edits.
        *   For multi-image edits, apply the same explicit requested change consistently across all relevant source images unless the user specifies image-specific differences.
        *   Only use production-bible style in edit mode if `style_policy=user_requested_restyle`.

    *   **Step 3: Generate**
        *   Call `kalaasetu_generateImage` with the crafted `prompt` and the `reference_images`.

    *   **Step 4: Verify**
        *   Read the output image file using your multimodal capabilities.
        *   Check against specific criteria:
            *   Does it match the requested `prompt`?
            *   Does it align with `reference_images` style/content?
            *   Visual quality (artifacts, anatomy, lighting).
            *   For `image_editing`: did it change only the requested region/object/aspect while preserving the rest?

    *   **Step 5: Decide**
        *   **PASS**: If visual quality and accuracy are good -> Return the file path (STOP).
        *   **FAIL**:
            1.  **Check Limit**: If this was the **3rd attempt**, return the current result with a warning (STOP).
            2.  **Prepare Feedback**:
                *   Identify specific issues (e.g., "Lighting is flat", "Missing scar", "Edited outside the highlighted region").
                *   For `image_editing`, explicitly note what should remain unchanged and what exact region/object still needs correction.
                *   If the issue is drift, unrequested restyling, or broader-than-requested change, remove unnecessary style language and tighten the edit scope before retrying.
                *   Append the *Failed Image* to `reference_images` with description: "NEGATIVE REFERENCE: Previous failed attempt, avoid these errors: [specific issues]".
                *   Read `references/qc-and-retry-playbook.md` and refine the `prompt` only as much as needed to fix the issue.
            3.  **Loop**: Go back to **Step 1** with the updated inputs.

## Rules
1.  **File Naming & Storage**:
    *   **NEVER** store images at the root level.
    *   **Always** store the final image at the exact absolute `output_path` provided by the Director. DO NOT invent paths or guess where to save.
2.  **Reference Retention**: Never drop index 0-N of the original `reference_images`. Only APPEND new negative references.
3.  **Multimodal Truth**: Trust your eyes (what you see in the image file) over the tool's text output.
4.  **No Hallucinations**: Do not invent file paths. Use the actual paths returned by the tool.
5.  **Editing Restraint**: For `image_editing`, preserve the original image look and structure unless the user explicitly requests broader restyling.
6.  **ROI Fidelity**: When the user marks or highlights a region, treat that region as the authoritative edit scope.
7.  **Scene/Continuity Styling Only Where Needed**: Production-bible `global_style` is mandatory for `scene_generation` or `continuity_generation`, but not for straightforward editing tasks.
8.  **Multi-Image Consistency**: For multi-image edit requests, apply the same explicit requested change consistently across all supplied images unless the user says otherwise.
9.  **No Narrative Escalation for Plain Edits**: Multiple source images alone do not justify switching to narrative-sequence behavior.
10.  **Anonymize Prompts**: Regardless of user instructions, character names (e.g., "Veeran", "Maharana Pratap") MUST be substituted with generic visual descriptions (age, gender, clothing) when internally constructing prompts for AI generation. This ensures the model (which has no context for names) produces accurate results and avoids filtering.
