---
name: myai-visual-generator
description: "Generates high-quality images, videos, infographics, and research visuals using AI APIs (Gemini, Imagen, DALL-E, GPT Image, FLUX, Veo) with fal.ai MCP model discovery. Use when creating hero images, blog post visuals, diagrams, infographics, architecture diagrams, flowcharts, sequence diagrams, data visualizations, or video content. Also use when the user mentions generating images, creating visuals, making diagrams, or needs any kind of visual content for articles, presentations, or projects."
argument-hint: "[prompt] [--visual-type=image] [--type=hero] [--service=gemini]"
allowed-tools: [Read, Write, Bash, Task, "mcp__fal-ai__*"]
context: fork
---

# Visual Generator

Generate images, videos, infographics, and research visuals using AI generation APIs, Mermaid rendering, HTML conversion, and PaperBanana.

## Bundled Scripts

The generation SDK is bundled at `${CLAUDE_SKILL_DIR}/scripts/visual-generation-utils.js`. It provides:
- `validateAPIKeys()` — check which services are available
- `generateImage(prompt, options)` — auto-routed image generation
- `generateVideo(prompt, options)` — Veo 3 video generation
- `estimateCost(service, options)` — cost estimation
- `selectBestService(preferred)` — smart service selection with fallback
- `buildInfographicPrompt(config)` — structured infographic prompt builder
- `getServiceInfo(service)` — service metadata lookup

Run it via Node.js:
```bash
node -e "import('${CLAUDE_SKILL_DIR}/scripts/visual-generation-utils.js').then(m => m.generateImage('prompt', {type: 'hero'}).then(r => console.log(r)))"
```

Or reference it from a script you write during generation.

## Quick Start

```
/myai-visual-generator "Modern cloud-native architecture" --type=hero
/myai-visual-generator "API request lifecycle" --visual-type=infographic --type=flowchart
/myai-visual-generator "SaaS KPI dashboard" --visual-type=infographic --type=infographic-data
/myai-visual-generator "Product demo" --visual-type=video
```

## Arguments

Parse from: `$ARGUMENTS`

| Parameter | Description | Default |
|-----------|-------------|---------|
| `prompt` | What the visual should depict | Required |
| `--visual-type` | `image`, `video`, `infographic`, `research-visuals` | image |
| `--type` | Sub-type (see routing below) | hero |
| `--service` | Preferred generation service | auto |
| `--quality` | low, medium, high | medium |
| `--size` | 1024x1024, 1792x1024, 1024x1792 | 1024x1024 |

## Visual Type Routing

1. Ask **visual type** first: `image`, `video`, `infographic`, or `research-visuals`
2. Then ask the relevant **sub-type** only:
   - **Image**: `hero` or `illustration`
   - **Video**: user intent (demo, tutorial, walkthrough)
   - **Infographic**: `flowchart` or `sequence-diagram` (quick mode)
   - **Research**: `research-diagram`, `research-plot`, or `research-evaluation`
3. Collect only mandatory inputs; use defaults for everything else

### Routing Table

| Visual Type + Sub-type | Path |
|----------------------|------|
| Image (hero, illustration) | AI generation via bundled SDK |
| Video | Veo 3 via bundled SDK |
| Infographic (from article `*.md`) | Mermaid analyze → design → render → insert |
| Infographic (standalone) | Mermaid/Gemini diagram path or HTML conversion |
| Research visuals | PaperBanana CLI |
| Text-heavy data visuals | HTML-to-screenshot (Playwright) |

**Critical routing rule**: When the user needs pixel-perfect text/numbers (metrics, data, labels), use HTML conversion instead of AI generation. AI models can hallucinate text. See [services reference](references/services.md) for the full decision matrix.

## Workflow

1. **Check config**: Run `validateAPIKeys()` from the bundled SDK to see what's available
2. **Gather requirements**: Visual type, sub-type, prompt, preferences (minimum questions)
3. **Route**: Pick the generation path based on the routing table
4. **Estimate cost**: Show estimated cost and check against budget before generating
5. **Generate**: Execute via the appropriate path
6. **Save**: Organize into `content-assets/images/YYYY-MM-DD/` or `content-assets/videos/YYYY-MM-DD/`
7. **Return**: Provide markdown embed code and generation report

## File Organization

```
content-assets/
  ├── images/
  │   └── YYYY-MM-DD/
  │       └── {type}-{description}-{id}.png
  └── videos/
      └── YYYY-MM-DD/
          └── video-{description}-{id}.mp4
```

## Additional Resources

- **Service catalog, pricing, MCP tools**: See [references/services.md](references/services.md)
- **Infographic/Mermaid pipeline details**: See [references/infographic-pipeline.md](references/infographic-pipeline.md)
- **Research visuals (PaperBanana)**: See [references/research-visuals.md](references/research-visuals.md)
- **Generation SDK source**: See [scripts/visual-generation-utils.js](scripts/visual-generation-utils.js) for the full API

## Quality Guardrails

- Verify all visible text in generated images is free from typos
- For text-heavy visuals, prefer GPT Image 1.5 or HTML conversion
- If generated text has mistakes, regenerate with clearer constraints or switch to HTML conversion
- Validate rendered Mermaid outputs exist and are >1KB before inserting references

## Error Handling

- **No API keys**: Show which keys are missing and how to configure them (`/myai-configure visual`)
- **Rate limiting**: Wait 60s and retry, or switch service
- **Budget exceeded**: Show usage and suggest increasing limits in `.env`
- **MCP unavailable**: Fall back to SDK-only approach automatically
- **Mermaid render fails**: Save source `.mmd` files and provide manual render instructions

## Integration

- Works with `content-writer` skill via `--with-images` flag
- Automatic asset organization and cost tracking
- Markdown reference generation for easy embedding
- fal.ai MCP model discovery when `FAL_KEY` is configured
