---
name: talking-head-edit
description: Plan a polished talking-head edit from word timestamps, preserve natural breath and speech boundaries, add justified B-roll and measured BGM, and hand a generic EDL to pi-media. Use when the user asks to cut pauses, tighten spoken delivery, remove dead air, edit a monologue, add B-roll, or add background music to a talking-head video.
---

# Talking-head edit

Use `pi-speech` for word evidence, this package for editorial decisions, and `pi-media` for deterministic rendering.

1. Treat a request to cut breaths/pauses or edit a talking head as a **full expression edit** by default: review the entire source, compare takes, remove redundant retakes and false starts, preserve unique information, then adjust pauses. Use `scope: pauses-only` only when the user explicitly restricts the task to gaps; record their restriction in `scopeReason`. Never invent a restriction such as “do not remove repeated sentences.”
2. Obtain a `pi-speech`-compatible transcript JSON containing word-level `beginMs` and `endMs`. Do not infer frame-accurate cuts from sentence text alone. Probe the source video, then call `talking_head_create` with the exact `sourceDurationMs`. Its initial timeline is a draft, returned as `previewMediaOperation`, not a completed edit. The 500ms cut threshold and 50ms head/80ms tail padding are starting suggestions, not universal rhythm rules. Adjust each join from source evidence; never standardize one video's durations or fades.
3. Read the complete transcript (all sentence pages or `includeTranscriptText`) and inspect the full source to identify complete takes, abandoned starts and information unique to each take. Review the full sentence context, filler candidates, and repetition candidates together with `contentCandidates`, pauses and `diagnostics`. Candidate lists are incomplete hints: missing candidates never mean no retakes. ASR segment boundaries are not proven semantic boundaries. Contiguous word timestamps do not prove absence of silence; inspect original audio/waveforms. Treat `safe` and every recommendation as evidence, not an instruction. Preserve pauses that carry emphasis, emotion, topic boundaries, or a deliberate breath.
4. Choose coherent, complete expressions before shortening gaps. Compare whole takes, remove complete false starts at word boundaries, and reinsert unique information from a discarded take where it belongs. Never delete a filler token automatically. A text-only delivery cue is low-confidence evidence; review audio and picture before relying on emotion or performance intent. Do not convert uncertain ASR spellings into confirmed spoken errors. Keep uncertain content and record the issue instead of silently skipping the whole category. Read [cut craft](references/cut-craft.md) for join decisions.
5. Summarize the actual proposed expression and rhythm changes, including retained unique information. Continue when editorial judgment is already delegated by the user's request; otherwise obtain approval for the concrete proposed changes. Use `talking_head_get` with `includeWords` for indexes, then call `talking_head_editorial_plan` **before applying the timeline** with the planned A-roll and structured `plan`: every reviewed sentence index; honest source-review method/notes; decisions for every pause/filler/repetition/content candidate; keep/remove/review word ranges for takes and false starts; every removed word's disposition; retained `uniqueInformation`; a reason for each output segment; and explicit unresolved issues. `candidateIds` can be empty for an issue discovered by full-source review. An empty `uniqueInformation` list means the full-take comparison found none, not that it was skipped.
6. Keep the returned `editorialPlanReceipt` unchanged and pass it to `talking_head_apply` with the exact planned A-roll. Missing plans remain `draft`; unresolved issues remain `needs-review`; explicit gap-only work is `pauses-only-reviewed`. These states expose only `previewMediaOperation`. A full reviewed plan exposes `mediaOperation` with `content-reviewed` status, while output listening still remains pending. You may render a clearly labeled preview at any stage; never call a draft or an unresolved edit completed. An A-roll change requires a new plan; B-roll/BGM-only changes retain the review of unchanged A-roll.
7. Stabilize A-roll before B-roll. A-roll carries the primary message. During the first analysis, prefer A-roll for the opening, conclusion, opinion, emotion, and transition; do not create a B-roll need merely because a noun has a matching asset. Create a need only when the picture adds visible information about an object, action, place, comparison, or evidence. For every candidate, keep the complete sentence context in `speechText`, put only the exact visual cue in `visualCueText`, choose one purpose, add 1-20 search terms, and explain the visual information in `reason`. Use `necessity: essential` only when the claim cannot be understood or verified as well without the picture; otherwise use `supporting`. The visual cue determines the shortest useful window. A complete spoken idea is the maximum context boundary, not a requirement to cover the entire sentence; never use a fixed three-second rule.
8. Call `talking_head_broll_plan` with the stable A-roll and all candidates before matching assets. The planner removes supporting transitions that add no visible information, keeps each viewing run at or below 8000ms, and reserves at least 3000ms of ending A-roll. An A-roll island shorter than 3000ms does not reset the viewing run; this three-second threshold protects meaningful presenter visibility and is not a fixed B-roll duration. Let the topic and visible information determine total B-roll coverage; do not optimize for a fixed percentage. If essential needs violate the viewing-run or ending-A-roll rules, shorten, move, or split their exact visual cues and plan again. Inspect `plan.selection.omittedNeeds`; Never match an omitted need. Match only `plan.needs`, and keep the returned `continuityPlanReceipt` unchanged. Its v4 hash binds the selected semantics and A-roll visibility policy. The planner may extend a selected B-roll across an A-roll gap of 500ms or less so selected clips hand off directly. A jump with `status: review` only asks for picture review. Keep A-roll by default; create `purpose: mask-cut` only after the actual picture is visibly objectionable. Save the reviewed image or video inside the workspace and submit `visualReview: { decision: "mask-with-broll", reviewedJumpCutOutputMs, artifactPath }`; apply rejects changed evidence.
9. For every selected `plan.needs` item, call `talking_head_broll_match` before generating contact sheets, passing the unchanged `continuityPlanReceipt`. Keep its returned `matchReceipt` unchanged; selection rejects assets outside that matched candidate page. Follow the bounded result:
   - `filename-direct`: probe and visually inspect only the selected asset to choose its source window. Do not call `media_contact_sheet` for unselected assets.
   - `filename-shortlist`: inspect only the returned shortlist and stop as soon as one asset satisfies the need. If none does and `nextCandidateOffset` is non-null, request that next page; do not repeat the first batch.
   - `visual-fallback`: use low-cost contact sheets on the bounded shortlist because filenames supplied no useful evidence. Exhaust that batch before requesting a non-null `nextCandidateOffset`.
   - If the scan is truncated, do not paginate it: narrow the asset directory and rescan so the candidate inventory is stable.
10. After selecting one asset, call `media_probe`, then use `media_contact_sheet` only as densely as timing requires: one high-density pass for a short ambiguous clip, medium for a normal clip, or low over a long clip followed by high density over the narrowed range. Read `contact_sheet_manifest.json` timestamps instead of guessing times from the PNG. Verify a continuous source window at least as long as the selected exact visual cue, preferably with 500ms handles on both sides. If it is too short, choose another source moment or asset; do not expand the output window back to the full sentence.
11. Build a shot group when one spoken idea benefits from several views: use genuinely different subjects, actions, scales, or angles. Never select the same asset twice, even from non-overlapping source moments. Review the actual selected window. Call `talking_head_broll_select` with the unchanged `continuityPlanReceipt`, the corresponding unchanged `matchReceipt`, selected asset, manifest path, exact source start, reviewed manifest timestamps, and a concrete `visualIdentity` containing subject, action, shotScale, and angle. The returned selection hash binds the match, identity, asset, manifest, and timestamps. Use the complete returned `placement`, including `selectionReceipt`; do not shorten bridged output ranges or hand-author selection fields.
12. When BGM is requested, call pi-media `media_audio_analyze` twice before applying the timeline: analyze the source video with `ranges` exactly matching the final A-roll in output order, then analyze the complete BGM with no ranges. Open the generated PNGs; they show a waveform, short-term LUFS, and absolute `HH:MM:SS.mmm` timestamps. A waveform does not prove emotion by itself, so combine its visible energy/section changes with the requested mood and available metadata.
13. Call `talking_head_bgm_plan` with those two unchanged `audio_analysis_manifest.json` files. Default to 12 LU below dialogue for ordinary dialogue-led videos: dialogue sidechain ducking provides additional protection while 12 LU keeps the music audible. Increase toward 14–18 LU only when the user explicitly wants a subtler bed or the selected music is unusually dense or bright; never treat 18 LU as a generically safer default. The supported range remains 6–30 LU. Do not type a gain percentage. The tool binds both manifests and the exact A-roll, then tells you whether the source must be trimmed or looped. Read start/end timestamps from the full BGM waveform PNG, call `media_audio_analyze` a third time with exactly that selected BGM range, then call `talking_head_bgm_select` with the timestamps and this selected-window manifest. A long BGM must provide one exact output-length source window; a short BGM uses a visually reviewed loop window. Final gain uses the selected window's Integrated LUFS, not the whole song's average. Keep the complete returned `bgm`, including the selection receipt. Loop joins are crossfaded, speech triggers sidechain ducking, and the final fade-out prevents an abrupt ending.
14. Call `talking_head_apply` using the same planned A-roll, every validated placement for `plan.needs`, the unchanged `continuityPlanReceipt`, and the complete selected `bgm` when present. The tool rejects omitted or otherwise unplanned ranges, placements without matching provenance, repeated asset SHA-256 values, duplicate visual identities across different files, output overlaps, viewing runs broken only by A-roll islands shorter than 3000ms, ending A-roll shorter than 3000ms, A-roll flashes of 500ms or less, and changed BGM assets/manifests/waveform evidence. Placements created before 0.1.9 must be matched and selected again if they lack `matchReceipt`. Legacy v1-v3 B-roll receipts remain readable, but any legacy receipt containing B-roll must be replanned as v4 before apply. Every B-roll window must have a concrete visual purpose in `reason`; keep `audio: keep-primary`.
15. Create/read a `pi-media` project for the same source. Pass the returned `mediaOperation` unchanged to `edit_apply`, then `render` the exact new revision and call `review`. Use the official renderer with the source frame rate by default; retain the returned frame/sample-aligned timeline mapping and render receipt for subtitles and review. Check the actual output joins, meaning, rhythm, tail sounds and both audio channels. If output audio cannot be heard directly, explicitly report **pending listening review** and give output-time join positions. Waveforms, contact sheets, decoding success and technical `review.accepted` cannot establish listening approval. Never claim “listened while cutting” from visual evidence alone.

Never overwrite source media or outputs. If a revision conflict occurs, re-read both projects and reconcile intent. If the source, transcript, B-roll, BGM, analysis manifest, or waveform artifact hash changed, stop and ask whether to create a new project rather than silently adopting new bytes.

Read [cut craft](references/cut-craft.md) when deciding whether a pause is natural, abrupt, or suitable for B-roll.

Read [tool handoffs](references/tool-handoffs.md) before B-roll matching/selection or adding BGM to an existing B-roll timeline. Prefer `needId` over copying a need; use `inheritBroll: true` with unchanged A-roll instead of repeating match/select. Recover from structured window errors using their actual timestamps, not source-code inspection or ad hoc scripts.
