# Cut craft

## Full expression edit comes first

Review the entire source before trusting pause candidates. Group repeated takes, choose a coherent complete version, remove abandoned starts, and preserve information that occurs only in the rejected take. Reordering source ranges is supported: a useful sentence can be inserted into the chosen take. Record these decisions and their exact word ranges in the editorial plan before applying it.

Do not equate ASR segments with semantic sentences, or touching word timestamps with continuous sound. Text matching can locate a repeated phrase but cannot decide whether it is intentional emphasis or redundant content.

## Cutting a breath

- Cut only between word timestamps. Never cut through a word or punctuation tail.
- The automatic padding policy accepts 30–200ms on both sides. Its default 50ms before and 80ms after is asymmetric because endings often need more decay. Final word-safe A-roll ranges can retain longer natural pauses; decide the duration per join, not from a fixed target.
- A long waveform gap is evidence of silence, not proof that it should disappear. Keep pauses that signal emphasis, emotion, a paragraph boundary, or a visible gesture.
- Prefer one clean removal over many micro-cuts. Dense sub-150ms cuts create robotic cadence and visible jump cuts.
- `pi-media` adds a 30ms audio fade at every primary-segment edge to suppress clicks. This does not repair a semantically bad cut.

## Editing spoken content

- Read the complete sentence before removing a filler or repeated word. `啊`, `嗯`, `那个`, and `就是` can carry hesitation, emphasis, transition, or actual meaning.
- Treat transcript punctuation as a low-confidence delivery hint, not acoustic emotion detection. Confirm expressive decisions against the source audio and picture.
- Prefer removing a complete false start at word boundaries over deleting an isolated sound that leaves an unnatural join.
- If a necessary content cut creates a hard join, first adjust the retained source handles or choose another take. Preserve room tone and use a short audio transition when the renderer supports it. This package does not yet expose inserted pauses, room-tone tracks or per-join crossfades: never claim those repairs were performed. A picture-only B-roll overlay cannot repair an audio seam. Edge fades suppress clicks but are not crossfades or listening approval.

## Using B-roll

- Place B-roll on the output timeline after A-roll cuts have stabilized.
- Use it to show the object, action, place, comparison, or evidence currently being discussed—not as random decoration.
- During the first analysis, keep the opening, conclusion, opinion, emotion, and transition on A-roll unless the picture adds indispensable information.
- Record the full sentence as context, but place the output window only on the exact words whose information must be seen. The complete thought is an upper semantic boundary, not a coverage target.
- Treat a candidate as essential only when the spoken claim cannot be understood or verified as well without the picture. Let the planner omit supporting candidates before any asset search.
- Treat a filename match as asset-level evidence, never as proof that a particular source window is usable. Inspect continuous motion before selecting `assetStartMs`.
- Start with filename and directory metadata. Generate contact sheets only for the selected asset, a bounded ambiguous shortlist, or a bounded fallback batch whose names carry no meaning.
- For a long selected asset, locate a broad range cheaply and then generate a denser contact sheet only for that range. Use manifest timestamps, not visual timestamp transcription.
- Ensure the selected source window covers the planned exact visual cue. Prefer extra source handles so a later timing adjustment does not require rescanning.
- Let meaning and delivery rhythm set the duration and total B-roll coverage; do not optimize for a fixed percentage. Keep viewing runs at or below eight seconds, count A-roll islands under three seconds inside the same run, and reserve at least three seconds of A-roll at the ending.
- For one semantic beat, use a shot group only when its subject, action, scale, or angle is visibly different. Never reuse the same asset SHA-256, and give every reviewed source window a structured visual identity so duplicated shots in different files are rejected.
- If neighboring B-roll windows expose 500ms or less of A-roll, plan a direct handoff by extending the previous B-roll need before source-window selection. Do not let a few frames of A-roll flash between overlays.
- Treat every uncovered jump cut as a picture-review candidate, not a request for B-roll. Add a mask-cut only after confirming the actual jump is objectionable, recording its output timestamp, and binding the reviewed workspace image or video into the continuity receipt.
- Keep the primary voice track. B-roll audio replacement is outside the current contract.
- Record the search query and editorial reason so a later agent can replace the asset without guessing intent.
