# Progress Tracking — ideation & design

**Status:** proposal. Goal: a reliable, accurate progress readout for every
generation, surfaced natively in the pi TUI, with zero new burden on the
end user.

## The accuracy problem: ComfyUI's progress number is a lie (for video)

ComfyUI exposes `GET /progress` → `{progress, max, queue_remaining}` and a
websocket with `progress` messages (`value/max`). That counter covers **only
sampling steps**. For this factory's default lane — H3 Turbo at 4–8 steps —
sampling is the *short* part of the wall clock:

```
download weights (first run only) → load into VRAM → encode → sample (4-8 steps) → decode + assemble
        minutes (dominates)                          seconds
```

A raw sampling bar would sit at 0% for most of the job, then jump 0→100% in
four ticks. "Accurate" progress therefore has to be **phase-weighted**, and
the phase weights have to come from *measured* jobs, not guesses.

## The phases

| Phase | What happens | Progress signal | Weight source |
| --- | --- | --- | --- |
| setup | weight download (first run on a pod) | none — wall clock | first-run flag |
| load | weights into VRAM, first-run quantization | none | measured per (model, gpu) |
| encode | prompt/image → latents | none | measured |
| sample | diffusion steps | `/progress` value/max (the only live signal) | measured per-step |
| decode | VAE decode, video assembly, save | output file appears in `/history` | measured |

The agent already observes every boundary: `t0` = POST `/prompt` accepted,
`t1` = first `progress` message, `t2` = last progress, `t3` = `/history`
shows outputs. Recording `load_s / encode_s / sample_s / decode_s` on the job
ledger entry turns **every completed job into calibration data** — after two
or three jobs the ETA model for (model, gpu, workflow) is real, not guessed.

## What pi gives us natively (researched against the extension API)

- `ctx.ui.setWidget(key, string[] | component)` — a persistent panel above the
  editor. **This is the live generation dashboard**: job id, phase bar, ETA,
  cost-so-far, queue position. Any string-array content, refreshed on demand.
- Tool `_upd` callback + `ToolExecutionUpdateEvent` + `renderCall` /
  `renderResult` with per-call shared state — a tool can stream updates and
  draw a real bar under its own row in the TUI.
- `setWorkingIndicator({frames})` — animated indicator during tool execution.
- `ui.notify` — coarse one-shot messages (already used).

Ecosystem scan (npm `pi-package`): pi-spark, pi-ui-tweaks, pi-footer,
pi-subagents, pi-ask — all UI polish/agent helpers. **Nothing tracks an
external job's progress**; this is greenfield, which is fine — the surfaces
above are all we need.

## The design

1. **Ledger records phase timings.** `Job` gains `load_s, encode_s, sample_s,
   decode_s` (agent records the observed boundaries at `lvrged_factory_job action=finish`).
   The workflow manifest gains a calibrated ETA model:
   `{ load_s, encode_s, per_step_s, decode_s }` per (model, gpu) — updated
   from every completed job (rolling average; drop outliers).
2. **New tool `lvrged_factory_job action=progress`.** Params: `deployment_id`,
   `prompt_id`, `job_id`. It GETs `/progress`, `/history`, `/queue` on the
   deployment endpoint and returns a structured readout:
   ```
   { phase: "load"|"encode"|"sample"|"decode"|"done"|"stuck",
     pct: 0-100, eta_s: number|null, steps_s: number,
     queue_remaining: number, cost_so_far_usd: number }
   ```
   ETA ladder: calibrated weights → running estimate (elapsed × completed
   weight fraction) → `null` (show phase + steps only). While executing, the
   tool also refreshes the widget via `ctx.ui.setWidget`.
3. **Agent polling loop (skill-level, gpu-ops skill).** After
   `lvrged_factory_job action=run`, the agent POSTs the workflow (as today), captures
   `prompt_id`, then polls `lvrged_factory_job action=progress` every ~5–10s until
   the job is done. Each poll updates the widget; on completion the widget
   shows the summary (duration, cost, artifact URI) for a few seconds, then
   clears to a last-run line.
4. **Hung-job detection for free.** If phase+pct don't advance across N polls
   (configurable, default ~3 min), the tool returns `phase: "stuck"` and the
   agent runs the existing failure protocol (classify → fix once → retry once)
   instead of silently burning pod hours. This is the money-saver.
5. **Batch view.** `queue_remaining` from `/queue` → widget shows
   "job 3 of 7 · 2 ahead" so a queue drain reads as a pipeline, not a mystery.

## Options considered

- **A. Agent-driven polling + progress tool (recommended).** Zero daemon
  code; matches the package's core rule ("extension is dumb, agent executes");
  the poll loop lives in the skill, the tool is a thin stateless fetch+compute.
- **B. Extension background watcher.** Autonomous widget updates without
  agent turns, but adds a polling daemon with session-lifecycle and stale-job
  edge cases. Keep as a v2 if the agent-loop cadence feels slow.
- **C. Synchronous `gpu_run` with streamed updates.** Would give the
  nicest inline renderer, but restructures the whole ledger-first flow.
  Rejected — the architecture is deliberately non-blocking.
- **D. Websocket instead of polling.** Lower latency, per-prompt_id accuracy,
  but adds a WS client to the toolchain. Not worth it until a batch runner
  exists; `/progress` with one job per deployment (the norm) is unambiguous.

## Concurrency caveat

`/progress` is *global* — it reports the currently-executing item, not a
specific prompt. With one job per deployment (the factory's normal shape) it's
correct. If a user ever runs two jobs on one pod, the widget can briefly show
the other job's phase; documented, not fixed — the fix (per-prompt WS tagging)
is v2 material.

## Implementation sketch

- `extension/registry.ts` — `Job` phase fields; `WorkflowManifest` gains
  `calibration?: {load_s, encode_s, per_step_s, decode_s}`.
- `extension/tools.ts` — `lvrged_factory_job action=progress` tool; pure
  `computeProgress(calibration, observed)` helper (unit-testable); widget key
  constant `"lvrged-factory-progress"`.
- `extension/commands.ts` — `/lvrged-factory gpu progress <job>` manual readout.
- `skills/lvrged-factory-gpu-ops/SKILL.md` — the polling loop, phase boundary
  recording, stuck detection, calibration update; `comfyui` skill notes the
  `/progress` + `/queue` endpoints.
- First job per pod: mark as setup run (weight download) — show phase, no ETA.

**Bottom line:** the widget + phase-weighted ETA makes every generation read
as a progress bar that's actually right, and the same machinery turns "job
didn't move for 3 minutes" into an automatic money-saving alert. The cost is
one small tool, five fields in the ledger, and a polling loop in the skill —
no daemon, no new runtime, no end-user setup.
