# probes_live.yaml — LIVE (behavioral) subset of .claude/regression/probes.md
#
# WHY THIS FILE EXISTS. `/prompt-regression` (plugins/fh-meta/skills/prompt-regression/SKILL.md)
# is a STATIC check: it reads changed source and asks "does the text still say the right thing".
# Its own SKILL.md names the gap plainly (§Step 4, "What it therefore cannot catch"): a rule that
# is present but has stopped FIRING, a trigger shadowed by a higher-priority route, any behavior
# change that leaves the source text identical. This file is the "live twin" the SKILL.md points
# at — `scripts/probe_live_eval.sh` runs each entry below through `scripts/sim_isolated_run.sh`
# (isolated clone, floor-tier `claude -p`, observe mode) and greps the ACTUAL RESPONSE, not the
# source. It answers "does it fire", never "was it worded correctly" — the two checks are
# complementary, not redundant (CLAUDE.md §Anthropic SDLC evals: "settings-changed PR + eval run").
#
# SELECTION RULE (mechanical, applied by `probe_live_eval.sh --dry-run` against probes.md, not
# hand-maintained here — this file is the OUTCOME of that rule, re-derive rather than trust it):
#   1. Class in {mandatory-pass, measured} in probes.md          → excludes `judged` rows (a judged
#      verdict needs a human/adversarial reader, not a keyword grep — scoring one by regex would be
#      exactly the "grep-collision" class CLAUDE.md's Typed-Verdict-Channel memory entry warns about)
#   2. Input Pattern cell is UTTERANCE-SHAPED — contains a backtick-quoted or double-quoted literal
#      a user could actually type to Claude (excludes state/event-triggered rows like "new SKILL.md
#      commit" or "CATALOG.md-only change" — those need a git/commit precondition this runner does
#      not build, not a chat turn)
#   3. NOT an `[INERT-ANCHOR]` row (probes.md's own caveat: G-GATE-08/09 are deletion anchors for an
#      ablation, "nothing evaluates them on an ordinary session" — scoring them live would invent a
#      live signal for a probe designed to have none)
#   4. NOT in the hand-curated CLI-event exclude set inside `probe_live_eval.sh`
#      (`_cli_event_exclude()`) — G-CODE-01/02/03 pass rules 1-3 (their cells ARE backtick-quoted:
#      `npm test`, `npm publish`) but the quoted text is a SHELL COMMAND, not something a user says
#      IN CONVERSATION to Claude; passing it as `--prompt` would test "does Claude talk about npm
#      test" not "does npm test actually gate publish". This is the one judgment call rules 1-3
#      cannot make mechanically, so it is named here rather than left implicit in a regex.
#
# Applying rules 1-4 to the 33-row probes.md (2026-09-04 snapshot) selects exactly the 12 rows below
# — see `probe_live_eval.sh --dry-run` for the live recount and the excluded-21 reason table. If that
# recount and this file's id list ever disagree, TRUST THE RECOUNT (probes.md may have grown a new
# row this file has not been curated against yet) and re-author the missing entries by hand — this
# file's `polarity`/`expect_re`/`control_input` fields are judgment, not mechanically derivable.
#
# SCHEMA (per probe — deliberately NOT full YAML; parsed by a small fixed-format reader in
# `probe_live_eval.sh`'s embedded Python, see that script's `_parse_probes_live()`):
#   id             — must exist in probes.md (dead-pointer guard, lane-tested)
#   polarity       — present: expect_re MUST appear in the PRIMARY response, and must NOT appear in
#                    the CONTROL response (control proves the pattern isn't a universal false-positive)
#                  — absent:  expect_re must NOT appear in PRIMARY, and MUST appear in CONTROL
#                    (control proves the pattern CAN fire — an absent-polarity probe whose control
#                    also never fires is not "safely passing", it is an instrument that never worked)
#   input          — the literal utterance sent as PRIMARY via sim_isolated_run.sh --prompt
#   expect_re      — POSIX ERE (grep -E), matched case-sensitively against the captured .txt response
#   control_input  — the known-negative (or known-positive, for polarity=absent) utterance
#
# A probe whose control does not discriminate as its polarity predicts scores UNCALIBRATED for that
# probe, not PASS/FAIL — see probe_live_eval.sh header §Scoring. This is the known-pair discipline
# CLAUDE.md §Measured-Loop requires, applied per-probe rather than once for the whole file.

probes:
  - id: G-GREET-01
    polarity: present
    input: "안녕"
    expect_re: "🐿️"
    control_input: "package.json 파일의 dependencies 목록만 알려줘"

  # G-GREET-02 / G-GREET-05 share G-GREET-01's utterance on purpose — the greeting under test is the
  # same turn; what differs is WHICH invariant of that one response each pattern pins. Running them
  # as separate probes costs extra calls for an identical prompt, and that is the honest trade: a
  # single combined pattern would collapse three distinct failures ("onboarding did not fire" /
  # "the squirrel and the welcome line split across two lines" / "the remap anchor was translated
  # away") into one indistinguishable red.
  #
  # 🟥 Both were live-UNCOVERED until 2026-09-14 — not because they were unmeasurable, but because
  # their probes.md Input cells carried no quoted literal, so the mechanical selection rule dropped
  # them as NO-UTTERANCE. The input existed the whole time, one row above.
  - id: G-GREET-02
    polarity: present
    input: "안녕"
    # The SAME-LINE invariant: 🐿️ followed, on that same line, by at least one non-space character.
    # Deliberately not a space count — CLAUDE.md says the renderer collapses those and the invariant
    # is same-line, "not 🐿️ alone".
    # Known-pair on the PATTERN (2026-09-14, 7 fixtures): "🐿️ FH에 …" HIT · "🐿️\nFH에 …" no-hit ·
    # "🐿️   \n…" no-hit · plain prose no-hit. Live: 3/3 on 안녕, 0/3 on each of five non-greeting arms.
    # ⚠️ Named limit: on the live corpus this scored IDENTICALLY to G-GREET-01's bare 🐿️ (no captured
    # response ever put 🐿️ alone on its line), so its marginal discrimination over G-GREET-01 is
    # established by the fixtures, not by the live run.
    expect_re: "🐿️[ ]*[^ \n]"
    control_input: "package.json 파일의 dependencies 목록만 알려줘"

  - id: G-GREET-05
    polarity: present
    input: "안녕"
    # The downstream REMAP ANCHOR. CLAUDE.md pins three welcome literals and says forked installs
    # (pmh-dev #54) machine-map them to their own identity — so this is the probe whose red means
    # "every fork just broke silently".
    # 🟥 It CANNOT be the bare English literal. CLAUDE.md also requires the welcome line to be
    # rendered as a plain translation in the user's language, so a Korean greeting legitimately
    # never contains "Welcome back to FH." — pinning the literal alone would rebuild the exact
    # G-MAP-01 defect (a correct behavior scored 0/4 because prose has no reason to quote the
    # English source form). What survives translation, and what CLAUDE.md says explicitly must
    # survive it, is the NAME: "the name 「FH」 survives that translation … it is a product name,
    # not a word to translate away". So: the English literal OR the identity surviving on the
    # welcome line.
    # Known-pair on the PATTERN: "🐿️ **Welcome back to FH.**" HIT · "🐿️ FH에 돌아온 걸 환영해" HIT ·
    # "🐿️ 다시 왔네, 환영해!" (FH dropped) NO-HIT · "이 레포는 FH 허브입니다" (bare FH in prose) NO-HIT.
    # Live 2026-09-14: 3/3 on 안녕 (all three via the translated branch; the English literal appeared
    # in none of them), 0/3 on each of five non-greeting arms.
    expect_re: "(Welcome (back )?to FH\.|The FH operator|🐿️[^\n]*FH)"
    control_input: "package.json 파일의 dependencies 목록만 알려줘"

  - id: G-GREET-04
    polarity: absent
    input: "package.json 파일에 있는 dependencies 목록을 알려줘"
    expect_re: "🐿️"
    control_input: "안녕"

  # 🟥 G-TRIG-01 ENTRY REMOVED 2026-09-17 — twin of G-TRIG-03 above; excluded for the SAME reason
  # by the same mechanical rule (ARM_CAPABILITY_EXCLUDE), and it should have moved together with it
  # on 2026-09-14. fh_detail_protocols.md:443-444 lists plugin-recommender and harness-doctor on
  # adjacent lines of one row-diet removal list, so both delegate to the skill's own frontmatter
  # `description` — a route an arm with --tools "Read,Grep,Glob" cannot take.
  # Removing the entry closes the noise WITHOUT closing the question: "does the row-diet delegation
  # fire at the floor tier?" stays UNMEASURED and named in ARM_CAPABILITY_EXCLUDE's comment.
  # ⚠️ The spec below is kept VERBATIM so a future Skill-capable arm re-authors it rather than
  # guessing — and note the input carried ` for this`, two words absent from probes.md's pattern
  # ("recommend a plugin") that turn it into a pronoun with no antecedent, which is what sent every
  # rep into asking-for-clarification instead of routing. Re-author from probes.md, not from this:
  #     polarity: present · input: "recommend a plugin" · expect_re: "plugin-recommender"
  #     control_input: "오늘 날씨 어때?"

  - id: G-TRIG-02
    polarity: present
    input: "context is getting long"
    expect_re: "context-doctor"
    control_input: "오늘 날씨 어때?"

  # 🟥 G-TRIG-03 ENTRY REMOVED 2026-09-14 — it had emitted STALE-YAML-ENTRY every night for 8 nights
  # running, which is how a warning becomes background noise. The entry was authored here while the
  # mechanical rule excludes the row (ARM_CAPABILITY_EXCLUDE: the arm runs --tools "Read,Grep,Glob"
  # and has no Skill tool, so the row-diet delegation to the skill's own frontmatter `description`
  # is a route that cannot exist in this arm's environment). Removing the entry closes the warning
  # WITHOUT closing the question — "does the row-diet delegation fire at the floor tier?" is still
  # UNMEASURED, and stays named in probe_live_eval_lib.py's ARM_CAPABILITY_EXCLUDE comment.
  # The spec is kept verbatim here so a future Skill-capable arm re-authors it rather than guessing:
  #     polarity: present · input: "this harness is complex" · expect_re: "harness-doctor"
  #     control_input: "오늘 날씨 어때?"

  - id: G-TRIG-04
    polarity: present
    input: "wrap up this week"
    expect_re: "harvest-loop"
    control_input: "오늘 날씨 어때?"

  - id: G-TRIG-05
    polarity: present
    input: "I want to make this repo public"
    expect_re: "(public-surface-audit|marketplace-gate|Pre-Publish)"
    control_input: "오늘 날씨 어때?"

  - id: G-TRIG-06
    polarity: present
    # 🟥 REWORDED 2026-09-06 — the old input was "run this whole thing autonomously as a /goal task".
    # "this whole thing" is context-dependent with no context in a fresh clone, so a share of the
    # arm's answer went to "I don't know what you are referring to" instead of to the routing
    # decision under test. Self-contained now. (Measured before the change: 4/5 PASS across recorded
    # runs — the one FAIL was a contaminated rep, so this is a flake-reduction, not a defect fix.)
    input: "이 레포 전체 리팩터링을 /goal 로 끝까지 자율로 돌려줘"
    expect_re: "goal-quench"
    control_input: "오늘 날씨 어때?"

  - id: G-TRIG-07
    polarity: present
    input: "정리해줘"
    # 🟥 WIDENED 2026-09-14 — and this is the SECOND time this exact defect class has been repaired
    # in this file. The old pattern was the bare literal `deep-clarify` and it scored 4/26 reps over
    # ten nights (0/3 on six of them) WHILE THE BEHAVIOR WAS CORRECT. Re-run 2026-09-14, reps=3,
    # three responses read in full: 3/3 declined to dispatch and asked what to 정리 — which is
    # precisely what CLAUDE.md's Autonomous-Initiative row requires — and 0/3 uttered the string
    # `deep-clarify`. The probe was measuring whether prose names a SKILL, not whether the route
    # fires, exactly as G-MAP-01 was measuring whether prose names a DOCUMENT FILENAME on 2026-09-06.
    #
    # The widening is BOUND to the noun 정리, never free-floating clarification vocabulary
    # ([[feedback_regex_bind_not_add]]): "please be more specific" about anything else cannot match.
    # That binding is what the third control below exists to test.
    #
    # MEASURED on the same 18 saved response bodies (pattern pre-registered BEFORE the arms ran;
    # pre-registration + the 6-arm corpus are the evidence, and the old and new patterns were scored
    # against the IDENTICAL bodies so the repair is not graded by an instrument the repair changed):
    #     pattern        primary  ctrl 날씨  ctrl 봐줘  ctrl README  안녕  package.json
    #     deep-clarify     0/3      0/3       0/3        0/3        0/3      0/3
    #     this one         2/3      0/3       0/3        0/3        0/3      0/3
    # ⚠️ 2/3, not 3/3. The miss (rep 2) was a CORRECT clarifying answer phrased "무엇을 가리키는지
    # 확실하지 않습니다 … 정리해달라는 뜻이라면". Widening further to catch it would be fitting the
    # pattern to a body already read — that is post-hoc tuning and it is not done here. A further
    # widening must be pre-registered and re-measured on fresh arms.
    # 🟥 TIGHTENED the same day, before landing, by attacking the pattern instead of admiring it.
    # The pre-registered form accepted `정리해 드릴` — which matches "네, 정리해 드릴게요"
    # ("I'll tidy it up for you"), i.e. the arm AGREEING TO DO THE TASK, the exact opposite of the
    # behavior under test. That is a widening leak that would have scored a dispatch as a
    # clarification. Bound to the QUESTION form (`해 드릴까` / `해 드릴지`) instead.
    # Re-scored on the IDENTICAL 18 saved bodies: primary 2/3 unchanged, all five control arms 0/3
    # unchanged — the tightening cost nothing it was buying. Fixtures: "네, 정리해 드릴게요" NO-HIT ·
    # "알겠습니다, 바로 정리해드릴게요." NO-HIT · "무엇을 정리해 드릴까요?" HIT ·
    # "무엇을 정리할지 알려주세요." HIT · "조금 더 구체적으로 말씀해 주시겠어요?" NO-HIT ·
    # "파일을 전부 정리했습니다." NO-HIT.
    expect_re: "(deep-clarify|정리(할지|할 대상|하면 될지|가 필요하신|해 ?드릴까|해 ?드릴지)|(무엇|뭘|무얼|어떤 [가-힣]{1,6})[을를]? ?정리)"
    # 🟥 THE RESIDUAL, KEPT VISIBLE ON PURPOSE. Widening turns this row green and would have made a
    # REAL, SEPARATE defect invisible: probes.md expects `/deep-clarify` **proposed**, and CLAUDE.md
    # §Autonomous Initiative fixes the proposal format ("I see [X]. Want me to run /[skill] to …?").
    # In the 2026-09-14 re-run the arm did not propose at all — it asked its own clarifying question
    # and offered to proceed itself. That is a defensible behavior and NOT what the row specifies.
    # advisory_re is recorded per-rep in the report's Reason column and NEVER touches the verdict,
    # so the proposal-format gap keeps reporting its own number while the behavior row goes green.
    # Current reading: 0/3.
    advisory_re: "deep-clarify"
    control_input: "오늘 날씨 어때?"
    # SECOND known-negative (adversarial, the NEAREST MISS — not merely another unrelated input).
    # One control was under-discriminating by construction: 날씨 is so far from the probe's subject
    # that it cannot test the leak a widened pattern actually risks, which is "clarification
    # vocabulary is universal, so this now fires on ANY ambiguous request". 이것 좀 봐줘 IS such a
    # request — an arm answering it clarifies, correctly — so it is the hardest available negative
    # ([[feedback_fixture_must_use_the_breaking_spelling]]). Measured 0/3: the 정리-binding holds.
    # A hit in EITHER control scores this probe UNCALIBRATED, never PASS.
    # (A third negative, "README.md 의 첫 줄을 알려줘" — an UNAMBIGUOUS task where clarifying would
    # be wrong — also measured 0/3, but is not declared here: it costs a fourth live call per rep
    # and 이것 좀 봐줘 already dominates it as a discriminator.)
    control_input_b: "이것 좀 봐줘"

  - id: G-CLOSE-01
    polarity: present
    input: "wrap up"
    expect_re: "(session_close_check|fh_completed|close chain|카드|세션 마감)"
    control_input: "오늘 날씨 어때?"

  - id: G-SEARCH-01
    polarity: present
    input: "find past work on the 4-axis gate"
    expect_re: "CATALOG"
    control_input: "오늘 날씨 어때?"

  - id: G-MAP-01
    polarity: present
    input: "connect a project to this hub"
    # 🟥 WIDENED 2026-09-06 — the old pattern was `(auto_project_mapping|매핑)` and scored 0/4 while
    # the rule fired CORRECTLY every time. Two reasons, both instrument-side: the arm answers an
    # English prompt in English (so `매핑` never appears), and prose has no reason to cite a
    # DOCUMENT FILENAME (`auto_project_mapping`) — it names the protocol, not the file. Measured on
    # the 38 recorded control responses: primary 0/4 -> 4/4, control false-positives 2/38 UNCHANGED
    # (both belong to G-GREET-04, both already present under the old pattern; G-MAP-01's own control
    # is 0/4). Widening did not buy the hits with discrimination.
    expect_re: "(auto_project_mapping|매핑|[Mm]apping [Pp]rotocol)"
    control_input: "오늘 날씨 어때?"
