openapi: 3.0.3
info:
  title: Repzo API - AI Object Detection Inference
  version: 1.0.0
  description: |
    **Per-task inference** runs the Ultralytics object-detection model on a
    single `ai-object-detection-task` image and (optionally) persists the
    predictions as a fresh `auto` annotation group on that task. It is the
    server-side proxy in front of the detection lambda.

    **Engines.** `engine` selects the detector:
    - `trained` — the Ultralytics lambda with a trained model (the classic path).
    - `zero_shot` — a hosted vision-language model (Qwen-VL over OpenRouter)
      prompted with the namespace's label names as the open-vocabulary
      detection prompt. For datasets/namespaces with NO trained model yet.
      The VLM's reply is normalized into the exact lambda result shape, so
      mapping/placement/analysis work unchanged; the saved annotation group
      carries `engine: zero_shot` + `zero_shot_model`. Needs at least one
      label in the namespace (their names are the vocabulary) — else `400`.
    - `auto` (default) — trained when a model resolves; when the classic
      dataset path finds nothing trained, falls back to zero-shot (when
      configured) instead of failing. Explicit `model` / `model_version`
      requests stay strict (no fallback).

    **Model resolution (trained).** Three paths, in priority order:
    1. `model_version` (and/or `model`) supplied in the body — used directly.
       This **bypasses** the task's dataset/default-model requirement and is the
       path used to place AR session frames. The lambda model id is
       `model` if given, else the version's `model`; the YOLO class-index →
       label fallback uses the version's `dataset_labels` order.
    2. `model` alone — used directly as the lambda model id; the model's
       `current_model_version` (when set) supplies the label order.
    3. Neither — the classic path: the task's `task_dataset` must point at a
       dataset with a `default_model` configured (else the zero-shot fallback
       under `engine: auto`, or `400`).

    **Label mapping.** Each prediction is mapped to a label by NAME
    (case-insensitive, unique per namespace), falling back to the YOLO class
    index into the version's `dataset_labels`. Unmapped predictions are
    counted in `unmapped` and never become annotations.

    **AR back-projection (session frames).** When the task carries `frame_meta`
    (pose + intrinsics) and a `depth_media` blob, each detection is back-projected
    into world coordinates: `world dims = bbox × depth × intrinsics × pose`.
    Placed annotations get `world_position`, `world_size`, `depth_at_center`,
    `depth_confidence`, `placement_confidence` and a `placed_box` snapshot;
    un-placeable ones get an `ignore_reason` (`no_pose` | `no_intrinsics` |
    `no_depth` | `empty_depth_region` | `insufficient_depth_pixels` |
    `behind_shelf`) naming exactly why. Both are persisted so downstream
    session analysis sees all of them — the kept/merged/ignored CONCLUSION is
    made there, not here. For non-session tasks (no `frame_meta`) the
    placement step is skipped. After placement the same DIMS RECLASSIFIER
    (`config.reclassify_labels`, off by default — re-labels a detection whose
    measured size fits a sibling label better, keeping `original_label` +
    `reclassification_reason`) and SIZE GATE (`config.size_gate`, on by
    default — flags `size_rejected` + `size_reject_detail` when the measured
    size exceeds the label's expected dims beyond the allowance) run as in
    the analysis stage.

    **Explainability (XAI).** `explain: true` makes this service call the
    dedicated DIAGNOSIS lambda (a separate function with its own memory
    budget, so per-request inference stays lean) after the prediction. It
    re-runs the identical forward pass with introspection hooks and returns an
    `xai` object, passed through unchanged in this response: per-detection
    top-k class scores ("why label A and not B"), an EigenCAM heatmap ("where
    did the model look", an RGBA PNG whose alpha channel is the activation —
    overlay it on the photo with an opacity slider), a TRUE gradient Grad-CAM
    heatmap ("what evidence drove the detections" — class-discriminative,
    computed via a gradient-enabled second pass), per-stage feature-map grids,
    2D UMAP/t-SNE/PCA embeddings of the detections, and the model version's
    training-time confusion-matrix image URLs. When the group is saved, the
    XAI results are persisted on it — heatmap/Grad-CAM/feature-map images are
    uploaded to media storage (refs + publicUrl snapshots), never inline in
    Mongo. Trained engine only; ignored for zero-shot. Each part degrades
    independently into `xai.notes`, and a diagnosis-lambda failure (or a
    missing `aiObjectDetectionDiagnosisUrl` config) degrades to `xai: null` —
    XAI never fails the prediction. When the diagnosis pass's detections do
    not align with this response's `predictions`, the index-bearing parts
    (`class_scores`, `gradcam.per_detection`, `embeddings.points`) are
    dropped with a note rather than mis-attributed.

    **XAI backfill (`explain_only: true`).** Explains the EXISTING detections:
    the detector is skipped, only the diagnosis lambda runs, and the XAI is
    grafted onto the task's already-saved unconfirmed auto group of the
    resolved model version — annotation `_id`s and provenance
    (`original_label`, size gating, cluster ids) survive, so stored analysis
    ledgers keep resolving. Falls through to a normal full inference when no
    such group exists (or when `save` is false). Rejected with `400` for the
    zero-shot engine (including the auto→zero-shot fallback): a VLM has no
    trained internals to diagnose, and falling through would re-detect —
    mutating the very group the caller asked to explain. This is how the
    dashboard's per-group Explain action computes XAI for an already-annotated
    task without churning its detections. (Session analysis itself does NOT
    take explain options — explaining is a per-task, per-group action.)

    **Saving.** With `save: true` (default) the mapped predictions become a
    new `auto`, unconfirmed group (`time` = now, `usable` left to its default).
    Only the unconfirmed auto group(s) from the SAME source are replaced
    (trained: same `model_version`; zero-shot: same VLM) — auto groups from
    other versions and every confirmed / manual group are kept. A re-run that
    maps zero annotations never wipes a populated same-source group (fresh
    XAI is grafted onto it instead). The write is a compare-and-swap on
    `updatedAt` (5 attempts); losing every attempt answers `saved: false`.
    `editor`, `edit_time` and `annotated` are restamped like a task update.

    **Who calls it.** Admins / reps with a valid JWT or api-key. The tenant
    key (`company_namespace`) is taken from the caller's token, not the body,
    and the caller's own credential is forwarded to the lambdas. When
    predictions are saved, the service emits `update-object-detection-task`
    so the review canvas live-updates.

    **Only `create` (POST) is allowed.** `find`, `get`, `update`, `patch`, and
    `remove` all reject with `400`. Cross-frame fusion lives in
    `ai-object-detection-session-analysis`. There is no stored document for
    this service — the response is computed per call.
servers:
  - url: https://sv.api.repzo.me
security:
  - ApiKeyAuth: []
  - JwtAuth: []
paths:
  /ai-object-detection-inference:
    post:
      summary: Run detection on one task
      description: |
        Resolves the model, calls the detection lambda, maps predictions to
        labels, optionally back-projects each detection into world coordinates
        for AR session frames, and (by default) saves them as a fresh `auto`
        annotation group on the task.
      operationId: createAiObjectDetectionInference
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: "#/components/schemas/InferenceRequest"
      responses:
        "201":
          description: The inference result, including mapped annotations and counts.
          content:
            application/json:
              schema:
                $ref: "#/components/schemas/InferenceResult"
        "400":
          description: |
            `task_id` missing; the task has no media / public URL; the given
            `model_version` / `model` was not found in the namespace; the task
            is not assigned to a dataset with a default model (and no zero-shot
            fallback applies); `explain_only` with the zero-shot engine; an
            unsupported `zero_shot_model`; or zero-shot requested with no labels
            in the namespace.
        "404":
          description: The task was not found in the caller's namespace.
        "500":
          description: The detection / zero-shot request failed, or the inference URL is not configured.
components:
  securitySchemes:
    ApiKeyAuth:
      type: apiKey
      in: header
      name: api-key
      description: |
        Server-issued API key. Also accepted via the `x-api-key` header or the
        `?apiKey=` query parameter as fallbacks.
    JwtAuth:
      type: apiKey
      in: header
      name: Authorization
      description: |
        Raw JWT in the `Authorization` header — **no `Bearer ` prefix**.
        Obtained from `POST /authenticate` (admin / rep / client login).
  schemas:
    InferenceRequest:
      type: object
      required:
        - task_id
      properties:
        task_id:
          type: string
          description: The `ai-object-detection-task` to infer on.
        model_version:
          type: string
          description: |
            `ai-object-detection-model-version` _id. When set, resolves the
            lambda model + label order directly and bypasses the dataset
            requirement.
        model:
          type: string
          description: |
            Object-detection model `_id` to infer with. Used directly as the
            lambda model id when given.
        engine:
          type: string
          enum: [auto, trained, zero_shot]
          default: auto
          description: |
            Detector selection — see the intro. `zero_shot` needs no trained
            model; `auto` falls back to zero-shot only when the classic dataset
            path resolves nothing trained.
        zero_shot_model:
          type: string
          enum:
            - qwen/qwen3-vl-8b-instruct
            - qwen/qwen3-vl-32b-instruct
            - qwen/qwen3-vl-235b-a22b-instruct
            - qwen/qwen2.5-vl-72b-instruct
          description: |
            Zero-shot VLM override. Defaults to the server-configured model,
            else `qwen/qwen3-vl-8b-instruct`. Unsupported names are rejected
            with `400`.
        conf:
          type: number
          description: |
            Detection confidence threshold. Defaults to the model's
            `predict_settings[0].conf`, else `0.25`.
          example: 0.25
        iou:
          type: number
          minimum: 0
          maximum: 1
          description: |
            NMS intersection-over-union threshold. Defaults to the model's
            `predict_settings[0].iou`, else `0.7`.
          example: 0.45
        agnostic_nms:
          type: boolean
          description: |
            Class-agnostic NMS. When true, non-max suppression merges
            overlapping boxes across all classes instead of per-class. Defaults
            to the model's `predict_settings[0].agnostic_nms`, else `false`.
            Applies to the `trained` (YOLO) engine only; ignored for zero-shot.
          example: true
        save:
          type: boolean
          default: true
          description: |
            When true (default), the mapped predictions are persisted as a new
            `auto` annotation group on the task (replacing only the unconfirmed
            auto group from the same source).
        explain:
          type: boolean
          default: false
          description: |
            Explainable-AI opt-in. When true (trained engine only), the
            diagnosis lambda is called after the prediction, the response
            gains an `xai` object and the saved annotation group persists it
            (images as media refs + URL snapshots). Adds a second lambda
            round-trip of latency; see the intro.
        explain_parts:
          type: array
          items:
            type: string
            enum:
              [
                class_scores,
                heatmap,
                gradcam,
                feature_maps,
                embeddings,
                confusion_matrix,
                all,
              ]
          description: |
            Subset of XAI parts to compute. Omit (or send empty / `["all"]`)
            for every part. Unknown names are dropped server-side.
        explain_topk:
          type: integer
          default: 5
          minimum: 1
          description: Candidate classes returned per detection in `class_scores`.
        explain_gradcam_detections:
          type: integer
          default: 20
          minimum: 0
          description: |
            Top-K cap on the per-detection Grad-CAM maps
            (`xai.gradcam.per_detection`), highest-confidence first. Omitted,
            the top 20 detections get their own map — each is a cheap
            head-only backward, and the lambda sheds lowest-confidence maps
            only if the response exceeds its ~6 MB payload limit. 0 disables
            per-detection maps.
        explain_embed_method:
          type: string
          enum: [auto, umap, tsne, pca]
          default: auto
          description: |
            2D projection for the embeddings scatter. `auto` picks by
            detection count (UMAP from 10, t-SNE from 5, PCA below); an
            explicit choice degrades down the umap -> tsne -> pca chain when
            infeasible, and `xai.embeddings.method` reports what actually ran.
        explain_only:
          type: boolean
          default: false
          description: |
            XAI backfill for the EXISTING detections. When true (implies
            `explain`; trained engine + `save` only), the detector is skipped:
            only the diagnosis lambda runs and its XAI is grafted onto the
            task's already-saved unconfirmed auto group of the resolved model
            version — annotation `_id`s and provenance stay untouched, and
            the response carries `explained_existing: true`. Falls through to
            a normal full inference (+explain) when no such group exists.
            Rejected with `400` when the resolved engine is zero-shot
            (explicit `engine: zero_shot`, or `auto` falling back to it) —
            a VLM cannot be diagnosed, and falling through would re-detect
            and replace the group instead of explaining it.
        config:
          $ref: "#/components/schemas/SceneMathConfig"
        company_namespace:
          type: array
          items: { type: string }
          description: Optional tenant namespace override for SDK callers. Not read by this endpoint — the namespace is taken from the caller's token.
    SceneMathConfig:
      type: object
      description: |
        Optional scene-engine tuning for AR session frames. Omit to use the
        defaults. Inference honours the depth-sampling knobs, the dims
        reclassifier and the size gate; any other `SceneMathConfig` key
        (clustering, plane merge, shelf composition, point cloud, RANSAC — the
        session-analysis knobs) is accepted and ignored here.
      properties:
        min_depth_m:
          type: number
          default: 0.05
          description: Ignore depth samples below this (meters).
        max_depth_m:
          type: number
          default: 6
          description: Ignore depth samples above this (meters).
        conf_threshold:
          type: number
          default: 1
          description: Minimum ARKit depth-confidence (0/1/2) to keep a pixel.
        front_percentile:
          type: number
          default: 30
          description: Percentile of bbox depths to take (front-biased).
        shelf_tolerance_m:
          type: number
          default: 0.3
          description: Depth-gate slack beyond `distance_to_shelf_m`.
        reclassify_labels:
          type: boolean
          default: false
          description: Dims reclassifier master switch — re-label a placed detection to a sibling label (same label group) whose expected dims fit its measured size better.
        reclassify_keep_dev:
          type: number
          default: 0.15
          description: Fit deviation at/below which the detected label is kept outright.
        reclassify_target_dev:
          type: number
          default: 0.15
          description: A sibling must fit within this deviation to steal the detection.
        reclassify_min_margin:
          type: number
          default: 0.06
          description: Base margin the sibling's fit must beat the original's by.
        reclassify_conf_margin_scale:
          type: number
          default: 0.5
          description: "Margin × (1 + scale · detector_conf); 0 = ignore confidence."
        reclassify_weight_scale:
          type: number
          default: 1.0
          description: Weight of the size mismatch in the fit score.
        reclassify_weight_aspect:
          type: number
          default: 0.5
          description: Weight of the aspect mismatch in the fit score.
        reclassify_min_depth_confidence:
          type: number
          default: 0.5
          description: Skip detections whose depth confidence is below this.
        size_gate:
          type: boolean
          default: true
          description: "Size gate master switch — flag detections whose measured size exceeds the label's expected dims beyond the allowance (`size_rejected` + `size_reject_detail`)."
        size_gate_dims_allowance:
          type: number
          default: 0.35
          description: "Per-axis allowance: reject when measured w or h > (1 + this) × expected."
        size_gate_area_allowance:
          type: number
          default: 0.35
          description: "Area allowance: reject when measured w·h > (1 + this) × expected area."
        size_gate_min_depth_confidence:
          type: number
          default: 0.5
          description: Skip gating detections whose depth confidence is below this.
      additionalProperties: true
    Box:
      type: object
      description: "Normalized box stored YOLO-style: `x1` = center x, `y1` = center y, `x2` = width, `y2` = height."
      properties:
        x1: { type: number }
        y1: { type: number }
        x2: { type: number }
        y2: { type: number }
    Annotation:
      type: object
      description: A mapped detection (same schema as a task annotation). World-dim fields appear only for placed AR frames.
      properties:
        _id:
          type: string
          description: "Present only on an `explained_existing` response (the stored group's annotations are echoed)."
        box:
          $ref: "#/components/schemas/Box"
        confidence:
          type: number
        label_id:
          type: string
        label_state:
          type: string
          enum: [auto, manual]
        placed_box:
          allOf:
            - $ref: "#/components/schemas/Box"
          description: Snapshot of the box geometry placement ran for (set whenever placement was attempted on a session frame).
        ignore_reason:
          type: string
          enum:
            - no_pose
            - no_intrinsics
            - no_depth
            - empty_depth_region
            - insufficient_depth_pixels
            - behind_shelf
          description: |
            Present when an AR session-frame detection could NOT be placed in
            world coordinates — names the exact failing stage. Unset for placed
            detections and for non-session tasks.
        world_position:
          type: object
          description: Back-projected centroid in world coordinates, metres (placed only).
          properties:
            x: { type: number }
            y: { type: number }
            z: { type: number }
        world_size:
          type: object
          description: "Physical front-face size, centimetres (`w` = world-horizontal, `h` = vertical; placed only)."
          properties:
            w: { type: number }
            h: { type: number }
        depth_at_center:
          type: number
          description: Sampled depth at the box (meters, placed only).
        depth_confidence:
          type: number
          description: Normalized 0..1 depth confidence (placed only).
        placement_confidence:
          type: number
          description: detection_conf × depth_conf × tracking (placed only).
        original_label:
          type: string
          description: Set when the dims reclassifier moved this detection to a sibling label — the label the detector originally produced.
        reclassification_reason:
          type: string
          enum: [dims_match_sibling, group_consensus, manual]
        size_rejected:
          type: boolean
          description: Flagged by the size gate when the measured size exceeded the label's expected dims beyond the allowance.
        size_reject_detail:
          type: object
          properties:
            exceeded: { type: string, enum: [width, height, area] }
            measured_w_cm: { type: number }
            measured_h_cm: { type: number }
            expected_w_cm: { type: number }
            expected_h_cm: { type: number }
            ratio:
              type: number
              description: measured / (expected × (1 + allowance)) for the tripped check.
        cluster_id:
          type: string
          description: "Present only on an `explained_existing` response — the analysis object this stored detection was clustered into."
    Prediction:
      type: object
      description: |
        One raw detector result (`images[0].results[]` of the Ultralytics
        lambda; zero-shot replies are normalized to the same shape with
        `class: -1`). Passed through unchanged.
      properties:
        name: { type: string }
        class: { type: integer }
        confidence: { type: number }
        box:
          type: object
          properties:
            x1:
              { type: number, description: Pixel corner on the original image. }
            y1: { type: number }
            x2: { type: number }
            y2: { type: number }
            x_center: { type: number, description: Normalized 0..1. }
            y_center: { type: number }
            width: { type: number }
            height: { type: number }
      additionalProperties: true
    InferenceResult:
      type: object
      properties:
        task_id:
          type: string
        engine:
          type: string
          enum: [trained, zero_shot]
          description: The detector that actually ran.
        zero_shot_model:
          type: string
          nullable: true
          description: The VLM used, when `engine` is `zero_shot`.
        model:
          type: string
          nullable: true
          description: "`null` for zero-shot runs."
        model_version:
          type: string
          nullable: true
        image_url:
          type: string
        shape:
          type: array
          items: { type: number }
          nullable: true
          description: |
            [height, width] of the inferred image (detector-reported). `null`
            on an `explained_existing` response — no detector ran.
        conf:
          type: number
        saved:
          type: boolean
        explained_existing:
          type: boolean
          description: |
            Present (true) when `explain_only` grafted the XAI onto the
            existing annotation group — no detector ran: `predictions` is
            empty, `shape` / `placed_count` / `unplaced_count` / `unmapped`
            are null, and `annotations` echoes the existing group's
            annotations.
        annotations_count:
          type: number
        placed_count:
          type: number
          nullable: true
          description: Annotations placed in world coordinates (0 for non-session tasks).
        unplaced_count:
          type: number
          nullable: true
          description: Annotations left 2D-only — each carries an `ignore_reason`.
        unmapped:
          type: number
          nullable: true
          description: Predictions that could not be mapped to a label.
        annotations:
          type: array
          items:
            $ref: "#/components/schemas/Annotation"
        predictions:
          type: array
          items:
            $ref: "#/components/schemas/Prediction"
          description: Raw lambda predictions (passthrough).
        xai:
          nullable: true
          allOf:
            - $ref: "#/components/schemas/Xai"
          description: |
            Raw DIAGNOSIS-lambda Explainable-AI payload (passthrough) —
            present only when the request set `explain` / `explain_only` on
            the trained engine and the diagnosis lambda succeeded, else `null`.
        task:
          type: object
          nullable: true
          description: "The saved `ai-object-detection-task` document when the group (or grafted XAI) was persisted, else `null`."
    Xai:
      type: object
      description: |
        Explainable-AI results, all derived from the SAME forward pass as the
        prediction. Any part can be missing — its failure reason is then
        appended to `notes` (XAI never fails the prediction). `class_scores`
        and `embeddings.points` reference detections by `index` into
        `predictions` and are self-describing (box/name/confidence), because
        unmapped predictions never become annotations.
      properties:
        parts:
          type: array
          items: { type: string }
          description: The parts that were requested.
        class_names:
          type: object
          additionalProperties: { type: string }
          description: "YOLO class index (stringified) to class name, from the model weights."
        notes:
          type: array
          items: { type: string }
          description: Human-readable reasons for any part that could not be produced.
        class_scores:
          type: array
          items:
            $ref: "#/components/schemas/XaiClassScore"
        heatmap:
          $ref: "#/components/schemas/XaiHeatmap"
        gradcam:
          $ref: "#/components/schemas/XaiGradcam"
        feature_maps:
          type: array
          items:
            $ref: "#/components/schemas/XaiFeatureMap"
        embeddings:
          $ref: "#/components/schemas/XaiEmbeddings"
        confusion_matrix:
          nullable: true
          allOf:
            - $ref: "#/components/schemas/XaiConfusionMatrix"
    XaiClassScore:
      type: object
      description: |
        Per-detection class-score inspection — the full class-score row of the
        pre-NMS anchor that produced the detection, top-k. A near-tie between
        the top candidates means the model could not distinguish those classes.
      properties:
        index:
          type: integer
          description: "Position of the detection in `predictions`."
        class:
          type: integer
          description: Predicted YOLO class index.
        name:
          type: string
          description: Predicted class name.
        confidence:
          type: number
        box:
          type: object
          description: Detection box in original-image pixels (xyxy).
          properties:
            x1: { type: number }
            y1: { type: number }
            x2: { type: number }
            y2: { type: number }
        candidates:
          type: array
          description: Top-k candidate classes of the winning anchor, best first.
          items:
            type: object
            properties:
              class: { type: integer }
              name: { type: string }
              score: { type: number }
        match_iou:
          type: number
          description: |
            IoU between the detection box and the matched pre-NMS anchor box —
            values near 1 mean the anchor match is exact.
    XaiHeatmap:
      type: object
      description: |
        EigenCAM activation heatmap ("where did the model look") sized to the
        original image aspect. The PNG is RGBA with alpha = activation, so the
        dashboard overlays it directly on the photo with an opacity slider.
      properties:
        png_base64:
          type: string
          description: Base64 RGBA PNG (JET-colored, alpha-carrying).
        method:
          type: string
          enum: [eigencam]
        layers:
          type: string
          description: Which activations were used (detect-head inputs, averaged).
        width: { type: integer }
        height: { type: integer }
    XaiGradcam:
      type: object
      description: |
        TRUE gradient-based Grad-CAM ("what evidence drove the detections").
        A second, gradient-enabled forward pass backpropagates the summed
        confident class scores and weights each activation channel by its
        pooled gradient — class-discriminative, unlike EigenCAM. Same overlay
        contract: RGBA PNG with alpha = activation. Computed by the diagnosis
        lambda only (the backward pass needs its memory budget).
      properties:
        png_base64:
          type: string
          description: Base64 RGBA PNG (JET-colored, alpha-carrying).
        method:
          type: string
          enum: [gradcam]
        layers:
          type: string
        target:
          type: string
          description: What was backpropagated (the Grad-CAM objective).
        imgsz:
          type: integer
          description: |
            Input size of the gradient pass — capped (default 960) so the
            backward pass stays within the diagnosis lambda's memory,
            independent of the inference imgsz.
        width: { type: integer }
        height: { type: integer }
        per_detection:
          type: array
          description: |
            One map PER DETECTION ("why did the model call THIS box this
            class") — each backpropagates ONLY the matched anchor's score for
            the detection's own class, with pixel-wise (LayerCAM-style)
            gradient weighting so the evidence localizes to that instance
            instead of lighting up every look-alike. The top
            `explain_gradcam_detections` detections by confidence get a map
            (default 20; `per_detection_note` reports truncation; the lambda
            also sheds lowest-confidence maps if the response exceeds its
            ~6 MB payload limit). `index` references the `predictions` array
            like class_scores.
          items:
            type: object
            properties:
              index: { type: integer }
              class: { type: integer }
              name: { type: string }
              confidence: { type: number }
              box:
                type: object
                description: Detection box in original-image pixels (xyxy).
                properties:
                  x1: { type: number }
                  y1: { type: number }
                  x2: { type: number }
                  y2: { type: number }
              match_iou:
                type: number
                description: |
                  IoU between the detection and the anchor of the gradient
                  pass it was matched to — low values mean the (size-capped)
                  gradient pass barely saw this object and the map is weak
                  evidence.
              png_base64:
                type: string
                description: Base64 RGBA PNG (JET-colored, alpha-carrying).
              method: { type: string, enum: [gradcam] }
              layers: { type: string }
              target: { type: string }
              width: { type: integer }
              height: { type: integer }
        per_detection_note:
          type: string
          description: Present when detections were truncated to the top-K.
    XaiFeatureMap:
      type: object
      description: One per-stage feature-map grid from ultralytics visualize=True.
      properties:
        stage:
          type: string
          description: "Network stage, e.g. `stage12_C2f`."
        jpg_base64:
          type: string
          description: Base64 JPEG of the stage's channel grid, downscaled.
    XaiEmbeddings:
      type: object
      description: |
        Per-detection embeddings (pooled from the finest detect-head input
        scale) projected to 2D for a scatter plot — detections the model sees
        as similar land near each other.
      properties:
        method:
          type: string
          enum: [umap, tsne, pca, none]
          description: |
            The projection that actually ran. With `explain_embed_method:
            auto` — UMAP for 10+ detections, t-SNE for 5-9, PCA for 2-4,
            `none` for a single point; explicit requests degrade down the
            umap -> tsne -> pca chain when infeasible.
        requested_method:
          type: string
          enum: [auto, umap, tsne, pca]
          description: What the caller asked for (differs from `method` on fallback).
        layer:
          type: string
        embedding_dim:
          type: integer
          description: Dimensionality of the pooled embedding before projection.
        points:
          type: array
          items:
            $ref: "#/components/schemas/XaiEmbeddingPoint"
    XaiEmbeddingPoint:
      type: object
      properties:
        index:
          type: integer
          description: "Position of the detection in `predictions`."
        class: { type: integer }
        name: { type: string }
        confidence: { type: number }
        x:
          type: number
          description: Projected coordinate, min-max normalized to [0, 1].
        y:
          type: number
          description: Projected coordinate, min-max normalized to [0, 1].
    XaiConfusionMatrix:
      type: object
      description: |
        Training-time confusion-matrix images of the model version (computed on
        the validation set during training — inference has no ground truth to
        build one from). URLs point at the stored media artifacts.
      properties:
        url:
          type: string
          nullable: true
          description: Rendered confusion-matrix image (absolute counts).
        normalized_url:
          type: string
          nullable: true
          description: Row-normalized variant.
        source:
          type: string
          enum: [training_artifacts]
