# Setup guide (end-to-end)

`pi-model-swap` is the **last 10%** of a local llama-swap stack. It does not install, configure,
start, stop, or kill llama.cpp, llama-swap, or any model server — it reads their state and moves
*pi's* binding between models. This guide is the stack underneath it.

**Platform: Linux only.** The census uses `ss -tanp` / `ss -ltnp` (iproute2) and `/proc/<pid>/stat`
(procfs). On macOS/Windows the census cannot attribute sockets and every swap refuses.

## The chain

```
llama.cpp llama-server  (one process per model, its own port)
        ^  proxy: http://127.0.0.1:${PORT}
llama-swap controller   (one endpoint, e.g. :1235; starts/stops the servers above)
        ^  baseUrl http://127.0.0.1:1235/v1
pi provider entry       (models.json: provider id + model ids)
        ^  provider: "<that provider id>"
pi-model-swap.json      (this extension's config)
```

Every id must line up: **roster id == llama-swap model id == models.json model id == the id you
pass to `/swap`.**

## 1. llama.cpp

Any `llama-server` build that supports your models. Note the flags that matter for swapping:

- `--host 127.0.0.1 --port ${PORT}` — the port comes from llama-swap's proxy macro
- `--chat-template-kwargs '{"reasoning_effort": "…"}'` — thinking levels only work if the chat
  template accepts `reasoning_effort`
- `--ctx-size`, `-ngl`, quantisation: your choice; cold-start cost scales with model size

## 2. llama-swap

Tested against llama-swap v262. Config file (point `llamaSwapConfigPath` at it):

```yaml
startPort: 1236          # ${PORT} base — the extension recomputes startPort + roster slot index
healthCheckTimeout: 900  # seconds; a model start can take minutes
globalTTL: 0             # keep models loaded; TTL eviction mid-session breaks swaps
logLevel: info
```

Bind to localhost and **do not set `apiKeys`**: the controller is loopback-only, and llama-swap's
own docs note `apiKeys` is not a firewall substitute (it also locks the web UI out). No secret in
any file. Remote machine? Tunnel it:

```bash
ssh -L 1235:127.0.0.1:1235 user@your-box
```

Run it with `--listen 127.0.0.1:1235`.

## 3. The roster (the extension's port map)

The roster is the **only** source of upstream ports. Each model needs a `proxy:` line:

```yaml
models:
  orchestrator:
    cmd: /path/to/llama-server -m /path/to/model.gguf --alias orchestrator \
         --host 127.0.0.1 --port ${PORT} --chat-template-kwargs '{"reasoning_effort":"medium"}'
    proxy: http://127.0.0.1:${PORT}
  coder:
    cmd: …
    proxy: http://127.0.0.1:${PORT}
```

**Order matters.** `${PORT}` resolves to `startPort + index-in-roster`: the first model is
`startPort`, the second `startPort + 1`, and so on. The extension parses `proxy:` and applies the
same rule, so a reordered roster silently points the census at the wrong ports. If you use
explicit ports instead of `${PORT}`, the extension reads them literally.

## 4. pi's `models.json` provider

`provider` in the extension config must be the key of this entry, and the model ids must match the
roster ids:

```json
{
  "providers": {
    "llama-swap": {
      "baseUrl": "http://127.0.0.1:1235/v1",
      "api": "openai-completions",
      "apiKey": "not-needed",
      "compat": {
        "supportsUsageInStreaming": true,
        "maxTokensField": "max_tokens",
        "thinkingFormat": "chat-template",
        "chatTemplateKwargs": { "reasoning_effort": { "$var": "thinking.effort" } }
      },
      "models": [
        {
          "id": "orchestrator",
          "name": "Orchestrator",
          "reasoning": true,
          "thinkingLevelMap": {
            "off": null, "minimal": null, "low": "low",
            "medium": "medium", "high": "xhigh", "xhigh": "xhigh", "max": "xhigh"
          },
          "compat": { "supportsReasoningEffort": true, "thinkingFormat": "chat-template" },
          "contextWindow": 130000,
          "maxTokens": 65536
        }
      ]
    }
  }
}
```

`thinkingLevelMap` is how a model declares which levels it honours — the extension does **not**
validate levels per model, so map the levels you intend to use.

## 5. `~/.pi/agent/pi-model-swap.json`

All keys are required (no fallbacks; with no config the extension refuses and says so):

```json
{
  "controllerUrl": "http://127.0.0.1:1235",
  "controllerPort": 1235,
  "rosterPath": "/path/to/llama-swap/models.yaml",
  "llamaSwapConfigPath": "/path/to/llama-swap/config.yaml",
  "startPort": 1236,
  "provider": "llama-swap",
  "sessionDefault": "orchestrator",
  "strataPort": 1301,
  "verifyTimeoutMs": 960000,
  "probeTimeoutMs": 5000,
  "claimStalenessS": 1800,
  "boundaryRestore": true
}
```

- `verifyTimeoutMs` ≥ `healthCheckTimeout` × 1000 + margin (900 s + 60 s = 960 000 ms). The verify
  request is what starts a cold model; a shorter client bound fails a swap that would have worked.
- `probeTimeoutMs` stays small (5 s) — a status probe must never inherit the 16-minute verify bound.
- `strataPort` is **corroboration only**: the plan prints `:<port>/slots` if your model server
  serves it. If it does not, the line says `corroboration skipped` and nothing else changes.
- `sessionDefault` is the restore target and must exist in the provider's model list.

## 6. Verify, in this order

```bash
/swap --status              # config path + controller + roster + startPort + my claim
/swap --dry-run <model>     # plan only: resident ids, eviction set E, census, would: steps
/swap <model>               # live swap
/swap --restore             # back to sessionDefault (announce at a boundary, execute at the next)
```

`/swap --status` must show your config path. If it shows a path you did not write,
`PI_MODEL_SWAP_CONFIG` is set somewhere.

## Reading the refusals

| Report line | What it means |
|---|---|
| `config … is unreadable/unparseable` | your JSON is broken — fix it, nothing runs |
| `cannot read …/running` | controller down, wrong port, tunnel closed |
| `body … not the documented shape` | llama-swap version/response shape the census cannot read — eviction set unknown, default-deny |
| `cannot read the roster …` | `rosterPath` wrong — the port map is unknown, default-deny |
| `roster … yielded 0 proxy: entries` | roster has no `proxy:` lines |
| `swap would unload a model a live session is bound to` | another session's claim file is live on a model in `E` |
| `N foreign ESTAB` on a model's port | a process that is neither the model's listener owner nor the controller is connected — default-deny |

## Worked example: Strata + `swift-1.5` (the stack this extension was built for)

The reference deployment is a **hybrid** roster: llama.cpp models *and* a MoE model served by
Strata, with the session's default being the Strata one. This is the case the refusal gate and the
cold-start warning exist for.

```
llama-swap :1235  (controller)
   ├── orchestrator  → llama-server :1236   (${PORT} = startPort + 0)
   ├── subagent      → llama-server :1237
   ├── …             → llama-server :1238 …
   └── swift-1.5     → Strata :1301         (literal proxy port, NOT ${PORT})
            └── server.py --engine strata  → spawns `strata --serve` over pipes
```

**swift-1.5 is not a llama-server process.** Its roster entry starts the Strata python HTTP parent,
which spawns the engine over pipes — call the binary directly and llama-swap and the parent
disagree about who owns the model:

```yaml
  swift-1.5:
    cmd: >
      /path/to/Strata/.venv/bin/python /path/to/Strata/serve/server.py
      --engine strata
      --config /path/to/Strata/strata-swift-iq3_xxs.json
      --port 1301
    proxy: http://127.0.0.1:1301   # literal port: the census reads it as-is (no ${PORT})
    ttl: 0
    unloadTimeout: 90
```

The Strata config behind it (`strata-swift-iq3_xxs.json`) is what makes the cold start expensive:
a large MoE pack with an expert cache, MTP draft, 131 072-token context and int8 KV.

```json
{
  "exe": "/path/to/Strata/engine/strata",
  "port": 1301,
  "model_name": "swift-1.5-iq3_xxs",
  "args": ["--pack", "/path/to/packs/swift-iq3_xxs",
           "--native", "/path/to/Swift-…-IQ3_XXS-00001-of-00002.gguf",
           "--expert-profile", "/path/to/expert-profile.bin", "--expert-cache", "auto",
           "--mtp", "/path/to/mtp/rt", "--max-context", "131072",
           "--kv", "int8", "--kv-resident", "32768"]
}
```

pi's `models.json` entry for it (note `contextWindow` matching `--max-context`, and the level map):

```json
{
  "id": "swift-1.5",
  "name": "Swift 1.5 IQ3_XXS (Strata, local)",
  "reasoning": true,
  "thinkingLevelMap": { "off": null, "minimal": null, "low": "low", "medium": "medium",
                        "high": "xhigh", "xhigh": "xhigh", "max": "xhigh" },
  "input": ["text"],
  "contextWindow": 131072,
  "maxTokens": 32768
}
```

Config for that stack — `examples/pi-model-swap.strata.json`, copy it and fix the paths:

```json
{
  "controllerUrl": "http://127.0.0.1:1235",
  "controllerPort": 1235,
  "rosterPath": "/path/to/llama-swap/models-harness/harness.yaml",
  "llamaSwapConfigPath": "/path/to/llama-swap/config-harness.yaml",
  "startPort": 1236,
  "provider": "llama-swap",
  "sessionDefault": "swift-1.5",
  "strataPort": 1301,
  "verifyTimeoutMs": 960000,
  "probeTimeoutMs": 5000,
  "claimStalenessS": 1800,
  "boundaryRestore": true
}
```

`strataPort: 1301` is the corroboration line only — the plan prints
`:1301/slots (model server) [{"id":0,"n_ctx":131072,"is_processing":false}]` so you can see the
engine is idle before you swap. It is never a port source; the ports come from the roster.

**Why the gate matters here.** `swift-1.5` is the session default, so a swap to any other model has
`E = {swift-1.5}` and the plan warns: *"This swap may evict swift-1.5. Recovery is a full cold start
(minutes), not a fast rollback."* If a **live** session is bound to `swift-1.5` (its claim file), the
swap is refused — that is the intended workflow: move your other sessions off Strata (or exit them),
then swap, do the work on the cheap model, and `/swap --restore` back to `swift-1.5` at a turn
boundary. Restoring is the expensive direction; the extension announces it and executes it at the
next boundary (`boundaryRestore: true`).

## Multi-session rule

Claims live in `~/.pi/agent/state/swap-state/<pid>.json`. A swap is refused when its eviction set
contains a model a **live** session is bound to. Before a big swap, move your other sessions off
the model they are using (or exit them) — that is the intended workflow, not an error to route
around.
