# Test Prompt 003 — Group 3: Agent Behaviors (Config Enforcement)

**Phase:** 3
**Focus:** Verify that agent configuration is actually enforced at runtime — tool restrictions, mode gating, iteration limits, and live config reload
**Agents running tasks:** Yes

---

## Your Role

You are a test engineer executing Phase 3 of the VeilCLI test suite. This phase tests whether VeilCLI actually honors agent configuration. These are silent failure territory: the HTTP response is 200, the task status is `finished`, but the agent did something it wasn't supposed to — used a forbidden tool, ran in a disabled mode, ignored iteration limits. A basic status-code check would pass all of these while the system is broken.

The primary verification mechanism in this phase is **task event inspection** — `GET /tasks/:id/events` — not the agent's text output.

---

## Environment

- **VeilCLI source:** `/home/ixi/khacloud/drive/Plugins/VeilCli`
- **How to invoke:** `veil` (globally linked)
- **Reference auth.json:** `/home/ixi/khacloud/drive/Plugins/VeilCli/.veil/auth.json`
- **API DOCS:** `/home/ixi/khacloud/drive/Plugins/VeilCli/docs/api/` — check them to ensure your tests are right
- **Agent schema & examples:** `/home/ixi/khacloud/drive/Plugins/VeilCli/schemas/agent.json` and `/home/ixi/khacloud/drive/Plugins/VeilCli/examples/`

---

## ⚠️ Non-Blocking Execution Rule

**Any command that takes time must be run in the background and monitored by polling — never block waiting for it.**

This applies to:
- Starting the server (`veil start` → run in background, poll `GET /health` in a loop until ready)
- Running test scripts — if a script can hang, run it with a timeout or in background and tail its output
- Waiting for tasks to finish → always poll `GET /tasks/:id` in a loop with a sleep interval, never a blocking wait

If a command blocks and gets stuck, you get stuck too. Background + poll is the pattern for everything time-sensitive in this phase.

---

## Step 1 — Create the Test Workspace

Create the workspace at:
```
/home/ixi/khacloud/drive/Plugins/VeilCli_TESTS/workspace-test-003
```

`.veil/settings.json`:
- Port **5353**
- A test secret of your choice
- Permissive execution-level permissions (allow all) — the restrictions being tested are at the agent config level, not the global settings level
- Reasonable iteration and duration limits for a test environment

Copy the reference `auth.json` into `.veil/auth.json`.

---

## Step 2 — Create the Test Agents

You need four agents. Check the schema and docs for exact valid fields. Pay close attention to the difference between:
- `tools` — a whitelist of tool names the LLM is shown (if set, LLM only sees these tools)
- `disallowedTools` — a blacklist that hides specific tools from the LLM
- `permissions.allow/deny` — execution-level gatekeeping (different layer from LLM visibility)

Read the docs and guide on permissions and tool configuration before creating agents — the distinction between these fields matters for writing correct test configs.

**Agent 1 — `restricted-agent`**
- Task mode enabled
- Configure it to explicitly deny/hide `bash` from the LLM (use the appropriate field)
- Has other tools available (file I/O tools) so it can still do useful work
- Keep memory disabled

**Agent 2 — `readonly-agent`**
- Task mode enabled
- Whitelist only `read_file` as an available tool (LLM sees nothing else)
- Keep memory disabled

**Agent 3 — `chat-disabled-agent`**
- Task mode enabled
- Chat mode explicitly disabled (`modes.chat.enabled: false`)
- Keep memory disabled

**Agent 4 — `iteration-agent`**
- Task mode enabled
- Set `maxIterations` very low (3) at the agent's task mode level
- Has file I/O tools so it can attempt multi-step work
- Keep memory disabled

---

## Step 3 — Start the Server

From inside `workspace-test-003`, start `veil` in the background. Poll `GET http://localhost:5353/health` until it responds (20 second timeout). Fatal failure if it doesn't come up.

---

## Step 4 — Run the Tests

---

### Test 01 — Deny List Enforcement (`disallowedTools`)

**What this catches:** If the deny/hidden tool list isn't applied when loading the agent's tool set, the LLM gets access to tools it shouldn't have. Since the LLM won't call a tool it can't see, this manifests as the forbidden tool appearing in task events — which a passing HTTP status would never reveal.

**Setup:** Use `restricted-agent`. Give it a task that explicitly requires bash — something like "run the shell command `echo BASH_WAS_CALLED` and report what it output."

**Verify:**
- Poll the task to a terminal state.
- `GET /tasks/:id/events` — scan all events. There must be **zero** `tool.start` events with the denied tool's name. If even one appears, the restriction is broken.
- The task outcome (finished or failed) is secondary — what matters is that the forbidden tool was never invoked. The agent should either work around it or explain it can't do the task without that tool.
- Also check `GET /agents/restricted-agent/skills` — the denied tool must not appear in the skills/tools list returned for this agent.

---

### Test 02 — Whitelist Enforcement (`tools` allow list)

**What this catches:** Mirror of Test 01. If the tools whitelist doesn't actually restrict what the LLM sees, the agent can use any available tool regardless of config.

**Setup:** Use `readonly-agent` (only `read_file` whitelisted). First, place a file in the workspace with known content. Then give the agent a task that requires both reading AND writing: "Read the file at [path], then write its content in reverse to a new file called reversed.txt."

**Verify:**
- Poll to terminal state.
- `GET /tasks/:id/events` — `write_file` must never appear in any `tool.start` event. The agent may read the file successfully, but it cannot write.
- Check that `reversed.txt` does **not** exist on disk. If it does, the whitelist failed.
- `GET /agents/readonly-agent/skills` — only `read_file` (and possibly system tools) should appear.

---

### Test 03 — Mode Enforcement

**What this catches:** If `modes.chat.enabled: false` doesn't actually gate the chat endpoint, disabled modes are meaningless. This is a simple but important contract.

**Setup:** Use `chat-disabled-agent`.

**Verify:**
- `POST /agents/chat-disabled-agent/chat` — must return an error response (not HTTP 200). Check the docs for what status code and error shape VeilCLI returns for disabled modes.
- `POST /agents/chat-disabled-agent/task` — must still work (task mode is enabled). Create a simple task and verify it gets a 202 with a taskId. The agent isn't broken — just its chat endpoint is gated.

---

### Test 04 — `maxIterations` and `onExhausted` Enforcement

**What this catches:** Without iteration limits, a broken or confused agent loop runs forever and burns tokens. `onExhausted` determines the final behavior — if the wrong state is reached, downstream systems (callers polling for `waiting` or `failed`) behave incorrectly.

**Setup:** Use `iteration-agent` (maxIterations: 3). Give it a task designed to require many steps — something like: "Research and write a comprehensive 10-section report on Node.js best practices, writing each section to its own file." This is intentionally too big to finish in 3 iterations.

**Test onExhausted: "fail" behavior:**
- Submit the task with the agent configured for `onExhausted: "fail"` (set this at test time, either via agent config or task parameters — check docs for how to override per-task).
- Poll to terminal state.
- Final status must be `failed`. Not `finished`, not `waiting` — `failed`.
- The `iterations` field on the task record must be ≤ 3.

**Test onExhausted: "wait" behavior:**
- Reconfigure or create a variant of the agent with `onExhausted: "wait"`.
- Submit the same large task.
- Poll until status is `waiting` (not `finished` or `failed`). This confirms the agent paused instead of failing.
- `POST /tasks/:id/respond` with a message like "Continue where you left off." — verify the task resumes (status goes back to `processing` then eventually terminates).
- Check docs for the exact `respond` endpoint payload format.

---

### Test 05 — Hot-Reload from Disk

**What this catches:** `POST /agents/:name/reload` must actually re-read the agent's `agent.json` from disk. If it's a no-op, live config changes never take effect without a server restart.

**Setup:** Use any of the agents created above (e.g. `restricted-agent`). Note its current `description` and `temperature` values.

**Verify:**

*Change takes effect after reload:*
- Directly edit `restricted-agent`'s `agent.json` on disk — change `description` to something recognizable like `"HOT_RELOAD_TEST_VALUE"` and change `temperature` to an unusual value like `0.42`.
- `GET /agents/restricted-agent` — verify the OLD values are still returned (confirms the API serves cached config, not live disk reads).
- `POST /agents/restricted-agent/reload` — expect HTTP 200.
- `GET /agents/restricted-agent` — now the description and temperature must match what you wrote to disk. If they don't, reload is broken.

*Change does NOT take effect without reload:*
- Make another disk edit (change description again to a different recognizable value).
- Do NOT call reload.
- `GET /agents/restricted-agent` — must still show the previous value (the one from the first reload). This confirms the server doesn't auto-watch for file changes — reload is required.

---

## Step 5 — Stop the Server

Stop the server cleanly from the workspace directory.

---

## Step 6 — Write the Summary

Create `summary.md` in `workspace-test-003/`:

```markdown
# Test Run Summary — workspace-test-003
**Date:** <date>
**Port:** 5353
**Phase:** Group 3 — Agent Behaviors (Config Enforcement)
**Test Path:** /home/ixi/khacloud/drive/Plugins/VeilCli_TESTS/workspace-test-003/

## Results

| Test | Status | Notes |
|------|--------|-------|
| Test 01: Deny List Enforcement | PASS / FAIL | |
| Test 02: Whitelist Enforcement | PASS / FAIL | |
| Test 03: Mode Enforcement | PASS / FAIL | |
| Test 04a: maxIterations + onExhausted fail | PASS / FAIL | |
| Test 04b: maxIterations + onExhausted wait | PASS / FAIL | |
| Test 05: Hot-Reload from Disk | PASS / FAIL | |

## Failures

For each FAIL:
- Which test and specific assertion
- Expected vs actual
- Whether it looks like a VeilCLI bug or test setup issue

## Observations

Unexpected behavior, config fields that didn't behave as documented, or anything worth flagging.
```
