Configure .claude/settings.json hooks and start a Claude Code session to see events here in real time.
Agent Recon™ server: http://localhost:3131
Network calls, sensitive file access, privilege escalation, credential exposure, and injection patterns will appear here.
npm install -g agent-recon agent-recon install
The guided installer configures hooks, optional API keys, and optional auto-start service. It detects your platform (Windows, macOS, Linux, WSL) and sets up the appropriate hook script and service template automatically.
Windows users: Install Claude Code first. Disable the Python Store alias and set ExecutionPolicy RemoteSigned — see Troubleshooting for details.
agent-recon start # or: cd server && node start.js
Open http://localhost:3131 in your browser. If TLS is enabled (via mkcert), the server redirects to HTTPS automatically with a green padlock. The server stores all data locally in SQLite — nothing leaves your machine.
Start any Claude Code session. Events appear in the dashboard automatically — the hook scripts forward every tool call, prompt, completion, and agent spawn in real time. No additional configuration is needed once the hooks are installed.
Also use Cursor or GitHub Copilot CLI? Agent Recon can watch those too. Re-run agent-recon install — it auto-detects them and registers their hooks — then fully restart the tool. See the Multi-Tool tab for step-by-step setup and troubleshooting.
The Sessions tab shows one column per active Claude Code session. Each column contains event cards that arrive in real time as Claude Code runs tools. The column header shows the session label, a color stripe, an environment icon indicating the originating platform (e.g. WSL, Windows, macOS, tmux, VS Code), and an accumulated token cost chip once any tokens have been tracked.
Each session gets a persistent label in the format project:hex — for example agent-recon:3a7f. The project name is derived from the git remote repo name, or falls back to the project_name in the hook payload, then the last component of the working directory. The hex suffix is the last 4 characters of the session UUID, extended automatically if needed to keep every label globally unique. Labels are stored in the database and never re-derived, so they stay consistent across server restarts. Each session also gets a unique color stripe so you can tell them apart at a glance.
Each collapsed card leads with a plain-English description of what the agent was doing on that step — a friendly intent phrase (e.g. “Reading a file”, “Installing dependencies”, “Searching the project”). Common commands are recognised by pattern, so git commit reads as “Committing changes” and npm test as “Running tests”; anything unrecognised falls back to a clear generic action like “Ran a terminal command”.
Detail line: directly beneath the phrase, in monospace, the card shows the exact raw argument — the file path, command, search pattern, or URL. It is trimmed to two lines; the complete, un-trimmed value is one click away in the Details panel (below). You can still select and copy the visible text without expanding the card.
Status chip: a small plain-English chip shows the step’s state — Done, Failed, or Waiting — instead of a developer event name. The precise hook and tool (e.g. PreToolUse / Bash) are shown in the Details panel for anyone who wants them.
Details & raw data: click anywhere on a card (or the ▾ Details control on the right) to open a panel with the full value plus the underlying event, tool, and session. Inside that panel, a View raw data (JSON) button reveals the complete raw event — including the working directory — for power users. The card is keyboard-accessible: tab to the Details control and press Enter or Space.
Background badge: events from tools invoked with run_in_background: true (Bash commands or Agent/Task spawns) display an amber BG badge, along with a duration indicator showing elapsed time (for active tasks) or total duration (for completed tasks). This makes it easy to identify which operations were fired asynchronously and track their lifecycle.
Environment, agent & model badges: each card shows a small icon for the environment it came from (e.g. penguin for WSL/Linux, window frame for Windows, terminal prompt for PowerShell, code brackets for VS Code), a badge for the source agent (Claude Code, Cursor, or GitHub Copilot CLI), and — when known — the model in use (e.g. Opus 4.8, Gemini 3.5 Flash). The agent badge reflects the session’s source tool, so it stays consistent even on sub-agent steps; the model badge is resolved per session and shown on every card of that session (it stays hidden only for sessions where no model was ever reported). GitHub Copilot CLI does not include its model in hook events, so Agent Recon reads it from Copilot’s own session files. Hover any badge for details. The session label is not repeated on each card because it already appears in the swimlane header.
Each session column shows a sticky bar at the top of its card list displaying the most recent user prompt (excluding sub-agent prompts). Collapsed (default): shows a Last Prompt: prefix followed by the text truncated to one line. Expanded: click the bar to toggle full text in a scrollable area with whitespace preserved.
Directly below the prompt bar, a green Context bar shows what Claude appears to be operating from in that session — the facts established, the constraints it must respect, the source files / data it has looked at, and the assumptions it seems to be taking for granted. Collapsed (default): a one-line digest of the item counts. Expanded: click the bar to see all four lists in full.
This snapshot is inferred by an LLM reading the session's observed activity — every user prompt, the tools Claude ran and their outcomes, and sub-agent dispatches. Agent Recon never sees Claude's actual context window. Scan it after a prompt: if an assumption is wrong or a constraint is missing, correct Claude in your next message before it acts on the mistake. Requires an Anthropic API key; configure the model and rate limit, or turn the feature off, under Settings → LLM Models → Context Observability.
Change tracking: the bar is recomputed on each user prompt, and each item carries provenance. Items the recompute kept unchanged stay quiet; a freshly inferred item shows a green new badge and a revised one a yellow changed badge, each annotated with the prompt it appeared or changed at (the collapsed digest also shows an N updated count). This lets you see at a glance what shifted since your last prompt rather than re-reading the whole bar. Each recompute is still a fresh inference from the evidence — provenance is matched in afterward — so a stale or hallucinated item is corrected on the next pass rather than accumulating.
Context-window gauge: when the statusline integration is set up, the collapsed bar also shows a stoplight dot and a percentage for how full Claude's context window is — green below 70%, yellow at 70–89%, red at 90%+. Expanded, it adds the token count and a recommendation (e.g. proactively run /compact). Unlike the inferred context above, this percentage is ground truth reported by Claude Code itself. It requires the Agent Recon statusline wrapper installed and Claude Code v2.1.141 or newer (older versions report inaccurate context-window counts); without those the gauge simply does not appear.
The row below the canvas shows category chips — one per event type observed. The number on each chip is the count of matching events. Click a chip to filter the feed to that category. Click the chip again (or the "All" chip) to clear the filter. Tool categories include: Bash ⚡, PowerShell 💠, Write ✍️, Edit ✏️, Notebook 📓, Read 👁️, Glob 📁, Grep 🔎, Web 🌐, Tasks 📋, Cron ⏰, Plan 🗺️, Worktree 🌿, MCP Res 🗄️, System 🧩.
BG badge: when an agent launches a tool call in the background (e.g. a long-running shell with run_in_background), its event card carries a small BG badge alongside the agent and model badges. That’s the honest signal — it marks that the work was started in the background. Agent Recon does not track background completion: a backgrounded launch returns immediately (a shell handle, not the result), and no agent host emits a background-finished hook, so a live “running/finished” count would be misleading. To see what a backgrounded command later produced, the agent re-reads it with a follow-up tool call, which appears as its own event.
The scrolling waveform at the top of the page shows activity level over time. Each active session gets its own labeled row. The waveform rises to the thinking level when Claude is processing a prompt and to the active level when it is executing tools, then drops back to the idle baseline when a turn completes.
Idle persistence: session rows remain visible for up to 15 minutes after the last event, so a session waiting for your next prompt stays on screen between turns. Only truly orphaned sessions — those stuck at the thinking or active level with no new events for 2 minutes — are automatically removed.
Restore on refresh: when the browser loads with the Today's sessions display setting, any session that had activity within the past 15 minutes is automatically restored as a flat idle row. The waveform animates normally from the moment the next live event arrives.
Session labels: each canvas row is labelled with the session's persistent project:hex label (e.g. agent-recon:3a7f). Labels are guaranteed unique across all sessions in the database, so rows are always unambiguous even when multiple sessions share the same project name.
Scrolling and height: the canvas expands as new sessions and sub-agents appear, up to a configurable maximum height (default 25% of the screen). When rows overflow the visible area a subtle scroll indicator appears at the bottom — scroll the canvas strip to reveal additional rows. Click the ▲ button in the top-right corner of the canvas to collapse it entirely; click ▼ to expand it again. The maximum height can be adjusted in Settings → Session Display → Canvas height.
Pinned icons: when a tool call or special state is in progress, a small icon is pinned just to the right of the session label and pulses until the state resolves. The icon tells you what Claude is currently doing:
Pinned icons appear only after their corresponding floating icon has scrolled past the label margin, so the same icon is never shown twice simultaneously.
Every event is classified by a YARA-X rule engine running server-side in under 2 ms. Rules cover 22 detection patterns across injection, credential exposure, privilege escalation, network activity, and sensitive file access — including 7 credential patterns derived from the Gitleaks open-source ruleset (GitHub PATs, Anthropic/OpenAI keys, AWS access keys, JWTs, and private keys). If an Anthropic API key is configured, a second LLM enrichment pass runs asynchronously, batching events for up to 60 seconds after session activity, identifying correlated event chains with agentic context awareness. The LLM model used is configurable in Settings → LLM Models; the default is claude-opus-4-6 for best threat correlation accuracy.
The LLM enrichment pass uses session context to reduce false positives and produce more accurate risk levels:
Calibration rules: a security event within 60 seconds of a user prompt or sub-agent dispatch that authorized similar activity is capped at Medium unless the action materially exceeds stated scope. Critical-level threats (code injection from external input, credential exfiltration, confirmed prompt injection) are never auto-downgraded by intent signals — they require explicit evidence of authorization to reduce from critical.
For best calibration accuracy, use claude-opus-4-6 as the Security LLM model (configurable in Settings). Haiku may miss causal links between user intent and multi-step sub-agent actions.
Semantic compression: For large sessions, consecutive runs of 3+ identical low/medium-risk events (same tool, category, and target) are grouped into a single compressed entry before sending to the LLM. Critical and high-risk events are never compressed. This improves token efficiency while preserving accuracy — all original event IDs are still tracked and stored.
| Level | Meaning |
|---|---|
| 🔴 Critical | Immediate confirmed threat: inline code execution (eval, python -c, node -e), active credential exfiltration to an external host, successful prompt injection redirecting agent behavior. |
| 🟠 High | Serious concern: pipe-to-shell injection, privilege escalation, credential exposure, exfiltration patterns, writes to sensitive paths. |
| 🟡 Medium | Noteworthy but often legitimate in agentic context: reading SSH keys, auth tooling, secret references in prompts, bash accessing sensitive paths. |
| 🟢 Low | Routine events with security context: outbound HTTP fetches (WebFetch), generic network tool usage, agent spawning sub-agents. |
Each security event in the expanded chain view shows up to two risk levels:
Example: a file read classified as 🔴 critical by the YARA engine (sensitive path pattern) may show 🔴 critical → 🟢 low after the LLM determines it was a routine sub-agent output retrieval consistent with the authorized task. The chain card's overall risk level reflects the worst event in the chain — individual event assessed risks give finer-grained context.
| Category | Description |
|---|---|
inject | Shell injection — Critical: eval/exec of literal strings, inline python/node/perl -c/-e. High: pipe to bash/sh, large command substitutions. |
secret | Credentials or API keys referenced in commands, files, or prompts |
data | Potential exfiltration to external destinations (curl POST, dd, base64 encode) |
exec | Privilege escalation attempts (sudo, chmod 777, chown root, setuid) |
auth | Authentication tool usage (ssh-keygen, gpg, openssl genrsa, passwd) |
net | SSH/SCP/SFTP connections, outbound HTTP/network requests (curl, wget, nc) |
file | Access to sensitive file paths (~/.ssh/, .env, /etc/shadow, .pem, .key) |
secret (write) | Writing credential or secret content to a file |
agent-spawn | Agent spawning a sub-agent via the Task or Agent tool — logged at low risk; LLM enrichment may elevate if the task description is anomalous. |
proc-exec | OS-level process spawns tracked by the process monitor (if enabled) |
chain | LLM-identified correlated event sequence across multiple tool calls |
pii | Personally identifiable information detected in event fields — see PII Detection section below |
After the LLM enrichment pass completes, chain cards appear in the Security tab with an AI badge. Chain types include: exfiltration, injection, escalation, credential-harvest, agent-spawn, prompt-inject, recon, and normal-operation (benign correlated events). The LLM framing is tuned to distinguish legitimate AI agent activity (npm install calling curl) from genuine threats.
Each session chain card shows a Live or Final status pip:
Stop hook).SessionEnd hook fired) or the session was auto-finalized after 4 hours of inactivity. The analysis shown is the definitive pass over the complete session.Each security card carries a small status pip reflecting LLM analysis progress:
| Pip | Meaning |
|---|---|
⋯ LLM | YARA-flagged; awaiting LLM analysis pass |
✓ LLM | LLM analyzed — no specific verdict; original risk level stands |
✓ Safe | LLM confirmed benign — card hidden by default (click Show all to reveal) |
↓ LLM | LLM downgraded the risk level — updated risk chip shown on card |
🔗 | Part of an LLM-identified threat chain (chain card) |
Once LLM analysis completes, each analyzed card shows a color-coded recommendation badge alongside the pip:
| Badge | Meaning |
|---|---|
| Safe to Ignore | LLM determined the event is benign in this agentic context |
| Review Recommended | Context-dependent; worth a quick check before proceeding |
| Immediate Action | LLM identified a confirmed or high-confidence threat |
Expand a card (click it) to see the expanded analysis detail — including the LLM's reason, any risk reclassification from/to, and for chain cards the full LLM summary and chained event IDs.
The Show all button in the Security tab header reveals LLM-confirmed safe events (those with a ✓ Safe pip) that are hidden by default. Click it again (Actionable only) to re-hide them, keeping the view focused on events that may still need attention.
Security cards marked with a 🤖 robot badge originated from a sub-agent — a nested Claude instance spawned via the Task tool. This badge provides context that the activity was delegated (indirect) rather than initiated directly by the top-level session, which is relevant when assessing threat chains in multi-agent workflows.
The LLM enrichment pass is sub-agent aware: when a sub-agent performs a security-relevant action, the task description the main agent used to spawn that sub-agent is included in the analysis context. If the sub-agent's actions are consistent with its authorized task, the risk level reflects that authorization. If a sub-agent's actions materially exceed what the main agent specified, risk is escalated accordingly.
Agent Recon™ scans every event for personally identifiable information using a local regex engine only — no PII is sent to any external service, including the Anthropic API. Detection is entirely server-side and runs synchronously before the event is written to the database.
What is detected:
| PII Type | Base Risk | Detection Strategy |
|---|---|---|
| SSN | Critical | Formatted (XXX-XX-XXXX) or bare 9-digit near ssn keyword |
| Credit Card | Critical | Formatted 4-group or known-BIN pattern; Luhn check boosts confidence |
| Passport Number | Critical | Alphanumeric pattern near passport keyword |
| Bank Account / IBAN | High | IBAN format or number near bank account keyword |
| Email Address | High | RFC-5321 pattern (high inherent specificity) |
| Phone Number | High | NANP / E.164 format; keyword proximity boosts confidence |
| Date of Birth | High | ISO or US date near dob / birth / born keyword |
| Driver's License | High | Format pattern near driver's license / dl# keyword |
| Personal Name | Medium | Capitalized name pattern only near explicit name field keywords (e.g. first name:, patient_name=) — requires keyword context to avoid false positives |
Confidence scoring: Each detection is assigned a confidence score (0–1) based on pattern completeness, keyword proximity, and whether multiple PII types appear together in the same event (composite boost). Only detections above a 0.55 threshold are promoted. The effective risk level is computed from both the base risk and the confidence score, so a partial match for a high-risk type is reported at a lower risk level.
Privacy guarantees:
["ssn","email"]) and the event field names where PII was found.[REDACTED:TYPE] tokens. The LLM receives a notice of which types were detected and in which fields, but never the actual values.🪪 PII category cards. The trigger text shows only the detected type names.Note: In "Raw" storage mode, the event payload column contains the full original JSON, which may include PII. In "Redacted" mode, PII values are replaced with [REDACTED:TYPE] tokens before storage — original values are permanently removed. In "Encrypted Raw" mode, the payload column stores redacted text while the original is preserved in an AES-256-GCM encrypted column (payload_encrypted) recoverable via the archive encryption key. All modes prevent PII from being transmitted to external services.
curl downloading a package, sudo apt-get install, or reading a .env.example file will be flagged. Use the risk level, subcategory tag, and the expanded event JSON — and especially the LLM chain summary — to judge each case in context.
The allowlist is a false-positive filter that reduces noise in the Security tab. When enabled, each security-classified event is evaluated against a set of rules before it is displayed. Rules can either suppress an event (hide it from the Security tab entirely) or downgrade its risk level (e.g. from high to low). Suppressed and downgraded events are still stored in the database for forensic analysis — the allowlist only affects what appears in the live dashboard.
Three built-in profiles are available in Settings → Security Allowlist → Active profile:
| Profile | Rules | What it does |
|---|---|---|
| Standard Development | 14 | Suppresses routine dev workflow noise: package manager sudo calls, git commit heredocs, credential manager patterns (pass, Azure CLI), tool path resolution subshells, standard chmod modes, localhost API testing. Downgrades node -e one-liners, general curl, base64 -d, and PII in test fixtures to lower severity. |
| DevOps / Infrastructure | 18 | Includes all Standard Development rules plus: SSH connections downgraded to low, all chmod calls suppressed, rsync / scp downgraded to low, all sudo usage downgraded to medium. |
| Strict | 0 | No suppressions — every event is shown at its original YARA-classified risk level. Use this when auditing a session or investigating a suspected incident. |
Clicking Apply seeds the selected profile's rules into the database. Switching profiles replaces the previous profile's rules but preserves any custom rules you have added.
In addition to profile rules, you can create your own rules for project-specific patterns. In Settings → Security Allowlist, click + Add Rule to open the rule form. Each rule has:
| Field | Description |
|---|---|
| Pattern | The text to match against the event's command, file path, URL, or prompt content. |
| Match mode | Contains (case-insensitive substring), Regex (case-insensitive regular expression), or Exact (full string match). |
| Action | Suppress hides the event from the Security tab. Downgrade changes the risk level to the specified target (low or medium). |
| Category / Subcategory | Optional scope filters. When set, the rule only matches events in that security category (e.g. exec) or subcategory (e.g. priv-escalation). |
| Priority | Lower numbers are evaluated first. The first matching rule wins. Profile rules typically use priorities 10–50; custom rules default to 100. |
| Reason | A note explaining why the rule exists — shown in the rule list for reference. |
Custom rules are evaluated alongside profile rules. They persist across profile changes — switching from Standard Development to DevOps replaces the profile rules but leaves your custom rules intact.
The Security tab header shows a suppressed chip with a count. Click it to toggle visibility of allowlist-suppressed events. Suppressed rows appear at reduced opacity with a dashed border so they are visually distinct from active events.
The pre-built profiles cut Security tab noise on day one, but they can't anticipate every project's quirks. Auto-Learning Baselines watches your actual workflow and surfaces patterns that recur across sessions, agents, and projects without ever escalating to high or critical risk — so the long tail of false-positives can become suppression rules tailored to your stack rather than yours-and-everyone-else's.
Patterns are tracked as (tool, classification, path-prefix-or-verb) tuples. A pattern becomes a candidate when it crosses all of:
auto_baseline_min_occurrences / _min_sessions)Patterns in five categories are never auto-baselined regardless of frequency: secret, inject, priv-escalation, prompt-secret, prompt-injection. Events that triggered a PII match are also skipped. A single session contributes at most three observations per pattern, so a runaway loop can't dominate the count.
Candidates appear as a panel at the top of the Security tab. Each card shows the pattern, occurrence count, session labels, and agent/project diversity. Choose:
source: auto-learned at priority 200 (lower than profile=100 and manual=50, so user-curated rules always win)auto_baseline_reject_cooldown_days (default 30) it re-enters observation, so a pattern you didn't want suppressed today can re-surface if conditions change. Set the cooldown to 0 to make rejections permanent.Toggle the whole feature off via auto_baseline_enabled = false in Settings if you'd rather keep using only manual / profile rules.
A per-token dollar cost is only real when you pay Anthropic per token with an API key. The chips reflect your billing mode (set it in Settings → Cost tracking; it is auto-detected from the hook when possible):
Cursor and Copilot CLI don't expose per-token usage or cost locally (they're subscription / credits-metered). Their session chips show usage n/a / credits-metered rather than a misleading $0.00, and their cost is shown as — in the Insights and Session History views.
Each session column header shows a token chip once the agent reports usage. For Claude Code it shows accumulated input + output tokens and — in API-key mode — an estimated cost (from the model in the Stop event); in subscription/unknown modes it shows tokens only.
The chip in the top navigation bar shows today's and this month's totals. In API-key mode two data sources contribute:
Token counts always come from SQLite estimates (the Usage API does not return per-period token totals). In subscription, Bedrock/Vertex, or unset modes the chip never shows a ✓ or a real billed $ — see Honest cost above.
Cost estimates use the model name reported in each Stop event. If the model is not in the pricing table the server falls back to Sonnet pricing. The pricing table is defined in server/tokens.js.
The Insights tab provides a per-session retrospective analysis of your Claude Code usage. Each completed session appears as a collapsible row. Click any row to expand it and see five sections: a session gist, usage stats, prompt coaching, key observations, and hallucination risk.
A 2–3 sentence plain-English summary of what was being built or fixed during the session, generated by Claude Haiku after the session ends. Useful for quickly recalling what a past session was about.
Computed locally from your SQLite data — no API call needed. Shows session duration, turn count (number of Stop events), user prompt count, total tool calls, token cost, tokens in/out, cache hit rate, and a bar chart of tool usage frequency.
Cache hit rate is the percentage of input tokens served from Claude's prompt cache. Higher is better — it means more context was reused, reducing cost.
Each UserPromptSubmit event is rated on a 1–5 clarity scale. For prompts scoring below 4, Agent Recon™ identifies specific issues (e.g. "too vague", "missing file context") and shows a rewritten example of how the prompt could be improved. Prompts scoring 4 or 5 are marked as good.
Privacy note: Only the raw text of your prompts is sent to the Anthropic API for analysis — no file contents, no tool outputs. Prompts are truncated at 600 characters before sending.
2–5 actionable observations about the session as a whole: communication patterns, tool diversity, iteration habits, cache efficiency, and other developer-facing insights that help you get more out of future sessions.
An overall 1–10 rating of prompt quality across the session. 8–10 (green) means prompts were specific, contextual, and effective. 5–7 (yellow) means there is room for improvement. 1–4 (red) means most prompts were too vague or required significant back-and-forth.
Each session includes a hallucination risk score (0–100%) based on observable evidence from tool activity. The score appears as an H:XX% chip in the collapsed row and as a detailed section when expanded.
Two-tier evidence model: Agent Recon™ distinguishes direct evidence (tool failures, user corrections) from circumstantial patterns (long chains, complexity). Sessions with only circumstantial signals are capped at Moderate risk to avoid false alarms.
Tier 1 — Direct evidence (strong weight):
Tier 2 — Circumstantial (lower weight, capped):
The score is computed from event patterns alone (no API key needed). When an API key is available, the LLM refines the assessment using actual tool output evidence. A clean session scores 0%.
Risk levels: 0–15% Low (green), 16–35% Moderate (yellow), 36–60% Elevated (orange), 61–100% High (red).
Labels: "Evidence-based" means direct tool failures or user corrections were found. "Pattern-based" means the score reflects session patterns only — no direct hallucination evidence was observed.
For sessions with moderate or higher risk, Agent Recon™ shows beginner-friendly remediation guidance explaining what was detected and what you can do differently next time.
Each row has a refresh button (↻) that deletes the cached analysis and re-queues the session for a fresh LLM pass. Useful if you update your API key after a session has already been processed.
The search bar filters insights by session label, gist text, or observation content. Category filters (from the Sessions view) carry over — filtering to "security" events shows only sessions with security-relevant insights. Use the header search field (/ shortcut) for cross-tab searching, or the Insights-specific filter within the tab for narrower results.
Click any insight card to expand the full analysis. The collapsed view shows the session gist and key metrics (efficiency score, hallucination risk, prompt count, tool count, cost). The expanded view adds prompt coaching details with per-prompt ratings, individual observations with actionable recommendations, and the hallucination risk breakdown with evidence tiers and remediation guidance.
The Environment tab provides a cross-session view of your overall Claude Code usage patterns over a configurable analysis window (default: 7 days, configurable in Settings up to your retention period). It shows LLM-generated recommendations, aggregate usage statistics, cost estimates, session duration metrics, an activity heatmap, and a history of resolved recommendations.
Claude Haiku analyzes your event data, cost data, session durations, prompt quality signals, and current Agent Recon™ settings to generate 3–6 actionable recommendations across four categories:
| Category | Description |
|---|---|
settings | Suggestions about Agent Recon™ or Claude Code configuration changes |
workflow | Session patterns, prompt habits, session length optimizations |
security | Observations based on security event patterns or risky behaviors seen |
performance | Ways to reduce token usage, cost, or latency |
Each recommendation shows a priority badge (high / medium / low), a category chip, a title, and a description. You can dismiss recommendations you don't want to see (hover to reveal the dismiss button). Dismissed recommendations can be shown again via the "Show dismissed" toggle. Settings-category recommendations include a link to open the Settings modal. Recommendations require an Anthropic Inference API key (set in Settings).
When a recommendation is no longer generated by the LLM (e.g., you fixed the issue), it is automatically moved to the Recently Resolved section at the bottom of the tab.
Recommendations are scoped to your development workflow and configuration — the LLM is explicitly instructed not to comment on Claude Code's internal AI mechanics (e.g. tool loading, plan mode transitions, task tracking). Those are system internals outside your control.
The stats panel is computed directly from SQLite — no API call needed. It shows:
ToolSearch, EnterPlanMode, TaskCreate, and similar internal tools) are excluded so the chart reflects your actual coding activity.The analysis window is configurable in Settings (7, 14, 30, 60, or 90 days), capped at your data retention period.
Analysis runs automatically 45 seconds after the server starts, then every 60 minutes. To avoid redundant broadcasts, a fingerprint of the recommendation titles and key stats is stored after each run. If the fingerprint matches the previous run, the results are not re-broadcast — this prevents the tab from flickering on reconnect when nothing has changed.
Each session in the Insights tab includes a Prompt Quality section that identifies the most significant prompting issues — patterns that caused unnecessary iterations, bugs, or misalignment. The analysis surfaces at most 3 issues per session, focused only on the worst offenders.
| Badge | Meaning |
|---|---|
Ambiguous | The prompt lacked enough specificity, causing the agent to guess intent — even accounting for prior session context. Common cause of off-target output. |
Misleading | The prompt contained incorrect assumptions or constraints. The agent followed instructions faithfully but produced the wrong outcome. |
Over-literal | The agent interpreted the prompt too literally and missed obvious implicit requirements. |
Cavalier planning | The agent made sweeping changes or assumptions beyond what was asked — unexpected refactors, scope creep, or unrelated file edits. |
Context assumption | The user relied on Claude retaining a constraint or decision from early in a long session without restating it. Claude Code auto-compacts context at ~95% capacity, which can silently drop early instructions — especially in sessions exceeding ~20 prompts or ~80K tokens. The fix: restate key constraints when sessions run long, or use /clear to start fresh between tasks. |
Session scope creep | Multiple unrelated workstreams were mixed into one session, increasing the risk that Claude applied patterns from one task to another. The fix: use /clear between distinct tasks or start a new session when switching topics. |
Claude Code sessions have cumulative memory — each prompt is sent with the full prior conversation. The analysis accounts for this: short referential prompts later in a session ("do the same for the other files", "now fix that too") are treated as effective, not penalized for brevity. Prompts at position 5 or later are only flagged as ambiguous if the reference would be genuinely unclear even with full session history.
For long sessions (20+ user prompts), the analysis also checks for context assumption patterns — corrections that reference something established very early, which may have been lost during context auto-compaction.
Analysis is triggered automatically 5 seconds after a session ends (requires at least 3 user prompts and 2 detected correction loops as a signal that something went wrong). The engine uses Claude Haiku for initial analysis and escalates to Claude Sonnet if 2 or more high-severity issues are found.
Correction loops are detected when a user prompt follows a Stop event within 3 minutes and contains corrective language ("actually", "wait", "undo", "that's not what I meant", etc.). Each correction is labeled with its position in the session and elapsed time from session start, helping distinguish early prompt quality issues from late-session context drift.
When Settings → Prompt Quality → GitHub Issue Correlation is enabled, the engine queries GitHub for issues opened or closed during the session time window. Bug reports filed during a session are strong signal that something went wrong, and are included as context in the LLM prompt.
The repo is auto-detected from the session's working directory (git remote get-url origin). If the project is not a git repo or does not use GitHub, enrichment is silently skipped. GitLab and Bitbucket support will be added in a future release.
Each recommendation has thumbs-up / thumbs-down buttons. Your ratings are stored locally and will be used in a future release to personalize the analysis — suppressing failure modes that are consistently unhelpful for your sessions and steering the LLM toward better recommendations using your past examples.
Agent Recon observes Claude Code by default, but it can also observe sessions from other AI coding agents. When a supported agent is detected, the installer registers a lightweight hook forwarder so that agent's lifecycle events stream into the same dashboard — the live feed, security classification, and PII scanning all work identically. Events carry a source-agent badge so you can tell them apart, and clicking a badge filters the feed to that agent.
| Agent | Hook config | Notes |
|---|---|---|
Claude Code | ~/.claude/settings.json | Full support, including token cost and statusline rate limits. |
Cursor | ~/.cursor/hooks.json | Lifecycle and tool events. Token cost not available — see below. |
GitHub Copilot CLI | ~/.copilot/hooks/agent-recon.json | Lifecycle and tool events. Token cost not available — see below. |
Each agent fires hooks at lifecycle points — session start and end, prompt submission, tool use. A per-agent forwarder script reads the event, tags it with its source agent, and hands it to a detached background worker that POSTs it to the local Agent Recon server. Because the network call runs in the detached worker, even a blocking pre-tool hook returns to the host agent in milliseconds — Agent Recon never slows down your editor. A server-side mapper rewrites each agent's event and tool names into one common vocabulary, so every dashboard feature behaves the same regardless of which agent produced the event.
When a session delegates to a sub-agent, Agent Recon records the spawn as a parent→child edge and captures the task the sub-agent was given and the result it returned (both PII-redacted at rest). Expand a session in the History tab to see its delegation tree. The hosts differ in what they expose:
SubagentStart carries the delegation prompt and SubagentStop the returned result. The parent→child link is reconstructed best-effort from spawn order within the session (hooks don't carry a parent id; only OpenTelemetry spans do, and Agent Recon reads trace ids but does not ingest spans).parent_conversation_id) plus the task and a summary, so its delegation edges are exact.Claude Code's experimental Agent Teams (enabled with CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1) let multiple agents share a task list and message each other directly. Agent Recon reads the team roster and task list from disk (~/.claude/teams and ~/.claude/tasks), exposed at GET /api/teams.
Limitation: the messages teammates send each other (the Mailbox / SendMessage payloads) are not exposed by any hook, so Agent Recon can show that coordination happened — who is on the team and how tasks flow — but not the message contents. GitHub Copilot CLI has no teams concept. Cursor's cloud Background Agents run inside an isolated VM that can't reach your local server, so they are observable only through Cursor's own dashboard / REST API, never local hooks.
The easy way (recommended) — the installer does it all:
agent-recon install — or just re-run it if Agent Recon is already installed. It detects each tool and registers its hooks automatically, preserving any hooks you already have.Agent Recon installs its hook scripts to ~/.agent-recon/hooks/ — a tool-neutral folder shared by Claude Code, Cursor, and Copilot. A Cursor or Copilot config that points at this folder is expected and correct; it is not a Claude-only path.
If you can't run the installer, register the hook yourself. For Cursor, create or edit ~/.cursor/hooks.json and add the command under each event you want to watch (sessionStart, beforeSubmitPrompt, preToolUse, postToolUse, stop, …):
{
"version": 1,
"hooks": {
"preToolUse": [
{ "command": "C:/Program Files/nodejs/node.exe C:/Users/you/.agent-recon/hooks/send-event-cursor.js" }
]
}
}
Use the full path to node — run where node (Windows) or which node (macOS/Linux) to find it. Cursor and Copilot are launched from your desktop and may not have node on their PATH, so a bare node can silently fail. The installer handles this for you; a hand-written config must spell it out. Save, then fully restart the tool. Copilot CLI uses the same idea in ~/.copilot/hooks/agent-recon.json with send-event-copilot.js.
http://localhost:3131; start it with agent-recon start.node, not a bare node.CLAUDE_CODE_SAFE_MODE is set in your environment — Safe Mode disables all hooks, so Agent Recon never receives events. Unset it (unset CLAUDE_CODE_SAFE_MODE) and restart Claude Code. A completely empty dashboard with a running server is the tell-tale signature./env inside Copilot CLI and confirm the Agent Recon hooks are listed under its hook configuration. If they're absent, the config didn't load — fully restart Copilot, and verify ~/.copilot/hooks/agent-recon.json uses the full node path.Re-running the installer is idempotent and never duplicates entries. Uninstalling removes only Agent Recon's own hook entries from each tool's config and deletes the scripts from ~/.agent-recon/hooks/, leaving any third-party hooks untouched.
Cursor and GitHub Copilot CLI do not expose token usage in their hook payloads, so the Tokens tab and per-event cost badges stay Claude-Code-only. Every other feature — the live feed, the Security tab, PII redaction, session labels, environment detection — works fully for all supported agents.
The History tab shows a searchable list of all recorded Claude Code sessions with LLM-generated summaries. Each session card shows a title, the project name, date, event count, and cost. Click a card to expand it and see the full session summary and usage stats.
| Pip | Meaning |
|---|---|
⋯ pending | LLM analysis has not yet run for this session. Analysis is triggered automatically after a session ends. |
✓ AI | LLM analysis complete. The title and detail text were generated by Claude. |
✗ | Analysis failed (e.g. no API key was configured when the session ended). |
When a session spawns sub-agents (Claude Code's Agent/Task tool, Cursor sub-agents, or Copilot delegated agents), an expanded card shows a Delegation tree: each sub-agent as a parent→child node with the task it was given, the result it returned, its duration, and whether it finished. The task and result text is PII-redacted at rest, like every other stored field.
The parent→child edge is reconstructed best-effort from event ordering (hooks don't carry a parent id on every host) — Cursor supplies an authoritative parent via parent_conversation_id; Claude Code and Copilot fall back to the spawn order within the session. Copilot exposes no sub-agent prompt or result inline, so those nodes show the sub-agent name only. Sessions with no sub-agents omit the section entirely.
Each session card has a Download button that exports the full session event log as a JSON file via GET /api/sessions/:id/export. The file contains all hook events recorded for that session.
The model used to generate session history summaries is configurable in Settings → LLM Models → Session History LLM. Defaults to claude-opus-4-6. Requires an Anthropic Inference API key.
⋯ pending indefinitely. Set a key in Settings to enable background analysis.
At the bottom of the History tab, a collapsible Prompt Trends panel shows aggregate prompt quality data across the last 30 days (configurable in Settings). Expand it to see a bar chart of how often each failure mode appeared, trend arrows showing whether each mode is improving or worsening over the period, and a top recommendation summarising the most common issue across recent sessions.
The panel loads lazily — data is fetched the first time you expand it. The trend window can be adjusted via Settings → Prompt Quality → Trend window.
Search by session label, narrative title, or summary text using the header search field (/ shortcut). Sessions are sorted by last activity (most recent first). The status pip next to each session card shows generation state: green means complete, yellow means pending, red means failed. Use category filters from the Sessions view to narrow results to sessions with specific event types.
Summaries are generated by the configured LLM model (default: claude-opus-4-6, configurable in Settings). Click any session card to expand the full narrative including a structured summary, key decisions, and outcome assessment. Use the Download button to export session data as JSON for external analysis. Click the refresh icon to regenerate a narrative with the current model — useful after changing the LLM model in Settings or if the original generation used incomplete data.
All events are stored in SQLite and persist across server restarts and Claude Code restarts. The database path is shown in Settings → Database → Storage file.
The server checkpoints the SQLite WAL file on graceful shutdown (Ctrl+C / SIGTERM) so that all data written since the last automatic checkpoint is safely flushed to the main database file before the process exits. A periodic checkpoint also runs every 5 minutes while the server is running.
NTFS / WSL path advisory: If the DB path starts with /mnt/ the database is on an NTFS volume accessed via the WSL 9P filesystem. SQLite WAL mode is less reliable there because mmap()-based shared memory is not fully supported. For maximum durability, set the DB_PATH environment variable to a WSL-native path before starting the server:
export DB_PATH=~/agent-recon/agent-recon.db node start.js
A yellow advisory banner is shown in Settings when the active DB path is on an NTFS mount.
Controls how many days of events are kept in the main database (valid range 1–365 days; default 90). Events older than the retention window are first archived (encrypted) and then removed by the hourly archiver — so lowering retention moves old events into recoverable archive files rather than deleting them outright.
The archiver is the single authority for removing aged-out events: it always writes an encrypted archive before deleting from the main database. Archives are AES-256-GCM encrypted; the key is auto-generated on first run and stored in data/archive.key. Use the "Download decrypt script" button to get a Python helper that decrypts any archive file back to a JSON-lines file.
The disk policy controls what happens when available disk space is low: Drop oldest archive removes the oldest archive file to free space, while Pause logging stops accepting new events until space is available.
Agent Recon™ uses two separate Anthropic API keys for different purposes:
sk-ant-api03-…) — used for LLM-powered features: security chain analysis and session Insights (gist, prompt coaching, observations). Can also be set via the ANTHROPIC_API_KEY environment variable before starting the server. Standard inference keys are created at console.anthropic.com → API Keys.
sk-ant-admin03-…) — required to query the Anthropic Admin Usage API (GET /v1/organizations/cost_report), which provides org-level verified cost data. Standard inference keys return 404 for this endpoint. Admin keys are created at console.anthropic.com → Settings → API Keys → Admin Keys.
Both keys are stored in the local SQLite database and never transmitted anywhere other than Anthropic's own APIs. If only an Inference Key is set, Agent Recon™ falls back to SQLite cost estimates (no ✓ badge) and logs a warning.
Your stored API keys and tokens are encrypted at rest with a master key held in your operating system's keystore (Windows DPAPI, macOS Keychain, Linux libsecret) or supplied via the AGENT_RECON_MASTER_KEY environment variable. The server refuses to start if secrets would land unencrypted, unless you deliberately opt in with AGENT_RECON_ALLOW_PLAINTEXT=1. Your saved activity uses a separate archive key, not this one.
Rotating the key (Settings → Master Encryption Key → Rotate now) generates a fresh key and re-encrypts your stored secrets in one atomic step; an interrupted rotation is finished automatically on the next restart, so no secret is lost. Rotation is unavailable when the key comes from the AGENT_RECON_MASTER_KEY environment variable (rotate that out-of-band by updating the env file) or when secure storage isn't available — the button shows why.
A 🔍 button in the header toolbar opens an inline search field. Type three or more characters to filter content across all tabs simultaneously. The filter persists as you switch tabs. Press / anywhere to open the search field. Press Escape to collapse the field (the filter remains active while the icon is highlighted blue). Click ✕ inside the field to clear the query.
Raw JSON search — controlled via ⚙️ Settings → Search raw JSON. When enabled (default), the search also matches against the full raw event JSON payload in the Sessions tab. Useful for finding specific file paths, tool arguments, or data not shown in the card preview. Disable this setting if search feels slow on large sessions.
On connect, the dashboard loads today’s sessions — all events from 12:00:01 AM in your local time zone. This keeps the feed focused on current work without being overwhelmed by historical data. The Canvas timeline is always live-only. Use the History tab to browse older sessions from the full database.
When enabled, Agent Recon™ tracks OS-level process spawns and classifies them for security relevance (requires server restart to take effect). This is a security signal — it catches a risky child process an agent spawns that the tool-call hook alone wouldn’t reveal (e.g. a curl | sh grandchild, a miner, a reverse shell). Spawns of known-safe processes are silently skipped; everything else appears in the Security tab under the proc-exec category.
Spawns are attributed best-effort to the session and agent that likely launched them (via a recent background-shell launch, falling back to the most-recently-active session) so a flagged process isn’t orphaned. It is not a background-task tracker — the OS monitor runs on Linux/WSL and Windows only (not macOS), polls periodically (so very short-lived processes can be missed), and no host emits a process-exit hook, so completion isn’t tracked. For whether an agent ran something in the background, see the BG badge on event cards instead.
Reduces false-positive noise in the Security tab by suppressing or downgrading events that match known-safe patterns. Choose a built-in profile (Standard Development, DevOps, or Strict) and add your own custom rules for project-specific patterns. See the Security docs tab for full details on profiles, custom rules, and how suppression works.
Push Agent Recon™ data to any OpenTelemetry-compatible collector (Grafana, Jaeger, Zipkin, or any OTLP/HTTP endpoint) via three signal types:
Set the Endpoint URL to your collector's OTLP/HTTP base URL (e.g. http://localhost:4318), then toggle Export on. Signals go to /v1/logs, /v1/traces, and /v1/metrics under that base. No additional packages required.
Choose an Auth mode to secure the connection to managed OTEL collectors:
| Mode | Sends | Used by |
|---|---|---|
| None | — | Local / open collectors (localhost Jaeger, OTC, etc.) |
| Bearer token | Authorization: Bearer <token> | Generic OAuth2, Elastic, Dynatrace |
| API key header | <Header>: <key> | Honeycomb (x-honeycomb-team), Datadog (DD-API-KEY), New Relic (api-key), Signoz (signoz-access-token) |
| Basic auth | Authorization: Basic base64(user:pass) | Grafana Cloud (instance ID + API key) |
Custom headers (one Header: value per line) are merged after the auth mode header, mirroring the OTEL_EXPORTER_OTLP_HEADERS environment variable. Custom headers take precedence on name collision. Credentials are stored in the local SQLite database and never logged.
Enable HTTPS to encrypt traffic between the browser, hook scripts, and the Agent Recon™ server. TLS is available to all tiers (Community, Personal, and Professional). When TLS is enabled, the HTTP server on the configured HTTP port (default 3131) redirects all requests to the HTTPS server on the configured HTTPS port (default 3132). A server restart is required after changing TLS settings.
brew install mkcert on macOS, choco install mkcert on Windows, apt install mkcert on Linux), then run mkcert -install (requires admin/sudo — installs a local root CA). Certificates are stored in data/tls/ and auto-regenerate when within 30 days of expiry.send-event.js probes HTTPS first, then falls back to HTTP. Set the AGENT_RECON_HTTPS_PORT environment variable to override the configured HTTPS port (visible in Settings).
When TLS is enabled with mkcert, open the HTTPS dashboard URL shown in Settings — it will display a green padlock immediately. No browser exceptions needed.
Use Export DB to download the full SQLite database file. Use Clear all data to permanently remove all events and sessions from the database. Both actions take effect immediately.
Each analysis pipeline (Security, Insights, Environment, Session History) can use a different Claude model. Defaults: Claude Opus 4.6 for Security and Session History (requires highest accuracy), Claude Haiku 4.5 for Environment and Prompt Quality (cost-efficient for lighter analysis). Change models in Settings → LLM Configuration. Model changes take effect on the next analysis run — previously generated results are not re-analyzed automatically.
Enter your license key in Settings → License. The dashboard shows your current tier (Community, Personal, Professional) and activation status. Community tier is free with core features. Personal and Professional unlock LLM-powered analysis, extended retention, and priority support. License keys support 3 device activations — deactivate a device from the Settings panel to free a slot.
When enabled, analyzes your prompts against common failure modes (vague instructions, missing context, overly broad scope). Results appear in the Insights tab under "Prompt Coaching." Each prompt is rated on a 1–5 clarity scale with specific improvement suggestions for prompts scoring below 4. Optionally correlates with GitHub Issues for project-specific recommendations. Configure in Settings → Prompt Quality.
The scrolling waveform at the top of the Sessions view. Height is configurable via Settings → Canvas. Shows real-time event density with color-coded activity bands per session. Pinned event icons appear for significant events (agent spawns, errors, security flags). The canvas is always live-only and is unaffected by the Session Display setting. Drag to scroll back through recent activity; the canvas auto-scrolls to the latest event when new events arrive.
Expand any item below for diagnosis steps and solutions.
Verify hooks are installed: check that ~/.claude/settings.json lists send-event.js for all 24 hook events. Verify the server is running: curl http://localhost:3131/health should return {"status":"ok"}. Check the browser console for WebSocket connection errors. On WSL2, ensure localhost forwarding is working — try curl http://$(cat /proc/net/route | awk 'NR==2{print $3}' | sed 's/../0x&\n/g' | tac | printf "%d.%d.%d.%d" $(cat)):3131/health as a fallback.
Check that your firewall allows connections to localhost:3131 (or your configured port). On WSL2, ensure localhost forwarding is enabled in .wslconfig. Browser tabs throttle WebSocket connections when backgrounded — keep the Agent Recon tab visible or use a pinned tab for uninterrupted monitoring. Agent Recon reconnects automatically with exponential back-off, so brief disconnects are handled gracefully.
Token cost tracking requires an Anthropic Admin API key (not an Inference key). Go to Settings → API Keys and enter your Admin key (sk-ant-admin03-...). Admin keys are created at console.anthropic.com → Settings → API Keys → Admin Keys. The Usage API polls every 10 minutes by default. SQLite-based cost estimates from Stop events work without an Admin key but do not show the verified checkmark badge.
This is normal if your session had no security-relevant events (file writes to sensitive paths, shell commands with injection patterns, credential access, network calls, etc.). The Security tab only shows events that match the built-in classification rules. Check that security_llm_enabled is on in Settings for LLM chain analysis — without it, only regex-based classification runs.
Insights analysis requires an Anthropic Inference API key (sk-ant-api03-...). Go to Settings → API Keys and enter your key. The analysis runs automatically when a session has enough events (at least one Stop event). Click the refresh button on any insight card to force a retry. Check the server console for API errors — common causes are invalid keys, rate limits, or network issues.
Another process is using port 3131. Either stop it or set a custom port: PORT=3132 node start.js. On Windows, use netstat -ano | findstr :3131 to find the process. On macOS/Linux, use lsof -i :3131. If a previous Agent Recon™ instance is still running, stop it first or use the service management commands for your platform.
The native SQLite addon requires C/C++ build tools. Windows: npm install -g windows-build-tools. macOS: xcode-select --install. Linux: sudo apt install build-essential python3 (Debian/Ubuntu) or sudo dnf groupinstall "Development Tools" (Fedora). The start.js launcher detects platform mismatches and runs npm rebuild better-sqlite3 automatically, but the build tools must be present.
Session labels derive from the git remote name of the working directory where Claude Code is running. If your working directory is not a git repo or has no remote configured, the label falls back to the directory name or a hex suffix. This is cosmetic only and does not affect functionality. To get meaningful labels, ensure your project has a git remote: git remote add origin <url>.
Claude Code on Windows runs hook commands through Git Bash, which sometimes doesn't share PATH with the shell that installed Agent Recon (nvm-windows / Volta / fnm shims). The installer probes for node --version at install time and falls back to the absolute path of the current Node runtime when the probe fails. If you switch Node versions after installing Agent Recon, re-run agent-recon install (or setup.ps1) to refresh the hook command.
The default Restricted execution policy prevents npm.ps1 from running in PowerShell. Run once: Set-ExecutionPolicy RemoteSigned -Scope CurrentUser. This allows locally-created scripts and signed remote scripts to run.
Windows Defender Firewall blocks inbound connections to node.exe on dynamic ports. When prompted during server start, allow access for private networks. Or create a rule manually in Windows Security → Firewall & network protection → Allow an app through firewall.
If the server crashes with EACCES: permission denied, mkdir '...data', the data directory resolved to a root-owned path. Upgrade to v1.0.3+ which stores data in ~/.config/agent-recon/data/ (Linux), ~/Library/Application Support/agent-recon/data/ (macOS), or %APPDATA%\agent-recon\data\ (Windows). Workaround: AGENT_RECON_DATA_DIR=~/.config/agent-recon/data agent-recon start.
On headless servers or SSH sessions, libsecret needs a D-Bus session and unlocked keyring. The server falls back to PBKDF2 file-based encryption automatically. If credential errors persist, set a master key: export AGENT_RECON_MASTER_KEY=$(openssl rand -hex 32) before starting.
The top navigation bar runs across the full width of the Agent Recon™ page. From left to right it contains: the logo, a connection status indicator, a spacer, a token cost chip, and a set of control buttons.
The Agent Recon™ logo on the far left. Clicking it does nothing — it is a label only.
A colored dot (● green / ● red / ● grey) and a short text label showing the current WebSocket connection state:
| State | Dot | Label |
|---|---|---|
| Connected | ● | Connected |
| Connecting / reconnecting | ● | Connecting… / Reconnecting… |
| Disconnected (max retries) | ● | Disconnected |
Agent Recon™ automatically reconnects on drop with exponential back-off. If the server is stopped and restarted, the browser will reconnect and replay the full event history automatically.
The cost chip (e.g. 12.3K $0.042 today · $1.23 mo) shows accumulated token usage and estimated cost for today and this month. It is invisible until at least one session has reported token data.
Hover the chip for a tooltip with full detail. See the Tokens docs tab for more.
Filters the Sessions tab (and Security tab) to show only events from the selected session. Defaults to All Sessions. Select a specific session label to focus on one swimlane. The dropdown populates automatically as new sessions are seen.
Opens this documentation modal. Press Escape or click outside the modal panel to close it.
The speaker button mutes and unmutes the audio cues that play when events arrive. The slider next to it controls the playback volume (0–100%). Sound cues use the Web Audio API — no audio files are loaded.
On page load or refresh, the speaker icon pulses amber — this is normal browser behaviour. Browsers require a user gesture before audio can play (the Web Audio API Autoplay Policy). The amber pulse means audio is ready but waiting for your first interaction. Hover over the icon to see the tooltip "Audio locked — click anywhere to activate". The pulse clears automatically the moment you click or press any key.
Opens a centered modal with accessibility settings: high-contrast mode, reduced motion, font size adjustment, and a visual input-wait alert for users who may not be able to hear the audio cues. Press Escape or click outside the modal to close it.
Opens a centered Settings modal with all Agent Recon™ configuration options. Press Escape or click outside the modal panel to close it.
Agent Recon™ supports keyboard navigation for accessibility:
Modal dialogs trap keyboard focus so Tab cycles only within the dialog. Focus returns to the element that opened the dialog when it closes.
When telemetry is enabled, Agent Recon™ sends a single heartbeat on startup and once daily. Each heartbeat contains exactly four fields:
| Field | Example | Description |
|---|---|---|
install_id | a1b2c3d4-... | Random UUID — not derived from any personal information |
version | 1.5.0 | Agent Recon version string |
platform | darwin-arm64 | OS and CPU architecture |
timestamp | 2026-03-24T12:00:00Z | ISO-8601 datetime of the heartbeat |
Your source IP address is captured server-side by Azure infrastructure. The client never sends IP information.
Everything else. The following data never leaves your machine:
YOUR MACHINE AZURE (telemetry.agent-recon.net) ──────────── ───────────────────────────────── Claude Code events → SQLite DB (never sent) Session content → SQLite DB (never sent) API keys → encrypted DB (never sent) Security analysis → local LLM calls (never sent to us) Heartbeat (if enabled): install_id + version ──HTTPS──→ Table Storage (90-day retention) + platform + timestamp Source IP captured server-side
Agent Recon™ periodically checks for new versions by fetching a static JSON file from agent-recon.net. This check follows the telemetry setting — when telemetry is off, automatic update checks are also off. You can always check manually via Settings → Updates → Check now regardless of the telemetry setting.
When an update is available, you'll see a green badge next to the version number in the header, a recommendation card in the Environment tab, and the current status in Settings.
Open Settings (gear icon) and toggle "Send anonymous usage statistics" to disabled. When disabled, zero telemetry is sent — no heartbeat, no version check, nothing. The change takes effect immediately.
Open Settings (gear icon) and click "Delete My Data" in the telemetry section. This permanently deletes all telemetry records associated with your install ID from our servers and automatically disables telemetry.
Deletion is immediate and permanent — deleted data cannot be recovered.
Heartbeat records and IP-to-ASN mapping data are retained for 90 days, then permanently deleted.
GDPR (EEA residents): right to access, erasure, objection, portability, and rectification. Use the Delete My Data button in Settings or open a GitHub Issue.
CCPA (California residents): right to know, right to delete, right to opt out of sale (we do not sell data), and non-discrimination.
See PRIVACY.md in the Agent Recon repository for the complete privacy policy, including data security details, children's privacy, and change history.
Agent Recon™ is a real-time observability dashboard for AI coding agent sessions. It intercepts every lifecycle event that Claude Code emits — tool calls, prompts, completions, agent activity — and displays them in a live feed so you can see exactly what Claude is doing, how many tokens it is spending, and whether any security-relevant patterns are occurring.
Claude Code → Node.js hook → POST /event → SQLite → WebSocket → browser
The server is a Node.js Express application (server/server.js) that receives hook events, deduplicates them, stores them in SQLite via better-sqlite3, and broadcasts them to all connected browser clients over WebSocket in real time.
send-event.js — single cross-platform Node.js forwarder. Runtime-detects Windows / macOS / Linux / WSL and adapts (WSL adds gateway-IP fallback).The script forwards every Claude Code lifecycle event to http://localhost:3131/event. Run agent-recon install, or the platform setup script (setup.ps1, setup-wsl.sh, setup-macos.sh, setup-linux.sh) to install it into ~/.claude/hooks/.
Supported hook events (24): SessionStart, SessionEnd, UserPromptSubmit, PreToolUse, PostToolUse, PostToolUseFailure, SubagentStart, SubagentStop, Stop, StopFailure, TaskCompleted, TaskCreated, Notification, TeammateIdle, PreCompact, PostCompact, PermissionRequest, PermissionDenied, InstructionsLoaded, ConfigChange, CwdChanged, FileChanged, Elicitation, ElicitationResult.
Enriched metadata fields: When available in the hook payload, the dashboard displays additional context: SessionStart.source (startup/resume/clear/compact), StopFailure.error_type (rate_limit/authentication_failed/etc.), FileChanged.change_type (created/modified/deleted), PreCompact.trigger / PostCompact.trigger (manual/auto), and permission_mode on tool events. Fields display only when present — older Claude Code versions that don't emit them are fully supported.
Known gaps: The --bare CLI flag skips all hooks — sessions started with claude --bare will not appear in the dashboard. WorktreeCreate / WorktreeRemove hooks are not registered due to a Claude Code bug where they break built-in worktree isolation.
Tracked tools (35): Bash, PowerShell, Read, Write, Edit, MultiEdit, NotebookEdit, Glob, Grep, WebFetch, WebSearch, Task, Agent, TaskCreate, TaskUpdate, TaskGet, TaskList, TaskOutput, TaskStop, TodoWrite, CronCreate, CronDelete, CronList, EnterPlanMode, ExitPlanMode, EnterWorktree, ExitWorktree, ListMcpResourcesTool, ReadMcpResourceTool, AskUserQuestion, ToolSearch, Skill, LSP, Monitor, PushNotification.
Large result badge: Claude Code v2.1.91+ allows MCP servers to set _meta["anthropic/maxResultSizeChars"] on PostToolUse when a tool response is permitted to exceed the default cap (up to 500K chars). Affected events display a [large result] prefix in the dashboard so oversized payloads stand out at a glance.
PreCompact decision:"block": Claude Code v2.1.105 lets a hook block context compaction by returning {"decision":"block"}. Agent Recon hooks are telemetry-only and never return decisions, so this is observation only — when compaction is blocked, the absence of a paired PostCompact in the timeline is the visible signal.
Statusline widget: Agent Recon's wrapped statusline command forwards three subobjects from Claude Code's per-tick statusline JSON to POST /api/statusline: rate_limits (Pro/Max only — 5-hour and 7-day usage windows, F11.3), effort (current reasoning effort level on supported models, F13, v2.1.119+), and thinking (extended-thinking on/off, F13, v2.1.119+). A compact indicator in the header shows whichever fields are present, and a Mode subsection on the Environment tab spells out effort + thinking. When Claude Code goes idle the statusline stops ticking; rather than disappear, the indicator keeps its last-known values but dims them and appends an “as of HH:MM” marker — quota windows can reset, so aged figures are flagged rather than shown as current. Community / Team API users don't receive rate_limits but still see effort and thinking when the active model emits them.
MCP-tool hook type (v2.1.118): Claude Code now lets hooks invoke MCP tools directly via {"type":"mcp_tool","tool":"<mcp-tool-name>"} entries in settings.json, in addition to the existing command hook type. This is a Claude Code-side configuration feature — Agent Recon's hooks remain type:"command" POSTing to localhost:3131/event, and the wire-format hook events (PreToolUse, PostToolUse, etc.) are unchanged. No Agent Recon code change is required to coexist with users who configure their own mcp_tool hooks alongside ours.
OTEL trace correlation: when TRACEPARENT / TRACESTATE are set in the hook subprocess environment (e.g. by an upstream OTEL collector), send-event.js forwards them as HTTP headers on every event POST. The server validates the W3C format (55 chars, 00-<32hex>-<16hex>-<2hex>), splits into otel_trace_id and otel_span_id, and persists them on the event row. Malformed or oversized headers are silently dropped. CLAUDE_CODE_SUBPROCESS_ENV_SCRUB users are unaffected — the env vars never reach the hook subprocess.
Per-tool execution duration: Claude Code v2.1.119+ ships a duration_ms field on every PostToolUse and PostToolUseFailure hook (tool execution time only — excludes permission prompts and PreToolUse hook time). Agent Recon validates the value, persists it as tool_duration_ms on the event row, and appends a · 150ms / · 2s suffix to the tool-card preview so latency is visible at a glance.
Agent Recon™ uses both deterministic classifiers and LLM-powered analysis. Accuracy has been independently verified:
LLM-powered features (security chain analysis, session insights, environment recommendations) use AI models that may produce inaccurate results. These features display disclaimers in the UI. Always review AI-generated classifications and recommendations manually before acting on them.
Full metrics: see ACCURACY-REPORT.md in the repository.
Agent Recon™ exposes a local HTTP + WebSocket API for integration and scripting. All endpoints are served on http://localhost:PORT (default 3131). The server only accepts connections from private IP ranges (127.x, 10.x, 172.16–31.x, 192.168.x), and additionally requires the Host header to be localhost, a loopback address, or one of the addresses the server actually bound (a WSL / Hyper-V gateway IP, for example). That second check is what stops a web page you visit from reaching this API by pointing its own domain name at your machine. In practice it means the dashboard must be opened at localhost or by IP address — reaching it by machine name will return 403. Event ingestion (POST /event) is exempt so hook forwarding can never break silently.
| Method | Path | Description |
|---|---|---|
| POST | /event | Ingest a hook event from Claude Code |
| GET | /events | Retrieve stored events (supports limit, offset, session_id, since) |
| DELETE | /events | Clear all events, sessions, and counters |
| GET | /api/events/:id/decrypt | Decrypt an event's encrypted-raw payload (requires encrypted-raw PII storage mode) |
| Method | Path | Description |
|---|---|---|
| GET | /health | Health check (event count, DB count, WS clients) |
| GET | /diag | Server diagnostics (uptime, dedup stats, env breakdowns) |
| GET | /api/version | App version from package.json |
| GET | /api/db/stats | Database stats and Usage API poll status |
| Method | Path | Description |
|---|---|---|
| POST | /api/statusline | Ingest Claude Code statusline state (rate_limits / effort / thinking) from the wrapped statusline command. Requires a localhost Host header (DNS-rebind defense), 4 KB body cap, strict per-field shape validation. Server coalesces unchanged values and broadcasts rate_limits_update over WebSocket. |
| Method | Path | Description |
|---|---|---|
| GET | /api/tokens | Per-session token usage summaries |
| GET | /api/tokens/period | Today + monthly cost (estimated and verified) |
| Method | Path | Description |
|---|---|---|
| GET | /api/sessions/history | All sessions with stats and LLM-generated history |
| POST | /api/sessions/history/reanalyze | Re-queue all sessions for history analysis (paid) |
| GET | /api/sessions/:id/export | Full forensic export of a session as JSON |
| Method | Path | Description |
|---|---|---|
| GET | /api/security/status | Security LLM pipeline diagnostics |
| GET | /api/security/chains | Security chains for a session (?session_id=...) |
| GET | /api/security/events/:session_id | Raw security events for a session |
| POST | /api/security/reanalyze | Re-analyze all sessions for security chains (paid) |
| Method | Path | Description |
|---|---|---|
| GET | /api/insights | All sessions with LLM-generated insights |
| GET | /api/insights/:session_id | Full insights for a single session |
| POST | /api/insights/:session_id/refresh | Force re-analysis of session insights (paid) |
| Method | Path | Description |
|---|---|---|
| GET | /api/sessions/:id/prompt-analysis | Prompt quality analysis for a session |
| GET | /api/prompt-analysis/trends | Prompt quality trends over time (?days=30) |
| GET | /api/prompt-analysis/summary | Status summary across all sessions |
| Method | Path | Description |
|---|---|---|
| GET | /api/environment | Recommendations, usage stats, dismissed list |
| POST | /api/environment/reanalyze | Trigger fresh environment analysis (paid) |
| GET | /api/environment/resolved | Recently auto-resolved recommendations |
| POST | /api/environment/dismiss | Dismiss a recommendation by title |
| Method | Path | Description |
|---|---|---|
| GET | /api/settings | All settings (API keys sanitized) |
| PATCH | /api/settings | Update a single setting ({ key, value }) |
| POST | /api/settings/test-provider | Test LLM provider connectivity |
| Method | Path | Description |
|---|---|---|
| GET | /api/archives | List encrypted archives and key fingerprint |
| GET | /api/archives/:filename/decrypt | Download and decrypt a named archive |
| GET | /api/archives/decrypt-script | Download offline decryption helper script |
| Method | Path | Description |
|---|---|---|
| GET | /api/license | Current license tier and feature availability |
| POST | /api/license | Activate a license key |
| DELETE | /api/license | Deactivate the current license |
| POST | /api/license/validate | Force re-validation against Lemonsqueezy API |
Connect to ws://localhost:PORT. The browser appends ?since=ISO_TIMESTAMP (today’s midnight) to request today’s events. On connect, the server sends a history message with matching events, session metadata, and security chains.
| Type | Description |
|---|---|
history | Full event history, sessions, and security chains (sent on connect) |
event | New hook event arrived via POST /event |
clear | All events cleared (DELETE /events was called) |
security_chains | LLM security analysis completed for a session |
usage_update | Anthropic Usage API poll completed |
session_insights | LLM insights pipeline completed for a session |
environment_update | Environment analysis pipeline completed |
session_history_update | Session history LLM analysis completed |
prompt_quality_update | Prompt quality analysis completed for a session |
context_update | Context-observability snapshot recomputed for a session |
context_window_update | Per-session context-window fill % from the statusline (F4.2) |
context_window_clear | A session's statusline went idle — hide its context-window gauge |
license_update | License activated, deactivated, or re-validated |
rate_limits_update | Statusline state change — payload data carries any of five_hour / seven_day (Pro/Max), effort.level, thinking.enabled. Broadcast only when at least one tracked value changed. |
rate_limits_clear | 5+ minutes have passed without a statusline tick — widget should hide. |
# Send a test event
curl -X POST http://localhost:3131/event \
-H "Content-Type: application/json" \
-d '{"session_id":"test-123","hook_event_name":"PreToolUse","tool_name":"Bash","tool_input":{"command":"echo hello"}}'
# Retrieve the last 50 events
curl "http://localhost:3131/events?limit=50"
# Get security chains for a session
curl "http://localhost:3131/api/security/chains?session_id=abc-123"
For complete API documentation including request/response schemas, see docs/API.md in the project repository.