/** * What a thread remembers about its own size, when that size is a problem, and * what it does about it. * * Four small pieces, one job: keep a long conversation inside the model's window * without anybody having to notice. The state codec is what survives between * turns; the estimate is how big the loop believes this turn's prompt is; the * trigger is the line it must not cross; and the summarizer is what happens when * it does — one pass, at the start of a turn, invisible to the user. * * The estimate measures ONE thing, THIS turn: the prompt that is about to be * sent — system prompt, tools block, projected messages — in characters, over a * single pessimistic characters-per-token ratio. Nothing is carried between * turns, and that is the whole of a bug class this shipped four times. * * The engine used to carry the provider's reported count forward (pi-mono's * estimator does, `packages/agent/src/harness/compaction/compaction.ts`, MIT, * Mario Zechner) and guess only the delta on top of it. The count describes what * the last turn SENT; the trigger decides about what the thread STORES; and after * a compaction those are different quantities — the prompt was a summary and a * tail, and the transcript was never truncated. One variable held both, so the * trigger read a compacted turn's small count as a fact about a large thread. It * went blind for the life of the thread when the figure was carried, and it * ALTERNATED — compact, ship the whole paste, compact, ship it again — when the * figure was dropped and the guess underneath was too optimistic. Attributing the * count to the messages it covered is the same confusion wearing bookkeeping: a * seed turn's 9,483 was credited with covering two 100,000-character statements. * A measurement taken fresh, of the thing in front of it, cannot be wrong about * which thing it measured. * * No tokenizer: a per-provider vocabulary is megabytes, is wrong for every model * it was not built for, and would have to load before the first turn. So the * ratio is pessimistic instead of accurate ({@link PESSIMISTIC_CHARS_PER_TOKEN}), * and the trigger sits at 81% so the margin is not the only thing standing * between a thread and a 400. */ import { type LanguageModel, type ModelMessage, type ToolSet, type UIMessage } from "ai"; export interface CompactionState { version: 1; summary?: string; /** Tools searched in over the life of this thread — vendo()'s loadout memory * (`find_tools` adds here; the loadout offers these on every later turn). * Additive: rows written before this field read fine without it, and it * rides the same slot so §1.3's clearing rules govern it too. */ loadedTools?: string[]; /** `id` of the newest UIMessage {@link summary} ABSORBED. Everything after it is * the verbatim tail, so the next turn rebuilds the same projection — summary, * then the messages the summary never read — instead of re-summarizing the * whole transcript. That is what makes compaction converge: without it a thread * whose bulk is one huge paste pays a summarizer pass on EVERY turn, reading * the whole paste each time, which costs about what simply sending it would. * * Ids, not indexes, because the store's rows are id-keyed and an index means * nothing after an edit. Not found in `turn.messages` = the thread has been * rewound or edited past what this state describes = the whole state is * DISCARDED and the turn measures the full transcript, which errs toward * compacting. Never carried into a branch that no longer holds the history it * was built from. */ boundaryMessageId?: string; } /** * Decode the thread's slot. * * The slot is opaque by contract (`turn.state`, build contract §1.3) and can hold * anything: a string written by a future version of this file, a foreign * harness's native session id, half a row a store lost. Every one of those reads * as "no state" rather than as a shape the loop then trusts — losing the state * costs one un-compacted turn, and trusting a bad one costs a prompt nobody can * predict. */ export declare function readCompactionState(slot: string | undefined): CompactionState | undefined; export declare function writeCompactionState(state: CompactionState): string; /** * Ported from cline `sdk/packages/core/src/extensions/context/compaction-shared.ts:15,17` * (Apache-2.0): `CONTEXT_WINDOW_INPUT_RATIO = 0.9` × `COMPACTION_TRIGGER_RATIO = 0.9`. * * Two multiplied margins, and both are load-bearing. The first keeps the ANSWER's * room: a prompt that fills the window leaves nowhere for the model to reply. The * second is the compaction headroom: the summarizer pass itself is a call against * the same window, so a trigger that waits for the window to be full has already * lost — the turn that discovers the problem is the turn that 400s. */ export declare const TRIGGER_RATIO = 0.81; /** * Characters per token, chosen to OVER-count rather than to be right. * * The engine assumed four for years, which is within a few percent on English * prose and JSON and is where the whole dead path started: the one shape * compaction exists for is a pasted statement or a 300KB tool result, and dense * text is nothing like prose. Measured on this repo's own walker thread — * 308,000 characters of statement text that the provider billed at 142,890 * tokens — the real figure is 2.156 characters per token, 1.83x denser than the * assumption. Two is the round number at least as conservative as the worst case * anything here has measured: it prices that same text at 154,000, above what it * actually cost, and it prices ordinary prose at roughly twice its cost. * * That second half is the price of the deletion, stated plainly: a prose thread * compacts at around 40% of the real window instead of 81%, so it summarizes * earlier and keeps less verbatim history than it strictly has to. That is the * cheap direction. The expensive direction is a prompt the provider rejects, and * an over-count cannot produce one. */ export declare const PESSIMISTIC_CHARS_PER_TOKEN = 2; /** * Ported from cline `sdk/packages/core/src/extensions/context/compaction-shared.ts:19` * (Apache-2.0): the verbatim tail a compaction always preserves. * * Declared here with the ratio it belongs beside; the cut point that reads it * arrives with the summarizer. */ export declare const PRESERVE_RECENT_TOKENS = 20000; /** The ceiling on one summarizer pass. Declared here, read by the summarizer. */ export declare const SUMMARY_MAX_OUTPUT_TOKENS = 2000; export interface CompactionConfig { contextWindowTokens: number; triggerRatio?: number; preserveRecentTokens?: number; } /** THE conversion, and the only one. Exported because `loop.ts`'s shed floor * charges its candidates with it too: two rails over one prompt, denominated * differently, is how the trigger came to say "over budget" and the floor "fits" * about the same 308,000 characters and neither of them acted. */ export declare const tokensFor: (chars: number) => number; /** * What THIS turn's prompt costs: the system prompt, the messages it projects and * the tools block it carries, all of it measured now. * * There is no history in this function and there is no state behind it. That is * deliberate — see the file header. */ export declare function estimatePromptTokens(input: { system: string; messages: readonly ModelMessage[]; tools: ToolSet; }): number; /** The estimate at which a turn must act. */ export declare function triggerTokens(config: CompactionConfig): number; export declare function shouldCompact(promptTokens: number, config: CompactionConfig): boolean; /** * Where the verbatim tail starts: everything below this index becomes summary, * everything from it survives word for word. * * Ported from cline `compaction-shared.ts:326-359` (Apache-2.0). Two rules, * applied in order, and each is a bug somebody already shipped: * 1. walk back from the newest message taking every one that still FITS inside * `preserveRecentTokens` — the tail is a token budget, not a message count, * because one tool result can outweigh forty exchanges; * 2. never cut past the newest user turn's start, so the ask the user is in the * middle of survives verbatim however small the budget is. * * cline has a third — walk back to a boundary that cannot orphan half of a * tool-call/tool-result pair — and this cuts in UIMessage space precisely so that * rule has nothing left to do. A tool call and its result are PARTS OF ONE * UIMessage here, so no boundary between two of them can separate them; the same * argument the host's `historyWindow` slice already runs on. Cutting in * ModelMessage space needed the rule because `ai` splits one UIMessage into an * assistant message and a `role: "tool"` message that must not be divided. * * The coordinate system is load-bearing for a second reason: the cut is what the * thread PERSISTS ({@link CompactionState.boundaryMessageId}), so it has to be * expressible as a stable id. UIMessages have ids and the store's rows are * id-keyed; converted ModelMessages have neither, which is why a boundary derived * from them could not be re-resolved on the next turn — and without that the * projection cannot be rebuilt and every turn pays a summarizer pass. * * Rule 1 stops BEFORE the message that tips the budget, and that word is the * whole of a defect this shipped with. The walk used to absorb the tipping * message, so one message bigger than the entire tail — a pasted statement, a * 300KB tool result — put the cut on index 0 and the caller read that as * "nothing to summarize". On a thread whose bulk is one message, which is the * single shape compaction exists for, the trigger then fired every turn and the * projection never changed. An oversized message belongs to the SUMMARY. * * The tail is never empty: the newest message is kept verbatim even when it * alone is oversized, because a cut past it would summarize the ask the turn is * answering. A thread that is nothing BUT one oversized message therefore cuts * at 0 — there is nothing above it to summarize, and the shed floor underneath * (which cannot drop a thread's last message either) sends it and lets the * provider's own refusal be the honest answer. */ export declare function findCutIndex(messages: readonly UIMessage[], preserveRecentTokens: number): number; /** * The summary as the model sees it. Ported from pi-mono * `packages/agent/src/harness/messages.ts:4-10` (MIT, Mario Zechner). * * Two properties, both load-bearing. It is a USER message, which is what makes * every projection assistant-first-safe by construction — the one prompt shape a * provider rejects outright cannot occur when the first non-system message is * always this one. And it is FENCED: the summary is a record of what happened, * not a directive, so a summarizer that copied an injected imperative into its * output hands the resident a quoted string rather than an order. * * The closing line is the silence rule (Design D3): the user never asked for * this and must never be told it happened. */ export declare function summaryMessage(summary: string): ModelMessage; export interface CompactionRequest { /** The BAND to absorb — already cut, because the cut is the caller's: it decides * the boundary in UIMessage space so the thread can persist it, and this is * handed the converted result. */ messages: readonly ModelMessage[]; /** The summary this thread already carries — fed back so ONE pass UPDATES it * rather than re-reading history it no longer holds (pi `compaction.ts:545`). */ summary?: string; /** D1: the thread's own resident seat. */ model: LanguageModel; config: CompactionConfig; signal?: AbortSignal; } export interface CompactionResult { summary: string; usage: { inputTokens: number; outputTokens: number; }; } /** * ONE summarizer pass. * * The isolation is the interesting part. This call carries NO tools and NO cache * breakpoint, which is pi's `cacheRetention: "none"` plus a fresh session id * (`compaction.ts:110-114`, MIT) expressed in our stack. No tools because a * summarizer that can act is an injected tool result away from acting; no cache * marker because the turn's own cached prefix is worth more than this one-off * request, and a breakpoint here would evict it for a prompt nothing will ever * send again. * * The history goes in as TEXT inside one fenced user message rather than as real * messages, following pi (`compaction.ts:549-563`): what the summarizer receives * is a transcript to read, not a conversation it is party to. */ export declare function compactContext(request: CompactionRequest): Promise; //# sourceMappingURL=compaction.d.ts.map