/** * product-kb/chunker — H2-section-aware markdown chunker. * * Why this exists * * The corpus is mkdocs Material markdown. The dominant structure is: * ### :material-info: Overview * * ??? tenx-overview "What is X" * * * * ### :material-cog: Capabilities * ... * * Naive line-based chunking breaks this in three ways: * * 1. The `??? tenx-overview "..."` admonition opens an indented block. * A naive splitter would cut the heading away from its body. * 2. Mermaid fences (```mermaid ... ```) carry headings inside the diagram * that look like prose H3 headings to a regex. * 3. Jinja includes ({% include "_partial.md" %}) are opaque — we never * want to split mid-tag. * * The chunker splits on H2 (## ) AND H3 (### ) when they appear at column 0 * AND are NOT inside a fenced code block / admonition body / jinja include. * H3 is treated as a sub-section split because the mkdocs Material pages * in this corpus dominantly use ### as the top-level user-visible heading * (the H1 lives in frontmatter `title:`). * * Each emitted chunk also carries a brief overlap tail from the next * section (~200 chars max, trimmed at the next paragraph or sentence * boundary). The overlap reduces "phrase straddles a section break" * recall misses without inflating the chunk count. */ import type { Chunk } from './types.js'; /** * Split a markdown body into H2/H3-bounded chunks, with a brief tail * overlap from the following section. * * The body is the post-frontmatter portion of the page. The caller is * responsible for stripping the frontmatter before calling this. * * @param body Markdown body (no frontmatter). * @param topic Page topic slug, used to build chunk_id. */ export declare function chunkMarkdown(body: string, topic: string): Chunk[];