import { Document } from "@langchain/core/documents"; import { ConfigService } from "@nestjs/config"; import { TokenUsageRecorderInterface } from "../../../common/tokens"; import { BaseConfigInterface } from "../../../config/interfaces/base.config.interface"; import { ModelService } from "../../../core/llm/services/model.service"; import { S3Service } from "../../s3/services/s3.service"; import { TokenUsageService } from "../../tokenusage/services/tokenusage.service"; import { DocXService } from "./types/docx.service"; import { EmailParserService } from "./types/email.service"; import { MarkdownChunkingService } from "./types/markdownchunking.service"; import { PdfService } from "./types/pdf.service"; import { PptxService } from "./types/pptx.service"; import { SemanticSplitterService } from "./types/semanticsplitter.service"; import { XlsxService } from "./types/xlsx.service"; /** * Opt-in cost attribution for the LLM calls the chunker makes while reading a * file (today: image analysis). Without both `relationshipId` and * `relationshipType` no usage record is written — the package stays * domain-agnostic and the caller decides what the usage is billed against. * Mirrors `EmbedderAttribution`. */ export interface ChunkerAttribution { relationshipId: string; relationshipType: string; } export declare class ChunkerService { private readonly markdownChunkingService; private readonly semanticSplitterService; private readonly docxService; private readonly pptxService; private readonly pdfService; private readonly xlsxService; private readonly modelService; private readonly s3Service; private readonly emailParserService; private readonly config; private readonly tokenUsageService?; private readonly tokenUsageRecorder?; private logger; private readonly splitter; private readonly targetChars; constructor(markdownChunkingService: MarkdownChunkingService, semanticSplitterService: SemanticSplitterService, docxService: DocXService, pptxService: PptxService, pdfService: PdfService, xlsxService: XlsxService, modelService: ModelService, s3Service: S3Service, emailParserService: EmailParserService, config: ConfigService, tokenUsageService?: TokenUsageService, tokenUsageRecorder?: TokenUsageRecorderInterface); private _logChunkResult; private _downloadFileAsBuffer; private _downloadFile; /** * Page count for formats that carry no pagination of their own (everything * except PDFs and images). Deliberately coarse — it exists so per-page costs * (OCR, document AI) can be attributed to a document of any format. */ private _estimatePageCountFromText; private _estimatePageCount; /** * Stamps the document's page count on every chunk it produced. The count * travels on metadata because the return type of the public entry point is * `Document[]` and must stay that way for existing callers. `extra` carries the * optional read-quality flags (today: `ocrTruncated`) alongside it. */ private _stampTotalPages; /** * Enforces `MAX_CHUNK_CHARS` on the chunks about to leave the chunker. * * Splitters are best-effort: given text with no separator near the target size * they emit a single huge chunk, which only fails much later (embedding call, * LLM context). Splitting here is the last point where the oversize is still * cheap to fix. `splitDocuments` carries the metadata onto the pieces, so the * page count stamped upstream survives; the slice is a guarantee for text with * no boundary to split on at all. */ private _capChunkChars; generateContentStructureFromFile(params: { fileType: string; filePath: string; attribution?: ChunkerAttribution; /** * Human-readable name of the document, used as the markdown title. Pass it whenever * the caller knows it: `filePath` is often a PRESIGNED URL, and deriving a title from * it put the signature and credential into the chunk text (and into every prompt and * embedding built from it). Falls back to the file name when omitted. */ title?: string; }): Promise; private _processLocalFile; private _loadLocalFileDocuments; private _createFromEmail; private _extractAttachmentContent; generateContentStructureFromMarkdown(params: { content: string; title?: string; }): Promise; extractContentFromUrl(params: { url: string; }): Promise; /** * Image analysis is a vision LLM call made inside the package, so — exactly * like `EmbedderService.persistUsage` — its usage is written through the * optional `TOKEN_USAGE_RECORDER` seam, attribution is opt-in, and a failure * to record never fails the chunking itself. */ private _persistImageUsage; private _createChunksFromImage; private _createFromMarkdown; private _createFromDocX; private _createFromPptx; private _createFromXlsx; /** * A PDF's own page count is authoritative — except when it clearly is not. * Broken or unreadable document metadata surfaces as "1 page", and a 50-page * PDF billed as one page under-reports per-page cost by a factor of 50, so a * reported single page carrying pages of text is replaced by the estimate. */ private _resolvePdfPageCount; private _createFromPdf; } /** * The file name of a path or URL, with any query string or fragment removed. * * `filePath` reaching the chunker is frequently a presigned URL whose last segment is * `.?X-Amz-Algorithm=…&X-Amz-Credential=…&X-Amz-Signature=…`. Splitting on * "/" alone therefore yields a "file name" carrying the credential and signature, and * no extension strip matches it because the string no longer ends in the extension. */ export declare function fileNameOf(filePath: string): string; //# sourceMappingURL=chunker.service.d.ts.map