/** * K-means clustering — Lloyd's algorithm with greedy k-means++ seeding. * * Faithful port of scikit-learn's `sklearn.cluster.KMeans` (algorithm="lloyd"): * - greedy k-means++ initialization with 2 + floor(log k) local trials * (Arthur & Vassilvitskii 2007; sklearn `_kmeans_plusplus`, _kmeans.py l.180) * - Lloyd expectation-maximization with strict-label and center-shift * tolerance convergence (sklearn `_kmeans_single_lloyd`, _kmeans.py l.630) * - `nInit` independent restarts, keeping the lowest-inertia result * - dataset-scaled tolerance `mean(var(X, axis=0)) * tol` (sklearn `_tolerance`) * * Determinism is total: the run is driven by a seeded mulberry32 PRNG. With a * fixed `seed` the labels, centers and inertia are reproducible bit-for-bit; * there is deliberately NO `Date.now()`/`Math.random()` fallback. * * Validated against committed reference fixtures * (three separable blobs, k=3, generated by sklearn.cluster.KMeans). * * @param {Array>} X2d * Observations, shape (nSamples, nFeatures). Each row is a plain array or a * typed array; all rows must share the same length. * @param {number} k - number of clusters (1 ≤ k ≤ nSamples). * @param {Object} [options] * @param {number} [options.nInit=10] - number of k-means++ restarts (≥ 1). * @param {number} [options.maxIter=300] - max Lloyd iterations per restart (≥ 1). * @param {number} [options.seed=0] - PRNG seed for reproducible seeding. * @param {number} [options.tol=1e-4] - relative center-shift tolerance * (scaled by the mean feature variance, matching sklearn). * @returns {{ labels: Int32Array, centers: number[][], inertia: number }} * `labels[i]` is the cluster index of observation i, `centers[c]` is the * centroid of cluster c, and `inertia` is the summed squared distance of * every observation to its assigned centroid. */ export function kmeans(X2d: Array>, k: number, { nInit, maxIter, seed, tol }?: { nInit?: number; maxIter?: number; seed?: number; tol?: number; }): { labels: Int32Array; centers: number[][]; inertia: number; };