Global

Members

WFetchWeb

Description:
  • 網頁抓取

Source:

網頁抓取

Example
詳見fetchWeb、fetchWebByCurl、fetchWebByPlaywrightHeadless、fetchWebByPlaywrightHead、fetchWebByCamofox範例

Methods

adapt(r) → {Object}

Description:
  • 把抓取函數之對外結果轉為內部流程用之結構

    四個fetchWebByXxx回傳{status:'success'|'error',...},而執行計畫內部以{success:Boolean,...}判斷, 由本函數轉換以隔離兩者

Source:
Example
import { adapt } from './src/finalizeResult.mjs'

console.log(adapt({ status: 'success', html: '<p>a</p>', method: 'curl' }))
// => { success: true, html: '<p>a</p>', method: 'curl', snapshot: undefined, contentKind: undefined }
Parameters:
Name Type Description
r Object

輸入抓取函數之結果物件

Returns:

回傳內部結構物件,成功時為{success:true,html,method,snapshot,contentKind},失敗時為{success:false,method,reason,message}

Type
Object

buildPlan(url, method, doInspect, hasFetchAdapteropt) → {Object}

Description:
  • 依網址與選項產生執行計畫

    四條網域分流與指定方法,差別僅在於「要跑哪幾階」,故一律表達為step陣列, 交由同一個runPlan執行,不再各自複製一份「抓取至判識至解析」流程。

    本函數只回傳純資料:不捕捉url與opt、不建立closure、不修改任何共享狀態。 轉址旗標之初值以redirect欄位回報,由runPlan擁有其後續變化

Source:
Example
import buildPlan from './src/buildPlan.mjs'

console.log(buildPlan('https://example.com/', 'auto', true).plan.map((s) => s.key))
// => ['curl', 'headless', 'headed', 'camofox']

console.log(buildPlan('https://mp.weixin.qq.com/s/a', 'auto', true).plan.map((s) => s.key))
// => ['camofox']
Parameters:
Name Type Attributes Default Description
url String

輸入待抓取網址字串

method String

輸入方法字串,'auto'代表自動階梯升級

doInspect Boolean

輸入是否做原始內容判識布林值,false時各階之inspect一律關閉

hasFetchAdapter Boolean <optional>
false

輸入命中之adapter是否具fetch掛點布林值,true時於計畫最前插入adapter階,預設false

Returns:

回傳{plan,redirect,log}物件,plan為step描述陣列,redirect為轉址旗標初值,log為分流說明字串;method不合法時回傳{error}

Type
Object

checkFetched(r, reason, pre) → {Object|null}

Description:
  • 檢核抓取器之回傳值是否符合輸出契約

    不合契約者轉為結構化失敗結果而非拋錯,使fetchWeb之「不會reject」契約於任何抓取器實作下皆成立

Source:
Example
import { checkFetched } from './src/fetchContract.mjs'

console.log(checkFetched({ status: 'success', html: '<p>a</p>' }, 'fetcher-error', 'curl '))
// => null

console.log(checkFetched('oops', 'fetcher-error', 'curl '))
// => { status: 'error', reason: 'fetcher-error', message: 'curl returned invalid result' }
Parameters:
Name Type Description
r *

輸入抓取器之回傳值

reason String

輸入不合契約時採用之失敗歸因字串

pre String

輸入錯誤訊息前綴字串

Returns:

合契約時回傳null,不合時回傳{status:'error',reason,message}

Type
Object | null

checkHttpStatus(httpCode) → {Object|null}

Description:
  • 依HTTP狀態碼判定是否為失敗,並給出失敗歸因與可重試性

    四個抓取器共用同一判準,不各自實作

Source:
Example
import { checkHttpStatus } from './src/httpStatus.mjs'

console.log(checkHttpStatus(200))
// => null

console.log(checkHttpStatus(404))
// => { reason: 'http-error', message: 'HTTP 404', httpCode: 404, retryable: false }

console.log(checkHttpStatus(503).retryable)
// => true
Parameters:
Name Type Description
httpCode Integer

輸入HTTP狀態碼整數,取不到時傳0

Returns:

非失敗時回傳null;失敗時回傳{reason,message,httpCode,retryable}

Type
Object | null

estimateVisibleText(html) → {String}

Description:
  • 由網頁HTML估算可見文字內容

    取內容後移除script、style、標籤與HTML實體,再壓縮空白,用以粗估頁面實質可見字數, 供轉址殼頁、空內容頁之判識使用

Source:
Example
import estimateVisibleText from './src/estimateVisibleText.mjs'

console.log(estimateVisibleText('<html><body><p>abc</p> <p>def</p></body></html>'))
// => 'abc def'

console.log(estimateVisibleText('<body><script>var a=1</script><div>xyz</div></body>'))
// => 'xyz'
Parameters:
Name Type Description
html String

輸入網頁HTML字串

Returns:

回傳估算之可見文字字串

Type
String

extractHtmlTitle(html) → {String}

Description:
  • 由網頁HTML取出標題,並剝除尾端之「-站名」

Source:
Example
import extractHtmlTitle from './src/extractHtmlTitle.mjs'

console.log(extractHtmlTitle('<html><head><title>流動性溢價的真相 - 格隆匯</title></head></html>'))
// => '流動性溢價的真相'
Parameters:
Name Type Description
html String

輸入網頁HTML字串

Returns:

回傳標題字串,取不到時回傳空字串

Type
String

(async) extractPageContent(page) → {Promise}

Description:
  • 由Playwright之page提取網頁HTML,可見文字過少時穿透Shadow DOM

    先取page.content(),若估算可見文字已達門檻則直接回傳;否則遞迴穿透Shadow DOM取得 各層innerText,再依換行切段重組為含article之簡易HTML,供Readability解析。

    回傳一併標明內容形態:合成內容之標籤結構已被剝除,其可信的判識證據與原始文件不同, 故由本函數(唯一知道走了哪個出口者)標明,而非由下游猜測

Source:
Example
import { chromium } from 'playwright'
import extractPageContent from './src/extractPageContent.mjs'

let test = async () => {

    let browser = await chromium.launch({ headless: true, channel: 'chrome' })
    let page = await browser.newPage()
    await page.goto('https://example.com/')
    let { html, contentKind } = await extractPageContent(page)
    console.log(html.length, contentKind)
    // => 234 'raw'
    await browser.close()

}
await test()
    .catch((err) => {
        console.log(err)
    })
Parameters:
Name Type Description
page Object

輸入Playwright之page物件

Returns:

回傳Promise,resolve回傳{html,contentKind}物件,contentKind為'raw'或'synthesized'

Type
Promise

extractRedirectTarget(url) → {String|null}

Description:
  • 由轉址服務之網址中提取其query參數所帶之真實網址

    僅處理已知會把目標網址放在query參數之服務;取出後之值已由URLSearchParams解碼一次, 再解一次以處理雙重編碼,若該值非合法百分比序列則退回僅解碼一次之結果。

    提取出之目標須通過內網位址檢核,指向迴環、私有網段、link-local或非http/https者一律不提取

Source:
Example
import { extractRedirectTarget } from './src/routeByUrl.mjs'

console.log(extractRedirectTarget('https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fa.com%2Fb'))
// => 'https://a.com/b'

console.log(extractRedirectTarget('https://www.linkedin.com/redir/redirect?url=http%3A%2F%2F169.254.169.254%2F'))
// => null
Parameters:
Name Type Description
url String

輸入網址字串

Returns:

命中且目標可用時回傳目標網址字串,未命中或目標指向內網時回傳null

Type
String | null

(async) fetchMsn(url, opt, ctx) → {Promise}

Description:
  • 經 msn 內容 API 取得內容,重組為 HTML 文件

    作為內建 msn adapter 之 fetch 掛點。API 之請求走本套件之 curl 抓取器, 故 User-Agent、重試與 HTTP 狀態判準與其餘抓取一致;opt 原樣轉傳,並沿用 opt._fetchers.curl 測試接縫

Source:
Example
import { matchMsn, fetchMsn } from './src/fetchMsn.mjs'

let test = async () => {
    let url = 'https://www.msn.com/zh-tw/news/other/abc/ar-AA2bZm9d'
    let r = await fetchMsn(url, {}, matchMsn(url))
    console.log(r.status, r.contentKind)
    // => 'success' 'synthesized'
}
await test()
    .catch((err) => {
        console.log(err)
    })
Parameters:
Name Type Description
url String

輸入內容頁網址字串,本函數僅用於訊息

opt Object

輸入設定物件,轉傳給 curl 抓取器

ctx Object

輸入 matchMsn 之回傳物件{kind,id}

Returns:

回傳Promise,resolve回傳結果物件,成功時為{status:'success',html,contentKind:'synthesized'},失敗時為{status:'error',reason,message},本函數不會reject

Type
Promise

(async) fetchWeb(url, optopt) → {Promise}

Description:
  • 抓取網頁文章內容,支援四種抓取方法自動階梯升級

    抓取方法: 方法①curl(預設,繞過TLS指紋),委派fetchWebByCurl; 方法②Playwright無頭(SPA動態渲染頁面),委派fetchWebByPlaywrightHeadless; 方法③Playwright有頭(反自動化偵測),委派fetchWebByPlaywrightHead; 方法④Camofox反偵測瀏覽器(Cloudflare等),委派fetchWebByCamofox

    本函數僅負責階梯升級、內容判識、文章解析之調度,實際抓取由4個抓取函數執行, 流程為fetch(委派)至inspectHtml(原始內容檢測)至Readability解析(可選); 對已知網站另有轉址提取、跳過特定方法之判識規則,且方法④會額外回傳snapshot欄位

    並行呼叫之限制:方法④之Camofox server綁定固定埠(預設19377),同一埠號同時只能有一個抓取, 詳見fetchWebByCamofox之說明。auto模式可能升級至該階,故同時發動多個fetchWeb時, 若其中一個以上走到方法④即會互相破壞(先完成者殺掉server,其餘回'camofox-error')。 需要並行時須為每個呼叫指定互不相同的opt.port;方法①②③則無此限制

Source:
Example
import fetchWeb from './src/fetchWeb.mjs'

let test = async () => {

    //auto模式, 解析出文章標題與內文
    let r = await fetchWeb('https://example.com/')
    console.log(r.status, r.method, r.title, r.contentLength)
    // => 'success' 'curl' 'Example Domain' 111

    //指定curl且不解析, 直接取原始HTML
    let rh = await fetchWeb('https://example.com/', { method: 'curl', parse: false })
    console.log(rh.status, rh.html.length)
    // => 'success' 559

    //失敗時回傳error結果物件, 不會reject
    let re = await fetchWeb('abc')
    console.log(re.status, re.message)
    // => 'error' 'invalid url (must be http/https)'

}
await test()
    .catch((err) => {
        console.log(err)
    })
Parameters:
Name Type Attributes Default Description
url String

輸入待抓取網址字串

opt Object <optional>
{}

輸入設定物件,其餘鍵值會轉傳給實際執行抓取之函數,預設{}

Properties
Name Type Attributes Default Description
method String <optional>
'auto'

輸入指定抓取方法字串,可為'auto'、'curl'、'playwright'、'playwright-headed'、'camofox','auto'代表自動階梯升級,預設'auto'。本選項塑形的是「要爬的時候用哪一種爬法」,不影響adapter之fetch掛點是否執行(後者回答的是「這次要不要爬」)

parse Boolean <optional>
true

輸入是否以Readability解析出文章標題與內文布林值,false時直接回傳原始HTML,預設true

adapters Array <optional>
[]

輸入站台adapter物件陣列,用於覆寫特定站台之取得、判識與解析方式,預設[]。形狀為{id,match,fetch,parse,inspect,fallback},三個掛點分屬管線的三個階段且各自獨立:fetch取代爬取(呼叫端有官方API、快取或已登入session時)、inspect為false表此adapter命中時不做原始內容判識、parse覆寫解析;fallback為false表此adapter階失敗時(fetch失敗、被判識擋下或解析未取得足量正文)不落回階梯,以該階之reason收攤。parse另收第四參數meta(含requestUrl、finalUrl、httpCode、method、contentKind),供其分辨內容實際來自哪裡;其完整契約(輸入形狀、輸出檢核、錯誤邊界)以src/adapterContract.mjs為唯一事實來源。使用端adapters排於內建adapters(gelonghui、bloomberg、msn)之前,同網域以先命中者勝出,故可逐站覆寫之

useDefaultAdapters Boolean <optional>
true

輸入是否附加內建adapter清單布林值,預設true。false時只用opt.adapters所給者;配合對外匯出之defaultAdapters陣列,可自行剔除某一個(如adapters:defaultAdapters.filter((a)=>a.id!=='msn'))或重排

detectors Array <optional>
[]

輸入使用端判識器物件陣列,用於補充內建判識器所不涵蓋之攔阻頁形態(如中文與其他語系之挑戰頁),排於內建之前故優先命中,預設[]。形狀為{type,message,test},其完整契約以src/detectorContract.mjs為唯一事實來源

inspect Boolean <optional>
true

輸入是否以inspectHtml對抓取結果做原始內容判識布林值,預設true。關閉後不因判定為挑戰頁或空內容而升級。此為整次呼叫之總開關;若只想豁免特定站台,改以該站台adapter之inspect:false表達,範圍較精確且不必在呼叫端重寫一次網址判斷

_fetchers Object <optional>
null

輸入置換抓取函數之物件,僅供測試使用,鍵可為'curl'、'playwrightHeadless'、'playwrightHead'、'camofox',值為與對應fetchWebByXxx同簽章之函數,未給之鍵沿用實際實作,預設null

useShowLog Boolean <optional>
true

輸入是否顯示階梯升級過程訊息布林值,預設true

maxRetries Integer <optional>
5

輸入各抓取函數失敗時最大重試次數整數,含初始共執行maxRetries+1次,預設5

Returns:

回傳Promise,resolve回傳結果物件,其中attempts為各階嘗試紀錄陣列,成功之紀錄為{method,status:'success',htmlLength}(htmlLength為原始HTML長度,與頂層contentLength之正文長度不同),失敗為{method,status:'failed',reason,message},被判識或解析失敗為{method,status:'blocked',type,reason,message}(判識所致者reason與type同值);經adapter之fetch掛點者其紀錄之method為'adapter'且另帶adapterId,成功結果與於該階收攤之失敗結果其頂層亦帶adapterId(落回後階梯耗盡者則無);parse=true成功時為{status:'success',url,title,content,contentLength,method,fetchedAt,attempts},parse=false成功時為{status:'success',url,html,method,fetchedAt,attempts};內容來自轉址後之另一網址時另帶finalUrl(url為本套件最後實際發出請求之網址,finalUrl為內容實際來源,兩者相同時不輸出該欄),失敗時為{status:'error',url,message,fetchedAt,attempts},本函數不會reject

Type
Promise

(async) fetchWebByCamofox(url, optopt) → {Promise}

Description:
  • 使用Camofox反偵測瀏覽器抓取網頁原始HTML,透過accessibility snapshot取得內容

    流程: 以Node模組解析機制取得已安裝之@askjo/camofox-browser之server.js位置; spawn node <server.js> 啟動Camofox server; POST /tabs 建立tab; GET /tabs/:id/snapshot 取accessibility snapshot(含內部重試); DELETE /tabs/:id 關閉tab; 殺整棵server進程樹(Windows用taskkill /F /T,Unix以負PID對spawn時建立之行程群組送SIGTERM)

    對server之每次HTTP請求皆有15秒硬上限,逾時即abort,不因server無回應而永久等待; 單次嘗試之server啟動與清理皆由runCamofoxAttempt完成,故重試退避期間不會佔用該埠

    該套件之server.js無任何export且於top-level即無條件listen,故只能spawn為子行程執行, 不可直接import;解析不到安裝位置時回傳reason='camofox-not-found'。

    失敗歸因分三種:'camofox-not-found'為未安裝、'camofox-empty'為頁面確實無足量內容(重試無益)、 'camofox-error'為server啟動失敗、tab建立失敗或snapshot傳輸失敗(可重試)

    同一埠號同時只能有一個抓取在進行,本函數不可並行呼叫。 Camofox server綁定固定埠(預設19377),且就緒探測只確認該埠有服務回應、不驗證是否為自己啟動者。 故並行呼叫時:後啟動者因埠被占用而啟動失敗,卻會探到前者的server而誤認就緒並借用之; 待任一方先完成,其清理程序會殺掉該server,另一方隨即全數失敗(回'camofox-error')。 需要並行時,必須為每個同時進行的呼叫指定互不相同的opt.port。 此限制同樣適用於auto模式的fetchWeb——其階梯末階即為本函數。

Source:
Example
import fetchWebByCamofox from './src/fetchWebByCamofox.mjs'

let test = async () => {

    let r = await fetchWebByCamofox('https://mp.weixin.qq.com/s/xxxxxx')
    console.log(r.status, r.snapshotChars, r.htmlLength)
    // => 'success' 3215 4102

    let re = await fetchWebByCamofox('abc')
    console.log(re.status, re.reason)
    // => 'error' 'invalid-url'

}
await test()
    .catch((err) => {
        console.log(err)
    })
Parameters:
Name Type Attributes Default Description
url String

輸入待抓取網址字串

opt Object <optional>
{}

輸入設定物件,預設{}

Properties
Name Type Attributes Default Description
port Integer <optional>
19377

輸入Camofox server監聽埠號整數,預設19377。並行呼叫時每個呼叫須給不同埠號

serverStartTimeoutMs Integer <optional>
30000

輸入等待Camofox server啟動最長毫秒整數,預設30000

snapshotRetries Integer <optional>
3

輸入snapshot內容不足時之重取次數整數,預設3

snapshotWaitMs Integer <optional>
5000

輸入snapshot重取間隔毫秒整數,預設5000

maxRetries Integer <optional>
5

輸入失敗時最大重試次數整數,含初始共執行maxRetries+1次,預設5

Returns:

回傳Promise,resolve回傳結果物件,成功時為{status:'success',url,html,htmlLength,contentKind,snapshot,snapshotChars,method,fetchedAt,attempts}(contentKind恆為'synthesized',因內容由accessibility snapshot合成),失敗時為{status:'error',url,message,reason,method,fetchedAt,attempts},本函數不會reject

Type
Promise

(async) fetchWebByCurl(url, optopt) → {Promise}

Description:
  • 使用系統curl抓取網頁原始HTML

    特點: 純curl抓取直接回傳原始HTML字串不解析; HTTP 5xx與429及curl錯誤會自動重試(線性退避3至15秒),HTTP 4xx(429除外)則不重試; 網址由execFile以參數陣列傳遞,無命令注入風險,且採非同步執行不阻塞node event loop

Source:
Example
import fetchWebByCurl from './src/fetchWebByCurl.mjs'

let test = async () => {

    let r = await fetchWebByCurl('https://example.com/')
    console.log(r.status, r.httpCode, r.htmlLength)
    // => 'success' 200 559

    let re = await fetchWebByCurl('abc')
    console.log(re.status, re.reason)
    // => 'error' 'invalid-url'

}
await test()
    .catch((err) => {
        console.log(err)
    })
Parameters:
Name Type Attributes Default Description
url String

輸入待抓取網址字串

opt Object <optional>
{}

輸入設定物件,預設{}

Properties
Name Type Attributes Default Description
timeoutMs Integer <optional>
15000

輸入curl最長等待毫秒整數,預設15000

maxRetries Integer <optional>
5

輸入失敗時最大重試次數整數,含初始共執行maxRetries+1次,預設5

userAgent String <optional>

輸入自訂User-Agent字串,未給時採requestIdentity.mjs之DEFAULT_UA(偽裝為Chrome,否則curl會送出自己的UA而被擋)

referer String <optional>

輸入自訂Referer字串,未給時採requestIdentity.mjs之DEFAULT_REFERER

acceptLanguage String <optional>

輸入自訂Accept-Language字串,未給時採requestIdentity.mjs之DEFAULT_ACCEPT_LANG

Returns:

回傳Promise,resolve回傳結果物件,成功時為{status:'success',url,html,htmlLength,httpCode,contentKind,method,fetchedAt,attempts}(contentKind恆為'raw'),失敗時為{status:'error',url,message,reason,httpCode,method,fetchedAt,attempts},本函數不會reject

Type
Promise

(async) fetchWebByPlaywrightHead(url, optopt) → {Promise}

Description:
  • 使用Playwright有頭Chrome抓取網頁原始HTML,含驗證checkbox自動點擊

    特點: 有頭模式(實體視窗)並加反自動化偽裝(隱藏webdriver、disable-blink-features); 自動偵測並點擊Cloudflare Turnstile與hCaptcha等驗證checkbox(模擬人類滑鼠軌跡); 可見文字過少時自動穿透Shadow DOM取得內文並重組為簡易HTML; 失敗時自動重試(線性退避3至15秒); 使用playwright之chromium並指定channel='chrome',故執行環境須已安裝Chrome瀏覽器,且執行時會開啟實體瀏覽器視窗

Source:
Example
import fetchWebByPlaywrightHead from './src/fetchWebByPlaywrightHead.mjs'

let test = async () => {

    let r = await fetchWebByPlaywrightHead('https://example.com/')
    console.log(r.status, r.verificationClicked, r.htmlLength)
    // => 'success' false 234

    let re = await fetchWebByPlaywrightHead('abc')
    console.log(re.status, re.reason)
    // => 'error' 'invalid-url'

}
await test()
    .catch((err) => {
        console.log(err)
    })
Parameters:
Name Type Attributes Default Description
url String

輸入待抓取網址字串

opt Object <optional>
{}

輸入設定物件,預設{}

Properties
Name Type Attributes Default Description
navigationTimeoutMs Integer <optional>
15000

輸入頁面導航最長等待毫秒整數,預設15000

postNavigationWaitMs Integer <optional>
5000

輸入導航後額外等待毫秒整數,預設5000

waitForRedirect Boolean <optional>
false

輸入是否等待JS轉址完成布林值,預設false

skipVerificationClick Boolean <optional>
false

輸入是否跳過驗證checkbox自動點擊布林值,預設false

maxRetries Integer <optional>
5

輸入失敗時最大重試次數整數,含初始共執行maxRetries+1次,預設5

Returns:

回傳Promise,resolve回傳結果物件,成功時為{status:'success',url,html,htmlLength,contentKind,verificationClicked,method,fetchedAt,attempts}(contentKind為'raw'或'synthesized',後者代表內容由Shadow DOM穿透後合成),失敗時為{status:'error',url,message,reason,method,fetchedAt,attempts},本函數不會reject

Type
Promise

(async) fetchWebByPlaywrightHeadless(url, optopt) → {Promise}

Description:
  • 使用Playwright無頭Chrome抓取網頁原始HTML

    特點: 無頭模式適用SPA等須執行JS渲染之頁面; 可見文字過少時自動穿透Shadow DOM取得內文並重組為簡易HTML; 失敗時自動重試(線性退避3至15秒); 使用playwright之chromium並指定channel='chrome',故執行環境須已安裝Chrome瀏覽器

Source:
Example
import fetchWebByPlaywrightHeadless from './src/fetchWebByPlaywrightHeadless.mjs'

let test = async () => {

    let r = await fetchWebByPlaywrightHeadless('https://example.com/')
    console.log(r.status, r.htmlLength)
    // => 'success' 234

    let re = await fetchWebByPlaywrightHeadless('abc')
    console.log(re.status, re.reason)
    // => 'error' 'invalid-url'

}
await test()
    .catch((err) => {
        console.log(err)
    })
Parameters:
Name Type Attributes Default Description
url String

輸入待抓取網址字串

opt Object <optional>
{}

輸入設定物件,預設{}

Properties
Name Type Attributes Default Description
navigationTimeoutMs Integer <optional>
15000

輸入頁面導航最長等待毫秒整數,預設15000

postNavigationWaitMs Integer <optional>
3000

輸入導航後額外等待毫秒整數,預設3000

waitForRedirect Boolean <optional>
false

輸入是否等待JS轉址完成布林值,預設false

maxRetries Integer <optional>
5

輸入失敗時最大重試次數整數,含初始共執行maxRetries+1次,預設5

Returns:

回傳Promise,resolve回傳結果物件,成功時為{status:'success',url,html,htmlLength,contentKind,method,fetchedAt,attempts}(contentKind為'raw'或'synthesized',後者代表內容由Shadow DOM穿透後合成),失敗時為{status:'error',url,message,reason,method,fetchedAt,attempts},本函數不會reject

Type
Promise

fetchedAtIso() → {String}

Description:
  • 取得ISO 8601 UTC格式時間戳

    供四個fetchWebByXxx使用,於進入函數時取值,代表「開始嘗試」的時間

Source:
Example
import { fetchedAtIso } from './src/fetchedAt.mjs'

console.log(fetchedAtIso())
// => '2026-09-09T08:00:00.000Z'
Returns:

回傳ISO 8601 UTC格式時間字串

Type
String

fetchedAtLocal() → {String}

Description:
  • 取得本地時間格式時間戳

    供fetchWeb之finalize使用,於彙整結果時取值,代表「完成」的時間

Source:
Example
import { fetchedAtLocal } from './src/fetchedAt.mjs'

console.log(fetchedAtLocal())
// => '2026-09-09 16:00:00'
Returns:

回傳'YYYY-MM-DD HH:mm:ss'格式之本地時間字串

Type
String

fetcherOf(opt, fetcherKey, real) → {function}

Description:
  • 依測試接縫取出應使用之抓取函數

Source:
Example
import fetcherOf from './src/fetcherSeam.mjs'

let real = async () => 'real'
let fake = async () => 'fake'
console.log(fetcherOf({ _fetchers: { curl: fake } }, 'curl', real) === fake, fetcherOf({}, 'curl', real) === real)
// => true true
Parameters:
Name Type Description
opt Object

輸入設定物件

fetcherKey String

輸入接縫鍵名字串,可為'curl'、'playwrightHeadless'、'playwrightHead'、'camofox'

real function

輸入真實抓取函數

Returns:

回傳應使用之抓取函數:opt._fetchers[fetcherKey]為函數時用它,否則用real

Type
function

finalize(url, result, attempts) → {Object}

Description:
  • 彙整fetchWeb之最終回傳結果

    成功且已解析時輸出title與content,成功但未解析時輸出html; 失敗時輸出message,並於有失敗歸因時附上reason;內容經adapter之fetch掛點取得、或於該階收攤時,另帶adapterId

Source:
Example
import { finalize } from './src/finalizeResult.mjs'

console.log(finalize('https://a.com/', { success: false, reason: 'empty-content', message: 'too short' }, []))
// => { status: 'error', url: 'https://a.com/', message: 'too short', fetchedAt: '2026-09-09 16:00:00', attempts: [], reason: 'empty-content' }
Parameters:
Name Type Description
url String

輸入網址字串

result Object

輸入內部結構之結果物件

attempts Array

輸入各階嘗試紀錄陣列

Returns:

回傳對外之結果物件

Type
Object

(async) findAdapter(url, adapters) → {Promise}

Description:
  • 由adapter清單中找出第一個命中網址者

    adapter之形狀與合法性判準以src/adapterContract.mjs為唯一事實來源,本函數只負責挑選。

    依序試各adapter,第一個命中者勝出;條目不合法(非物件、缺id、match型別不符、缺parse)一律略過。 match執行拋錯時回傳type='error'而非視為未命中——呼叫端註冊了adapter即代表選定該解析階段, 若靜默改用預設解析器,等於讓呼叫端在不知情下經歷未選擇的處理管線階段,故一律顯性回報。 match之回傳值一律await,故async match可正常運作,其reject亦被同一錯誤邊界攔下

Source:
Example
import findAdapter from './src/findAdapter.mjs'

let adapters = [
    { id: 'a', match: /^https?:\/\/a\.com\//, parse: () => ({ success: true }) },
    { id: 'b', match: (u) => { let m = u.match(/\/ar-(\w+)/); return m ? { id: m[1] } : null }, parse: () => ({ success: true }) },
]

console.log(findAdapter('https://a.com/x', adapters).adapter.id)
// => 'a'

console.log(findAdapter('https://b.com/ar-AA1X', adapters).ctx)
// => { id: 'AA1X' }

console.log(findAdapter('https://c.com/', adapters))
// => { type: 'miss' }
Parameters:
Name Type Description
url String

輸入待比對網址字串

adapters Array

輸入adapter物件陣列,非陣列時視為空陣列

Returns:

回傳Promise,resolve回傳結果物件,命中時為{type:'hit',adapter,ctx},其中ctx為match函數之回傳值(RegExp或回傳true時為null);未命中時為{type:'miss'};match執行拋錯時為{type:'error',id,message},本函數不會reject

Type
Promise

getBrowserPageOptions(opt) → {Object}

Description:
  • 由請求身分組出Playwright之newPage選項

    只帶入使用者明確指定者,全未指定時回傳空物件使瀏覽器沿用自身身分

Source:
Example
import { getBrowserPageOptions } from './src/requestIdentity.mjs'

console.log(getBrowserPageOptions({ userAgent: 'X/1.0' }))
// => { userAgent: 'X/1.0' }

console.log(getBrowserPageOptions({}))
// => {}
Parameters:
Name Type Description
opt Object

輸入設定物件

Returns:

回傳Playwright之newPage選項物件

Type
Object

getCurlIdentity(opt) → {Object}

Description:
  • 取得curl階之HTTP請求身分,未指定者採預設值

Source:
Example
import { getCurlIdentity } from './src/requestIdentity.mjs'

console.log(getCurlIdentity({}).referer)
// => 'https://www.google.com/'
Parameters:
Name Type Description
opt Object

輸入設定物件

Returns:

回傳{userAgent,referer,acceptLanguage}物件

Type
Object

getOptArr(opt, key, def) → {Array}

Description:
  • 由設定物件取陣列選項,型別不符時採預設值

Source:
Example
import { getOptArr } from './src/getOpt.mjs'

console.log(getOptArr({ adapters: [1] }, 'adapters', []), getOptArr({ adapters: 'x' }, 'adapters', []))
// => [1] []
Parameters:
Name Type Description
opt Object

輸入設定物件

key String

輸入鍵名字串

def Array

輸入預設值陣列

Returns:

回傳選項陣列

Type
Array

getOptBool(opt, key, def) → {Boolean}

Description:
  • 由設定物件取布林選項,型別不符時採預設值

Source:
Example
import { getOptBool } from './src/getOpt.mjs'

console.log(getOptBool({ parse: false }, 'parse', true), getOptBool({ parse: 'x' }, 'parse', true))
// => false true
Parameters:
Name Type Description
opt Object

輸入設定物件

key String

輸入鍵名字串

def Boolean

輸入預設值布林值

Returns:

回傳選項布林值

Type
Boolean

getOptP0Int(opt, key, def) → {Integer}

Description:
  • 由設定物件取非負整數選項,型別不符時採預設值

Source:
Example
import { getOptP0Int } from './src/getOpt.mjs'

console.log(getOptP0Int({ maxRetries: 0 }, 'maxRetries', 5), getOptP0Int({ maxRetries: -1 }, 'maxRetries', 5))
// => 0 5
Parameters:
Name Type Description
opt Object

輸入設定物件

key String

輸入鍵名字串

def Integer

輸入預設值整數

Returns:

回傳選項整數

Type
Integer

getOptPInt(opt, key, def) → {Integer}

Description:
  • 由設定物件取正整數選項,型別不符時採預設值

Source:
Example
import { getOptPInt } from './src/getOpt.mjs'

console.log(getOptPInt({ timeoutMs: 3000 }, 'timeoutMs', 15000), getOptPInt({ timeoutMs: 0 }, 'timeoutMs', 15000))
// => 3000 15000
Parameters:
Name Type Description
opt Object

輸入設定物件

key String

輸入鍵名字串

def Integer

輸入預設值整數

Returns:

回傳選項整數

Type
Integer

getOptStr(opt, key, def) → {String}

Description:
  • 由設定物件取非空字串選項,型別不符時採預設值

Source:
Example
import { getOptStr } from './src/getOpt.mjs'

console.log(getOptStr({ method: 'curl' }, 'method', 'auto'), getOptStr({ method: '' }, 'method', 'auto'))
// => 'curl' 'auto'
Parameters:
Name Type Description
opt Object

輸入設定物件

key String

輸入鍵名字串

def String

輸入預設值字串

Returns:

回傳選項字串

Type
String

getRequestIdentity(opt) → {Object}

Description:
  • 取得使用者明確指定之HTTP請求身分,未指定者為空字串

    供瀏覽器階使用:只有非空者才覆寫瀏覽器自身之身分

Source:
Example
import { getRequestIdentity } from './src/requestIdentity.mjs'

console.log(getRequestIdentity({ userAgent: 'X/1.0' }))
// => { userAgent: 'X/1.0', referer: '', acceptLanguage: '' }
Parameters:
Name Type Description
opt Object

輸入設定物件

Returns:

回傳{userAgent,referer,acceptLanguage}物件,未指定之項目為空字串

Type
Object

getRetryWaitMs(attempt) → {Integer}

Description:
  • 取得重試前之線性退避等待毫秒

    第n次失敗後等待n*3000毫秒,上限15000毫秒,即3000、6000、9000、12000、15000、15000...

Source:
Example
import getRetryWaitMs from './src/getRetryWaitMs.mjs'

console.log(getRetryWaitMs(1), getRetryWaitMs(2), getRetryWaitMs(5), getRetryWaitMs(99))
// => 3000 6000 15000 15000
Parameters:
Name Type Description
attempt Integer

輸入第幾次嘗試之正整數,由1起算

Returns:

回傳等待毫秒整數

Type
Integer

getUrlErrorResult(url, method, fetchedAt) → {Object|null}

Description:
  • 檢核網址並取得錯誤結果物件

    供各抓取函數於入口統一檢核網址,網址有效時回傳null代表可繼續執行

Source:
Example
import getUrlErrorResult from './src/getUrlErrorResult.mjs'

console.log(getUrlErrorResult('https://a.com/', 'curl', 'now'))
// => null

console.log(getUrlErrorResult('abc', 'curl', 'now'))
// => { status: 'error', url: 'abc', message: 'invalid url (must be http/https)', reason: 'invalid-url', method: 'curl', fetchedAt: 'now', attempts: 0 }
Parameters:
Name Type Description
url String

輸入待檢核網址字串

method String

輸入抓取方法名稱字串,將寫入結果物件之method欄位

fetchedAt String

輸入抓取時間字串,將寫入結果物件之fetchedAt欄位

Returns:

網址無效時回傳錯誤結果物件{status:'error',url,message,reason:'invalid-url',method,fetchedAt,attempts:0},網址有效時回傳null

Type
Object | null

inspectHtml(html, optopt) → {Object}

Description:
  • 由網頁原始HTML檢測頁面是否為有效內容

    純粹基於原始HTML結構判斷,不依賴Readability,用於判識CAPTCHA與反爬蟲挑戰頁、驗證頁、 轉址包裝頁、空內容頁,供階梯升級決策使用。 判識器以資料表定義並依序比對,首個命中者勝出;順序本身即語意,調動順序會改變判定結果。

    contentKind標明待測內容是抓取器取回的原始文件,或由已渲染DOM萃取後合成者。 合成內容之標籤結構已被剝除,故只比對semantic類判識器;預設為'raw', 亦即未指定時行為與原先一致

Source:
Example
import inspectHtml from './src/inspectHtml.mjs'

console.log(inspectHtml('<html><head><title>Just a moment</title></head><body></body></html>'))
// => { pass: false, type: 'captcha', message: 'Cloudflare/anti-bot challenge' }

console.log(inspectHtml('<html><head><title>abc</title></head><body><p>' + 'x'.repeat(300) + '</p></body></html>'))
// => { pass: true, type: 'pass', message: 'ok' }
Parameters:
Name Type Attributes Default Description
html String

輸入網頁HTML字串

opt Object <optional>
{}

輸入設定物件,預設{}

Properties
Name Type Attributes Default Description
contentKind String <optional>
'raw'

輸入內容形態字串,'raw'為原始文件,'synthesized'為合成內容,預設'raw'

detectors Array <optional>
[]

輸入使用端判識器陣列,排於內建判識器之前故優先命中,其契約以src/detectorContract.mjs為唯一事實來源,預設[]

builtin Boolean <optional>
true

輸入是否比對內建判識器布林值,預設true。false時只比對opt.detectors所註冊者,供adapter之inspect:false豁免「套件對該站台之通用猜測」而不動呼叫端自己的判準

Returns:

回傳檢測結果物件,格式為{pass,type,message},其中pass為是否通過布林值,type為'pass'、'captcha'、'verify'、'redirect'、'empty'之一,message為說明字串

Type
Object

isInternalHost(hostname) → {Boolean}

Description:
  • 判別主機名是否指向內網、迴環或保留位址

    僅供把關本套件自行推導出之網址(如由轉址服務query參數提取者), 不可用於把關呼叫端明確給定之網址——後者是呼叫端的決定

Source:
Example
import isInternalHost from './src/isInternalHost.mjs'

console.log(isInternalHost('169.254.169.254'), isInternalHost('127.0.0.1'), isInternalHost('localhost'))
// => true true true

console.log(isInternalHost('example.com'), isInternalHost('8.8.8.8'))
// => false false
Parameters:
Name Type Description
hostname String

輸入主機名字串,不含協定與埠

Returns:

回傳是否為內網或保留位址之布林值

Type
Boolean

isValidAdapter(adapter) → {Boolean}

Description:
  • 檢核adapter條目是否符合輸入契約

Source:
Example
import { isValidAdapter } from './src/adapterContract.mjs'

console.log(isValidAdapter({ id: 'a', match: /x/, parse: () => ({}) }), isValidAdapter({ id: 'a', match: 'x', parse: () => ({}) }))
// => true false
Parameters:
Name Type Description
adapter *

輸入待檢核之adapter

Returns:

回傳是否合法之布林值

Type
Boolean

isValidDetector(detector) → {Boolean}

Description:
  • 檢核使用端判識器是否符合契約

    不合法者一律略過而非拋錯,與adapter之處置一致

Source:
Example
import { isValidDetector } from './src/detectorContract.mjs'

console.log(isValidDetector({ type: 'captcha', message: 'x', test: () => true }))
// => true

console.log(isValidDetector({ type: 'unknown', message: 'x', test: () => true }))
// => false
Parameters:
Name Type Description
detector *

輸入待檢核之判識器

Returns:

回傳是否合法之布林值

Type
Boolean

isValidUrl(url) → {Boolean}

Description:
  • 檢核是否為有效的http或https網址字串

Source:
Example
import isValidUrl from './src/isValidUrl.mjs'

console.log(isValidUrl('https://www.google.com/'))
// => true

console.log(isValidUrl('http://127.0.0.1:8080/abc'))
// => true

console.log(isValidUrl('ftp://www.google.com/'))
// => false

console.log(isValidUrl('www.google.com'))
// => false
Parameters:
Name Type Description
url String

輸入待檢核網址字串

Returns:

回傳是否為有效http或https網址之布林值

Type
Boolean

matchMsn(url) → {Object|null}

Description:
  • 比對 msn 內容頁網址並取出內容型別與 id

    作為內建 msn adapter 之 match 掛點。回傳之物件即為 ctx,傳入 fetch 掛點

Source:
Example
import { matchMsn } from './src/fetchMsn.mjs'

console.log(matchMsn('https://www.msn.com/zh-tw/news/other/abc/ar-AA2bZm9d'))
// => { kind: 'ar', id: 'AA2bZm9d' }

console.log(matchMsn('https://www.msn.com/en-us/video/news/abc/vi-AA2bYtCB'))
// => { kind: 'vi', id: 'AA2bYtCB' }

console.log(matchMsn('https://www.msn.com/zh-tw/news'))
// => null
Parameters:
Name Type Description
url String

輸入網址字串

Returns:

命中時回傳{kind,id},kind為'ar'(文章)或'vi'(影片);未命中時回傳null

Type
Object | null

meetsMinContent(content) → {Boolean}

Description:
  • 判定正文長度是否達最低門檻

    adapter路徑與Readability路徑刻意採同一標準,故兩處皆呼叫本函數而非各自比較, 避免同一規則手寫兩處而日後分歧

Source:
Example
import { meetsMinContent } from './src/adapterContract.mjs'

console.log(meetsMinContent('x'.repeat(50)), meetsMinContent('x'.repeat(49)))
// => true false
Parameters:
Name Type Description
content String

輸入正文字串

Returns:

回傳是否達門檻之布林值

Type
Boolean
Description:
  • 導航至網址並等待JS轉址完成

    先以domcontentloaded導航,再等待網址host脫離原host(代表已轉址),最後等待networkidle; 等待轉址與networkidle皆為盡力而為,逾時不視為失敗

Source:
Example
import { chromium } from 'playwright'
import navigateWithRedirectWait from './src/navigateWithRedirectWait.mjs'

let test = async () => {

    let browser = await chromium.launch({ headless: true, channel: 'chrome' })
    let page = await browser.newPage()
    await navigateWithRedirectWait(page, 'https://news.google.com/articles/xxxxxx', 15000)
    console.log(page.url())
    // => 'https://www.example-news.com/real-article'
    await browser.close()

}
await test()
    .catch((err) => {
        console.log(err)
    })
Parameters:
Name Type Description
page Object

輸入Playwright之page物件

url String

輸入待導航網址字串

navTimeout Integer

輸入導航最長等待毫秒整數

Returns:

回傳Promise,resolve回傳首次導航之Response物件(供呼叫端檢核HTTP狀態),同頁錨點導航等情形可能為null

Type
Promise

normalizeDetector(detector) → {Object}

Description:
  • 把使用端判識器正規化為內部形狀,補上預設之evidence並標記來源

    origin標記為'custom',供判識流程據以略過內容量閘門,理由見本檔檔頭

Source:
Example
import { normalizeDetector } from './src/detectorContract.mjs'

console.log(normalizeDetector({ type: 'captcha', message: 'x', test: () => true }).evidence)
// => 'semantic'
Parameters:
Name Type Description
detector Object

輸入已通過檢核之判識器

Returns:

回傳正規化後之判識器物件

Type
Object

normalizeParsed(parsed, pre) → {Object}

Description:
  • 把adapter之parse回傳值檢核並投影為固定形狀

Source:
Example
import { normalizeParsed } from './src/adapterContract.mjs'

console.log(normalizeParsed({ success: 'true', content: 'x' }, 'adapter a '))
// => { success: false, reason: 'parse-error', message: 'adapter a returned non-boolean success' }
Parameters:
Name Type Description
parsed *

輸入adapter.parse之回傳值

pre String

輸入錯誤訊息前綴字串

Returns:

回傳內部結構結果物件,成功時為{success:true,title,content,contentLength},失敗時為{success:false,reason,message}

Type
Object

(async) parseArticle(html, url, hit, metaopt) → {Promise}

Description:
  • 依已解析之adapter命中結果解析文章,未命中者走Readability

    adapter之挑選(findAdapter)刻意不在本函數內:它只取決於網址,與抓回之HTML無關, 故由runPlan於計畫執行前解析一次後傳入,詳見runPlan之_resolveAdapter

Source:
Parameters:
Name Type Attributes Default Description
html String

輸入網頁HTML字串

url String

輸入網址字串

hit Object | null

輸入findAdapter之結果物件,null視為未命中

meta Object | null <optional>
null

輸入本次抓取之附帶資訊物件,含finalUrl等,供adapter判斷與JSDOM之base url使用,預設null

Returns:

回傳Promise,resolve回傳解析結果物件,本函數不會reject

Type
Promise

parseBloomberg(html, url) → {Object}

Description:
  • 解析Bloomberg之文章內文

    由