Members
WFetchWeb
- Description:
網頁抓取
- Source:
網頁抓取
Example
詳見fetchWeb、fetchWebByCurl、fetchWebByPlaywrightHeadless、fetchWebByPlaywrightHead、fetchWebByCamofox範例
Methods
adapt(r) → {Object}
- Description:
把抓取函數之對外結果轉為內部流程用之結構
四個fetchWebByXxx回傳{status:'success'|'error',...},而執行計畫內部以{success:Boolean,...}判斷, 由本函數轉換以隔離兩者
- Source:
Example
import { adapt } from './src/finalizeResult.mjs'
console.log(adapt({ status: 'success', html: '<p>a</p>', method: 'curl' }))
// => { success: true, html: '<p>a</p>', method: 'curl', snapshot: undefined, contentKind: undefined }
Parameters:
| Name | Type | Description |
|---|---|---|
r |
Object | 輸入抓取函數之結果物件 |
Returns:
回傳內部結構物件,成功時為{success:true,html,method,snapshot,contentKind},失敗時為{success:false,method,reason,message}
- Type
- Object
buildPlan(url, method, doInspect, hasFetchAdapteropt) → {Object}
- Description:
依網址與選項產生執行計畫
四條網域分流與指定方法,差別僅在於「要跑哪幾階」,故一律表達為step陣列, 交由同一個runPlan執行,不再各自複製一份「抓取至判識至解析」流程。
本函數只回傳純資料:不捕捉url與opt、不建立closure、不修改任何共享狀態。 轉址旗標之初值以redirect欄位回報,由runPlan擁有其後續變化
- Source:
Example
import buildPlan from './src/buildPlan.mjs'
console.log(buildPlan('https://example.com/', 'auto', true).plan.map((s) => s.key))
// => ['curl', 'headless', 'headed', 'camofox']
console.log(buildPlan('https://mp.weixin.qq.com/s/a', 'auto', true).plan.map((s) => s.key))
// => ['camofox']
Parameters:
| Name | Type | Attributes | Default | Description |
|---|---|---|---|---|
url |
String | 輸入待抓取網址字串 |
||
method |
String | 輸入方法字串,'auto'代表自動階梯升級 |
||
doInspect |
Boolean | 輸入是否做原始內容判識布林值,false時各階之inspect一律關閉 |
||
hasFetchAdapter |
Boolean |
<optional> |
false
|
輸入命中之adapter是否具fetch掛點布林值,true時於計畫最前插入adapter階,預設false |
Returns:
回傳{plan,redirect,log}物件,plan為step描述陣列,redirect為轉址旗標初值,log為分流說明字串;method不合法時回傳{error}
- Type
- Object
checkFetched(r, reason, pre) → {Object|null}
- Description:
檢核抓取器之回傳值是否符合輸出契約
不合契約者轉為結構化失敗結果而非拋錯,使fetchWeb之「不會reject」契約於任何抓取器實作下皆成立
- Source:
Example
import { checkFetched } from './src/fetchContract.mjs'
console.log(checkFetched({ status: 'success', html: '<p>a</p>' }, 'fetcher-error', 'curl '))
// => null
console.log(checkFetched('oops', 'fetcher-error', 'curl '))
// => { status: 'error', reason: 'fetcher-error', message: 'curl returned invalid result' }
Parameters:
| Name | Type | Description |
|---|---|---|
r |
* | 輸入抓取器之回傳值 |
reason |
String | 輸入不合契約時採用之失敗歸因字串 |
pre |
String | 輸入錯誤訊息前綴字串 |
Returns:
合契約時回傳null,不合時回傳{status:'error',reason,message}
- Type
- Object | null
checkHttpStatus(httpCode) → {Object|null}
- Description:
依HTTP狀態碼判定是否為失敗,並給出失敗歸因與可重試性
四個抓取器共用同一判準,不各自實作
- Source:
Example
import { checkHttpStatus } from './src/httpStatus.mjs'
console.log(checkHttpStatus(200))
// => null
console.log(checkHttpStatus(404))
// => { reason: 'http-error', message: 'HTTP 404', httpCode: 404, retryable: false }
console.log(checkHttpStatus(503).retryable)
// => true
Parameters:
| Name | Type | Description |
|---|---|---|
httpCode |
Integer | 輸入HTTP狀態碼整數,取不到時傳0 |
Returns:
非失敗時回傳null;失敗時回傳{reason,message,httpCode,retryable}
- Type
- Object | null
estimateVisibleText(html) → {String}
- Description:
由網頁HTML估算可見文字內容
取
內容後移除script、style、標籤與HTML實體,再壓縮空白,用以粗估頁面實質可見字數, 供轉址殼頁、空內容頁之判識使用
- Source:
Example
import estimateVisibleText from './src/estimateVisibleText.mjs'
console.log(estimateVisibleText('<html><body><p>abc</p> <p>def</p></body></html>'))
// => 'abc def'
console.log(estimateVisibleText('<body><script>var a=1</script><div>xyz</div></body>'))
// => 'xyz'
Parameters:
| Name | Type | Description |
|---|---|---|
html |
String | 輸入網頁HTML字串 |
Returns:
回傳估算之可見文字字串
- Type
- String
extractHtmlTitle(html) → {String}
- Description:
由網頁HTML取出標題,並剝除尾端之「-站名」
- Source:
Example
import extractHtmlTitle from './src/extractHtmlTitle.mjs'
console.log(extractHtmlTitle('<html><head><title>流動性溢價的真相 - 格隆匯</title></head></html>'))
// => '流動性溢價的真相'
Parameters:
| Name | Type | Description |
|---|---|---|
html |
String | 輸入網頁HTML字串 |
Returns:
回傳標題字串,取不到時回傳空字串
- Type
- String
(async) extractPageContent(page) → {Promise}
- Description:
由Playwright之page提取網頁HTML,可見文字過少時穿透Shadow DOM
先取page.content(),若估算可見文字已達門檻則直接回傳;否則遞迴穿透Shadow DOM取得 各層innerText,再依換行切段重組為含article之簡易HTML,供Readability解析。
回傳一併標明內容形態:合成內容之標籤結構已被剝除,其可信的判識證據與原始文件不同, 故由本函數(唯一知道走了哪個出口者)標明,而非由下游猜測
- Source:
Example
import { chromium } from 'playwright'
import extractPageContent from './src/extractPageContent.mjs'
let test = async () => {
let browser = await chromium.launch({ headless: true, channel: 'chrome' })
let page = await browser.newPage()
await page.goto('https://example.com/')
let { html, contentKind } = await extractPageContent(page)
console.log(html.length, contentKind)
// => 234 'raw'
await browser.close()
}
await test()
.catch((err) => {
console.log(err)
})
Parameters:
| Name | Type | Description |
|---|---|---|
page |
Object | 輸入Playwright之page物件 |
Returns:
回傳Promise,resolve回傳{html,contentKind}物件,contentKind為'raw'或'synthesized'
- Type
- Promise
extractRedirectTarget(url) → {String|null}
- Description:
由轉址服務之網址中提取其query參數所帶之真實網址
僅處理已知會把目標網址放在query參數之服務;取出後之值已由URLSearchParams解碼一次, 再解一次以處理雙重編碼,若該值非合法百分比序列則退回僅解碼一次之結果。
提取出之目標須通過內網位址檢核,指向迴環、私有網段、link-local或非http/https者一律不提取
- Source:
Example
import { extractRedirectTarget } from './src/routeByUrl.mjs'
console.log(extractRedirectTarget('https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Fa.com%2Fb'))
// => 'https://a.com/b'
console.log(extractRedirectTarget('https://www.linkedin.com/redir/redirect?url=http%3A%2F%2F169.254.169.254%2F'))
// => null
Parameters:
| Name | Type | Description |
|---|---|---|
url |
String | 輸入網址字串 |
Returns:
命中且目標可用時回傳目標網址字串,未命中或目標指向內網時回傳null
- Type
- String | null
(async) fetchMsn(url, opt, ctx) → {Promise}
- Description:
經 msn 內容 API 取得內容,重組為 HTML 文件
作為內建 msn adapter 之 fetch 掛點。API 之請求走本套件之 curl 抓取器, 故 User-Agent、重試與 HTTP 狀態判準與其餘抓取一致;opt 原樣轉傳,並沿用 opt._fetchers.curl 測試接縫
- Source:
Example
import { matchMsn, fetchMsn } from './src/fetchMsn.mjs'
let test = async () => {
let url = 'https://www.msn.com/zh-tw/news/other/abc/ar-AA2bZm9d'
let r = await fetchMsn(url, {}, matchMsn(url))
console.log(r.status, r.contentKind)
// => 'success' 'synthesized'
}
await test()
.catch((err) => {
console.log(err)
})
Parameters:
| Name | Type | Description |
|---|---|---|
url |
String | 輸入內容頁網址字串,本函數僅用於訊息 |
opt |
Object | 輸入設定物件,轉傳給 curl 抓取器 |
ctx |
Object | 輸入 matchMsn 之回傳物件{kind,id} |
Returns:
回傳Promise,resolve回傳結果物件,成功時為{status:'success',html,contentKind:'synthesized'},失敗時為{status:'error',reason,message},本函數不會reject
- Type
- Promise
(async) fetchWeb(url, optopt) → {Promise}
- Description:
抓取網頁文章內容,支援四種抓取方法自動階梯升級
抓取方法: 方法①curl(預設,繞過TLS指紋),委派fetchWebByCurl; 方法②Playwright無頭(SPA動態渲染頁面),委派fetchWebByPlaywrightHeadless; 方法③Playwright有頭(反自動化偵測),委派fetchWebByPlaywrightHead; 方法④Camofox反偵測瀏覽器(Cloudflare等),委派fetchWebByCamofox
本函數僅負責階梯升級、內容判識、文章解析之調度,實際抓取由4個抓取函數執行, 流程為fetch(委派)至inspectHtml(原始內容檢測)至Readability解析(可選); 對已知網站另有轉址提取、跳過特定方法之判識規則,且方法④會額外回傳snapshot欄位
並行呼叫之限制:方法④之Camofox server綁定固定埠(預設19377),同一埠號同時只能有一個抓取, 詳見fetchWebByCamofox之說明。auto模式可能升級至該階,故同時發動多個fetchWeb時, 若其中一個以上走到方法④即會互相破壞(先完成者殺掉server,其餘回'camofox-error')。 需要並行時須為每個呼叫指定互不相同的opt.port;方法①②③則無此限制
- Source:
Example
import fetchWeb from './src/fetchWeb.mjs'
let test = async () => {
//auto模式, 解析出文章標題與內文
let r = await fetchWeb('https://example.com/')
console.log(r.status, r.method, r.title, r.contentLength)
// => 'success' 'curl' 'Example Domain' 111
//指定curl且不解析, 直接取原始HTML
let rh = await fetchWeb('https://example.com/', { method: 'curl', parse: false })
console.log(rh.status, rh.html.length)
// => 'success' 559
//失敗時回傳error結果物件, 不會reject
let re = await fetchWeb('abc')
console.log(re.status, re.message)
// => 'error' 'invalid url (must be http/https)'
}
await test()
.catch((err) => {
console.log(err)
})
Parameters:
| Name | Type | Attributes | Default | Description | ||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
url |
String | 輸入待抓取網址字串 |
||||||||||||||||||||||||||||||||||||||||||||||||||||
opt |
Object |
<optional> |
{}
|
輸入設定物件,其餘鍵值會轉傳給實際執行抓取之函數,預設{} Properties
|
Returns:
回傳Promise,resolve回傳結果物件,其中attempts為各階嘗試紀錄陣列,成功之紀錄為{method,status:'success',htmlLength}(htmlLength為原始HTML長度,與頂層contentLength之正文長度不同),失敗為{method,status:'failed',reason,message},被判識或解析失敗為{method,status:'blocked',type,reason,message}(判識所致者reason與type同值);經adapter之fetch掛點者其紀錄之method為'adapter'且另帶adapterId,成功結果與於該階收攤之失敗結果其頂層亦帶adapterId(落回後階梯耗盡者則無);parse=true成功時為{status:'success',url,title,content,contentLength,method,fetchedAt,attempts},parse=false成功時為{status:'success',url,html,method,fetchedAt,attempts};內容來自轉址後之另一網址時另帶finalUrl(url為本套件最後實際發出請求之網址,finalUrl為內容實際來源,兩者相同時不輸出該欄),失敗時為{status:'error',url,message,fetchedAt,attempts},本函數不會reject
- Type
- Promise
(async) fetchWebByCamofox(url, optopt) → {Promise}
- Description:
使用Camofox反偵測瀏覽器抓取網頁原始HTML,透過accessibility snapshot取得內容
流程: 以Node模組解析機制取得已安裝之@askjo/camofox-browser之server.js位置; spawn
node <server.js>啟動Camofox server; POST /tabs 建立tab; GET /tabs/:id/snapshot 取accessibility snapshot(含內部重試); DELETE /tabs/:id 關閉tab; 殺整棵server進程樹(Windows用taskkill /F /T,Unix以負PID對spawn時建立之行程群組送SIGTERM)對server之每次HTTP請求皆有15秒硬上限,逾時即abort,不因server無回應而永久等待; 單次嘗試之server啟動與清理皆由runCamofoxAttempt完成,故重試退避期間不會佔用該埠
該套件之server.js無任何export且於top-level即無條件listen,故只能spawn為子行程執行, 不可直接import;解析不到安裝位置時回傳reason='camofox-not-found'。
失敗歸因分三種:'camofox-not-found'為未安裝、'camofox-empty'為頁面確實無足量內容(重試無益)、 'camofox-error'為server啟動失敗、tab建立失敗或snapshot傳輸失敗(可重試)
同一埠號同時只能有一個抓取在進行,本函數不可並行呼叫。 Camofox server綁定固定埠(預設19377),且就緒探測只確認該埠有服務回應、不驗證是否為自己啟動者。 故並行呼叫時:後啟動者因埠被占用而啟動失敗,卻會探到前者的server而誤認就緒並借用之; 待任一方先完成,其清理程序會殺掉該server,另一方隨即全數失敗(回'camofox-error')。 需要並行時,必須為每個同時進行的呼叫指定互不相同的opt.port。 此限制同樣適用於auto模式的fetchWeb——其階梯末階即為本函數。
- Source:
Example
import fetchWebByCamofox from './src/fetchWebByCamofox.mjs'
let test = async () => {
let r = await fetchWebByCamofox('https://mp.weixin.qq.com/s/xxxxxx')
console.log(r.status, r.snapshotChars, r.htmlLength)
// => 'success' 3215 4102
let re = await fetchWebByCamofox('abc')
console.log(re.status, re.reason)
// => 'error' 'invalid-url'
}
await test()
.catch((err) => {
console.log(err)
})
Parameters:
| Name | Type | Attributes | Default | Description | ||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
url |
String | 輸入待抓取網址字串 |
||||||||||||||||||||||||||||||||
opt |
Object |
<optional> |
{}
|
輸入設定物件,預設{} Properties
|
Returns:
回傳Promise,resolve回傳結果物件,成功時為{status:'success',url,html,htmlLength,contentKind,snapshot,snapshotChars,method,fetchedAt,attempts}(contentKind恆為'synthesized',因內容由accessibility snapshot合成),失敗時為{status:'error',url,message,reason,method,fetchedAt,attempts},本函數不會reject
- Type
- Promise
(async) fetchWebByCurl(url, optopt) → {Promise}
- Description:
使用系統curl抓取網頁原始HTML
特點: 純curl抓取直接回傳原始HTML字串不解析; HTTP 5xx與429及curl錯誤會自動重試(線性退避3至15秒),HTTP 4xx(429除外)則不重試; 網址由execFile以參數陣列傳遞,無命令注入風險,且採非同步執行不阻塞node event loop
- Source:
Example
import fetchWebByCurl from './src/fetchWebByCurl.mjs'
let test = async () => {
let r = await fetchWebByCurl('https://example.com/')
console.log(r.status, r.httpCode, r.htmlLength)
// => 'success' 200 559
let re = await fetchWebByCurl('abc')
console.log(re.status, re.reason)
// => 'error' 'invalid-url'
}
await test()
.catch((err) => {
console.log(err)
})
Parameters:
| Name | Type | Attributes | Default | Description | ||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
url |
String | 輸入待抓取網址字串 |
||||||||||||||||||||||||||||||||
opt |
Object |
<optional> |
{}
|
輸入設定物件,預設{} Properties
|
Returns:
回傳Promise,resolve回傳結果物件,成功時為{status:'success',url,html,htmlLength,httpCode,contentKind,method,fetchedAt,attempts}(contentKind恆為'raw'),失敗時為{status:'error',url,message,reason,httpCode,method,fetchedAt,attempts},本函數不會reject
- Type
- Promise
(async) fetchWebByPlaywrightHead(url, optopt) → {Promise}
- Description:
使用Playwright有頭Chrome抓取網頁原始HTML,含驗證checkbox自動點擊
特點: 有頭模式(實體視窗)並加反自動化偽裝(隱藏webdriver、disable-blink-features); 自動偵測並點擊Cloudflare Turnstile與hCaptcha等驗證checkbox(模擬人類滑鼠軌跡); 可見文字過少時自動穿透Shadow DOM取得內文並重組為簡易HTML; 失敗時自動重試(線性退避3至15秒); 使用playwright之chromium並指定channel='chrome',故執行環境須已安裝Chrome瀏覽器,且執行時會開啟實體瀏覽器視窗
- Source:
Example
import fetchWebByPlaywrightHead from './src/fetchWebByPlaywrightHead.mjs'
let test = async () => {
let r = await fetchWebByPlaywrightHead('https://example.com/')
console.log(r.status, r.verificationClicked, r.htmlLength)
// => 'success' false 234
let re = await fetchWebByPlaywrightHead('abc')
console.log(re.status, re.reason)
// => 'error' 'invalid-url'
}
await test()
.catch((err) => {
console.log(err)
})
Parameters:
| Name | Type | Attributes | Default | Description | ||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
url |
String | 輸入待抓取網址字串 |
||||||||||||||||||||||||||||||||
opt |
Object |
<optional> |
{}
|
輸入設定物件,預設{} Properties
|
Returns:
回傳Promise,resolve回傳結果物件,成功時為{status:'success',url,html,htmlLength,contentKind,verificationClicked,method,fetchedAt,attempts}(contentKind為'raw'或'synthesized',後者代表內容由Shadow DOM穿透後合成),失敗時為{status:'error',url,message,reason,method,fetchedAt,attempts},本函數不會reject
- Type
- Promise
(async) fetchWebByPlaywrightHeadless(url, optopt) → {Promise}
- Description:
使用Playwright無頭Chrome抓取網頁原始HTML
特點: 無頭模式適用SPA等須執行JS渲染之頁面; 可見文字過少時自動穿透Shadow DOM取得內文並重組為簡易HTML; 失敗時自動重試(線性退避3至15秒); 使用playwright之chromium並指定channel='chrome',故執行環境須已安裝Chrome瀏覽器
- Source:
Example
import fetchWebByPlaywrightHeadless from './src/fetchWebByPlaywrightHeadless.mjs'
let test = async () => {
let r = await fetchWebByPlaywrightHeadless('https://example.com/')
console.log(r.status, r.htmlLength)
// => 'success' 234
let re = await fetchWebByPlaywrightHeadless('abc')
console.log(re.status, re.reason)
// => 'error' 'invalid-url'
}
await test()
.catch((err) => {
console.log(err)
})
Parameters:
| Name | Type | Attributes | Default | Description | |||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
url |
String | 輸入待抓取網址字串 |
|||||||||||||||||||||||||||
opt |
Object |
<optional> |
{}
|
輸入設定物件,預設{} Properties
|
Returns:
回傳Promise,resolve回傳結果物件,成功時為{status:'success',url,html,htmlLength,contentKind,method,fetchedAt,attempts}(contentKind為'raw'或'synthesized',後者代表內容由Shadow DOM穿透後合成),失敗時為{status:'error',url,message,reason,method,fetchedAt,attempts},本函數不會reject
- Type
- Promise
fetchedAtIso() → {String}
- Description:
取得ISO 8601 UTC格式時間戳
供四個fetchWebByXxx使用,於進入函數時取值,代表「開始嘗試」的時間
- Source:
Example
import { fetchedAtIso } from './src/fetchedAt.mjs'
console.log(fetchedAtIso())
// => '2026-09-09T08:00:00.000Z'
Returns:
回傳ISO 8601 UTC格式時間字串
- Type
- String
fetchedAtLocal() → {String}
- Description:
取得本地時間格式時間戳
供fetchWeb之finalize使用,於彙整結果時取值,代表「完成」的時間
- Source:
Example
import { fetchedAtLocal } from './src/fetchedAt.mjs'
console.log(fetchedAtLocal())
// => '2026-09-09 16:00:00'
Returns:
回傳'YYYY-MM-DD HH:mm:ss'格式之本地時間字串
- Type
- String
fetcherOf(opt, fetcherKey, real) → {function}
- Description:
依測試接縫取出應使用之抓取函數
- Source:
Example
import fetcherOf from './src/fetcherSeam.mjs'
let real = async () => 'real'
let fake = async () => 'fake'
console.log(fetcherOf({ _fetchers: { curl: fake } }, 'curl', real) === fake, fetcherOf({}, 'curl', real) === real)
// => true true
Parameters:
| Name | Type | Description |
|---|---|---|
opt |
Object | 輸入設定物件 |
fetcherKey |
String | 輸入接縫鍵名字串,可為'curl'、'playwrightHeadless'、'playwrightHead'、'camofox' |
real |
function | 輸入真實抓取函數 |
Returns:
回傳應使用之抓取函數:opt._fetchers[fetcherKey]為函數時用它,否則用real
- Type
- function
finalize(url, result, attempts) → {Object}
- Description:
彙整fetchWeb之最終回傳結果
成功且已解析時輸出title與content,成功但未解析時輸出html; 失敗時輸出message,並於有失敗歸因時附上reason;內容經adapter之fetch掛點取得、或於該階收攤時,另帶adapterId
- Source:
Example
import { finalize } from './src/finalizeResult.mjs'
console.log(finalize('https://a.com/', { success: false, reason: 'empty-content', message: 'too short' }, []))
// => { status: 'error', url: 'https://a.com/', message: 'too short', fetchedAt: '2026-09-09 16:00:00', attempts: [], reason: 'empty-content' }
Parameters:
| Name | Type | Description |
|---|---|---|
url |
String | 輸入網址字串 |
result |
Object | 輸入內部結構之結果物件 |
attempts |
Array | 輸入各階嘗試紀錄陣列 |
Returns:
回傳對外之結果物件
- Type
- Object
(async) findAdapter(url, adapters) → {Promise}
- Description:
由adapter清單中找出第一個命中網址者
adapter之形狀與合法性判準以src/adapterContract.mjs為唯一事實來源,本函數只負責挑選。
依序試各adapter,第一個命中者勝出;條目不合法(非物件、缺id、match型別不符、缺parse)一律略過。 match執行拋錯時回傳type='error'而非視為未命中——呼叫端註冊了adapter即代表選定該解析階段, 若靜默改用預設解析器,等於讓呼叫端在不知情下經歷未選擇的處理管線階段,故一律顯性回報。 match之回傳值一律await,故async match可正常運作,其reject亦被同一錯誤邊界攔下
- Source:
Example
import findAdapter from './src/findAdapter.mjs'
let adapters = [
{ id: 'a', match: /^https?:\/\/a\.com\//, parse: () => ({ success: true }) },
{ id: 'b', match: (u) => { let m = u.match(/\/ar-(\w+)/); return m ? { id: m[1] } : null }, parse: () => ({ success: true }) },
]
console.log(findAdapter('https://a.com/x', adapters).adapter.id)
// => 'a'
console.log(findAdapter('https://b.com/ar-AA1X', adapters).ctx)
// => { id: 'AA1X' }
console.log(findAdapter('https://c.com/', adapters))
// => { type: 'miss' }
Parameters:
| Name | Type | Description |
|---|---|---|
url |
String | 輸入待比對網址字串 |
adapters |
Array | 輸入adapter物件陣列,非陣列時視為空陣列 |
Returns:
回傳Promise,resolve回傳結果物件,命中時為{type:'hit',adapter,ctx},其中ctx為match函數之回傳值(RegExp或回傳true時為null);未命中時為{type:'miss'};match執行拋錯時為{type:'error',id,message},本函數不會reject
- Type
- Promise
getBrowserPageOptions(opt) → {Object}
- Description:
由請求身分組出Playwright之newPage選項
只帶入使用者明確指定者,全未指定時回傳空物件使瀏覽器沿用自身身分
- Source:
Example
import { getBrowserPageOptions } from './src/requestIdentity.mjs'
console.log(getBrowserPageOptions({ userAgent: 'X/1.0' }))
// => { userAgent: 'X/1.0' }
console.log(getBrowserPageOptions({}))
// => {}
Parameters:
| Name | Type | Description |
|---|---|---|
opt |
Object | 輸入設定物件 |
Returns:
回傳Playwright之newPage選項物件
- Type
- Object
getCurlIdentity(opt) → {Object}
- Description:
取得curl階之HTTP請求身分,未指定者採預設值
- Source:
Example
import { getCurlIdentity } from './src/requestIdentity.mjs'
console.log(getCurlIdentity({}).referer)
// => 'https://www.google.com/'
Parameters:
| Name | Type | Description |
|---|---|---|
opt |
Object | 輸入設定物件 |
Returns:
回傳{userAgent,referer,acceptLanguage}物件
- Type
- Object
getOptArr(opt, key, def) → {Array}
- Description:
由設定物件取陣列選項,型別不符時採預設值
- Source:
Example
import { getOptArr } from './src/getOpt.mjs'
console.log(getOptArr({ adapters: [1] }, 'adapters', []), getOptArr({ adapters: 'x' }, 'adapters', []))
// => [1] []
Parameters:
| Name | Type | Description |
|---|---|---|
opt |
Object | 輸入設定物件 |
key |
String | 輸入鍵名字串 |
def |
Array | 輸入預設值陣列 |
Returns:
回傳選項陣列
- Type
- Array
getOptBool(opt, key, def) → {Boolean}
- Description:
由設定物件取布林選項,型別不符時採預設值
- Source:
Example
import { getOptBool } from './src/getOpt.mjs'
console.log(getOptBool({ parse: false }, 'parse', true), getOptBool({ parse: 'x' }, 'parse', true))
// => false true
Parameters:
| Name | Type | Description |
|---|---|---|
opt |
Object | 輸入設定物件 |
key |
String | 輸入鍵名字串 |
def |
Boolean | 輸入預設值布林值 |
Returns:
回傳選項布林值
- Type
- Boolean
getOptP0Int(opt, key, def) → {Integer}
- Description:
由設定物件取非負整數選項,型別不符時採預設值
- Source:
Example
import { getOptP0Int } from './src/getOpt.mjs'
console.log(getOptP0Int({ maxRetries: 0 }, 'maxRetries', 5), getOptP0Int({ maxRetries: -1 }, 'maxRetries', 5))
// => 0 5
Parameters:
| Name | Type | Description |
|---|---|---|
opt |
Object | 輸入設定物件 |
key |
String | 輸入鍵名字串 |
def |
Integer | 輸入預設值整數 |
Returns:
回傳選項整數
- Type
- Integer
getOptPInt(opt, key, def) → {Integer}
- Description:
由設定物件取正整數選項,型別不符時採預設值
- Source:
Example
import { getOptPInt } from './src/getOpt.mjs'
console.log(getOptPInt({ timeoutMs: 3000 }, 'timeoutMs', 15000), getOptPInt({ timeoutMs: 0 }, 'timeoutMs', 15000))
// => 3000 15000
Parameters:
| Name | Type | Description |
|---|---|---|
opt |
Object | 輸入設定物件 |
key |
String | 輸入鍵名字串 |
def |
Integer | 輸入預設值整數 |
Returns:
回傳選項整數
- Type
- Integer
getOptStr(opt, key, def) → {String}
- Description:
由設定物件取非空字串選項,型別不符時採預設值
- Source:
Example
import { getOptStr } from './src/getOpt.mjs'
console.log(getOptStr({ method: 'curl' }, 'method', 'auto'), getOptStr({ method: '' }, 'method', 'auto'))
// => 'curl' 'auto'
Parameters:
| Name | Type | Description |
|---|---|---|
opt |
Object | 輸入設定物件 |
key |
String | 輸入鍵名字串 |
def |
String | 輸入預設值字串 |
Returns:
回傳選項字串
- Type
- String
getRequestIdentity(opt) → {Object}
- Description:
取得使用者明確指定之HTTP請求身分,未指定者為空字串
供瀏覽器階使用:只有非空者才覆寫瀏覽器自身之身分
- Source:
Example
import { getRequestIdentity } from './src/requestIdentity.mjs'
console.log(getRequestIdentity({ userAgent: 'X/1.0' }))
// => { userAgent: 'X/1.0', referer: '', acceptLanguage: '' }
Parameters:
| Name | Type | Description |
|---|---|---|
opt |
Object | 輸入設定物件 |
Returns:
回傳{userAgent,referer,acceptLanguage}物件,未指定之項目為空字串
- Type
- Object
getRetryWaitMs(attempt) → {Integer}
- Description:
取得重試前之線性退避等待毫秒
第n次失敗後等待n*3000毫秒,上限15000毫秒,即3000、6000、9000、12000、15000、15000...
- Source:
Example
import getRetryWaitMs from './src/getRetryWaitMs.mjs'
console.log(getRetryWaitMs(1), getRetryWaitMs(2), getRetryWaitMs(5), getRetryWaitMs(99))
// => 3000 6000 15000 15000
Parameters:
| Name | Type | Description |
|---|---|---|
attempt |
Integer | 輸入第幾次嘗試之正整數,由1起算 |
Returns:
回傳等待毫秒整數
- Type
- Integer
getUrlErrorResult(url, method, fetchedAt) → {Object|null}
- Description:
檢核網址並取得錯誤結果物件
供各抓取函數於入口統一檢核網址,網址有效時回傳null代表可繼續執行
- Source:
Example
import getUrlErrorResult from './src/getUrlErrorResult.mjs'
console.log(getUrlErrorResult('https://a.com/', 'curl', 'now'))
// => null
console.log(getUrlErrorResult('abc', 'curl', 'now'))
// => { status: 'error', url: 'abc', message: 'invalid url (must be http/https)', reason: 'invalid-url', method: 'curl', fetchedAt: 'now', attempts: 0 }
Parameters:
| Name | Type | Description |
|---|---|---|
url |
String | 輸入待檢核網址字串 |
method |
String | 輸入抓取方法名稱字串,將寫入結果物件之method欄位 |
fetchedAt |
String | 輸入抓取時間字串,將寫入結果物件之fetchedAt欄位 |
Returns:
網址無效時回傳錯誤結果物件{status:'error',url,message,reason:'invalid-url',method,fetchedAt,attempts:0},網址有效時回傳null
- Type
- Object | null
inspectHtml(html, optopt) → {Object}
- Description:
由網頁原始HTML檢測頁面是否為有效內容
純粹基於原始HTML結構判斷,不依賴Readability,用於判識CAPTCHA與反爬蟲挑戰頁、驗證頁、 轉址包裝頁、空內容頁,供階梯升級決策使用。 判識器以資料表定義並依序比對,首個命中者勝出;順序本身即語意,調動順序會改變判定結果。
contentKind標明待測內容是抓取器取回的原始文件,或由已渲染DOM萃取後合成者。 合成內容之標籤結構已被剝除,故只比對semantic類判識器;預設為'raw', 亦即未指定時行為與原先一致
- Source:
Example
import inspectHtml from './src/inspectHtml.mjs'
console.log(inspectHtml('<html><head><title>Just a moment</title></head><body></body></html>'))
// => { pass: false, type: 'captcha', message: 'Cloudflare/anti-bot challenge' }
console.log(inspectHtml('<html><head><title>abc</title></head><body><p>' + 'x'.repeat(300) + '</p></body></html>'))
// => { pass: true, type: 'pass', message: 'ok' }
Parameters:
| Name | Type | Attributes | Default | Description | ||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
html |
String | 輸入網頁HTML字串 |
||||||||||||||||||||||
opt |
Object |
<optional> |
{}
|
輸入設定物件,預設{} Properties
|
Returns:
回傳檢測結果物件,格式為{pass,type,message},其中pass為是否通過布林值,type為'pass'、'captcha'、'verify'、'redirect'、'empty'之一,message為說明字串
- Type
- Object
isInternalHost(hostname) → {Boolean}
- Description:
判別主機名是否指向內網、迴環或保留位址
僅供把關本套件自行推導出之網址(如由轉址服務query參數提取者), 不可用於把關呼叫端明確給定之網址——後者是呼叫端的決定
- Source:
Example
import isInternalHost from './src/isInternalHost.mjs'
console.log(isInternalHost('169.254.169.254'), isInternalHost('127.0.0.1'), isInternalHost('localhost'))
// => true true true
console.log(isInternalHost('example.com'), isInternalHost('8.8.8.8'))
// => false false
Parameters:
| Name | Type | Description |
|---|---|---|
hostname |
String | 輸入主機名字串,不含協定與埠 |
Returns:
回傳是否為內網或保留位址之布林值
- Type
- Boolean
isValidAdapter(adapter) → {Boolean}
- Description:
檢核adapter條目是否符合輸入契約
- Source:
Example
import { isValidAdapter } from './src/adapterContract.mjs'
console.log(isValidAdapter({ id: 'a', match: /x/, parse: () => ({}) }), isValidAdapter({ id: 'a', match: 'x', parse: () => ({}) }))
// => true false
Parameters:
| Name | Type | Description |
|---|---|---|
adapter |
* | 輸入待檢核之adapter |
Returns:
回傳是否合法之布林值
- Type
- Boolean
isValidDetector(detector) → {Boolean}
- Description:
檢核使用端判識器是否符合契約
不合法者一律略過而非拋錯,與adapter之處置一致
- Source:
Example
import { isValidDetector } from './src/detectorContract.mjs'
console.log(isValidDetector({ type: 'captcha', message: 'x', test: () => true }))
// => true
console.log(isValidDetector({ type: 'unknown', message: 'x', test: () => true }))
// => false
Parameters:
| Name | Type | Description |
|---|---|---|
detector |
* | 輸入待檢核之判識器 |
Returns:
回傳是否合法之布林值
- Type
- Boolean
isValidUrl(url) → {Boolean}
- Description:
檢核是否為有效的http或https網址字串
- Source:
Example
import isValidUrl from './src/isValidUrl.mjs'
console.log(isValidUrl('https://www.google.com/'))
// => true
console.log(isValidUrl('http://127.0.0.1:8080/abc'))
// => true
console.log(isValidUrl('ftp://www.google.com/'))
// => false
console.log(isValidUrl('www.google.com'))
// => false
Parameters:
| Name | Type | Description |
|---|---|---|
url |
String | 輸入待檢核網址字串 |
Returns:
回傳是否為有效http或https網址之布林值
- Type
- Boolean
matchMsn(url) → {Object|null}
- Description:
比對 msn 內容頁網址並取出內容型別與 id
作為內建 msn adapter 之 match 掛點。回傳之物件即為 ctx,傳入 fetch 掛點
- Source:
Example
import { matchMsn } from './src/fetchMsn.mjs'
console.log(matchMsn('https://www.msn.com/zh-tw/news/other/abc/ar-AA2bZm9d'))
// => { kind: 'ar', id: 'AA2bZm9d' }
console.log(matchMsn('https://www.msn.com/en-us/video/news/abc/vi-AA2bYtCB'))
// => { kind: 'vi', id: 'AA2bYtCB' }
console.log(matchMsn('https://www.msn.com/zh-tw/news'))
// => null
Parameters:
| Name | Type | Description |
|---|---|---|
url |
String | 輸入網址字串 |
Returns:
命中時回傳{kind,id},kind為'ar'(文章)或'vi'(影片);未命中時回傳null
- Type
- Object | null
meetsMinContent(content) → {Boolean}
- Description:
判定正文長度是否達最低門檻
adapter路徑與Readability路徑刻意採同一標準,故兩處皆呼叫本函數而非各自比較, 避免同一規則手寫兩處而日後分歧
- Source:
Example
import { meetsMinContent } from './src/adapterContract.mjs'
console.log(meetsMinContent('x'.repeat(50)), meetsMinContent('x'.repeat(49)))
// => true false
Parameters:
| Name | Type | Description |
|---|---|---|
content |
String | 輸入正文字串 |
Returns:
回傳是否達門檻之布林值
- Type
- Boolean
(async) navigateWithRedirectWait(page, url, navTimeout) → {Promise}
- Description:
導航至網址並等待JS轉址完成
先以domcontentloaded導航,再等待網址host脫離原host(代表已轉址),最後等待networkidle; 等待轉址與networkidle皆為盡力而為,逾時不視為失敗
- Source:
Example
import { chromium } from 'playwright'
import navigateWithRedirectWait from './src/navigateWithRedirectWait.mjs'
let test = async () => {
let browser = await chromium.launch({ headless: true, channel: 'chrome' })
let page = await browser.newPage()
await navigateWithRedirectWait(page, 'https://news.google.com/articles/xxxxxx', 15000)
console.log(page.url())
// => 'https://www.example-news.com/real-article'
await browser.close()
}
await test()
.catch((err) => {
console.log(err)
})
Parameters:
| Name | Type | Description |
|---|---|---|
page |
Object | 輸入Playwright之page物件 |
url |
String | 輸入待導航網址字串 |
navTimeout |
Integer | 輸入導航最長等待毫秒整數 |
Returns:
回傳Promise,resolve回傳首次導航之Response物件(供呼叫端檢核HTTP狀態),同頁錨點導航等情形可能為null
- Type
- Promise
normalizeDetector(detector) → {Object}
- Description:
把使用端判識器正規化為內部形狀,補上預設之evidence並標記來源
origin標記為'custom',供判識流程據以略過內容量閘門,理由見本檔檔頭
- Source:
Example
import { normalizeDetector } from './src/detectorContract.mjs'
console.log(normalizeDetector({ type: 'captcha', message: 'x', test: () => true }).evidence)
// => 'semantic'
Parameters:
| Name | Type | Description |
|---|---|---|
detector |
Object | 輸入已通過檢核之判識器 |
Returns:
回傳正規化後之判識器物件
- Type
- Object
normalizeParsed(parsed, pre) → {Object}
- Description:
把adapter之parse回傳值檢核並投影為固定形狀
- Source:
Example
import { normalizeParsed } from './src/adapterContract.mjs'
console.log(normalizeParsed({ success: 'true', content: 'x' }, 'adapter a '))
// => { success: false, reason: 'parse-error', message: 'adapter a returned non-boolean success' }
Parameters:
| Name | Type | Description |
|---|---|---|
parsed |
* | 輸入adapter.parse之回傳值 |
pre |
String | 輸入錯誤訊息前綴字串 |
Returns:
回傳內部結構結果物件,成功時為{success:true,title,content,contentLength},失敗時為{success:false,reason,message}
- Type
- Object
(async) parseArticle(html, url, hit, metaopt) → {Promise}
- Description:
依已解析之adapter命中結果解析文章,未命中者走Readability
adapter之挑選(findAdapter)刻意不在本函數內:它只取決於網址,與抓回之HTML無關, 故由runPlan於計畫執行前解析一次後傳入,詳見runPlan之_resolveAdapter
- Source:
Parameters:
| Name | Type | Attributes | Default | Description |
|---|---|---|---|---|
html |
String | 輸入網頁HTML字串 |
||
url |
String | 輸入網址字串 |
||
hit |
Object | null | 輸入findAdapter之結果物件,null視為未命中 |
||
meta |
Object | null |
<optional> |
null
|
輸入本次抓取之附帶資訊物件,含finalUrl等,供adapter判斷與JSDOM之base url使用,預設null |
Returns:
回傳Promise,resolve回傳解析結果物件,本函數不會reject
- Type
- Promise
parseBloomberg(html, url) → {Object}
- Description:
解析Bloomberg之文章內文
由