開始使用
Web Data Extraction Pipeline

Web Data Extraction Pipeline

Turn any public webpage or API response into a clean, validated, deduplicated structured dataset (JSON/CSV/SQLite): schema design with Pydantic, CSS/XPath extraction, cleaning rules, and storage, plus a built-in legal/ethics checklist (robots.txt, ToS, rate limiting). --- 把任何公開網頁或API回應轉成乾淨、已驗證、已去重的結構化資料集(JSON/CSV/SQLite):Pydantic schema設計、CSS/XPath萃取、清洗規則、儲存,並內建法遵/倫理檢查清單(robots.txt、服務條款、速率限制)。中英文版本皆完整獨立撰寫。
#工程#研究#生產力
評分
需要更多評價
已售出
0
使用方式
在 Capafy 上執行
也可用於外部應用程式
由創作者提供
Claude Sonnet 5

What this is

A complete pipeline for turning a messy public webpage or JSON API response into a clean, typed, deduplicated dataset. Not a scraping bot that runs on a schedule -- a methodology plus ready-to-run code for the fetch -> parse -> validate -> dedupe -> store steps, with a Pydantic schema at the center so bad data fails loudly instead of silently corrupting your dataset.

Why this is different from a generic scraper prompt

Most scraping code that gets generated ad-hoc skips validation and dedup entirely, and just dumps raw dicts to a file. This skill treats those as first-class steps: every record goes through a typed schema before it's saved, duplicates are caught by both URL and content hash, and there's a built-in legal/ethics checklist (robots.txt, ToS, rate limiting) run before any code gets written.

What's included

  • Legal/ethics checklist (robots.txt, ToS, public-data-only, rate limiting) -- run before writing any code
  • API-first strategy: check DevTools Network tab before parsing HTML
  • Pydantic schema template with common validators (price parsing, rating parsing)
  • CSS Selector and XPath extraction patterns
  • Pagination with required rate limiting
  • Deduplication (URL + content hash)
  • Storage in JSON Lines, CSV, or SQLite
  • English and Traditional Chinese versions, both fully independent (not machine-translated from each other)

FAQ

Does this bypass CAPTCHAs or anti-bot protection?
No. This skill is scoped to publicly accessible pages and APIs, with realistic headers -- it does not include CAPTCHA-solving or fingerprint-spoofing techniques.

What if the site is JS-rendered?
The skill flags this case and tells you to switch to a browser-automation approach instead of plain requests.

Can I get SQLite output instead of CSV?
Yes, all three storage backends are included; pick whichever fits your downstream use.

中文版

這是什麼

把雜亂的公開網頁或JSON API回應,轉成乾淨、有型別、已去重的資料集。核心是用Pydantic schema把關,讓壞資料直接報錯,而不是悄悄污染你的資料集。

包含什麼

  • 法遵/倫理檢查清單(robots.txt、服務條款、只收公開資料、速率限制)
  • API優先策略:解析HTML前先查DevTools Network分頁
  • Pydantic schema範本,含常見驗證器
  • CSS Selector與XPath萃取範例
  • 去重(URL + 內容hash)
  • JSON Lines / CSV / SQLite三種儲存方式
  • 中文版與英文版皆完整獨立撰寫,不是互相機器翻譯

常見問題

這會繞過CAPTCHA或反機器人防護嗎?
不會。這個skill的範圍限定在公開可存取的頁面與API。

如果網站是JS動態渲染怎麼辦?
skill會標記出這種狀況,並提示你改用瀏覽器自動化方案。