
Web Data Extraction Pipeline
What this is
A complete pipeline for turning a messy public webpage or JSON API response into a clean, typed, deduplicated dataset. Not a scraping bot that runs on a schedule -- a methodology plus ready-to-run code for the fetch -> parse -> validate -> dedupe -> store steps, with a Pydantic schema at the center so bad data fails loudly instead of silently corrupting your dataset.
Why this is different from a generic scraper prompt
Most scraping code that gets generated ad-hoc skips validation and dedup entirely, and just dumps raw dicts to a file. This skill treats those as first-class steps: every record goes through a typed schema before it's saved, duplicates are caught by both URL and content hash, and there's a built-in legal/ethics checklist (robots.txt, ToS, rate limiting) run before any code gets written.
What's included
- Legal/ethics checklist (robots.txt, ToS, public-data-only, rate limiting) -- run before writing any code
- API-first strategy: check DevTools Network tab before parsing HTML
- Pydantic schema template with common validators (price parsing, rating parsing)
- CSS Selector and XPath extraction patterns
- Pagination with required rate limiting
- Deduplication (URL + content hash)
- Storage in JSON Lines, CSV, or SQLite
- English and Traditional Chinese versions, both fully independent (not machine-translated from each other)
FAQ
Does this bypass CAPTCHAs or anti-bot protection?
No. This skill is scoped to publicly accessible pages and APIs, with realistic headers -- it does not include CAPTCHA-solving or fingerprint-spoofing techniques.
What if the site is JS-rendered?
The skill flags this case and tells you to switch to a browser-automation approach instead of plain requests.
Can I get SQLite output instead of CSV?
Yes, all three storage backends are included; pick whichever fits your downstream use.
中文版
這是什麼
把雜亂的公開網頁或JSON API回應,轉成乾淨、有型別、已去重的資料集。核心是用Pydantic schema把關,讓壞資料直接報錯,而不是悄悄污染你的資料集。
包含什麼
- 法遵/倫理檢查清單(robots.txt、服務條款、只收公開資料、速率限制)
- API優先策略:解析HTML前先查DevTools Network分頁
- Pydantic schema範本,含常見驗證器
- CSS Selector與XPath萃取範例
- 去重(URL + 內容hash)
- JSON Lines / CSV / SQLite三種儲存方式
- 中文版與英文版皆完整獨立撰寫,不是互相機器翻譯
常見問題
這會繞過CAPTCHA或反機器人防護嗎?
不會。這個skill的範圍限定在公開可存取的頁面與API。
如果網站是JS動態渲染怎麼辦?
skill會標記出這種狀況,並提示你改用瀏覽器自動化方案。


