跳到主内容
智客 ZICQ

技能库 智客分类:其他 playwright-scraper-skill OpenClaw

Playwright Scraper Skill

Playwright网络刮出OpenClaw Skill 有反机器人保护. 在Lore.com.hk等复杂地段成功测试.

1079 安装量 · 57 星标

官方网址:ClawHub

技能介绍

先看中文介绍;官方 description 原文单独保留,不改写 SKILL.md。

做什么

Playwright网络刮出OpenClaw Skill 有反机器人保护. 在Lore.com.hk等复杂地段成功测试.

何时用

官方 description 未单独写出 Use when。按规范,代理会在用户任务与这段 description 的关键词匹配时激活本技能。

代理如何加载

按 Agent Skills 渐进披露:启动时只加载 name 与 description(约 100 token);任务匹配后才读入整份 SKILL.md 正文;scripts/、references/、assets/ 仅在需要时再读。 本文件正文结构:Playwright Scraper Skill、🎯 Use Case Matrix、📦 Installation、🚀 Quick Start、1️⃣ Simple Sites (No Anti-Bot)、Invoke directly in OpenClaw。

文件分析

文件分析:除 SKILL.md 外,正文引用了 scripts/playwright-simple.js、scripts/playwright-stealth.js、assets/youtube_handler.js,属于带资源的技能包,这些文件按需再读。

官方 description(原文)

Playwright-based web scraping OpenClaw Skill with anti-bot protection. Successfully tested on complex sites like Discuss.com.hk.

Playwright Scraper Skill🎯 Use Case Matrix📦 Installation🚀 Quick Start1️⃣ Simple Sites (No Anti-Bot)Invoke directly in OpenClaw2️⃣ Dynamic Sites (Requires JavaScript)3️⃣ Anti-Bot Protected Sites (Cloudflare etc.)4️⃣ YouTube Video TranscriptsInstall deep-scraper skillUse it📖 Script Descriptions

来源分类:ClawHub Playwright

SKILL.md 与 Agent 调用

官方规范 ↗
name
playwright-scraper-skill
description
Playwright-based web scraping OpenClaw Skill with anti-bot protection. Successfully tested on complex sites like Discuss.com.hk.
  1. 发现技能客户端向 Agent 提供名称与描述目录。
  2. 匹配与调用用户指定或任务匹配后,载入 SKILL.md 指令。
  3. 按需加载按步骤读取参考文档、使用脚本与素材。
指令中引用的文件 · 3
  • scripts/playwright-simple.js
  • scripts/playwright-stealth.js
  • assets/youtube_handler.js

以下路径提取自原文;文件是否齐全请以来源仓库中的完整目录为准。

具体调用语法与可用工具以目标 Agent 客户端为准。 查看调用机制说明 ↗

安装这个技能

Skills CLI ↗

先选择目标 Agent 和安装范围,保留技能包的附属文件,安装后检查客户端能否发现该技能。

该技能引用了附属文件,请从来源获取完整目录;仅复制 SKILL.md 可能缺少依赖。

交给 Agent 安装

复制安装指令给支持 Agent Skills 的代理,确认其中的目标目录与客户端匹配。

把 Agent Skill「playwright-scraper-skill」安装到我的项目:SKILL.md 原文与官方 description 见 https://zicq.com/zh/skills/skl-da78303e9ab49d70-Playwright-Scraper-Skill.html
请存为 .cursor/skills/playwright-scraper-skill/SKILL.md 或 .claude/skills/playwright-scraper-skill/SKILL.md,frontmatter 的 name 与 description 保持原样,不要改写。
该技能还带 scripts/、references/、assets/ 等文件,请从 https://clawhub.ai/skills/playwright-scraper-skill 取完整目录,不要只建一个 SKILL.md。

当前没有明确的 GitHub 技能包地址,请按来源页面的安装器说明操作。

ClawHub ↗

阅读排版

name: playwright-scraper-skill description: Playwright-based web scraping OpenClaw Skill with anti-bot protection. Successfully tested on complex sites like Discuss.com.hk. version: 1.2.0 author: Simon Chan

Playwright Scraper Skill

A Playwright-based web scraping OpenClaw Skill with anti-bot protection. Choose the best approach based on the target website's anti-bot level.


🎯 Use Case Matrix

| Target Website | Anti-Bot Level | Recommended Method | Script | |---------------|----------------|-------------------|--------| | Regular Sites | Low | web_fetch tool | N/A (built-in) | | Dynamic Sites | Medium | Playwright Simple | scripts/playwright-simple.js | | Cloudflare Protected | High | Playwright Stealth ⭐ | scripts/playwright-stealth.js | | YouTube | Special | deep-scraper | Install separately | | Reddit | Special | reddit-scraper | Install separately |


📦 Installation

cd playwright-scraper-skill
npm install
npx playwright install chromium

🚀 Quick Start

1️⃣ Simple Sites (No Anti-Bot)

Use OpenClaw's built-in web_fetch tool:

# Invoke directly in OpenClaw
Hey, fetch me the content from https://example.com

2️⃣ Dynamic Sites (Requires JavaScript)

Use Playwright Simple:

node scripts/playwright-simple.js "https://example.com"

Example output:

{
  "url": "https://example.com",
  "title": "Example Domain",
  "content": "...",
  "elapsedSeconds": "3.45"
}

3️⃣ Anti-Bot Protected Sites (Cloudflare etc.)

Use Playwright Stealth:

node scripts/playwright-stealth.js "https://m.discuss.com.hk/#hot"

Features:

  • Hide automation markers (navigator.webdriver = false)
  • Realistic User-Agent (iPhone, Android)
  • Random delays to mimic human behavior
  • Screenshot and HTML saving support

4️⃣ YouTube Video Transcripts

Use deep-scraper (install separately):

# Install deep-scraper skill
npx clawhub install deep-scraper

# Use it
cd skills/deep-scraper
node assets/youtube_handler.js "https://www.youtube.com/watch?v=VIDEO_ID"

📖 Script Descriptions

scripts/playwright-simple.js

  • Use Case: Regular dynamic websites
  • Speed: Fast (3-5 seconds)
  • Anti-Bot: None
  • Output: JSON (title, content, URL)

scripts/playwright-stealth.js ⭐

  • Use Case: Sites with Cloudflare or anti-bot protection
  • Speed: Medium (5-20 seconds)
  • Anti-Bot: Medium-High (hides automation, realistic UA)
  • Output: JSON + Screenshot + HTML file
  • Verified: 100% success on Discuss.com.hk

🎓 Best Practices

1. Try web_fetch First

If the site doesn't have dynamic loading, use OpenClaw's web_fetch tool—it's fastest.

2. Need JavaScript? Use Playwright Simple

If you need to wait for JavaScript rendering, use playwright-simple.js.

3. Getting Blocked? Use Stealth

If you encounter 403 or Cloudflare challenges, use playwright-stealth.js.

4. Special Sites Need Specialized Skills

  • YouTube → deep-scraper
  • Reddit → reddit-scraper
  • Twitter → bird skill

🔧 Customization

All scripts support environment variables:

# Set screenshot path
SCREENSHOT_PATH=/path/to/screenshot.png node scripts/playwright-stealth.js URL

# Set wait time (milliseconds)
WAIT_TIME=10000 node scripts/playwright-simple.js URL

# Enable headful mode (show browser)
HEADLESS=false node scripts/playwright-stealth.js URL

# Save HTML
SAVE_HTML=true node scripts/playwright-stealth.js URL

# Custom User-Agent
USER_AGENT="Mozilla/5.0 ..." node scripts/playwright-stealth.js URL

📊 Performance Comparison

| Method | Speed | Anti-Bot | Success Rate (Discuss.com.hk) | |--------|-------|----------|-------------------------------| | web_fetch | ⚡ Fastest | ❌ None | 0% | | Playwright Simple | 🚀 Fast | ⚠️ Low | 20% | | Playwright Stealth | ⏱️ Medium | ✅ Medium | 100% ✅ | | Puppeteer Stealth | ⏱️ Medium | ✅ Medium-High | ~80% | | Crawlee (deep-scraper) | 🐢 Slow | ❌ Detected | 0% | | Chaser (Rust) | ⏱️ Medium | ❌ Detected | 0% |


🛡️ Anti-Bot Techniques Summary

Lessons learned from our testing:

✅ Effective Anti-Bot Measures

  1. Hide navigator.webdriver — Essential
  2. Realistic User-Agent — Use real devices (iPhone, Android)
  3. Mimic Human Behavior — Random delays, scrolling
  4. Avoid Framework Signatures — Crawlee, Selenium are easily detected
  5. Use addInitScript (Playwright) — Inject before page load

❌ Ineffective Anti-Bot Measures

  1. Only changing User-Agent — Not enough
  2. Using high-level frameworks (Crawlee) — More easily detected
  3. Docker isolation — Doesn't help with Cloudflare

🔍 Troubleshooting

Issue: 403 Forbidden

Solution: Use playwright-stealth.js

Issue: Cloudflare Challenge Page

Solution:

  1. Increase wait time (10-15 seconds)
  2. Try headless: false (headful mode sometimes has higher success rate)
  3. Consider using proxy IPs

Issue: Blank Page

Solution:

  1. Increase waitForTimeout
  2. Use waitUntil: 'networkidle' or 'domcontentloaded'
  3. Check if login is required

📝 Memory & Experience

2026-02-07 Discuss.com.hk Test Conclusions

  • ✅ Pure Playwright + Stealth succeeded (5s, 200 OK)
  • ❌ Crawlee (deep-scraper) failed (403)
  • ❌ Chaser (Rust) failed (Cloudflare)
  • ❌ Puppeteer standard failed (403)

Best Solution: Pure Playwright + anti-bot techniques (framework-independent)


🚧 Future Improvements

  • [ ] Add proxy IP rotation
  • [ ] Implement cookie management (maintain login state)
  • [ ] Add CAPTCHA handling (2captcha / Anti-Captcha)
  • [ ] Batch scraping (parallel URLs)
  • [ ] Integration with OpenClaw's browser tool

📚 References

相关技能

其他

Self Improving Agent

掌握学习、错误和校正,以便不断改进。 使用时间:(1) 命令或操作意外失败, (2) 用户更正 Claude ('No, that's wrong...', 'Actually...'), (3)用户请求具备不存在的能力,(4)外部API或工具失败,(5)克洛德意识到其知识已经过…

其他

Self Improving + Proactive Agent

自我反省+自我批评+自我学习+自我组织记忆. 特工评估自己的工作,发现错误,并永久改进. 当 (1) 命令,工具, API, 或操作失败时使用; (2) 用户纠正或拒绝您的工作; (3) 您意识到您的知识已经过时或不正确; (4) 您发现更好的方法; (5) 用户明确安装或引用当…

其他

Proactive Agent

将AI代理人从任务跟踪者转变为预先预测需要并不断改进的主动伙伴. 现在WAL协议,工作缓冲,自主克龙, 和战斗测试模式。 哈尔·斯塔克的一部分

其他

Multi Search Engine

多搜索引擎集成16个引擎(7个CN+9 Global). 支持高级搜索操作员,时间过滤器,站点搜索,隐私引擎,和WolframAlpha知识查询. 不需要 API 密钥 .