做什么
Playwright网络刮出OpenClaw Skill 有反机器人保护. 在Lore.com.hk等复杂地段成功测试.
技能库 智客分类:其他 playwright-scraper-skill OpenClaw
Playwright网络刮出OpenClaw Skill 有反机器人保护. 在Lore.com.hk等复杂地段成功测试.
官方网址:ClawHub
先看中文介绍;官方 description 原文单独保留,不改写 SKILL.md。
Playwright网络刮出OpenClaw Skill 有反机器人保护. 在Lore.com.hk等复杂地段成功测试.
官方 description 未单独写出 Use when。按规范,代理会在用户任务与这段 description 的关键词匹配时激活本技能。
按 Agent Skills 渐进披露:启动时只加载 name 与 description(约 100 token);任务匹配后才读入整份 SKILL.md 正文;scripts/、references/、assets/ 仅在需要时再读。 本文件正文结构:Playwright Scraper Skill、🎯 Use Case Matrix、📦 Installation、🚀 Quick Start、1️⃣ Simple Sites (No Anti-Bot)、Invoke directly in OpenClaw。
文件分析:除 SKILL.md 外,正文引用了 scripts/playwright-simple.js、scripts/playwright-stealth.js、assets/youtube_handler.js,属于带资源的技能包,这些文件按需再读。
Playwright-based web scraping OpenClaw Skill with anti-bot protection. Successfully tested on complex sites like Discuss.com.hk.
Playwright Scraper Skill🎯 Use Case Matrix📦 Installation🚀 Quick Start1️⃣ Simple Sites (No Anti-Bot)Invoke directly in OpenClaw2️⃣ Dynamic Sites (Requires JavaScript)3️⃣ Anti-Bot Protected Sites (Cloudflare etc.)4️⃣ YouTube Video TranscriptsInstall deep-scraper skillUse it📖 Script Descriptions
来源分类:ClawHub Playwright
市场来源:ClawHub
nameplaywright-scraper-skilldescriptionscripts/playwright-simple.jsscripts/playwright-stealth.jsassets/youtube_handler.js以下路径提取自原文;文件是否齐全请以来源仓库中的完整目录为准。
具体调用语法与可用工具以目标 Agent 客户端为准。 查看调用机制说明 ↗
先选择目标 Agent 和安装范围,保留技能包的附属文件,安装后检查客户端能否发现该技能。
该技能引用了附属文件,请从来源获取完整目录;仅复制 SKILL.md 可能缺少依赖。
复制安装指令给支持 Agent Skills 的代理,确认其中的目标目录与客户端匹配。
把 Agent Skill「playwright-scraper-skill」安装到我的项目:SKILL.md 原文与官方 description 见 https://zicq.com/zh/skills/skl-da78303e9ab49d70-Playwright-Scraper-Skill.html 请存为 .cursor/skills/playwright-scraper-skill/SKILL.md 或 .claude/skills/playwright-scraper-skill/SKILL.md,frontmatter 的 name 与 description 保持原样,不要改写。 该技能还带 scripts/、references/、assets/ 等文件,请从 https://clawhub.ai/skills/playwright-scraper-skill 取完整目录,不要只建一个 SKILL.md。
当前没有明确的 GitHub 技能包地址,请按来源页面的安装器说明操作。
A Playwright-based web scraping OpenClaw Skill with anti-bot protection. Choose the best approach based on the target website's anti-bot level.
| Target Website | Anti-Bot Level | Recommended Method | Script |
|---------------|----------------|-------------------|--------|
| Regular Sites | Low | web_fetch tool | N/A (built-in) |
| Dynamic Sites | Medium | Playwright Simple | scripts/playwright-simple.js |
| Cloudflare Protected | High | Playwright Stealth ⭐ | scripts/playwright-stealth.js |
| YouTube | Special | deep-scraper | Install separately |
| Reddit | Special | reddit-scraper | Install separately |
cd playwright-scraper-skill
npm install
npx playwright install chromium
Use OpenClaw's built-in web_fetch tool:
# Invoke directly in OpenClaw
Hey, fetch me the content from https://example.com
Use Playwright Simple:
node scripts/playwright-simple.js "https://example.com"
Example output:
{
"url": "https://example.com",
"title": "Example Domain",
"content": "...",
"elapsedSeconds": "3.45"
}
Use Playwright Stealth:
node scripts/playwright-stealth.js "https://m.discuss.com.hk/#hot"
Features:
navigator.webdriver = false)Use deep-scraper (install separately):
# Install deep-scraper skill
npx clawhub install deep-scraper
# Use it
cd skills/deep-scraper
node assets/youtube_handler.js "https://www.youtube.com/watch?v=VIDEO_ID"
scripts/playwright-simple.jsscripts/playwright-stealth.js ⭐If the site doesn't have dynamic loading, use OpenClaw's web_fetch tool—it's fastest.
If you need to wait for JavaScript rendering, use playwright-simple.js.
If you encounter 403 or Cloudflare challenges, use playwright-stealth.js.
All scripts support environment variables:
# Set screenshot path
SCREENSHOT_PATH=/path/to/screenshot.png node scripts/playwright-stealth.js URL
# Set wait time (milliseconds)
WAIT_TIME=10000 node scripts/playwright-simple.js URL
# Enable headful mode (show browser)
HEADLESS=false node scripts/playwright-stealth.js URL
# Save HTML
SAVE_HTML=true node scripts/playwright-stealth.js URL
# Custom User-Agent
USER_AGENT="Mozilla/5.0 ..." node scripts/playwright-stealth.js URL
| Method | Speed | Anti-Bot | Success Rate (Discuss.com.hk) | |--------|-------|----------|-------------------------------| | web_fetch | ⚡ Fastest | ❌ None | 0% | | Playwright Simple | 🚀 Fast | ⚠️ Low | 20% | | Playwright Stealth | ⏱️ Medium | ✅ Medium | 100% ✅ | | Puppeteer Stealth | ⏱️ Medium | ✅ Medium-High | ~80% | | Crawlee (deep-scraper) | 🐢 Slow | ❌ Detected | 0% | | Chaser (Rust) | ⏱️ Medium | ❌ Detected | 0% |
Lessons learned from our testing:
navigator.webdriver — EssentialaddInitScript (Playwright) — Inject before page loadSolution: Use playwright-stealth.js
Solution:
headless: false (headful mode sometimes has higher success rate)Solution:
waitForTimeoutwaitUntil: 'networkidle' or 'domcontentloaded'Best Solution: Pure Playwright + anti-bot techniques (framework-independent)
browser tool其他
掌握学习、错误和校正,以便不断改进。 使用时间:(1) 命令或操作意外失败, (2) 用户更正 Claude ('No, that's wrong...', 'Actually...'), (3)用户请求具备不存在的能力,(4)外部API或工具失败,(5)克洛德意识到其知识已经过…
其他
自我反省+自我批评+自我学习+自我组织记忆. 特工评估自己的工作,发现错误,并永久改进. 当 (1) 命令,工具, API, 或操作失败时使用; (2) 用户纠正或拒绝您的工作; (3) 您意识到您的知识已经过时或不正确; (4) 您发现更好的方法; (5) 用户明确安装或引用当…
其他
将AI代理人从任务跟踪者转变为预先预测需要并不断改进的主动伙伴. 现在WAL协议,工作缓冲,自主克龙, 和战斗测试模式。 哈尔·斯塔克的一部分
其他
多搜索引擎集成16个引擎(7个CN+9 Global). 支持高级搜索操作员,时间过滤器,站点搜索,隐私引擎,和WolframAlpha知识查询. 不需要 API 密钥 .