跳到主内容
智客 ZICQ

技能库 智客分类:Agent 工作流 firecrawl-scrape

Firecrawl Scrape

读取已知的网页或执行已发现的工作流程或数据提供者的能力。 一旦选择了URL或工具,则用于页面内容或结构化结果.

82719 安装量

官方网址:skills.sh

技能介绍

先看中文介绍;官方 description 原文单独保留,不改写 SKILL.md。

做什么

读取已知的网页或执行已发现的工作流程或数据提供者的能力。 一旦选择了URL或工具,则用于页面内容或结构化结果.

何时用

官方 description 未单独写出 Use when。按规范,代理会在用户任务与这段 description 的关键词匹配时激活本技能。

代理如何加载

按 Agent Skills 渐进披露:启动时只加载 name 与 description(约 100 token);任务匹配后才读入整份 SKILL.md 正文;scripts/、references/、assets/ 仅在需要时再读。 本文件正文结构:firecrawl scrape、Quick start、Basic markdown extraction、Main content only, no nav/footer、Wait for JS to render, then scrape、Multiple URLs (markdown only; each saved to .firecrawl/; -o is ignored)。

文件分析

文件分析:除 SKILL.md 外,正文引用了 references/large-results.md,属于带资源的技能包,这些文件按需再读。

官方 description(原文)

Read a known webpage or execute a discovered workflow or data-provider capability. Use for page content or structured results once the URL or tool is selected.

firecrawl scrapeQuick startBasic markdown extractionMain content only, no nav/footerWait for JS to render, then scrapeMultiple URLs (markdown only; each saved to .firecrawl/; -o is ignored)Get markdown and links togetherAsk a question about the pageFind tools, inspect inputs, and get helpWeb + domain matching + semantic toolsSemantic tools onlyCategories → providers → tools → contract

来源分类:skills.sh agent-skill

SKILL.md 与 Agent 调用

官方规范 ↗
name
firecrawl-scrape
description
Read a known webpage or execute a discovered workflow or data-provider capability. Use for page content or structured results once the URL or tool is selected.
  1. 发现技能客户端向 Agent 提供名称与描述目录。
  2. 匹配与调用用户指定或任务匹配后,载入 SKILL.md 指令。
  3. 按需加载按步骤读取参考文档、使用脚本与素材。
指令中引用的文件 · 1
  • references/large-results.md

以下路径提取自原文;文件是否齐全请以来源仓库中的完整目录为准。

具体调用语法与可用工具以目标 Agent 客户端为准。 查看调用机制说明 ↗

安装这个技能

Skills CLI ↗

先选择目标 Agent 和安装范围,保留技能包的附属文件,安装后检查客户端能否发现该技能。

该技能引用了附属文件,请从来源获取完整目录;仅复制 SKILL.md 可能缺少依赖。

交给 Agent 安装

复制安装指令给支持 Agent Skills 的代理,确认其中的目标目录与客户端匹配。

把 Agent Skill「firecrawl-scrape」安装到我的项目:SKILL.md 原文与官方 description 见 https://zicq.com/zh/skills/skl-7fc08ec64a7b5e31-Firecrawl-Scrape.html
请存为 .cursor/skills/firecrawl-scrape/SKILL.md 或 .claude/skills/firecrawl-scrape/SKILL.md,frontmatter 的 name 与 description 保持原样,不要改写。
该技能还带 scripts/、references/、assets/ 等文件,请从 https://github.com/firecrawl/cli 取完整目录,不要只建一个 SKILL.md。

GitHub 完整包 ↗

终端安装 · Skills CLI

需要 Node.js 与 npx。先查看仓库技能列表,确认实际名称。

npx skills add 'https://github.com/firecrawl/cli' --list

npx skills add 'https://github.com/firecrawl/cli' --skill 'firecrawl-scrape'

CLI 会交互选择目标 Agent,默认安装到项目;用户级安装使用 -g。先通过查看命令核对仓库内容,再用 npx skills list 检查已安装技能。

阅读排版
--- name: firecrawl-scrape description: Read a known webpage or execute a discovered workflow or data-provider capability. Use for page content or structured results once the URL or tool is selected. allowed-tools: - Bash(firecrawl *) - Bash(npx firecrawl-cli *) --- # firecrawl scrape Read a URL for page content, or execute a selected provider tool for structured data. Discover tools with `search` and inspect their inputs with `list` before execution. Multiple URLs can be scraped concurrently. For structured datasets, first check for a suitable workflow or data provider using the [search skill](../firecrawl-search/SKILL.md). Read a known page directly; reuse a selected contract instead of repeating discovery. ## Quick start ```bash # Basic markdown extraction firecrawl scrape "" -o .firecrawl/page.md # Main content only, no nav/footer firecrawl scrape "" --only-main-content -o .firecrawl/page.md # Wait for JS to render, then scrape firecrawl scrape "" --wait-for 3000 -o .firecrawl/page.md # Multiple URLs (markdown only; each saved to .firecrawl/; -o is ignored) firecrawl scrape https://example.com https://example.com/blog https://example.com/docs # Get markdown and links together firecrawl scrape "" --format markdown,links -o .firecrawl/page.json # Ask a question about the page firecrawl scrape "https://example.com/pricing" --query "What is the enterprise plan price?" ``` Run `firecrawl scrape --help` for the full option list. **Done when:** the page content or provider result has been checked for errors and inspected in bounded sections to answer the request. Preserve source links and disclose partial results. ## Find tools, inspect inputs, and get help Use the CLI help to check supported options rather than guessing: ```bash firecrawl search --help firecrawl list --help firecrawl scrape --help ``` Domain discovery with `--domain-tools` returns tool summaries by default. Add `--tool-detail full` for contracts upfront, or inspect one selected tool with `list` as shown below. Use `--tool-detail compact` for only provider, capability and description; inspect by those two IDs with `list`. Summary remains the default. Prefer full when several related contracts will be needed immediately. For structured data, search for the task, inspect a matching tool's contract, then execute with the exact input fields it declares: ```bash # Web + domain matching + semantic tools firecrawl search '' # Semantic tools only firecrawl search alexandria '' # Categories → providers → tools → contract firecrawl list firecrawl list --category firecrawl list firecrawl list --pretty # Execute a tool firecrawl scrape / --options '' ``` Normal search includes web results and tool matches; `search alexandria` searches tools only. `list --pretty` shows the selected contract; use `--json` for machine-readable output. To browse progressively, use `list`, then `list --category`, then `list `. Search and list do not execute the selected provider tool. Read only the contracts needed for the task; use returned identifiers rather than guessing them. Read the expanded contract before building inputs or parsing results: - `required: true` requires that input; each `requiresOneOf` group requires at least one member, not all of them. - Selected-contract inspection already requests examples. Read the singular `example.request` and `example.response` when present; an empty request can be valid for tools with optional inputs. - `response.key` identifies the records field inside `data.alexandria[i].data`; an empty key means that data object itself. Do not assume every provider returns `records`. - Provider pagination differs from catalogue `next`: use the contract's continuation input and the returned page/cursor, preserve filters, and stop at its exhaustion signal. `paginated: true` alone does not specify that mapping. ## Execution and large results URL scraping does not execute provider tools automatically. Use exact discovered input fields and resolve record IDs with lookup tools rather than inventing them. Check each `data.alexandria[]` result for errors, not just the outer success flag. If the client reports an output/context limit, the upstream request may have succeeded. Preserve the request or scrape ID and recover the retained result before repeating the provider call. For large datasets and PDFs, save output with `--json -o` when a local filesystem is available and inspect bounded sections with `jq` or other file tools. Keep stderr separate from JSON stdout; do not merge streams with `2>&1` when piping to a JSON parser. Where remote processing is preferable, use `firecrawl scrape firecrawl/bash` to select from a retained result. Read [large-result recovery](references/large-results.md) for IDs, command examples, expiry, and errors. This is explicit recovery, not automatic overflow detection. ## PDFs and page budgets PDFs cost 1 credit per parsed page. Use `--max-pages` (an integer from 1 to 10000) to limit PDF parsing, especially for large or unknown documents: ```bash firecrawl scrape "https://example.com/report.pdf" --max-pages 5 --json -o .firecrawl/report.json ``` The cap applies to each PDF, not the whole command or total credits. Extra formats and options can add charges. The CLI does not quote page counts or costs before execution. Use JSON output to inspect the returned `metadata.numPages` (parsed), `metadata.totalPages` (document total), and `metadata.creditsUsed` when present; a smaller parsed count means the result is partial. ## Tips - **Prefer plain scrape over `--query`.** Scrape to a file, then use `grep`, `head`, or read the markdown directly — you can search and reason over the full content yourself. Use `--query` only when you want a single targeted answer without saving the page (costs 5 extra credits). - **Scrape handles static pages and JS-rendered SPAs.** Escalate to `interact` when the page needs interaction (clicks, form fills, pagination) or scrape misses content. - Multiple URLs are scraped concurrently — check `firecrawl --status` for your concurrency limit. This mode saves markdown only and ignores `-o`; other requested formats are dropped. If markdown wasn't requested, the whole JSON response is written into the `.md` file. - Single format outputs raw content. Multiple formats (e.g., `--format markdown,links`) output JSON. - Always quote URLs — shell interprets `?` and `&` as special characters. - Naming convention: `.firecrawl/{site}-{path}.md` ## See also - [firecrawl-search](../firecrawl-search/SKILL.md) — find pages when you don't have a URL - [firecrawl-interact](../firecrawl-interact/SKILL.md) — when scrape can't get the content, use `interact` to click, fill forms, etc. - [firecrawl-download](../firecrawl-download/SKILL.md) — bulk download an entire site to local files - [firecrawl-build-scrape](https://github.com/firecrawl/skills/tree/main/skills/build/firecrawl-build-scrape) — building scrape into an app instead of running it here

相关技能

Agent 工作流

Skill Creator

创造有效技能指南。 当用户想创造出新的技能(或更新现有的技能),以专业知识,工作流程,或工具集成来扩展克洛德的能力时,应该使用这种技能.

Agent 工作流

Clawdhub

使用ClawdHub CLI搜索,安装,更新并发布从taladhub.com的代理技能. 需要获取苍蝇上的新技能时使用,将安装的技能同步到最新版本或特定版本,或者发布 npm-instainddhub CLI 的新/更新的技能文件夹.

Agent 工作流

Agent Team Orchestration

管弦乐团多代理团队,任务设定周期,交接协议,审查工作流程. 使用时间: (1)建立2+特派员队伍,具有不同专业,(2)确定任务路线和生命周期(收录框_ spec_建设_审查_完成),(3)在特派员之间制定交接协议,(4)建立审查和质量关口,(5)管理特派员之间的交流和文物共享.

Agent 工作流

Superpowers

Spec-first,TDD,子代理驱动的软件开发工作流程. 当:(1)构建任何新功能或应用——触发脑暴_计划_子代理执行回路,(2)调试出一个bug或测试失败——触发系统性的根起过程,(3)用户说"让我们构建","帮助我计划","我想添加X",或"这个被打破",(4)完成一个功…