跳到主内容
智客 ZICQ

技能库 智客分类:数据与分析 apify-ultimate-scraper

Apify Ultimate Scraper

任何平台都使用通用AI驱动的网络刮刀. 来自Instagram,Facebook,TikTok,YouTube,LinkedIn,X/Twitter,Google地图,Google搜索,Google趋势,Reddit,Airbnb,Yelp等15+平台的搜索数据. 用于铅生成,品牌监测,竞争者分析,影响者发现,趋势研究,内容分析,受众分析,审查分析,SIO智能,招聘,或任何数据提取任务.

15516 安装量

官方网址:skills.sh

技能介绍

先看中文介绍;官方 description 原文单独保留,不改写 SKILL.md。

做什么

任何平台都使用通用AI驱动的网络刮刀. 来自Instagram,Facebook,TikTok,YouTube,LinkedIn,X/Twitter,Google地图,Google搜索,Google趋势,Reddit,Airbnb,Yelp等15+平台的搜索数据. 用于铅生成,品牌监测,竞争者分析,影响者发现,趋势研究,内容分析,受众分析,审查分析,SIO智能,招聘,或任何数据提取任务.

何时用

官方 description 未单独写出 Use when。按规范,代理会在用户任务与这段 description 的关键词匹配时激活本技能。

代理如何加载

按 Agent Skills 渐进披露:启动时只加载 name 与 description(约 100 token);任务匹配后才读入整份 SKILL.md 正文;scripts/、references/、assets/ 仅在需要时再读。 本文件正文结构:Universal web scraper、Prerequisites、Authentication、Workflow、Step 1: Understand goal and select Actor、Step 2: Fetch Actor schema and check gotchas。 其中含规范建议的小节:分步指令。

文件分析

文件分析:除 SKILL.md 外,正文引用了 references/actor-index.md、references/workflows/lead-generation.md、references/workflows/competitive-intel.md、references/workflows/influencer-vetting.md、references/workflows/brand-monitoring.md、references/workflows/review-analysis.md,属于带资源的技能包,这些文件按需再读。

官方 description(原文)

Universal AI-powered web scraper for any platform. Scrape data from Instagram, Facebook, TikTok, YouTube, LinkedIn, X/Twitter, Google Maps, Google Search, Google Trends, Reddit, Airbnb, Yelp, and 15+ more platforms. Use for lead generation, brand monitoring, competitor analysis, influencer discovery, trend research, content analytics, audience analysis, review analysis, SEO intelligence, recruitment, or any data extraction task.

Universal web scraperPrerequisitesAuthenticationWorkflowStep 1: Understand goal and select ActorStep 2: Fetch Actor schema and check gotchasStep 3: Configure and runStep 4: Deliver resultsTroubleshooting

来源分类:skills.sh agent-skill

SKILL.md 与 Agent 调用

官方规范 ↗
name
apify-ultimate-scraper
description
Universal AI-powered web scraper for any platform. Scrape data from Instagram, Facebook, TikTok, YouTube, LinkedIn, X/Twitter, Google Maps, Google Search, Google Trends, Reddit, Airbnb, Yelp, and 15+ more platforms. Use for lead generation, brand monitoring, competitor analysis, influencer discovery, trend research, content analytics, audience analysis, review analysis, SEO intelligence, recruitment, or any data extraction task.
  1. 发现技能客户端向 Agent 提供名称与描述目录。
  2. 匹配与调用用户指定或任务匹配后,载入 SKILL.md 指令。
  3. 按需加载按步骤读取参考文档、使用脚本与素材。
指令中引用的文件 · 12
  • references/actor-index.md
  • references/workflows/lead-generation.md
  • references/workflows/competitive-intel.md
  • references/workflows/influencer-vetting.md
  • references/workflows/brand-monitoring.md
  • references/workflows/review-analysis.md
  • references/workflows/content-and-seo.md
  • references/workflows/social-media-analytics.md
  • references/workflows/trend-research.md
  • references/workflows/job-market-and-recruitment.md
  • references/workflows/real-estate-and-hospitality.md
  • references/workflows/ecommerce-price-monitoring.md

以下路径提取自原文;文件是否齐全请以来源仓库中的完整目录为准。

具体调用语法与可用工具以目标 Agent 客户端为准。 查看调用机制说明 ↗

安装这个技能

Skills CLI ↗

先选择目标 Agent 和安装范围,保留技能包的附属文件,安装后检查客户端能否发现该技能。

该技能引用了附属文件,请从来源获取完整目录;仅复制 SKILL.md 可能缺少依赖。

交给 Agent 安装

复制安装指令给支持 Agent Skills 的代理,确认其中的目标目录与客户端匹配。

把 Agent Skill「apify-ultimate-scraper」安装到我的项目:SKILL.md 原文与官方 description 见 https://zicq.com/zh/skills/skl-93380fa233ceffce-Apify-Ultimate-Scraper.html
请存为 .cursor/skills/apify-ultimate-scraper/SKILL.md 或 .claude/skills/apify-ultimate-scraper/SKILL.md,frontmatter 的 name 与 description 保持原样,不要改写。
该技能还带 scripts/、references/、assets/ 等文件,请从 https://github.com/apify/agent-skills 取完整目录,不要只建一个 SKILL.md。

GitHub 完整包 ↗

终端安装 · Skills CLI

需要 Node.js 与 npx。先查看仓库技能列表,确认实际名称。

npx skills add 'https://github.com/apify/agent-skills' --list

npx skills add 'https://github.com/apify/agent-skills' --skill 'apify-ultimate-scraper'

CLI 会交互选择目标 Agent,默认安装到项目;用户级安装使用 -g。先通过查看命令核对仓库内容,再用 npx skills list 检查已安装技能。

阅读排版

name: apify-ultimate-scraper description: Universal AI-powered web scraper for any platform. Scrape data from Instagram, Facebook, TikTok, YouTube, LinkedIn, X/Twitter, Google Maps, Google Search, Google Trends, Reddit, Airbnb, Yelp, and 15+ more platforms. Use for lead generation, brand monitoring, competitor analysis, influencer discovery, trend research, content analytics, audience analysis, review analysis, SEO intelligence, recruitment, or any data extraction task.

Universal web scraper

AI-driven data extraction from ~100 Actors across 15+ platforms via the Apify CLI.

Rules for every apify command:

  1. Pass --json for machine-readable output (stable across CLI versions).
  2. Pass --user-agent apify-agent-skills/apify-ultimate-scraper for telemetry attribution.
  3. Redirect stderr with 2>/dev/null (stderr contains progress messages that break JSON parsers).

Prerequisites

  • Apify CLI v1.5.0+ (npm install -g apify-cli)
  • Authenticated session (see below)

Authentication

If a CLI command fails with an auth error, authenticate using one of these methods:

  1. OAuth (interactive): apify login (opens browser)
  2. Environment variable: export APIFY_TOKEN=your_token_here
  3. From .env file: source .env (if the file contains APIFY_TOKEN=...)

Generate token: https://console.apify.com/settings/integrations

Workflow

Step 1: Understand goal and select Actor

Identify the target platform and use case. Read references/actor-index.md to find the right Actor.

If the task involves a multi-step pipeline, also read the matching workflow guide:

| Task involves... | Read | |-----------------|------| | leads, contacts, emails, B2B | references/workflows/lead-generation.md | | competitor, ads, pricing | references/workflows/competitive-intel.md | | influencer, creator | references/workflows/influencer-vetting.md | | brand, mentions, sentiment | references/workflows/brand-monitoring.md | | reviews, ratings, reputation | references/workflows/review-analysis.md | | SEO, SERP, crawl, content, RAG | references/workflows/content-and-seo.md | | analytics, engagement, performance | references/workflows/social-media-analytics.md | | trends, keywords, hashtags | references/workflows/trend-research.md | | jobs, recruiting, candidates | references/workflows/job-market-and-recruitment.md | | real estate, listings, hotels | references/workflows/real-estate-and-hospitality.md | | price monitoring, e-commerce, products | references/workflows/ecommerce-price-monitoring.md | | contact enrichment, email extraction | references/workflows/contact-enrichment.md | | knowledge base, RAG, LLM data feed | references/workflows/knowledge-base-and-rag.md | | company research, due diligence | references/workflows/company-research.md |

If no Actor matches in the index, search dynamically:

apify actors search "KEYWORDS" --user-agent apify-agent-skills/apify-ultimate-scraper --json --limit 10 2>/dev/null

From results: items[].username/items[].name (Actor ID), items[].title, items[].stats.totalUsers30Days, items[].currentPricingInfo.pricingModel.

Step 2: Fetch Actor schema and check gotchas

Fetch the input schema dynamically:

apify actors info "ACTOR_ID" --user-agent apify-agent-skills/apify-ultimate-scraper --input --json 2>/dev/null

Also read references/gotchas.md to check for common pitfalls for the selected Actor.

For Actor documentation: apify actors info "ACTOR_ID" --user-agent apify-agent-skills/apify-ultimate-scraper --readme

Step 3: Configure and run

Skip user preferences for simple lookups (e.g., "Nike's follower count"). Go straight to running with quick answer mode.

For larger tasks, confirm output format (quick answer / CSV / JSON) and result count.

Standard run (blocking):

apify actors call "ACTOR_ID" --input-file input.json --user-agent apify-agent-skills/apify-ultimate-scraper --json 2>/dev/null

Prefer --input-file input.json for large or complex inputs. For tiny inputs, inline JSON is acceptable with shell quoting: --input '{"maxItems":10}'.

From output: .id (run ID), .status, .defaultDatasetId, .stats.durationMillis

Fetch results:

apify datasets get-items DATASET_ID --user-agent apify-agent-skills/apify-ultimate-scraper --format json

For CSV: apify datasets get-items DATASET_ID --user-agent apify-agent-skills/apify-ultimate-scraper --format csv

Quick answer mode: Fetch results as JSON, pick top 5, present formatted in chat.

Save to file: Fetch results, use Write tool to save as YYYY-MM-DD_descriptive-name.csv or .json.

Large/long-running scrapes:

apify actors start "ACTOR_ID" --input-file input.json --user-agent apify-agent-skills/apify-ultimate-scraper --json 2>/dev/null

Poll: apify runs info RUN_ID --user-agent apify-agent-skills/apify-ultimate-scraper --json 2>/dev/null (check .status for SUCCEEDED).

Step 4: Deliver results

Report: result count, file location (if saved), key data fields, and links:

  • Dataset: https://console.apify.com/storage/datasets/DATASET_ID
  • Run: https://console.apify.com/actors/runs/RUN_ID

For multi-step workflows: suggest the next pipeline step from the workflow guide.

Troubleshooting

Common errors and pitfalls are documented in references/gotchas.md. Read it before running PPE (pay-per-event) Actors.

相关技能

数据与分析

Marketing Skills

TL; DR:23个营销剧本(CRO,SEO,拷贝,分析,实验,定价,发布,广告,社交). 用于快速获取清单+副本/粘贴可交付品.

数据与分析

Crypto & Stock Market Data (Node.js)

免费等级不需要 API 关键值 。 专业级别的密码货币和股票市场数据集成,用于实时价格、公司概况和全球分析。 由Node.js提供动力,外部依赖性为零.

数据与分析

Azure Kusto

Azure Data Explorer(Kusto/ADX)中使用KQL进行日志分析,遥测和时间序列分析的查询并分析数据. When: KQL 查询, Kusto 数据库查询, Azure Data Explorer, ADX 集群, 日志分析, 时间序列数据, IoT 遥测, …

数据与分析

Azure Storage

Azure存储服务包括Blob存储,文件共享,等式存储,表存储,和数据湖. 解答关于存储访问级别(热,凉,冷,存档),何时使用每个级别,以及级别比较的问题. 提供对象存储,SMB文件共享,async消息,NoSQL密钥-值,和大数据分析. 包括生命周期管理。 USE FOR: b…