跳到主内容
智客 ZICQ

技能库 智客分类:Agent 工作流 Browser Use

浏览器使用

AI代理自动浏览器自动化. 两个工具:代理浏览器(CLI Playwright用于分步控制)和浏览器使用(Python自主代理在页面上决定该做什么). 导航,点击,填表,刮取数据,管理会话,并运行复杂的多步浏览器任务.

193 安装量

官方网址:skills.sh

技能介绍

先看中文介绍;官方 description 原文单独保留,不改写 SKILL.md。

做什么

AI代理自动浏览器自动化. 两个工具:代理浏览器(CLI Playwright用于分步控制)和浏览器使用(Python自主代理在页面上决定该做什么). 导航,点击,填表,刮取数据,管理会话,并运行复杂的多步浏览器任务.

何时用

官方 description 未单独写出 Use when。按规范,代理会在用户任务与这段 description 的关键词匹配时激活本技能。

代理如何加载

按 Agent Skills 渐进披露:启动时只加载 name 与 description(约 100 token);任务匹配后才读入整份 SKILL.md 正文;scripts/、references/、assets/ 仅在需要时再读。 本文件正文结构:Browser Use — Autonomous Browser Automation、Quick Start、agent-browser (recommended for most tasks)、Navigate and inspect、Interact using refs、Extract data。

文件分析

文件分析:这是一份仅含 SKILL.md 的指令型技能,代理激活后整份正文进入上下文。

官方 description(原文)

Autonomous browser automation for AI agents. Two tools: agent-browser (CLI Playwright for step-by-step control) and browser-use (Python autonomous agent that decides what to do on pages). Navigate, click, fill forms, scrape data, manage sessions, and run complex multi-step browser tasks.

Browser Use — Autonomous Browser AutomationQuick Startagent-browser (recommended for most tasks)Navigate and inspectInteract using refsExtract dataDonebrowser-use (autonomous agent)Run a full autonomous browsing taskagent-browser — Full ReferenceNavigationSnapshot (page analysis)

· allowed-tools:Bash(agent-browser:*,browser-use-agent:*,xvfb-run:*)

来源分类:skills.sh agent-skill

SKILL.md 与 Agent 调用

官方规范 ↗
name
Browser Use
description
Autonomous browser automation for AI agents. Two tools: agent-browser (CLI Playwright for step-by-step control) and browser-use (Python autonomous agent that decides what to do on pages). Navigate, click, fill forms, scrape data, manage sessions, and run complex multi-step browser tasks.
allowed-tools
Bash(agent-browser:*,browser-use-agent:*,xvfb-run:*)实验字段,支持情况取决于客户端;字段声明本身不会授予工具权限。
  1. 发现技能客户端向 Agent 提供名称与描述目录。
  2. 匹配与调用用户指定或任务匹配后,载入 SKILL.md 指令。
  3. 按需加载按步骤读取参考文档、使用脚本与素材。

具体调用语法与可用工具以目标 Agent 客户端为准。 查看调用机制说明 ↗

安装这个技能

Skills CLI ↗

先选择目标 Agent 和安装范围,保留技能包的附属文件,安装后检查客户端能否发现该技能。

交给 Agent 安装

复制安装指令给支持 Agent Skills 的代理,确认其中的目标目录与客户端匹配。

把 Agent Skill「Browser Use」安装到我的项目:SKILL.md 原文与官方 description 见 https://zicq.com/zh/skills/skl-8123db009875faf7-%E6%B5%8F%E8%A7%88%E5%99%A8%E4%BD%BF%E7%94%A8.html
请存为 .cursor/skills/browser-use/SKILL.md 或 .claude/skills/browser-use/SKILL.md,frontmatter 的 name 与 description 保持原样,不要改写。

GitHub 完整包 ↗

终端安装 · Skills CLI

需要 Node.js 与 npx。先查看仓库技能列表,确认实际名称。

npx skills add 'https://github.com/quentintou/openclaw-skill-browser-use' --list

CLI 会交互选择目标 Agent,默认安装到项目;用户级安装使用 -g。先通过查看命令核对仓库内容,再用 npx skills list 检查已安装技能。

阅读排版
--- name: Browser Use description: > Autonomous browser automation for AI agents. Two tools: agent-browser (CLI Playwright for step-by-step control) and browser-use (Python autonomous agent that decides what to do on pages). Navigate, click, fill forms, scrape data, manage sessions, and run complex multi-step browser tasks. read_when: - Automating web interactions beyond simple fetch - Filling forms or completing multi-step web flows - Scraping structured data from dynamic pages - Running an autonomous browsing agent for complex tasks - Testing or interacting with authenticated web apps - Taking screenshots or recording browser sessions metadata: clawdbot: emoji: "🌐" requires: bins: ["node", "npm", "python3"] system: ["chromium", "xvfb"] allowed-tools: Bash(agent-browser:*,browser-use-agent:*,xvfb-run:*) --- # Browser Use — Autonomous Browser Automation Two complementary tools for browser automation: | Tool | Best for | How it works | |------|----------|-------------| | **agent-browser** | Step-by-step control, scraping, form filling | CLI commands, you drive each action | | **browser-use** | Complex autonomous tasks | Python agent that decides actions itself | ## Quick Start ### agent-browser (recommended for most tasks) ```bash # Navigate and inspect agent-browser open "https://example.com" agent-browser snapshot -i # Get interactive elements with @refs # Interact using refs agent-browser click @e3 # Click element agent-browser fill @e2 "text" # Fill input (clears first) agent-browser press Enter # Press key # Extract data agent-browser get text @e1 # Get element text agent-browser get attr @e1 href # Get attribute agent-browser screenshot /tmp/p.png # Screenshot # Done agent-browser close ``` ### browser-use (autonomous agent) ```bash # Run a full autonomous browsing task browser-use-agent "Find the pricing for Notion and compare plans" ``` The agent will navigate, click, read pages, and return a structured result. ## agent-browser — Full Reference ### Navigation ```bash agent-browser open # Navigate to URL agent-browser back # Go back agent-browser forward # Go forward agent-browser reload # Reload page agent-browser close # Close browser ``` ### Snapshot (page analysis) ```bash agent-browser snapshot # Full accessibility tree agent-browser snapshot -i # Interactive elements only (recommended) agent-browser snapshot -c # Compact output agent-browser snapshot -d 3 # Limit depth to 3 agent-browser snapshot -s "#main" # Scope to CSS selector agent-browser snapshot -i --json # JSON output for parsing ``` ### Interactions (use @refs from snapshot) ```bash agent-browser click @e1 # Click agent-browser dblclick @e1 # Double-click agent-browser fill @e2 "text" # Clear and type (use this for inputs) agent-browser type @e2 "text" # Type without clearing agent-browser press Enter # Press key agent-browser press Control+a # Key combination agent-browser hover @e1 # Hover agent-browser check @e1 # Check checkbox agent-browser uncheck @e1 # Uncheck checkbox agent-browser select @e1 "value" # Select dropdown option agent-browser scroll down 500 # Scroll page agent-browser scrollintoview @e1 # Scroll element into view agent-browser drag @e1 @e2 # Drag and drop agent-browser upload @e1 file.pdf # Upload files ``` ### Extract Data ```bash agent-browser get text @e1 # Get element text agent-browser get html @e1 # Get innerHTML agent-browser get value @e1 # Get input value agent-browser get attr @e1 href # Get attribute agent-browser get title # Page title agent-browser get url # Current URL agent-browser get count ".item" # Count matching elements ``` ### Wait ```bash agent-browser wait @e1 # Wait for element agent-browser wait 2000 # Wait milliseconds agent-browser wait --text "Done" # Wait for text to appear agent-browser wait --url "/dash" # Wait for URL pattern agent-browser wait --load networkidle # Wait for network idle ``` ### Screenshots, PDF & Recording ```bash agent-browser screenshot path.png # Save screenshot agent-browser screenshot --full # Full page screenshot agent-browser pdf output.pdf # Save as PDF agent-browser record start ./demo.webm # Start recording agent-browser record stop # Stop and save ``` ### Sessions (parallel browsers) ```bash agent-browser --session s1 open "https://site1.com" agent-browser --session s2 open "https://site2.com" agent-browser session list ``` ### State (persist auth/cookies) ```bash agent-browser state save auth.json # Save session (cookies, storage) agent-browser state load auth.json # Restore session ``` ### Cookies & Storage ```bash agent-browser cookies # Get all cookies agent-browser cookies set name value # Set cookie agent-browser cookies clear # Clear cookies agent-browser storage local # Get all localStorage agent-browser storage local set k v # Set value ``` ### Tabs & Frames ```bash agent-browser tab # List tabs agent-browser tab new [url] # New tab agent-browser tab 2 # Switch to tab agent-browser frame "#iframe" # Switch to iframe agent-browser frame main # Back to main frame ``` ### Browser Settings ```bash agent-browser set viewport 1920 1080 agent-browser set device "iPhone 14" agent-browser set geo 37.7749 -122.4194 agent-browser set offline on agent-browser set media dark ``` ### JavaScript ```bash agent-browser eval "document.title" # Run JS in page context ``` ## browser-use — Autonomous Agent For complex tasks where you want the agent to figure out the browsing steps: ```bash browser-use-agent "Your task description here" ``` ### Custom Script (advanced) ```python # Run via: /opt/browser-use/bin/python3 script.py import asyncio, os from browser_use import Agent, Browser from langchain_anthropic import ChatAnthropic async def run(): browser = Browser() llm = ChatAnthropic( model='claude-sonnet-4-20250514', api_key=os.environ['ANTHROPIC_API_KEY'] ) agent = Agent( task="Compare pricing on 3 competitor sites", llm=llm, browser=browser, ) result = await agent.run(max_steps=15) await browser.close() return result asyncio.run(run()) ``` You can swap the LLM for any langchain-compatible model (OpenAI, Anthropic, etc). ## Standard Workflow ```bash # 1. Open page agent-browser open "https://example.com" # 2. Snapshot to see what's on the page agent-browser snapshot -i # 3. Interact with elements using @refs from snapshot agent-browser fill @e1 "search query" agent-browser click @e2 # 4. Wait for new page to load agent-browser wait --load networkidle # 5. Re-snapshot (refs change after navigation!) agent-browser snapshot -i # 6. Extract what you need agent-browser get text @e5 # 7. Close when done agent-browser close ``` ## Important Rules 1. **Always `snapshot -i` after navigation** — refs change on every page load 2. **Use `fill` not `type`** for inputs — fill clears existing text first 3. **Wait after clicks that trigger navigation** — `wait --load networkidle` 4. **Close the browser when done** — `agent-browser close` 5. **Google/Bing block headless browsers** (CAPTCHA) — use DuckDuckGo or `web_search` instead 6. **Save auth state** for sites requiring login — `state save/load` 7. **Use `--json`** when you need machine-parseable output 8. **Use sessions** for parallel browsing — `--session ` ## Troubleshooting - **Element not found**: Re-run `snapshot -i` to get current refs - **Page not loaded**: Add `wait --load networkidle` after navigation - **CAPTCHA on search engines**: Use DuckDuckGo or the `web_search` tool instead - **Auth expired**: Re-login and `state save` again - **Display errors**: The install script sets up Xvfb for headless rendering

相关技能

Agent 工作流

技能创建者Skill Creator

创造有效技能指南。 当用户想创造出新的技能(或更新现有的技能),以专业知识,工作流程,或工具集成来扩展克洛德的能力时,应该使用这种技能.

Agent 工作流

克劳德胡布Clawdhub

使用ClawdHub CLI搜索,安装,更新并发布从taladhub.com的代理技能. 需要获取苍蝇上的新技能时使用,将安装的技能同步到最新版本或特定版本,或者发布 npm-instainddhub CLI 的新/更新的技能文件夹.

Agent 工作流

团队指挥Agent Team Orchestration

管弦乐团多代理团队,任务设定周期,交接协议,审查工作流程. 使用时间: (1)建立2+特派员队伍,具有不同专业,(2)确定任务路线和生命周期(收录框_ spec_建设_审查_完成),(3)在特派员之间制定交接协议,(4)建立审查和质量关口,(5)管理特派员之间的交流和文物共享.

Agent 工作流

超级力量Superpowers

Spec-first,TDD,子代理驱动的软件开发工作流程. 当:(1)构建任何新功能或应用——触发脑暴_计划_子代理执行回路,(2)调试出一个bug或测试失败——触发系统性的根起过程,(3)用户说"让我们构建","帮助我计划","我想添加X",或"这个被打破",(4)完成一个功…