跳到主内容
智客 ZICQ

技能库 智客分类:运维与云 tao-run-on-brev

Tao Run On Brev

在NVIDIA Brev GPU实例上运行一个TAO培训/评价/参考容器。 实例提供(创建/搜索/停止/删除/登录)被授权给正式的brev-cli代理技能或Brev MCP服务器;这一技能仅涵盖TAO特定部分——通过四动词插头合同在 " brev exec " 上运行容器。 触发短语包括"运行在Brev上","Brev GPU实例","TAO on Brev","向Brev提交工作"等.

1581 安装量

官方网址:skills.sh

技能介绍

先看中文介绍;官方 description 原文单独保留,不改写 SKILL.md。

做什么

在NVIDIA Brev GPU实例上运行一个TAO培训/评价/参考容器。 实例提供(创建/搜索/停止/删除/登录)被授权给正式的brev-cli代理技能或Brev MCP服务器;这一技能仅涵盖TAO特定部分——通过四动词插头合同在 " brev exec " 上运行容器。 触发短语包括"运行在Brev上","Brev GPU实例","TAO on Brev","向Brev提交工作"等.

何时用

官方 description 未单独写出 Use when。按规范,代理会在用户任务与这段 description 的关键词匹配时激活本技能。

代理如何加载

按 Agent Skills 渐进披露:启动时只加载 name 与 description(约 100 token);任务匹配后才读入整份 SKILL.md 正文;scripts/、references/、assets/ 仅在需要时再读。 本文件正文结构:Brev — TAO execution glue、Provisioning: use the official Brev skill or MCP、installs to ~/.claude/skills/brev-cli/ , ~/.codex/skills/brev-cli/ , ~/.agents/skills/brev-cli/、Storage、Execution — the four verbs (a compound over Docker)、`brev exec` argument form。

文件分析

文件分析:除 SKILL.md 外,正文引用了 scripts/install-agent-skill.sh、scripts/tao_job_record.py,属于带资源的技能包,这些文件按需再读。

官方 description(原文)

Run a TAO training/evaluation/inference container on an NVIDIA Brev GPU instance. Instance provisioning (create/search/stop/delete/login) is delegated to the official brev-cli agent skill or the Brev MCP server; this skill covers only the TAO-specific part — running the container over `brev exec` via the four-verb docker contract. Trigger phrases include "run on Brev", "Brev GPU instance", "TAO on Brev", "submit job to Brev".

Brev — TAO execution glueProvisioning: use the official Brev skill or MCPinstalls to ~/.claude/skills/brev-cli/ , ~/.codex/skills/brev-cli/ , ~/.agents/skills/brev-cli/StorageExecution — the four verbs (a compound over Docker)`brev exec` argument formNGC auth (one-time per instance) — value never on argv.Single-quoted locally so $NGC_KEY expands in the instance's shell; export itthere first (or pipe it in from the local shell, if the instance has no copy).Verify auth without reading ~/.docker/config.json. Failure before a successfullogin = not authenticated; failure after = the key's org lacks entitlement.Pull BEFORE the GPU run. `docker run` would pull implicitly, but the instance

兼容:Requires the brev CLI (https://github.com/brevdev/brev-cli) and an active brev login. Instance provisioning is handled by the official brev-cli agent skill or the Brev MCP server. · 许可:Apache-2.0 · allowed-tools:Read Bash

来源分类:skills.sh agent-skill

SKILL.md 与 Agent 调用

官方规范 ↗
name
tao-run-on-brev
description
Run a TAO training/evaluation/inference container on an NVIDIA Brev GPU instance. Instance provisioning (create/search/stop/delete/login) is delegated to the official brev-cli agent skill or the Brev MCP server; this skill covers only the TAO-specific part — running the container over `brev exec` via the four-verb docker contract. Trigger phrases include "run on Brev", "Brev GPU instance", "TAO on Brev", "submit job to Brev".
compatibility
Requires the brev CLI (https://github.com/brevdev/brev-cli) and an active brev login. Instance provisioning is handled by the official brev-cli agent skill or the Brev MCP server.
allowed-tools
Read Bash实验字段,支持情况取决于客户端;字段声明本身不会授予工具权限。
许可
Apache-2.0
  1. 发现技能客户端向 Agent 提供名称与描述目录。
  2. 匹配与调用用户指定或任务匹配后,载入 SKILL.md 指令。
  3. 按需加载按步骤读取参考文档、使用脚本与素材。
指令中引用的文件 · 2
  • scripts/install-agent-skill.sh
  • scripts/tao_job_record.py

以下路径提取自原文;文件是否齐全请以来源仓库中的完整目录为准。

具体调用语法与可用工具以目标 Agent 客户端为准。 查看调用机制说明 ↗

安装这个技能

Skills CLI ↗

先选择目标 Agent 和安装范围,保留技能包的附属文件,安装后检查客户端能否发现该技能。

该技能引用了附属文件,请从来源获取完整目录;仅复制 SKILL.md 可能缺少依赖。

交给 Agent 安装

复制安装指令给支持 Agent Skills 的代理,确认其中的目标目录与客户端匹配。

把 Agent Skill「tao-run-on-brev」安装到我的项目:SKILL.md 原文与官方 description 见 https://zicq.com/zh/skills/skl-abc9363b952babd2-Tao-Run-On-Brev.html
请存为 .cursor/skills/tao-run-on-brev/SKILL.md 或 .claude/skills/tao-run-on-brev/SKILL.md,frontmatter 的 name 与 description 保持原样,不要改写。
该技能还带 scripts/、references/、assets/ 等文件,请从 https://github.com/nvidia/skills 取完整目录,不要只建一个 SKILL.md。

GitHub 完整包 ↗

终端安装 · Skills CLI

需要 Node.js 与 npx。先查看仓库技能列表,确认实际名称。

npx skills add 'https://github.com/nvidia/skills' --list

npx skills add 'https://github.com/nvidia/skills' --skill 'tao-run-on-brev'

CLI 会交互选择目标 Agent,默认安装到项目;用户级安装使用 -g。先通过查看命令核对仓库内容,再用 npx skills list 检查已安装技能。

阅读排版
--- name: tao-run-on-brev description: Run a TAO training/evaluation/inference container on an NVIDIA Brev GPU instance. Instance provisioning (create/search/stop/delete/login) is delegated to the official brev-cli agent skill or the Brev MCP server; this skill covers only the TAO-specific part — running the container over `brev exec` via the four-verb docker contract. Trigger phrases include "run on Brev", "Brev GPU instance", "TAO on Brev", "submit job to Brev". license: Apache-2.0 compatibility: Requires the brev CLI (https://github.com/brevdev/brev-cli) and an active brev login. Instance provisioning is handled by the official brev-cli agent skill or the Brev MCP server. metadata: author: NVIDIA Corporation version: "0.2.0" allowed-tools: Read Bash tags: - gpu - compute - instance-based - brev --- # Brev — TAO execution glue > **Standalone install?** If this session was not initialized by the TAO skill bank plugin, run the `tao-setup` skill first (host preflight, credentials, cross-skill discovery). NVIDIA Brev provides on-demand GPU instances (pre-loaded with NVIDIA drivers, CUDA, Docker, and the NVIDIA Container Toolkit). Brev is **instance-based**: you provision an instance, run commands on it over `brev exec`, and delete it when done. This skill is deliberately thin. **Provisioning and managing instances — create, search by GPU/price, start/stop, delete, login — is owned by NVIDIA Brev's own agent skill, not duplicated here.** This skill covers only the TAO-specific part: running a TAO container on a reached instance through the **four-verb docker contract**, deferring the container-how to `tao-run-on-docker` over `brev exec`. ## Provisioning: use the official Brev skill or MCP NVIDIA Brev publishes an agent skill that manages instances in natural language ("create an A100 instance", "search for GPUs under $3/hr", "stop all my instances"). Install it once — it self-registers into your agent's skills dir and is discovered at runtime: ```bash curl -fsSL https://raw.githubusercontent.com/brevdev/brev-cli/main/scripts/install-agent-skill.sh | bash # installs to ~/.claude/skills/brev-cli/ , ~/.codex/skills/brev-cli/ , ~/.agents/skills/brev-cli/ ``` Or connect the **Brev MCP server** (`https://docs.nvidia.com/brev/_mcp/server`). Either one owns login/auth quirks, placement IDs, GPU search, and teardown flags. It does **not** cover container execution on the instance — that is this skill. **Preflight for this skill:** the `brev` CLI is on `PATH` and logged in (headless: `brev login --token "$BREV_API_TOKEN"` before any other call), and you can reach a target instance — poll with a **two-word** command until it succeeds before issuing real work (a fresh instance reports `RUNNING` before sshd is up): ```bash for i in $(seq 1 60); do brev exec "echo ok" 2>/dev/null | grep -qx ok && break; sleep 5; done brev exec "echo ok" 2>/dev/null | grep -qx ok || { echo "instance not exec-ready"; exit 1; } ``` The probe must be **two words, quoted as one argument**. A single-token probe (`brev exec -- true`) passes even when every real command is broken, because `brev exec [instance...] ` treats only the LAST positional as the command — so a lone `true` lands in the right slot by accident while `docker run ...` does not. See *`brev exec` argument form* below. Allow **≥ 600 s** for the first `brev exec` on a new instance (SSH bring-up + first container pull); a 60–120 s wrapper timeout truncates startup and looks like a spurious `exec failed`. ## Storage No shared NFS/Lustre — storage tier **B/C** via `tao-data-io`: stage inputs from S3 to the instance's local disk (or fetch in-container) and **upload results to S3 before deleting the instance**. Instance-local `~/` persists across stop/start but **not** across delete/create, so the results upload must precede teardown. ## Execution — the four verbs (a compound over Docker) Brev is a **compound consumer**: `submit` reaches an instance, then **defers the container-how to the four docker verbs** (`tao-run-on-docker`) run over `brev exec`. It is not a symmetric peer — teardown must additionally delete the instance to stop billing. `$BANK` = `${TAO_SKILL_BANK_PATH}`. - **submit** — reach an instance (provision/reuse via the official Brev skill or MCP; reuse an existing instance by its `instance_id`; wait for readiness, above). Lint the assembled command, open the record to mint `$JOB_ID` **before** launch, then run the docker `submit` verb *inside* the instance and mark RUNNING: ```bash redact_secrets.py lint <<<"$REMOTE_CMD" # no inline secrets; creds as -e VAR JOB_ID=$("$BANK/scripts/tao_job_record.py" open \ --platform brev --image "$IMG" \ --network-arch "$ARCH" --action "$ACTION" \ --storage-tier "$TIER" --results-root "$RESULTS_ROOT") brev exec "docker inspect '$JOB_ID' >/dev/null 2>&1 && { echo '$JOB_ID already submitted'; exit 0; }; docker run -d --name '$JOB_ID' --label 'tao-job=$JOB_ID' ..." "$BANK/scripts/tao_job_record.py" mark "$JOB_ID" --state RUNNING \ --backend-ref "/$JOB_ID" # instance is part of the ref: the # container is unreachable without it ``` - **status / logs** — `brev exec "docker inspect $JOB_ID"` / `brev exec "docker logs $JOB_ID"`, mapped to the vocab exactly as the docker verbs do. Recover `` from the record's `backend-ref`. - **cancel / teardown** — remove the container, then for an ephemeral instance delete it (stops billing), then mark the record. Never leave an ephemeral instance running: ```bash brev exec "docker rm -f $JOB_ID" brev delete # ephemeral instances only "$BANK/scripts/tao_job_record.py" mark "$JOB_ID" --state CANCELED --source agent ``` ### `brev exec` argument form The remote command is **one argument**. The CLI signature is `brev exec [instance...] `: every positional except the last is an instance name, and `--` only ends flag parsing — it does not group the words after it. So `brev exec -- docker inspect "$JOB_ID"` is read as instances ` docker inspect` plus command `"$JOB_ID"`, and fails with `could not look up instance "docker"` / `ssh: illegal option -- -` (exit 255) — an error that reads like an instance or SSH fault but is a syntax fault. Quote every remote command as a single string, exactly as `brev exec --help` shows. NGC auth once per instance — **never put `NGC_KEY` on argv** (it lands in the remote process table); pipe it to `--password-stdin`: ```bash IMG=nvcr.io/nvidia/tao/tao-toolkit:7.2.0-pyt # versions-key: images.tao_toolkit.pyt # NGC auth (one-time per instance) — value never on argv. # Single-quoted locally so $NGC_KEY expands in the instance's shell; export it # there first (or pipe it in from the local shell, if the instance has no copy). brev exec 'printf %s "$NGC_KEY" | docker login nvcr.io -u "$oauthtoken" --password-stdin' # Verify auth without reading ~/.docker/config.json. Failure before a successful # login = not authenticated; failure after = the key's org lacks entitlement. brev exec "docker manifest inspect $IMG >/dev/null && echo AUTH_OK || echo AUTH_FAIL" # Pull BEFORE the GPU run. `docker run` would pull implicitly, but the instance # bills from boot, so a multi-GB first-time TAO pull is billed GPU-idle time. # Pulling as its own step also separates a pull failure (auth/entitlement) from # a training failure in the logs. brev exec "docker image inspect $IMG >/dev/null 2>&1 || docker pull $IMG" # Run a TAO job (the docker `submit` verb, over brev exec) brev exec "docker inspect '$JOB_ID' >/dev/null 2>&1 && { echo '$JOB_ID already submitted'; exit 0; }; docker run -d --name '$JOB_ID' --label 'tao-job=$JOB_ID' --gpus all -v ~/data:/data -e NGC_KEY '$IMG' visual_changenet train -e /data/spec.yaml" ``` ## Multi-GPU and multi-node **Multi-node is not supported on Brev** — instance-based, no cross-instance coordination. Multi-GPU **on a single instance** is supported (up to 8× H100 / A100 / L40S); `torchrun --nproc-per-node=N` or PyTorch DDP work within the instance.

相关技能

运维与云

Docker Essentials

用于容器管理,图像操作,调试的基本道克命令和工作流程.

运维与云

Find Skills

从开放的代理技能生态系统中发现并安装技能. 使用时:(1)用户问"我如何做X",X可能拥有现有技能,(2)用户说"为X找到技能"或"是否为X有技能",(3)用户问"你能否做X",X是专门能力,(4)用户想扩展代理能力,(5)用户想搜索工具,模板,或工作流程,(6)用户提到他们希望…

运维与云

Azure Diagnostics

Azure上使用AppLens,AzureMonitor,资源健康,安全分型的调试Azure生产问题. 当:调试生产问题,故障解答应用服务,应用服务高CPU,应用服务部署失败,故障解答容器应用,故障解答功能,故障解答AKS,VM RDP,Linux SSH,VM黑屏幕,无法连接到…

运维与云

Azure Prepare

准备 azd 用于部署的Azure项目:为Azure开发者CLI(azd)工作流程生成azure.yaml,基础设施(Bicep/Terraform)和多克文件. 仅当用户明确想要使用 azd 作为部署工具时使用, 或项目已经有一个 azure 。 雅姆尔文件。 不使用: 非az…