跳到主内容
智客 ZICQ

技能库 智客分类:运维与云 google-agents-cli-observability

Google Agents Cli Observability

当用户想"设置追踪","监控我的代理","配置伐木","增加可观察性","调试出产流量",或者需要指导监测部署的代理,包括ADK(Agent Development Kit)代理时,应该使用这种技能. 封面有"云迹","即时反应记录","大查询代理分析","第三方整合"(AgentOps,凤凰,MLflow等)等,并进行故障排除. 部分特工-Cli技能套房. 不要用于部署设置(使用google-agents-cli-depload)或API代码模式(使用google-agents-cli-adk-code).

187824 安装量

官方网址:skills.sh

技能介绍

先看中文介绍;官方 description 原文单独保留,不改写 SKILL.md。

做什么

当用户想"设置追踪","监控我的代理","配置伐木","增加可观察性","调试出产流量",或者需要指导监测部署的代理,包括ADK(Agent Development Kit)代理时,应该使用这种技能. 封面有"云迹","即时反应记录","大查询代理分析","第三方整合"(AgentOps,凤凰,MLflow等)等,并进行故障排除. 部分特工-Cli技能套房. 不要用于部署设置(使用google-agents-cli-depload)或API代码模式(使用google-agents-cli-adk-code).

何时用

官方 description 未单独写出 Use when。按规范,代理会在用户任务与这段 description 的关键词匹配时激活本技能。

代理如何加载

按 Agent Skills 渐进披露:启动时只加载 name 与 description(约 100 token);任务匹配后才读入整份 SKILL.md 正文;scripts/、references/、assets/ 仅在需要时再读。 本文件正文结构:Observability Guide、Order of operations for `agent_runtime` deployments、Reference Files、Observability Tiers、Cloud Trace、Span Hierarchy。

文件分析

文件分析:除 SKILL.md 外,正文引用了 references/cloud-trace-and-logging.md、references/bigquery-agent-analytics.md、references/adk-docs.md、references/feedback-mechanism.md,属于带资源的技能包,这些文件按需再读。

官方 description(原文)

This skill should be used when the user wants to "set up tracing", "monitor my agent", "configure logging", "add observability", "debug production traffic", or needs guidance on monitoring deployed agents, including ADK (Agent Development Kit) agents. Covers Cloud Trace, prompt-response logging, BigQuery Agent Analytics, third-party integrations (AgentOps, Phoenix, MLflow, etc.), and troubleshooting. Part of the agents-cli skills suite. Do NOT use for deployment setup (use google-agents-cli-deploy) or API code patterns (use google-agents-cli-adk-code).

Observability GuideOrder of operations for `agent_runtime` deploymentsReference FilesObservability TiersCloud TraceSpan HierarchySetup by Deployment TypePrompt-Response LoggingBigQuery Agent Analytics PluginThird-Party IntegrationsTroubleshootingRelated Skills

来源分类:skills.sh agent-skill

SKILL.md 与 Agent 调用

官方规范 ↗
name
google-agents-cli-observability
description
This skill should be used when the user wants to "set up tracing", "monitor my agent", "configure logging", "add observability", "debug production traffic", or needs guidance on monitoring deployed agents, including ADK (Agent Development Kit) agents. Covers Cloud Trace, prompt-response logging, BigQuery Agent Analytics, third-party integrations (AgentOps, Phoenix, MLflow, etc.), and troubleshooting. Part of the agents-cli skills suite. Do NOT use for deployment setup (use google-agents-cli-deploy) or API code patterns (use google-agents-cli-adk-code).
  1. 发现技能客户端向 Agent 提供名称与描述目录。
  2. 匹配与调用用户指定或任务匹配后,载入 SKILL.md 指令。
  3. 按需加载按步骤读取参考文档、使用脚本与素材。
指令中引用的文件 · 4
  • references/cloud-trace-and-logging.md
  • references/bigquery-agent-analytics.md
  • references/adk-docs.md
  • references/feedback-mechanism.md

以下路径提取自原文;文件是否齐全请以来源仓库中的完整目录为准。

具体调用语法与可用工具以目标 Agent 客户端为准。 查看调用机制说明 ↗

安装这个技能

Skills CLI ↗

先选择目标 Agent 和安装范围,保留技能包的附属文件,安装后检查客户端能否发现该技能。

该技能引用了附属文件,请从来源获取完整目录;仅复制 SKILL.md 可能缺少依赖。

交给 Agent 安装

复制安装指令给支持 Agent Skills 的代理,确认其中的目标目录与客户端匹配。

把 Agent Skill「google-agents-cli-observability」安装到我的项目:SKILL.md 原文与官方 description 见 https://zicq.com/zh/skills/skl-a010414671bc4dee-Google-Agents-Cli-Observability.html
请存为 .cursor/skills/google-agents-cli-observability/SKILL.md 或 .claude/skills/google-agents-cli-observability/SKILL.md,frontmatter 的 name 与 description 保持原样,不要改写。
该技能还带 scripts/、references/、assets/ 等文件,请从 https://github.com/google/agents-cli 取完整目录,不要只建一个 SKILL.md。

GitHub 完整包 ↗

终端安装 · Skills CLI

需要 Node.js 与 npx。先查看仓库技能列表,确认实际名称。

npx skills add 'https://github.com/google/agents-cli' --list

npx skills add 'https://github.com/google/agents-cli' --skill 'google-agents-cli-observability'

CLI 会交互选择目标 Agent,默认安装到项目;用户级安装使用 -g。先通过查看命令核对仓库内容,再用 npx skills list 检查已安装技能。

阅读排版
--- name: google-agents-cli-observability description: > This skill should be used when the user wants to "set up tracing", "monitor my agent", "configure logging", "add observability", "debug production traffic", or needs guidance on monitoring deployed agents, including ADK (Agent Development Kit) agents. Covers Cloud Trace, prompt-response logging, BigQuery Agent Analytics, third-party integrations (AgentOps, Phoenix, MLflow, etc.), and troubleshooting. Part of the agents-cli skills suite. Do NOT use for deployment setup (use google-agents-cli-deploy) or API code patterns (use google-agents-cli-adk-code). metadata: author: Google license: Apache-2.0 version: 1.5.0 requires: bins: - agents-cli install: "uv tool install google-agents-cli" --- # Observability Guide > **Cloud Trace** works out of the box — no infrastructure needed. **Prompt-response logging** and **BigQuery Agent Analytics** require Terraform-provisioned infrastructure (service account, GCS bucket, BigQuery dataset). Run `agents-cli infra single-project --project PROJECT_ID` to provision these resources. See `references/cloud-trace-and-logging.md` for details, env vars, and verification commands. If your project isn't scaffolded yet, see `/google-agents-cli-scaffold` first. ### Order of operations for `agent_runtime` deployments For `deployment_target = agent_runtime`, run `agents-cli infra single-project` **before** the first `agents-cli deploy`. The Terraform module owns the entire Reasoning Engine resource (service account, deployment spec, env vars), so applying it after an SDK-based deploy creates a state mismatch Terraform can't reconcile without taking ownership of the whole resource. Already ran `agents-cli deploy`? Two options: 1. **Switch to Terraform-managed** — delete the SDK-deployed Reasoning Engine, then run `agents-cli infra single-project` and `agents-cli deploy` (sessions and in-flight state are lost). 2. **Keep the SDK-deployed instance** — skip `infra single-project` and set the observability env vars by re-running `agents-cli deploy --update-env-vars "KEY=VALUE,..."`; deploy matches the existing Reasoning Engine by display name and updates it in place, preserving env vars set outside the deploy. You must also grant its service account the telemetry IAM roles the Terraform module would otherwise provision: `roles/storage.admin` (write completions to the logs bucket), `roles/logging.logWriter`, `roles/cloudtrace.agent`, plus `roles/bigquery.dataOwner` + `roles/bigquery.jobUser` when scaffolded with `--bq-analytics`. The full set lives in `deployment/terraform/single-project/iam.tf` (from `app_sa_roles`) and `telemetry.tf`. Terraform-managed env vars aren't available in this mode. ### Reference Files | File | Contents | |------|----------| | `references/cloud-trace-and-logging.md` | Scaffolded project details — Terraform-provisioned resources, environment variables, verification commands, enabling/disabling locally | | `references/bigquery-agent-analytics.md` | BQ Agent Analytics plugin — enabling, key features, GCS offloading, tool provenance | | `references/adk-docs.md` | **ADK:** adk.dev pages to fetch for detail beyond this skill | | `references/feedback-mechanism.md` | Adding a user-feedback endpoint — request model, structured logging, log sink → BigQuery | --- ## Observability Tiers Choose the right level of observability based on your needs: | Tier | What It Does | Scope | Default State | Best For | |------|-------------|-------|---------------|----------| | **Cloud Trace** | Distributed tracing — execution flow, latency, errors via OpenTelemetry spans | All templates, all environments | Always enabled | Debugging latency, understanding agent execution flow | | **Prompt-Response Logging** | GenAI interactions exported to GCS, BigQuery, and Cloud Logging | Scaffolded projects | Disabled locally, enabled when deployed | Auditing LLM interactions, compliance | | **BigQuery Agent Analytics** | Structured agent events (LLM calls, tool use, outcomes) to BigQuery | ADK agents with the plugin enabled | Opt-in (`--bq-analytics` at scaffold time) | Conversational analytics, custom dashboards, LLM-as-judge evals | | **Third-Party Integrations** | External observability platforms (AgentOps, Phoenix, MLflow, etc.) | Any OpenTelemetry-instrumented agent | Opt-in, per-provider setup | Team collaboration, specialized visualization, prompt management | **Ask the user** which tier(s) they need — they can be combined. Cloud Trace is always on; the others are additive. --- ## Cloud Trace Scaffolded agents use OpenTelemetry to emit distributed traces. Every agent invocation produces spans that track the full execution flow. ### Span Hierarchy > **ADK projects.** These are ADK's span names; other frameworks emit their own (`generate_content` comes from the shared google-genai instrumentor either way). ``` invoke_workflow (top-level run) └── invoke_agent (one per agent in the chain) ├── call_llm (model request) │ └── generate_content (underlying GenAI model call) └── execute_tool (tool execution) ``` ### Setup by Deployment Type | Deployment | Setup | |-----------|-------| | **Agent Runtime** | Automatic — exporters wired at startup, gated on `GOOGLE_CLOUD_AGENT_ENGINE_ENABLE_TELEMETRY` (set by deploy); exports to Cloud Trace/Logging + Agent Engine console | | **Cloud Run / GKE (scaffolded)** | Automatic — exporters wired at startup, exports to Cloud Trace/Logging | | **Cloud Run / GKE (manual)** | Configure OpenTelemetry exporter in your app | | **Local dev** | Works with `agents-cli playground`; traces visible in Cloud Console | **ADK:** the wiring is `get_fast_api_app(otel_to_cloud=True)` in `app/fast_api_app.py`. Other templates call their own setup at startup (e.g. `app/app_utils/telemetry.py`). View traces: **Cloud Console → Trace → Trace explorer** **ADK:** for detailed setup instructions (Agent Runtime CLI/SDK, Cloud Run, custom deployments), fetch `https://adk.dev/integrations/cloud-trace/index.md`. --- ## Prompt-Response Logging Captures GenAI interactions and exports to GCS (JSONL) and BigQuery (via log sinks + external tables). Content is governed by **two independent tiers**; the net Terraform-deploy default is **full content in GCS/BigQuery, none in traces**: | Tier | Captures | Controlled by | Default (Terraform deploy) | |------|----------|---------------|----------------------------| | **GCS/BigQuery completions** | Full prompts/responses (the prompt-response logging feature) | `OTEL_INSTRUMENTATION_GENAI_COMPLETION_HOOK=upload` + `LOGS_BUCKET_NAME` | **On** — full content | | **Trace spans / Cloud Logging events** | Span/event content | `OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT` (plus `ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS=false`, **ADK only**) | **Off** — `NO_CONTENT` | The tiers are independent: GCS/BigQuery uploads capture full content whenever their upload vars are set and do **not** honor `OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT`, which governs the traces/events tier only. Its valid (experimental-semconv) values: - `NO_CONTENT` — no content in spans/events (scaffolded default) - `EVENT_ONLY` — content in Cloud Logging events - `SPAN_ONLY` / `SPAN_AND_EVENT` — content in trace spans - `true` / `false` — **invalid**; fall back to `NO_CONTENT` For the full mechanics (semconv opt-in, declarative Terraform config, env-var table, enabling/disabling, verification commands), see `references/cloud-trace-and-logging.md`. For ADK logging docs (log levels, configuration, debugging), fetch `https://adk.dev/observability/logging/index.md`. --- ## BigQuery Agent Analytics Plugin > **ADK projects.** Optional ADK plugin that logs structured agent events to BigQuery. Enable with `--bq-analytics` at scaffold time. See `references/bigquery-agent-analytics.md` for details. --- ## Third-Party Integrations Many third-party observability platforms can ingest agent telemetry (via OpenTelemetry or custom instrumentation). The table below covers common ones; the full list is larger (see the pointer below it). | Platform | Key Differentiator | Setup Complexity | Self-Hosted Option | |----------|-------------------|-----------------|-------------------| | **AgentOps** | Session replays, 2-line setup, replaces native telemetry | Minimal | No (SaaS) | | **Arize AX** | Commercial platform, production monitoring, evaluation dashboards | Low | No (SaaS) | | **Phoenix** | Open-source, custom evaluators, experiment testing | Low | Yes | | **MLflow** | OTel traces to MLflow Tracking Server, span tree visualization | Medium (needs SQL backend) | Yes | | **Monocle** | 1-call setup, VS Code Gantt chart visualizer | Minimal | Yes (local files) | | **Weave** | W&B platform, team collaboration, timeline views | Low | No (SaaS) | | **Freeplay** | Prompt management + evals + observability in one platform | Low | No (SaaS) | **Ask the user** which platform they prefer — present the trade-offs and let them choose. **ADK:** fetch a platform's setup page at `https://adk.dev/integrations//index.md` (slugs for the table above: `agentops`, `arize-ax`, `phoenix`, `mlflow-tracing`, `monocle`, `weave`, `freeplay`); ADK has more observability integrations (Datadog, Galileo, LangWatch, Latitude, Future AGI, Respan, Zespan, …) — browse the complete, current list at `https://adk.dev/integrations/` (observability topic). On other frameworks the OpenTelemetry-based platforms still work, but follow the platform's own setup docs. --- ## Troubleshooting | Issue | Solution | |-------|----------| | No traces in Cloud Trace | Verify telemetry setup runs at startup (**ADK:** `fast_api_app.py` uses `get_fast_api_app(otel_to_cloud=True)`; Agent Runtime gates it on `GOOGLE_CLOUD_AGENT_ENGINE_ENABLE_TELEMETRY`) and the SA has the `cloudtrace.agent` role | | Prompt-response data not appearing | Check `LOGS_BUCKET_NAME` is set; verify SA has `storage.objectCreator` on the bucket; check app logs for telemetry setup warnings | | Content in traces/events (unwanted) | `OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=NO_CONTENT` keeps content out of spans/events. NOTE: GCS/BigQuery completions still capture full content — to stop that, remove `LOGS_BUCKET_NAME`/`OTEL_INSTRUMENTATION_GENAI_COMPLETION_HOOK` (drop the upload block in `service.tf`) | | BigQuery Analytics not logging | **ADK:** verify the plugin is configured in `app/agent.py`; check `BQ_ANALYTICS_DATASET_ID` env var is set | | Third-party integration not capturing spans | Check provider-specific env vars (API keys, endpoints); some providers (AgentOps) replace native telemetry | | Traces missing tool spans | **ADK:** tool execution spans appear under `execute_tool` (other frameworks use their own span names) — check trace explorer filters | | High telemetry costs | Switch to `NO_CONTENT` mode; reduce BigQuery retention; disable unused tiers | --- ## Related Skills - `/google-agents-cli-deploy` — Deployment targets, CI/CD pipelines, and production workflows - `/google-agents-cli-workflow` — Development workflow, coding guidelines, and operational rules - `/google-agents-cli-adk-code` — ADK Python API quick reference for writing agent code

相关技能

运维与云

Docker Essentials

用于容器管理,图像操作,调试的基本道克命令和工作流程.

运维与云

Find Skills

从开放的代理技能生态系统中发现并安装技能. 使用时:(1)用户问"我如何做X",X可能拥有现有技能,(2)用户说"为X找到技能"或"是否为X有技能",(3)用户问"你能否做X",X是专门能力,(4)用户想扩展代理能力,(5)用户想搜索工具,模板,或工作流程,(6)用户提到他们希望…

运维与云

Azure Diagnostics

Azure上使用AppLens,AzureMonitor,资源健康,安全分型的调试Azure生产问题. 当:调试生产问题,故障解答应用服务,应用服务高CPU,应用服务部署失败,故障解答容器应用,故障解答功能,故障解答AKS,VM RDP,Linux SSH,VM黑屏幕,无法连接到…

运维与云

Azure Prepare

准备 azd 用于部署的Azure项目:为Azure开发者CLI(azd)工作流程生成azure.yaml,基础设施(Bicep/Terraform)和多克文件. 仅当用户明确想要使用 azd 作为部署工具时使用, 或项目已经有一个 azure 。 雅姆尔文件。 不使用: 非az…