跳到主内容
智客 ZICQ

技能库 智客分类:运维与云 devops-automation

开发自动化

开发与信息技术 业务自动化-CI/CD、监测、事件管理和基础设施工作流程

4144 安装量

官方网址:skills.sh

技能介绍

先看中文介绍;官方 description 原文单独保留,不改写 SKILL.md。

做什么

开发与信息技术 业务自动化-CI/CD、监测、事件管理和基础设施工作流程

何时用

官方 description 未单独写出 Use when。按规范,代理会在用户任务与这段 description 的关键词匹配时激活本技能。

代理如何加载

按 Agent Skills 渐进披露:启动时只加载 name 与 description(约 100 token);任务匹配后才读入整份 SKILL.md 正文;scripts/、references/、assets/ 仅在需要时再读。 本文件正文结构:DevOps Automation、Overview、CI/CD Automation、GitHub Actions Integration、Deployment Pipeline、Monitoring & Alerting。 其中含规范建议的小节:分步指令。

文件分析

文件分析:这是一份仅含 SKILL.md 的指令型技能,代理激活后整份正文进入上下文。

官方 description(原文)

DevOps and IT Ops automation - CI/CD, monitoring, incident management, and infrastructure workflows

DevOps AutomationOverviewCI/CD AutomationGitHub Actions IntegrationDeployment PipelineMonitoring & AlertingAlert RoutingAlert TemplatesIncident ManagementIncident WorkflowPostmortem TemplateInfrastructure Automation

· 许可:MIT

来源分类:skills.sh agent-skill

SKILL.md 与 Agent 调用

官方规范 ↗
name
devops-automation
description
DevOps and IT Ops automation - CI/CD, monitoring, incident management, and infrastructure workflows
许可
MIT
  1. 发现技能客户端向 Agent 提供名称与描述目录。
  2. 匹配与调用用户指定或任务匹配后,载入 SKILL.md 指令。
  3. 按需加载按步骤读取参考文档、使用脚本与素材。

具体调用语法与可用工具以目标 Agent 客户端为准。 查看调用机制说明 ↗

安装这个技能

Skills CLI ↗

先选择目标 Agent 和安装范围,保留技能包的附属文件,安装后检查客户端能否发现该技能。

交给 Agent 安装

复制安装指令给支持 Agent Skills 的代理,确认其中的目标目录与客户端匹配。

把 Agent Skill「devops-automation」安装到我的项目:SKILL.md 原文与官方 description 见 https://zicq.com/zh/skills/skl-0b4904cda9044f3d-%E5%BC%80%E5%8F%91%E8%87%AA%E5%8A%A8%E5%8C%96.html
请存为 .cursor/skills/devops-automation/SKILL.md 或 .claude/skills/devops-automation/SKILL.md,frontmatter 的 name 与 description 保持原样,不要改写。

GitHub 完整包 ↗

终端安装 · Skills CLI

需要 Node.js 与 npx。先查看仓库技能列表,确认实际名称。

npx skills add 'https://github.com/claude-office-skills/skills' --list

npx skills add 'https://github.com/claude-office-skills/skills' --skill 'devops-automation'

CLI 会交互选择目标 Agent,默认安装到项目;用户级安装使用 -g。先通过查看命令核对仓库内容,再用 npx skills list 检查已安装技能。

阅读排版

name: devops-automation description: "DevOps and IT Ops automation - CI/CD, monitoring, incident management, and infrastructure workflows" version: "1.0.0" author: claude-office-skills license: MIT

category: devops tags:

  • devops
  • ci-cd
  • monitoring
  • incident
  • automation department: Engineering

models: recommended: - claude-sonnet-4

mcp: server: devops-mcp tools: - github_api - jenkins_trigger - aws_cli - kubernetes_api

capabilities:

  • ci_cd_pipelines
  • monitoring_alerts
  • incident_management
  • infrastructure_automation
  • deployment_workflows

languages:

  • en
  • zh

related_skills:

  • slack-workflows
  • telegram-bot
  • ai-agent-builder

DevOps Automation

Automate DevOps workflows including CI/CD pipelines, monitoring, incident management, and infrastructure operations. Based on n8n's IT Ops workflow templates.

Overview

This skill covers:

  • CI/CD pipeline automation
  • Monitoring and alerting
  • Incident management
  • Infrastructure automation
  • Deployment workflows

CI/CD Automation

GitHub Actions Integration

workflow: "GitHub CI/CD Notifications"

triggers:
  - github_push
  - github_pull_request
  - github_workflow_run
  
on_push:
  action:
    - trigger_ci: if_main_branch
    - notify_slack:
        channel: "#deployments"
        message: |
          📦 *New Push to {branch}*
          
          Commit: `{commit_sha_short}`
          Author: {author}
          Message: {commit_message}
          
          [View Diff]({compare_url})

on_pr_opened:
  action:
    - notify_slack:
        channel: "#code-review"
        message: |
          🔀 *New Pull Request*
          
          Title: {pr_title}
          Author: {author}
          Branch: {head} → {base}
          
          [Review PR]({pr_url})
    - assign_reviewers: based_on_codeowners
    - run_ci_checks

on_workflow_complete:
  action:
    - notify_slack:
        message: |
          {status_emoji} *Build {status}*
          
          Workflow: {workflow_name}
          Branch: {branch}
          Duration: {duration}
          
          {if_failed: [View Logs]({logs_url})}

Deployment Pipeline

deployment_pipeline:
  stages:
    build:
      trigger: push_to_main
      steps:
        - checkout_code
        - install_dependencies
        - run_tests
        - build_artifact
        - push_to_registry
        
    staging:
      trigger: build_success
      steps:
        - deploy_to_staging
        - run_integration_tests
        - notify_qa
        
    production:
      trigger: manual_approval
      steps:
        - create_backup
        - deploy_to_production
        - run_smoke_tests
        - notify_team
        
  rollback:
    trigger: deployment_failed OR manual
    steps:
      - revert_to_previous
      - notify_team
      - create_incident

Monitoring & Alerting

Alert Routing

alert_routing:
  sources:
    - prometheus
    - datadog
    - cloudwatch
    - new_relic
    
  severity_levels:
    critical:
      response_time: 5_minutes
      channels: [pagerduty, slack_urgent, sms]
      escalation: immediate
      
    high:
      response_time: 15_minutes
      channels: [slack_alerts, email]
      escalation: after_15_minutes
      
    medium:
      response_time: 1_hour
      channels: [slack_alerts]
      
    low:
      response_time: 24_hours
      channels: [slack_logging]
      
  routing_rules:
    - if: service == "payments"
      team: payments_oncall
      severity_boost: +1
      
    - if: service == "auth"
      team: security_oncall
      
    - default:
      team: platform_oncall

Alert Templates

alert_templates:
  infrastructure:
    cpu_high:
      title: "🔥 High CPU Usage"
      body: |
        Server: {host}
        CPU: {cpu_percent}%
        Duration: {duration}
        
        Threshold: {threshold}%
        
        [View Dashboard]({grafana_url})
        
    memory_critical:
      title: "💾 Critical Memory"
      body: |
        Server: {host}
        Memory: {memory_percent}%
        Available: {available_mb}MB
        
        [SSH to Server]({ssh_link})
        
    disk_full:
      title: "💿 Disk Space Critical"
      body: |
        Server: {host}
        Disk: {disk_percent}%
        Available: {available_gb}GB
        
        Suggestion: Clean logs or expand volume
        
  application:
    error_spike:
      title: "📈 Error Rate Spike"
      body: |
        Service: {service}
        Error Rate: {error_rate}%
        Normal: {baseline}%
        
        Top Errors:
        {top_errors}
        
    latency_high:
      title: "🐢 High Latency"
      body: |
        Service: {service}
        P99 Latency: {p99_ms}ms
        Threshold: {threshold_ms}ms

Incident Management

Incident Workflow

incident_workflow:
  detection:
    sources: [monitoring, user_report, automated_check]
    
  triage:
    auto_severity:
      - if: affects_payments
        severity: critical
      - if: affects_auth
        severity: critical
      - if: affects_api AND error_rate > 10%
        severity: high
        
  response:
    critical:
      - create_incident_channel: "#inc-{timestamp}"
      - page_oncall: immediately
      - notify_stakeholders: [engineering_lead, product]
      - start_war_room: zoom_link
      - create_status_page: incident
      
    high:
      - create_incident_channel
      - notify_oncall: slack
      - create_ticket: jira
      
  communication:
    internal:
      frequency: every_30_minutes
      channel: incident_channel
      template: |
        📊 *Incident Update*
        
        Status: {status}
        Impact: {impact}
        Next update: {next_update_time}
        
        Current actions:
        {action_items}
        
    external:
      channel: status_page
      template: customer_facing_update
      
  resolution:
    steps:
      - confirm_resolution
      - update_status_page: resolved
      - notify_stakeholders
      - schedule_postmortem
      - close_incident_channel: after_24h

Postmortem Template

postmortem_template:
  sections:
    summary:
      - incident_title
      - duration
      - severity
      - impact
      
    timeline:
      format: |
        | Time | Event |
        |------|-------|
        | {time} | {event} |
        
    root_cause:
      - what_happened
      - why_it_happened
      - contributing_factors
      
    impact:
      - users_affected
      - revenue_impact
      - sla_breach
      
    resolution:
      - how_it_was_fixed
      - time_to_detect
      - time_to_resolve
      
    action_items:
      format: |
        | Action | Owner | Due Date | Status |
        |--------|-------|----------|--------|
        
    lessons_learned:
      - what_went_well
      - what_went_poorly
      - lucky_breaks

Infrastructure Automation

Server Provisioning

provisioning_workflow:
  trigger: jira_ticket OR slack_request
  
  steps:
    1. validate_request:
        check: [budget_approval, security_review]
        
    2. create_infrastructure:
        terraform:
          - vpc
          - security_groups
          - ec2_instances
          - load_balancer
          
    3. configure_server:
        ansible:
          - base_configuration
          - security_hardening
          - monitoring_agent
          - application_setup
          
    4. validate:
        - health_check
        - security_scan
        - performance_baseline
        
    5. notify:
        slack: "✅ Server {hostname} is ready"
        include: [ssh_access, dashboard_link]

Scheduled Maintenance

maintenance_automation:
  tasks:
    certificate_renewal:
      schedule: "30 days before expiry"
      action:
        - request_new_cert: letsencrypt
        - deploy_cert
        - verify_ssl
        - notify: if_failure
        
    security_patching:
      schedule: "weekly"
      action:
        - check_updates
        - if_critical: immediate_patch
        - else: schedule_maintenance_window
        
    log_rotation:
      schedule: "daily"
      action:
        - rotate_logs
        - compress_old
        - upload_to_s3
        - delete_local: older_than_7_days
        
    backup_verification:
      schedule: "weekly"
      action:
        - restore_to_test_env
        - run_integrity_checks
        - report_status

Kubernetes Automation

K8s Workflows

kubernetes_automation:
  deployment:
    trigger: docker_image_pushed
    steps:
      - update_manifest: with_new_image_tag
      - apply_to_staging
      - run_tests
      - if_success: apply_to_production
      
  scaling:
    trigger: metric_threshold
    rules:
      - if: cpu > 80%
        action: scale_up
        max_replicas: 10
      - if: cpu < 20%
        action: scale_down
        min_replicas: 2
        
  rollback:
    trigger: health_check_failed
    action:
      - kubectl_rollout_undo
      - notify_team
      - create_incident

Output Example

Request: "Set up deployment notifications for GitHub Actions"

Output:

# GitHub Actions Deployment Notifications

## n8n Workflow

```yaml
trigger: GitHub Webhook
events: [workflow_run]

Notification Templates

Build Started:

🚀 *Deployment Started*

Branch: main
Commit: abc1234
Author: @developer
Triggered by: Push

[View Workflow](https://github.com/...)

Build Success:

✅ *Deployment Successful*

Environment: Production
Duration: 3m 42s
Version: v1.2.3

Changes:
• Feature X
• Bug fix Y

[View Deployment](https://app.example.com)

Build Failed:

❌ *Deployment Failed*

Stage: Test
Error: npm test failed

[View Logs](https://github.com/...)
[Retry](https://github.com/...)

Slack Integration

channel: "#deployments"
mention_on_failure: "@oncall"
thread_replies: true

---

*DevOps Automation Skill - Part of Claude Office Skills*

相关技能

运维与云

Docker 要点Docker Essentials

用于容器管理,图像操作,调试的基本道克命令和工作流程.

运维与云

查找技能Find Skills

从开放的代理技能生态系统中发现并安装技能. 使用时:(1)用户问"我如何做X",X可能拥有现有技能,(2)用户说"为X找到技能"或"是否为X有技能",(3)用户问"你能否做X",X是专门能力,(4)用户想扩展代理能力,(5)用户想搜索工具,模板,或工作流程,(6)用户提到他们希望…

运维与云

微软 创建Microsoft Foundry

构建,部署,评价,优化,微调,管理微软创始代理,模型,资源最终结束. 用法:铸造、azd ai代理、azd 提供/部署、托管代理脚手架/开发/运行/部署/拖放、即时代理创建、创建代理、更新代理、添加代理工具、援引代理、代理.yaml、评估代理、批量活化、连续活化、持续监测、代理C…

运维与云

Azure 部署Azure Deploy

执行已有.azure/部署计划.md和基础设施文件的ALREADY-PREPARED应用程序的Azure部署。 当用户向 CREATE 请求新应用程序时, 请不要使用此技能 。 这种技能运行了 azd上行, azd 部署, trarraform 应用, 和 az 部署命令带有内置…