跳到主内容
智客 ZICQ

技能库 智客分类:媒体内容 music-to-video

音乐到视频

将一首音乐曲目(一个音频文件,一个用来拉出音频的视频,或者一个由心情简介所生成的曲目)转换为一首节奏视频——歌词视频,幻灯片,或者动能宣传. 音乐驱动所有节奏;任何用户提供的图像/视频被剪接到同一个beat网格上,完整的视频需要零资产. 叙述片段 – 输入相匹配的工作流程(见/hyperframes). 不清楚 / 假肢 .

266012 安装量

官方网址:skills.sh

技能介绍

先看中文介绍;官方 description 原文单独保留,不改写 SKILL.md。

做什么

将一首音乐曲目(一个音频文件,一个用来拉出音频的视频,或者一个由心情简介所生成的曲目)转换为一首节奏视频——歌词视频,幻灯片,或者动能宣传. 音乐驱动所有节奏;任何用户提供的图像/视频被剪接到同一个beat网格上,完整的视频需要零资产. 叙述片段 – 输入相匹配的工作流程(见/hyperframes). 不清楚 / 假肢 .

何时用

官方 description 未单独写出 Use when。按规范,代理会在用户任务与这段 description 的关键词匹配时激活本技能。

代理如何加载

按 Agent Skills 渐进披露:启动时只加载 name 与 description(约 100 token);任务匹配后才读入整份 SKILL.md 正文;scripts/、references/、assets/ 仅在需要时再读。 本文件正文结构:music-to-video — one music-grounded, beat-synced video workflow、Two ideas that shape everything、Step 0: Setup, BGM, and inputs、only if the user gave you images/videos:、Step 1: Analyze the music、Step 2: Frame skeleton (structure only)。 其中含规范建议的小节:分步指令。

文件分析

文件分析:除 SKILL.md 外,正文引用了 references/plugin-installation.md、references/brief-contract.md、assets/bgm.mp3、references/intent-interview.md、references/routes/music-to-video.md、references/bgm.md,属于带资源的技能包,这些文件按需再读。

官方 description(原文)

Turn a music track (an audio file, a video to pull audio from, or a track generated from a mood brief) into a beat-synced video — lyric video, slideshow, or kinetic promo. The music drives all pacing; any user-supplied images/videos are cut onto the same beat grid, and a complete video needs zero assets. Narrated pieces → the input-matched workflow (see /hyperframes). Unclear → /hyperframes.

music-to-video — one music-grounded, beat-synced video workflowTwo ideas that shape everythingStep 0: Setup, BGM, and inputsonly if the user gave you images/videos:Step 1: Analyze the musicStep 2: Frame skeleton (structure only)Step 3: Fill the plan (user-gated)Step 4: Build frames from the planStep 5: AssembleStep 6: Verify and renderResume tableQuick Reference

来源分类:skills.sh agent-skill

SKILL.md 与 Agent 调用

官方规范 ↗
name
music-to-video
description
Turn a music track (an audio file, a video to pull audio from, or a track generated from a mood brief) into a beat-synced video — lyric video, slideshow, or kinetic promo. The music drives all pacing; any user-supplied images/videos are cut onto the same beat grid, and a complete video needs zero assets. Narrated pieces → the input-matched workflow (see /hyperframes). Unclear → /hyperframes.
  1. 发现技能客户端向 Agent 提供名称与描述目录。
  2. 匹配与调用用户指定或任务匹配后,载入 SKILL.md 指令。
  3. 按需加载按步骤读取参考文档、使用脚本与素材。
指令中引用的文件 · 9
  • references/plugin-installation.md
  • references/brief-contract.md
  • assets/bgm.mp3
  • references/intent-interview.md
  • references/routes/music-to-video.md
  • references/bgm.md
  • references/setup-providers.md
  • scripts/stage-assets.mjs
  • scripts/analyze-beatgrid.py

以下路径提取自原文;文件是否齐全请以来源仓库中的完整目录为准。

具体调用语法与可用工具以目标 Agent 客户端为准。 查看调用机制说明 ↗

安装这个技能

Skills CLI ↗

先选择目标 Agent 和安装范围,保留技能包的附属文件,安装后检查客户端能否发现该技能。

该技能引用了附属文件,请从来源获取完整目录;仅复制 SKILL.md 可能缺少依赖。

交给 Agent 安装

复制安装指令给支持 Agent Skills 的代理,确认其中的目标目录与客户端匹配。

把 Agent Skill「music-to-video」安装到我的项目:SKILL.md 原文与官方 description 见 https://zicq.com/zh/skills/skl-deb52633deb73812-%E9%9F%B3%E4%B9%90%E5%88%B0%E8%A7%86%E9%A2%91.html
请存为 .cursor/skills/music-to-video/SKILL.md 或 .claude/skills/music-to-video/SKILL.md,frontmatter 的 name 与 description 保持原样,不要改写。
该技能还带 scripts/、references/、assets/ 等文件,请从 https://github.com/heygen-com/hyperframes 取完整目录,不要只建一个 SKILL.md。

GitHub 完整包 ↗

终端安装 · Skills CLI

需要 Node.js 与 npx。先查看仓库技能列表,确认实际名称。

npx skills add 'https://github.com/heygen-com/hyperframes' --list

npx skills add 'https://github.com/heygen-com/hyperframes' --skill 'music-to-video'

CLI 会交互选择目标 Agent,默认安装到项目;用户级安装使用 -g。先通过查看命令核对仓库内容,再用 npx skills list 检查已安装技能。

阅读排版
--- name: music-to-video description: "Turn a music track (an audio file, a video to pull audio from, or a track generated from a mood brief) into a beat-synced video — lyric video, slideshow, or kinetic promo. The music drives all pacing; any user-supplied images/videos are cut onto the same beat grid, and a complete video needs zero assets. Narrated pieces → the input-matched workflow (see /hyperframes). Unclear → /hyperframes." --- **Plugin installs:** Before setup or freshness commands, follow [plugin execution rules](../hyperframes/references/plugin-installation.md) when this skill is inside a HyperFrames plugin. Standalone installs keep the update instructions below. > **First, keep this skill fresh — confirm with the user before running:** `npx hyperframes skills update music-to-video`. A fast no-op when everything is current; otherwise it refreshes this skill plus the core domain skills it depends on before you rely on them. # music-to-video — one music-grounded, beat-synced video workflow Use this skill to turn a **music track** into a beat-synced HyperFrames video. You analyze the track once, lay out the frames, fill in a per-frame plan, and build each frame as a composition. The input is a music track plus optional user images or videos — there is **no narration and no website capture**. Typography and templates are the floor (a complete video needs zero assets); any media the user supplies is cut in on the same beat grid. You are the **orchestrator**. Work in `videos//`. Run the steps in order and pass each **Gate** before moving on. Two steps need the user: **Step 3** (plan approval) and **Step 6** (render approval) — both are checkpoint gates per `../hyperframes/references/brief-contract.md` (read it before Step 0): in autonomous mode, post the summary as a heads-up and proceed instead of waiting. Do every step yourself except **Step 4**, where you dispatch **one sub-agent per frame**. Keep design and motion rules out of this file — they live in `references/` and the `frame-worker` sub-agent. `SKILL_DIR` = this skill directory. `PROJECT_DIR` = `videos//`. Workflow: Step 0 setup → `hyperframes.json` + `assets/bgm.mp3`; Step 1 analyze → `audiomap.json`; Step 2 skeleton → `STORYBOARD.md` (frames, groups `TBD`); Step 3 plan → complete `STORYBOARD.md` + `frame.md`; Step 4 build → `compositions/frames/NN-*.html`; Step 5 assemble → `index.html`; Step 6 render → `renders/video.mp4`. ## Two ideas that shape everything - **One analyzer, and you trust it.** `analyze-beatgrid.py` is the only beat analyzer — never re-measure beats with another tool or by ear. Its energy / density / rolls / onsets / silences are always reliable. Its `bpm` and `beats_sec` are reliable **only when the music is genuinely rhythmic**; on calm music the grid is a metronome the tracker imposed, so pace by phrases and energy instead and never hard-cut to it. Deciding which case you're in is each frame's `pacing` (Step 2). - **One frame = one file; groups live inside.** Step 2 cuts the track into **frames**, and each frame becomes one composition file `compositions/frames/NN-.html`, built by one frame-worker. A frame can subdivide into **groups** (each a template or a motion-primitives combo). Extra density goes _inside_ a group, so **frame count tracks distinct treatments, not beats** — a fast track does not blow up the number of sub-agents. --- ## Step 0: Setup, BGM, and inputs Goal: Establish the music source, create the HyperFrames project, and note any user-supplied media. **The brief starts at the intent layer.** Opening rule, in order: **(1)** `BRIEF.md` exists → read it and ask nothing it answers — its `flow`/`storyboard` derive the mode (brief contract § 1). **(2)** No `BRIEF.md` but the project exists → resume from what's on disk; never re-interrogate. **(3)** A fresh creation request that arrived here directly → read `/hyperframes` and run its intent layer (`references/intent-interview.md`): it confirms this route's must-haves (the music source, destination → aspect — `../hyperframes/references/routes/music-to-video.md`) and announces what stays deferred — brand and genre are chosen at Step 3 by design. Write `BRIEF.md` immediately after init (never before — `init` refuses a non-empty directory) and record the preference-backed answers (`brief-format.md`). Edit requests skip all of this. The **music is the spine** — establish one track before anything else. This skill is tuned for **fast, high-energy BGM**: a strong beat grid drives the cuts (calm tracks work, but pace by phrase rather than beat). If the user supplied audio — a music file, or a video to pull audio from — use it. Otherwise choose the mood from the request and generate a track through `/media-use` (`references/bgm.md`). Before the first authenticated provider action, run `npx hyperframes auth status` and relay its output verbatim. If signed out, apply one branch: - **Collaborative:** wait for sign-in or an explicit choice to continue offline with the local provider. - **Autonomous:** state the status and continue through the available local provider. If no offline provider can satisfy the required music capability, surface the blocker. Never write keys into a per-repo `.env`. Auth ownership and offline fallbacks live in `/media-use` `references/setup-providers.md` § Providers. The resulting track lands at `assets/bgm.mp3`. Stage supplied images or videos so frames can use them on the beat grid; otherwise typography carries the video. **Lyric videos:** for lyrics synced to the vocals, get word/line timing by transcribing the track via `/media-use`, or ask the user for the lyrics text and place lines on the beat grid. Initialize only if `hyperframes.json` is missing. Name `` from the brief in kebab-case, such as `midnight-drive-loop` — never a timestamp. `init` checks the installed skills against the latest on GitHub and updates the global set if any are out of date. ```bash npx hyperframes init "videos/" --non-interactive --example=blank --skill=music-to-video mkdir -p "$PROJECT_DIR/assets" "$PROJECT_DIR/renders" cp "" "$PROJECT_DIR/assets/bgm.mp3" # extract from a video first if needed # only if the user gave you images/videos: node /scripts/stage-assets.mjs --from --hyperframes "$PROJECT_DIR" --into public ``` The **brand** (font + palette) is chosen at Step 3, not here. Don't pick a genre or a track type up front — assets are just an optional ingredient, and the genre emerges from the per-frame choices. **Gate:** `hyperframes.json` + `assets/bgm.mp3` exist; aspect / length / fps and (if any) the asset inventory are noted. --- ## Step 1: Analyze the music Goal: Produce the one canonical timing analysis the whole video is built on. `analyze-beatgrid.py` is the **only** beat analyzer — never re-measure beats with another tool or by ear. It reads the track once and writes `audiomap.json`: energy phases (level / density / feel), onsets + `onset_rate`, rolls, silences, `hard_stops`, `key_moments`, phrases, tempo / grid, and `audio.duration_sec`. It's deterministic — the same file always gives the same map. Most fields are reliable on any music; `bpm` and `beats_sec` are reliable only when the music is genuinely rhythmic, and judging that is the call you make at Step 2. Prerequisites: Python 3 with `librosa`, `numpy`, and `soundfile` available. If import fails, install them into the active Python environment before running the analyzer: ```bash python3 -m pip install librosa numpy soundfile ``` ```bash python3 /scripts/analyze-beatgrid.py "$PROJECT_DIR/assets/bgm.mp3" \ -o "$PROJECT_DIR/audiomap.json" --print ``` **Gate:** `audiomap.json` exists; `audio.duration_sec` is known. --- ## Step 2: Frame skeleton (structure only) Goal: Read the music and lay out the frames — the skeleton of `STORYBOARD.md`. Read [`references/frame-skeleton.md`](references/frame-skeleton.md). Turn `audiomap.json` into the **skeleton** of `STORYBOARD.md` yourself — there is no intermediate JSON. Cut the track into **frames** at real musical changes (`hard_stops`, SURGE / DROP `key_moments`, the edges of a roll, a stretch with no onsets, a big energy jump), snapping every boundary to an audiomap anchor. For each frame set `span_sec`, `pacing` (the verdict from Step 1's trust call — `beat_cut` when the grid is real, `phrase_flow` when it's a metronome imposed on calm music), `mood`, and a one-line `feel` (the plain music situation Step 3 matches a template against). Only classify and lay out here: leave every frame's `### Groups` as `TBD (Step 3)` and the frontmatter `style` blank — no templates, copy, color, or fonts. Expect ~1–6 frames. **Gate:** frames tile the track (first at 0, last at `duration_s`); each carries `span_sec` + `pacing` + `mood` + `feel`; every `### Groups` is `TBD`; no content anywhere. --- ## Step 3: Fill the plan (user-gated) Goal: Turn the skeleton into an approved, complete `STORYBOARD.md`. Read [`references/planning.md`](references/planning.md), [`storyboard-format.md`](references/storyboard-format.md), [`template-catalog.md`](references/template-catalog.md), [`motion-primitive-catalog.md`](references/motion-primitive-catalog.md), and [`montage.md`](references/montage.md) (only if the user supplied assets). Editing the same file in place, do two things: 1. **Pick the brand.** Choose one preset from `../hyperframes-creative/frame-presets/` using the table in `../hyperframes-creative/references/design-spec.md` (match the track's mood; **only its fonts and colors matter** — templates own composition). Copy it into `frame.md` **unmodified** and fill the frontmatter `style` (font + a ≤4–6 swatch palette) from it. 2. **Fill every frame.** Decide its groups and give each a treatment: a matched template from the catalog (with bound params and real audiomap anchors), a free-compose from the primitive catalog, or an asset treatment that **obeys `pacing`**. **Before you free-compose a named look, search the live catalog for it**: for every look, effect, treatment or transition the user asked for — "CRT scanlines", "glitch", "film grain", "shimmer sweep" — run `npx hyperframes catalog --query "" --json` and read the top results. `template-catalog.md` and `motion-primitive-catalog.md` list only this skill's own local materials; the search ranks the whole hosted registry (~400 blocks and components) and needs **nothing installed** — no project, no prior `add`, no account. Free-compose a look only after a search for it came back with nothing that fits. Write the copy. You own WHAT (template / primitives + content + anchors); the frame-worker owns HOW — **never write millisecond tweens into the storyboard**. ```bash node /scripts/validate-plan.mjs --storyboard "$PROJECT_DIR/STORYBOARD.md" \ --audiomap "$PROJECT_DIR/audiomap.json" --templates /references/templates ``` Fix every `✗` (hard errors: duration mismatch, frames not tiling the track, a missing `src`); warnings are best-effort. Then present the frame-by-frame summary in chat as a proposal (`../hyperframes/references/review-loop.md` § 1) and iterate on the user's replies until they approve; for `storyboard: yes`, also write it as `storyboard.html` (`../hyperframes-creative/references/storyboard-recipe.md` § 3) for them to open. In autonomous mode this is a checkpoint gate: post the summary as a heads-up and proceed (the `validate-plan.mjs` pass is a quality gate and still blocks). **Gate:** `frame.md` is a verbatim preset copy; `validate-plan.mjs` exits 0; the user approved the plan (autonomous: the summary was posted as a heads-up). --- ## Step 4: Build frames from the plan Goal: Build every frame as a self-contained composition file. Create `compositions/frames/`. Read [`sub-agents/frame-worker.md`](sub-agents/frame-worker.md) and `../hyperframes/references/subagent-dispatch.md`. Dispatch **one frame-worker per frame**, in parallel where possible (otherwise in waves). Each worker gets exactly one frame and this context: ```text PROJECT_DIR: frame_id: # = the frame file stem, e.g. 02-f2; the composition id Your block: the `## Frame N — ` block in PROJECT_DIR/STORYBOARD.md audiomap: PROJECT_DIR/audiomap.json frame.md: PROJECT_DIR/frame.md Materials: for each group, /references/templates//index.html (templates) and /references/motion-primitives// (free); staged assets/ (asset groups) Contracts: ../hyperframes-core/references/sub-compositions.md + determinism-rules.md Canvas: × Pacing: Write to: PROJECT_DIR/compositions/frames/.html ``` The worker forks the cited materials, converts every anchor to frame-local seconds (`local_t = track_t − span_sec[0]`), gates its groups with 0ms cuts, and writes one seek-safe frame file. **The worker never runs the `hyperframes` CLI** — those commands operate on the assembled project, which doesn't exist yet, so they'd report on the wrong files. The worker just writes to the contract and stops; you verify after assembly (Step 6). As each worker returns, you can confirm its file landed on disk. **Gate:** every frame has its `compositions/frames/NN-*.html` on disk. --- ## Step 5: Assemble Goal: Wire the built frames + BGM into the playable `index.html`. `assemble-index.mjs` is deterministic — no subagent, no judgment. It references each frame file at its cumulative `data-start`, mounts `assets/bgm.mp3` on track 11, and hard-cuts frame → frame (frames tile the track with no gaps, so there is **no transition injector**). ```bash node /scripts/assemble-index.mjs --storyboard "$PROJECT_DIR/STORYBOARD.md" \ --hyperframes "$PROJECT_DIR" --audiomap "$PROJECT_DIR/audiomap.json" ``` Fix any `✗` it reports — a missing or blank frame file means that worker wrote a partial file; re-dispatch it (Step 4) and re-assemble. **Gate:** `index.html` exists; total duration == `audiomap.audio.duration_sec`. --- ## Step 6: Verify and render Goal: Verify the assembled video, get user approval, and render the final MP4. Run the CLI on the **assembled project** — that's the correct unit (the per-frame workers couldn't run it). `check` runs structural lint and the headless-browser runtime, layout, motion, and contrast gate in one pass; `--snapshots` also emits the review frames. ```bash ( cd "$PROJECT_DIR" && npx hyperframes check . --snapshots ) ``` Inspect at `t=0`, each frame start, the strongest DROP / SURGE, every `hard_stops[].t`, and the final frame. On failure, make the **cheapest safe fix** yourself: edit the offending `compositions/frames/NN-*.html`. Never change duration or audio timing to hide a sync issue. Once the gates pass, pause for user review, then render only on approval (autonomous mode: ask the one kept question — "preview first, or render?" — then deliver the MP4 with the contact sheet): ```bash ( cd "$PROJECT_DIR" && npx hyperframes render . --skill=music-to-video -q draft -o renders/video.mp4 --fps 30 ) ``` **Gate:** `check` passed and the snapshots were inspected; the user approved (autonomous: checks passed and the delivery includes the contact sheet); `renders/video.mp4` exists with audio, duration == `audiomap.audio.duration_sec`. The final reply states the MP4 path and duration. --- ## Resume table | You have | Continue from | | -------------------------- | ------------- | | `assets/bgm.mp3` only | Step 1 | | `audiomap.json` | Step 2 | | `STORYBOARD.md` (skeleton) | Step 3 | | `STORYBOARD.md` (complete) | Step 4 | | all frame files | Step 5 | | `index.html` | Step 6 | ## Quick Reference **Formats:** landscape `1920x1080` by default; portrait `1080x1920`; square `1080x1080`. Set the canvas once in the storyboard frontmatter (`canvas: { w, h, fps }`). **Scripts** under `scripts/`: `analyze-beatgrid.py` (the one analyzer), `validate-plan.mjs` (plan check), `assemble-index.mjs` (index assembly), `stage-assets.mjs` (stage user media), `lib/storyboard.mjs` (vendored parser). Everything else is the `hyperframes` CLI. | Read | When | | -------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------- | | [`references/frame-skeleton.md`](references/frame-skeleton.md) | Step 2: read the music, lay out the frames, set pacing | | [`references/planning.md`](references/planning.md) · [`storyboard-format.md`](references/storyboard-format.md) | Step 3: pick the brand, fill each frame, write the plan | | [`references/template-catalog.md`](references/template-catalog.md) | Step 3: pick a template per group | | [`references/motion-primitive-catalog.md`](references/motion-primitive-catalog.md) | Step 3/4: L0 recipes for free-compose | | [`references/montage.md`](references/montage.md) | Step 3/4: asset treatments (beat-cut / ken-burns) | | [`sub-agents/frame-worker.md`](sub-agents/frame-worker.md) | Step 4: dispatch + build one frame | | `../hyperframes/references/subagent-dispatch.md` | Step 4: dispatch sub-agents safely | | `../hyperframes-creative/references/design-spec.md` | Step 3: pick the preset (the brand) | ## Directory layout ``` music-to-video/ SKILL.md references/ frame-skeleton.md · planning.md · storyboard-format.md template-catalog.md · motion-primitive-catalog.md · montage.md templates// { index.html (+ assets/ · program.json) } ← L1 catalog impls motion-primitives// { index.html (mounts the scene), scene.html (the sub-composition) } (+ ../assets/gsap.min.js shared by recipes) ← L0 catalog impls scripts/ analyze-beatgrid.py · assemble-index.mjs · validate-plan.mjs · stage-assets.mjs · lib/storyboard.mjs sub-agents/ frame-worker.md ← the one subagent (one per frame) ```

相关技能

媒体内容

纳米香蕉 赞成Nano Banana Pro

以纳米·香蕉Pro(Gemini 3 Pro Image)来生成/编辑图像. 用于创建/修改请求,包括编辑。 支持文本到图像+图像到图像; 1K/2K/4K; 使用 -- input-image.

媒体内容

Youtube 监视器Youtube Watcher

从YouTube视频中获取并读取文字记录. 需要汇总视频时使用,回答有关视频内容的问题,或者从中提取信息.

媒体内容

Atxp 密码Atxp

Access ATXP支付API工具用于网页搜索,AI图像生成,音乐创建,视频生成,X/Twitter搜索,电子邮件,代理账户管理. 当用户需要实时网络搜索时使用,AI生成的媒体(图像,音乐,视频),X/Twitter搜索,发送/接收电子邮件,或创建和资助代理账户. 需要通过“ …

媒体内容

YouTube (英语)Youtube

YouTube Data API与管理的OAuth的集成. 搜索视频,管理播放列表,访问频道数据,并与评论互动. 用户想与YouTube互动时使用此技能. 对于其他第三方应用,使用api-gateway技能(https://clawhub.ai/byungkyu/api-gate…