跳到主内容
智客 ZICQ

技能库 智客分类:媒体内容 video-use

Video Use

通过对话编辑任何视频 。 转写,剪接,配色分级,生成叠加动画,烧制字幕——用于说话头,蒙太奇,辅导,出行,采访. 没有预设,没有菜单. 问问题,确认计划,执行,执行,坚持. 生产正确规则很困难;其他一切都是艺术自由.

3708 安装量

官方网址:skills.sh

技能介绍

先看中文介绍;官方 description 原文单独保留,不改写 SKILL.md。

做什么

通过对话编辑任何视频 。 转写,剪接,配色分级,生成叠加动画,烧制字幕——用于说话头,蒙太奇,辅导,出行,采访. 没有预设,没有菜单. 问问题,确认计划,执行,执行,坚持. 生产正确规则很困难;其他一切都是艺术自由.

何时用

官方 description 未单独写出 Use when。按规范,代理会在用户任务与这段 description 的关键词匹配时激活本技能。

代理如何加载

按 Agent Skills 渐进披露:启动时只加载 name 与 description(约 100 token);任务匹配后才读入整份 SKILL.md 正文;scripts/、references/、assets/ 仅在需要时再读。 本文件正文结构:Video Use、Principle、Hard Rules (production correctness — non-negotiable)、Directory layout、Setup、Helpers。

文件分析

文件分析:除 SKILL.md 外,正文还引用了 scripts/,属于带资源的技能包,代理会按需再读这些文件。

官方 description(原文)

Edit any video by conversation. Transcribe, cut, color grade, generate overlay animations, burn subtitles — for talking heads, montages, tutorials, travel, interviews. No presets, no menus. Ask questions, confirm the plan, execute, iterate, persist. Production-correctness rules are hard; everything else is artistic freedom.

Video UsePrincipleHard Rules (production correctness — non-negotiable)Directory layoutSetupHelpersThe processCut craft (techniques)The packed transcript (primary reading view)C0103 (duration: 43.0s, 8 phrases)Editor sub-agent brief (for multi-take selection)Color grade (when requested)

来源分类:skills.sh agent-skill

SKILL.md 与 Agent 调用

官方规范 ↗
name
video-use
description
Edit any video by conversation. Transcribe, cut, color grade, generate overlay animations, burn subtitles — for talking heads, montages, tutorials, travel, interviews. No presets, no menus. Ask questions, confirm the plan, execute, iterate, persist. Production-correctness rules are hard; everything else is artistic freedom.
  1. 发现技能客户端向 Agent 提供名称与描述目录。
  2. 匹配与调用用户指定或任务匹配后,载入 SKILL.md 指令。
  3. 按需加载按步骤读取参考文档、使用脚本与素材。

具体调用语法与可用工具以目标 Agent 客户端为准。 查看调用机制说明 ↗

安装这个技能

Skills CLI ↗

先选择目标 Agent 和安装范围,保留技能包的附属文件,安装后检查客户端能否发现该技能。

该技能引用了附属文件,请从来源获取完整目录;仅复制 SKILL.md 可能缺少依赖。

交给 Agent 安装

复制安装指令给支持 Agent Skills 的代理,确认其中的目标目录与客户端匹配。

把 Agent Skill「video-use」安装到我的项目:SKILL.md 原文与官方 description 见 https://zicq.com/zh/skills/skl-86a13e7b5b5e6238-Video-Use.html
请存为 .cursor/skills/video-use/SKILL.md 或 .claude/skills/video-use/SKILL.md,frontmatter 的 name 与 description 保持原样,不要改写。
该技能还带 scripts/、references/、assets/ 等文件,请从 https://github.com/browser-use/video-use 取完整目录,不要只建一个 SKILL.md。

GitHub 完整包 ↗

终端安装 · Skills CLI

需要 Node.js 与 npx。先查看仓库技能列表,确认实际名称。

npx skills add 'https://github.com/browser-use/video-use' --list

npx skills add 'https://github.com/browser-use/video-use' --skill 'video-use'

CLI 会交互选择目标 Agent,默认安装到项目;用户级安装使用 -g。先通过查看命令核对仓库内容,再用 npx skills list 检查已安装技能。

阅读排版
--- name: video-use description: Edit any video by conversation. Transcribe, cut, color grade, generate overlay animations, burn subtitles — for talking heads, montages, tutorials, travel, interviews. No presets, no menus. Ask questions, confirm the plan, execute, iterate, persist. Production-correctness rules are hard; everything else is artistic freedom. --- # Video Use ## Principle 1. **LLM reasons from raw transcript + on-demand visuals.** The only derived artifact that earns its keep is a packed phrase-level transcript (`takes_packed.md`). Everything else — filler tagging, retake detection, shot classification, emphasis scoring — you derive at decision time. 2. **Audio is primary, visuals follow.** Cut candidates come from speech boundaries and silence gaps. Drill into visuals only at decision points. 3. **Ask → confirm → execute → iterate → persist.** Never touch the cut until the user has confirmed the strategy in plain English. 4. **Generalize.** Do not assume what kind of video this is. Look at the material, ask the user, then edit. 5. **Artistic freedom is the default.** Every specific value, preset, font, color, duration, pitch structure, and technique in this document is a *worked example* from one proven video — not a mandate. Read them to understand what's possible and why each worked. Then make your own taste calls based on what the material actually is and what the user actually wants. **The only things you MUST do are in the Hard Rules section below.** Everything else is yours. 6. **Invent freely.** If the material calls for a technique not described here — split-screen, picture-in-picture, lower-third identity cards, reaction cuts, speed ramps, freeze frames, crossfades, match cuts, L-cuts, J-cuts, speed ramps over breath, whatever — build it. The helpers are ffmpeg and PIL. They can do anything the format supports. Do not wait for permission. 7. **Verify your own output before showing it to the user.** If you wouldn't ship it, don't present it. ## Hard Rules (production correctness — non-negotiable) These are the things where deviation produces silent failures or broken output. They are not taste, they are correctness. Memorize them. 1. **Subtitles are applied LAST in the filter chain**, after every overlay. Otherwise overlays hide captions. Silent failure. 2. **Per-segment extract → lossless `-c copy` concat**, not single-pass filtergraph. Otherwise you double-encode every segment when overlays are added. 3. **30ms audio fades at every segment boundary** (`afade=t=in:st=0:d=0.03,afade=t=out:st={dur-0.03}:d=0.03`). Otherwise audible pops at every cut. 4. **Overlays use `setpts=PTS-STARTPTS+T/TB`** to shift the overlay's frame 0 to its window start. Otherwise you see the middle of the animation during the overlay window. 5. **Master SRT uses output-timeline offsets**: `output_time = word.start - segment_start + segment_offset`. Otherwise captions misalign after segment concat. 6. **Never cut inside a word.** Snap every cut edge to a word boundary from the Scribe transcript. 7. **Pad every cut edge.** Working window: 30–200ms. Scribe timestamps drift 50–100ms — padding absorbs the drift. Tighter for fast-paced, looser for cinematic. 8. **Word-level verbatim ASR only.** Never SRT/phrase mode (loses sub-second gap data). Never normalized fillers (loses editorial signal). 9. **Cache transcripts per source.** Never re-transcribe unless the source file itself changed. 10. **Parallel sub-agents for multiple animations.** Never sequential. Spawn N at once via the `Agent` tool; total wall time ≈ slowest one. 11. **Strategy confirmation before execution.** Never touch the cut until the user has approved the plain-English plan. 12. **All session outputs in `/edit/`.** Never write inside the `video-use/` project directory. Everything else in this document is a worked example. Deviate whenever the material calls for it. ## Directory layout The skill lives in `video-use/`. User footage lives wherever they put it. All session outputs go into `/edit/`. ``` / ├── └── edit/ ├── project.md ← memory; appended every session ├── takes_packed.md ← phrase-level transcripts, the LLM's primary reading view ├── edl.json ← cut decisions ├── transcripts/.json ← cached raw Scribe JSON ├── animations/slot_/ ← per-animation source + render + reasoning ├── clips_graded/ ← per-segment extracts with grade + fades ├── master.srt ← output-timeline subtitles ├── downloads/ ← yt-dlp outputs ├── verify/ ← debug frames / timeline PNGs ├── preview.mp4 └── final.mp4 ``` ## Setup First-time install lives in `install.md` (clone, deps, ffmpeg, skill registration, API key). Don't re-run it every session; on cold start just verify: - `ELEVENLABS_API_KEY` resolves — either in the environment or in `.env` at the video-use repo root. If missing, ask the user to paste one and write it to `.env` (never to the user's ``). - `ffmpeg` + `ffprobe` on PATH. - Python deps installed (`uv sync` or `pip install -e .` inside the repo). - Node.js + npm available if the session needs HyperFrames or Remotion slots. HyperFrames currently requires Node.js 22+. - `yt-dlp`, HyperFrames, Remotion, Manim installed only on first use. - First-use animation setup happens inside the slot directory, never at the video-use repo root. HyperFrames can be invoked with `npx --yes hyperframes ...`; Remotion can be scaffolded with `npx create-video@latest` or installed as a project-local dependency before using its `remotion render` command. - This skill vendors `skills/manim-video/`. Read its SKILL.md when building a Manim slot. Helpers (`helpers/transcribe.py`, `helpers/render.py`, etc.) live alongside this SKILL.md. Resolve their paths relative to the directory containing this file — the skill is typically symlinked at `~/.claude/skills/video-use/` or `~/.codex/skills/video-use/`. ## Helpers - **`transcribe.py `** — single-file Scribe call. `--num-speakers N` optional. Cached. - **`transcribe_batch.py `** — 4-worker parallel transcription. Use for multi-take. - **`pack_transcripts.py --edit-dir `** — `transcripts/*.json` → `takes_packed.md` (phrase-level, break on silence ≥ 0.5s). - **`timeline_view.py `** — filmstrip + waveform PNG. On-demand visual drill-down. **Not a scan tool** — use it at decision points, not constantly. - **`render.py -o `** — per-segment extract → concat → overlays (PTS-shifted) → subtitles LAST. `--preview` for 720p fast. `--build-subtitles` to generate master.srt inline. - **`grade.py -o `** — ffmpeg filter chain grade. Presets + `--filter ''` for custom. For animations, create `/animations/slot_/` with `Bash` and spawn a sub-agent via the `Agent` tool. ## The process 1. **Inventory.** `ffprobe` every source. `transcribe_batch.py` on the directory. `pack_transcripts.py` to produce `takes_packed.md`. Sample one or two `timeline_view`s for a visual first impression. 2. **Pre-scan for problems.** One pass over `takes_packed.md` to note verbal slips, obvious mis-speaks, or phrasings to avoid. Plain list, feed into the editor brief. 3. **Converse.** Describe what you see in plain English. Ask questions *shaped by the material*. Collect: content type, target length/aspect, aesthetic/brand direction, pacing feel, must-preserve moments, must-cut moments, animation and grade preferences, subtitle needs. Do not use a fixed checklist — the right questions are different every time. 4. **Propose strategy.** 4–8 sentences: shape, take choices, cut direction, animation plan, grade direction, subtitle style, length estimate. **Wait for confirmation.** 5. **Execute.** Produce `edl.json` via the editor sub-agent brief. Drill into `timeline_view` at ambiguous moments. Build animations in parallel sub-agents. Apply grade per-segment. Compose via `render.py`. 6. **Preview.** `render.py --preview`. 7. **Self-eval (before showing the user).** Run `timeline_view` on the **rendered output** (not the sources) at every cut boundary (±1.5s window). Check each image for: - Visual discontinuity / flash / jump at the cut - Waveform spike at the boundary (audio pop that slipped past the 30ms fade) - Subtitle hidden behind an overlay (Rule 1 violation) - Overlay misaligned or showing wrong frames (Rule 4 violation) Also sample: first 2s, last 2s, and 2–3 mid-points — check grade consistency, subtitle readability, overall coherence. Run `ffprobe` on the output to verify duration matches the EDL expectation. If anything fails: fix → re-render → re-eval. **Cap at 3 self-eval passes** — if issues remain after 3, flag them to the user rather than looping forever. Only present the preview once the self-eval passes. 8. **Iterate + persist.** Natural-language feedback, re-plan, re-render. Never re-transcribe. Final render on confirmation. Append to `project.md`. ## Cut craft (techniques) - **Audio-first.** Candidate cuts from word boundaries and silence gaps. - **Preserve peaks.** Laughs, punchlines, emphasis beats. Extend past punchlines to include reactions — the laugh IS the beat. - **Speaker handoffs** benefit from air between utterances. Common values: 400–600ms. Less for fast-paced, more for cinematic. Taste call. - **Audio events as signals.** `(laughs)`, `(sighs)`, `(applause)` mark beats. Extend past them. - **Silence gaps are cut candidates.** Silences ≥400ms are usually the cleanest. 150–400ms phrase boundaries are usable with a visual check. <150ms is unsafe (mid-phrase). - **Example cut padding** (the launch video shipped with this): 50ms before the first kept word, 80ms after the last. Tighter for montage energy, looser for documentary. Stay in the 30–200ms working window (Hard Rule 7). - **Never reason audio and video independently.** Every cut must work on both tracks. ## The packed transcript (primary reading view) `pack_transcripts.py` reads all `transcripts/*.json` and produces one markdown file where each take is a list of phrase-level lines, each prefixed with its `[start-end]` time range. Phrases break on any silence ≥ 0.5s OR speaker change. This is the artifact the editor sub-agent reads to pick cuts — it gives word-boundary precision from text alone at 1/10 the tokens of raw JSON. Example line: ``` ## C0103 (duration: 43.0s, 8 phrases) [002.52-005.36] S0 Ninety percent of what a web agent does is completely wasted. [006.08-006.74] S0 We fixed this. ``` ## Editor sub-agent brief (for multi-take selection) When the task is "pick the best take of each beat across many clips," spawn a dedicated sub-agent with a brief shaped like this. The structure is load-bearing; the pitch-shape example is not. ``` You are editing a video. Pick the best take of each beat and assemble them chronologically by beat, not by source clip order. INPUTS: - takes_packed.md (time-annotated phrase-level transcripts of all takes) - Product/narrative context: <2 sentences from the user> - Speaker(s): - Expected structure: - Verbal slips to avoid: - Target runtime: Common structural archetypes (pick, adapt, or invent): - Tech launch / demo: HOOK → PROBLEM → SOLUTION → BENEFIT → EXAMPLE → CTA - Tutorial: INTRO → SETUP → STEPS → GOTCHAS → RECAP - Interview: (QUESTION → ANSWER → FOLLOWUP) repeat - Travel / event: ARRIVAL → HIGHLIGHTS → QUIET MOMENTS → DEPARTURE - Documentary: THESIS → EVIDENCE → COUNTERPOINT → CONCLUSION - Music / performance: INTRO → VERSE → CHORUS → BRIDGE → OUTRO - Or invent your own. RULES: - Start/end times must fall on word boundaries from the transcript. - Pad cut boundaries (working window 30–200ms). - Prefer silences ≥ 400ms as cut targets. - Unavoidable slips are kept if no better take exists. Note them in "reason". - If over budget, revise: drop a beat or trim tails. Report total and self-correct. OUTPUT (JSON array, no prose): [{"source": "C0103", "start": 2.42, "end": 6.85, "beat": "HOOK", "quote": "...", "reason": "..."}, ...] Return the final EDL and a one-line total runtime check. ``` ## Color grade (when requested) Your job is to **reason about the image**, not apply a preset. Look at a frame (via `timeline_view`), decide what's wrong, adjust one thing, look again. Mental model is ASC CDL. Per channel: `out = (in * slope + offset) ** power`, then global saturation. `slope` → highlights, `offset` → shadows, `power` → midtones. **Example filter chains** (`grade.py` has `--list-presets`; use them as starting points or mix your own): - **`warm_cinematic`** — retro/technical, subtle teal/orange split, desaturated. Shipped in a real launch video. Safe for talking heads. - **`neutral_punch`** — minimal corrective: contrast bump + gentle S-curve. No hue shifts. - **`none`** — straight copy. Default when the user hasn't asked. For anything else — portraiture, nature, product, music video, documentary — invent your own chain. `grade.py --filter ''` accepts any filter string. Hard rules: apply **per-segment during extraction** (not post-concat, which re-encodes twice). Never go aggressive without testing skin tones. ## Subtitles (when requested) Subtitles have three dimensions worth reasoning about: **chunking** (1/2/3/sentence per line), **case** (UPPER/Title/Natural), and **placement** (margin from bottom). The right combo depends on content. **Worked styles** — pick, adapt, or invent: **`bold-overlay`** — short-form tech launch, fast-paced social. 2-word chunks, UPPERCASE, break on punctuation, Helvetica 18 Bold, white-on-outline, `MarginV=35`. `render.py` ships with this as `SUB_FORCE_STYLE`. ``` FontName=Helvetica,FontSize=18,Bold=1, PrimaryColour=&H00FFFFFF,OutlineColour=&H00000000,BackColour=&H00000000, BorderStyle=1,Outline=2,Shadow=0, Alignment=2,MarginV=35 ``` **`natural-sentence`** (if you invent this mode) — narrative, documentary, education. 4–7 word chunks, sentence case, break on natural pauses, `MarginV=60–80`, larger font for readability, slightly wider max-width. No shipped force_style — design one if you need it. Invent a third style if neither fits. Hard rules: subtitles LAST (Rule 1), output-timeline offsets (Rule 5). ## Animations (when requested) Animations match the content and the brand. **Get the palette, font, and visual language from the conversation** — never assume a default. If the user hasn't told you, propose a palette in the strategy phase and wait for confirmation before building anything. **Tool options:** Pick the engine per animation slot. Do not default to Remotion just because the animation is web-adjacent. - **HyperFrames** — Browser-native HTML/CSS/GSAP video compositions: product UI motion, website-to-video or mockup-to-video captures, kinetic typography, landing-page/storyboard promos, data-driven UI states, transparent WebM overlays, and clips that need deterministic frame capture plus HyperFrames lint/validate/render checks. Best when the animation should be authored and verified like a web composition instead of a React component tree. - **Remotion** — React/CSS compositions with component state, reusable React primitives, or an existing Remotion brand system. Best when the user specifically asks for React/Remotion or when React composition is the simpler authoring model. - **Manim** — formal diagrams, state machines, equation derivations, graph morphs. Read `skills/manim-video/SKILL.md` and its references for depth. - **PIL + PNG sequence + ffmpeg** — simple overlay cards: counters, typewriter text, single bar reveals, progressive draws. Fast to iterate, any aesthetic you want. The launch video used this. For HyperFrames slots, scaffold the slot inside `edit/animations/slot_/` with `npx --yes hyperframes init . --example blank --non-interactive --skip-skills`, build the HTML composition there, run the HyperFrames checks that fit the slot (`lint`, `validate`, and a draft render when practical), then produce the final overlay video with `npx --yes hyperframes render . -o render.mp4` or `--format webm -o render.webm` when alpha is required. Point the EDL overlay `file` at the actual rendered path. For Remotion slots, keep the Remotion project isolated inside the same slot directory, scaffold with `npx create-video@latest` or install Remotion locally there, render the composition to `render.mp4` with the project-local `remotion render` command, and verify duration and dimensions with `ffprobe`. None is mandatory. Invent hybrids if useful (e.g., PIL background with a HyperFrames or Remotion layer on top). **Duration rules of thumb, context-dependent:** - **Sync-to-narration explanations.** A viewer needs to parse the content at 1×. Rough floor 3s, typical 5–7s for simple cards, 8–14s for complex diagrams. The launch video shipped at 5–7s per simple card. - **Beat-synced accents** (music video, fast montage). 0.5–2s is fine — they're visual accents, not information. The "readable at 1×" rule becomes *"recognizable at 1×"*, not *"fully parseable."* - **Hold the final frame ≥ 1s** before the cut (universal). - **Over voiceover:** total duration ≥ `narration_length + 1s` (universal). - **Never parallel-reveal independent elements** — the eye can't track two new things at once. One thing, pause, next thing. **Animation payoff timing (rule for sync-to-narration):** get the payoff word's timestamp. Start the overlay `reveal_duration` seconds earlier so the landing frame coincides with the spoken payoff word. Without this sync the animation feels disconnected. **Easing** (universal — never `linear`, it looks robotic): ```python def ease_out_cubic(t): return 1 - (1 - t) ** 3 def ease_in_out_cubic(t): if t < 0.5: return 4 * t ** 3 return 1 - (-2 * t + 2) ** 3 / 2 ``` `ease_out_cubic` for single reveals (slow landing). `ease_in_out_cubic` for continuous draws. **Typing text anchor trick:** center on the FULL string's width, not the partial-string width — otherwise text slides left during reveal. **Example palette** (the launch video — one aesthetic among infinite): - Background `(10, 10, 10)` near-black - Accent `#FF5A00` / `(255, 90, 0)` orange - Labels `(110, 110, 110)` dim gray - Font: Menlo Bold at `/System/Library/Fonts/Menlo.ttc` (index 1) - ≤ 2 accent colors, ~40% empty space, minimal chrome - Result: terminal / retro tech feel This is one style. If the brand is warm and serif, use that. If it's colorful and playful, use that. If the user handed you a style guide, follow it. If they didn't, propose one and confirm. **Parallel sub-agent brief** — each animation is one sub-agent spawned via the `Agent` tool. Each prompt is self-contained (sub-agents have no parent context). Include: 1. One-sentence goal: *"Build ONE animation: [spec]. Nothing else."* 2. Absolute output path (`/animations/slot_/render.mp4`) 3. Exact technical spec: resolution, fps, codec, pix_fmt, CRF, duration 4. Style palette as concrete values (RGB tuples, hex, or reference to a design system) 5. Font path with index 6. Frame-by-frame timeline (what happens when, with easing) 7. Anti-list ("no chrome, no extras, no titles unless specified") 8. Code pattern reference (copy helpers inline, don't import across slots) 9. Deliverable checklist (script, render, verify duration via ffprobe, report) 10. **"Do not ask questions. If anything is ambiguous, pick the most obvious interpretation and proceed."** One sub-agent = one file (unique filenames, parallel agents don't overwrite each other). ## Output spec Match the source unless the user asked for something specific. Common targets: `1920×1080@24` cinematic, `1920×1080@30` screen content, `1080×1920@30` vertical social, `3840×2160@24` 4K cinema, `1080×1080@30` square. `render.py` defaults the scale to 1080p from any source; pass `--filter` or edit the extract command for other targets. Worth asking the user which delivery format matters. ## EDL format ```json { "version": 1, "sources": {"C0103": "/abs/path/C0103.MP4", "C0108": "/abs/path/C0108.MP4"}, "ranges": [ {"source": "C0103", "start": 2.42, "end": 6.85, "beat": "HOOK", "quote": "...", "reason": "Cleanest delivery, stops before slip at 38.46."}, {"source": "C0108", "start": 14.30, "end": 28.90, "beat": "SOLUTION", "quote": "...", "reason": "Only take without the false start."} ], "grade": "warm_cinematic", "overlays": [ {"file": "edit/animations/slot_1/render.mp4", "start_in_output": 0.0, "duration": 5.0} ], "subtitles": "edit/master.srt", "total_duration_s": 87.4 } ``` `grade` is a preset name or raw ffmpeg filter. `overlays` are rendered animation clips. `subtitles` is optional and applied LAST. ## Memory — `project.md` Append one section per session at `/project.md`: ```markdown ## Session N — YYYY-MM-DD **Strategy:** one paragraph describing the approach **Decisions:** take choices, cuts, grades, animations + why **Reasoning log:** one-line rationale for non-obvious decisions **Outstanding:** deferred items ``` On startup, read `project.md` if it exists and summarize the last session in one sentence before asking whether to continue. ## Anti-patterns Things that consistently fail regardless of style: - **Hierarchical pre-computed codec formats** with USABILITY / tone tags / shot layers. Over-engineering. Derive from the transcript at decision time. - **Hand-tuned moment-scoring functions.** The LLM picks better than any heuristic you'll write. - **Whisper SRT / phrase-level output.** Loses sub-second gap data. Always word-level verbatim. - **Running Whisper locally on CPU.** Slow and it normalizes fillers. Use hosted Scribe. - **Burning subtitles into base before compositing overlays.** Overlays hide them. (Hard Rule 1.) - **Single-pass filtergraph when you have overlays.** Double re-encodes. Use per-segment extract → concat. - **Linear animation easing.** Looks robotic. Always cubic. - **Hard audio cuts at segment boundaries.** Audible pops. (Hard Rule 3.) - **Typing text centered on the partial string.** Text slides left as it grows. - **Sequential sub-agents for multiple animations.** Always parallel. - **Editing before confirming the strategy.** Never. - **Re-transcribing cached sources.** Immutable outputs of immutable inputs. - **Assuming what kind of video it is.** Look first, ask second, edit last.

相关技能

媒体内容

Nano Banana Pro

以纳米·香蕉Pro(Gemini 3 Pro Image)来生成/编辑图像. 用于创建/修改请求,包括编辑。 支持文本到图像+图像到图像; 1K/2K/4K; 使用 -- input-image.

媒体内容

Youtube Watcher

从YouTube视频中获取并读取文字记录. 需要汇总视频时使用,回答有关视频内容的问题,或者从中提取信息.

媒体内容

Atxp

Access ATXP支付API工具用于网页搜索,AI图像生成,音乐创建,视频生成,X/Twitter搜索,电子邮件,代理账户管理. 当用户需要实时网络搜索时使用,AI生成的媒体(图像,音乐,视频),X/Twitter搜索,发送/接收电子邮件,或创建和资助代理账户. 需要通过“ …

媒体内容

Youtube

YouTube Data API与管理的OAuth的集成. 搜索视频,管理播放列表,访问频道数据,并与评论互动. 用户想与YouTube互动时使用此技能. 对于其他第三方应用,使用api-gateway技能(https://clawhub.ai/byungkyu/api-gate…