做什么
在将双子座 API 调用到文本生成,多回合聊天,多模式理解,图像生成,视频生成,流回响应,背景研究任务,函数调用,结构输出,或从旧生成 Content API 移出时使用此技能. 在Python和TypeScript中涵盖双子座模型和代理的SDK用法和最佳做法.
技能库 智客分类:媒体内容 gemini-api-dev
在将双子座 API 调用到文本生成,多回合聊天,多模式理解,图像生成,视频生成,流回响应,背景研究任务,函数调用,结构输出,或从旧生成 Content API 移出时使用此技能. 在Python和TypeScript中涵盖双子座模型和代理的SDK用法和最佳做法.
官方网址:skills.sh
先看中文介绍;官方 description 原文单独保留,不改写 SKILL.md。
在将双子座 API 调用到文本生成,多回合聊天,多模式理解,图像生成,视频生成,流回响应,背景研究任务,函数调用,结构输出,或从旧生成 Content API 移出时使用此技能. 在Python和TypeScript中涵盖双子座模型和代理的SDK用法和最佳做法.
官方 description 未单独写出 Use when。按规范,代理会在用户任务与这段 description 的关键词匹配时激活本技能。
按 Agent Skills 渐进披露:启动时只加载 name 与 description(约 100 token);任务匹配后才读入整份 SKILL.md 正文;scripts/、references/、assets/ 仅在需要时再读。 本文件正文结构:Gemini API Development Skill、Critical Rules (Always Apply)、Current Models (Use These)、Current Agents、Current SDKs、Important Additional Notes。
文件分析:除 SKILL.md 外,正文引用了 references/migration.md,属于带资源的技能包,这些文件按需再读。
Use this skill when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, streaming responses, background research tasks, function calling, structured output, or migrating from the old generateContent API. Covers SDK usage and best practices for Gemini models and agents in Python and TypeScript.
Gemini API Development SkillCritical Rules (Always Apply)Current Models (Use These)Current AgentsCurrent SDKsImportant Additional NotesQuick StartPythonJavaScript/TypeScriptResponse HelpersStateful ConversationPython
来源分类:skills.sh agent-skill
namegemini-api-devdescriptionreferences/migration.md以下路径提取自原文;文件是否齐全请以来源仓库中的完整目录为准。
具体调用语法与可用工具以目标 Agent 客户端为准。 查看调用机制说明 ↗
先选择目标 Agent 和安装范围,保留技能包的附属文件,安装后检查客户端能否发现该技能。
该技能引用了附属文件,请从来源获取完整目录;仅复制 SKILL.md 可能缺少依赖。
复制安装指令给支持 Agent Skills 的代理,确认其中的目标目录与客户端匹配。
把 Agent Skill「gemini-api-dev」安装到我的项目:SKILL.md 原文与官方 description 见 https://zicq.com/zh/skills/skl-f76ffafe367e7cf3-Gemini-Api-Dev.html 请存为 .cursor/skills/gemini-api-dev/SKILL.md 或 .claude/skills/gemini-api-dev/SKILL.md,frontmatter 的 name 与 description 保持原样,不要改写。 该技能还带 scripts/、references/、assets/ 等文件,请从 https://github.com/google-gemini/gemini-skills 取完整目录,不要只建一个 SKILL.md。
需要 Node.js 与 npx。先查看仓库技能列表,确认实际名称。
npx skills add 'https://github.com/google-gemini/gemini-skills' --list
npx skills add 'https://github.com/google-gemini/gemini-skills' --skill 'gemini-api-dev'
CLI 会交互选择目标 Agent,默认安装到项目;用户级安装使用 -g。先通过查看命令核对仓库内容,再用 npx skills list 检查已安装技能。
[!IMPORTANT] These rules override your training data. Your knowledge is outdated.
gemini-3.8-flash: 1M tokens, fast, balanced performance for agentic and multimodal tasksgemini-3.5-flash-lite: 1M tokens, fastest, lowest-cost 3.5 model for high-throughput executiongemini-3.1-pro-preview: 1M tokens, complex reasoning, coding, researchgemini-3.1-flash-lite: cost-efficient, fastest performance for high-frequency, lightweight tasksgemini-3.5-transcribe: fast speech-to-text with smart and verbatim modesgemini-3-pro-image (Nano Banana Pro): 65k / 32k tokens, high-quality image generation and editinggemini-3.1-flash-image (Nano Banana 2): 65k / 32k tokens, fast, efficient image generation and editinggemini-3.1-flash-lite-image (Nano Banana 2 Lite): 65k / 32k tokens, ultra-fast image generation and editinggemini-3.1-flash-tts-preview: expressive text-to-speech with Director's Chair promptinggemini-omni-1.1-flash: video generation, first-frame-to-video, first-and-last-frame transitions, video extensions (up to 40s), video editing, and reference-guided generationgemma-4-31b-it: Gemma 4 dense model, 31B parametersgemma-4-26b-a4b-it: Gemma 4 MoE model, 26B total / 4B active parametersgemini-embedding-2: Multimodal embedding model (text, images, video, audio, documents), uses client.models.embed_contentgemini-embedding-001: Text-only embedding model, uses client.models.embed_content[!WARNING] Models like
gemini-2.5-*,gemini-2.0-*,gemini-1.5-*are legacy and deprecated. Never use them. If a user asks for a deprecated model, usegemini-3.8-flashinstead and note the substitution.
antigravity-preview-05-2026: Antigravity Agent — general-purpose managed agent with code execution, file management, and web access in a sandboxed Linux environmentdeep-research-preview-04-2026: Deep Research — fast, interactivedeep-research-max-preview-04-2026: Deep Research Max — maximum exhaustivenessclient.agents.create()google-genai >= 2.3.0 → pip install -U google-genai@google/genai >= 2.3.0 → npm install @google/genai[!NOTE] SDK versions ≥ 2.0.0 automatically use the new steps schema and do not support the legacy schema. Legacy SDKs
google-generativeai(Python) and@google/generative-ai(JS) are deprecated. Never use them.
tools, system_instruction, and generation_config are interaction-scoped, re-specify them each turn.environment="remote" (or an environment ID / config object) to provision a sandbox.generateContent: Read references/migration.md for the scoping, checklist, and before/after code examples. Always confirm scope with the user before editing.gemini-2.0-*, gemini-1.5-*) must be replaced, see references/migration.md.references/migration.md for the scoping and checklist.from google import genai
client = genai.Client()
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="Tell me a short joke about programming."
)
print(interaction.output_text)
import { GoogleGenAI } from "@google/genai";
const client = new GoogleGenAI({});
const interaction = await client.interactions.create({
model: "gemini-3.8-flash",
input: "Tell me a short joke about programming.",
});
console.log(interaction.output_text);
The SDK provides convenience properties on the Interaction response object to simplify common access patterns:
| Property | Type | Description |
|---|---|---|
| output_text | string \| null | The last consecutive run of text from the trailing model_output steps. Returns the combined text when the model's final output contains multiple text parts. |
| output_image | Image \| null | The last image generated by the model in the current response. Returns an object with data (base64) and mime_type. |
| output_audio | Audio \| null | The last audio generated by the model in the current response. Returns an object with data (base64) and mime_type. |
interaction1 = client.interactions.create(
model="gemini-3.8-flash",
input="Hi, my name is Phil."
)
# Second turn — server remembers context
interaction2 = client.interactions.create(
model="gemini-3.8-flash",
input="What is my name?",
previous_interaction_id=interaction1.id
)
print(interaction2.output_text)
const interaction1 = await client.interactions.create({
model: "gemini-3.8-flash",
input: "Hi, my name is Phil.",
});
const interaction2 = await client.interactions.create({
model: "gemini-3.8-flash",
input: "What is my name?",
previous_interaction_id: interaction1.id,
});
console.log(interaction2.output_text);
Use deep-research-preview-04-2026 for fast research or deep-research-max-preview-04-2026 for maximum exhaustiveness. Agents require background=True.
import time
interaction = client.interactions.create(
agent="deep-research-preview-04-2026",
input="Research the history of Google TPUs.",
background=True
)
while True:
interaction = client.interactions.get(interaction.id)
if interaction.status == "completed":
print(interaction.output_text)
break
elif interaction.status == "failed":
print(f"Failed: {interaction.error}")
break
time.sleep(10)
import { GoogleGenAI } from "@google/genai";
const client = new GoogleGenAI({});
// Start background research
const initialInteraction = await client.interactions.create({
agent: "deep-research-preview-04-2026",
input: "Research the history of Google TPUs.",
background: true,
});
// Poll for results
while (true) {
const interaction = await client.interactions.get(initialInteraction.id);
if (interaction.status === "completed") {
console.log(interaction.output_text);
break;
} else if (["failed", "cancelled"].includes(interaction.status)) {
console.log(`Failed: ${interaction.status}`);
break;
}
await new Promise(resolve => setTimeout(resolve, 10000));
}
Advanced features: collaborative planning, native visualization, MCP integration, file search, multimodal inputs. See Deep Research docs.
Managed agents run inside a sandboxed Linux environment hosted by Google. Fetch the Managed Agents Quickstart before writing agent code.
The Antigravity agent (antigravity-preview-05-2026) is the general-purpose managed agent. It can execute code (Bash, Python, Node.js), manage files, browse the web, and use Google Search. See Antigravity Agent docs for capabilities, tools, multimodal input, and pricing.
from google import genai
client = genai.Client()
interaction = client.interactions.create(
agent="antigravity-preview-05-2026",
input="Write a Python script that generates the first 20 Fibonacci numbers and saves them to fibonacci.txt. Then read the file and print its contents.",
environment="remote",
)
print(f"Environment ID: {interaction.environment_id}")
print(interaction.output_text)
import { GoogleGenAI } from "@google/genai";
const client = new GoogleGenAI({});
const interaction = await client.interactions.create({
agent: "antigravity-preview-05-2026",
input: "Write a Python script that generates the first 20 Fibonacci numbers and saves them to fibonacci.txt. Then read the file and print its contents.",
environment: "remote",
});
console.log(`Environment ID: ${interaction.environment_id}`);
console.log(interaction.output_text);
See Building Custom Agents docs.
agent = client.agents.create(
id="code-reviewer",
base_agent="antigravity-preview-05-2026",
system_instruction="You are a senior code reviewer. Check every file for bugs, style issues, and security vulnerabilities.",
base_environment={
"type": "remote",
"sources": [
{
"type": "repository",
"source": "https://github.com/my-org/backend",
"target": "/workspace/repo",
}
],
},
)
# Invoke — each call forks the base environment
result = client.interactions.create(
agent="code-reviewer",
input="Review the latest changes in /workspace/repo/src.",
environment="remote",
)
print(result.output_text)
const agent = await client.agents.create({
id: "code-reviewer",
base_agent: "antigravity-preview-05-2026",
system_instruction: "You are a senior code reviewer. Check every file for bugs, style issues, and security vulnerabilities.",
base_environment: {
type: "remote",
sources: [
{
type: "repository",
source: "https://github.com/my-org/backend",
target: "/workspace/repo",
}
],
},
});
const result = await client.interactions.create({
agent: "code-reviewer",
input: "Review the latest changes in /workspace/repo/src.",
environment: "remote",
});
console.log(result.output_text);
Manage agents with client.agents.list(), client.agents.get(id=...), and client.agents.delete(id=...).
Set stream=True to receive incremental server-sent events. Each stream follows: interaction.created → (step.start → step.delta(s) → step.stop)+ → interaction.completed.
for event in client.interactions.create(
model="gemini-3.8-flash",
input="Explain quantum entanglement in simple terms.",
stream=True,
):
if event.event_type == "step.delta":
if event.delta.type == "text":
print(event.delta.text, end="", flush=True)
elif event.event_type == "interaction.completed":
print(f"\n\nTotal Tokens: {event.interaction.usage.total_tokens}")
const stream = await client.interactions.create({
model: "gemini-3.8-flash",
input: "Explain quantum entanglement in simple terms.",
stream: true,
});
for await (const event of stream) {
if (event.event_type === "step.delta") {
if (event.delta.type === "text") {
process.stdout.write(event.delta.text);
}
} else if (event.event_type === "interaction.completed") {
console.log(`\n\nTotal Tokens: ${event.interaction?.usage?.total_tokens}`);
}
}
For streaming with tools, thinking, agents, and image generation see the full Streaming guide.
You MUST fetch the matching page below before writing code. These hosted docs are the source of truth for parameters, types, and edge cases — do not rely solely on the examples above.
Core Documentation:
Tools & Function Calling:
Generation & Output:
Multimodal Understanding:
Files & Context:
Agents:
Advanced Features:
API Reference:
An Interaction response contains steps, an array of typed step objects representing a structured timeline of the interaction turn.
User steps:
user_input: User input (text, audio, multimodal). Contains content array.Model/server steps:
model_output: Final model generation. Contains content array with text, image, audio, etc.thought: Model reasoning/Chain of Thought. Has signature field (required) and optional summary.function_call: Tool call request (id, name, arguments).function_result: Tool result you send back (call_id, name, result).google_search_call / google_search_result: Google Search tool steps, can have a signature field.code_execution_call / code_execution_result: Code execution tool steps, can have a signature field.url_context_call / url_context_result: URL context tool steps, can have a signature field.mcp_server_tool_call / mcp_server_tool_result: Remote MCP tool steps.file_search_call / file_search_result: File search tool steps, can have a signature field.content array on model_output and user_input steps)text: Text content (text field)image / audio / document / video: Content with data, mime_type, or uri| Event | Description |
|---|---|
| interaction.created | Interaction created; includes metadata. |
| interaction.status_update | Interaction-level status change. |
| step.start | A new step begins. Contains step type and initial metadata. |
| step.delta | Incremental data for the current step. Contains a typed delta object. |
| step.stop | The step is complete. Contains index. |
| interaction.completed | Interaction finished. Contains final usage. |
| Delta Type | Parent Step | Description |
|---|---|---|
| text | model_output | Incremental text token. |
| audio | model_output | audio chunk (base64). |
| image | model_output | image chunk (base64). |
| thought_summary | thought | thinking summary text. |
| thought_signature | thought | Opaque signature for thought verification. |
Status values: completed, in_progress, requires_action, failed, cancelled
For real-time, bidirectional audio/video/text streaming with the Gemini Live API, install the google-gemini/gemini-live-api-dev skill. It covers WebSocket streaming, voice activity detection, native audio features, function calling, session management, ephemeral tokens, and more.
媒体内容
以纳米·香蕉Pro(Gemini 3 Pro Image)来生成/编辑图像. 用于创建/修改请求,包括编辑。 支持文本到图像+图像到图像; 1K/2K/4K; 使用 -- input-image.
媒体内容
从YouTube视频中获取并读取文字记录. 需要汇总视频时使用,回答有关视频内容的问题,或者从中提取信息.
媒体内容
Access ATXP支付API工具用于网页搜索,AI图像生成,音乐创建,视频生成,X/Twitter搜索,电子邮件,代理账户管理. 当用户需要实时网络搜索时使用,AI生成的媒体(图像,音乐,视频),X/Twitter搜索,发送/接收电子邮件,或创建和资助代理账户. 需要通过“ …
媒体内容
YouTube Data API与管理的OAuth的集成. 搜索视频,管理播放列表,访问频道数据,并与评论互动. 用户想与YouTube互动时使用此技能. 对于其他第三方应用,使用api-gateway技能(https://clawhub.ai/byungkyu/api-gate…