跳到主内容
智客 ZICQ

技能库 智客分类:写作与研究 autoresearch

Autoresearch

任意编程任务的自主迭代实验循环. 引导用户通过定义目标,可计量的度量和范围限制,然后运行一个自动循环的代码修改,测试,测量,并保存/丢弃结果. 受卡多斯自发研究的启发. USE FOR:自主改进,迭代优化,实验循环,汽车研究,性能调谐,自动化实验,登山,自动尝试事物,优化代码,运行实验,自主编码循环. 不使用 : 一发任务、 简单的错误修正、 代码检讨, 或没有可测量的衡量标准的任务 .

3378 安装量

官方网址:skills.sh

技能介绍

先看中文介绍;官方 description 原文单独保留,不改写 SKILL.md。

做什么

任意编程任务的自主迭代实验循环. 引导用户通过定义目标,可计量的度量和范围限制,然后运行一个自动循环的代码修改,测试,测量,并保存/丢弃结果. 受卡多斯自发研究的启发. USE FOR:自主改进,迭代优化,实验循环,汽车研究,性能调谐,自动化实验,登山,自动尝试事物,优化代码,运行实验,自主编码循环. 不使用 : 一发任务、 简单的错误修正、 代码检讨, 或没有可测量的衡量标准的任务 .

何时用

官方 description 未单独写出 Use when。按规范,代理会在用户任务与这段 description 的关键词匹配时激活本技能。

代理如何加载

按 Agent Skills 渐进披露:启动时只加载 name 与 description(约 100 token);任务匹配后才读入整份 SKILL.md 正文;scripts/、references/、assets/ 仅在需要时再读。 本文件正文结构:Autoresearch: Autonomous Iterative Experimentation、Agent Behavior Rules、Phase 1: Setup (Interactive)、1.1 Define the Goal、1.2 Define the Metric、1.3 Define the Scope。

文件分析

文件分析:这是一份仅含 SKILL.md 的指令型技能,代理激活后整份正文进入上下文。

官方 description(原文)

Autonomous iterative experimentation loop for any programming task. Guides the user through defining goals, measurable metrics, and scope constraints, then runs an autonomous loop of code changes, testing, measuring, and keeping/discarding results. Inspired by Karpathy''s autoresearch. USE FOR: autonomous improvement, iterative optimization, experiment loop, auto research, performance tuning, automated experimentation, hill climbing, try things automatically, optimize code, run experiments, autonomous coding loop. DO NOT USE FOR: one-shot tasks, simple bug fixes, code review, or tasks without a measurable metric.

Autoresearch: Autonomous Iterative ExperimentationAgent Behavior RulesPhase 1: Setup (Interactive)1.1 Define the Goal1.2 Define the Metric1.3 Define the Scope1.4 Define Constraints1.5 Define the Experiment Budget (Optional)1.6 Simplicity Criterion1.7 Confirm SetupPhase 2: Branch & BaselinePhase 3: Experiment Loop

兼容:Requires git. The project must be a git repository. Requires terminal access to run commands. · 许可:MIT

来源分类:skills.sh agent-skill

SKILL.md 与 Agent 调用

官方规范 ↗
name
autoresearch
description
Autonomous iterative experimentation loop for any programming task. Guides the user through defining goals, measurable metrics, and scope constraints, then runs an autonomous loop of code changes, testing, measuring, and keeping/discarding results. Inspired by Karpathy''s autoresearch. USE FOR: autonomous improvement, iterative optimization, experiment loop, auto research, performance tuning, automated experimentation, hill climbing, try things automatically, optimize code, run experiments, autonomous coding loop. DO NOT USE FOR: one-shot tasks, simple bug fixes, code review, or tasks without a measurable metric.
compatibility
Requires git. The project must be a git repository. Requires terminal access to run commands.
许可
MIT
  1. 发现技能客户端向 Agent 提供名称与描述目录。
  2. 匹配与调用用户指定或任务匹配后,载入 SKILL.md 指令。
  3. 按需加载按步骤读取参考文档、使用脚本与素材。

具体调用语法与可用工具以目标 Agent 客户端为准。 查看调用机制说明 ↗

安装这个技能

Skills CLI ↗

先选择目标 Agent 和安装范围,保留技能包的附属文件,安装后检查客户端能否发现该技能。

交给 Agent 安装

复制安装指令给支持 Agent Skills 的代理,确认其中的目标目录与客户端匹配。

把 Agent Skill「autoresearch」安装到我的项目:SKILL.md 原文与官方 description 见 https://zicq.com/zh/skills/skl-99207d3815adfa70-Autoresearch.html
请存为 .cursor/skills/autoresearch/SKILL.md 或 .claude/skills/autoresearch/SKILL.md,frontmatter 的 name 与 description 保持原样,不要改写。

GitHub 完整包 ↗

终端安装 · Skills CLI

需要 Node.js 与 npx。先查看仓库技能列表,确认实际名称。

npx skills add 'https://github.com/github/awesome-copilot' --list

npx skills add 'https://github.com/github/awesome-copilot' --skill 'autoresearch'

CLI 会交互选择目标 Agent,默认安装到项目;用户级安装使用 -g。先通过查看命令核对仓库内容,再用 npx skills list 检查已安装技能。

阅读排版
--- name: autoresearch description: 'Autonomous iterative experimentation loop for any programming task. Guides the user through defining goals, measurable metrics, and scope constraints, then runs an autonomous loop of code changes, testing, measuring, and keeping/discarding results. Inspired by Karpathy''s autoresearch. USE FOR: autonomous improvement, iterative optimization, experiment loop, auto research, performance tuning, automated experimentation, hill climbing, try things automatically, optimize code, run experiments, autonomous coding loop. DO NOT USE FOR: one-shot tasks, simple bug fixes, code review, or tasks without a measurable metric.' license: MIT compatibility: Requires git. The project must be a git repository. Requires terminal access to run commands. metadata: author: luiscantero inspired-by: https://github.com/karpathy/autoresearch --- # Autoresearch: Autonomous Iterative Experimentation An autonomous experimentation loop for any programming task. You define the goal and how to measure it; the agent iterates autonomously -- modifying code, running experiments, measuring results, and keeping or discarding changes -- until interrupted. This skill is inspired by [Karpathy's autoresearch](https://github.com/karpathy/autoresearch), generalized from ML training to **any programming task with a measurable outcome**. --- ## Agent Behavior Rules 1. **DO** guide the user through the Setup phase interactively before starting the loop. 2. **DO** establish a baseline measurement before making any changes. 3. **DO** commit every experiment attempt before running it (so it can be reverted cleanly). 4. **DO** keep a results log (TSV) tracking every experiment. 5. **DO** revert changes that do not improve the metric (git reset to last known good). 6. **DO** run autonomously once the loop starts -- never pause to ask "should I continue?". 7. **DO NOT** modify files the user marked as out-of-scope. 8. **DO NOT** skip the measurement step -- every experiment must be measured. 9. **DO NOT** keep changes that regress the metric unless the user explicitly allowed trade-offs. 10. **DO NOT** install new dependencies or make environment changes unless the user approved it. --- ## Phase 1: Setup (Interactive) Before any experimentation begins, work with the user to establish these parameters. Ask the user directly for each item. Do not assume or skip any. ### 1.1 Define the Goal Ask the user: > **What are you trying to improve or optimize?** > > Examples: execution time, memory usage, binary size, test pass rate, code coverage, > API response latency, throughput, error rate, benchmark score, build time, bundle size, > lines of code, cyclomatic complexity, etc. Record the user's answer as the **goal**. ### 1.2 Define the Metric Ask the user: > **How do we measure success? What exact command produces the metric?** > > I need: > 1. **The command** to run (e.g., `dotnet test`, `npm run benchmark`, `time ./build.sh`, `pytest --tb=short`) > 2. **How to extract the metric** from the output (e.g., a regex pattern, a specific line, a JSON field) > 3. **Direction**: Is lower better or higher better? > > Example: "Run `dotnet test --logger trx`, count passing tests. Higher is better." > Example: "Run `hyperfine './my-program'`, extract mean time. Lower is better." Record: - `METRIC_COMMAND`: the command to run - `METRIC_EXTRACTION`: how to extract the numeric metric from output - `METRIC_DIRECTION`: `lower_is_better` or `higher_is_better` ### 1.3 Define the Scope Ask the user: > **Which files or directories am I allowed to modify?** > > And which files are OFF LIMITS (read-only)? Record: - `IN_SCOPE_FILES`: files/dirs the agent may edit - `OUT_OF_SCOPE_FILES`: files/dirs that must not be modified ### 1.4 Define Constraints Ask the user: > **Are there any constraints I should respect?** > > Examples: > - Time budget per experiment (e.g., "each run should take < 2 minutes") > - No new dependencies > - Must keep all existing tests passing > - Must not change the public API > - Must maintain backward compatibility > - VRAM/memory limit > - Code complexity limits (prefer simpler solutions) Record as `CONSTRAINTS`. ### 1.5 Define the Experiment Budget (Optional) Ask the user: > **How many experiments should I run, or should I just keep going until you stop me?** > > You can say a number (e.g., "try 20 experiments") or "unlimited" (I'll run until you interrupt). Record as `MAX_EXPERIMENTS` (number or `unlimited`). ### 1.6 Simplicity Criterion Inform the user of the default simplicity policy: > **Simplicity policy (default):** All else being equal, simpler is better. A small improvement > that adds ugly complexity is not worth it. Removing code while maintaining or improving > the metric is a great outcome. I'll weigh the complexity cost against the improvement > magnitude. Does this policy work for you, or do you want to adjust it? Record any adjustments as `SIMPLICITY_POLICY`. ### 1.7 Confirm Setup Summarize all parameters back to the user in a clear table: | Parameter | Value | | ------------------ | ---------------------------- | | Goal | ... | | Metric command | ... | | Metric extraction | ... | | Direction | lower is better / higher ... | | In-scope files | ... | | Out-of-scope files | ... | | Constraints | ... | | Max experiments | ... | | Simplicity policy | ... | Ask the user to confirm. Do not proceed until confirmed. --- ## Phase 2: Branch & Baseline Once the user confirms: 1. **Create a branch**: Propose a tag based on today's date (e.g., `autoresearch/mar17`). Create the branch: `git checkout -b autoresearch/`. 2. **Read in-scope files**: Read all files that are in scope to build full context of the current state. 3. **Initialize results.tsv**: Create `results.tsv` in the repo root with the header row: ``` experiment commit metric status description ``` Add `results.tsv` and `run.log` to `.git/info/exclude` (append if not already present) so they stay untracked without modifying any tracked files. 4. **Run the baseline**: Execute the metric command on the current unmodified code. Record the result as experiment `0` with status `baseline` in `results.tsv`. 5. **Report baseline** to the user: > Baseline established: **[metric_name] = [value]** > Starting autonomous experimentation loop. --- ## Phase 3: Experiment Loop Run this loop continuously. Do not stop to ask the user. Run until: - `MAX_EXPERIMENTS` is reached, OR - The user manually interrupts ### For each experiment: ``` LOOP: 1. THINK - Analyze previous results and the current code. Generate an experiment hypothesis. Consider: what worked, what didn't, what hasn't been tried. 2. EDIT - Modify the in-scope file(s) to implement the idea. Keep changes focused and minimal per experiment. 3. COMMIT - git add + git commit with a short descriptive message. Format: "experiment: " 4. RUN - Execute the metric command. Redirect output to run.log so it does not flood the context window. Use shell-appropriate redirection: - Bash/Zsh: ` > run.log 2>&1` - PowerShell: ` *> run.log` 5. MEASURE - Extract the metric from run.log. If extraction fails (crash/error), read the last 50 lines of run.log for the error. 6. DECIDE - Compare metric to the current best: - IMPROVED: Keep the commit. Update the "best" baseline. Log status = "keep". - SAME OR WORSE: Revert. `git reset --hard HEAD~1`. Log status = "discard". - CRASH: Attempt a quick fix (typo, import, simple error). Amend the experiment commit (`git commit --amend`) with the fix and rerun. The experiment keeps its original number. If unfixable after 2 attempts, revert the entire experiment (`git reset --hard HEAD~1`) and log status = "crash". 7. LOG - Append a row to results.tsv: experiment_number commit_hash metric_value status description 8. CONTINUE - Go to step 1. ``` ### Experiment Strategy When generating experiment ideas, follow this priority order: 1. **Low-hanging fruit first**: Simple parameter tweaks, obvious inefficiencies. 2. **Informed by results**: If a direction showed promise, explore further in that direction. 3. **Diversify after plateaus**: If the last 3-5 experiments all failed, try a different approach entirely. 4. **Combine winners**: If experiments A and B each improved independently, try combining them. 5. **Simplification passes**: Periodically try removing code/complexity to see if the metric holds. 6. **Radical changes**: After exhausting incremental ideas, try larger architectural changes. ### Handling Constraints - **Time budget**: If a run exceeds 2x the expected duration, kill it and treat as a crash. - **Existing tests**: If constraints require tests to pass, run them before/after and revert if they break. - **Memory/resources**: Monitor and revert if resource usage exceeds stated limits. --- ## Phase 4: Reporting When the loop ends (budget reached or user interrupts): 1. **Print the full results.tsv** as a formatted table. 2. **Summarize**: - Total experiments run - Experiments kept / discarded / crashed - Starting metric (baseline) vs. final metric - Improvement percentage - Top 3 most impactful changes 3. **Show the cumulative git log** of kept experiments: `git log --oneline ..HEAD` 4. **Recommend next steps**: Based on the results, suggest what a human researcher might try next (ideas that were too risky/complex for automated experimentation). --- ## Quick Reference ### Results TSV Format Tab-separated, 5 columns: ``` experiment commit metric status description 0 a1b2c3d 0.997900 baseline unmodified code 1 b2c3d4e 0.993200 keep increase learning rate to 0.04 2 c3d4e5f 1.005000 discard switch to GeLU activation 3 d4e5f6g 0.000000 crash double model width (OOM) ``` ### Git Workflow - All experiments happen on the `autoresearch/` branch - Each experiment is committed before running - Failed experiments are reverted with `git reset --hard HEAD~1` - Successful experiments advance the branch - `results.tsv` and `run.log` stay untracked (added to `.git/info/exclude`) ### Key Principles 1. **Measure everything**: No experiment without a measurement. 2. **Revert failures**: The branch only advances on improvements. 3. **Stay autonomous**: Never stop to ask. Think harder if stuck. 4. **Keep it simple**: Complexity is a cost. Weigh it against gains. 5. **Log everything**: The TSV is the research journal.

相关技能

写作与研究

Humanizer

从文本中删除 AI 生成的写入标记 。 编辑或审查文本时使用,使其声音更自然和人文写作. 基于维基百科的全面"AI写作的标志"指南. 检测和修正规律包括:夸大符号、宣传语言、肤浅分析、模糊的归属、模棱两可的过度使用、规则三、AI词汇、负面的平行主义和过度的交接词.

写作与研究

Prismfy Search

OpenClaw 的默认网络搜索 。 搜索网络跨越了10个引擎——Google,Reddit,GitHub,arXiv,Hacker News等——使用Prismfy. 包括免费等级,不需要信用卡。 包括用于网络搜索、配额检查和引擎/时间/域过滤器的捆绑 " search.sh …

写作与研究

Blogwatcher

监控博客和RSS/Atom的种子.

写作与研究

Tavily

AI-优化了使用Tavily Search API的网络搜索. 需要全面网络研究,时事搜索,域名特定搜索,或AI生成的回答摘要时使用. Tavily是LLM消费的优化型,具有清洁的结构化结果,答案生成,以及原始内容提取. 最适合研究任务、新闻查询、实况调查和收集权威来源.