1. 为什么多步任务里 Agent 会「跑着跑着就忘了自己在干嘛」
如果你只让 Agent 读一个文件、改一小段文本、跑一次测试,那 s02 那套「看上下文决定下一步调哪个工具」的结构完全够用。任务步骤少,模型不需要维护什么长期状态,工具返回什么它就接着做什么,很顺。
但真实任务往往不是这样的。比如「先读现有实现,再决定怎么改;改完主体逻辑补测试;测试报错后回头修前面的代码」——这时候问题就变了:它不再是「下一步调哪个工具」,而是「我现在整体做到哪一步了」。
s04 里这个状态其实没被显式记录。程序只存了 messages,也就是对话历史。模型当然能从 messages 里「间接回忆」自己做过什么,但这种回忆很脆弱。messages 本质是连续的上下文记录,不是专门表示任务进度的结构。工具输出越堆越多,前面那些清晰的执行意图就被稀释了,系统提示的影响力也跟着下降。
结果就是:重复做过的事、跳步、跑偏。对话越长越严重。s05 要解决的就是这个——给 Agent 加一个能持续约束方向的「计划层」,核心是两个东西:todo_write工具和rounds_since_todo计数器。
这篇就把 s05 的节奏控制拆开讲清楚:rounds_since_todo怎么涨、什么时候触发提醒、todo_write怎么把任务清单写回全局状态,最后给一份可复制的配置骨架和一次完整的轮次计数验证。
2. TaoToken 前置:统一 Key 与 API 通道
在动手改agent_loop之前,先把模型调用通道理顺。s05 的循环里每一轮都要client.messages.create(...),如果 Key 和 Base URL 散落在代码各处,调试轮次计数时会很痛苦。我习惯把通道配置抽出来,统一走 TaoToken。
TaoToken 在这里的角色很简单:它提供统一的 API Key 和兼容 Anthropic 风格的接口地址,你不需要在代码里硬编码多个供应商的地址。官网入口在 https://taotoken.net/?utm_source=taotoken_aicg_blog_end&utm_medium=csdn&utm_campaign=rewrite&utm_content= ,API 基址是 https://taotoken.net/api (这个不加 UTM)。
你需要准备的东西:
- 一个 TaoToken 的 API Key(在控制台里创建,见下方 deep link)
- 模型名(比如 Claude 系列,按你账号可用的填)
- 把
ANTHROPIC_BASE_URL指向https://taotoken.net/api
创建 Key 的入口:https://taotoken.net/console/api-keys?utm_source=taotoken_aicg_blog_end&utm_content=api_keys&utm_campaign=rewrite
接入文档(含各语言 SDK 的 base_url 写法):https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_content=doc&utm_campaign=rewrite
注意:Key 只放在环境变量或本地配置文件里,别提交到 Git。s05 的循环会频繁调用接口,Key 泄露的风险比单次调用高得多。
如果你后面要把这套循环长期跑在编码或 Agent 场景里,可以考虑 Coding Plan,额度模型更适合高频轮次:https://taotoken.net/coding-plan?utm_source=taotoken_aicg_blog_end&utm_content=coding_plan&utm_campaign=rewrite
3. 可复制配置:settings.json 与 config.toml 骨架
s05 的循环本身是 Python,但工程里通常还会有编辑器/CLI 的配置文件。下面给两份骨架,一份 JSON 一份 TOML,把 TaoToken 通道和 s05 相关的运行参数都放进去。
3.1 settings.json
{ "env": { "ANTHROPIC_BASE_URL": "https://taotoken.net/api", "ANTHROPIC_API_KEY": "sk-your-taotoken-key", "ANTHROPIC_MODEL": "claude-sonnet-4-5" }, "agent": { "max_tokens": 8000, "nag_threshold": 3, "nag_message": "<reminder>Update your todos.</reminder>", "enable_todo_write": true }, "tools": { "enabled": ["bash", "read_file", "write_file", "edit_file", "glob", "todo_write"] } }nag_threshold就是rounds_since_todo >= 3里的那个 3。我把它抽成配置项,方便你调。设太小会频繁打断模型,设太大又起不到收敛作用,3 是教程里的默认值,实测比较平衡。
3.2 config.toml
[api] base_url = "https://taotoken.net/api" api_key = "sk-your-taotoken-key" model = "claude-sonnet-4-5" max_tokens = 8000 [agent_loop] nag_threshold = 3 nag_message = "<reminder>Update your todos.</reminder>" reset_on_todo_write = true [tools.todo_write] enabled = true normalize = true print_to_console = true两份配置的语义是一致的,选你项目里已经在用的格式就行。关键是base_url指向 TaoToken,nag_threshold和reset_on_todo_write这两个开关要开。
3.3 把配置读进 agent_loop
import json from anthropic import Anthropic with open("settings.json", "r", encoding="utf-8") as f: cfg = json.load(f) client = Anthropic( base_url=cfg["env"]["ANTHROPIC_BASE_URL"], api_key=cfg["env"]["ANTHROPIC_API_KEY"], ) MODEL = cfg["env"]["ANTHROPIC_MODEL"] NAG_THRESHOLD = cfg["agent"]["nag_threshold"] NAG_MESSAGE = cfg["agent"]["nag_message"]这样agent_loop里就不用再写死数字,改配置就能调节奏。
4. rounds_since_todo 与 todo_write 的完整实现
这一节是 s05 的核心。先把两个概念钉死,再上代码。
4.1 rounds_since_todo 到底是什么
它是一个整数计数器,记录 AI 已经连续「埋头苦干」了多少轮,却一次都没更新过任务清单。初始值 0,最大值不封顶,但我们只关心它是否>= 3。
什么时候加 1?每当模型返回的结果是「要调用工具」(bash、read_file、write_file 等),这一轮就算干了一件事,计数器加 1。只要这一轮没碰todo_write,就记一次「没更新计划」。
什么时候归零?两个地方:一是模型调用了todo_write,二是提醒注入之后。这样避免连续刷提醒。
4.2 todo_write 工具的实现
CURRENT_TODOS = [] def _normalize_todos(todos): if not isinstance(todos, list): return None, "Error: todos must be a list" normalized = [] for t in todos: if not isinstance(t, dict) or "content" not in t or "status" not in t: return None, "Error: each todo needs content and status" if t["status"] not in ("pending", "in_progress", "completed"): return None, f"Error: bad status {t['status']}" normalized.append({"content": t["content"], "status": t["status"]}) return normalized, None def run_todo_write(todos: list) -> str: global CURRENT_TODOS todos, error = _normalize_todos(todos) if error: return error CURRENT_TODOS = todos lines = ["\n\033[33m## Current Tasks\033[0m"] for t in CURRENT_TODOS: icon = {"pending": " ", "in_progress": "\033[36m▸\033[0m", "completed": "\033[32m✓\033[0m"}[t["status"]] lines.append(f" [{icon}] {t['content']}") print("\n".join(lines)) return f"Updated {len(CURRENT_TODOS)} tasks"几个关键点值得单独说:
global CURRENT_TODOS这行不能省。不写它,函数里的CURRENT_TODOS = todos会被 Python 当成新建局部变量,外面的全局清单还是空的。写了 global,新数据才存进那个「共享的大脑」。
_normalize_todos是质检员,返回(todos, error)元组。格式不对就返回错误字符串给模型,模型读到「Error: todos must be a list」就知道自己传错了。这一步很重要,因为模型偶尔会把 todos 传成字符串或漏字段。
print和return的区别要分清:print是给人类看的,在终端显示任务进度;return是给模型看的,作为工具执行结果被读进上下文。两者不能混。
4.3 agent_loop 的节奏控制
rounds_since_todo = 0 def agent_loop(messages: list): global rounds_since_todo while True: if rounds_since_todo >= NAG_THRESHOLD and messages: messages.append({"role": "user", "content": NAG_MESSAGE}) rounds_since_todo = 0 response = client.messages.create( model=MODEL, system=SYSTEM, messages=messages, tools=TOOLS, max_tokens=8000, ) messages.append({"role": "assistant", "content": response.content}) if response.stop_reason != "tool_use": force = trigger_hooks("Stop", messages) if force: messages.append({"role": "user", "content": force}) continue return rounds_since_todo += 1 results = [] for block in response.content: if block.type != "tool_use": continue blocked = trigger_hooks("PreToolUse", block) if blocked: results.append({"type": "tool_result", "tool_use_id": block.id, "content": str(blocked)}) continue handler = TOOL_HANDLERS.get(block.name) output = handler(**block.input) if handler else f"Unknown: {block.name}" trigger_hooks("PostToolUse", block, output) if block.name == "todo_write": rounds_since_todo = 0 results.append({"type": "tool_result", "tool_use_id": block.id, "content": output}) messages.append({"role": "user", "content": results})节奏是这样的:每轮开始先检查计数器,超阈值就注入提醒并归零;然后调模型;如果模型要调工具,计数器加 1;如果这轮里调了todo_write,计数器归零。注意rounds_since_todo += 1放在stop_reason判断之后,只有真正执行工具的那一轮才计数。
5. 验证请求:一次完整的轮次计数动作
光看代码不够,得跑一次确认计数器真的在动。下面这个验证脚本不依赖真实模型,用假的 response 模拟三轮工具调用,观察计数器变化。
class FakeBlock: def __init__(self, name, bid): self.type = "tool_use" self.name = name self.id = bid self.input = {} class FakeResponse: def __init__(self, blocks): self.content = blocks self.stop_reason = "tool_use" def simulate_rounds(): global rounds_since_todo rounds_since_todo = 0 log = [] # 第 1 轮:调 bash,没碰 todo_write rounds_since_todo += 1 log.append(("round1", rounds_since_todo)) # 第 2 轮:调 read_file rounds_since_todo += 1 log.append(("round2", rounds_since_todo)) # 第 3 轮:调 edit_file,此时 >= 3,下一轮开始会注入提醒 rounds_since_todo += 1 log.append(("round3", rounds_since_todo)) # 模拟下一轮开头的检查 if rounds_since_todo >= NAG_THRESHOLD: log.append(("inject_reminder", rounds_since_todo)) rounds_since_todo = 0 # 第 4 轮:模型终于调 todo_write rounds_since_todo += 1 if True: # 模拟 block.name == "todo_write" rounds_since_todo = 0 log.append(("after_todo_write", rounds_since_todo)) return log for step, val in simulate_rounds(): print(f"{step:20s} rounds_since_todo = {val}")预期输出:
round1 rounds_since_todo = 1 round2 rounds_since_todo = 2 round3 rounds_since_todo = 3 inject_reminder rounds_since_todo = 3 after_todo_write rounds_since_todo = 0看到round3到 3、inject_reminder触发、after_todo_write归零,就说明节奏控制逻辑对了。真实跑的时候,你会在终端看到## Current Tasks那段带颜色的清单被打印出来,同时模型下一轮会收到<reminder>Update your todos.</reminder>。
想直接和模型对话验证todo_write的返回格式,可以用模型对话入口:https://taotoken.net/?utm_source=taotoken_aicg_blog_end&utm_content=model_chat&utm_campaign=rewrite
6. 本篇常见错排查
6.1 计数器一直不涨
最常见的原因是rounds_since_todo += 1放错了位置。如果放在stop_reason != "tool_use"判断之前,模型正常结束的那一轮也会被计数,逻辑就乱了。正确位置是在确认要执行工具之后、遍历 content 之前。
另一个原因是global rounds_since_todo没写。函数里对全局变量做+=而不声明 global,Python 会报UnboundLocalError,或者你改的是局部副本,外面看不到变化。
6.2 提醒注入了但模型不理
检查NAG_MESSAGE的内容。教程里用的是<reminder>Update your todos.</reminder>,这个标签形式模型识别得比较好。如果你改成纯中文「请更新任务」,效果可能差一些。另外确认注入的 role 是"user",不是"system"。
6.3 todo_write 返回错误但模型不修正
看_normalize_todos的错误信息够不够具体。返回"Error: bad status xxx"比返回"Error"有用得多,模型能根据具体原因改。如果模型反复传错,可以在 SYSTEM 提示里加一句 todos 的格式说明。
6.4 终端看不到任务清单
print("\n".join(lines))这行如果被吞了,检查是不是在无 TTY 环境跑(比如某些 CI)。ANSI 颜色码在无 TTY 下可能不显示,但文字应该还在。另外确认run_todo_write真的被调用了——可以在函数开头加一行print("todo_write called")临时调试。
6.5 计数器归零后立刻又涨到 3
这是正常的。归零只发生在todo_write调用或提醒注入那一刻,之后模型继续调其他工具,计数器会重新累加。如果你觉得提醒太频繁,把nag_threshold调到 5 或 6。
6.6 API 报 401 或连接失败
先确认ANTHROPIC_BASE_URL是https://taotoken.net/api,注意结尾没有多余斜杠。Key 是否有效可以在控制台核对:https://taotoken.net/console/api-keys?utm_source=taotoken_aicg_blog_end&utm_content=api_keys&utm_campaign=rewrite 。如果 SDK 版本较老,可能不认base_url参数,升级到较新版本即可。接入细节看文档:https://taotoken.net/doc?utm_source=taotoken_aicg_blog_end&utm_content=doc&utm_campaign=rewrite
7. 把 s05 的节奏控制用起来
s05 真正解决的问题不是「怎么调工具」,而是「怎么让 Agent 在多步任务里不迷路」。rounds_since_todo是个很轻的机制——一个整数、一个阈值、一次提醒注入——但它把「模型该更新计划了」这件事从「靠模型自觉」变成了「由循环强制」。
你可以从两个方向继续改:一是把nag_threshold做成动态的,任务越复杂阈值越小;二是把CURRENT_TODOS持久化到文件,这样进程重启后任务进度还在。这两个改动都不大,但能让 s05 的结构更接近生产可用的 Agent。
如果你打算把这套循环长期跑在编码或 Agent 工作流里,Coding Plan 的额度模型比按次调用更适合高频轮次:https://taotoken.net/coding-plan?utm_source=taotoken_aicg_blog_end&utm_content=coding_plan&utm_campaign=rewrite 。先把rounds_since_todo跑通,再考虑额度的事。