Code Mode collapses the executor, not just the wire
Code Mode 塌缩执行器而非仅通告面
`mode: 'code'` collapsed only the announcement surface, not the execution surface. `wireSchemas()` sent the model exactly one tool — `run_code` — but the executor resolved every call through `get()`, which returns the full visible map plus the reserved transport. A model that emitted a native tool name (`write`, `read`, `bash`, `subagent`, …) bypassed `run_code` entirely: the call traversed the normal pipeline and ex
English
Problem
mode: 'code' collapsed only the announcement surface, not the execution surface. wireSchemas() sent the model exactly one tool — run_code — but the executor resolved every call through get(), which returns the full visible map plus the reserved transport. A model that emitted a native tool name (write, read, bash, subagent, …) bypassed run_code entirely: the call traversed the normal pipeline and executed, even though no schema for it had ever been advertised. Providers do not intercept unadvertised tool names, so schema omission enforced nothing.
The package contract names this exact anti-pattern: schema omission is not enforcement when a direct caller can bypass it; denial must be tested through the executor.
Decision
ToolRuntime resolves callable definitions through a new private resolveExecution(name, scope, nested) that applies the mode collapse at the operation boundary that owns it. When modeFor(scope) resolves to code, a model-direct call (nested = false) may only name the reserved run_code transport; every native name resolves to undefined and surfaces as the executor's existing UNKNOWN_TOOL error, whose message names the route back through run_code because the name IS declared to this model (an already-aborted caller signal keeps the cancellation contract: ABORTED_BEFORE_DISPATCH, with the visible tool's finalizer applied). The effective scope mode includes declarations inherited from an agent preset, so its wire schema and execution permissions remain aligned. A collapsed call terminates at createExecution — the first stage of prepare — BEFORE the extensible policy pipeline, so tools/pre-execute listeners, approval ask, and guards never observe a call that is deterministically denied; a human is never prompted to approve it. A nested sub-dispatch (nested = true — a parent token set, which only the run_code SDK binding sets in production code) may call any visible tool, so programs keep every binding the generated SDK declared.
Four execution-path lookups — executionMode, dispatchToolBody, postExecute, normalizeDispatchResult — go through resolveExecution. createExecution applies the same collapse via the shared collapses(name, nested) predicate so it can distinguish a collapsed call from a genuinely unknown name before the policy pipeline. The public registry view (get) and SDK projection (schemas) keep their semantics: presentation, inspection, and binding enumeration still see the full visible set. The wire (wireSchemas) and the executor now agree. A collapsed call with non-JSON-serializable arguments reports the parameter TypeError (the invalid-args contract), not UNKNOWN_TOOL — the body still never runs and policy still does not.
The collapse is a security-relevant invariant, so acceptance is pinned through the executor: a model-direct native call under code returns UNKNOWN_TOOL, the same tool via an SDK sub-dispatch succeeds, and native/both direct calls plus run_code itself are unchanged. The base Code Mode foundation owns the transport design this note layers the execution boundary onto.
Alternatives considered
Filter get() / the registry view by mode
The view is consumed by presenters, tool-cordis inspection, and the SDK binder; collapsing it would hide from the program surface tools that must still bind, and would change the public resolution contract for every consumer, not just the executor.
Filter at the agent-loop entry
The loop is not the only executor caller, and the distinction that matters (model-direct vs transport sub-dispatch) rides on the execution input, not at the loop boundary. An entry filter would also re-encode mode semantics the registry already owns.
Reject via a shipped guard
Guards are an optional plugin extension; a security invariant must not depend on a deployment composing the right plugin. The registry owns the mode decision and must enforce it itself.
Keep schema omission only (status quo)
No provider guarantees interception of unadvertised names; the reported session proves it does not happen.
Consequences
mode: 'code'now enforces what it announces: a model-direct native call becomesUNKNOWN_TOOL, which the model can correct by routing throughrun_code(a pre-aborted call still resolvesABORTED_BEFORE_DISPATCH, per the cancellation contract).bothandnativebehavior is unchanged; SDK sub-dispatches are unchanged (theparenttoken is the discriminator).- A collapsed call is rejected at
prepare, BEFORE the extensible policy pipeline: pre-execute listeners, approvalask, and guards never observe it.executionModealso fails closed (exclusive), so scheduling has no observable difference. - Native-tool guidance sections (
tool:read,tool:write,tool:bash, etc.) remain in the system prompt because they describe capabilities available through the generated SDK as well as native function calls, and several carry cross-tool routing policy (readoverbash cat,readbeforewritefor the default fs-observation-policy,subagentoverworkflow) that no single tool description can hold. The executor collapse, not prompt filtering, prevents model-direct native calls. - The prompt STATES the collapse, in the
tools:code-onlysection ordered ahead of the 100-199 guidance band. Those sections name their tool without qualifying how it is reached, so a model that read only them emitted a native call, receivedUNKNOWN_TOOLfor a tool the same prompt declared, and concluded the deployment was inconsistent rather than correcting itself. The denial carries the route for the same reason.bothrenders the rule empty: its native calls do execute, so stating it there would be false — which is whyboth-mode-turnno longer sharescode-mode-turn's expected prompt. - Any future composite transport that sets a
parenttoken opts its sub-dispatches into the full table, matching the nested-call semantics the token already documents.
中文
问题
mode: 'code' 只塌缩了通告面,没有塌缩执行面。wireSchemas() 只向模型发送一个工具——run_code——但执行器通过 get() 解析所有调用,而 get() 返回完整的可见工具表外加保留的传输工具。模型一旦发出原生工具名(write、read、bash、subagent 等),就能完全绕过 run_code:调用照常走完整流水线并执行成功,尽管它的 schema 从未被通告过。模型提供方不拦截未通告的工具名,因此不发 schema 等于没有约束。
包契约点名了这个反模式:当直接调用方可以绕过时,schema 省略不算强制执行;拒绝必须经执行器验证。
决策
ToolRuntime 通过新增的私有 resolveExecution(name, scope, nested) 解析可执行定义,在拥有该决策的操作边界上应用模式塌缩。当 modeFor(scope) 解析为 code 时,模型直呼(nested = false)只允许命名保留的 run_code 传输工具;任何原生名字都解析为 undefined,并以执行器既有的 UNKNOWN_TOOL 错误呈现,其消息会指出改走 run_code 的正确路径——因为这个名字对当前模型而言是已声明过的(已中止的调用方 signal 保留取消契约:ABORTED_BEFORE_DISPATCH,并应用可见工具的 finalizer)。有效的 scope 模式包括从 agent preset 继承的声明,因此其 wire schema 与执行权限保持一致。被塌缩的调用在 createExecution(prepare 的第一阶段)即终止——在可扩展策略流水线之前,因此 tools/pre-execute 监听器、approval ask 与 guard 永远不会观察到一个注定被拒绝的调用,人类也不会被提示去批准它。嵌套子调用(nested = true——即设置了 parent token,生产代码中只有 run_code SDK 绑定会设置)可以调用任意可见工具,因此程序保留生成 SDK 声明的全部绑定。
执行链路的四处查表——executionMode、dispatchToolBody、postExecute、normalizeDispatchResult——改走 resolveExecution。createExecution 通过共享的 collapses(name, nested) 谓词应用同一塌缩,以便在策略流水线之前区分被塌缩的调用与真正未知的名字。公共注册表视图(get)与 SDK 投影(schemas)语义不变:展示、检查与绑定枚举仍看到完整可见集合。通告(wireSchemas)与执行器现在一致。带非 JSON 可序列化参数的塌缩调用报告参数 TypeError(invalid-args 契约),而非 UNKNOWN_TOOL——函数体仍不会运行,策略也不会执行。
塌缩是安全相关的不变量,因此验收经执行器钉死:code 模式下模型直呼原生工具返回 UNKNOWN_TOOL;同一工具经 SDK 子调用成功;native/both 模式直呼与 run_code 本身行为不变。本 note 把执行边界叠加在基础 Code Mode 基础 之上,传输设计由后者拥有。
备选方案
按模式过滤 get() / 注册表视图
视图被展示方、tool-cordis 检查与 SDK 绑定消费;塌缩视图会从程序表面隐藏仍必须绑定的工具,并改变所有消费者的公共解析契约,而不只是执行器。
在 agent-loop 入口过滤
loop 不是唯一的执行器调用方,且真正要紧的区分(模型直呼 vs 传输子调用)挂在执行输入上,不在 loop 边界。入口过滤还会重复编码注册表已经拥有的模式语义。
通过内置 guard 拒绝
guard 是可选的插件扩展;安全不变量不能依赖部署恰好组装了正确的插件。模式决策归注册表所有,必须由它自己执行。
只保留 schema 省略(维持现状)
没有提供方保证拦截未通告的名字;被报告的会话证明拦截不会发生。
后果
mode: 'code'现在兑现其通告:模型直呼原生工具变为UNKNOWN_TOOL,模型可以通过改走run_code自行纠正(已中止的调用仍按取消契约解析为ABORTED_BEFORE_DISPATCH)。both与native行为不变;SDK 子调用不变(判别信号是parenttoken)。- 被塌缩的调用在
prepare阶段即被拒绝——在可扩展策略流水线之前:pre-execute 监听器、approvalask与 guard 永远不会观察到它。executionMode同样 fail-closed(exclusive),调度无可观察差异。 - 原生工具指引段(
tool:read、tool:write、tool:bash等)保留在系统提示词中,因为它们同时描述了通过生成 SDK 及原生函数调用可用的能力,其中若干段还承载着任何单个工具描述都装不下的跨工具路由策略(read优先于bash cat、默认 fs-observation-policy 要求先read再write、一两个委派用subagent而非workflow)。防止模型直呼原生工具的是执行器塌缩,而非提示词过滤。 - 提示词会声明这条塌缩,位于排在 100–199 指导段之前的
tools:code-only段。那些段只写出工具名而不限定其可达方式,因此只读到它们的模型会发出原生调用,为一个同一份提示词刚刚声明过的工具收到UNKNOWN_TOOL,进而判定部署不一致,而不是自行纠正。拒绝信息给出正确路径也是同一原因。both下该规则渲染为空:它的原生调用确实会执行,在那里声明就是假话——这也是both-mode-turn不再与code-mode-turn共用期望提示词的原因。 - 未来任何设置
parenttoken 的组合传输,其子调用自动走全表,与该 token 已有的嵌套调用语义一致。