Product / Client
用户交互、clarification、brief、background job、引用与错误呈现。
一次 request 为什么能完成搜索、阅读、代码分析和再次搜索?从最小手写 loop 开始,把每一层的 owner 找出来。
整页只用一个案例:查询东盟十国首都坐标,并用 Python 找出距离最近的一对。每遇到一个新名词,我们都回到四个问题:谁决定?谁执行?谁保存 observation?谁管理 lifecycle?
核心哲学:实验 1-3 最容易让人误解的地方,是把四个不同层级压缩成一句「模型原生具备 Deep Research 能力」。真正发生的事情是一条逐层迁移:Model 学会选择 action;Responses API runtime 维持 hosted inner loop;Web Search 与 Code Interpreter 执行真实动作;Product / Client 负责澄清需求和管理任务生命周期。 循环没有消失,搜索引擎也没有进入 model weights;变化的是 decision policy 和 loop ownership。
实验 1-3 的原文在很短的篇幅内连续做了六次跳跃:
模型有原生 Agent 能力
↓
模型能自主调用工具
↓
Responses API 有 Web Search / Code Interpreter
↓
因此可以 Deep Research
↓
还能先做 Intent Clarification
↓
所以无需外部编排、无需手写 ReAct loop
每个箭头都需要额外解释。因为这六句话并不属于同一层:
| 书中说法 | 实际在谈哪一层 | 不能误解成什么 |
|---|---|---|
| 原生 Agent / ReAct 能力 | Model 的 tool-use decision policy | 模型权重里装了浏览器和 Python |
| Responses API | Provider 的 API + agent runtime | 单纯的文本生成 endpoint |
| Built-in tools | Provider 托管的 execution infrastructure | 模型自己访问互联网 |
| Deep Research | Model + runtime + tools + context 的 system capability | 一个单独参数或一个工具名称 |
| Intent Clarification | 通常是 Product / Client 的 preliminary flow | Raw Deep Research API 自动追问 |
| 不用手写 loop | Hosted tools 的 inner loop 由 service 管 | 整个应用不再需要 orchestration |
关键洞察:真正需要回答的不是「GPT 到底有没有 Agent 能力」,而是:一次任务中的每一个 decision、execution、observation 和 lifecycle action,分别由谁拥有?
Learning goal:先意识到困惑不是你的理解问题,而是原文跨层跳跃太快。接下来我们逐个补上箭头。
点「下一步」观察:Model 只提出 action;Client 执行并把 observation 写回 Context。右下角会显示此刻模型下一轮真正能看到什么。
这里的 ReAct 是
Reasoning + Acting,不是前端框架 React。
它描述一个闭环:
观察当前 Context
↓
决定下一步 action
↓
执行 action
↓
把 observation 放回 Context
↓
根据新 Context 再决定
最重要的字不是 Reasoning,也不是 Acting,而是 loop。只要下一步必须依赖刚刚发生的结果,就一定需要某个 runtime 维持这个闭环。
假设我们给模型一个 custom function:
{
"type": "function",
"name": "search_web",
"description": "Search the public web",
"parameters": {
"type": "object",
"properties": {
"query": { "type": "string" }
},
"required": ["query"]
}
}用户问:
找出东盟 10 国首都之间距离最近的一对。
模型可能输出:
{
"type": "function_call",
"name": "search_web",
"arguments": "{\"query\":\"ASEAN capitals official coordinates\"}",
"call_id": "call_01"
}这时什么都还没有被搜索。
模型只生成了一段结构化 token,意思相当于:
「应用,请帮我执行
search_web(query=...)。」
真实网络请求必须由模型之外的程序完成。
conversation = [user_message]
while True:
response = call_model(
input=conversation,
tools=[search_web, run_python],
)
conversation.extend(response.output)
tool_calls = find_tool_calls(response.output)
if not tool_calls:
return response.output_text
for call in tool_calls:
result = execute_in_my_application(call)
conversation.append({
"type": "function_call_output",
"call_id": call.call_id,
"output": result,
})逐行看职责:
| 步骤 | 谁负责 | 发生了什么 |
|---|---|---|
call_model(...) |
Client → Model | 把 Context 和 tools 交给模型 |
function_call |
Model | 选择下一步 action |
execute_in_my_application |
Client / Tool | 真正执行搜索或代码 |
function_call_output |
Client | 把 observation 放回 Context |
再次 call_model |
Client → Model | 让模型基于新 observation 决定 |
while True |
Client | 维持 loop,直到没有 tool call |
这就是「手写 ReAct loop」最具体的含义。
在工具执行之前,模型不知道工具会返回什么。
Model inference #1
输出:请搜索 A
现实世界发生搜索
返回:B
Model inference #2
输入:刚才返回了 B
输出:根据 B,下一步执行 C
B 是第一次 inference
结束之后才产生的新信息。因此必须有某个模型外部 runtime:
所以:
“Model-native ReAct” 不可能表示整个循环都存在于 weights 内。它最多表示:模型已经学会在循环中做下一步 decision。
Learning goal:看到一个 tool call 时,立即区分「提出 action」和「执行 action」。这条边界是后面全部概念的地基。
高亮层表示该能力真实依赖的组成部分。没有任何一个复杂 Agent 能力只靠一个框完成。
用户交互、clarification、brief、background job、引用与错误呈现。
保存 trajectory、识别 hosted call、路由工具、把 observation 接回循环。
真实搜索索引、页面读取、Python container 与文件执行。
读取 Context、选择 next action、生成 query / code、综合报告。
以后遇到任何 Agent 宣传语,先把它放进下面四层。
┌──────────────────────────────────────────────┐
│ Layer 4 · Product / Client │
│ 用户交互、Intent Clarification、job 管理、引用展示 │
├──────────────────────────────────────────────┤
│ Layer 3 · Responses API Runtime / Service │
│ 保存 trajectory、路由 hosted calls、维持 inner loop │
├──────────────────────────────────────────────┤
│ Layer 2 · Hosted Tools │
│ Web Search、Code Interpreter、File Search ... │
├──────────────────────────────────────────────┤
│ Layer 1 · Model │
│ 读取 Context,选择 next action,生成最终综合 │
└──────────────────────────────────────────────┘
模型负责:
模型不负责:
Hosted Tool 是 Provider 运行的真实服务:
web_search 背后有搜索、页面读取与 citation 数据;code_interpreter 背后有 container、Python
runtime、文件系统与资源限制。它们是 infrastructure,不是 weights。
Responses API 不只是「把 prompt 传给模型」。它还是一个 hosted agent runtime,可以:
Product / Client 决定完整用户体验:
直观理解:Model 像研究员,Hosted Tools 像数据库和计算实验室,Responses runtime 像负责派单与传递实验结果的研究机构,Product / Client 像接待用户、定义项目和交付报告的咨询团队。
Learning goal:任何“能力”都要问:是 model decision、tool execution、runtime orchestration,还是 product workflow?
早期 Agent 常把流程写成确定性规则:
search("ASEAN capitals")
open(first_result)
extract_coordinates()
run_python(haversine_code)
write_report()每一步由开发者预先规定。模型主要负责填内容。
缺点是任务一变,流程就容易失效:
经过 tool-use training / reinforcement learning 后,模型可以学习一个 policy:
π(next_action | goal, available_tools, trajectory)
不用记公式。它只表达:
给定目标、可用工具和已经发生的 trajectory,模型学习选择下一步 action。
被写进模型能力的是:
没有被写进模型 weights 的是:
因为开发者不再需要用大量 if/else 规定每个 decision:
以前:Developer policy 决定下一步
现在:Model policy 决定下一步
但 loop host 仍然存在:
Model policy 决定 next action
Runtime / Harness 维持 action → observation → next decision
Tool infrastructure 执行 action
准确说法应该是:
模型原生具备参与 ReAct loop 的 decision policy,而不是模型单独拥有整个 ReAct system。
Learning goal:理解 “native” 描述 decision policy 的来源,不描述真实工具的物理位置。
Freeform 只是 payload format;Hosted 才说明 execution infrastructure 由 Provider 提供。切换三类工具看不变与变化。
快速判断:哪一种声明会让 Python 在 OpenAI 托管的 sandbox 中真实执行?
哪一个才是 runtime 能识别并 dispatch 的 Custom Tool call?
用户最容易把三种 tool 混在一起。它们看起来都放进
tools=[...],但 execution ownership 完全不同。
| 类型 | 输入格式 | 谁执行 | Client 要不要回传结果 | 示例 |
|---|---|---|---|---|
| Function tool | JSON Schema arguments | 你的应用 | 要 | get_weather({city}) |
| Custom tool | 外层是结构化 custom_tool_call;内层 input 是 freeform text,可加 grammar |
你的应用 | 要 | 原始 Python / SQL / shell 文本 |
| Hosted tool | Provider 定义的固定类型 | OpenAI infrastructure | 通常不用你逐次执行 | web_search、code_interpreter |
定义:
{
"type": "function",
"name": "lookup_coordinates",
"parameters": {
"type": "object",
"properties": {
"city": { "type": "string" },
"country": { "type": "string" }
},
"required": ["city", "country"],
"additionalProperties": false
}
}模型输出:
{
"type": "function_call",
"name": "lookup_coordinates",
"arguments": "{\"city\":\"Bangkok\",\"country\":\"Thailand\"}",
"call_id": "call_42"
}你的应用执行并回传:
{
"type": "function_call_output",
"call_id": "call_42",
"output": "{\"lat\":13.7563,\"lon\":100.5018}"
}先把最容易误解的一句话修正掉:
不是“模型随口说一句自然语言,API 就猜测它要调用工具”。Custom Tool call 仍然有 machine-readable 的严格外层协议;freeform 的只有其中
input字段。
工具定义本身首先是结构化的:
{
"type": "custom",
"name": "python_runner",
"description": "Run Python supplied as plain text"
}模型决定调用后,Responses API 返回的不是普通 assistant
prose,而是一个有明确 type 的 output item:
{
"id": "ctc_123",
"type": "custom_tool_call",
"status": "completed",
"call_id": "call_42",
"name": "python_runner",
"input": "print(2 + 2)"
}把它拆成两层就不会混淆:
Responses API output item
┌────────────────────────────────────────────┐
│ Strict call envelope │
│ │
│ type: "custom_tool_call" │
│ name: "python_runner" │
│ call_id: "call_42" │
│ input: ───────────────────────────────┐ │
└─────────────────────────────────────────│──┘
│
▼
Freeform payload string
"print(2 + 2)"
| 层 | 谁规定格式 | 是否严格 |
|---|---|---|
| Call envelope | Responses API protocol | 严格。必须能识别
type、name、call_id、input |
| Payload slot | Custom Tool contract | API 不要求它是 JSON object,只要求它是一个 string |
| Payload language | 你的 executor | 仍可能很严格:Python、SQL、shell、DSL 或 natural language |
所以,Freeform 不等于 No Format:
Freeform
≠ 整个调用没有结构
≠ 随便写什么工具都能理解
Freeform
= Responses API 不再要求 input 遵循 JSON Schema
= input 可以直接承载工具原生需要的文本
只有当工具实现本身就接受 natural language 时,input
才能是自然语言。
| Custom Tool | 合法 payload | “帮我算一下”是否有效 |
|---|---|---|
python_runner |
Python source code | 通常无效;executor 需要 Python |
sql_runner |
SQL statement | 通常无效;executor 需要 SQL |
shell_runner |
shell command | 通常无效;executor 需要 shell syntax |
instruction_router |
natural-language instruction | 可以,因为工具被设计成读取自然语言 |
换句话说,Custom Tool 并没有让所有工具突然“听懂人话”。它只是让工具不必先把原生文本塞进 JSON arguments。
Custom Tool 可以附加
grammar。例如只允许一个极小的算术表达式语言:
{
"type": "custom",
"name": "calculator",
"description": "Evaluate a basic arithmetic expression",
"format": {
"type": "grammar",
"syntax": "regex",
"definition": "[0-9]+([+*][0-9]+)*"
}
}此时:
custom_tool_call;input 仍是 string;如果模型只是输出普通文本:
请调用 python_runner 帮我计算 2 + 2。
这只是 assistant message,不是
custom_tool_call;runtime 没有机器可读的 call
item,就不会把它当成工具调用 dispatch。
你的 Client 真正处理的是下面这条闭环:
custom_tool_call envelope
↓ extract name + input
Client validates input
↓
Client dispatches python_runner("print(2 + 2)")
↓
{"type":"custom_tool_call_output","call_id":"call_42","output":"4"}
↓
Model reads observation and continues
type: "custom"仍然是你的工具。Responses API 只把“这是一次工具调用”编码成结构化 item,不会因为input里出现 Python 就自动执行它。
Learning goal:看到 Freeform Custom Tool 时,能立即说出:outer envelope is structured; only inner payload is freeform; executor syntax and Client validation still apply.
{
"type": "web_search"
}或:
{
"type": "code_interpreter",
"container": { "type": "auto" }
}这里不需要提供你的函数地址或 executor。因为 OpenAI 已经实现并运营执行环境。
Learning goal:看到 tools
数组时,不要只看“都是工具”;第一眼先看 type,再判断
execution ownership。
如果工具参数是一大段 Python:
print("hello")
path = "C:\\temp\\data.csv"硬塞进 JSON arguments,需要大量 escaping:
{
"code": "print(\"hello\")\npath = \"C:\\\\temp\\\\data.csv\""
}Freeform Custom Tool 允许 custom_tool_call.input
直接承载原始文本:
{
"type": "custom_tool_call",
"name": "python_runner",
"call_id": "call_42",
"input": "print(\"hello\")\npath = \"C:\\temp\\data.csv\""
}这里 wire-level output item 仍然是结构化 JSON;减少的只是
input 内部再包一层 {"code": ...} 的
serialization friction。Python、SQL、shell、DSL 仍须符合各自 executor
的语法。
Freeform 并不意味着:
它只改变:
工具调用的 payload 格式
JSON object
↓
raw text
它没有改变:
工具归谁所有
谁执行
谁回传结果
谁维持 loop
Freeform Tool Calling
= 参数表示方式
Hosted Tool
= 执行基础设施归 Provider
Managed Inner Loop
= action / observation 循环归 Service
三个概念可以同时出现,也可以完全独立。
例如:
python_runner:raw text,但 Client
执行、Client loop;get_weather:JSON arguments,Client
执行、Client loop。Learning goal:能用一句话解释:Freeform 是 syntax / serialization 特性,不是 Agent architecture 特性。
你的请求:
response = client.responses.create(
model="gpt-5.6-sol",
input="找一条最近的正面新闻并引用来源。",
tools=[{"type": "web_search"}],
){"type": "web_search"}
不是把搜索代码发送给模型。它更像打开一个 capability flag:
本次 response 允许模型申请使用 OpenAI 托管的 Web Search。
OpenAI 同时控制:
因此 service 可以在内部完成:
Model 生成 web_search_call
↓
Responses runtime 识别 hosted call
↓
OpenAI Web Search 执行 search / open_page / find_in_page
↓
结果回到 response trajectory
↓
Model 基于结果继续
Client 不需要收到 query 后自己去调用搜索 API。
概念化的 response.output:
[
{
"type": "web_search_call",
"action": {
"type": "search",
"query": "ASEAN capitals official coordinates"
}
},
{
"type": "web_search_call",
"action": {
"type": "open_page",
"url": "https://..."
}
},
{
"type": "message",
"content": [
{
"type": "output_text",
"text": "...",
"annotations": ["url_citation ..."]
}
]
}
]具体字段会随 SDK 版本变化,但 ownership 不变:
准确说法:
Built into the Responses platform
而不是:
Built into model weights
Learning goal:理解 API 之所以能内置搜索,是因为 Provider 不只提供模型,还运营搜索与 orchestration runtime。
请求声明:
response = client.responses.create(
model="gpt-5.6-sol",
input="用 Python 计算所有东盟首都对的大圆距离。",
tools=[{
"type": "code_interpreter",
"container": {"type": "auto"}
}],
)这里的 container: auto 表示 Provider 创建或复用一个受控
container。
模型会做两件事:
Hosted infrastructure 做另外几件事:
python_runner 的本质区别| 问题 | Custom python_runner |
Hosted code_interpreter |
|---|---|---|
| 谁定义工具 | 你 | OpenAI |
| 谁提供 sandbox | 你 | OpenAI |
| 谁执行代码 | 你的 Client / backend | OpenAI infrastructure |
| 谁回传 result | 你的 Client | Responses runtime |
| Client 是否手写 inner loop | 是 | 对该 hosted loop 通常不需要 |
| 参数能否是代码文本 | 可以 | 模型也会生成代码,但接口由平台定义 |
如果让模型直接心算所有城市组合,它可能:
Code Interpreter 把 deterministic computation 交给真实 runtime:
Model 决定算法与生成代码
↓
Python runtime 负责确定性执行
↓
Model 解释结果与限制
Learning goal:把「会写代码」和「有地方执行代码」分开。
左侧是完整顺序;右侧解释当前 step 的 owner、可观察 item,以及为什么它会引出下一步。Client 没有逐步 dispatch hosted tools。
这是实验 1-3 最核心、也最容易被一句话带过的地方。
Client 可能只写一次:
response = client.responses.create(...)但 Service 内部可以产生一条长 trajectory:
1. Model: 先搜索首都名单
2. Tool: web_search.search
3. Model: 打开权威来源
4. Tool: web_search.open_page
5. Model: 在页面中找坐标
6. Tool: web_search.find_in_page
7. Model: 坐标不全,再搜索
8. Tool: web_search.search
9. Model: 数据足够,生成 Python
10. Tool: code_interpreter_call
11. Model: 检查结果并写报告
这并不矛盾:
Client 视角:一次 request / 一个 background job
Service 视角:多次 decision + tool action + observation
手写 custom loop:
Client
├─ call model
├─ inspect tool call
├─ execute tool
├─ append output
└─ repeat
Hosted tools:
Client
├─ submit request
├─ poll / wait webhook
└─ render result
Responses Service
├─ get model decision
├─ execute hosted tool
├─ append observation
└─ continue inner trajectory
因此准确结论是:
Client-written inner loop 被 service-managed inner loop 替代。Loop 的 ownership 迁移了,loop 本身没有消失。
对 Client 来说,不需要把它实现成多次 API call。Provider 内部如何组织 inference、continuation 与 tool round,是服务实现细节。
我们能可靠讨论的是 observable contract:
不要把看不到的 provider internals 伪装成确定事实。
Learning goal:理解「一次 request」与「一次 action / 一次 decision」不是同一粒度。
一个模型调用一次
web_search,然后根据第一条结果回答,只能叫 web-enabled
answering,不能自动叫 Deep Research。
Deep Research 通常需要:
Deep Research capability
=
long-horizon research policy
+ persistent trajectory / context
+ retrieval tools
+ analysis runtime
+ loop host
+ source-aware synthesis
拿掉任何一项:
| 拿掉什么 | 会发生什么 |
|---|---|
| Research policy | 可能只搜一次就草率回答 |
| Persistent trajectory | 忘记查过什么、重复搜索 |
| Web / file / MCP sources | 只能依赖训练时知识 |
| Code Interpreter | 数据计算变得脆弱或不可复核 |
| Loop host | 模型提出第一次 tool call 后就停住 |
| Citation synthesis | 报告难以追溯证据 |
模型经过训练,具有更强的长程 research policy:会规划、搜集、检查、再搜索和综合。
Responses API 直接提供 compatible retrieval / analysis tools,并托管 inner trajectory。
ChatGPT 等产品把 clarification、progress、background execution、report presentation 组成完整体验。
书中一句「GPT-5.6 原生 Deep Research」实际上把这三种 native 叠在了一起。
更准确的说法是:
GPT-5.6 的 research policy,加上 Responses API 的 hosted loop 和工具基础设施,共同形成了可直接调用的 Deep Research system capability。
Learning goal:以后看到“原生 Deep Research”,马上追问它是在说 model、platform,还是 product。
切换视角。Raw API 从你给的 input 开始研究;ChatGPT product 或你自己的 app 可以在它之前增加 preliminary flow。
用户说:
分析最近一个月的比特币。
至少缺少:
Deep Research 越努力,模糊目标造成的浪费越大。
产品可以实现:
Raw user request
↓
Model identifies missing requirements
↓
Product asks user 2–4 questions
↓
User answers
↓
Model rewrites a research brief
↓
Product starts Deep Research
这看起来像「Deep Research 会澄清意图」,因为用户只看到一个统一产品。
当前官方 guide 明确指出:
Deep research via the Responses API does not include
clarification or prompt rewriting.
Raw API 的行为更接近:
你给什么 research instructions
↓
它就从这些 instructions 开始研究
如果你希望 API 应用也有 clarification,需要自己增加 preliminary flow:
questions = fast_model.find_missing_requirements(user_request)
answers = ask_user(questions)
research_brief = fast_model.rewrite(user_request, answers)
research_job = start_deep_research(research_brief)Intent Clarification 不需要一个神秘的新神经网络模块。它可以来自:
也就是说:
Language reasoning capability
+ clarification prompt
+ product state
+ user interaction
=
Intent Clarification experience
它通常是 workflow capability,不是 raw Deep Research endpoint 自动携带的阶段。
Learning goal:明确区分「ChatGPT 产品表现出来的能力」和「Responses API contract 保证的能力」。
| 场景 | Inner loop 谁写 | 工具谁执行 | Client 仍负责什么 |
|---|---|---|---|
| Function tool | Client | Client | 全部 loop + lifecycle |
| Freeform custom tool | Client | Client | 全部 loop + lifecycle |
| Responses hosted tools | Responses Service | OpenAI | outer lifecycle |
| ChatGPT Deep Research product | Product + Responses Service | OpenAI | 最终用户只使用产品 |
你通常不再需要写:
while tool_call_exists:
inspect_tool_call()
execute_search_or_python()
append_tool_output()
call_model_again()因为这部分由 Responses Service 与 hosted tools 协作完成。
如果你在构建自己的应用,通常仍需:
输入与权限校验
Intent Clarification(如需要)
Research brief 构建
background job 提交
polling 或 webhook
timeout / retry / cancel
max_tool_calls / cost budget
错误和部分结果处理
citation rendering
日志、安全与 prompt injection 防护
所以书中「无需外部编排代码」如果不加限定,会误导。
准确版本:
使用 Responses API 的 hosted tools 时,开发者无需为这些 hosted tools 手写 search–read–analyze 的 inner ReAct loop;但仍需编排任务之前和任务之外的 product lifecycle。Custom tools 仍需 Client 执行并回传结果。
Learning goal:能把 orchestration 切成 inner trajectory 和 outer lifecycle,而不是笼统说“有”或“没有”。
现在把所有概念放进同一个案例。
找出东盟 10 国首都之间距离最近的一对。
应用发现以下不确定性:
“首都”是否指当前国家首都?
距离是否指经纬度大圆距离?
坐标来源需要什么可信度?
结果需要方法、代码和 citations 吗?
用户回答后,应用生成 brief:
目标:比较东盟 10 个成员国当前首都之间的大圆距离。
来源:优先政府、国际组织或权威地理来源;保留每个坐标的 citation。
方法:统一使用城市中心经纬度和 Haversine formula;用 Python 枚举所有城市对。
输出:最近的一对、距离、计算方法、来源、数据冲突与限制。
这一步不是 Raw Deep Research API 自动完成的。
response = client.responses.create(
model="gpt-5.6-sol",
background=True,
input=research_brief,
tools=[
{"type": "web_search"},
{
"type": "code_interpreter",
"container": {"type": "auto"},
},
],
max_tool_calls=40,
)这段代码只声明:
它没有手写研究步骤。
| Step | Decision / Action | Owner | 可观察对象 |
|---|---|---|---|
| 1 | 解析 brief,决定先收集首都名单 | Model | reasoning summary(如启用) |
| 2 | 搜索东盟成员与首都权威来源 | Model → Service | web_search_call: search |
| 3 | 真正运行搜索 | Hosted Web Search | search results |
| 4 | 决定打开候选来源 | Model | next action |
| 5 | 打开与读取页面 | Hosted Web Search | open_page |
| 6 | 在页面定位坐标字段 | Hosted Web Search | find_in_page |
| 7 | 检查缺失或冲突坐标,决定补搜 | Model | revised search action |
| 8 | 获取足够的十组坐标 | Service + Tool | accumulated trajectory |
| 9 | 决定用 Haversine 枚举城市对 | Model | code plan / call |
| 10 | 生成并运行 Python | Hosted Code Interpreter | code_interpreter_call |
| 11 | 返回距离列表或计算文件 | Code Interpreter | execution output |
| 12 | 检查单位、异常与来源一致性 | Model | summary / possible new call |
| 13 | 生成最终报告与 citations | Model + Service | final message |
Client 不必在每个 Step 之间发起自定义 executor,但仍需管理 job:
while response.status in {"queued", "in_progress"}:
sleep(2)
response = client.responses.retrieve(response.id)
render_report(response.output_text)
render_citations(response.output)这里的 while 是 job lifecycle
polling,不是 search / code 的 ReAct inner loop。
如果你不用 hosted web_search,而是声明自己的
search_web function:
Model 输出 function_call
↓
你的 Client 调搜索 API
↓
你的 Client 回传 function_call_output
↓
你的 Client 再次调用 Responses API
这时 inner loop 又回到 Client。
关键洞察:是不是要手写 loop,不由“模型聪不聪明”单独决定,而由 tool execution ownership + API runtime contract 决定。
Learning goal:能够逐步指认完整 trajectory 中每一步的 owner,而不是只记住一句“API 自动研究”。
GPT-5.6 具有原生 Deep Research 能力。
GPT-5.6 具备经过训练的长程 research / tool-use policy;当它运行在 Responses API 的 hosted runtime 中,并获得 Web Search、Code Interpreter 等工具时,整套系统可以执行 multi-step Deep Research。
Responses API 有网络搜索和代码解释器内置工具。
Responses API 识别
web_search和code_interpreter这类 provider-defined tool types。模型可以选择调用它们,OpenAI 托管的搜索与 container infrastructure 负责真实执行,Service 将结果接回同一 response trajectory。
模型支持自由格式工具调用。
对
type: "custom"的工具,模型可以把 raw text 而不是 JSON object 作为 tool input。这降低了代码、SQL、DSL 的 serialization friction;工具仍由 Client 执行,也仍需要回传 output。
模型引入了意图澄清过程。
ChatGPT Deep Research product 可以在研究前运行 clarification 与 prompt rewriting workflow;Raw Deep Research API 当前不自动包含这两个阶段。API 应用如需同样体验,应在 research request 之前自行实现 preliminary flow。
无需外部编排代码,无需手写 ReAct 循环。
对 Responses API 托管的 tools,Client 无需手写每一次 tool dispatch、result append 和 next-decision inner loop;Responses Service 承担这部分。但 Client 仍负责 clarification、background lifecycle、limits、errors、security 与 citation presentation。Custom tools 仍需 Client loop。
Learning goal:学会把营销式或高度压缩的表述改写成带 ownership 条件的工程语言。
先预测,再展开。 每个答案都要求你用四层图判断,不是回忆页面原句。
0 / 6 已揭示不要马上看答案。先在脑中画四层图。
模型输出:
{"type":"custom_tool_call","name":"python_runner","input":"print(2+2)"}Responses API 会自动执行这段 Python 吗?
不会。type: "custom" 的 freeform input 只是 raw text
payload。Client 必须验证、执行并回传
custom_tool_call_output。如果希望 OpenAI
执行,应使用平台定义的 hosted
code_interpreter,并遵循其配置接口。
使用:
{"type":"web_search"}是不是说明搜索引擎存在于 model weights?
不是。Model 负责生成搜索 action;OpenAI 托管的 Web Search infrastructure 执行搜索。Built-in 指 platform built-in,不是 model-weight built-in。
Client 只调用一次
responses.create(...),是否说明内部只发生一个 action?
不是。一个 response / background job 可以包含多步 model decision 与 hosted tool trajectory。Client request 粒度和内部 action 粒度不同。
Deep Research API 是否会自动问用户偏好的数据源和报告格式?
Raw API 当前不会自动进行 clarification / prompt rewriting。ChatGPT product 可以有这样的 workflow;自建 API 产品需要在研究前实现。
既然 hosted tools 不需要手写 inner loop,应用是不是只剩一行代码?
不是。你仍需处理 task preparation、background job、polling/webhook、timeout、budget、security、prompt injection、citations 与用户体验。
你的公司有私有数据库查询工具
query_revenue(sql)。你只把它声明为
type: "custom",模型产生 SQL。谁应该执行 SQL?需要 Client
loop 吗?
你的 backend 执行,而且必须实施权限、SQL validation、审计和结果回传。因为它不是 Provider 托管工具,Client 仍需参与 inner loop。Freeform 并没有改变 execution ownership。
这些会随着 API 版本变化。稳定的是 ownership 和 information flow。
如果你能自然说出下面这句话,就抓住了实验 1-3:
模型学会决定下一步,Responses Service 托管闭环,Hosted Tools 执行真实动作,Product / Client 管研究前后。
最初:只有一个 LLM call
↓
模型无法接触实时世界
↓
加入 Function Calling
↓
模型能提出 action,但 Client 必须执行并回传
↓
Client 手写 ReAct loop
↓
Tool-use training 让 model policy 更自主
↓
“Native Agent capability” = decision policy 更原生
↓
Provider 提供 Hosted Tools
↓
Responses runtime 托管 action → observation → next decision
↓
Client 不再手写 hosted inner loop
↓
Deep Research model 学会长程搜索、核验与综合
↓
Model + Runtime + Tools + Context = Deep Research system
↓
ChatGPT Product 再加 Clarification / Rewrite / Progress UX
↓
完整的 Deep Research product experience
用 numbered story 再说一遍: