← 全部课程

From ReAct to Hosted Deep Research

一次 request 为什么能完成搜索、阅读、代码分析和再次搜索?从最小手写 loop 开始,把每一层的 owner 找出来。

35–45 MIN · ONE CASE · FROM FIRST PRINCIPLES

不是背六个名词,而是追踪一条 loop 的 ownership 如何迁移。

整页只用一个案例:查询东盟十国首都坐标,并用 Python 找出距离最近的一对。每遇到一个新名词,我们都回到四个问题:谁决定?谁执行?谁保存 observation?谁管理 lifecycle?

核心哲学:实验 1-3 最容易让人误解的地方,是把四个不同层级压缩成一句「模型原生具备 Deep Research 能力」。真正发生的事情是一条逐层迁移:Model 学会选择 action;Responses API runtime 维持 hosted inner loop;Web Search 与 Code Interpreter 执行真实动作;Product / Client 负责澄清需求和管理任务生命周期。 循环没有消失,搜索引擎也没有进入 model weights;变化的是 decision policy 和 loop ownership。


1. 先拆开书中压缩在一起的六句话

实验 1-3 的原文在很短的篇幅内连续做了六次跳跃:

模型有原生 Agent 能力
    ↓
模型能自主调用工具
    ↓
Responses API 有 Web Search / Code Interpreter
    ↓
因此可以 Deep Research
    ↓
还能先做 Intent Clarification
    ↓
所以无需外部编排、无需手写 ReAct loop

每个箭头都需要额外解释。因为这六句话并不属于同一层:

书中说法 实际在谈哪一层 不能误解成什么
原生 Agent / ReAct 能力 Model 的 tool-use decision policy 模型权重里装了浏览器和 Python
Responses API Provider 的 API + agent runtime 单纯的文本生成 endpoint
Built-in tools Provider 托管的 execution infrastructure 模型自己访问互联网
Deep Research Model + runtime + tools + context 的 system capability 一个单独参数或一个工具名称
Intent Clarification 通常是 Product / Client 的 preliminary flow Raw Deep Research API 自动追问
不用手写 loop Hosted tools 的 inner loop 由 service 管 整个应用不再需要 orchestration

关键洞察:真正需要回答的不是「GPT 到底有没有 Agent 能力」,而是:一次任务中的每一个 decision、execution、observation 和 lifecycle action,分别由谁拥有?

Learning goal:先意识到困惑不是你的理解问题,而是原文跨层跳跃太快。接下来我们逐个补上箭头。


2. Baseline:先亲手写一个最小 ReAct Loop

Interactive Baseline · 不懂这一段,后面都只是口号

亲手走完一次 Client-side ReAct loop

点「下一步」观察:Model 只提出 action;Client 执行并把 observation 写回 Context。右下角会显示此刻模型下一轮真正能看到什么。

CLIENTcall / inspect / append
MODELchoose next action
YOUR TOOLsearch or execute
CONTEXTtrajectory + observation
CLIENT

模型下一轮可见状态

2.1 ReAct 不是 React

这里的 ReAct 是 Reasoning + Acting,不是前端框架 React。

它描述一个闭环:

观察当前 Context
    ↓
决定下一步 action
    ↓
执行 action
    ↓
把 observation 放回 Context
    ↓
根据新 Context 再决定

最重要的字不是 Reasoning,也不是 Acting,而是 loop。只要下一步必须依赖刚刚发生的结果,就一定需要某个 runtime 维持这个闭环。

2.2 模型单独调用工具时,实际上只会“提出调用”

假设我们给模型一个 custom function:

{
  "type": "function",
  "name": "search_web",
  "description": "Search the public web",
  "parameters": {
    "type": "object",
    "properties": {
      "query": { "type": "string" }
    },
    "required": ["query"]
  }
}

用户问:

找出东盟 10 国首都之间距离最近的一对。

模型可能输出:

{
  "type": "function_call",
  "name": "search_web",
  "arguments": "{\"query\":\"ASEAN capitals official coordinates\"}",
  "call_id": "call_01"
}

这时什么都还没有被搜索。

模型只生成了一段结构化 token,意思相当于:

「应用,请帮我执行 search_web(query=...)。」

真实网络请求必须由模型之外的程序完成。

2.3 最小的 Client-side ReAct loop

conversation = [user_message]

while True:
    response = call_model(
        input=conversation,
        tools=[search_web, run_python],
    )

    conversation.extend(response.output)

    tool_calls = find_tool_calls(response.output)
    if not tool_calls:
        return response.output_text

    for call in tool_calls:
        result = execute_in_my_application(call)
        conversation.append({
            "type": "function_call_output",
            "call_id": call.call_id,
            "output": result,
        })

逐行看职责:

步骤 谁负责 发生了什么
call_model(...) Client → Model 把 Context 和 tools 交给模型
function_call Model 选择下一步 action
execute_in_my_application Client / Tool 真正执行搜索或代码
function_call_output Client 把 observation 放回 Context
再次 call_model Client → Model 让模型基于新 observation 决定
while True Client 维持 loop,直到没有 tool call

这就是「手写 ReAct loop」最具体的含义。

2.4 为什么模型权重不可能独自完成这个 loop

在工具执行之前,模型不知道工具会返回什么。

Model inference #1
  输出:请搜索 A

现实世界发生搜索
  返回:B

Model inference #2
  输入:刚才返回了 B
  输出:根据 B,下一步执行 C

B 是第一次 inference 结束之后才产生的新信息。因此必须有某个模型外部 runtime:

  1. 截获 action;
  2. 执行 action;
  3. 把 observation 加入 Context;
  4. 继续模型决策。

所以:

“Model-native ReAct” 不可能表示整个循环都存在于 weights 内。它最多表示:模型已经学会在循环中做下一步 decision。

Learning goal:看到一个 tool call 时,立即区分「提出 action」和「执行 action」。这条边界是后面全部概念的地基。


3. The Four-Layer Map:四层不能混在一起

Responsibility Lens · 点击一句宣传语

一句“原生能力”到底跨了几层?

高亮层表示该能力真实依赖的组成部分。没有任何一个复杂 Agent 能力只靠一个框完成。

Layer 4

Product / Client

用户交互、clarification、brief、background job、引用与错误呈现。

Layer 3

Responses Runtime

保存 trajectory、识别 hosted call、路由工具、把 observation 接回循环。

Layer 2

Hosted Tools

真实搜索索引、页面读取、Python container 与文件执行。

Layer 1

Model

读取 Context、选择 next action、生成 query / code、综合报告。

以后遇到任何 Agent 宣传语,先把它放进下面四层。

┌──────────────────────────────────────────────┐
│ Layer 4 · Product / Client                   │
│ 用户交互、Intent Clarification、job 管理、引用展示 │
├──────────────────────────────────────────────┤
│ Layer 3 · Responses API Runtime / Service    │
│ 保存 trajectory、路由 hosted calls、维持 inner loop │
├──────────────────────────────────────────────┤
│ Layer 2 · Hosted Tools                       │
│ Web Search、Code Interpreter、File Search ... │
├──────────────────────────────────────────────┤
│ Layer 1 · Model                              │
│ 读取 Context,选择 next action,生成最终综合       │
└──────────────────────────────────────────────┘

Layer 1 · Model

模型负责:

模型不负责:

Layer 2 · Hosted Tools

Hosted Tool 是 Provider 运行的真实服务:

它们是 infrastructure,不是 weights。

Layer 3 · Responses API Runtime

Responses API 不只是「把 prompt 传给模型」。它还是一个 hosted agent runtime,可以:

Layer 4 · Product / Client

Product / Client 决定完整用户体验:

直观理解:Model 像研究员,Hosted Tools 像数据库和计算实验室,Responses runtime 像负责派单与传递实验结果的研究机构,Product / Client 像接待用户、定义项目和交付报告的咨询团队。

Learning goal:任何“能力”都要问:是 model decision、tool execution、runtime orchestration,还是 product workflow?


4. “Native ReAct” 到底 native 在哪里

4.1 早期做法:开发者写死 workflow

早期 Agent 常把流程写成确定性规则:

search("ASEAN capitals")
open(first_result)
extract_coordinates()
run_python(haversine_code)
write_report()

每一步由开发者预先规定。模型主要负责填内容。

缺点是任务一变,流程就容易失效:

4.2 Model-native tool-use policy

经过 tool-use training / reinforcement learning 后,模型可以学习一个 policy:

π(next_action | goal, available_tools, trajectory)

不用记公式。它只表达:

给定目标、可用工具和已经发生的 trajectory,模型学习选择下一步 action。

被写进模型能力的是:

没有被写进模型 weights 的是:

4.3 为什么有人叫它“原生 ReAct”

因为开发者不再需要用大量 if/else 规定每个 decision:

以前:Developer policy 决定下一步
现在:Model policy 决定下一步

但 loop host 仍然存在:

Model policy         决定 next action
Runtime / Harness    维持 action → observation → next decision
Tool infrastructure  执行 action

准确说法应该是:

模型原生具备参与 ReAct loop 的 decision policy,而不是模型单独拥有整个 ReAct system。

Learning goal:理解 “native” 描述 decision policy 的来源,不描述真实工具的物理位置。


5. Three Tool Types:Function、Custom、Hosted

Tool Taxonomy · 同在 tools[],ownership 完全不同

Function、Custom、Hosted:先问谁执行

Freeform 只是 payload format;Hosted 才说明 execution infrastructure 由 Provider 提供。切换三类工具看不变与变化。

Call envelope
Inner payload
Executor
Inner loop

快速判断:哪一种声明会让 Python 在 OpenAI 托管的 sandbox 中真实执行?

先选一个。关键不是输入看起来像不像代码,而是谁拥有 executor。

哪一个才是 runtime 能识别并 dispatch 的 Custom Tool call?

先区分 ordinary assistant text 与 typed output item。

用户最容易把三种 tool 混在一起。它们看起来都放进 tools=[...],但 execution ownership 完全不同。

类型 输入格式 谁执行 Client 要不要回传结果 示例
Function tool JSON Schema arguments 你的应用 要 get_weather({city})
Custom tool 外层是结构化 custom_tool_call;内层 input 是 freeform text,可加 grammar 你的应用 要 原始 Python / SQL / shell 文本
Hosted tool Provider 定义的固定类型 OpenAI infrastructure 通常不用你逐次执行 web_search、code_interpreter

5.1 Function tool:结构化 JSON

定义:

{
  "type": "function",
  "name": "lookup_coordinates",
  "parameters": {
    "type": "object",
    "properties": {
      "city": { "type": "string" },
      "country": { "type": "string" }
    },
    "required": ["city", "country"],
    "additionalProperties": false
  }
}

模型输出:

{
  "type": "function_call",
  "name": "lookup_coordinates",
  "arguments": "{\"city\":\"Bangkok\",\"country\":\"Thailand\"}",
  "call_id": "call_42"
}

你的应用执行并回传:

{
  "type": "function_call_output",
  "call_id": "call_42",
  "output": "{\"lat\":13.7563,\"lon\":100.5018}"
}

5.2 Custom tool:外层严格,内层 Freeform

先把最容易误解的一句话修正掉:

不是“模型随口说一句自然语言,API 就猜测它要调用工具”。Custom Tool call 仍然有 machine-readable 的严格外层协议;freeform 的只有其中 input 字段。

工具定义本身首先是结构化的:

{
  "type": "custom",
  "name": "python_runner",
  "description": "Run Python supplied as plain text"
}

模型决定调用后,Responses API 返回的不是普通 assistant prose,而是一个有明确 type 的 output item:

{
  "id": "ctc_123",
  "type": "custom_tool_call",
  "status": "completed",
  "call_id": "call_42",
  "name": "python_runner",
  "input": "print(2 + 2)"
}

把它拆成两层就不会混淆:

Responses API output item
┌────────────────────────────────────────────┐
│ Strict call envelope                      │
│                                            │
│ type:    "custom_tool_call"               │
│ name:    "python_runner"                  │
│ call_id: "call_42"                        │
│ input:   ───────────────────────────────┐  │
└─────────────────────────────────────────│──┘
                                          │
                                          ▼
                            Freeform payload string
                            "print(2 + 2)"
层 谁规定格式 是否严格
Call envelope Responses API protocol 严格。必须能识别 type、name、call_id、input
Payload slot Custom Tool contract API 不要求它是 JSON object,只要求它是一个 string
Payload language 你的 executor 仍可能很严格:Python、SQL、shell、DSL 或 natural language

所以,Freeform 不等于 No Format:

Freeform
≠ 整个调用没有结构
≠ 随便写什么工具都能理解

Freeform
= Responses API 不再要求 input 遵循 JSON Schema
= input 可以直接承载工具原生需要的文本

自然语言什么时候才可以?

只有当工具实现本身就接受 natural language 时,input 才能是自然语言。

Custom Tool 合法 payload “帮我算一下”是否有效
python_runner Python source code 通常无效;executor 需要 Python
sql_runner SQL statement 通常无效;executor 需要 SQL
shell_runner shell command 通常无效;executor 需要 shell syntax
instruction_router natural-language instruction 可以,因为工具被设计成读取自然语言

换句话说,Custom Tool 并没有让所有工具突然“听懂人话”。它只是让工具不必先把原生文本塞进 JSON arguments。

如果确实需要强制 payload 格式

Custom Tool 可以附加 grammar。例如只允许一个极小的算术表达式语言:

{
  "type": "custom",
  "name": "calculator",
  "description": "Evaluate a basic arithmetic expression",
  "format": {
    "type": "grammar",
    "syntax": "regex",
    "definition": "[0-9]+([+*][0-9]+)*"
  }
}

此时:

如果模型只是输出普通文本:

请调用 python_runner 帮我计算 2 + 2。

这只是 assistant message,不是 custom_tool_call;runtime 没有机器可读的 call item,就不会把它当成工具调用 dispatch。

你的 Client 真正处理的是下面这条闭环:

custom_tool_call envelope
        ↓ extract name + input
Client validates input
        ↓
Client dispatches python_runner("print(2 + 2)")
        ↓
{"type":"custom_tool_call_output","call_id":"call_42","output":"4"}
        ↓
Model reads observation and continues

type: "custom" 仍然是你的工具。Responses API 只把“这是一次工具调用”编码成结构化 item,不会因为 input 里出现 Python 就自动执行它。

Learning goal:看到 Freeform Custom Tool 时,能立即说出:outer envelope is structured; only inner payload is freeform; executor syntax and Client validation still apply.

5.3 Hosted tool:Provider 托管

{
  "type": "web_search"
}

或:

{
  "type": "code_interpreter",
  "container": { "type": "auto" }
}

这里不需要提供你的函数地址或 executor。因为 OpenAI 已经实现并运营执行环境。

Learning goal:看到 tools 数组时,不要只看“都是工具”;第一眼先看 type,再判断 execution ownership。


6. Freeform Tool Calling 只改变了什么

6.1 它解决的是 serialization friction

如果工具参数是一大段 Python:

print("hello")
path = "C:\\temp\\data.csv"

硬塞进 JSON arguments,需要大量 escaping:

{
  "code": "print(\"hello\")\npath = \"C:\\\\temp\\\\data.csv\""
}

Freeform Custom Tool 允许 custom_tool_call.input 直接承载原始文本:

{
  "type": "custom_tool_call",
  "name": "python_runner",
  "call_id": "call_42",
  "input": "print(\"hello\")\npath = \"C:\\temp\\data.csv\""
}

这里 wire-level output item 仍然是结构化 JSON;减少的只是 input 内部再包一层 {"code": ...} 的 serialization friction。Python、SQL、shell、DSL 仍须符合各自 executor 的语法。

6.2 它没有改变的四件事

Freeform 并不意味着:

它只改变:

工具调用的 payload 格式

JSON object
    ↓
raw text

它没有改变:

工具归谁所有
谁执行
谁回传结果
谁维持 loop

6.3 三个概念不要串联错

Freeform Tool Calling
    = 参数表示方式

Hosted Tool
    = 执行基础设施归 Provider

Managed Inner Loop
    = action / observation 循环归 Service

三个概念可以同时出现,也可以完全独立。

例如:

Learning goal:能用一句话解释:Freeform 是 syntax / serialization 特性,不是 Agent architecture 特性。


7.1 因为 Responses API 是 Provider-controlled runtime

你的请求:

response = client.responses.create(
    model="gpt-5.6-sol",
    input="找一条最近的正面新闻并引用来源。",
    tools=[{"type": "web_search"}],
)

{"type": "web_search"} 不是把搜索代码发送给模型。它更像打开一个 capability flag:

本次 response 允许模型申请使用 OpenAI 托管的 Web Search。

OpenAI 同时控制:

因此 service 可以在内部完成:

Model 生成 web_search_call
        ↓
Responses runtime 识别 hosted call
        ↓
OpenAI Web Search 执行 search / open_page / find_in_page
        ↓
结果回到 response trajectory
        ↓
Model 基于结果继续

Client 不需要收到 query 后自己去调用搜索 API。

7.2 你能看到的 output items

概念化的 response.output:

[
  {
    "type": "web_search_call",
    "action": {
      "type": "search",
      "query": "ASEAN capitals official coordinates"
    }
  },
  {
    "type": "web_search_call",
    "action": {
      "type": "open_page",
      "url": "https://..."
    }
  },
  {
    "type": "message",
    "content": [
      {
        "type": "output_text",
        "text": "...",
        "annotations": ["url_citation ..."]
      }
    ]
  }
]

具体字段会随 SDK 版本变化,但 ownership 不变:

7.3 “内置”不等于“模型内置”

准确说法:

Built into the Responses platform

而不是:

Built into model weights

Learning goal:理解 API 之所以能内置搜索,是因为 Provider 不只提供模型,还运营搜索与 orchestration runtime。


8. Responses API 为什么能“内置” Code Interpreter

8.1 Code Interpreter 是 hosted sandbox

请求声明:

response = client.responses.create(
    model="gpt-5.6-sol",
    input="用 Python 计算所有东盟首都对的大圆距离。",
    tools=[{
        "type": "code_interpreter",
        "container": {"type": "auto"}
    }],
)

这里的 container: auto 表示 Provider 创建或复用一个受控 container。

模型会做两件事:

  1. 决定需要用 Python;
  2. 生成要运行的代码。

Hosted infrastructure 做另外几件事:

  1. 准备隔离环境;
  2. 执行代码;
  3. 捕获 stdout、错误和文件;
  4. 将 execution result 交回 trajectory。

8.2 它和 custom python_runner 的本质区别

问题 Custom python_runner Hosted code_interpreter
谁定义工具 你 OpenAI
谁提供 sandbox 你 OpenAI
谁执行代码 你的 Client / backend OpenAI infrastructure
谁回传 result 你的 Client Responses runtime
Client 是否手写 inner loop 是 对该 hosted loop 通常不需要
参数能否是代码文本 可以 模型也会生成代码,但接口由平台定义

8.3 Code Interpreter 不是“模型会算”

如果让模型直接心算所有城市组合,它可能:

Code Interpreter 把 deterministic computation 交给真实 runtime:

Model 决定算法与生成代码
        ↓
Python runtime 负责确定性执行
        ↓
Model 解释结果与限制

Learning goal:把「会写代码」和「有地方执行代码」分开。


9. One Request, Many Steps:托管的 Inner Loop

One Request · Thirteen Observable Responsibilities

把一次 Deep Research request 展开

左侧是完整顺序;右侧解释当前 step 的 owner、可观察 item,以及为什么它会引出下一步。Client 没有逐步 dispatch hosted tools。

Observable contract
Loop boundary
为什么下一步成立:

9.1 一个 HTTP request 不等于一个 action

这是实验 1-3 最核心、也最容易被一句话带过的地方。

Client 可能只写一次:

response = client.responses.create(...)

但 Service 内部可以产生一条长 trajectory:

1. Model: 先搜索首都名单
2. Tool:  web_search.search
3. Model: 打开权威来源
4. Tool:  web_search.open_page
5. Model: 在页面中找坐标
6. Tool:  web_search.find_in_page
7. Model: 坐标不全,再搜索
8. Tool:  web_search.search
9. Model: 数据足够,生成 Python
10. Tool: code_interpreter_call
11. Model: 检查结果并写报告

这并不矛盾:

Client 视角:一次 request / 一个 background job
Service 视角:多次 decision + tool action + observation

9.2 Loop 到底移到了哪里

手写 custom loop:

Client
  ├─ call model
  ├─ inspect tool call
  ├─ execute tool
  ├─ append output
  └─ repeat

Hosted tools:

Client
  ├─ submit request
  ├─ poll / wait webhook
  └─ render result

Responses Service
  ├─ get model decision
  ├─ execute hosted tool
  ├─ append observation
  └─ continue inner trajectory

因此准确结论是:

Client-written inner loop 被 service-managed inner loop 替代。Loop 的 ownership 迁移了,loop 本身没有消失。

9.3 这是不是一定意味着多次独立 model API call?

对 Client 来说,不需要把它实现成多次 API call。Provider 内部如何组织 inference、continuation 与 tool round,是服务实现细节。

我们能可靠讨论的是 observable contract:

不要把看不到的 provider internals 伪装成确定事实。

Learning goal:理解「一次 request」与「一次 action / 一次 decision」不是同一粒度。


10. Native Deep Research 是一个 System Capability

10.1 普通工具调用不自动等于 Deep Research

一个模型调用一次 web_search,然后根据第一条结果回答,只能叫 web-enabled answering,不能自动叫 Deep Research。

Deep Research 通常需要:

10.2 能力公式

Deep Research capability
=
long-horizon research policy
+ persistent trajectory / context
+ retrieval tools
+ analysis runtime
+ loop host
+ source-aware synthesis

拿掉任何一项:

拿掉什么 会发生什么
Research policy 可能只搜一次就草率回答
Persistent trajectory 忘记查过什么、重复搜索
Web / file / MCP sources 只能依赖训练时知识
Code Interpreter 数据计算变得脆弱或不可复核
Loop host 模型提出第一次 tool call 后就停住
Citation synthesis 报告难以追溯证据

10.3 “Native Deep Research”的三个可能含义

Model-native

模型经过训练,具有更强的长程 research policy:会规划、搜集、检查、再搜索和综合。

Platform-native

Responses API 直接提供 compatible retrieval / analysis tools,并托管 inner trajectory。

Product-native

ChatGPT 等产品把 clarification、progress、background execution、report presentation 组成完整体验。

书中一句「GPT-5.6 原生 Deep Research」实际上把这三种 native 叠在了一起。

更准确的说法是:

GPT-5.6 的 research policy,加上 Responses API 的 hosted loop 和工具基础设施,共同形成了可直接调用的 Deep Research system capability。

Learning goal:以后看到“原生 Deep Research”,马上追问它是在说 model、platform,还是 product。


11. Intent Clarification 到底从哪里来

Product ≠ Raw API · 同一句“会澄清”有三种实现

Intent Clarification 放在哪一层?

切换视角。Raw API 从你给的 input 开始研究;ChatGPT product 或你自己的 app 可以在它之前增加 preliminary flow。

11.1 为什么研究前需要澄清

用户说:

分析最近一个月的比特币。

至少缺少:

Deep Research 越努力,模糊目标造成的浪费越大。

11.2 ChatGPT Product Flow

产品可以实现:

Raw user request
    ↓
Model identifies missing requirements
    ↓
Product asks user 2–4 questions
    ↓
User answers
    ↓
Model rewrites a research brief
    ↓
Product starts Deep Research

这看起来像「Deep Research 会澄清意图」,因为用户只看到一个统一产品。

11.3 Raw Deep Research API

当前官方 guide 明确指出:

Deep research via the Responses API does not include
clarification or prompt rewriting.

Raw API 的行为更接近:

你给什么 research instructions
        ↓
它就从这些 instructions 开始研究

如果你希望 API 应用也有 clarification,需要自己增加 preliminary flow:

questions = fast_model.find_missing_requirements(user_request)
answers = ask_user(questions)
research_brief = fast_model.rewrite(user_request, answers)
research_job = start_deep_research(research_brief)

11.4 这种能力“怎么会出现”

Intent Clarification 不需要一个神秘的新神经网络模块。它可以来自:

  1. 通用模型理解自然语言与发现歧义的能力;
  2. 一段专门的 developer prompt;
  3. Product / Client 状态机;
  4. 用户回答被保存并合并;
  5. 再把完整 brief 交给 Deep Research model。

也就是说:

Language reasoning capability
+ clarification prompt
+ product state
+ user interaction
=
Intent Clarification experience

它通常是 workflow capability,不是 raw Deep Research endpoint 自动携带的阶段。

Learning goal:明确区分「ChatGPT 产品表现出来的能力」和「Responses API contract 保证的能力」。


12. “无需外部编排代码”成立的准确条件

12.1 三种场景对比

场景 Inner loop 谁写 工具谁执行 Client 仍负责什么
Function tool Client Client 全部 loop + lifecycle
Freeform custom tool Client Client 全部 loop + lifecycle
Responses hosted tools Responses Service OpenAI outer lifecycle
ChatGPT Deep Research product Product + Responses Service OpenAI 最终用户只使用产品

12.2 Hosted tools 省掉了哪些代码

你通常不再需要写:

while tool_call_exists:
    inspect_tool_call()
    execute_search_or_python()
    append_tool_output()
    call_model_again()

因为这部分由 Responses Service 与 hosted tools 协作完成。

12.3 仍然需要哪些代码

如果你在构建自己的应用,通常仍需:

输入与权限校验
Intent Clarification(如需要)
Research brief 构建
background job 提交
polling 或 webhook
timeout / retry / cancel
max_tool_calls / cost budget
错误和部分结果处理
citation rendering
日志、安全与 prompt injection 防护

所以书中「无需外部编排代码」如果不加限定,会误导。

准确版本:

使用 Responses API 的 hosted tools 时,开发者无需为这些 hosted tools 手写 search–read–analyze 的 inner ReAct loop;但仍需编排任务之前和任务之外的 product lifecycle。Custom tools 仍需 Client 执行并回传结果。

Learning goal:能把 orchestration 切成 inner trajectory 和 outer lifecycle,而不是笼统说“有”或“没有”。


13. End-to-End Trace:东盟首都距离案例

现在把所有概念放进同一个案例。

13.1 Raw request

找出东盟 10 国首都之间距离最近的一对。

13.2 Optional clarification — Product / Client

应用发现以下不确定性:

“首都”是否指当前国家首都?
距离是否指经纬度大圆距离?
坐标来源需要什么可信度?
结果需要方法、代码和 citations 吗?

用户回答后,应用生成 brief:

目标:比较东盟 10 个成员国当前首都之间的大圆距离。
来源:优先政府、国际组织或权威地理来源;保留每个坐标的 citation。
方法:统一使用城市中心经纬度和 Haversine formula;用 Python 枚举所有城市对。
输出:最近的一对、距离、计算方法、来源、数据冲突与限制。

这一步不是 Raw Deep Research API 自动完成的。

13.3 Submit request — Client → Responses API

response = client.responses.create(
    model="gpt-5.6-sol",
    background=True,
    input=research_brief,
    tools=[
        {"type": "web_search"},
        {
            "type": "code_interpreter",
            "container": {"type": "auto"},
        },
    ],
    max_tool_calls=40,
)

这段代码只声明:

它没有手写研究步骤。

13.4 Hosted trajectory

Step Decision / Action Owner 可观察对象
1 解析 brief,决定先收集首都名单 Model reasoning summary(如启用)
2 搜索东盟成员与首都权威来源 Model → Service web_search_call: search
3 真正运行搜索 Hosted Web Search search results
4 决定打开候选来源 Model next action
5 打开与读取页面 Hosted Web Search open_page
6 在页面定位坐标字段 Hosted Web Search find_in_page
7 检查缺失或冲突坐标,决定补搜 Model revised search action
8 获取足够的十组坐标 Service + Tool accumulated trajectory
9 决定用 Haversine 枚举城市对 Model code plan / call
10 生成并运行 Python Hosted Code Interpreter code_interpreter_call
11 返回距离列表或计算文件 Code Interpreter execution output
12 检查单位、异常与来源一致性 Model summary / possible new call
13 生成最终报告与 citations Model + Service final message

13.5 Client 看到什么

Client 不必在每个 Step 之间发起自定义 executor,但仍需管理 job:

while response.status in {"queued", "in_progress"}:
    sleep(2)
    response = client.responses.retrieve(response.id)

render_report(response.output_text)
render_citations(response.output)

这里的 while 是 job lifecycle polling,不是 search / code 的 ReAct inner loop。

13.6 如果换成 custom tools 会怎样

如果你不用 hosted web_search,而是声明自己的 search_web function:

Model 输出 function_call
    ↓
你的 Client 调搜索 API
    ↓
你的 Client 回传 function_call_output
    ↓
你的 Client 再次调用 Responses API

这时 inner loop 又回到 Client。

关键洞察:是不是要手写 loop,不由“模型聪不聪明”单独决定,而由 tool execution ownership + API runtime contract 决定。

Learning goal:能够逐步指认完整 trajectory 中每一步的 owner,而不是只记住一句“API 自动研究”。


14. 把书中的关键表述重新翻译成准确语言

原表述 1

GPT-5.6 具有原生 Deep Research 能力。

更准确

GPT-5.6 具备经过训练的长程 research / tool-use policy;当它运行在 Responses API 的 hosted runtime 中,并获得 Web Search、Code Interpreter 等工具时,整套系统可以执行 multi-step Deep Research。


原表述 2

Responses API 有网络搜索和代码解释器内置工具。

更准确

Responses API 识别 web_search 和 code_interpreter 这类 provider-defined tool types。模型可以选择调用它们,OpenAI 托管的搜索与 container infrastructure 负责真实执行,Service 将结果接回同一 response trajectory。


原表述 3

模型支持自由格式工具调用。

更准确

对 type: "custom" 的工具,模型可以把 raw text 而不是 JSON object 作为 tool input。这降低了代码、SQL、DSL 的 serialization friction;工具仍由 Client 执行,也仍需要回传 output。


原表述 4

模型引入了意图澄清过程。

更准确

ChatGPT Deep Research product 可以在研究前运行 clarification 与 prompt rewriting workflow;Raw Deep Research API 当前不自动包含这两个阶段。API 应用如需同样体验,应在 research request 之前自行实现 preliminary flow。


原表述 5

无需外部编排代码,无需手写 ReAct 循环。

更准确

对 Responses API 托管的 tools,Client 无需手写每一次 tool dispatch、result append 和 next-decision inner loop;Responses Service 承担这部分。但 Client 仍负责 clarification、background lifecycle、limits、errors、security 与 citation presentation。Custom tools 仍需 Client loop。

Learning goal:学会把营销式或高度压缩的表述改写成带 ownership 条件的工程语言。


15. Prediction Checks

先预测,再展开。 每个答案都要求你用四层图判断,不是回忆页面原句。

0 / 6 已揭示

不要马上看答案。先在脑中画四层图。

Q1

模型输出:

{"type":"custom_tool_call","name":"python_runner","input":"print(2+2)"}

Responses API 会自动执行这段 Python 吗?

查看答案

不会。type: "custom" 的 freeform input 只是 raw text payload。Client 必须验证、执行并回传 custom_tool_call_output。如果希望 OpenAI 执行,应使用平台定义的 hosted code_interpreter,并遵循其配置接口。

Q2

使用:

{"type":"web_search"}

是不是说明搜索引擎存在于 model weights?

查看答案

不是。Model 负责生成搜索 action;OpenAI 托管的 Web Search infrastructure 执行搜索。Built-in 指 platform built-in,不是 model-weight built-in。

Q3

Client 只调用一次 responses.create(...),是否说明内部只发生一个 action?

查看答案

不是。一个 response / background job 可以包含多步 model decision 与 hosted tool trajectory。Client request 粒度和内部 action 粒度不同。

Q4

Deep Research API 是否会自动问用户偏好的数据源和报告格式?

查看答案

Raw API 当前不会自动进行 clarification / prompt rewriting。ChatGPT product 可以有这样的 workflow;自建 API 产品需要在研究前实现。

Q5

既然 hosted tools 不需要手写 inner loop,应用是不是只剩一行代码?

查看答案

不是。你仍需处理 task preparation、background job、polling/webhook、timeout、budget、security、prompt injection、citations 与用户体验。

Q6 · Transfer

你的公司有私有数据库查询工具 query_revenue(sql)。你只把它声明为 type: "custom",模型产生 SQL。谁应该执行 SQL?需要 Client loop 吗?

查看答案

你的 backend 执行,而且必须实施权限、SQL validation、审计和结果回传。因为它不是 Provider 托管工具,Client 仍需参与 inner loop。Freeform 并没有改变 execution ownership。


16. What You Actually Need to Remember

Must understand

Understand conceptually, but do not over-memorize

这些会随着 API 版本变化。稳定的是 ownership 和 information flow。

用一句话自测

如果你能自然说出下面这句话,就抓住了实验 1-3:

模型学会决定下一步,Responses Service 托管闭环,Hosted Tools 执行真实动作,Product / Client 管研究前后。


17. The Full Conceptual Line

最初:只有一个 LLM call
    ↓
模型无法接触实时世界
    ↓
加入 Function Calling
    ↓
模型能提出 action,但 Client 必须执行并回传
    ↓
Client 手写 ReAct loop
    ↓
Tool-use training 让 model policy 更自主
    ↓
“Native Agent capability” = decision policy 更原生
    ↓
Provider 提供 Hosted Tools
    ↓
Responses runtime 托管 action → observation → next decision
    ↓
Client 不再手写 hosted inner loop
    ↓
Deep Research model 学会长程搜索、核验与综合
    ↓
Model + Runtime + Tools + Context = Deep Research system
    ↓
ChatGPT Product 再加 Clarification / Rewrite / Progress UX
    ↓
完整的 Deep Research product experience

用 numbered story 再说一遍:

  1. Function Calling 先让模型能够表达「我希望外部做什么」。
  2. Client-side ReAct loop 负责真的执行并把结果送回去。
  3. Tool-use training 把 next-action policy 更多地交给模型,而不是开发者写死 workflow。
  4. Hosted Tools 把具体 execution infrastructure 交给 Provider。
  5. Responses runtime 把 hosted tool 的 inner loop 也托管起来。
  6. Deep Research model 在这个环境里执行更长、更系统的研究策略。
  7. Intent Clarification 是研究开始之前的 product/client workflow,并非 Raw API 自动步骤。
  8. 因此“无需手写 loop”只对 hosted inner loop 成立,不对整个应用成立。

18. Key Takeaways


Primary Sources