docs: LLM 错误码对齐实施计划(4 任务 TDD 分解)

This commit is contained in:
lhl
2026-08-09 13:28:45 +08:00
parent 1a3c829d4c
commit a0a40afd45
@@ -0,0 +1,391 @@
# LLM 错误码对齐(error_code)实施计划
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** 在 engine 层暴露结构化 `error_code`(对齐 api-design §7),使调用方/UI 可依据错误码执行 retry/skip/abort/配置决策。
**Architecture:** 单一事实源 = `LLMError.error_code` 类属性(五异常覆写);`ChatResult`/`StructuredResult` 增加 `error_code` 字段透传;engine 失败路径取最后一次异常的同源 error_codeapi-design §7 补 `LLM_NETWORK_ERROR` 行。
**Tech Stack:** Python 3.11+、pytest(覆盖红线 fail_under=99)、dataclass
## Global Constraints
- 项目为中文交流:注释、docstring、commit 消息用中文;标识符英文
- 零真实网络:单测一律注入 FakeLLMClient / MockTransport`DEEPSEEK_API_KEY` 不落代码
- 覆盖红线:`pyproject.toml` `fail_under = 99`(当前基线 119 passed / 100.00%),新增分支需足量测试
- 全量回归:`python -m pytest -v`;聚焦:`python -m pytest tests/<file> -v`
- 提交消息:`feat:` / `test:` / `docs:` + 简中文描述
- 每次修改后追加 `_AI_USAGE_LOG.md`(范式步骤:"Agent 实现"/"测试验证",模型:"deepseek-v4-flash-free"
- 本仓库工作在 master 分支(项目惯例);Windows/PowerShell 环境
- Spec: `docs/superpowers/specs/2026-08-09-llm-errorcode-alignment-design.md`
---
### Task 1: 异常树携带 error_codeexceptions.py
**Files:**
- Modify: `src/genesis/inference/exceptions.py`
- Test: `tests/test_inference_errors.py`
**Interfaces:**
- Consumes: 现有 `LLMError`/`LLMNetworkError`/`LLMTimeoutError`/`LLMNotConfiguredError`/`LLMResponseError`(已存在)
- Produces: 每个异常类新增类属性 `error_code: str | None`——`LLMNetworkError.error_code == "LLM_NETWORK_ERROR"``LLMTimeoutError.error_code == "LLM_TIMEOUT"``LLMNotConfiguredError.error_code == "LLM_NOT_CONFIGURED"``LLMResponseError.error_code == "LLM_PARSE_ERROR"`、基类 `LLMError.error_code is None`
- [ ] **Step 1: 写失败测试**
`tests/test_inference_errors.py` 追加(import 上方 `LLMError` 已引入):
```python
def test_error_code_class_attributes():
# error_code 与 api-design §7 错误码表一一对应(机器可读事实源)
assert LLMError.error_code is None
assert LLMTimeoutError.error_code == "LLM_TIMEOUT"
assert LLMNetworkError.error_code == "LLM_NETWORK_ERROR"
assert LLMNotConfiguredError.error_code == "LLM_NOT_CONFIGURED"
assert LLMResponseError.error_code == "LLM_PARSE_ERROR"
def test_error_code_inherited_by_instance():
e = LLMNetworkError("network")
assert e.error_code == "LLM_NETWORK_ERROR"
```
- [ ] **Step 2: 运行验证失败**
Run: `python -m pytest tests/test_inference_errors.py -v`
Expected: FAIL`AttributeError: type object 'LLMError' has no attribute 'error_code'`
- [ ] **Step 3: 实现**
`src/genesis/inference/exceptions.py` 中为每个类加类属性:
```python
class LLMError(Exception):
"""LLM 调用相关的异常基类(api-design §7 映射基底)"""
error_code: str | None = None # api §7 错误码(机器可读);新增子类必须覆写
class LLMNetworkError(LLMError):
"""网络失败 / 5xx 重试耗尽(可重试语义)"""
error_code = "LLM_NETWORK_ERROR"
class LLMTimeoutError(LLMError):
"""LLM 调用超时(api-error: LLM_TIMEOUT 502"""
error_code = "LLM_TIMEOUT"
class LLMNotConfiguredError(LLMError):
"""Key / 模型缺失(api-error: LLM_NOT_CONFIGURED 503"""
error_code = "LLM_NOT_CONFIGURED"
class LLMResponseError(LLMError):
"""响应结构损坏(JSON 解析失败等)"""
error_code = "LLM_PARSE_ERROR"
```
- [ ] **Step 4: 运行验证通过**
Run: `python -m pytest tests/test_inference_errors.py -v`
Expected: PASS4 passed
- [ ] **Step 5: 追加 `_AI_USAGE_LOG.md`**
范式步骤:"Agent 实现",摘要:异常树 error_code 类属性(五类型 + 基类 None)。
- [ ] **Step 6: 提交**
```bash
git add src/genesis/inference/exceptions.py tests/test_inference_errors.py _AI_USAGE_LOG.md
git commit -m "feat: 异常树 error_code 类属性(对齐 api-design §7 错误码)"
```
---
### Task 2: 结果对象 error_code 字段(types.py
**Files:**
- Modify: `src/genesis/inference/types.py`
- Test: `tests/test_inference_types.py`
**Interfaces:**
- Consumes: 现有 `ChatResult` / `StructuredResult` dataclassTask 1 的 error_code 常量不直接依赖,但语义对应)
- Produces: `ChatResult.error_code: str | None = None``StructuredResult.error_code: str | None = None`(均默认 None,向后兼容)
- [ ] **Step 1: 写失败测试**
`tests/test_inference_types.py` 追加:
```python
def test_chat_result_error_code_default_none():
r = ChatResult(
text="t", model="m", prompt_version="v1", usage=TokenUsage(),
duration_ms=10, status="ok",
)
assert r.error_code is None
def test_structured_result_error_code_default_none():
r = StructuredResult(
data={"a": 1}, raw_text='{"a":1}', parse_attempts=1, model="m",
prompt_version="v1", usage=TokenUsage(), duration_ms=10, status="ok",
)
assert r.error_code is None
```
- [ ] **Step 2: 运行验证失败**
Run: `python -m pytest tests/test_inference_types.py -v`
Expected: FAIL`TypeError: __init__() got an unexpected keyword argument 'error_code'` 或 AttributeError——取决于 dataclass 是否带默认;实际为 `AttributeError: 'ChatResult' object has no attribute 'error_code'`
- [ ] **Step 3: 实现**
`src/genesis/inference/types.py` 两个 dataclass 的 `error` 字段后追加:
```python
@dataclass
class ChatResult:
"""chat() 的返回值"""
text: str
model: str
prompt_version: str
usage: TokenUsage
duration_ms: int
status: Literal["ok", "fallback", "failed"]
error: str | None = None
error_code: str | None = None # 失败时的 api §7 错误码;成功为 None
@dataclass
class StructuredResult:
"""chat_structured() 的返回值(补丁 1:含 status 字段)"""
data: dict
raw_text: str
parse_attempts: int
model: str
prompt_version: str
usage: TokenUsage
duration_ms: int
status: Literal["ok", "fallback", "parse_error", "failed"]
error: str | None = None
error_code: str | None = None # 失败/parse_error 时的错误码;成功为 None
```
- [ ] **Step 4: 运行验证通过**
Run: `python -m pytest tests/test_inference_types.py -v`
Expected: PASS8 passed
- [ ] **Step 5: 追加 `_AI_USAGE_LOG.md`**
范式步骤:"Agent 实现",摘要:ChatResult/StructuredResult 增加 error_code 字段(默认 None 向后兼容)。
- [ ] **Step 6: 提交**
```bash
git add src/genesis/inference/types.py tests/test_inference_types.py _AI_USAGE_LOG.md
git commit -m "feat: 结果对象 error_code 字段(默认 None 向后兼容)"
```
---
### Task 3: engine 透传 error_code(同源语义)
**Files:**
- Modify: `src/genesis/inference/engine.py`
- Test: `tests/test_inference_engine.py`
**Interfaces:**
- Consumes: Task 1 `LLMError.error_code`Task 2 `ChatResult.error_code` / `StructuredResult.error_code`;现有 `FakeLLMClient``tests/inference_helpers.py`,脚本 `("ok"|"raise_timeout"|"raise_network"|"parse_fail", content)`
- Produces: `chat()` 失败路径 `ChatResult.error_code``error` 同源(取最后一次 LLMError);`chat_structured()` 失败/parse_error 路径同源透传
- [ ] **Step 1: 写失败测试**
`tests/test_inference_engine.py``test_chat_all_failed_returns_failed` 附近追加(文件已 import `FakeLLMClient`/`InferenceEngine`/`Models`):
```python
def test_chat_failed_error_code_timeout():
# 主模型超时 + 备用也失败:error_code 与 error 同源(取最后一次异常)
client = FakeLLMClient([("raise_timeout", ""), ("raise_timeout", "")])
eng = InferenceEngine(client=client, models=Models())
r = eng.chat(session_id="s1", prompt=Prompt(name="p", version="v1", template="t"), variables={})
assert r.status == "failed"
assert r.error_code == "LLM_TIMEOUT"
def test_chat_failed_error_code_network_last():
# 主模型超时(第一次)、备用网络失败(最后一次)→ error_code 取最后一次 = LLM_NETWORK_ERROR
client = FakeLLMClient([("raise_timeout", ""), ("raise_network", "")])
eng = InferenceEngine(client=client, models=Models())
r = eng.chat(session_id="s1", prompt=Prompt(name="p", version="v1", template="t"), variables={})
assert r.status == "failed"
assert r.error_code == "LLM_NETWORK_ERROR"
def test_chat_failed_error_code_not_configured():
# 备用模型 Key 缺失(最后一次)→ LLM_NOT_CONFIGURED
client = FakeLLMClient([("raise_network", ""), ("raise_timeout", "")])
# 用自定义异常客户端模拟 NotConfigured
class NotConfiguredClient:
def __init__(self):
self.calls = []
def chat(self, *, model, messages, temperature, max_tokens):
self.calls.append(model)
from genesis.inference.exceptions import LLMNotConfiguredError
raise LLMNotConfiguredError("no key")
eng = InferenceEngine(client=NotConfiguredClient(), models=Models())
r = eng.chat(session_id="s1", prompt=Prompt(name="p", version="v1", template="t"), variables={})
assert r.status == "failed"
assert r.error_code == "LLM_NOT_CONFIGURED"
def test_chat_structured_parse_error_code():
client = FakeLLMClient([("parse_fail", ""), ("parse_fail", "")])
eng = InferenceEngine(client=client, models=Models())
r = eng.chat_structured(
session_id="s1", prompt=Prompt(name="p", version="v1", template="提取"),
variables={}, schema={}, retry_count=1,
)
assert r.status == "parse_error"
assert r.error_code == "LLM_PARSE_ERROR"
def test_chat_structured_failed_error_code_network():
client = FakeLLMClient([("raise_network", "")])
eng = InferenceEngine(client=client, models=Models())
r = eng.chat_structured(
session_id="s1", prompt=Prompt(name="p", version="v1", template="提取"),
variables={}, schema={},
)
assert r.status == "failed"
assert r.error_code == "LLM_NETWORK_ERROR"
```
> 说明:`test_chat_failed_error_code_not_configured` 使用本地 NotConfiguredClientFakeLLMClient 脚本不含该异常);文件需已 import `Prompt``from genesis.inference.types import Prompt`——若 Task 2 前是 `ChatMessage, Prompt`,本计划沿用 `Prompt`)。
- [ ] **Step 2: 运行验证失败**
Run: `python -m pytest tests/test_inference_engine.py -v`
Expected: FAIL(新增用例中 `r.error_code``None` 或 AttributeError——取决于 engine 是否已透传;当前 engine 未透传,断言 `== "LLM_TIMEOUT"` 失败)
- [ ] **Step 3: 实现**
`src/genesis/inference/engine.py`
`chat()` 的失败循环增加 error_code 记录(第 103 行附近):
```python
start = time.monotonic()
last_error: str | None = None
last_error_code: str | None = None
for idx, name in enumerate(self._model_names(model)):
try:
...
except LLMError as exc:
last_error = str(exc)
last_error_code = exc.error_code # 同源:取最后一次失败异常
```
并更新失败返回值(第 119-124 行):
```python
return ChatResult(
text="", model=name,
prompt_version=getattr(prompt, "version", "inline"),
usage=TokenUsage(), duration_ms=int((time.monotonic() - start) * 1000),
status="failed", error=last_error, error_code=last_error_code,
)
```
`chat_structured()`
- 循环前 `last_error` 旁加 `last_error_code: str | None = None`
- `except json.JSONDecodeError` 分支:`last_error_code = "LLM_PARSE_ERROR"`
- `except LLMError as exc` 分支的返回值:`status="failed", error=str(exc), error_code=exc.error_code`
- 循环后 parse_error 返回值:`status="parse_error", error=last_error, error_code=last_error_code`
- [ ] **Step 4: 运行验证通过**
Run: `python -m pytest tests/test_inference_engine.py -v`
Expected: PASS22 passed
- [ ] **Step 5: 追加 `_AI_USAGE_LOG.md`**
范式步骤:"Agent 实现",摘要:engine 失败/parse_error 路径透传 error_code(与 error 同源取最后一次异常)。
- [ ] **Step 6: 提交**
```bash
git add src/genesis/inference/engine.py tests/test_inference_engine.py _AI_USAGE_LOG.md
git commit -m "feat: engine 透传 error_code(与 error 同源取最后一次异常)"
```
---
### Task 4: api-design §7 文档补丁
**Files:**
- Modify: `docs/api-design.md`(§7 错误码表)
- Test: 无(纯文档;回归验证全量)
**Interfaces:**
- Consumes: Task 1 的 `LLM_NETWORK_ERROR` / `LLM_TIMEOUT` / `LLM_PARSE_ERROR` / `LLM_NOT_CONFIGURED` 错误码
- Produces: api §7 表五条 LLM 错误码全齐,来源列与异常类型一一对应
- [ ] **Step 1: 修改错误码表**
`docs/api-design.md` §7 表格的 `LLM_TIMEOUT` 行后插入新行:
```markdown
| `LLM_TIMEOUT` | 502 | LLM 调用超时 | retry / skip / abort | exceptions.LLMTimeoutError |
| `LLM_NETWORK_ERROR` | 502 | 网络失败/5xx 重试耗尽 | retry / skip / abort | exceptions.LLMNetworkError |
```
(保留既有 `LLM_NOT_CONFIGURED` / `LLM_PARSE_ERROR` 行不动;`INTERNAL_ERROR` 行保留为 API 层兜底,来源列无需改。)
- [ ] **Step 2: 追加 `_AI_USAGE_LOG.md`**
范式步骤:"Agent 实现",摘要:api-design §7 补 LLM_NETWORK_ERROR 行(对齐异常树 error_code)。
- [ ] **Step 3: 全量回归**
Run: `python -m pytest -v`
Expected: PASS(≥ 124 passed / 100.00% 覆盖,fail_under=99 达标)
- [ ] **Step 4: 提交**
```bash
git add docs/api-design.md _AI_USAGE_LOG.md
git commit -m "docs: api §7 补 LLM_NETWORK_ERROR 行(对齐异常树错误码)"
```
---
## Self-Review
**1. Spec 覆盖**
- §3.1 异常树 error_code → Task 1 ✓
- §3.2 结果对象字段 → Task 2 ✓
- §3.3 engine 同源透传 → Task 3(含「主模型超时+备用网络失败→取最后一次」测试)✓
- §3.4 语义表 INTERNAL_ERROR 兜底 → 文档保留(Task 4 说明)✓
- §3.5 api §7 补行 → Task 4 ✓
- §4 测试策略 → Task 1errors/ Task 2types/ Task 3engine 五用例)✓
- §5 验收 1-6 → 各任务 Step 4 + Task 4 Step 3 全量回归 ✓
**2. 占位符检查**:无 TBD/TODO;每步含完整代码与预期输出。✓
**3. 类型一致性**
- `error_code` 字段名在 types.py / engine.py / 测试一致 ✓
- `FakeLLMClient` 脚本状态(raise_timeout/raise_network/parse_fail)与既有 helper 一致 ✓
- `Models()` 类(`primary`/`fallback`)在 engine 测试已定义,Task 3 直接复用 ✓
- 错误码字符串常量(`LLM_TIMEOUT` 等)与 api §7 表格一致 ✓
**4. 依赖顺序**Task 1(异常常量)→ Task 2(types 字段,不依赖常量)→ Task 3(engine 透传,依赖 1+2)→ Task 4(文档)。✓
## 提交消息汇总
- Task 1: `feat: 异常树 error_code 类属性(对齐 api-design §7 错误码)`
- Task 2: `feat: 结果对象 error_code 字段(默认 None 向后兼容)`
- Task 3: `feat: engine 透传 error_code(与 error 同源取最后一次异常)`
- Task 4: `docs: api §7 补 LLM_NETWORK_ERROR 行(对齐异常树错误码)`