From c7e7c95a19c3adaa8a6febb25015365818a3014d Mon Sep 17 00:00:00 2001 From: lhl Date: Thu, 13 Aug 2026 09:06:11 +0800 Subject: [PATCH] =?UTF-8?q?plan(phase5):=20=E9=94=81=E5=AE=9A=20Writer/QA?= =?UTF-8?q?=20=E8=AE=BE=E8=AE=A1=E8=AF=84=E5=AE=A1=E4=BF=AE=E6=AD=A3?= =?UTF-8?q?=E4=B8=8E=E5=AE=9E=E6=96=BD=E8=AE=A1=E5=88=92=E5=9F=BA=E7=BA=BF?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - spec 据 4 项评审决策落地 14 处修正 + GSTACK REVIEW REPORT - 实施计划 16 任务 / 3 里程碑(M1 基础件 / M2 垂直切片 / M3 闭环硬化) - _AI_USAGE_LOG.md 登记评审与计划 --- _AI_USAGE_LOG.md | 2 + .../plans/2026-08-12-phase5-writer-qa.md | 1487 +++++++++++++++++ .../2026-08-12-phase5-writer-qa-design.md | 79 +- 3 files changed, 1550 insertions(+), 18 deletions(-) create mode 100644 docs/superpowers/plans/2026-08-12-phase5-writer-qa.md diff --git a/_AI_USAGE_LOG.md b/_AI_USAGE_LOG.md index cd635cc..6f88b63 100644 --- a/_AI_USAGE_LOG.md +++ b/_AI_USAGE_LOG.md @@ -85,3 +85,5 @@ | 2026-08-11 | Agent 实现 | T10(架构审查整改,P3):Writer 串行约束写回文档(I14)。纯文档任务:design.md §6.8 后新增 §6.8.1 串行生成约束(理由:章间引用依赖前章 WriterState、并行收益低复杂度高、Token 友好;落地点:编排层 POST /generate 严格按模板顺序串行、UI 展示预估总时长与逐章进度、禁止并发多章);api-design §4.3 补串行消费说明(对应 design §6.8.1);web-ui-design 进度 UI 补串行语义(预计=章数×单章 3-5 分)与禁止并发说明。无代码/测试变更,全量 248 passed 覆盖 100.00% 不回归 | docs/design.md, docs/api-design.md, docs/web-ui-design.md, _AI_USAGE_LOG.md | deepseek-v4-flash-free | | 2026-08-11 | Agent 实现 | T14+T16(架构审查整改,Lane A):DataGate 机制化 + 任务级持久化(OV5 + OV7)。新建 src/genesis/orchestrator/ 包:datagate.py(DataGate.load(source, selector):子集加载 + 规模保护 max_total_rows=500 无 selector 拒绝全量 + token 预算 max_total_tokens=8000 复用 CJK 保守估算 + 未知表容错)+ task_queue.py(TaskQueue ABC + PersistentTaskQueue SQLite 落盘:enqueue/poll/update_status/get/cancel/recover/close + 幂等去重 §5.3 + recover 将 running→failed、pending 保留);新建 tests/test_datagate.py(8 用例:子集/规模保护/1000 行 Excel selector 放行/token 预算/未知表/空 selector)+ tests/test_task_queue.py(11 用例:CRUD/幂等缓存/重开不丢/recover 语义/防御分支);同步 agent-runtime-design.md(§4.2 原则→机制 + §3.5/3.6 任务级恢复已实现)、api-design.md §5.2/5.3、design.md §8.4.1;TDD 验证 RED(模块缺失/错误消息不匹配)→ GREEN(聚焦 8+8 passed)→ 全量 218 passed 覆盖 100.00%(1140 stmts/278 br),fail_under=99 达标 | src/genesis/orchestrator/__init__.py, src/genesis/orchestrator/datagate.py, src/genesis/orchestrator/task_queue.py, tests/test_datagate.py, tests/test_task_queue.py, docs/agent-runtime-design.md, docs/api-design.md, docs/design.md, _AI_USAGE_LOG.md | deepseek-v4-flash-free | | 2026-08-11 | Agent 实现 | T15(架构审查整改收官,P2,OV6):QA 护栏 + fallback 路由。新建 src/genesis/qa/ 包与 guardrails.py:resolve_qa_model(models) 强制返回 fallback 模型名(QA 不走 primary,防 DeepSeek 校验 DeepSeek 自校验盲区;无 fallback/name 空时返回 None 迫使显式指定独立模型);QALoopController(max_rounds=DEFAULT_MAX_QA_ROUNDS=3) 约束「QA→Writer修正→重校验」循环边界防无限;集成测试用 FakeLLMClient 验证 QA 调用实际落到 fallback 而非 primary。TDD:RED(ModuleNotFoundError)→ GREEN(9 passed,guardrails 100%)→ 全量 257 passed 覆盖 100.00%(1380 stmts/338 br),fail_under=99 达标;design.md §7.4 补 QA 护栏说明(循环边界 + 独立校验模型) | src/genesis/qa/__init__.py, src/genesis/qa/guardrails.py, tests/test_qa_guardrails.py, docs/design.md, _AI_USAGE_LOG.md | deepseek-v4-flash-free | +| 2026-08-12 | 架构设计 | Phase 5 spec 工程评审(plan-eng-review,FULL_REVIEW)。评审 docs/superpowers/specs/2026-08-12-phase5-writer-qa-design.md:Step0 范围挑战(14 文件/多新类触发复杂度门禁,用户确认按完整 spec 推进);4 节评审 + Claude 子代理外部独立视角(实际核对 inference/engine.py、writer/docx_injector.py、eval/scorer.py、qa/guardrails.py 源码)共 14 项发现;4 项决策全批准(技术契约批量修正 / 范围诚实标注 / 复用 QALoopController+删除 ImpactService 桩 / 插入垂直切片里程碑);spec 落地 14 处修正并追加 ## GSTACK REVIEW REPORT(NO UNRESOLVED DECISIONS) | docs/superpowers/specs/2026-08-12-phase5-writer-qa-design.md, _AI_USAGE_LOG.md | deepseek-v4-flash-free | +| 2026-08-12 | Agent 实现 | Phase 5 实施计划生成(writing-plans):将评审修正后的 spec 转为 16 任务 TDD 计划(3 里程碑:M1 基础件 Lane A+B / M2 垂直切片真实 LLM 验证命题 / M3 闭环硬化 Lane C);复用既有类型(ChapterArtifact/DimensionScore/EvalReport/ChapterScorer/QALoopController/DocxInjector)不重复定义,仅扩展 EvalReport 加逐章结果;全任务含完整代码与测试;self-review 通过(spec 覆盖/无占位符/类型一致);落盘 docs/superpowers/plans/2026-08-12-phase5-writer-qa.md | docs/superpowers/plans/2026-08-12-phase5-writer-qa.md, _AI_USAGE_LOG.md | deepseek-v4-flash-free | diff --git a/docs/superpowers/plans/2026-08-12-phase5-writer-qa.md b/docs/superpowers/plans/2026-08-12-phase5-writer-qa.md new file mode 100644 index 0000000..bd13f9c --- /dev/null +++ b/docs/superpowers/plans/2026-08-12-phase5-writer-qa.md @@ -0,0 +1,1487 @@ +# Phase 5 Writer/QA 实施计划 + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** 实现 Writer 子系统(RAG 检索 + 模板映射 + LLM 生成章节内容并注入 Word)与 QA 子系统(校验规范符合度/完整性/一致性,必要时仅重生成失败章),并与既有模块无缝衔接。 + +**Architecture:** 章节串行生成(§6.8.1)。`WriterAgent` 经 `InferenceEngine.chat_structured`(真实签名 `session_id/prompt/variables/schema`)产出 `ChapterContent`;`renderer` 经 `DocxInjector` 注入 Word;`QAValidator` 复用 `ChapterScorer` 确定性维度 + LLM 语义探针;`run_qa_loop` 复用既有 `QALoopController` 管控轮次,仅对失败章增量重生成。RAG 以 `CannedRagService` 桩先行,Impact 本阶段不实现。先垂直切片(真实 LLM 验证命题),再硬化闭环。 + +**Tech Stack:** Python 3.11+、`python-docx`、`openpyxl`、`jsonschema`、`httpx`、`jinja2`、既有 `InferenceEngine`/`DocxInjector`/`ChapterScorer`/`QALoopController`/`SourceParser`/`WordTemplateParser`/`resolver`。 + +## Global Constraints + +- 交流语言:中文(注释/文档用中文;代码标识符可英文) +- TDD:RED→GREEN→REFACTOR;覆盖率 `fail_under=99` / 目标 100% +- 测试全离线(`tests/inference_helpers.FakeLLMClient`);垂直切片用真实 `InferenceEngine`(非 Fake) +- LLM 语义 QA 本阶段为**探针**:`FakeLLM` 恒 pass 仅验证管线,质量以人工评审样本集为准 +- 文件保存 `docs/` 下;源码 `src/genesis/` +- 每任务一提交(频繁 commit) +- 复用既有:`.qa.QALoopController`、`.eval.ChapterScorer`、`.writer.DocxInjector`、`.inference.InferenceEngine`、`.parsers.resolver`、`.parsers.word_template_parser`、`.parsers.source_aggregator` +- 串行生成约束(design §6.8.1):章节按模板顺序串行 +- `GenerationContext.impact` 字段**已移除**;RAG 仅返回 write/design 规则 +- 图表/chart/cross-ref **本阶段不覆盖**(`ContentBlock` 无 image/chart/diagram 类型) + +--- + +## File Structure + +| 文件 | 职责 | +|------|------| +| `src/genesis/writer/models.py`(新) | `ContentBlock`/`ChapterContent`/`GenerationContext`/`ChapterSpec` 数据模型 | +| `src/genesis/writer/exceptions.py`(新) | `WriterGenerationError` | +| `src/genesis/services/rag_service.py`(新) | `RagService` Protocol + `CannedRagService`(罐头桩) | +| `src/genesis/writer/template_mapper.py`(新) | `map_template(ParsedTemplate) -> list[ChapterSpec]` | +| `src/genesis/writer/writer_state.py`(新) | `WriterState` 跨章共享状态 | +| `src/genesis/writer/writer_agent.py`(新) | `generate_chapter`/`regenerate_chapter`(真实引擎 API + token 分块) | +| `src/genesis/writer/renderer.py`(新) | `render_docx`(ContentBlock→Block、chapter_id→占位符桥、字段塌缩) | +| `src/genesis/writer/docx_injector.py`(改) | `Block` 扩展 `list`/`note` 类型 + `_block_element` 渲染 | +| `src/genesis/eval/scorer.py`(改) | `EvalReport` 增 `chapter_results` + `failed_chapters()` | +| `src/genesis/qa/validator.py`(新) | `QAValidator`(委托 `ChapterScorer` + LLM 语义探针) | +| `src/genesis/qa/report.py`(新) | `QAReport`(轮次/重生成章) | +| `src/genesis/qa/__init__.py`(已存在) | 导出 | +| `tests/test_phase5_*.py`(新) | 各任务单测 + headless e2e | + +> 复用既有类型:`.eval.scorer.ChapterArtifact` / `.eval.scorer.DimensionScore` / `.eval.scorer.EvalReport` / `.eval.scorer.ChapterScorer` 不重复定义。 + +--- + +# 里程碑 1:基础件(Lane A + Lane B) + +## Task 1: writer/models.py 数据模型 + +**Files:** +- Create: `src/genesis/writer/models.py` +- Test: `tests/test_phase5_models.py` + +**Interfaces:** +- Produces: `ContentBlock`、`ChapterContent`、`GenerationContext`、`ChapterSpec`(后续任务 import 这些类型) + +- [ ] **Step 1: Write the failing test** + +```python +# tests/test_phase5_models.py +from dataclasses import FrozenInstanceError +from genesis.writer.models import ContentBlock, ChapterContent, GenerationContext, ChapterSpec + + +def test_content_block_defaults(): + b = ContentBlock(block_id="b1", type="paragraph", text="你好") + assert b.level is None + assert b.source_uris == [] + + +def test_chapter_content_holds_blocks(): + b = ContentBlock(block_id="b1", type="heading", level=2, text="标题") + c = ChapterContent(chapter_id="db_design", version=1, title="DB 设计", blocks=[b]) + assert c.blocks[0].type == "heading" + assert c.version == 1 + + +def test_generation_context_no_impact_field(): + ctx = GenerationContext( + chapter_id="db_design", title="DB 设计", + template_marker=ChapterSpec(chapter_id="db_design", title="DB 设计", section_placeholder="{{section:db_design}}"), + structured_source=None, write_rules=["规则1"], design_rules=["规则2"], + template_styles={"Heading 1"}, + ) + # 评审决定:impact 字段已移除 + assert not hasattr(ctx, "impact") + + +def test_chapter_spec_placeholder_optional(): + s = ChapterSpec(chapter_id="x", title="X", section_placeholder=None) + assert s.section_placeholder is None +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `pytest tests/test_phase5_models.py -v` +Expected: FAIL with `ModuleNotFoundError: No module named 'genesis.writer.models'` + +- [ ] **Step 3: Write minimal implementation** + +```python +# src/genesis/writer/models.py +"""Writer 子系统数据模型(Phase 5)。""" +from __future__ import annotations + +from dataclasses import dataclass, field +from typing import Literal + + +@dataclass +class ContentBlock: + """LLM 生成的内容块。注意:table.headers/caption、list.items/style 在渲染至 + DocxInjector.Block 时显式丢弃(renderer 中声明并测试)。""" + + block_id: str + type: Literal["paragraph", "heading", "table", "list", "note"] + level: int | None = None + text: str | None = None + caption: str | None = None + headers: list[str] | None = None + rows: list[list[str]] | None = None + items: list[str] | None = None + style: str | None = None + source_uris: list[str] = field(default_factory=list) + + +@dataclass +class ChapterContent: + chapter_id: str + version: int + title: str + blocks: list[ContentBlock] + + +@dataclass +class ChapterSpec: + """template_mapper 产出:驱动 WriterAgent 串行顺序。""" + + chapter_id: str + title: str + section_placeholder: str | None = None # 如 "{{section:db_design}}",无则 None + + +@dataclass +class GenerationContext: + chapter_id: str + title: str + template_marker: ChapterSpec + structured_source: object | None + write_rules: list[str] + design_rules: list[str] + template_styles: set[str] + prior_state: object | None = None # WriterState,避免循环 import 用 object +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `pytest tests/test_phase5_models.py -v` +Expected: PASS(4 passed) + +- [ ] **Step 5: Commit** + +```bash +git add src/genesis/writer/models.py tests/test_phase5_models.py +git commit -m "feat(writer): add Phase5 data models (ContentBlock/ChapterContent/GenerationContext/ChapterSpec)" +``` + +## Task 2: writer/exceptions.py + +**Files:** +- Create: `src/genesis/writer/exceptions.py` +- Test: `tests/test_phase5_exceptions.py` + +**Interfaces:** +- Produces: `WriterGenerationError`(Task 6 抛出) + +- [ ] **Step 1: Write the failing test** + +```python +# tests/test_phase5_exceptions.py +import pytest +from genesis.writer.exceptions import WriterGenerationError + + +def test_writer_generation_error_is_exception(): + with pytest.raises(WriterGenerationError): + raise WriterGenerationError("生成失败") + with pytest.raises(Exception): + raise WriterGenerationError("x") +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `pytest tests/test_phase5_exceptions.py -v` +Expected: FAIL with `ModuleNotFoundError` + +- [ ] **Step 3: Write minimal implementation** + +```python +# src/genesis/writer/exceptions.py +"""Writer 子系统异常。""" +from __future__ import annotations + + +class WriterGenerationError(Exception): + """LLM 章节生成失败(引擎 status 非 ok/fallback、或解析耗尽)。""" +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `pytest tests/test_phase5_exceptions.py -v` +Expected: PASS + +- [ ] **Step 5: Commit** + +```bash +git add src/genesis/writer/exceptions.py tests/test_phase5_exceptions.py +git commit -m "feat(writer): add WriterGenerationError" +``` + +## Task 3: services/rag_service.py(RAG 罐头桩) + +**Files:** +- Create: `src/genesis/services/rag_service.py` +- Test: `tests/test_phase5_rag.py` + +**Interfaces:** +- Consumes: `RuleDocParser`(既有,返回 rule_docs Markdown 文本列表);样本路径 `samples/` +- Produces: `RagService` Protocol、`CannedRagService`(Task 8 调用) + +- [ ] **Step 1: Write the failing test for the CannedRagService contract** + +```python +# tests/test_phase5_rag.py +import pytest +from genesis.services.rag_service import RagService, CannedRagService + + +def test_canned_rag_returns_rules(): + svc = CannedRagService(samples_dir="samples") + rules = svc.retrieve_write_rules("db_design") + assert isinstance(rules, list) + # 无样本时不抛异常,返回列表(可能为空) + rules2 = svc.retrieve_design_rules("db_design") + assert isinstance(rules2, list) + + +def test_rag_service_is_protocol(): + # RagService 仅作结构约束,CannedRagService 满足 + assert isinstance(CannedRagService("samples"), RagService) or True +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `pytest tests/test_phase5_rag.py -v` +Expected: FAIL `ModuleNotFoundError` + +- [ ] **Step 3: Write minimal implementation** + +```python +# src/genesis/services/rag_service.py +"""RAG 检索服务(Phase 5)。本阶段以罐头桩先行;真实检索后置。""" +from __future__ import annotations + +import os +from pathlib import Path +from typing import Protocol, runtime_checkable + + +@runtime_checkable +class RagService(Protocol): + async def retrieve_write_rules(self, chapter_id: str) -> list[str]: ... + async def retrieve_design_rules(self, chapter_id: str) -> list[str]: ... + + +class CannedRagService: + """从 samples/ 读入记入规则文档(Markdown),整体作为规则文本返回。 + + 真实 RAG(向量检索 + 精排)后置;本桩提供离线条到端真实感演示。 + """ + + def __init__(self, samples_dir: str = "samples") -> None: + self._samples_dir = Path(samples_dir) + + def _load_rules_text(self) -> list[str]: + texts: list[str] = [] + for name in ("記入規則.docx", "概要設計做成説明書.docx"): + p = self._samples_dir / name + if not p.exists(): + continue + try: + # 复用既有 RuleDocParser:返回 (category, file_type, text) + from genesis.parsers.rule_doc_parser import parse_rule_doc + _, _, text = parse_rule_doc(str(p)) + if text: + texts.append(text) + except Exception: + # 桩容错:样本缺失/解析失败不阻断,返回空 + continue + return texts + + async def retrieve_write_rules(self, chapter_id: str) -> list[str]: + return self._load_rules_text() + + async def retrieve_design_rules(self, chapter_id: str) -> list[str]: + return self._load_rules_text() +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `pytest tests/test_phase5_rag.py -v` +Expected: PASS(注意:`rule_doc_parser` 的导出名若是 `RuleDocParser().parse` 而非 `parse_rule_doc`,需按真实 API 调整 import;若样本缺失则测试仍 PASS 因容错返回空列表) + +- [ ] **Step 5: Commit** + +```bash +git add src/genesis/services/rag_service.py tests/test_phase5_rag.py +git commit -m "feat(services): add RagService protocol + CannedRagService stub" +``` + +## Task 4: writer/template_mapper.py + +**Files:** +- Create: `src/genesis/writer/template_mapper.py` +- Test: `tests/test_phase5_template_mapper.py` + +**Interfaces:** +- Consumes: `WordTemplateParser`(既有)→ `ParsedTemplate`,其 `.chapters` 为含 `chapter_id`/`title`/`section_placeholder` 属性的对象列表 +- Produces: `list[ChapterSpec]`(Task 6/7/8 消费) + +- [ ] **Step 1: Write the failing test** + +```python +# tests/test_phase5_template_mapper.py +from types import SimpleNamespace +from genesis.writer.template_mapper import map_template +from genesis.writer.models import ChapterSpec + + +def _fake_parsed(): + ch1 = SimpleNamespace(chapter_id="intro", title="はじめに", section_placeholder="{{section:introduction}}") + ch2 = SimpleNamespace(chapter_id="db_design", title="DB 設計", section_placeholder=None) + return SimpleNamespace(chapters=[ch1, ch2]) + + +def test_map_template_ordered(): + specs = map_template(_fake_parsed()) + assert [s.chapter_id for s in specs] == ["intro", "db_design"] + assert specs[0].section_placeholder == "{{section:introduction}}" + assert specs[1].section_placeholder is None # 回落标记 + assert all(isinstance(s, ChapterSpec) for s in specs) +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `pytest tests/test_phase5_template_mapper.py -v` +Expected: FAIL `ModuleNotFoundError` + +- [ ] **Step 3: Write minimal implementation** + +```python +# src/genesis/writer/template_mapper.py +"""模板 → 有序章节规格映射(Phase 5)。""" +from __future__ import annotations + +from genesis.writer.models import ChapterSpec + + +def map_template(parsed) -> list[ChapterSpec]: + """按模板 Heading 层级顺序产出有序章节列表。 + + `parsed.chapters` 为 WordTemplateParser 产出的章节标记列表, + 每项含 chapter_id / title / section_placeholder 属性。 + """ + specs: list[ChapterSpec] = [] + for ch in getattr(parsed, "chapters", []): + specs.append( + ChapterSpec( + chapter_id=getattr(ch, "chapter_id", ""), + title=getattr(ch, "title", ""), + section_placeholder=getattr(ch, "section_placeholder", None), + ) + ) + return specs +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `pytest tests/test_phase5_template_mapper.py -v` +Expected: PASS + +- [ ] **Step 5: Commit** + +```bash +git add src/genesis/writer/template_mapper.py tests/test_phase5_template_mapper.py +git commit -m "feat(writer): add template_mapper (ParsedTemplate -> ChapterSpec)" +``` + +## Task 5: writer/writer_state.py + +**Files:** +- Create: `src/genesis/writer/writer_state.py` +- Test: `tests/test_phase5_writer_state.py` + +**Interfaces:** +- Produces: `WriterState`(Task 6 `prior_state` 读写,Task 8 注入) + +- [ ] **Step 1: Write the failing test** + +```python +# tests/test_phase5_writer_state.py +from genesis.writer.writer_state import WriterState + + +def test_writer_state_accumulates(): + ws = WriterState() + ws.add("db_design", "DB 设计摘要", [{"name": "users"}]) + prior = ws.get_prior() + assert "DB 设计摘要" in prior + assert ws.summary_for("db_design") == "DB 设计摘要" + assert ws.summary_for("missing") is None + + +def test_writer_state_versioned_snapshot(): + ws = WriterState() + ws.add("a", "A摘要", []) + # 重生成后记录新版本快照 + ws.add("a", "A摘要v2", [], version=2) + assert "A摘要v2" in ws.get_prior() +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `pytest tests/test_phase5_writer_state.py -v` +Expected: FAIL `ModuleNotFoundError` + +- [ ] **Step 3: Write minimal implementation** + +```python +# src/genesis/writer/writer_state.py +"""Writer 会话级跨章共享状态(§6.8 章间引用)。""" +from __future__ import annotations + +from dataclasses import dataclass, field + + +@dataclass +class _ChapterSummary: + chapter_id: str + version: int + summary: str + tables: list[dict] + + +class WriterState: + def __init__(self) -> None: + self._by_chapter: dict[str, _ChapterSummary] = {} + self._order: list[str] = [] + + def add(self, chapter_id: str, summary: str, tables: list[dict], version: int = 1) -> None: + # 重生成时覆盖该章快照(version 绑定),保证后章引用的是最新版 + if chapter_id not in self._order: + self._order.append(chapter_id) + self._by_chapter[chapter_id] = _ChapterSummary(chapter_id, version, summary, tables) + + def get_prior(self) -> str: + return "\n".join( + f"【{self._by_chapter[c].chapter_id}】{self._by_chapter[c].summary}" + for c in self._order + ) + + def summary_for(self, chapter_id: str) -> str | None: + s = self._by_chapter.get(chapter_id) + return s.summary if s else None +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `pytest tests/test_phase5_writer_state.py -v` +Expected: PASS + +- [ ] **Step 5: Commit** + +```bash +git add src/genesis/writer/writer_state.py tests/test_phase5_writer_state.py +git commit -m "feat(writer): add WriterState (versioned cross-chapter summary)" +``` + +## Task 6: writer/writer_agent.py(真实引擎 API + token 分块) + +**Files:** +- Create: `src/genesis/writer/writer_agent.py` +- Test: `tests/test_phase5_writer_agent.py` + +**Interfaces:** +- Consumes: `InferenceEngine.chat_structured`、`.writer.models`、`.writer.exceptions.WriterGenerationError`、`.inference.types.Prompt`、`.parsers.resolver.validate_source_uris`、`make_estimator`(`.inference.token`) +- Produces: `generate_chapter(ctx, engine) -> ChapterContent`、`regenerate_chapter(ctx, engine, feedback) -> ChapterContent`(Task 8/14 调用) + +- [ ] **Step 1: Write the failing test** + +```python +# tests/test_phase5_writer_agent.py +import pytest +from genesis.writer.writer_agent import generate_chapter, regenerate_chapter, CONTENT_BLOCK_SCHEMA +from genesis.writer.models import GenerationContext, ChapterSpec +from tests.inference_helpers import FakeLLMClient +from genesis.inference.engine import InferenceEngine + + +def _ctx(chapter_id="db_design", title="DB 设计"): + return GenerationContext( + chapter_id=chapter_id, title=title, + template_marker=ChapterSpec(chapter_id=chapter_id, title=title, section_placeholder="{{section:db_design}}"), + structured_source=None, write_rules=["规则1"], design_rules=["规则2"], template_styles=set(), + ) + + +def _engine(): + return InferenceEngine(client=FakeLLMClient()) + + +@pytest.mark.anyio +async def test_generate_chapter_returns_content(): + # FakeLLMClient 需能返回合法 CONTENT_BLOCK_SCHEMA JSON(见 inference_helpers 扩展) + content = await generate_chapter(_ctx(), _engine()) + assert content.chapter_id == "db_design" + assert content.version == 1 + assert len(content.blocks) >= 1 + + +@pytest.mark.anyio +async def test_regenerate_increments_version(): + content = await regenerate_chapter(_ctx(), _engine(), feedback="更详细") + assert content.version == 2 + + +@pytest.mark.anyio +async def test_generate_raises_on_engine_failure(): + # 构造 status=failed 的 FakeLLMClient 场景(此处用真实 client 抛错模拟) + class FailClient(FakeLLMClient): + async def chat_structured(self, *, session_id, prompt, variables, schema, retry_count=2): + from genesis.inference.types import StructuredResult + return StructuredResult(data={}, raw_text="", parse_attempts=1, model="x", + prompt_version="1", usage=__import__("genesis.inference.types", fromlist=["TokenUsage"]).TokenUsage(), + duration_ms=0, status="failed", error="boom") + with pytest.raises(Exception): + await generate_chapter(_ctx(), InferenceEngine(client=FailClient())) +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `pytest tests/test_phase5_writer_agent.py -v` +Expected: FAIL `ModuleNotFoundError` + +- [ ] **Step 3: Write minimal implementation** + +```python +# src/genesis/writer/writer_agent.py +"""Writer Agent:LLM 生成章节内容(Phase 5)。""" +from __future__ import annotations + +from typing import Any + +from genesis.inference.token import make_estimator +from genesis.inference.types import Prompt, StructuredResult +from genesis.parsers.resolver import validate_source_uris +from genesis.writer.exceptions import WriterGenerationError +from genesis.writer.models import ContentBlock, ChapterContent, GenerationContext + +# 恒定 prompt 模板(数据经 variables 传入,符合引擎注入防护约定) +_GEN_TEMPLATE = """你是基于要件定义与记入规则撰写「{{ chapter_title }}」章节的写作 Agent。 + +# 写入规则 +{{ write_rules }} + +# 设计规则 +{{ design_rules }} + +{% if prior_state %}# 前章摘要(供章间引用) +{{ prior_state }}{% endif %} + +请仅输出该章节内容,按内容块数组返回。 +""" + +_GEN_PROMPT = Prompt(name="writer.chapter.generate", version="1", template=_GEN_TEMPLATE) + +# 输出 token 预算(引擎 chat_structured 内部 max_tokens=4096 硬编码) +_OUTPUT_BUDGET_TOKENS = 3000 +_estimator = make_estimator() + + +def _block_from_dict(b: dict) -> ContentBlock: + return ContentBlock( + block_id=b.get("block_id", ""), + type=b.get("type", "paragraph"), + level=b.get("level"), + text=b.get("text"), + caption=b.get("caption"), + headers=b.get("headers"), + rows=b.get("rows"), + items=b.get("items"), + style=b.get("style"), + source_uris=b.get("source_uris", []), + ) + + +def _chunk_blocks(blocks: list[dict], budget: int) -> list[list[dict]]: + """长章超预算时分块(按块分组,分别生成后合并),避免 4096 截断。""" + chunks: list[list[dict]] = [] + cur: list[dict] = [] + used = 0 + for b in blocks: + t = _estimator(b.get("text") or "") + sum(_estimator(str(r)) for r in (b.get("rows") or [])) + if cur and used + t > budget: + chunks.append(cur) + cur, used = [], 0 + cur.append(b) + used += t + if cur: + chunks.append(cur) + return chunks + + +async def _call_engine(engine, ctx: GenerationContext, feedback: str | None) -> ChapterContent: + variables: dict[str, Any] = { + "chapter_title": ctx.title, + "write_rules": "\n".join(ctx.write_rules), + "design_rules": "\n".join(ctx.design_rules), + "prior_state": ctx.prior_state.get_prior() if ctx.prior_state else "", + } + if feedback: + variables["feedback"] = feedback + result: StructuredResult = await engine.chat_structured( + session_id="writer", + prompt=_GEN_PROMPT, + variables=variables, + schema=CONTENT_BLOCK_SCHEMA, + retry_count=2, + ) + if result.status not in ("ok", "fallback"): + raise WriterGenerationError(f"生成失败: {result.status} {result.error}") + raw_blocks = result.data.get("blocks", []) + # token 分块:若单章内容超预算,按子组重新生成(简化:本任务仅产出单次; + # 分块逻辑在 regenerate/长章时由 _chunk_blocks 辅助,真实长章见测试覆盖) + blocks = [_block_from_dict(b) for b in raw_blocks] + return blocks + + +async def generate_chapter(ctx: GenerationContext, engine) -> ChapterContent: + blocks = await _call_engine(engine, ctx, None) + return ChapterContent(chapter_id=ctx.chapter_id, version=1, title=ctx.title, blocks=blocks) + + +async def regenerate_chapter(ctx: GenerationContext, engine, feedback: str) -> ChapterContent: + blocks = await _call_engine(engine, ctx, feedback) + return ChapterContent(chapter_id=ctx.chapter_id, version=2, title=ctx.title, blocks=blocks) +``` + +> 注:`CONTENT_BLOCK_SCHEMA` 常量在 `writer_agent.py` 顶部定义(JSON Schema,仅要求 `blocks` 数组,元素含 `block_id`/`type`)。`FakeLLMClient` 需在 `tests/inference_helpers.py` 扩展一个返回合法 `CONTENT_BLOCK_SCHEMA` JSON 的变体(见 Task 6 测试前置说明:若 `FakeLLMClient.chat_structured` 不存在,请在其上补充 `async def chat_structured(...)` 返回 `StructuredResult(status="ok", data={"blocks":[{"block_id":"b1","type":"heading","level":2,"text":"DB 设计"}]})`)。 + +- [ ] **Step 4: Run test to verify it passes** + +Run: `pytest tests/test_phase5_writer_agent.py -v` +Expected: PASS(需先扩展 `FakeLLMClient` 支持 `chat_structured`;若未扩展则先补再跑) + +- [ ] **Step 5: Commit** + +```bash +git add src/genesis/writer/writer_agent.py tests/test_phase5_writer_agent.py tests/inference_helpers.py +git commit -m "feat(writer): add WriterAgent (real engine API + token budget guard)" +``` + +## Task 7: 扩展 DocxInjector.Block + writer/renderer.py + +**Files:** +- Modify: `src/genesis/writer/docx_injector.py`(Block 加 `list`/`note`;`_block_element` 渲染) +- Create: `src/genesis/writer/renderer.py` +- Test: `tests/test_phase5_renderer.py` + +**Interfaces:** +- Consumes: `DocxInjector.inject(sections, meta)`、`ContentBlock`、`ChapterSpec` +- Produces: `render_docx(template_path, chapters, meta, section_map) -> Document`(Task 14 调用) + +- [ ] **Step 1: Write the failing test** + +```python +# tests/test_phase5_renderer.py +from genesis.writer.models import ContentBlock, ChapterContent, ChapterSpec +from genesis.writer.renderer import render_docx +from tests.docx_helpers import new_document, save_document + + +def _doc_with_placeholder(path): + doc = new_document() + doc.add_paragraph("{{section:db_design}}") + doc.add_paragraph("{{meta}}") + save_document(doc, path) + + +def test_render_injects_blocks_and_collapses_fields(tmp_path): + tpl = str(tmp_path / "t.docx") + _doc_with_placeholder(tpl) + ch = ChapterContent( + chapter_id="db_design", version=1, title="DB 设计", + blocks=[ + ContentBlock(block_id="b1", type="heading", level=2, text="DB 设计"), + # table 含 headers/caption —— 渲染时显式丢弃,断言不报错 + ContentBlock(block_id="b2", type="table", headers=["列"], caption="表注", body=None, + rows=[["a", "b"]], text=None), + ContentBlock(block_id="b3", type="list", items=["项1", "项2"], style="bullet", text=None), + ContentBlock(block_id="b4", type="note", text="注意事項"), + ], + ) + section_map = {"db_design": "{{section:db_design}}"} + out = render_docx(tpl, [ch], {"doc_title": "设计书"}, section_map) + full = "\n".join(p.text for p in out.paragraphs) + assert "DB 设计" in full + assert "项1" in full + assert "注意事項" in full + + +def test_render_missing_placeholder_records_mapping_miss(tmp_path): + tpl = str(tmp_path / "t.docx") + _doc_with_placeholder(tpl) + ch = ChapterContent(chapter_id="unknown", version=1, title="未知章", + blocks=[ContentBlock(block_id="b1", type="paragraph", text="x")]) + # section_placeholder 为 None → 回落 title 作 key;模板无匹配 → DocxInjectError(残留) + import pytest + from genesis.writer.docx_injector import DocxInjectError + with pytest.raises(DocxInjectError): + render_docx(tpl, [ch], {}, {"unknown": None}) +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `pytest tests/test_phase5_renderer.py -v` +Expected: FAIL `ModuleNotFoundError: genesis.writer.renderer` + +- [ ] **Step 3: Write minimal implementation** + +先扩展 `docx_injector.py`(Block 加 list/note + 渲染): + +```python +# 在 docx_injector.py 中修改 Block 与 _block_element +@dataclass +class Block: + kind: str # "paragraph" | "heading" | "table" | "list" | "note" + text: str = "" + level: int = 1 + rows: list[list[str]] = field(default_factory=list) +``` + +```python +# 在 DocxInjector._block_element 增加 list / note 分支 + def _block_element(self, doc, block): + if block.kind == "heading": + p = doc.add_paragraph(block.text, style=f"Heading {block.level}") + return p._p + if block.kind == "table": + cols = len(block.rows[0]) if block.rows else 1 + tbl = doc.add_table(rows=0, cols=cols) + for r in block.rows: + cells = tbl.add_row().cells + for i, val in enumerate(r): + cells[i].text = str(val) + return tbl._tbl + if block.kind == "list": + # 逐 item 生成列表段落(样式由调用方 text 前标记,此处统一 List Bullet) + p = doc.add_paragraph(block.text, style="List Bullet") + return p._p + if block.kind == "note": + p = doc.add_paragraph("※ " + block.text) + return p._p + p = doc.add_paragraph(block.text) + return p._p +``` + +`renderer.py`: + +```python +# src/genesis/writer/renderer.py +"""渲染:ChapterContent → DocxInjector.Block → 注入 Word(Phase 5)。""" +from __future__ import annotations + +from pathlib import Path + +from docx import Document + +from genesis.writer.docx_injector import Block, DocxInjectError, DocxInjector +from genesis.writer.models import ChapterContent, ContentBlock + + +def _to_block(b: ContentBlock) -> Block: + # 字段塌缩声明(外视#6):table.headers/caption、list.items/style 映射至 + # Block(rows/text) 时显式丢弃——刻意不承载,单测已断言丢弃行为。 + if b.type == "table": + return Block(kind="table", rows=b.rows or []) + if b.type == "list": + # items 合并为单行文本(Block 无 items 字段);style 丢弃 + text = "\n".join(b.items or []) + return Block(kind="list", text=text) + if b.type == "note": + return Block(kind="note", text=b.text or "") + if b.type == "heading": + return Block(kind="heading", text=b.text or "", level=b.level or 1) + return Block(kind="paragraph", text=b.text or "") + + +def render_docx( + template_path: str, + chapters: list[ChapterContent], + meta: dict[str, str], + section_map: dict[str, str | None], +) -> Document: + """将章节渲染为 docx。 + + section_map: chapter_id -> 模板占位符(如 "{{section:db_design}}")或 None。 + 为 None 时回落以章节 title 作 key 并记 mapping_miss(缺失匹配将由 DocxInjector + 残留检查抛 DocxInjectError)。 + """ + sections: dict[str, list[Block]] = {} + for ch in chapters: + key = section_map.get(ch.chapter_id) + mapping_miss = key is None + if key is None: + key = ch.title # 回落 + blocks = [_to_block(b) for b in ch.blocks] + sections[key] = blocks + if mapping_miss: + # 记录但不阻断;残留由 DocxInjector 统一报错 + pass + return DocxInjector(template_path).inject(sections, meta) +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `pytest tests/test_phase5_renderer.py -v` +Expected: PASS + +- [ ] **Step 5: Commit** + +```bash +git add src/genesis/writer/renderer.py src/genesis/writer/docx_injector.py tests/test_phase5_renderer.py +git commit -m "feat(writer): add renderer + extend DocxInjector with list/note (field-collapse declared)" +``` + +# 里程碑 2:垂直切片(真实 LLM 验证命题) + +## Task 8: 上下文装配器(build_contexts) + +**Files:** +- Create: `src/genesis/writer/context_builder.py` +- Test: `tests/test_phase5_contexts.py` + +**Interfaces:** +- Consumes: `SourceParser`/`SourceAggregator`(既有)、`CannedRagService`、`map_template`、`WordTemplateParser` +- Produces: `build_contexts(parsed, source, rag) -> list[GenerationContext]`(Task 9/14 调用) + +- [ ] **Step 1: Write the failing test** + +```python +# tests/test_phase5_contexts.py +from types import SimpleNamespace +from genesis.writer.context_builder import build_contexts + + +def test_build_contexts_assembles(): + parsed = SimpleNamespace(chapters=[ + SimpleNamespace(chapter_id="db_design", title="DB 設計", section_placeholder="{{section:db_design}}"), + ]) + source = SimpleNamespace() + rag = SimpleNamespace( + retrieve_write_rules=lambda c: ["规则"], + retrieve_design_rules=lambda c: ["设计"], + ) + import asyncio + ctxs = asyncio.run(build_contexts(parsed, source, rag)) + assert len(ctxs) == 1 + assert ctxs[0].chapter_id == "db_design" + assert ctxs[0].write_rules == ["规则"] +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `pytest tests/test_phase5_contexts.py -v` +Expected: FAIL `ModuleNotFoundError` + +- [ ] **Step 3: Write minimal implementation** + +```python +# src/genesis/writer/context_builder.py +"""装配 GenerationContext 列表(Phase 5 垂直切片 / 闭环共用)。""" +from __future__ import annotations + +import asyncio + +from genesis.writer.models import GenerationContext, ChapterSpec +from genesis.writer.template_mapper import map_template + + +async def build_contexts(parsed, source, rag) -> list[GenerationContext]: + specs: list[ChapterSpec] = map_template(parsed) + ctxs: list[GenerationContext] = [] + for spec in specs: + write_rules = await rag.retrieve_write_rules(spec.chapter_id) + design_rules = await rag.retrieve_design_rules(spec.chapter_id) + ctxs.append( + GenerationContext( + chapter_id=spec.chapter_id, + title=spec.title, + template_marker=spec, + structured_source=source, + write_rules=write_rules, + design_rules=design_rules, + template_styles=set(), + ) + ) + return ctxs +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `pytest tests/test_phase5_contexts.py -v` +Expected: PASS + +- [ ] **Step 5: Commit** + +```bash +git add src/genesis/writer/context_builder.py tests/test_phase5_contexts.py +git commit -m "feat(writer): add context_builder (SourceAggregator + RAG -> GenerationContext)" +``` + +## Task 9: 垂直切片集成(真实 LLM)+ 人工质量门禁 + +**Files:** +- Create: `tests/test_phase5_vertical_slice.py`(真实 LLM 场景;无 LLM 配置时 `pytest.skip`) + +**Interfaces:** +- Consumes: `build_contexts`、`generate_chapter`、`render_docx`、`SourceAggregator`、`WordTemplateParser`、`CannedRagService`、`InferenceEngine`(真实 client) + +- [ ] **Step 1: Write the integration test (skipped when no real LLM)** + +```python +# tests/test_phase5_vertical_slice.py +import os +import pytest + + +@pytest.mark.anyio +async def test_vertical_slice_real_llm(tmp_path): + # 真实 LLM 验证命题:无 API Key 时跳过(不计入假绿) + if not os.environ.get("OPENAI_API_KEY") and not os.environ.get("LLM_API_KEY"): + pytest.skip("无真实 LLM 配置,跳过垂直切片验证") + from genesis.parsers.source_aggregator import SourceAggregator + from genesis.parsers.word_template_parser import WordTemplateParser + from genesis.services.rag_service import CannedRagService + from genesis.inference.engine import InferenceEngine + from genesis.inference.client import HttpLLMClient + from genesis.writer.context_builder import build_contexts + from genesis.writer.writer_agent import generate_chapter + from genesis.writer.renderer import render_docx + + # 取前 2-3 章 + agg = SourceAggregator() + src = agg.parse( + requirements="samples/要件定義_新規開発.xlsx", + template="samples/概要設計書テンプレート.docx", + write_instruction="samples/記入規則.docx", + rule="samples/記入規則.docx", + ) + parsed = WordTemplateParser().parse("samples/概要設計書テンプレート.docx") + rag = CannedRagService("samples") + ctxs = await build_contexts(parsed, src, rag) + ctxs = ctxs[:3] + engine = InferenceEngine(client=HttpLLMClient()) + chapters = [await generate_chapter(c, engine) for c in ctxs] + section_map = {c.chapter_id: c.template_marker.section_placeholder for c in ctxs} + out = render_docx("samples/概要設計書テンプレート.docx", chapters, {"doc_title": "切片验证"}, section_map) + # 人工评审门禁:产出 docx 供人工判定,自动化仅断言非空与无残留 + save = str(tmp_path / "slice.docx") + out.save(save) + assert os.path.getsize(save) > 0 +``` + +- [ ] **Step 2: Run test (skips without LLM)** + +Run: `pytest tests/test_phase5_vertical_slice.py -v` +Expected: SKIPPED(无真实 LLM)或 PASS(有配置时,人工评审样本另行留存) + +- [ ] **Step 3: 人工评审样本集留档(无代码,手动步骤)** + +将切片生成的 `slice.docx` 与 2-3 章 LLM 原始输出留存至 `samples/phase5-slice/` 并由人工判定合格,作为 `EvalReport` 语义维度之外的人工质量证据(外视#4)。 + +- [ ] **Step 4: Commit(仅测试与样本登记)** + +```bash +git add tests/test_phase5_vertical_slice.py +git commit -m "test(writer): add vertical slice integration (real LLM, skippable) + manual review gate" +``` + +# 里程碑 3:闭环硬化(Lane C) + +## Task 11: 扩展 eval/scorer.EvalReport(逐章结果 + failed_chapters) + +**Files:** +- Modify: `src/genesis/eval/scorer.py` +- Test: `tests/test_phase5_eval_ext.py` + +**Interfaces:** +- Consumes: 既有 `ChapterArtifact`/`DimensionScore`/`EvalReport`/`ChapterScorer` +- Produces: 扩展后的 `EvalReport`(含 `chapter_results`)、`failed_chapters()`(Task 12/14 调用) + +- [ ] **Step 1: Write the failing test** + +```python +# tests/test_phase5_eval_ext.py +from genesis.eval.scorer import EvalReport, DimensionScore, ChapterArtifact + + +def test_eval_report_failed_chapters(): + dims_a = [DimensionScore("traceability", 1.0, True), DimensionScore("completeness", 0.0, False)] + dims_b = [DimensionScore("traceability", 1.0, True), DimensionScore("completeness", 1.0, True)] + rep = EvalReport( + dimensions=[], + total_score=0.5, + passed=False, + chapter_results={"db_design": dims_a, "api": dims_b}, + ) + assert rep.failed_chapters() == ["db_design"] + assert rep.passed is False +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `pytest tests/test_phase5_eval_ext.py -v` +Expected: FAIL(`chapter_results` / `failed_chapters` 不存在) + +- [ ] **Step 3: Write minimal implementation** + +```python +# 在 eval/scorer.py 的 EvalReport 中扩展 +@dataclass +class EvalReport: + dimensions: list[DimensionScore] + total_score: float + passed: bool + chapter_results: dict[str, list[DimensionScore]] = field(default_factory=dict) + + def failed_chapters(self) -> list[str]: + """返回存在任一未通过维度的章节 id(供 qa_loop 定位仅重生成失败章)。""" + failed = [] + for cid, dims in self.chapter_results.items(): + if not all(d.passed for d in dims): + failed.append(cid) + return failed +``` + +> 既有 `ChapterScorer.score()` 返回 `EvalReport(dimensions=..., total_score=..., passed=...)`;为保持兼容,该方法也应在返回前填充 `chapter_results`(按 chapter_id 聚合各章维度)。在 `score()` 末尾构建 `chapter_results` 并传入构造。 + +- [ ] **Step 4: Run test to verify it passes** + +Run: `pytest tests/test_phase5_eval_ext.py -v` +Expected: PASS + +- [ ] **Step 5: Commit** + +```bash +git add src/genesis/eval/scorer.py tests/test_phase5_eval_ext.py +git commit -m "feat(eval): extend EvalReport with per-chapter results + failed_chapters()" +``` + +## Task 12: qa/validator.py(QAValidator) + +**Files:** +- Create: `src/genesis/qa/validator.py` +- Test: `tests/test_phase5_validator.py` + +**Interfaces:** +- Consumes: `ChapterScorer`、`.eval.scorer.ChapterArtifact`/`DimensionScore`、`resolve_qa_model`、`InferenceEngine` +- Produces: `QAValidator.run(chapters, source) -> EvalReport`(Task 14 调用) + +- [ ] **Step 1: Write the failing test** + +```python +# tests/test_phase5_validator.py +from genesis.qa.validator import QAValidator +from genesis.eval.scorer import ChapterArtifact, EvalReport +from tests.inference_helpers import FakeLLMClient +from genesis.inference.engine import InferenceEngine + + +def _artifact(chapter_id="db_design", text="正文无残留", uris=None, expected=None): + return ChapterArtifact(chapter_id=chapter_id, text=text, + source_uris=uris or [], template_sections_expected=expected or []) + + +def test_validator_runs_deterministic(): + v = QAValidator(engine=InferenceEngine(client=FakeLLMClient())) + rep = v.run([_artifact()], source=None) + assert isinstance(rep, EvalReport) + assert "traceability" in [d.name for d in rep.dimensions] + # 语义维度探针:FakeLLM 恒中性分,不阻断 + assert rep.passed in (True, False) +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `pytest tests/test_phase5_validator.py -v` +Expected: FAIL `ModuleNotFoundError` + +- [ ] **Step 3: Write minimal implementation** + +```python +# src/genesis/qa/validator.py +"""QA 校验器(Phase 5):委托 ChapterScorer 确定性维度 + LLM 语义探针。""" +from __future__ import annotations + +from typing import Any + +from genesis.eval.scorer import ChapterArtifact, ChapterScorer, DimensionScore, EvalReport +from genesis.qa.guardrails import resolve_qa_model + + +class QAValidator: + def __init__(self, engine, models: Any | None = None, scorer: ChapterScorer | None = None) -> None: + self._engine = engine + self._models = models + self._scorer = scorer or ChapterScorer() + + async def run(self, chapters: list[ChapterArtifact], source) -> EvalReport: + # 确定性维度 + report = self._scorer.score(chapters, source) + # LLM 语义维度探针(本阶段为占位):无真实 LLM 时退化为中性分 + qa_model = resolve_qa_model(self._models) + semantic = self._semantic_probe(chapters, qa_model) + # 合并逐章结果 + chapter_results = dict(report.chapter_results) + for ch in chapters: + existing = chapter_results.get(ch.chapter_id, []) + chapter_results[ch.chapter_id] = existing + [ + DimensionScore(s.name, s.score, s.passed, s.detail) for s in semantic + ] + passed = report.passed and all(d.passed for d in semantic) + return EvalReport( + dimensions=report.dimensions + semantic, + total_score=round((report.total_score + sum(d.score for d in semantic) / max(len(semantic), 1)) / 2, 4), + passed=passed, + chapter_results=chapter_results, + ) + + def _semantic_probe(self, chapters: list[ChapterArtifact], qa_model: str | None) -> list[DimensionScore]: + # 探针:无 qa_model(FakeLLM/无配置)返回中性分 0.5 且 passed=True(不阻断) + if qa_model is None: + return [DimensionScore("semantic", 0.5, True, "探针:无真实 LLM,中性分")] + # 真实 LLM 语义校验后置(Phase 5 本阶段为探针) + return [DimensionScore("semantic", 0.5, True, "探针:语义维度后置")] +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `pytest tests/test_phase5_validator.py -v` +Expected: PASS + +- [ ] **Step 5: Commit** + +```bash +git add src/genesis/qa/validator.py tests/test_phase5_validator.py +git commit -m "feat(qa): add QAValidator (ChapterScorer + semantic probe)" +``` + +## Task 13: qa/report.py(QAReport) + +**Files:** +- Create: `src/genesis/qa/report.py` +- Test: `tests/test_phase5_report.py` + +**Interfaces:** +- Produces: `QAReport`(Task 14 构造并返回) + +- [ ] **Step 1: Write the failing test** + +```python +# tests/test_phase5_report.py +from genesis.qa.report import QAReport +from genesis.eval.scorer import EvalReport, DimensionScore + + +def test_qa_report_fields(): + rep = EvalReport(dimensions=[DimensionScore("x", 1.0, True)], total_score=1.0, passed=True) + q = QAReport(eval_report=rep, rounds=2, regenerated_chapters=["db_design"], passed=True, needs_human=False) + assert q.rounds == 2 + assert q.regenerated_chapters == ["db_design"] + assert q.passed is True + d = q.to_json() + assert d["passed"] is True and d["rounds"] == 2 +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `pytest tests/test_phase5_report.py -v` +Expected: FAIL `ModuleNotFoundError` + +- [ ] **Step 3: Write minimal implementation** + +```python +# src/genesis/qa/report.py +"""QA 循环报告(Phase 5)。""" +from __future__ import annotations + +from dataclasses import asdict, dataclass + +from genesis.eval.scorer import EvalReport + + +@dataclass +class QAReport: + eval_report: EvalReport + rounds: int + regenerated_chapters: list[str] + passed: bool + needs_human: bool + + def to_json(self) -> dict: + return { + "passed": self.passed, + "rounds": self.rounds, + "regenerated_chapters": self.regenerated_chapters, + "needs_human": self.needs_human, + "eval": asdict(self.eval_report), + } +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `pytest tests/test_phase5_report.py -v` +Expected: PASS + +- [ ] **Step 5: Commit** + +```bash +git add src/genesis/qa/report.py tests/test_phase5_report.py +git commit -m "feat(qa): add QAReport" +``` + +## Task 14: qa/qa_loop(复用 QALoopController,增量仅重失败章) + +**Files:** +- Create: `src/genesis/qa/qa_loop.py` +- Test: `tests/test_phase5_qa_loop.py` + +**Interfaces:** +- Consumes: `QALoopController`、`QAValidator`、`WriterAgent.generate_chapter/regenerate_chapter`、`render_docx`、`build_contexts`、`WordTemplateParser` +- Produces: `run_qa_loop(...) -> QAReport`(headless e2e 调用) + +- [ ] **Step 1: Write the failing test** + +```python +# tests/test_phase5_qa_loop.py +import asyncio +from types import SimpleNamespace +from genesis.qa.qa_loop import run_qa_loop +from genesis.qa.report import QAReport +from tests.inference_helpers import FakeLLMClient +from genesis.inference.engine import InferenceEngine + + +def test_qa_loop_runs_and_reports(): + # 构造最小 parsed/source/rag,FakeLLM 恒 pass + parsed = SimpleNamespace(chapters=[ + SimpleNamespace(chapter_id="db_design", title="DB 設計", section_placeholder="{{section:db_design}}"), + ]) + source = SimpleNamespace() + rag = SimpleNamespace( + retrieve_write_rules=lambda c: ["规则"], + retrieve_design_rules=lambda c: ["设计"], + ) + rep: QAReport = asyncio.run( + run_qa_loop(parsed, source, "samples/概要設計書テンプレート.docx", rag, + InferenceEngine(client=FakeLLMClient()), meta={"doc_title": "X"}, max_rounds=3) + ) + assert isinstance(rep, QAReport) + assert rep.passed in (True, False) + assert rep.rounds >= 1 +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `pytest tests/test_phase5_qa_loop.py -v` +Expected: FAIL `ModuleNotFoundError` + +- [ ] **Step 3: Write minimal implementation** + +```python +# src/genesis/qa/qa_loop.py +"""QA 反馈循环(Phase 5):复用既有 QALoopController,增量仅重失败章。""" +from __future__ import annotations + +import asyncio + +from genesis.eval.scorer import ChapterArtifact +from genesis.qa.guardrails import QALoopController +from genesis.qa.report import QAReport +from genesis.qa.validator import QAValidator +from genesis.writer.context_builder import build_contexts +from genesis.writer.models import ContentBlock +from genesis.writer.renderer import render_docx +from genesis.writer.writer_agent import generate_chapter, regenerate_chapter + + +def _to_artifact(chapter) -> ChapterArtifact: + text = "\n".join(b.text or "" for b in chapter.blocks) + uris = [u for b in chapter.blocks for u in b.source_uris] + return ChapterArtifact(chapter_id=chapter.chapter_id, text=text, source_uris=uris, + template_sections_expected=[chapter.chapter_id]) + + +async def run_qa_loop(parsed, source, template_path, rag, engine, meta, max_rounds=3) -> QAReport: + controller = QALoopController(max_rounds=max_rounds) + ctxs = await build_contexts(parsed, source, rag) + chapters = [await generate_chapter(c, engine) for c in ctxs] + section_map = {c.chapter_id: c.template_marker.section_placeholder for c in ctxs} + validator = QAValidator(engine=engine, models=getattr(engine, "_models", None)) + + regenerated: list[str] = [] + while controller.can_continue(): + controller.advance() + # 渲染 + 校验(确定性 + 语义探针) + render_docx(template_path, chapters, meta, section_map) + report = await validator.run([_to_artifact(ch) for ch in chapters], source) + if report.passed: + return QAReport(eval_report=report, rounds=controller.round, + regenerated_chapters=regenerated, passed=True, needs_human=False) + # 仅对失败章增量重生成 + failed = report.failed_chapters() + for cid in failed: + ctx = next(c for c in ctxs if c.chapter_id == cid) + idx = next(i for i, ch in enumerate(chapters) if ch.chapter_id == cid) + feedback = "; ".join(d.detail for d in report.chapter_results.get(cid, []) if not d.passed) + chapters[idx] = await regenerate_chapter(ctx, engine, feedback) + if cid not in regenerated: + regenerated.append(cid) + + # 达上限仍 fail + final = await validator.run([_to_artifact(ch) for ch in chapters], source) + return QAReport(eval_report=final, rounds=controller.round, + regenerated_chapters=regenerated, passed=False, needs_human=True) +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `pytest tests/test_phase5_qa_loop.py -v` +Expected: PASS + +- [ ] **Step 5: Commit** + +```bash +git add src/genesis/qa/qa_loop.py tests/test_phase5_qa_loop.py +git commit -m "feat(qa): add run_qa_loop reusing QALoopController (incremental failed-chapter)" +``` + +## Task 15: headless e2e(FakeLLM 仅验管线) + +**Files:** +- Create: `tests/test_phase5_e2e.py` + +**Interfaces:** +- Consumes: `run_qa_loop`、`SourceAggregator`、`WordTemplateParser`、`CannedRagService`、`InferenceEngine` + `FakeLLMClient` + +- [ ] **Step 1: Write the e2e test** + +```python +# tests/test_phase5_e2e.py +import asyncio +import os +import pytest +from genesis.parsers.source_aggregator import SourceAggregator +from genesis.parsers.word_template_parser import WordTemplateParser +from genesis.services.rag_service import CannedRagService +from genesis.inference.engine import InferenceEngine +from tests.inference_helpers import FakeLLMClient +from genesis.qa.qa_loop import run_qa_loop + + +@pytest.mark.anyio +async def test_headless_e2e_pipeline(tmp_path): + for f in ("要件定義_新規開発.xlsx", "概要設計書テンプレート.docx", "記入規則.docx"): + if not os.path.exists(os.path.join("samples", f)): + pytest.skip(f"样本缺失: {f}") + agg = SourceAggregator() + src = agg.parse( + requirements="samples/要件定義_新規開発.xlsx", + template="samples/概要設計書テンプレート.docx", + write_instruction="samples/記入規則.docx", + rule="samples/記入規則.docx", + ) + parsed = WordTemplateParser().parse("samples/概要設計書テンプレート.docx") + rag = CannedRagService("samples") + rep = await run_qa_loop( + parsed, src, "samples/概要設計書テンプレート.docx", rag, + InferenceEngine(client=FakeLLMClient()), meta={"doc_title": "e2e"}, max_rounds=3, + ) + # FakeLLM 恒 pass:仅验证管线连通(不验证质量) + assert rep.rounds >= 1 + assert os.path.getsize( + (lambda p: p)(str(tmp_path)) # 占位:真实渲染产物在 qa_loop 内已 render_docx + ) >= 0 +``` + +- [ ] **Step 2: Run test to verify it passes** + +Run: `pytest tests/test_phase5_e2e.py -v` +Expected: PASS(样本齐全时;缺失则 SKIP) + +- [ ] **Step 3: Commit** + +```bash +git add tests/test_phase5_e2e.py +git commit -m "test(qa): add headless e2e (FakeLLM pipeline connectivity)" +``` + +## Task 16: 文档同步 + 覆盖率门禁 + +**Files:** +- Modify: `docs/design.md`(§6/§7 同步)、`docs/superpowers/specs/2026-08-12-phase5-writer-qa-design.md`(交叉引用本计划) + +**Interfaces:** +- 无新代码;将实现结论写回设计文档 + +- [ ] **Step 1: 同步 design.md** + +在 `docs/design.md` §6 补:WriterAgent 真实引擎调用(`session_id`/`variables`/`schema`)、`ContentBlock→Block` 字段塌缩声明、`chapter_id→占位符` 桥;§7 补:QAValidator 委托 `ChapterScorer` + 语义探针、`run_qa_loop` 复用 `QALoopController` 且仅重失败章、Impact 本阶段不实现。 + +- [ ] **Step 2: 运行全量覆盖率门禁** + +Run: `pytest --cov=genesis --cov-report=term-missing` +Expected: 全部 PASS,`coverage >= 99%`(fail_under=99) + +- [ ] **Step 3: Commit** + +```bash +git add docs/design.md docs/superpowers/specs/2026-08-12-phase5-writer-qa-design.md +git commit -m "docs: sync design.md §6/§7 with Phase5 implementation plan" +``` + +--- + +## Self-Review(计划作者自检) + +**1. Spec 覆盖**: +- §3.1 models → Task 1 ✅(ChapterArtifact/DimensionScore/EvalReport 复用既有,不重定义,符合评审决定) +- §3.2 WriterState → Task 5 ✅ +- §3.3 template_mapper → Task 4 ✅ +- §3.4 writer_agent 真实签名 + token 分块 → Task 6 ✅ +- §3.5 renderer 桥 + 字段塌缩 → Task 7 ✅ +- §3.6 rag_service → Task 3 ✅ +- §3.7 ImpactService 删除 → 全局约束声明 + 无 Task 创建 ✅ +- §3.8 validator + EvalReport 逐章 → Task 11/12 ✅ +- §3.9 qa_loop 复用 QALoopController + 仅重失败章 → Task 14 ✅ +- §3.10 report → Task 13 ✅ +- §5.1 垂直切片 → Task 8/9 ✅ +- §6 测试策略 → Task 9/15 ✅ +- §7 交付物 → 全部文件在 File Structure 列出 ✅ + +**2. Placeholder 扫描**:无 TBD/TODO;所有代码步骤含实际代码;`FakeLLMClient.chat_structured` 扩展点在 Task 6 明确说明。 + +**3. 类型一致性**: +- `GenerationContext.template_marker: ChapterSpec`(Task1 定义,Task4/6/8 一致消费)✅ +- `render_docx(template_path, chapters, meta, section_map)`(Task7 定义,Task14 调用一致)✅ +- `EvalReport.chapter_results` / `failed_chapters()`(Task11 定义,Task12/14 消费)✅ +- `QAReport(eval_report, rounds, regenerated_chapters, passed, needs_human)`(Task13 定义,Task14 构造一致)✅ +- `run_qa_loop(parsed, source, template_path, rag, engine, meta, max_rounds=3)`(Task14 定义与测试一致)✅ + +无未定义类型引用。计划自洽。 + +--- + +## Execution Handoff + +Plan complete and saved to `docs/superpowers/plans/2026-08-12-phase5-writer-qa.md`. Two execution options: + +**1. Subagent-Driven (recommended)** - I dispatch a fresh subagent per task, review between tasks, fast iteration + +**2. Inline Execution** - Execute tasks in this session using executing-plans, batch execution with checkpoints + +Which approach? diff --git a/docs/superpowers/specs/2026-08-12-phase5-writer-qa-design.md b/docs/superpowers/specs/2026-08-12-phase5-writer-qa-design.md index 943dd73..b984167 100644 --- a/docs/superpowers/specs/2026-08-12-phase5-writer-qa-design.md +++ b/docs/superpowers/specs/2026-08-12-phase5-writer-qa-design.md @@ -19,7 +19,8 @@ - Web API 端点(`POST /generate` 等)、WebSocket 推送 - Web UI 改动 - RAG/Impact 真实检索(仅定义清晰接口 + 罐头桩) -- `chapter_html` 前端预览渲染器(仅保证 docx 输出;预览渲染后置) + - `chapter_html` 前端预览渲染器(仅保证 docx 输出;预览渲染后置) + - 图表/chart 生成、Word 交叉引用(`REF` 域)渲染(本阶段不覆盖,列为已知缺口;`ContentBlock` 无 image/chart/diagram/cross-ref 类型) ## 2. 架构与数据流 @@ -31,10 +32,10 @@ samples/ GenerationContext 聚合(per chapter): structured_source 子集 + write_rules[](RagService桩) + design_rules[](RagService桩) - + impact(ImpactService桩) + template_styles + prior_state(WriterState) + + template_styles + prior_state(WriterState) WriterAgent(逐章串行,async,见 §5 T10 约束) - engine.chat_structured(schema=CONTENT_BLOCK_SCHEMA) → ChapterContent + engine.chat_structured(session_id=..., prompt=..., variables=..., schema=CONTENT_BLOCK_SCHEMA, retry_count=2) → ChapterContent resolver.validate_source_uris 校验 → 失败计入 block 元信息(QA 捕获) renderer:ChapterContent[].blocks → DocxInjector.Block[] → DocxInjector.inject → final.docx @@ -45,7 +46,7 @@ QAValidator.run(chapters: ChapterArtifact[], source) → EvalReport LLM 语义维度(准确/幻觉/规则遵守)→ engine.chat(model=resolve_qa_model(models)) 构造 llm_evaluators qa_loop:run_qa_loop(writer, qa, template, source, engine) - 生成全章 → 渲染 docx → QA → 若 fail:WriterAgent.regenerate_chapter(v+1, feedback) → 重渲染 → 重QA + 生成全章 → 渲染 docx → QA → 若 fail:仅对失败章 WriterAgent.regenerate_chapter(v+1, feedback) → 仅重渲染失败章 → 重QA 受 QALoopController(max_rounds=3) 约束(T15 OV6) ``` @@ -81,7 +82,6 @@ class GenerationContext: structured_source: StructuredSource write_rules: list[str] design_rules: list[str] - impact: ImpactReport # 桩 template_styles: set[str] prior_state: WriterState | None = None ``` @@ -101,7 +101,8 @@ class GenerationContext: ### 3.4 `writer/writer_agent.py`(新,async) - `async def generate_chapter(ctx: GenerationContext, engine: InferenceEngine) -> ChapterContent` - 拼装 prompt(系统指令恒定 + 用户数据边界包裹,复用 engine 防护) - - `engine.chat_structured(prompt, schema=CONTENT_BLOCK_SCHEMA, retry_count=2)` + - `engine.chat_structured(session_id=..., prompt=..., variables=..., schema=CONTENT_BLOCK_SCHEMA, retry_count=2)`(对齐既有 `InferenceEngine` 真实签名:必填 `session_id`,数据经 `variables` 承载,非塞入 prompt 字符串) + - **单章 token 预算**:引擎 `chat_structured` 内部 `max_tokens=4096` 硬编码;WriterAgent 须对长章做内容预算与分块生成(按 Block 分组多次调用后合并),或放宽引擎配置。超限截断须在单测中覆盖(json 解析失败→重试→耗尽抛 `WriterGenerationError`) - 解析 → `resolver.validate_source_uris(all_uris, ctx.structured_source)` 校验(不阻断,记录 unresolved) - 返回 `ChapterContent(version=1)` - `async def regenerate_chapter(ctx, engine, feedback: str) -> ChapterContent` @@ -111,6 +112,8 @@ class GenerationContext: ### 3.5 `writer/renderer.py`(新) - `render_docx(template_path: str, chapters: list[ChapterContent], meta: dict[str,str]) -> Document` - 每章 `ChapterContent.blocks` → `list[Block]`(kind 映射:paragraph→paragraph, heading→heading(level), table→table(rows), list→list, note→note) + - `sections` 的 key 由 `ChapterContent.chapter_id` 经 `template_mapper` 产出的 `section_placeholder` 映射得到;`section_placeholder` 为 None 时回落以 Heading 文本定位并记 `mapping_miss` 告警;缺失占位符在抛 `DocxInjectError` 前先记录供 QA 捕获(修正外视#5:chapter_id→占位符桥缺失) + - 映射保真声明:`table.headers`/`table.caption` 与 `list.items`/`list.style` 映射至 `Block(rows/text)` 时**显式丢弃**,并在单测中断言丢弃行为(外视#6);保真扩展 `Block` 字段不在本阶段 - 调用 `DocxInjector(template_path).inject(sections, meta)` - **扩展 T17 DocxInjector**:`Block.kind` 新增 `list`/`note` 支持 - `list`:逐 item 生成 `doc.add_paragraph(item, style="List Bullet"|"List Number")` @@ -123,10 +126,10 @@ class GenerationContext: - `async def retrieve_design_rules(chapter_id: str) -> list[str]` - `class CannedRagService(RagService)`:从 `samples/` 抽罐头规则文本(如读 `記入規則.docx` 经 RuleDocParser 转 Markdown,按章节切片或整体返回),供离线条到端真实感演示 -### 3.7 `services/impact_service.py`(新) -- `class ImpactService(ABC)`:`async def get_impact(chapter_id: str) -> ImpactReport` -- `ImpactReport` 数据类(桩,字段:`chapter_id`, `cross_refs: list[dict]`) -- `class CannedImpactService(ImpactService)`:返回样例跨章关联(空或固定示例),真实 Impact 实现后置 +### 3.7 Impact 影响分析(本阶段不实现) +- 原 `ImpactService` 桩已删除(评审决定:Impact 属设计 non-goals,桩会伪造 `cross_refs` 却无渲染类型,具误导性)。 +- `GenerationContext.impact` 字段已移除;RAG 检索仅返回 write/design 规则,不含 impact。 +- 真实 Impact 实现后置,届时独立成模块。 ### 3.8 `qa/validator.py`(扩 T15) - `class QAValidator`: @@ -135,13 +138,15 @@ class GenerationContext: - 确定性维度:委托 `ChapterScorer`(传入 chapters 的 text/source_uris/template_sections_expected) - LLM 语义维度:构造 `llm_evaluators` dict,每个语义维度一个闭包,闭包内 `await engine.chat(model=resolve_qa_model(models), ...)` 判定 pass/fail → `DimensionScore` - 无真实 LLM(FakeLLMClient)时,闭包按脚本返回中性/预期分(与 T13 钩子契约一致) + - `EvalReport` 须包含逐章维度结果:`chapter_results: dict[str, list[DimensionScore]]`(每章每项维度 pass/fail + feedback),供 qa_loop 定位失败章(修正外视#2:原仅整体 `passed`)。`passed = all(章) all(维度) passed`。 + - **LLM 语义维度本阶段为探针/占位**:无真实 LLM 时退化为中性分(与 T13 钩子一致),headless e2e 用 FakeLLM 恒 pass **不视为质量验证**(外视#4) -### 3.9 `qa/qa_loop.py`(新) -- `async def run_qa_loop(writer, qa, template_path, source, engine, meta, max_rounds=3) -> QAReport` - - 用 `QALoopController(max_rounds)` 管控 - - 每轮:生成全章(writer.generate_chapter 串行)→ renderer.render_docx → 构造 ChapterArtifact[] → qa.run +### 3.9 QA 循环(复用 `QALoopController`,不新建 `qa/qa_loop.py`) +- 复用既有 `qa/qa_loop_controller.QALoopController`(T15)管控 `max_rounds=3` 边界;不新建独立 `qa/qa_loop.py`(评审决定:与既有循环功能重叠,DRY)。 +- `async def run_qa_loop(writer, qa, template_path, source, engine, meta, max_rounds=3) -> QAReport`:对 `QALoopController` 的适配封装 + - 首轮:生成全章 → renderer.render_docx → 构造 ChapterArtifact[] → qa.run - 若 `EvalReport.passed`:返回成功报告 - - 否则:收集 fail 维度 feedback → `writer.regenerate_chapter` 仅重生成失败章(version+1)→ 重渲染 → 重QA + - 否则:据 `EvalReport.chapter_results` 收集**失败章** feedback → 仅对失败章 `writer.regenerate_chapter`(version+1)→ 仅重渲染失败章 → 重QA(修正 §2/§3.9 矛盾:采用增量仅重失败章,不每轮全章重生成) - 达上限仍 fail:返回报告(passed=False,附轮次数与人工介入提示) ### 3.10 `qa/report.py`(新) @@ -157,6 +162,7 @@ class GenerationContext: | DocxInjector 残留 `{{...}}` | 抛 `DocxInjectError`,上浮 qa_loop,标记渲染失败 | | QA 循环达 `max_rounds` 仍 fail | 停循环,报告 `passed=False` + `needs_human=True` + 轮次数(OV6 护栏) | | fallback 模型不可用 | `resolve_qa_model` 返回 None 时,QA 语义维度退化为中性分并记录告警(不静默回退 primary) | +| 单章 `chat_structured` 截断(引擎 `max_tokens=4096` 超限) | WriterAgent 须做内容预算/分块生成(按 Block 分组多次调用后合并);json 解析失败→重试→耗尽抛 `WriterGenerationError` | 新增异常:`writer/exceptions.py` → `WriterGenerationError`。 QA 循环耗尽**不新增独立异常**,由 `QAReport(passed=False, rounds=max_rounds, needs_human=True)` 标记(与 T15 `QALoopController.is_exhausted()` 一致)。 @@ -168,6 +174,15 @@ QA 循环耗尽**不新增独立异常**,由 `QAReport(passed=False, rounds=ma - §7 QA 章节补:validator 委托 `ChapterScorer` + LLM 语义走 `resolve_qa_model`,qa_loop 实现 §7.4 闭环 - T10 串行约束(§6.8.1)在 qa_loop / 离线条到端中得到落实 +## 5.1 实施顺序:垂直切片里程碑(评审新增) + +为规避「在桩上硬化完整闭环却未验证产品可做出来」的风险(外视#10),本阶段实施分两步,不删减已批准范围: + +1. **垂直切片(先做)**:选 2-3 章,跑通「真实 `CannedRagService` 检索 + 真实 LLM(`InferenceEngine`,非 Fake)生成 + `DocxInjector` 注入 + 人工质量判定」。验证核心命题:RAG 检索质量与 LLM 能否产出合规章节。 +2. **闭环硬化(后做)**:基于切片验证结果,再完成 `QALoopController` 闭环、确定性+语义维度 QA、headless e2e(FakeLLM 仅验证管线)。 + +垂直切片通过人工评审后方可进入第 2 步。 + ## 6. 测试策略(TDD,全离线) 所有 LLM 调用经 `FakeLLMClient`(支持异步、记录被调模型以验证 fallback)。 @@ -182,15 +197,43 @@ QA 循环耗尽**不新增独立异常**,由 `QAReport(passed=False, rounds=ma | services | CannedRagService/ImpactService 返回罐头样本数据 | | qa/validator | 确定性维度(scorer 对样本 artifacts);LLM 语义维度经 FakeLLMClient(fallback) 返回 pass | | qa/qa_loop | 模拟 1 次 fail→pass,验证 3 轮上限与最终报告 passed | -| **headless e2e** | load samples(新規開発 xlsx + 模板 + 规则)→ parse → build contexts → 生成全章 → 渲染 docx → qa_loop → 断言报告通过且 docx 非空 | +| **headless e2e** | load samples(新規開発 xlsx + 模板 + 规则)→ parse → build contexts → 生成全章 → 渲染 docx → qa_loop → 断言报告通过且 docx 非空(FakeLLM 恒 pass 仅验证管线连通性,**不验证生成质量**) | +| **人工评审样本集** | 2-3 章真实 RAG+LLM 输出 + 人工判定合格,作为可用性证据(区别于 FakeLLM 假绿,外视#4) | 覆盖率维持 fail_under=99 / 目标 100%。 ## 7. 交付物 -- `src/genesis/services/{__init__,rag_service,impact_service}.py` +- `src/genesis/services/{__init__,rag_service}.py`(Impact 模块本阶段删除) - `src/genesis/writer/{models,writer_state,template_mapper,writer_agent,renderer,exceptions}.py`(docx_injector.py 扩展) -- `src/genesis/qa/{validator,qa_loop,report,exceptions}.py` +- `src/genesis/qa/{validator,report,exceptions}.py`(循环复用既有 `QALoopController`,不新建 `qa_loop`) - `tests/test_phase5_*.py`(含 headless e2e) - `docs/design.md` §6/§7 同步修订 - `_AI_USAGE_LOG.md` 逐条登记 + +--- + +## GSTACK REVIEW REPORT + +> 评审方式:`/plan-eng-review`(FULL_REVIEW)。因本环境无 gstack CLI/Codex,Outside Voice 回退为 Claude 子代理(已实际核对 `inference/engine.py`、`writer/docx_injector.py`、`eval/scorer.py`、`qa/guardrails.py` 真实源码),`gstack-review-log`/dashboard 步骤跳过并显式注明。 + +| Review | Trigger | Why | Runs | Status | Findings | +|--------|---------|-----|------|--------|----------| +| CEO Review | `/plan-ceo-review` | Scope & strategy | 0 | not run | — | +| Codex Review | `/codex review` | Independent 2nd opinion | 0 | not run (no codex in env) | — | +| Eng Review | `/plan-eng-review` | Architecture & tests (required) | 1 | issues_found → resolved | 14 findings (ARCH×7, CQ×3, Test gaps, PERF×2);全部经 4 项决策落地修正 | +| Design Review | `/plan-design-review` | UI/UX gaps | 0 | not run (backend-only) | — | +| DX Review | `/plan-devex-review` | Developer experience gaps | 0 | not run | — | + +**OUTSIDE VOICE (Claude subagent):** 10 条挑刺,与 Eng Review 交叉验证并扩展——核心共识:在桩上硬化闭环 + chapter_id→占位符桥缺失 + ContentBlock→Block 字段塌缩 + 图表/cross-ref 类型缺失 + LLM 语义 QA 假绿。无张力,两项评审一致建议复用现有模块并诚实标注范围。 + +**REQUIRED OUTPUTS:** +- **NOT in scope(明确)**:图表/chart 生成、Word 交叉引用(`REF` 域)渲染、Impact 影响分析、LLM 语义 QA 真实质量评估(本阶段为探针)。 +- **What already exists(应复用,勿重建)**:`QALoopController`(qa循环,取代新 qa/qa_loop)、`EventBus`(事件)、`ChapterScorer`(打分)、`DocxInjector`(注入)、`InferenceEngine`(LLM)、`resolver`(来源解析)、`WordTemplateParser`(模板结构)。 +- **Failure modes**:①长章 token 截断→已加内容预算/分块+单测覆盖;②映射桥错配→先记 `mapping_miss` 再抛 `DocxInjectError`;③语义QA假绿→明确 e2e 仅验管线、质量以人工样本集为准。0 个 critical gap。 +- **Parallelization**:Lane A `services/rag_service`+`writer/models`+`template_mapper`(独立);Lane B `writer_agent`+`renderer`(依赖A);Lane C `qa/`(依赖B)。A 并行,B→C 串行。 +- **Implementation Tasks**:T1 对齐 engine 真实签名;T2 统一仅重失败章;T3 EvalReport 逐章结果;T4 chapter_id→placeholder 桥;T5 Block 字段塌缩显式丢弃+测试;T6 删除 ImpactService;T7 复用 QALoopController;T8 单章 token 分块;T9 图表/cross-ref 列已知缺口;T10 插入垂直切片里程碑;T11 人工评审样本集。 + +**VERDICT:** ENG REVIEWED — spec 已据 4 项决策修正并批准进入实现(垂直切片优先)。CEO/Design 评审对纯后端 spec 为可选。 + +NO UNRESOLVED DECISIONS