From 17cf385a6fd582718ad49e163694883b7b9c17e3 Mon Sep 17 00:00:00 2001 From: lhl Date: Sun, 9 Aug 2026 23:05:28 +0800 Subject: [PATCH] =?UTF-8?q?docs:=20Phase3=20Word=20=E8=A7=A3=E6=9E=90?= =?UTF-8?q?=E4=BC=98=E5=85=88=E8=AE=BE=E8=AE=A1=EF=BC=88=E6=A8=A1=E6=9D=BF?= =?UTF-8?q?/=E8=A7=84=E5=88=99/=E8=81=9A=E5=90=88=E4=B8=89=E6=A8=A1?= =?UTF-8?q?=E5=9D=97=EF=BC=89?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- _AI_USAGE_LOG.md | 1 + .../2026-08-09-phase3-word-parser-design.md | 160 ++++++++++++++++++ 2 files changed, 161 insertions(+) create mode 100644 docs/superpowers/specs/2026-08-09-phase3-word-parser-design.md diff --git a/_AI_USAGE_LOG.md b/_AI_USAGE_LOG.md index 67f8647..106f7da 100644 --- a/_AI_USAGE_LOG.md +++ b/_AI_USAGE_LOG.md @@ -59,3 +59,4 @@ | 2026-08-09 | 测试验证 | 里程碑3.1 LLM错误码对齐收尾改动(最终评审 Minor):① 删除 tests/test_inference_engine.py 中 test_chat_failed_error_code_not_configured 首行未使用变量 client = FakeLLMClient(...)(死代码,未引用,函数自洽性核查通过);② 补齐 src/genesis/inference/engine.py 文件末尾换行符(最后一字节 ) → \n)。覆盖测试 22 passed;pytest 全量 128 passed,覆盖 100.00%(803 stmts/188 br),fail_under=99 达标;提交见 git log | tests/test_inference_engine.py, src/genesis/inference/engine.py, _AI_USAGE_LOG.md | deepseek-v4-flash-free | | 2026-08-09 | 测试验证 | 新增 api §7 ↔ 异常树 error_code 一致性门禁测试(防文档漂移)。新建 tests/test_api_design_consistency.py(4 用例:树→文档 / 文档→树双向断言 + 行数护栏 / 空解析护栏);首次运行 RED(行数护栏 ==5 与事实 4 冲突,文档与异常各为 4 条/4 个子类,双向一致)→ 修正护栏为 4 → GREEN(聚焦 4 passed);pytest 全量 132 passed 覆盖 100.00%(803 stmts/188 br),fail_under=99 达标;提交见 git log | tests/test_api_design_consistency.py, _AI_USAGE_LOG.md | deepseek-v4-flash-free | | 2026-08-09 | 测试验证 | 一致性门禁最终评审收尾修复(Important x1 + Minor x3):① plan 文档 Interfaces 段 L311-314「五行 LLM 行」→「四行」(与 ==4 护栏一致);② 测试文件补尾随换行(末字节 " → \n);③ 断言消息改「异常类存在文档未登记或来源类名失配的 error_code」(覆盖双失败场景);④ _LLM_ROW_RE 加正则隐式限域注释。覆盖测试 4 passed;pytest 全量 132 passed 覆盖 100.00%,fail_under=99 达标 | docs/superpowers/plans/2026-08-09-error-code-consistency-test.md, tests/test_api_design_consistency.py, _AI_USAGE_LOG.md | deepseek-v4-flash-free | +| 2026-08-09 | 架构设计 | Phase3 Word 解析优先设计(brainstorming):范围澄清(Word 解析优先=3.1/3.2/3.6/3.7,PPT/现有系统探索后续;规则分类按来源映射零LLM;SourceAggregator 全量整合 StructuredSource;样式名级提取;方案A 三模块门面聚合);输出设计文档 docs/superpowers/specs/2026-08-09-phase3-word-parser-design.md | docs/superpowers/specs/2026-08-09-phase3-word-parser-design.md, _AI_USAGE_LOG.md | deepseek-v4-flash-free | diff --git a/docs/superpowers/specs/2026-08-09-phase3-word-parser-design.md b/docs/superpowers/specs/2026-08-09-phase3-word-parser-design.md new file mode 100644 index 0000000..37906ad --- /dev/null +++ b/docs/superpowers/specs/2026-08-09-phase3-word-parser-design.md @@ -0,0 +1,160 @@ +# Phase 3「Word 解析优先」设计(Parser Agent - Word 解析 + 聚合) + +> **状态:** 已批准(brainstorming + 范围澄清吸收) +> **日期:** 2026-08-09 +> **里程碑:** implementation-plan 阶段 3(Word 解析优先切片) + +## 1. 背景与目标 + +`implementation-plan.md` 阶段 3 定义 Parser Agent 的 Word/PPT 解析与现系统探索。经范围澄清(用户确认),本次迭代只做 **Word 解析优先** 切片:实现为后续 Writer(阶段 7)、RAG 规则检索(阶段 4)、QA(阶段 8)服务的 **Word 模板解析 / Word 规则文档 Markdown 化 / 全量输入聚合** 基础设施。 + +**现状:** +- `data_models.py` 已定义 `ParsedTemplate` / `ChapterMarker` / `RuleDocument` / `StructuredSource` / `UnifiedDocument` 等数据类,可直接复用 +- `parsers/` 已有完整 Excel 解析链(ExcelParser → ExcelParseResult),且 `python-docx>=1.0` 已在 `pyproject.toml` +- `samples/` 已有 3 个 Word 样本:`概要設計書テンプレート.docx`(模板)、`概要設計做成説明書.docx`(做成说明书)、`記入規則.docx`(记入规则),以及 5 个 Excel 样本 +- 全量测试基线 **132 passed / 100.00% 覆盖**(fail_under=99 红线) + +**目标:** 新增 WordTemplateParser(章构成/占位符/样式名提取)、RuleDocParser(docx → Markdown + 规则分类)、SourceParser 门面(按扩展名路由全量输入 → StructuredSource),并配套单元测试与真实样本集成测试,保持红线不回归。 + +## 2. 范围与不做的事 + +**范围(implementation-plan 3.1 / 3.2 / 3.6 / 3.7):** +- 3.1 WordTemplateParser:章构成(Heading 层级)提取、占位符(`{{section:xxx}}`)检测、样式名提取 +- 3.2 RuleDocParser(Word):规则文档 Markdown 化、规则分类(按来源映射) +- 3.6 SourceAggregator:全量输入(Excel 要件定义 + Word 模板 + Word 规则)统一组装 `StructuredSource` +- 3.7 Parser Agent 集成测试:真实样本(3 Word + 5 Excel)全链路 + +**不做(YAGNI / 后续里程碑):** +- ❌ 3.3 PPTXParser(无 .pptx 样本、python-pptx 未安装)→ 后续补样本后实施 +- ❌ 3.4 现有系统代码探索(CodeParser 未实现、无 .java 样本)→ 后续实施 +- ❌ 3.5 现有系统设计书探索(无现系统 Word/Excel 设计书样本)→ 后续实施 +- ❌ FileReader 统一读取层(UnifiedDocument)(方案 B 否决;现有 Excel 链直接 openpyxl,本轮不重构) +- ❌ 完整样式定义提取(字号/颜色等)(YAGNI:阶段 3 无样式消费方,阶段 7/8 再按需深挖) +- ❌ LLM 参与的规则分类(按来源映射,零网络依赖、离线可测) +- ❌ data_models.py 数据模型改动(现有类型完全够用) + +## 3. 设计 + +### 3.1 架构 + +在 `src/genesis/parsers/` 下新增 3 个模块,与现有 `excel_parser.py` 同构(扁平模块、dataclass 结果、纯函数拆分): + +``` +parsers/ +├── excel_parser.py # 现有,不动 +├── word_template_parser.py # 新:WordTemplateParser +├── rule_doc_parser.py # 新:RuleDocParser +└── source_aggregator.py # 新:SourceParser 门面 +``` + +### 3.2 WordTemplateParser + +**输入:** docx 路径(`str | Path`) +**输出:** `ParsedTemplate{file_name, sections, placeholders, styles}`(复用 data_models 类型) + +提取逻辑(`python-docx`): + +| 产物 | 提取方式 | 说明 | +|------|---------|------| +| `sections` | 遍历文档段落,识别 Heading 样式的段落 → `ChapterMarker(type="heading", name=段落文本, level=大纲级别)`;遍历书签(`bookmarkStart`)→ `type="bookmark"`;遍历占位符 → `type="placeholder"` | `ChapterMarker.type` 取值 `"heading" | "bookmark" | "placeholder"` | +| `placeholders` | 正则识别 `{{section:xxx}}` 与封面型 `{{doc_title}}` / `{{version}}` / `{{created_at}}` → `{占位符名: 出现位置或上下文}` | 占位符名去重 | +| `styles` | 收集文档命名样式 + 各段落实际使用的样式名 → 去重集合 | **样式名级**(不提取字号/颜色/字体明细) | + +### 3.3 RuleDocParser + +**输入:** docx 路径 + `category`(调用方按来源映射传入:`"write" | "design" | "ref"`) +**输出:** `RuleDocument{file_name, category, markdown_content, source_path, file_type="word", hash}` + +Markdown 化规则: + +| docx 元素 | Markdown 输出 | +|-----------|--------------| +| Heading 1/2/3 | `#` / `##` / `###` 标题行 | +| 普通段落 | 文本行 | +| 表格 | GFM 表格(表头 + 分隔行 + 数据行) | +| 列表(List Bullet/Number) | `- ` 项 / 编号项 | +| 空段 | 空行 | + +- `hash` = 文件内容 sha256 hex(供阶段 4 规则版本管理) +- `file_type` 固定 `"word"` + +**来源映射(分类策略,零 LLM):** + +| 输入文件 | category | +|---------|----------| +| `記入規則.docx` | `write` | +| `概要設計做成説明書.docx` | `ref` | + +**明确的策略:** `rule_paths` 仅接受 `.docx` 规则文档;类别由门面按文件名映射(`記入規則*` → `write`,`*説明書*` → `ref`,其余文件名 → `write` 兜底)。Excel 图表规则(`図表規則.xlsx`)本轮**不纳入** `rule_docs`(Excel 规则解析留给后续迭代)。 + +### 3.4 SourceParser 门面(3.6 SourceAggregator) + +**输入:** +- `requirement_paths: list[str | Path]` — 要件定义 Excel 列表 +- `template_path: str | Path | None` — 概要设计模板 docx +- `rule_paths: list[str | Path]` — 规则文档列表(**仅 docx**) + +**类别映射(内置于门面,零 LLM):** `rule_paths` 中文件名含 `記入規則` → `write`;含 `説明書` → `ref`;其它 → `write` 兜底。 + +**路由:** + +| 扩展名/角色 | 解析器 | 产物 | +|------------|--------|------| +| `.xlsx` | 现有 `ExcelParser` | `tables` + `comments` | +| 模板 docx(template_path) | `WordTemplateParser` | `template: ParsedTemplate` | +| 规则 docx(rule_paths) | `RuleDocParser` | `rule_docs: list[RuleDocument]` | + +**输出:** `StructuredSource{tables, template, rule_docs, image_analyses=[], existing_system=None, comments}`(`image_analyses`/`existing_system` 本轮恒为空/None,类型字段保留)。 + +### 3.5 数据流 + +``` +samples/*.xlsx ─────────────► ExcelParser ─────► tables / comments +samples/テンプレート.docx ──► WordTemplateParser ─► template: ParsedTemplate ─┐ +samples/記入規則.docx ──────► RuleDocParser ────► rule_docs: [RuleDocument] ──┼─► StructuredSource + ┘ +``` + +### 3.6 错误处理 + +| 场景 | 行为 | +|------|------| +| 文件不存在 | 抛出 `FileNotFoundError`(消息含路径,便于定位) | +| 模板无 Heading / 空文档 | 返回空 `sections`/`placeholders`,不崩溃(与 ExcelParser 的 skipped 哲学一致) | +| 非法占位符格式(非 `{{section:xxx}}` / `{{doc_*}}`) | 不识别为占位符,正文文本原样保留 | +| 未知扩展名(如 `.txt`/`.md`) | `SourceParser` 路由阶段报错:`ValueError("不支持的文件类型: ...")` | +| docx 损坏 | 透出 python-docx 异常(PackageNotFoundError),测试不吞 | + +### 3.7 与现有代码的关系 + +- **不改** `excel_parser.py`、`data_models.py`、`pyproject.toml`、任何现有测试 +- 只新增 `src` 3 个生产模块 + 3 个测试文件 + 扩展 `test_real_samples.py` +- 提交消息前缀:`feat:`(生产实现)/ `test:`(门禁测试)/ `docs:`(本文档与 AI 日志) + +## 4. 文件定位与测试先例 + +- 单元测试沿用 `tests/test_real_samples.py:7` 的定位方式:`Path(__file__).resolve().parents[1] / "samples"`,样本缺失 `pytest.skip` +- 临时 docx 构造(单元测试):用 `python-docx` 在 `tmp_path` 生成内存文档(Heading/表格/占位符/书签),不依赖真实样本 +- 全量回归命令:`python -m pytest -q`(保持 ≥ 132 passed / 100.00%) + +## 5. 验收标准 + +1. `python -m pytest tests/test_word_template_parser.py tests/test_rule_doc_parser.py tests/test_source_aggregator.py tests/test_real_samples.py -v` 全绿 +2. 全量回归 `python -m pytest -q` ≥ 132 passed(新增用例后 ≥ 现有基线 + 新增数)/ 100.00% 覆盖 +3. 真实样本:模板 docx 解析出 7 个 H1 章(はじめに…バッチ一覧)+ 占位符;记入规则 docx 产出 category="write" 的 markdown +4. 门面全量组装:`StructuredSource.tables` 非空、`template` 非 None、`rule_docs` 含 write/ref 两类 +5. 手工破坏验证(可选):删除模板占位符段落 → 占位符断言变红;恢复正常 + +## 6. 影响面 + +- 新增生产文件 3 个:`src/genesis/parsers/word_template_parser.py`、`rule_doc_parser.py`、`source_aggregator.py` +- 新增测试文件 3 个:`tests/test_word_template_parser.py`、`tests/test_rule_doc_parser.py`、`tests/test_source_aggregator.py` +- 修改文件:`tests/test_real_samples.py`(追加 3 个 Word 样本用例)、`_AI_USAGE_LOG.md`(追加日志行) +- 提交纪律:分任务提交(feat/test/docs/chore 各自独立 commit) + +## 7. 不做的事(YAGNI 最终确认) + +- ❌ PPTXParser(3.3)— 无样本、无依赖,后续里程碑 +- ❌ 现有系统代码/设计书探索(3.4/3.5)— CodeParser 未实现 +- ❌ FileReader 统一读取层 — 现有 Excel 链不重构 +- ❌ 完整样式定义 / LLM 规则分类 / 模型改动 \ No newline at end of file