docs: Phase3 Word 解析实施计划(5任务TDD)

This commit is contained in:
lhl
2026-08-10 08:50:00 +08:00
parent af3e756ba0
commit 129c75bb06
2 changed files with 831 additions and 0 deletions
+1
View File
@@ -61,3 +61,4 @@
| 2026-08-09 | 测试验证 | 一致性门禁最终评审收尾修复(Important x1 + Minor x3):① plan 文档 Interfaces 段 L311-314「五行 LLM 行」→「四行」(与 ==4 护栏一致);② 测试文件补尾随换行(末字节 " → \n);③ 断言消息改「异常类存在文档未登记或来源类名失配的 error_code」(覆盖双失败场景);④ _LLM_ROW_RE 加正则隐式限域注释。覆盖测试 4 passedpytest 全量 132 passed 覆盖 100.00%fail_under=99 达标 | docs/superpowers/plans/2026-08-09-error-code-consistency-test.md, tests/test_api_design_consistency.py, _AI_USAGE_LOG.md | deepseek-v4-flash-free |
| 2026-08-09 | 架构设计 | Phase3 Word 解析优先设计(brainstorming):范围澄清(Word 解析优先=3.1/3.2/3.6/3.7,PPT/现有系统探索后续;规则分类按来源映射零LLM;SourceAggregator 全量整合 StructuredSource;样式名级提取;方案A 三模块门面聚合);输出设计文档 docs/superpowers/specs/2026-08-09-phase3-word-parser-design.md | docs/superpowers/specs/2026-08-09-phase3-word-parser-design.md, _AI_USAGE_LOG.md | deepseek-v4-flash-free |
| 2026-08-09 | 架构设计 | Phase3 spec 工程完备性审阅(用户审阅关卡):实测 samples 3 个 docx(模板7H1+占位符+1书签、记入规则全 H1+・列表、说明书纯文本+H1);修正 P0-1 做成说明书分类 ref→write(对齐 api-design §2.2 write_instruction RAG 归类 Type A 写入规则);P1-1 列表检测双通道(・前缀+List Bullet 样式);P1-2 SourceParser 显式角色参数对齐 file_type 枚举;P1-3 fail_under=99 覆盖路径;统一占位符正则;新增 docx_helpers.py 测试基建 | docs/superpowers/specs/2026-08-09-phase3-word-parser-design.md, _AI_USAGE_LOG.md | deepseek-v4-flash-free |
| 2026-08-09 | Agent 实现 | Phase3 Word 解析优先实施计划(writing-plans):5 任务 TDDdocx_helpers 基建 / WordTemplateParser / RuleDocParser / SourceParser 门面 / 真实样本集成测试);实证验证 python-docx 关键点(body 级遍历 Paragraph/Table 构造、书签 XML、样式提取、List Bullet 可用);清理 Task2 占位写法;输出 docs/superpowers/plans/2026-08-09-phase3-word-parser.md | docs/superpowers/plans/2026-08-09-phase3-word-parser.md, _AI_USAGE_LOG.md | deepseek-v4-flash-free |
@@ -0,0 +1,830 @@
# Phase 3「Word 解析优先」实施计划
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** 实现 Word 模板解析(WordTemplateParser)、Word 规则文档 Markdown 化(RuleDocParser)与全量输入聚合(SourceParser 门面),配套单元测试与真实样本集成测试。
**Architecture:**`src/genesis/parsers/` 下新增 3 个扁平模块,复用 `data_models.py` 既有类型(ParsedTemplate/RuleDocument/StructuredSource),零模型改动、零新增依赖(python-docx 已装 1.2.0)。SourceParser 按 api-design `file_type` 语义显式接收角色参数(requirements/template/write_instruction/rules),按来源映射 category=write,组装 StructuredSource。
**Tech Stack:** Python 3.11+ / python-docx 1.2.0 / pytest 8 / openpyxl(现有 Excel 链路)
**Spec:** `docs/superpowers/specs/2026-08-09-phase3-word-parser-design.md`
## Global Constraints
- 覆盖率红线:`pyproject.toml` fail_under=99,新增模块需分支全覆盖(当前 132 passed / 100.00%
- 交流语言:所有注释、文档、断言消息、commit 消息**使用中文**
- 不修改:`excel_parser.py``data_models.py``pyproject.toml`、任何现有测试文件
- 零真实网络:不得引入 LLM 调用;规则分类按来源映射(零 LLM)
- 样本定位:`Path(__file__).resolve().parents[1] / "samples"`,样本缺失 `pytest.skip`
- 导入风格:`from tests.xxx import ...`tests 为包,先例 test_excel_parser.py:4
- 提交消息前缀:`feat:`(生产实现)/ `test:`(测试)/ `docs:`(文档)
- 提交纪律:每任务独立 commit,不混提交
---
### Task 1: tests/docx_helpers.py 测试基建
**Files:**
- Create: `tests/docx_helpers.py`
- Test: 由 Task 2 开始消费(本任务无可独立测试的交付物,直接作为后续 3 个任务的公共基建)
**Interfaces:**
- Produces:
- `new_document() -> docx.Document` — 新建内存 Document(标题/表格/列表等由调用方自由添加)
- `save_document(tmp_path, doc) -> str` — 落盘到 `tmp_path/source.docx` 并返回路径字符串
- `make_rule_doc(tmp_path, lines: list[tuple[str, str]]) -> str` — 快捷构造规则文档:每项 (样式名或空, 文本);`("H1", "1. 章")` → Heading 1、「"", "・リスト"」→ 普通段落(・ 前缀)。返回路径
**说明:** 与现有 `tests/excel_helpers.py`new_workbook/save_workbook)对称。python-docx 默认模板含 `List Bullet`/`Heading 1` 样式(已实证),`doc.add_paragraph(text, style=...)` 即可。
- [ ] **Step 1: 新建 docx_helpers.py**
```python
from pathlib import Path
from docx import Document
def new_document() -> Document:
"""新建内存 Word 文档(python-docx 默认模板)。"""
return Document()
def save_document(tmp_path, doc: Document) -> str:
"""落盘到 tmp_path 并返回路径字符串。"""
path = tmp_path / "source.docx"
doc.save(path)
return str(path)
def make_rule_doc(tmp_path, lines: list[tuple[str, str]]) -> str:
"""快捷构造规则文档。
lines 每项为 (样式标识, 文本):
("H1", "1. 章") / ("H2", "1.1 節") / ("H3", ...) → Heading 层级
("list", "・項目") → List Bullet 样式段落
("", "普通段落") → Normal 段落
返回落盘路径字符串。
"""
doc = new_document()
for kind, text in lines:
if kind == "H1":
doc.add_heading(text, level=1)
elif kind == "H2":
doc.add_heading(text, level=2)
elif kind == "H3":
doc.add_heading(text, level=3)
elif kind == "list":
doc.add_paragraph(text, style="List Bullet")
else:
doc.add_paragraph(text)
return save_document(tmp_path, doc)
```
- [ ] **Step 2: 验证可导入**
Run: `python -c "from tests.docx_helpers import new_document, save_document, make_rule_doc; print('ok')"`
Expected: `ok`(工作目录为项目根 `D:\00_project\Genesis`
- [ ] **Step 3: 提交**
```bash
git add tests/docx_helpers.py
git commit -m "chore: docx 测试基建(新建/落盘/规则文档快捷构造)"
```
---
### Task 2: WordTemplateParser 实现
**Files:**
- Create: `src/genesis/parsers/word_template_parser.py`
- Test: `tests/test_word_template_parser.py`
**Interfaces:**
- Consumes: `genesis.data_models.ChapterMarker` / `genesis.data_models.ParsedTemplate`
- Produces:
- `class WordTemplateParser`,方法 `parse(path: str | Path) -> ParsedTemplate`
- 内部正则 `PLACEHOLDER_RE``r"\{\{([a-z][a-z0-9_]*)(?::([a-z][a-z0-9_]*))?\}\}"`(两组:group1=键名,group2=可选章节名)
- 行为:Heading 段落 → `ChapterMarker(type="heading", name=段落文本, level=大纲级别 1/2/3)``bookmarkStart``type="bookmark"`;匹配占位符 → `type="placeholder"`name=`键名`(封面型)或 `键名:章节名`section 型)),同时写入 `placeholders` 字典 `{占位符名: 段落上下文文本}`
- `styles``{"defined": [文档命名样式名...], "used": [各段落实际样式名去重]}`(样式名级,不提取字号/颜色)
- 空文档/无 Heading → 空 `sections`/`placeholders`,不崩溃;无法打开 → 透出 python-docx 异常
- [ ] **Step 1: 写失败测试**
Create `tests/test_word_template_parser.py`:
```python
from genesis.data_models import ChapterMarker, ParsedTemplate
from genesis.parsers.word_template_parser import WordTemplateParser
from tests.docx_helpers import new_document, save_document
def test_parse_extracts_heading_levels(tmp_path):
doc = new_document()
doc.add_heading("1. はじめに", level=1)
doc.add_heading("2.1 画面遷移図", level=2)
doc.add_heading("2.1.1 詳細", level=3)
path = save_document(tmp_path, doc)
result = WordTemplateParser().parse(path)
assert isinstance(result, ParsedTemplate)
headings = [s for s in result.sections if s.type == "heading"]
assert [(s.name, s.level) for s in headings] == [
("1. はじめに", 1),
("2.1 画面遷移図", 2),
("2.1.1 詳細", 3),
]
def test_parse_extracts_bookmark(tmp_path):
from docx.oxml.ns import qn
doc = new_document()
para = doc.add_paragraph("アンカー")
bm_start = para._p.makeelement(qn("w:bookmarkStart"), {qn("w:id"): "0", qn("w:name"): "template_start"})
para._p.insert(0, bm_start)
path = save_document(tmp_path, doc)
result = WordTemplateParser().parse(path)
bookmarks = [s for s in result.sections if s.type == "bookmark"]
assert [s.name for s in bookmarks] == ["template_start"]
def test_parse_extracts_placeholders(tmp_path):
doc = new_document()
doc.add_paragraph("{{doc_title}}")
doc.add_paragraph("{{section:introduction}}")
doc.add_paragraph("{{section:function_list}}")
path = save_document(tmp_path, doc)
result = WordTemplateParser().parse(path)
assert result.placeholders == {
"doc_title": "{{doc_title}}",
"section:introduction": "{{section:introduction}}",
"section:function_list": "{{section:function_list}}",
}
ph = [s for s in result.sections if s.type == "placeholder"]
assert [s.name for s in ph] == ["doc_title", "section:introduction", "section:function_list"]
def test_parse_invalid_placeholder_kept_as_text(tmp_path):
doc = new_document()
doc.add_paragraph("{{ invalid }}")
doc.add_paragraph("ただの {text}")
path = save_document(tmp_path, doc)
result = WordTemplateParser().parse(path)
assert result.placeholders == {}
assert [s for s in result.sections if s.type == "placeholder"] == []
def test_parse_empty_document(tmp_path):
doc = new_document()
path = save_document(tmp_path, doc)
result = WordTemplateParser().parse(path)
assert result.sections == []
assert result.placeholders == {}
assert "Normal" in result.styles["defined"]
def test_parse_styles_collected(tmp_path):
doc = new_document()
doc.add_heading("", level=1)
doc.add_paragraph("本文")
path = save_document(tmp_path, doc)
result = WordTemplateParser().parse(path)
assert "Heading 1" in result.styles["used"]
assert "Normal" in result.styles["used"]
assert "Heading 1" in result.styles["defined"]
```
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_word_template_parser.py -v`
Expected: FAIL / ERROR`ModuleNotFoundError: No module named 'genesis.parsers.word_template_parser'`
- [ ] **Step 3: 写最小实现**
Create `src/genesis/parsers/word_template_parser.py`:
```python
from __future__ import annotations
import re
from pathlib import Path
from docx import Document
from docx.oxml.ns import qn
from genesis.data_models import ChapterMarker, ParsedTemplate
# 统一占位符正则:{{键名}} 或 {{键名:章节名}}spec §3.2
PLACEHOLDER_RE = re.compile(r"\{\{([a-z][a-z0-9_]*)(?::([a-z][a-z0-9_]*))?\}\}")
def _heading_level(style_name: str) -> int:
"""从 Heading N 样式名解析大纲级别;非数字兜底 1。"""
try:
return int(style_name.split()[-1])
except (ValueError, IndexError):
return 1
class WordTemplateParser:
"""概要设计模板 docx 解析:章构成 / 占位符 / 样式名提取。"""
def parse(self, path: str | Path) -> ParsedTemplate:
doc = Document(str(path))
sections: list[ChapterMarker] = []
placeholders: dict[str, str] = {}
used_styles: set[str] = set()
# 文档命名样式(定义集合)
defined = {s.name for s in doc.styles if s.name}
for para in doc.paragraphs:
style_name = para.style.name if para.style else "Normal"
used_styles.add(style_name)
text = para.text
if style_name.startswith("Heading"):
sections.append(ChapterMarker(
type="heading", name=text, level=_heading_level(style_name)
))
for m in PLACEHOLDER_RE.finditer(text):
if m.group(2):
key = f"{m.group(1)}:{m.group(2)}"
else:
key = m.group(1)
placeholders[key] = text
sections.append(ChapterMarker(type="placeholder", name=key, level=0))
# 书签:遍历 body 中全部 bookmarkStart
for bm in doc.element.body.iter(qn("w:bookmarkStart")):
name = bm.get(qn("w:name"))
if name:
sections.append(ChapterMarker(type="bookmark", name=name, level=0))
return ParsedTemplate(
file_name=Path(path).name,
sections=sections,
placeholders=placeholders,
styles={"defined": sorted(defined), "used": sorted(used_styles)},
)
```
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_word_template_parser.py -v`
Expected: 6 passed
- [ ] **Step 5: 全量回归**
Run: `python -m pytest -q`
Expected: 138 passed / 100.00%132 基线 + 6 新增),fail_under=99 达标
- [ ] **Step 6: 提交**
```bash
git add tests/test_word_template_parser.py src/genesis/parsers/word_template_parser.py
git commit -m "feat: WordTemplateParser 章构成/占位符/样式名提取"
```
---
### Task 3: RuleDocParser 实现
**Files:**
- Create: `src/genesis/parsers/rule_doc_parser.py`
- Test: `tests/test_rule_doc_parser.py`
**Interfaces:**
- Consumes: `genesis.data_models.RuleDocument``tests/docx_helpers.make_rule_doc`
- Produces:
- `class RuleDocParser`,方法 `parse(path: str | Path, category: str = "write") -> RuleDocument`
- Markdown 化:Heading N → `#`×N`・`/`-`/`•` 前缀文本或 List Bullet/Number 样式 → `- ` 项;普通段落 → 原文;表格 → GFM(表头+分隔行+数据行);空段 → 空行
- `file_type` 固定 `"word"``hash` = 文件内容 sha256 hex
- [ ] **Step 1: 写失败测试**
Create `tests/test_rule_doc_parser.py`:
```python
from genesis.data_models import RuleDocument
from genesis.parsers.rule_doc_parser import RuleDocParser
from tests.docx_helpers import make_rule_doc, new_document, save_document
def test_parse_heading_and_bullet(tmp_path):
path = make_rule_doc(tmp_path, [
("H1", "1. 機能一覧の書き方"),
("list", "・機能ID は F001 から連番で付与する。"),
("list", "・省略記号(~、…)は禁止し、正式名称を記載する。"),
])
result = RuleDocParser().parse(path, category="write")
assert isinstance(result, RuleDocument)
assert result.category == "write"
assert result.file_type == "word"
assert result.file_name == "source.docx"
assert "# 1. 機能一覧の書き方" in result.markdown_content
assert "- 機能ID は F001 から連番で付与する。" in result.markdown_content
assert "- 省略記号(~、…)は禁止し、正式名称を記載する。" in result.markdown_content
def test_parse_prefix_bullet_without_list_style(tmp_path):
# 真实样本的「・」前缀在 Normal 样式段落(非 List Bullet)——双通道检测
doc = new_document()
doc.add_paragraph("・表ヘッダーは太字とし、下線を付ける。")
path = save_document(tmp_path, doc)
result = RuleDocParser().parse(path)
assert "- 表ヘッダーは太字とし、下線を付ける。" in result.markdown_content
def test_parse_table_to_markdown(tmp_path):
doc = new_document()
table = doc.add_table(rows=3, cols=2)
headers = ["項目", "規則"]
for c, h in enumerate(headers):
table.cell(0, c).text = h
table.cell(1, 0).text = "表ヘッダー"
table.cell(1, 1).text = "太字のみ"
table.cell(2, 0).text = "枠線"
table.cell(2, 1).text = "付ける"
path = save_document(tmp_path, doc)
result = RuleDocParser().parse(path)
lines = result.markdown_content.splitlines()
assert "| 項目 | 規則 |" in lines
assert "| --- | --- |" in lines
assert "| 表ヘッダー | 太字のみ |" in lines
assert "| 枠線 | 付ける |" in lines
def test_parse_empty_document(tmp_path):
doc = new_document()
path = save_document(tmp_path, doc)
result = RuleDocParser().parse(path)
assert result.markdown_content == ""
def test_parse_hash_stable_and_sha256(tmp_path):
path = make_rule_doc(tmp_path, [("H1", ""), ("", "本文")])
r1 = RuleDocParser().parse(path)
r2 = RuleDocParser().parse(path)
assert r1.hash == r2.hash
assert len(r1.hash) == 64 # sha256 hex
def test_parse_design_category(tmp_path):
path = make_rule_doc(tmp_path, [("H1", "アーキテクチャ制約")])
result = RuleDocParser().parse(path, category="design")
assert result.category == "design"
```
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_rule_doc_parser.py -v`
Expected: FAIL / ERROR`ModuleNotFoundError: No module named 'genesis.parsers.rule_doc_parser'`
- [ ] **Step 3: 写最小实现**
Create `src/genesis/parsers/rule_doc_parser.py`:
```python
from __future__ import annotations
import hashlib
from pathlib import Path
from docx import Document
from docx.oxml.ns import qn
from genesis.data_models import RuleDocument
# 列表项前缀(真实样本为 Normal 样式 + ・ 前缀,样式名通道不足以命中)
LIST_PREFIXES = ("", "-", "")
def _heading_level(style_name: str) -> int:
try:
return int(style_name.split()[-1])
except (ValueError, IndexError):
return 1
def _is_list_item(text: str, style_name: str) -> bool:
"""双通道列表检测:List 样式或文本前缀(spec §3.3)。"""
if "List" in style_name:
return True
return text.startswith(LIST_PREFIXES)
def _table_to_markdown(table) -> list[str]:
"""docx 表格 → GFM 表格行。"""
lines: list[str] = []
for r_idx, row in enumerate(table.rows):
cells = [cell.text.replace("|", "\\|").strip() for cell in row.cells]
lines.append("| " + " | ".join(cells) + " |")
if r_idx == 0:
lines.append("| " + " | ".join(["---"] * len(cells)) + " |")
return lines
class RuleDocParser:
"""规则文档 docx 解析:Markdown 化 + 分类。"""
def parse(self, path: str | Path, category: str = "write") -> RuleDocument:
doc = Document(str(path))
md: list[str] = []
# 按文档顺序遍历段落与表格(body 级子元素)
for child in doc.element.body.iterchildren():
tag = child.tag
if tag == qn("w:p"):
from docx.text.paragraph import Paragraph
para = Paragraph(child, doc)
text = para.text
style_name = para.style.name if para.style else "Normal"
stripped = text.strip()
if not stripped:
md.append("")
elif style_name.startswith("Heading"):
md.append("#" * _heading_level(style_name) + " " + stripped)
elif _is_list_item(stripped, style_name):
md.append("- " + stripped.lstrip("・-•").strip())
else:
md.append(stripped)
elif tag == qn("w:tbl"):
from docx.table import Table
md.extend(_table_to_markdown(Table(child, doc)))
md.append("")
content = "\n".join(md).strip()
return RuleDocument(
file_name=Path(path).name,
category=category,
markdown_content=content,
source_path=str(path),
file_type="word",
hash=hashlib.sha256(Path(path).read_bytes()).hexdigest(),
)
```
> 说明:`Paragraph(child, doc)` / `Table(child, doc)` 的第二个参数(parent)只需提供 `.part` 属性;`Document` 实例满足,实测可用。若实现中直接从 `doc.paragraphs` + `doc.tables` 读取,将丢失**文档内交错顺序**(段落/表格混排),故用 body 级遍历保证顺序正确。
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_rule_doc_parser.py -v`
Expected: 6 passed
- [ ] **Step 5: 全量回归**
Run: `python -m pytest -q`
Expected: 144 passed / 100.00%138 + 6 新增),fail_under=99 达标
- [ ] **Step 6: 提交**
```bash
git add tests/test_rule_doc_parser.py src/genesis/parsers/rule_doc_parser.py
git commit -m "feat: RuleDocParser 规则文档 Markdown 化与分类"
```
---
### Task 4: SourceParser 门面实现
**Files:**
- Create: `src/genesis/parsers/source_aggregator.py`
- Test: `tests/test_source_aggregator.py`
**Interfaces:**
- Consumes: `genesis.parsers.excel_parser.ExcelParser`(现有)、`genesis.parsers.word_template_parser.WordTemplateParser`Task 2)、`genesis.parsers.rule_doc_parser.RuleDocParser`Task 3);`genesis.data_models.StructuredSource`
- Produces:
- `class SourceParser`,方法 `parse(requirement_paths: list[str | Path] | None = None, template_path: str | Path | None = None, write_instruction_paths: list[str | Path] | None = None, rule_paths: list[str | Path] | None = None) -> StructuredSource`
- 角色→解析器:`.xlsx`requirements)→ ExcelParser → tables/comments`template_path` → WordTemplateParser`write_instruction_paths`/`rule_paths`(仅 .docx)→ RuleDocParser(category="write")
- 错误:文件不存在 → `FileNotFoundError`;未知扩展名 → `ValueError("不支持的文件类型: ...")`
- `image_analyses=[]``existing_system=None` 固定(本轮无来源)
- [ ] **Step 1: 写失败测试**
Create `tests/test_source_aggregator.py`:
```python
import pytest
from genesis.data_models import StructuredSource
from genesis.parsers.source_aggregator import SourceParser
from tests.docx_helpers import make_rule_doc, new_document, save_document
from tests.excel_helpers import new_workbook, save_workbook
def _xlsx(tmp_path, name: str = "source.xlsx") -> str:
wb = new_workbook({"機能一覧": [["機能ID", "機能名"], ["A001", "社員登録"]]})
path = tmp_path / name
wb.save(path)
return str(path)
def test_parse_full_assembly(tmp_path):
xlsx = _xlsx(tmp_path)
template = save_document(tmp_path, new_document())
rule = make_rule_doc(tmp_path, [("H1", "1. 機能一覧の書き方")])
instr = make_rule_doc(tmp_path, [("H1", "2. 機能一覧の作成手順")])
result = SourceParser().parse(
requirement_paths=[xlsx],
template_path=template,
write_instruction_paths=[instr],
rule_paths=[rule],
)
assert isinstance(result, StructuredSource)
assert len(result.tables) == 1
assert result.template is not None
assert result.template.file_name == "source.docx"
assert len(result.rule_docs) == 2
assert all(r.category == "write" for r in result.rule_docs)
assert result.image_analyses == []
assert result.existing_system is None
def test_parse_without_template_and_rules(tmp_path):
xlsx = _xlsx(tmp_path)
result = SourceParser().parse(requirement_paths=[xlsx])
assert len(result.tables) == 1
assert result.template is None
assert result.rule_docs == []
def test_parse_missing_requirement_file(tmp_path):
with pytest.raises(FileNotFoundError):
SourceParser().parse(requirement_paths=[tmp_path / "missing.xlsx"])
def test_parse_missing_template_file(tmp_path):
xlsx = _xlsx(tmp_path)
with pytest.raises(FileNotFoundError):
SourceParser().parse(requirement_paths=[xlsx], template_path=tmp_path / "missing.docx")
def test_parse_unknown_extension_in_requirements(tmp_path):
bad = tmp_path / "note.txt"
bad.write_text("hello", encoding="utf-8")
with pytest.raises(ValueError, match="不支持的文件类型"):
SourceParser().parse(requirement_paths=[bad])
def test_parse_unknown_extension_in_rules(tmp_path):
bad = tmp_path / "note.txt"
bad.write_text("hello", encoding="utf-8")
with pytest.raises(ValueError, match="不支持的文件类型"):
SourceParser().parse(rule_paths=[bad])
```
- [ ] **Step 2: 运行测试确认失败**
Run: `python -m pytest tests/test_source_aggregator.py -v`
Expected: FAIL / ERROR`ModuleNotFoundError: No module named 'genesis.parsers.source_aggregator'`
- [ ] **Step 3: 写最小实现**
Create `src/genesis/parsers/source_aggregator.py`:
```python
from __future__ import annotations
from pathlib import Path
from genesis.data_models import StructuredSource
from genesis.parsers.excel_parser import ExcelParser
from genesis.parsers.rule_doc_parser import RuleDocParser
from genesis.parsers.word_template_parser import WordTemplateParser
XLSX_EXTS = (".xlsx", ".xls")
DOCX_EXT = ".docx"
class SourceParser:
"""全量输入门面:Excel 要件定义 + Word 模板 + Word 规则 → StructuredSource。
角色由调用方按 api-design file_type 语义显式传入(requirements/template/
write_instruction/rules),不做基于文件名的隐式猜测(spec §3.4)。
"""
def __init__(self) -> None:
self._excel = ExcelParser()
def parse(
self,
requirement_paths: list[str | Path] | None = None,
template_path: str | Path | None = None,
write_instruction_paths: list[str | Path] | None = None,
rule_paths: list[str | Path] | None = None,
) -> StructuredSource:
requirement_paths = requirement_paths or []
write_instruction_paths = write_instruction_paths or []
rule_paths = rule_paths or []
tables = []
comments = []
for p in requirement_paths:
path = Path(p)
if path.suffix.lower() not in XLSX_EXTS:
raise ValueError(f"不支持的文件类型: {path.suffix or '(无扩展名)'}")
if not path.exists():
raise FileNotFoundError(str(p))
result = self._excel.parse(path)
tables.extend(result.tables)
comments.extend(result.comments)
template = None
if template_path is not None:
tpath = Path(template_path)
if tpath.suffix.lower() != DOCX_EXT:
raise ValueError(f"不支持的文件类型: {tpath.suffix or '(无扩展名)'}")
if not tpath.exists():
raise FileNotFoundError(str(template_path))
template = WordTemplateParser().parse(tpath)
rule_docs = []
for p in [*write_instruction_paths, *rule_paths]:
path = Path(p)
if path.suffix.lower() != DOCX_EXT:
raise ValueError(f"不支持的文件类型: {path.suffix or '(无扩展名)'}")
if not path.exists():
raise FileNotFoundError(str(p))
# 做成说明书与记入规则均为 Type A 写入规则 → writeapi-design §2.2
rule_docs.append(RuleDocParser().parse(path, category="write"))
return StructuredSource(
tables=tables,
template=template,
rule_docs=rule_docs,
image_analyses=[],
existing_system=None,
comments=comments,
)
```
- [ ] **Step 4: 运行测试确认通过**
Run: `python -m pytest tests/test_source_aggregator.py -v`
Expected: 6 passed
- [ ] **Step 5: 全量回归**
Run: `python -m pytest -q`
Expected: 150 passed / 100.00%144 + 6 新增),fail_under=99 达标
- [ ] **Step 6: 提交**
```bash
git add tests/test_source_aggregator.py src/genesis/parsers/source_aggregator.py
git commit -m "feat: SourceParser 门面全量输入聚合 StructuredSource"
```
---
### Task 5: 真实样本集成测试(test_real_samples.py 扩展)
**Files:**
- Modify: `tests/test_real_samples.py`(追加 3 个 Word 样本用例 + 1 个全量组装用例;不删改现有 Excel 用例)
**Interfaces:**
- Consumes: `WordTemplateParser`Task 2)、`RuleDocParser`Task 3)、`SourceParser`Task 4);`SAMPLES = Path(__file__).resolve().parents[1] / "samples"`(现有)
- Produces: 真实样本验证——模板 7 个 H1 + 占位符 + 1 书签;记入规则 category=write markdown;说明书画属 write;全量组装双向字段
- [ ] **Step 1: 追加测试用例**
Append to `tests/test_real_samples.py`(保留现有 4 个 Excel 用例):
```python
from genesis.parsers.word_template_parser import WordTemplateParser
from genesis.parsers.rule_doc_parser import RuleDocParser
from genesis.parsers.source_aggregator import SourceParser
def _d(name: str) -> Path:
return SAMPLES / name
def test_word_template_sample_chapters_and_placeholders():
p = _d("概要設計書テンプレート.docx")
if not p.exists():
pytest.skip("样本缺失")
result = WordTemplateParser().parse(p)
headings = [s for s in result.sections if s.type == "heading"]
h1 = [h for h in headings if h.level == 1]
assert len(h1) == 7
assert h1[0].name == "1. はじめに"
assert h1[-1].name == "7. バッチ一覧"
assert "section:introduction" in result.placeholders
assert "doc_title" in result.placeholders
bookmarks = [s for s in result.sections if s.type == "bookmark"]
assert len(bookmarks) == 1
assert bookmarks[0].name == "template_start"
def test_rule_doc_sample_markdown_and_category():
p = _d("記入規則.docx")
if not p.exists():
pytest.skip("样本缺失")
result = RuleDocParser().parse(p, category="write")
assert result.category == "write"
assert result.file_type == "word"
assert "# 1. 機能一覧の書き方" in result.markdown_content
assert "- 機能ID は F001 から連番で付与する。" in result.markdown_content
def test_write_instruction_sample_category_write():
p = _d("概要設計做成説明書.docx")
if not p.exists():
pytest.skip("样本缺失")
result = RuleDocParser().parse(p, category="write")
assert result.category == "write"
assert "# 2. 機能一覧" in result.markdown_content
def test_source_parser_full_sample_assembly():
xlsx = _x("要件定義_新規開発.xlsx")
template = _d("概要設計書テンプレート.docx")
rule = _d("記入規則.docx")
instr = _d("概要設計做成説明書.docx")
if not all(p.exists() for p in [xlsx, template, rule, instr]):
pytest.skip("样本缺失")
result = SourceParser().parse(
requirement_paths=[xlsx],
template_path=template,
write_instruction_paths=[instr],
rule_paths=[rule],
)
assert result.tables
assert result.template is not None
assert len(result.rule_docs) == 2
assert all(r.category == "write" for r in result.rule_docs)
```
- [ ] **Step 2: 运行测试确认通过**
Run: `python -m pytest tests/test_real_samples.py -v`
Expected: 8 passed(现有 4 Excel + 新增 4 Word
- [ ] **Step 3: 全量回归(红线验证)**
Run: `python -m pytest -q`
Expected: 154 passed / 100.00%150 + 4 新增),fail_under=99 达标
- [ ] **Step 4: 提交**
```bash
git add tests/test_real_samples.py
git commit -m "test: 真实样本 Word 解析集成测试(模板/规则/说明书画属write)"
```
---
## Self-Review 对照
**Spec 覆盖:**
- §3.2 WordTemplateParser(章构成/占位符/书签/样式名)→ Task 2
- §3.3 RuleDocParserMarkdown 化 / 双通道列表 / hash / category)→ Task 3
- §3.4 SourceParser 门面(显式角色参数 / 扩展名路由 / 错误处理)→ Task 4
- §3.5 数据流(Excel→tables、模板→template、规则→rule_docs)→ Task 4 全量组装测试 + Task 5 真实样本组装
- §5 验收标准 1/2/3/4 → Task 2-5;验收标准 5(手工破坏)→ 可选不落自动化
- spec §4 测试基建 docx_helpers → Task 1
**占位符扫描:** 全部步骤含完整代码与精确命令,无 TBD/「后续处理」/「类似上文」等占位。
**类型一致性:**
- `WordTemplateParser.parse(path) -> ParsedTemplate` 在 Task 2 定义,Task 4/5 引用同名同型
- `RuleDocParser.parse(path, category="write") -> RuleDocument` 在 Task 3 定义,Task 4 传 `category="write"`、Task 5 传显式
- `SourceParser.parse(requirement_paths, template_path, write_instruction_paths, rule_paths) -> StructuredSource` Task 4 定义,Task 5 引用
- `docx_helpers.make_rule_doc(tmp_path, lines)` 在 Task 1 定义,Task 3/4/5 引用
- `ExcelParser().parse()` 返回 `.tables/.comments`(现有,test_excel_parser.py 先例)
**预期提交数:** 5 个(Task 1-5 各 1