feat: add unified-review workflow (scoring tools + skill)

Adds the unified-review integration that fuses CRG graph context with the
ai-code-review scoring methodology and gstack-review fix-first workflow:

- scoring.py: objective Layer-2 metrics (sql_risk, exception_coverage,
  redundancy_rate, high_risk_density, vulnerability_risk) with
  good/warn/fail grades, plus dedupe_findings (fingerprint merge,
  multi-source confidence boost, PR quality score) and report data builder
- tools/scoring_tools.py + main.py: three new MCP tools
  (score_review_tool, dedupe_findings_tool, generate_report_tool)
- assets/report-template.html: self-contained HTML report template
- skills.py + skills/unified-review/: new read-only unified-review skill
  with language/manual-review/specialist checklists
- docs and CHANGELOG updated; tests added (test_scoring, test_report,
  test_unified_review) and test_skills updated for 5 skills
This commit is contained in:
dev
2026-08-05 13:31:55 +08:00
parent 82b7c6dc9e
commit 84ae9b817e
32 changed files with 2229 additions and 19 deletions
+42
View File
@@ -21,6 +21,13 @@ Review a PR or branch diff.
- Full impact analysis across all PR commits
- Structured output with risk assessment
### `/code-review-graph:unified-review`
Three-layer unified code review (CRG graph context + ai-code-review scoring + gstack-review workflow).
- Read-only: every finding waits for a manual fix decision
- Objective Layer-2 metrics via `score_review_tool`
- Fingerprint merge + quality score via `dedupe_findings_tool`
- Standalone HTML report via `generate_report_tool`
## MCP Tools
### Core Tools
@@ -223,6 +230,41 @@ detail_level: str = "standard"
Primary tool for code review. Maps changed files to affected functions, flows, communities, and test coverage gaps. Returns risk scores and prioritized review items.
Relevant responses may include compact estimated `context_savings` metadata.
### Unified Review Tools
#### `score_review_tool`
```
changed_files: list[str] | None # Auto-detected from git diff if omitted
base: str = "HEAD~1"
include_churn: bool = True
repo_root: str | None
detail_level: str = "standard" # "minimal" for grades + values only
```
Computes objective Layer-2 metrics for changed files: `sql_risk`,
`exception_coverage`, `redundancy_rate`, `high_risk_density`,
`vulnerability_risk`. Each metric carries a `good`/`warn`/`fail` grade,
thresholds and evidence. LLM-judged metrics are listed in `llm_judged`.
#### `dedupe_findings_tool`
```
findings: list # [{path, category, severity, confidence, source?}]
suppress_prior: list | None # Previously user-skipped findings to suppress
repo_root: str | None
```
Merges findings by `path:line:category` fingerprint (highest confidence
wins), boosts multi-source confidence (+1, cap 10), routes low-confidence
findings to the appendix, and computes `PR quality score = max(0, 10 -
(critical*2 + informational*0.5))`.
#### `generate_report_tool`
```
review_data: dict # metrics + findings + verdict + tier + scope
output_path: str | None # Default: <repo_root>/code-review-report.html
repo_root: str | None
```
Renders the standalone HTML code review report (self-contained, no external
dependencies) from the bundled `report-template.html`.
#### `refactor_tool`
```
mode: str = "rename" # "rename", "dead_code", or "suggest"
+17 -10
View File
@@ -27,6 +27,8 @@
│ │ ├── Communities: list, get, architecture │ │
│ │ ├── Analysis: detect_changes, refactor, │ │
│ │ │ apply_refactor, hotspots, gaps │ │
│ │ ├── Scoring: score_review, dedupe_ │ │
│ │ │ findings, generate_report │ │
│ │ ├── Wiki: generate, get_page │ │
│ │ └── Multi-repo: list_repos, cross_search │ │
│ └────────────────┬───────────────────────────┘ │
@@ -34,16 +36,21 @@
┌───────────┼───────────────┐
▼ ▼ ▼
┌─────────┐ ┌─────────┐ ┌─────────────┐
│ Parser │ │ Graph │ │ Incremental │
│ │ │ Store │ │ Engine │
└────┬────┘ └────┬────┘ └──────┬──────┘
│ │ │
▼ ▼ ▼
Tree-sitter SQLite DB git/svn diff
grammars (.code-review- subprocess
graph/
graph.db)
┌─────────┐ ┌─────────┐ ┌─────────────┐
│ Parser │ │ Graph │ │ Incremental │
│ │ │ Store │ │ Engine │
└────┬────┘ └────┬────┘ └──────┬──────┘
│ │ │
▼ ▼ ▼
Tree-sitter SQLite DB git/svn diff
grammars (.code-review- subprocess
graph/
graph.db)
Scoring module (code_review_graph/scoring.py) sits on top of the
GraphStore and reads changed files + git history to compute the
objective review metrics consumed by the unified-review skill and the
score_review / dedupe_findings / generate_report MCP tools.
```
## Data Flow