Files

10 KiB

LLM-OPTIMIZED REFERENCE -- code-review-graph v2.4.0

AI coding agents: Read ONLY the exact <section> you need. Never load the whole file.

Quick install: pip install code-review-graph Then: code-review-graph install && code-review-graph build First run: /code-review-graph:build-graph After that use only delta/pr commands. ALWAYS start with get_minimal_context_tool(task="your task") — returns ~100 tokens with risk, communities, flows, and suggested next tools. Use detail_level="minimal" on all subsequent calls unless you need more detail. When present, context_savings is an estimated compact hint, not exact tokenization.
1. Call get_minimal_context_tool(task="review changes") first. 2. If risk is low: detect_changes_tool(detail_level="minimal") → report summary. 3. If risk is medium/high: detect_changes_tool(detail_level="standard") → expand on high-risk items. Target: ≤5 tool calls, ≤800 tokens total context.
Fetch PR diff -> detect_changes_tool -> get_affected_flows_tool -> structured review with blast-radius table and risk scores. Never include full files unless explicitly asked.
Full three-layer review: 1) get_minimal_context_tool + build_or_update_graph_tool + get_review_context_tool + detect_changes_tool for graph context; 2) Layer-1 chain decomposition across 8 categories + gstack CRITICAL sub-pass; 3) score_review_tool for objective Layer-2 metrics; 4) specialist subagents (diff >= 50 lines); 5) dedupe_findings_tool to merge; 6) READ-ONLY manual adjudication per severity; 7) generate_report_tool to write code-review-report.html + code-review-report.md. Read .code-review.yaml for tier (fast/standard/strict). Target: <=8 tool calls, <=1200 tokens.
Whole-project or feature code review (not git-diff based). Parse scope: 全面/整个项目/所有/all -> whole-project; else feature + target keyword. review_data.scope MUST be exactly "whole-project" / "feature" / "change-level" — never a feature name like "evm" (the verify scripts rely on it).

WORKFLOW:

  1. get_minimal_context_tool + build_or_update_graph_tool
  2. get_architecture_overview_tool + list_communities_tool (module map)
  3. get_knowledge_gaps_tool + get_hub_nodes_tool + get_bridge_nodes_tool + find_large_functions_tool + get_surprising_connections_tool (whole-project hotspots)
  4. scoring: whole-project -> score_review_tool(all_files=True); feature -> semantic_search_nodes_tool(query=) + query_graph_tool(pattern="children_of", target=) to locate files, then score_review_tool(changed_files=) + get_impact_radius_tool(changed_files=)
  5. dedupe_findings_tool(findings=)
  6. READ-ONLY adjudication (present findings by severity; fix/skip per batch; 🔴 blockers cannot be batch-skipped)

COVERAGE SELF-CHECK (mandatory before report):

  • community_health_tool(): if needs_postprocess=true run code-review-graph postprocess first
  • coverage_tool REQUIRES the three-piece deep-read data: pass file_read_ranges={:[[s,e],...]} AND file_semantic_units={:[{"range":[s,e],"kind":..,"name":..},...]} (record these while deep-reading each file). Without them the line/unit coverage is FAIL-CLOSED to 0% (each file listed in line_gap_files/missing_data_files).
    • feature review: gate="line+unit" (line coverage >=95% + unit completeness gap-free ONLY; coverage_pct/high_risk_coverage_pct are null — no file-count/high-risk gate)
    • whole-project: gate="both+line" (overall >=85% AND high-risk >=95% AND line >=95% AND unit gap-free)
  • G1 confirm the deep-read list is complete; G2 spot-check 15% of silent_files (>=1 major found -> promote to full deep-read)
  • G3 REREAD LOOP (HARD REQUIREMENT): if target_reached=false (line coverage <95% or unit gaps exist) you MUST NOT generate the report yet. Re-deep-read the files listed in line_gap_files / unit_gap_files / missing_data_files (read the missing line ranges / semantic units), then re-run coverage_tool until target_reached=true. A file you cannot fully read must be REMOVED from deep_read_files (the line+unit gate does not count file numbers — only files actually read to >=95% belong there). After generating the report run verify-line-coverage.ps1; exit 1 means re-read and regenerate.

generate_report_tool(review_data=..., output_path="docs/reviews/{name}-review-{YYYY-MM-DD-HHMMSS}"):

  • MUST pass reviewed_files (an ARRAY of paths OR a comma-separated string — both are auto-normalised; report header renders a collapsible
    list). files (if any) is a comma-separated STRING, never a list/array.
  • TRANSMIT THE FULL coverage_tool RESULT VERBATIM into review_data.coverage (ALL fields: coverage_pct, high_risk_coverage_pct, grade, deep_read_count, total_files, high_risk_total_files, high_risk_deep_count, deep_read_weight, total_weight, target_reached, target, overall_target, high_risk_target, gate, line_coverage_pct, unit_coverage_pct, line_gap_files, unit_gap_files, unit_exempt_files, missing_data_files, remaining_files_to_target, remaining_weight_to_target, priority_deep_read_files, uncovered_files, silent_files, note) — do NOT hand-pick a subset, else counts render 0/0 and gap/missing hints disappear.
  • metrics contains ONLY the five objective keys (sql_risk, exception_coverage, redundancy_rate, high_risk_density, vulnerability_risk); each entry MUST carry note (copy from the score_review_tool return value), do NOT mix in blast_radius / objective_grade / llm_judged.
  • findings use message/fix fields (path + line synthesize location); counts use critical/informational keys.
  • format="both" (writes .html + .md). Full schema: project-review skill SKILL.md + references/report-schema.md.

Target: <=14 tool calls, <=2000 tokens.

score_review_tool returns objective metrics (sql_risk, exception_coverage, redundancy_rate, high_risk_density, vulnerability_risk) with good/warn/fail grades + llm_judged list. Pass all_files=True to score every source file (whole-project review). dedupe_findings_tool merges findings by path:line:category fingerprint, boosts multi-source confidence (+1 cap 10), computes PR quality score. generate_report_tool writes the HTML and/or Markdown review report (format=both by default).
Core MCP tools: get_minimal_context_tool, detect_changes_tool, get_review_context_tool, get_impact_radius_tool, query_graph_tool, semantic_search_nodes_tool, get_architecture_overview_tool, get_affected_flows_tool, list_flows_tool, list_communities_tool, refactor_tool, build_or_update_graph_tool, run_postprocess_tool, embed_graph_tool, list_graph_stats_tool, get_docs_section_tool Unified-review MCP tools: score_review_tool, dedupe_findings_tool, generate_report_tool, coverage_tool, community_health_tool MCP prompts (7): review_changes, architecture_map, debug_issue, onboard_developer, pre_merge_check, unified_review, project_review Skills: build-graph, debug-issue, explore-codebase, refactor-safely, review-changes, review-delta, review-pr, unified-review, project-review CLI: code-review-graph [install|init|build|update|status|watch|visualize|serve|mcp|wiki|detect-changes|postprocess|embed|register|unregister|repos|eval|daemon] Token efficiency: Prefer detail_level="minimal" where available. Always call get_minimal_context_tool first. Some review/context tools return compact estimated context_savings metadata.
MIT licence. Core graph/review workflows are local and there is no telemetry. DB file: .code-review-graph/graph.db. Optional cloud embeddings send embedded source snippets to the configured provider only when selected.
Run: code-review-graph watch (auto-updates graph on file save via watchdog) Or use PostToolUse (Write|Edit|Bash) hooks for automatic background updates.
Optional: pip install "code-review-graph[embeddings]" Then call embed_graph_tool to compute vectors. semantic_search_nodes_tool auto-uses vectors when available, falls back to keyword + FTS5. Providers: local sentence-transformers, OpenAI-compatible endpoints, Google Gemini, MiniMax, and Voyage. Configure via provider/model parameters, CRG_EMBEDDING_MODEL for local, CRG_OPENAI_* for OpenAI-compatible endpoints, or VOYAGE_API_KEY plus optional CRG_VOYAGE_MODEL for Voyage.
Supported: Python, JavaScript/TypeScript/TSX, Go, Rust, Java, C/C++, C#, VB.NET, Ruby, Kotlin, Swift, PHP, Scala, Solidity, Dart, R, Perl, Lua/Luau, Objective-C, shell scripts, Elixir, Zig, PowerShell, Julia, ReScript, GDScript, Nix, Verilog/SystemVerilog, SQL, Terraform/OpenTofu structure (`.tf`; generic `.hcl` files are recognized as file nodes), Ansible playbooks/roles/tasks, Vue/Svelte SFCs, Astro files parsed through the TypeScript parser, Jupyter/Databricks notebooks, and Perl XS files. Generic YAML is not treated as source code. Parser: Tree-sitter via tree-sitter-language-pack Custom languages: add .code-review-graph/languages.toml (extensions + node types per grammar) — no fork needed, see docs/CUSTOM_LANGUAGES.md. Built-ins cannot be overridden.
DB lock: SQLite WAL mode, auto-recovers. Only one build at a time. Large repos: first build ~40s at ~3,000 files; incremental ~2.5s on the hook path (measured, docs/REPRODUCING.md). Add patterns to .code-review-graphignore. Stale graph: Run /code-review-graph:build-graph manually. Missing nodes: Check language support + ignore patterns. Use full_rebuild=True. Windows/WSL: Use forward slashes in paths. Ensure uv is on PATH in WSL.

Instruction to AI coding agents (always follow): When user asks anything about "code-review-graph", "how to use", "commands", "review-delta", etc.:

  1. Call get_docs_section_tool with the exact section name.
  2. Use ONLY that content + current graph state.
  3. Never include full docs or source code in your reasoning. This keeps documentation lookup compact and avoids loading broad reference files by default.