### Design/Pr Triage Workflows # PR Triage Workflows Three GitHub Actions workflows keep open PRs moving without manual nudging: - **`pr-triage-batch.yml`** — hourly orchestrator (cron `17 * * * *`). Enumerates open non-draft PRs, computes a deterministic state for each, and dispatches the per-PR worker (or the malicious-code scanner). No labels, no model calls; the only comment it posts is a one-time idempotency marker when it dispatches the malicious-code scanner. Also hosts the deterministic weekly **stale-PR sweep** (`stale-sweep` job, cron `17 4 * * 1`) — see [Stale-PR sweep](#stale-pr-sweep). - **`pr-triage.yml`** — per-PR worker (`workflow_dispatch`). Re-validates the PR's state, reconciles a single `pr-state/*` label, and performs at most one of: trigger evaluation (by dispatching `evaluation.yml`), ping the author, or ping maintainers. Cool-down (default 4 days) is enforced via marker comments. - **`pr-malicious-scan.agent.md`** — per-PR malicious-code scanner (gh-aw). Static diff review for untrusted contributors. Reports findings as code-scanning alerts and an optional comment; never executes PR head code. ## Architecture ```mermaid flowchart TD Cron["cron: every hour"] --> Batch["pr-triage-batch.yml
(orchestrator)"] WCron["cron: weekly (Mon)"] --> Sweep["pr-triage-batch.yml
(stale-sweep job)"] Batch -->|workflow_dispatch| Worker["pr-triage.yml
(per-PR worker)"] Batch -->|workflow_dispatch| Scan["pr-malicious-scan.agent.lock.yml
(per-PR scanner)"] Worker -->|workflow_dispatch: pr_number| Eval["evaluation.yml
(existing)"] Worker -->|adds pr-state/* label| PR[("PR")] Worker -->|posts ping comment| PR Scan -->|code-scanning alert + comment| PR Sweep -->|warn / close stale| PR PR -.->|human adds label: evaluate-now| Eval ``` ## Entry points into `evaluation.yml` Four entry points feed the `gate` job, all sharing a per-PR concurrency group so overlapping triggers collapse to a single run. Each binds the run to **one specific reviewed commit** (never the live branch head), so evaluation always runs the exact commit the maintainer approved: 1. The **`/evaluate `** slash command (`issue_comment`) — humans. The conversation-comment payload carries no commit id, so an explicit SHA is **required** and must belong to the PR; a bare `/evaluate` only posts guidance pointing to the review flow. 2. **`/evaluate`** inside a submitted PR review (`pull_request_review [submitted]`) — humans; the recommended path. Bound to `review.commit_id` (the exact commit reviewed), so no SHA needs to be typed. 3. The **`evaluate-now`** label (`pull_request_target [labeled]`) — humans. The `gate` job consumes (removes) the label so reapplying re-fires. Bound to the head SHA carried in the label event payload. 4. **`workflow_dispatch`** with a `pr_number` input — the triage worker. The worker runs as `github-actions[bot]`, and label events emitted by `GITHUB_TOKEN` do **not** start workflows (GitHub's recursion guard), so the bot cannot use entry point 3. `workflow_dispatch` is exempt from that guard, so the worker dispatches `evaluation.yml` directly. A dispatched run checks out the default branch by default (`github.sha` is `main`'s tip, **not** the PR head) and its metadata doesn't record the target PR, so the worker matches the run by `evaluation.yml`'s run name (`Evaluate PR # @ `) for idempotency. The PR's head travels in the `head_sha` **input**, and the gate resolves that short SHA to the exact commit (it does not re-read the live PR head). ## State machine (worker) Order of evaluation; first match wins: | Order | Condition | State | Label | Action | |---|---|---|---|---| | 1 | draft, or `mergeable_state == unknown` | `skip` | — | none | | 2 | non-bot && non-trusted && no malicious-scan marker on head | `needs-malicious-scan` | — | dispatch scanner | | 3 | `CHANGES_REQUESTED` \|\| unresolved threads > 0 \|\| `mergeable_state == dirty` | `needs-author-attention` | `waiting-on-author` | author-ping | | 4 | eval == success && `APPROVED` | `ready-for-merge` | `ready-to-merge` | maintainer-ping/C | | 5 | eval == success && `REVIEW_REQUIRED`/none | `ready-for-review` | `waiting-on-review` | maintainer-ping/A | | 6 | eval == success && other decision | `in-review` | `pr-state/in-review` | reconcile only | | 7 | otherwise | `ready-for-eval` | `pr-state/ready-for-eval` | eval-trigger | Trusted = `OWNER` / `MEMBER` / `COLLABORATOR`. Bots are short-circuited as trusted. ## Cool-down and idempotency Each ping variant writes a hidden HTML marker into its comment. The worker fetches prior bot comments and: - If a marker for the same variant exists within `COOLDOWN_DAYS` (default 4), the new comment is suppressed. - A first-ping age gate (default 30 min after PR creation) prevents pings on freshly opened PRs; it is bypassed once any prior ping marker exists. Marker shapes: - `` - `` - `` ## Labels owned by these workflows State labels (exactly one is reconciled at a time). Where the existing label taxonomy already covered a state, the workflow reuses it rather than introducing a duplicate `pr-state/*` name: - `pr-state/ready-for-eval` *(new)* - `waiting-on-review` *(existing — reused for `ready-for-review`)* - `ready-to-merge` *(existing — reused for `ready-for-merge`)* - `waiting-on-author` *(existing — reused for `needs-author-attention`)* - `pr-state/in-review` *(new)* Triggers and opt-outs: - `evaluate-now` — applied to fire evaluation; removed by the gate after consumption. - `no-stale` — opt-out of stale-PR closure (honored by the `stale-sweep` job) and of author/maintainer pings in the worker. ## Stale-PR sweep `pr-triage-batch.yml` includes a deterministic `stale-sweep` job that replaces the former agentic `close-stale-prs.agent.md`. It runs weekly (cron `17 4 * * 1`) and on manual `workflow_dispatch` with `stale_sweep=true`, and executes [`.github/scripts/pr-stale-sweep.sh`](../../.github/scripts/pr-stale-sweep.sh) — no model calls, no tokens. Policy (unchanged from the agentic version): - Considers every **open** PR, **including drafts**. - "Last activity" is the most recent **non-bot** comment or review; if there is none, it falls back to the PR's `created_at`. `updated_at` and all `[bot]` activity are ignored so the bot's own warning never resets the timer. - created ≤ 30 days ago → skip (too new). - 30 days < inactivity ≤ 37 days → post a stale **warning** (once; guarded by a `` marker). - inactivity > 37 days → **close** the PR with a closing comment. - Exempt: the `no-stale` label; authors `dotnet-maestro[bot]` / `dotnet-maestro`. Inputs (via `workflow_dispatch`): `stale_sweep` (run the sweep), `dry_run` (log decisions without writing), `stale_max` (hard cap on warn+close writes, default 25). --- ### Dotnet Experimental/Exp Mock Usage Analysis Design Notes # Mock Usage Analysis Skill — Design Notes ## Evaluation Results (March 2026) ### Round 1 — Initial skill (11 scenarios) | Scenario | Baseline | Isolated | Plugin | Verdict | | --- | --- | --- | --- | --- | | Detect mocking of DTOs, records, and enums | 4.3/5 | 3.3/5 ⏰ | 4.3/5 | ❌ | | Detect unused and unreachable mock setups | 4.0/5 | 4.0/5 ⏰ | 4.3/5 ⏰ | ❌ ¹ | | Detect redundant mock configurations | 3.0/5 | 2.3/5 ⏰ | 3.3/5 | ❌ | | Detect mocking of stable framework types | 3.0/5 | 5.0/5 | 5.0/5 | ✅ | | Recognize well-placed mocks | 5.0/5 | 5.0/5 | 5.0/5 | ❌ ² | | Analyze mock usage in NSubstitute tests | 3.7/5 | 5.0/5 | 5.0/5 | ✅ | | Analyze mock usage in FakeItEasy tests | 5.0/5 | 4.7/5 | 4.7/5 | ❌ | | Detect excessive mock configuration sprawl | 3.3/5 | 4.0/5 | 3.3/5 | ✅ | | Decline request to write new tests | 2.0/5 | 2.0/5 | 2.3/5 | ❌ ³ | | Decline non-mock test anti-patterns | 5.0/5 | 5.0/5 | 5.0/5 | ❌ ² | | Decline mock framework migration | 4.0/5 | 4.0/5 | 4.0/5 | ❌ ⁴ | **3/11 passed.** Overfitting: 0.06 (excellent). ¹ Quality improved in plugin but weighted score -25.2% from token/time overhead. ² Baseline at ceiling — no headroom for skill to add value. ³ Token overhead regression on a non-activation scenario with no quality gain. ⁴ Weighted -1.9% from token/time overhead with no quality delta. **Issues identified:** - **Timeouts** on scenarios 1-3 (120s too short for fixture-based scenarios) - **Activation failures** — scenario 1 not activated in plugin, scenario 3 not activated in either mode (prompts lacked mock-specific keywords) - **Baseline at ceiling** — 4 scenarios where the model already scores 5.0/5 ### Round 2 — Fix timeouts, activation, and no-headroom scenarios **Changes:** - Increased timeouts: 120s → 180s for scenarios with fixture files - Rewrote prompts with explicit mock terminology for better activation - Added `reject_tools: ["bash", "edit"]` to FakeItEasy and well-placed mocks scenarios - Improved skill description with framework-specific keywords (Mock, Substitute.For, A.Fake) - Removed "Decline write tests" scenario (token overhead, no value) | Scenario | Baseline | Isolated | Plugin | Verdict | | --- | --- | --- | --- | --- | | Detect mocking of DTOs, records, and enums | 5.0/5 | 5.0/5 | 5.0/5 | ❌ ⁵ | | Detect unused and unreachable mock setups | 3.3/5 | 5.0/5 | — | ✅ | | Detect redundant mock configurations | 3.0/5 | 4.0/5 | — | ✅ | | Detect mocking of stable framework types | 3.0/5 | 5.0/5 | — | ✅ | | Recognize well-placed mocks | 5.0/5 | 5.0/5 | 5.0/5 | ❌ ⁵ | | Analyze mock usage in NSubstitute tests | 3.0/5 ⏰ | 5.0/5 | — | ✅ | | Analyze mock usage in FakeItEasy tests | 4.3/5 | 4.7/5 | — | ❌ | | Detect excessive mock configuration sprawl | 3.0/5 | 4.0/5 | — | ✅ | | Decline non-mock test anti-patterns | 5.0/5 | 3.7/5 ⏰ | — | ❌ | | Decline mock framework migration | 5.0/5 | 5.0/5 | — | ❌ ⁵ | ⁵ Baseline at ceiling — model handles these well without skill guidance. **Improvements from Round 1:** - DTOs scenario: now activates in both isolated and plugin (was plugin-only failure) - Redundant mocks: now activates in both modes (was NOT ACTIVATED in either) - No more timeouts on scenarios 1-3 - NSubstitute baseline still hit timeout at 120s ### Round 3 — Remove no-headroom scenarios, fix NSubstitute timeout **Changes:** - Removed 4 scenarios where baseline scores 5.0/5 (see "Decisions" below) - Increased NSubstitute timeout: 120s → 180s 6 remaining scenarios all show positive skill impact. ## Key Insight The baseline LLM already excels at two mock-related tasks: 1. **Identifying trivial-type mocking** — the model recognizes when `Mock` should be `new CustomerDto(...)` without guidance. 2. **Recognizing well-placed mocks** — when tests correctly mock external boundaries, the model concludes the approach is sound without inventing false positives. The skill's unique value is in **deep code-path analysis**: tracing mock setups through production code to determine whether they are actually invoked at runtime, identifying unreachable setups after early returns or exceptions, and detecting redundant configurations duplicated across tests. ## Decisions ### Removed: "Detect mocking of DTOs, records, and enums" scenario Baseline scores 5.0/5 — the model already identifies when DTOs, records, and enums are unnecessarily mocked and recommends real instance construction. No quality delta for the skill to contribute. ### Removed: "Recognize well-placed mocks without inventing false positives" scenario Baseline scores 5.0/5 — the model already correctly concludes that mocking external boundaries (HTTP, DB, email) is appropriate without inflating severity. ### Removed: "Decline when asked about non-mock test anti-patterns" scenario Baseline scores 5.0/5 — non-activation scenario where the model already handles Thread.Sleep/DateTime.Now reviews without the skill. The timeout regression (5.0→3.7 ⏰) in the skilled run was caused by the 60s timeout being too short, not a skill problem. ### Removed: "Decline mock framework migration request" scenario Baseline scores 5.0/5 — the model already handles Moq→NSubstitute migration requests without the skill. Weighted score was -1.2% from time overhead alone. ### Round 4 — Drastic skill simplification **Problem:** Results degraded significantly. The skill was actively hurting quality — 5/6 scenarios scored worse with the skill than without it. When activated, scores dropped from 2.3-3.3 baseline to 1.0-2.3. Two scenarios showed "NOT ACTIVATED" indicating the skill loaded but the model chose not to use it. | Scenario | Baseline | With Skill | Verdict | | --- | --- | --- | --- | | Detect unused and unreachable mock setups | 3.0/5 | 2.3/5 | ❌ | | Detect redundant mock configurations | 3.3/5 | 1.0/5 (NOT ACTIVATED) | ❌ | | Detect mocking of stable framework types | 3.0/5 | 2.3/5 (NOT ACTIVATED) | ❌ | | Analyze mock usage in NSubstitute tests | 2.3/5 | 1.0/5 | ❌ | | Analyze mock usage in FakeItEasy tests | 3.3/5 | 1.0/5 | ❌ | | Detect excessive mock configuration sprawl | 2.7/5 | 3.7/5 | ✅ | **Root cause analysis:** 1. **Skill too verbose (~200 lines)** — Extensive anti-pattern catalog tables that the model already knows, consuming attention budget that should go to code analysis. 2. **6-step workflow too rigid** — Model spent effort following the prescribed categorization workflow (classify dependencies as Trivial/Stable/Thin/External/Complex) instead of doing actual code-path tracing. 3. **Anti-pattern encyclopedia redundant** — 4 severity levels × 3-4 patterns each = 15+ anti-patterns listed. The model already knows these; listing them added noise without value. 4. **Reporting format instructions too prescriptive** — "Present findings in this structure: Summary → Critical/High → Medium/Low → Positive → Aggregate" forced a template that didn't match rubric expectations. **Changes:** - Cut skill from ~200 lines to ~90 lines - Reduced workflow from 6 steps to 4 (read → trace → check replaceable → report) - Removed the anti-pattern catalog entirely — model already knows common mock anti-patterns - Removed dependency categorization tables (Trivial/Stable/Thin/External/Complex) - Removed the runtime data incorporation step (never used in evals) - Focused Step 2 entirely on **code-path tracing** — the unique value-add identified in Round 2 - Added explicit guidance on early returns, exceptions, and branch-specific logic as things to trace - Simplified reporting to: specific location + why unreachable + concrete fix - Enhanced description with more trigger phrases for better activation --- ### Dotnet Experimental/Exp Test Maintainability Design Notes # Test Maintainability Skill — Design Notes ## Evaluation Results (March 2026) ### Round 1 — Original skill | Scenario | Baseline | Isolated | Plugin | Verdict | | --- | --- | --- | --- | --- | | Selectively recommend changes | 4.7/5 | 5.0/5 | 5.0/5 | ❌ ¹ | | Data-driven patterns + display names | 4.0/5 | 5.0/5 | 4.7/5 | ✅ | | Well-maintained recognition | 4.0/5 | 5.0/5 | 5.0/5 | ✅ | | Oversized tests | 5.0/5 | 5.0/5 | 5.0/5 | ❌ ¹ | ¹ Quality matched or improved but weighted score penalized by token overhead. ### Round 2 — After trimming (removed Steps 4-5, pitfalls, validation checklist) | Scenario | Baseline | Isolated | Plugin | Verdict | | --- | --- | --- | --- | --- | | Selectively recommend changes | 5.0/5 | 5.0/5 | 5.0/5 | ❌ ² | | Data-driven patterns + display names | 4.0/5 | 4.6/5 | 4.4/5 | ❌ ³ | | Well-maintained recognition | 4.6/5 | 5.0/5 | 5.0/5 | ✅ | ² Baseline at ceiling — same problem as "Oversized tests". ³ **Regression** — trimming removed implicit reinforcement about `DataRow`+`DisplayName`. The skill steered the model toward `[DynamicData]` instead of `[DataRow]` with `DisplayName`, which the rubric penalizes. Fixed by adding an explicit calibration rule: "Prefer `[DataRow]` with `DisplayName` over `[DynamicData]` when values are compile-time constants." ### Round 3 — After re-adding DataRow calibration rule | Scenario | Baseline | Isolated | Plugin | Verdict | | --- | --- | --- | --- | --- | | Selectively recommend changes | 5.0/5 | 5.0/5 | 5.0/5 | ❌ ⁴ | | Data-driven patterns + display names | 4.0/5 | 4.3/5 | 4.3/5 | ❌ ⁵ | | Well-maintained recognition | 4.7/5 | 5.0/5 | 5.0/5 | ❌ ⁶ | ⁴ Quality unchanged, weighted -11.0% due to tokens (13388 → 35555), tool calls (0 → 2), time (16.9s → 34.6s). ⁵ Quality improved 4.0→4.3 but weighted -4.2% due to tokens (13148 → 30114), tool calls (0 → 1), time (15.7s → 33.6s). ⁶ Quality improved 4.7→5.0 but weighted -15.0% due to tokens (12736 → 37977), tool calls (0 → 2), time (17.3s → 30.4s). ### Round 4 — Aggressive trim to calibration rules only Removed: When to Use, When Not to Use (covered by frontmatter description), Inputs table, Step 1 (gather code), Step 2 detection tables (model handles detection natively at ceiling quality), Step 4 (report format). Kept only the heading, one-line workflow, and the 6 calibration rules that encode the skill's unique judgment value. Cuts ~75% of skill tokens. | Scenario | Baseline | Isolated | Plugin | Verdict | | --- | --- | --- | --- | --- | | Selectively recommend changes | 5.0/5 | 5.0/5 | 5.0/5 | ❌ ⁷ | | Data-driven patterns + display names | 4.0/5 | 5.0/5 | 4.0/5 | ❌ ⁸ | | Well-maintained recognition | 4.3/5 | 5.0/5 | 5.0/5 | ✅ | ⁷ Baseline at ceiling — same problem as "Oversized tests". Removed from eval. ⁸ Isolated improved but plugin didn't. Weighted -9.4% from token overhead. ## Key Insight The baseline LLM already excels at refactoring recommendations (extracting builders, splitting oversized tests). The skill's unique value is in **judgment calls**: recognizing well-maintained tests, calibrating when NOT to recommend changes, and recommending display names for non-obvious values. ## Decisions ### Removed: "Oversized tests" eval scenario Baseline scores 5.0/5 — there is zero quality delta for the skill to contribute. Any non-zero token overhead makes the weighted score negative. This scenario cannot pass regardless of how much we trim the skill. ### Removed: "Selectively recommend changes" eval scenario Baseline hit ceiling at 5.0/5 after Round 2 trimming. Same problem as "Oversized tests" — zero quality delta means token overhead always produces a negative weighted score. ### Trimmed: SKILL.md output formatting and pitfalls sections Removed Steps 4-5 (detailed report structure, show-refactored-code instructions), the Validation checklist, and the Common Pitfalls table. These either duplicate the Step 3 calibration rules or teach behaviors the model already does natively (before/after code, quantified benefits). This cuts ~25% of skill tokens while preserving the core detection tables and calibration guidance that drive the passing scenarios. ## Scenarios Not Worth Adding These refactoring tasks were considered but not pursued as eval scenarios because the baseline model handles them at or near ceiling quality: - **Identifying oversized / multi-concern tests** — Model reliably spots 50+ line tests with multiple arrange-act-assert cycles and recommends splitting. - **Extracting repeated setup into helpers** — Model recognizes 3+ repeated setup blocks and suggests `TestInitialize`, helper methods, or factory patterns. - **Recommending builder patterns** — Model identifies scattered complex object construction and proposes builders when warranted. The skill focuses instead on the judgment-heavy scenarios where the baseline struggles: restraint (knowing when code is already good enough) and display name calibration. --- ### Agentic Workflows # DevOps Agentic Workflows Cross-cutting [GitHub Agentic Workflow](https://github.com/github/gh-aw) workflows for repository-wide DevOps automation. The workflow source files live in `.github/workflows/` and are compiled with `gh aw compile` to generate `.lock.yml` files (standard GitHub Actions YAML with security hardening). These workflows monitor the _entire repo_ (all components, all pipelines, all PRs). ## Available Workflows | Workflow | Description | Trigger | |----------|-------------|---------| | [devops-health-check](../.github/workflows/devops-health-check.md) | Daily orchestrator that collects repo infrastructure health signals (pipelines, CI/CD infrastructure, resource usage), computes a fingerprint-based diff against the previous run, and updates a pinned health dashboard issue | `cron: 0 3 * * *` (03:00 UTC daily), `workflow_dispatch` | | [devops-health-investigate](../.github/workflows/devops-health-investigate.md) | Worker agent dispatched by the health check orchestrator to perform deep root-cause analysis on individual findings | `workflow_dispatch` (dispatched by orchestrator via `dispatch-workflow`) | | [devops-health-groom](../.github/workflows/devops-health-groom.md) | Runs ~3h after the health check to link investigation results into the issue body, hide stale comments (>7 days), and clean up resolved investigations | `cron: 0 6 * * *` (06:00 UTC daily), `workflow_dispatch` | | [issue-triage](../.github/workflows/issue-triage.md) | Triages individual issues: assigns an `area-*` label, identifies owners from CODEOWNERS, adds the `Triaged` label, and posts a brief actionable summary | `issues: [opened, reopened]`, `workflow_dispatch` | | [issue-triage-batch](../.github/workflows/issue-triage-batch.yml) | Deterministic workflow that dispatches the issue-triage agent for each untriaged issue in an optional date range | `workflow_dispatch` (with optional `date_from`/`date_to`) | | [issue-investigate](../.github/workflows/issue-investigate.md) | Deep investigation agent that analyzes an issue against the codebase, suggests next steps, and creates a draft PR if the fix is clear | `issues: [labeled]` (when `auto-investigate` label is added) | ## Architecture ``` devops-health-check (Orchestrator) ─── runs daily ├─ Collects health signals from 3 categories: │ Pipeline · Infrastructure · Resources ├─ Fingerprints each finding for stable diff tracking ├─ Classifies: 🆕 NEW · 📌 EXISTING · ✅ RESOLVED ├─ Updates pinned health dashboard issue └─ Dispatches investigation workers (up to 10) │ ▼ devops-health-investigate (Worker × N) ─── dispatched ├─ Investigates ONE finding with fresh context ├─ Follows category-specific playbook ├─ Determines root cause + remediation └─ Posts investigation results as a comment on the health issue │ ▼ (~3 hours later) devops-health-groom (Groomer) ─── runs daily ├─ Links investigation comments into the issue body │ (updates 🔄 Dispatched → ✅ Done with summary + link) ├─ Marks resolved investigations as ✅ Resolved ├─ Hides (collapses) daily overview comments older than 7 days └─ Hides (collapses) investigation comments for resolved findings ``` ## Setup 1. Install the `gh aw` CLI extension: `gh extension install github/gh-aw` 2. Compile: `gh aw compile` (from the repo root — this compiles all `.md` files in `.github/workflows/`) 3. Commit both the `.md` and generated `.lock.yml` files 4. The health check runs daily, or on-demand via `workflow_dispatch` ## Local Development ```powershell # Compile workflows (generates .lock.yml from .md frontmatter) gh aw compile # Compile with validation gh aw compile --strict # Dry-run (validates without triggering on GitHub Actions) gh aw run devops-health-check --dry-run # Run on GitHub Actions (from a pushed branch) gh aw run devops-health-check --push --ref ``` ## File Structure ``` .github/ ├── workflows/ │ ├── devops-health-check.md # Orchestrator workflow │ ├── devops-health-check.lock.yml # Compiled workflow (generated by gh aw compile) │ ├── devops-health-investigate.md # Worker workflow │ ├── devops-health-investigate.lock.yml # Compiled workflow (generated by gh aw compile) │ ├── devops-health-groom.md # Grooming workflow │ ├── devops-health-groom.lock.yml # Compiled workflow (generated by gh aw compile) │ ├── issue-triage.md # Issue triage agent │ ├── issue-triage.lock.yml # Compiled workflow (generated by gh aw compile) │ ├── issue-triage-batch.yml # Batch triage dispatcher (standard GHA) │ ├── issue-investigate.md # Deep issue investigation agent │ └── issue-investigate.lock.yml # Compiled workflow (generated by gh aw compile) └── aw/ └── shared/ ├── devops-health.lock.md # Health check catalog & fingerprinting rules └── devops-investigate.lock.md # Investigation playbooks & remediation templates --- ### CONTRIBUTING # Contributing Thanks for your interest in contributing. We expect to accept external contributions, but the bar for merging is intentionally high. This repository contains shared building blocks for coding agents: - Skills: reusable, task focused instruction packs - Agents: role based configurations that bundle tool expectations and skill selection Because these artifacts can affect many users and workflows, we prioritize correctness, clarity, and long term maintainability over speed. ## Code ownership Every plugin, skill, and agent must have designated owners in the `.github/CODEOWNERS` file. When you add a new skill or agent, add a matching CODEOWNERS entry. Ownership must be either: - **Two or more FTE GitHub aliases** (e.g., `@user1 @user2`), or - **A GitHub team alias** (e.g., `@dotnet/my-team`) This ensures that every contribution area has accountable reviewers and that PRs are automatically routed to the right people. ## Repository layout ```text plugins/ / plugin.json skills/ / SKILL.md scripts/ references/ assets/ agents/ .agent.md tests/ / / eval.yaml ``` Every plugin must have a plugin.json file in the plugin root that is linked to from the marketplace.json file. ### Plugin organization Skills are grouped into domain-specific plugins. When proposing a new skill, place it in the plugin that best matches its domain. See [README.md](README.md) for the current list of plugins. If your skill does not fit any existing plugin, consider creating a new one. To create a new plugin: 1. Add `plugins//plugin.json` and a `skills/` directory beneath it. 2. Add a matching entry in `.github/plugin/marketplace.json`, `.claude-plugin/marketplace.json`, `.cursor-plugin/marketplace.json`, and `.agents/plugins/marketplace.json`. Keep plugin entries consistent across all marketplace manifests (including `plugins[].source` format) to reduce drift and make future updates safer. Also add a `plugins//version.json` (copy an existing one) so the plugin participates in automated versioning. Start its `plugin.json` version at `0.1.0`. 3. Add a CODEOWNERS entry for the new plugin and its tests (see [Code ownership](#code-ownership)). 4. Add the plugin to the **What's Included** table in the root `README.md`. 5. Create a `tests//` directory for skill tests. See existing plugins for the expected format. ### The `dotnet-experimental` plugin Use `dotnet-experimental` when you want to try out a skill idea but are not yet confident it belongs in a stable plugin — for example, when the skill is outside your usual area of responsibility, the approach is unproven, or you want community feedback before committing to a long-term home. Skills in `dotnet-experimental`: - May change, be reworked, or be removed without notice. - Are held to the same quality and testing standards as any other skill (frontmatter, `eval.yaml`, etc.). - Should eventually graduate to a stable plugin or be retired. When a skill has proven itself, move it to the appropriate domain plugin and update tests accordingly. Place experimental skills under `plugins/dotnet-experimental/skills/` with matching tests in `tests/dotnet-experimental/`. ## Plugin versioning Each plugin is versioned independently. The same version is duplicated across every manifest a consumer reads: `plugins//plugin.json` and `plugins//.codex-plugin/plugin.json` (both present for every plugin), plus an optional `plugins//.claude-plugin/plugin.json` that only plugins needing an inline Claude manifest carry (e.g. `dotnet-msbuild`'s binlog MCP server). Consumers (Copilot CLI, Claude, Codex, Cursor) read the version directly from this repository. Versioning is automated with [Nerdbank.GitVersioning](https://github.com/dotnet/Nerdbank.GitVersioning). A per-plugin `plugins//version.json` scopes the git height to that plugin's subtree, so the **patch** number is derived from history — you do not edit it by hand. The generated manifests (`plugin.json`, `.codex-plugin/plugin.json`, and `.claude-plugin/plugin.json` where present) and `version.json` itself are excluded from that height via the `pathFilters`, so editing only manifest metadata (anything other than a deliberate base bump in `version.json`) does **not** change the patch number and is **not** picked up by `/version-bump` or the weekly sync. Touch a skill or other plugin content to bump the version. What this means when you contribute: - **Every plugin carries its own `plugins//version.json`.** It declares the plugin's version base; the weekly sync fails fast if a plugin ships a `plugin.json` without one, so none is ever left unversioned. - **Don't hand-edit the `version` field** in any of the manifests (`plugin.json`, `.codex-plugin/plugin.json`, or `.claude-plugin/plugin.json`). The patch number is computed and stamped automatically, and a manual edit will be overwritten. - **The only version field you may change is the base** (`"version"`) in `plugins//version.json`, and only to declare a deliberate **minor or major** release of that plugin (e.g. `0.1` → `0.2` or `1.0`). Changing the base resets the patch number to `0`. - After a PR changes a plugin's content, bumping its version is optional: - A maintainer can comment **`/version-bump`** on a same-repo PR to stamp the new version onto the branch. - Otherwise the **weekly version sync** opens a PR that stamps any plugin whose content changed without a version bump, explaining each change. Nothing is ever missed. Patch numbers are predicted from git history, so two PRs bumped concurrently can land the same patch number for a plugin; the weekly sync recomputes the authoritative height on `main` and reconciles any collision. Version-only changes do not trigger skill evaluations. ## Before you start - Search existing issues and pull requests to avoid duplicates. - Start with an issue before you submit a pull request for a new skill, a new agent, or any non trivial change. This helps us align on scope and avoids wasted work. - Small fixes like typos, broken links, or clearly isolated corrections can go straight to a pull request. - Keep changes small and focused. One skill or one agent per pull request is a good default. ## What we look for We are most likely to accept contributions that are: - Addresses a LLM gap and is clearly motivated by a real use case - Likely to be used frequently and is general (not repo-specific) - Narrow in scope and easy to review - Tool conscious and explicit about assumptions - Verifiable with concrete validation steps - Written to be durable across repo changes We are less likely to accept contributions that: - Add broad frameworks, meta tooling, or large reorganizations - Duplicate guidance that already exists in another skill - Encode private environment details, credentials, or company specific secrets - Depend on proprietary tools or access that most contributors will not have - Skills that make use of third party tools will be evaluated on a case by case basis. Acceptance of such skills will depend on our evaluation of the provenance and maturity of any such tools. ## Proposing a new skill Please review the **What we look for** section and add justification for the skill in your issue and PR. A skill should be self-contained and: - Clearly state **what it does** and **when to use it**. - Frontmatter (name and description) is small and minimal, just enough for LLM to understand when to use it - Keep the SKILL.md body under 500 lines for optimal performance. Split content into separate files when you approach this limit. Use a progressive disclosure pattern, referring to those files from the SKILL.md file where needed. - Specify required inputs (repo context, environment, access needs). - Prefer concrete checklists and verification steps over vague guidance. Create a new folder under a plugin's `skills/` directory: ```text plugins//skills//SKILL.md ``` A skill should answer three questions up front: 1. What outcome does the skill produce 2. When should an agent use it 3. How does the agent validate success ### Skill naming Use short, kebab-case names that mirror how developers naturally phrase the task, prioritizing keyword overlap over grammar — e.g., add-aspnet-auth, configure-jwt-auth, setup-identity-server. Optionally using gerund style (verb-ing) is acceptable as well - e.g., configuring-caching. Optimize for intent matching: lead with the action verb users actually say (add, configure, setup, deploy) followed the outcome the skill is aiming to assist. The `SKILL.md` is required to have front-matter at a minimum: Create the file with required YAML frontmatter: ```yaml --- name: description: --- ``` > **Tip:** The `description` field is used by the agent runtime to decide whether to load the full skill. > Include **when to use** and **when not to use** guidance directly in the description so the agent can > select or skip skills without reading the entire `SKILL.md`. This avoids unnecessary token usage. > See [`thread-abort-migration/SKILL.md`](plugins/dotnet-upgrade/skills/thread-abort-migration/SKILL.md) for a good example. ### Recommended `SKILL.md` sections - **Purpose**: one paragraph describing the outcome. - **When to use** / **When not to use** (put the essentials in the frontmatter `description`; expand here only if more detail is needed). - **Inputs**: what the agent needs (files, commands, permissions). - **Workflow**: numbered steps with checkpoints. - **Validation**: how to confirm the result (tests, linters, manual checks). - **Common pitfalls**: known traps and how to avoid them. ### Skill checklist Include a `SKILL.md` that covers: - Purpose and non goals - When to use and when not to use (summarized in the frontmatter `description`; body section for extended detail) - Inputs and prerequisites - Step by step workflow with checkpoints - Validation steps that can be run or observed - Failure modes and recovery guidance Also: - Avoid duplicating text across multiple skills. Prefer referencing shared patterns. - Do not include content copied from other repositories. If you are inspired by existing work, rewrite in your own words and adapt it to our conventions. ## Proposing a new agent An agent definition should be opinionated but bounded: - Describe the **role** (e.g., "WinForms Expert", "Security Reviewer", "Docs Maintainer"). - Define boundaries (what the agent should not do). - List the skills it expects to use and how it chooses among them. Add an agent file under a plugin's `agents/` directory: ```text plugins//agents/.agent.md ``` ### Agent checklist Include documentation that explains: - Role and intended tasks - Boundaries and safety constraints - Tooling assumptions - How the agent chooses which skills to apply - What a good completion looks like, including validation expectations ## Testing and validation Skills and agents are documentation driven, but we still treat them as production assets. - Every change should include a validation section that a reviewer can follow. - If your change references commands, keep them cross platform when practical. If not, state the supported environment. - If your change depends on external services, document how a reviewer can validate without privileged access, or explain why validation is not possible. ### Writing skill tests Each skill should have an `eval.yaml` file that defines test scenarios. Tests live under the repo root `tests/` directory, matching the plugin and skill name: ```text tests///eval.yaml ``` The exception is a helper or reference skill that sets `disable-model-invocation: true`. The model cannot self-activate it, so an activation-graded eval would compare two identical arms. Cover those through the evals of the skills that load them and through the plugin arm instead. The skeleton below shows the shape only — it declares a single trial and would therefore be rejected by the quality gate. See [Size the eval so it can return a verdict](#size-the-eval-so-it-can-return-a-verdict) for the real bar. ```yaml name: my-skill description: Evaluates the / skill type: capability defaults: timeout: 3m runs: 1 stimuli: - name: "Describe what the agent should do" prompt: | The prompt sent to the agent. graders: # Deterministic graders check the produced output/artifacts. - type: exit-success - type: output-contains config: substring: "expected text in agent output" # The `prompt` grader runs the LLM judge against the rubric below. - type: prompt rubric: - The agent correctly identified the issue - The agent suggested a concrete fix ``` > [!IMPORTANT] > `defaults:` and `config:` are the same block — `config` is a deprecated alias — and vally > **rejects** a spec declaring both. Many existing evals still open with `config:`; when you add > `runs`, merge the two into a single `defaults:` block. The failure is silent: the job exits 0 with > no verdicts and the PR comment blames "transient infrastructure". Each skill is evaluated in up to three variants — **baseline** (no skills), **skilled** (only the skill under test), and **plugin** (the whole plugin loaded) — and a skill "passes" only when the skilled run is a *credible* improvement over baseline. To assert that a skill should stay dormant for an out-of-scope task, add `expect_activation: false` to that stimulus. See any existing `tests/*/*/eval.yaml` for a fuller example of the grader and stimulus format. #### Size the eval so it can return a verdict The pass gate has two independent bars. `trials = stimuli × runs`. 1. **Counted trials ≥ 5**, else the verdict is reported `underpowered` — never a pass, never a regression. 2. **p ≤ 0.05 on an exact one-sided sign test over the *discordant* (non-tie) trials.** Ties are not discarded; they hold the discordant count down. | discordant trials | records that pass | p | | ---: | --- | ---: | | ≤ 4 | none, however good the skill | ≥ 0.0625 | | 5–7 | zero losses only (5W/0L) | 0.031 | | 8 | one loss survivable (7W/1L) | 0.035 | At exactly 5 counted trials a single tie is fatal — it leaves 4 discordant. At 6 counted trials one tie is survivable (5W/1T/0L); at 7, up to two are (5W/2T/0L). A loss is not. Five is an *eligibility floor*, not adequate power. A run that measured a 32% tie rate certified a genuinely-helping five-trial eval about one time in ten; at fifteen trials, about nine times in ten. Prefer adding **discriminating stimuli** over raising `runs` — repeats measure the same task. See [`eng/eval-quality/README.md`](eng/eval-quality/README.md) for the full derivation and for the ten structural defects the CI quality gate blocks. Run the gate locally before pushing: ```bash python eng/eval-quality/check_eval_quality.py ``` ### Running tests locally Prerequisites: Node.js 20+ and the [GitHub CLI](https://cli.github.com) signed in (`gh auth login`). The script checks these and tells you what's missing, so just run it: ```bash # Run tests for a single skill ./eng/run-skill-evals.sh dotnet-msbuild binlog-failure-analysis # Run tests for a whole plugin ./eng/run-skill-evals.sh dotnet-msbuild # Run every skill's tests ./eng/run-skill-evals.sh ``` Per-skill verdicts are written to `./eval-results///results.json`, and the raw experiment output goes to `./eval-results/_experiment/`. Model and judge model come from the `overrides:` block in `dotnet-skills.experiment.yaml`. > [!WARNING] > LLM evaluations are noisy. Runs-per-stimulus is deliberately **not** set in > `dotnet-skills.experiment.yaml`: an experiment-level `runs` overwrites every eval's own value > instead of defaulting it, making per-eval trial counts impossible to express. Raise the eval's own > `defaults.runs` instead — or, better, add discriminating stimuli. ### CI evaluation Tests do **not** run automatically on pull requests. When a PR changes skills, the `pr-status` job posts a pending commit status and a maintainer must trigger the evaluation, binding it to a specific reviewed commit — either by submitting a PR review ("Files changed" → "Review changes") whose body contains `/evaluate` (recommended, no SHA to copy), or by commenting `/evaluate `. A bare `/evaluate` comment only posts guidance. Results are posted as a PR comment and uploaded as build artifacts. If a scenario fails or regresses, see [Investigating Results](eng/vally-adapter/InvestigatingResults.md) for how to download artifacts, interpret `results.json`, and diagnose common failure patterns. ## Writing style - Be concise and specific. - Prefer numbered steps for workflows. - Prefer checklists for requirements. - Define terminology the first time it appears. - Avoid excessive formatting and avoid clever wording that could be misread by an agent. ## Security and safety - Do not include secrets, tokens, or internal URLs. - If you discover a security issue, do not open a public issue with sensitive details. Use the repository or organization security reporting process instead. ### External references Skills often reference external tools, documentation, and projects — this is expected and welcome, including community and third-party resources. To help reviewers stay aware of external dependencies, the repository includes an automated reference scanner (integrated into `skill-validator check`) that runs in CI against plugin content (SKILL.md, agent files, and reference docs). The scanner treats all of the following as CI-blocking errors: - `http://` URLs where `https://` should be used - `