Skip to content

[evals] Daily Evals Feature Report - 2026-07-29 #48821

Description

@github-actions

Executive Summary

24 evals-bearing runs across 17 workflows were analyzed. Evals jobs were reliable, but the overall YES rate was 32.7%, so the feature is DEGRADED. Only 1 of 24 runs passed every eval question. I counted UNKNOWN answers as non-YES for health math; 26 of 55 eval outputs were UNKNOWN.

Note

Status: DEGRADED - evals artifacts were produced for every evals run, but the overall YES rate is below the healthy threshold.

Key Metrics

Metric Value
Workflows with evals 17
Runs analyzed 24
Runs with evals results 24
Evals job success rate 100.0%
Overall YES rate 32.7%

Per-Workflow Pass Rates

Workflow Runs Evals Job Success Run Pass Rate Lowest-Scoring Question
Design Decision Gate 🏗️ 4 100.0% 0.0% "Did the agent add a PR comment, push a draft ADR, or call noop?" (0.0% YES)
PR Code Quality Reviewer 4 100.0% 0.0% "Does the agent output show that the review findings are limited to changes in the pull request diff rather than unrelated code?" (0.0% YES)
Smoke Copilot 2 100.0% 100.0% "Did the gh-aw binary build succeed? Look for a successful make build step or explicit mention that compilation succeeded." (100.0% YES)
Auto-Triage Issues 1 100.0% 0.0% "Was a summary discussion created listing the issues processed and the labels applied?" (0.0% YES)
Code Scanning Fixer 1 100.0% 0.0% "Was a pull request created with a remediation for a code scanning alert, or was noop used when no fixable alerts existed?" (0.0% YES)
Copilot Session Insights 1 100.0% 0.0% "Was a report produced with usage patterns, success rates, and performance metrics?" (0.0% YES)
Discussion Task Miner - Code Quality Improvement Agent 1 100.0% 0.0% "Does the agent output confirm that the created issues include the expected labels (code-quality, automation, task-mining)?" (0.0% YES)
Issue Monster 1 100.0% 0.0% "Does the agent output show that at most one issue was assigned to Copilot per run?" (0.0% YES)
PR Sous Chef 1 100.0% 0.0% "Did the agent add a comment to at least one pull request?" (0.0% YES)
PR Triage Agent 1 100.0% 0.0% "Does the agent output include a triage report summarizing the PRs processed?" (0.0% YES)
Smoke Antigravity 1 100.0% 0.0% "Does the agent output show that the objective for experiment sub_agent_strategy was successfully completed?" (0.0% YES)
Smoke Copilot - AOAI (Entra) 1 100.0% 0.0% "Does the agent output show that the objective for experiment caveman was successfully completed?" (0.0% YES)
Smoke Copilot - AOAI (apikey) 1 100.0% 0.0% "Does the agent output show that the objective for experiment caveman was successfully completed?" (0.0% YES)
Smoke Gemini 1 100.0% 0.0% "Does the agent output show that the objective for experiment sub_agent_strategy was successfully completed?" (0.0% YES)
Sub-Issue Closer 1 100.0% 0.0% "Did the agent check parent issues for the completion status of all their sub-issues?" (0.0% YES)
Tidy 1 100.0% 0.0% "Was a pull request created with formatting and tidying fixes, or was noop used when no changes were needed?" (0.0% YES)
[aw] Failure Investigator (6h) 1 100.0% 0.0% "Were fix sub-issues created for unresolved failures, or were resolved tracking issues closed?" (0.0% YES)
Per-Question Breakdown per Workflow

Design Decision Gate 🏗️

Question ID Question YES NO YES Rate
action-taken Did the agent add a PR comment, push a draft ADR, or call noop? 0 4 0.0%
adr-check-performed Does the agent output confirm that it checked for existing ADRs before deciding on an action? 0 4 0.0%
decision-justified Does the agent output explain why an ADR is required or why no ADR gate was triggered for this PR? 0 4 0.0%

PR Code Quality Reviewer

Question ID Question YES NO YES Rate
findings_scoped Does the agent output show that the review findings are limited to changes in the pull request diff rather than unrelated code? 0 4 0.0%
review_posted Did the agent post a code review comment on the pull request? 4 0 100.0%

Smoke Copilot

Question ID Question YES NO YES Rate
build-succeeded Did the gh-aw binary build succeed? Look for a successful make build step or explicit mention that compilation succeeded. 2 0 100.0%
issue-created Was a smoke test issue created with test results? Look for a create_issue output containing 'Smoke Test' in the title. 2 0 100.0%
smoke-passed Did all or most smoke tests pass? Look for an overall PASS status or the majority of tests showing ✅ in the agent output. 2 0 100.0%

Auto-Triage Issues

Question ID Question YES NO YES Rate
labels-applied Did the agent apply at least one label to an unlabeled issue, or correctly call noop when no unlabeled issues were found? 1 0 100.0%
report-created Was a summary discussion created listing the issues processed and the labels applied? 0 1 0.0%

Code Scanning Fixer

Question ID Question YES NO YES Rate
alerts_analyzed Did the agent analyze code scanning alerts and identify at least one fixable alert, or correctly skip when no fixable alerts were found? 1 0 100.0%
pr_created_or_noop Was a pull request created with a remediation for a code scanning alert, or was noop used when no fixable alerts existed? 0 1 0.0%

Copilot Session Insights

Question ID Question YES NO YES Rate
insights_report_produced Was a report produced with usage patterns, success rates, and performance metrics? 0 1 0.0%
sessions_analyzed Did the agent analyze GitHub Copilot coding agent sessions? 1 0 100.0%

Discussion Task Miner - Code Quality Improvement Agent

Question ID Question YES NO YES Rate
labels-applied Does the agent output confirm that the created issues include the expected labels (code-quality, automation, task-mining)? 0 1 0.0%
output-produced Did the agent create at least one code quality issue or add a comment? 1 0 100.0%
tasks-extracted Does the agent output show that actionable tasks were identified from the analyzed discussions? 0 1 0.0%

Issue Monster

Question ID Question YES NO YES Rate
issue_assigned Did the agent assign at least one issue to the Copilot coding agent, or correctly skip when no suitable issues were found? 1 0 100.0%
single_issue_scoped Does the agent output show that at most one issue was assigned to Copilot per run? 0 1 0.0%

PR Sous Chef

Question ID Question YES NO YES Rate
comment-added Did the agent add a comment to at least one pull request? 0 1 0.0%
nudge-targeted Does the agent output show a specific reason why the selected PR needs a nudge toward maintainer investigation? 0 1 0.0%
pr-evaluated Does the agent output confirm that it evaluated at least one open PR for nudge eligibility? 0 1 0.0%

PR Triage Agent

Question ID Question YES NO YES Rate
labels-applied Did the agent apply triage labels to at least one pull request? 1 0 100.0%
report-produced Does the agent output include a triage report summarizing the PRs processed? 0 1 0.0%
triage-data-set Does the agent output confirm that category, risk, and action data were determined for each processed PR? 0 1 0.0%

Smoke Antigravity

Question ID Question YES NO YES Rate
sub_agent_strategy_goal_met Does the agent output show that the objective for experiment sub_agent_strategy was successfully completed? 0 1 0.0%

Smoke Copilot - AOAI (Entra)

Question ID Question YES NO YES Rate
caveman_goal_met Does the agent output show that the objective for experiment caveman was successfully completed? 0 1 0.0%
subagent_model_goal_met Does the agent output show that the objective for experiment subagent_model was successfully completed? 0 1 0.0%

Smoke Copilot - AOAI (apikey)

Question ID Question YES NO YES Rate
caveman_goal_met Does the agent output show that the objective for experiment caveman was successfully completed? 0 1 0.0%
subagent_model_goal_met Does the agent output show that the objective for experiment subagent_model was successfully completed? 0 1 0.0%

Smoke Gemini

Question ID Question YES NO YES Rate
sub_agent_strategy_goal_met Does the agent output show that the objective for experiment sub_agent_strategy was successfully completed? 0 1 0.0%

Sub-Issue Closer

Question ID Question YES NO YES Rate
issues_checked Did the agent check parent issues for the completion status of all their sub-issues? 0 1 0.0%
issues_closed_or_noop Were completed parent issues closed with a comment, or does the agent output confirm no issues were ready to close? 1 0 100.0%

Tidy

Question ID Question YES NO YES Rate
pr_created_or_noop Was a pull request created with formatting and tidying fixes, or was noop used when no changes were needed? 0 1 0.0%
tidy_completed Did the agent run code formatting and tidying tools on the codebase? 0 1 0.0%

[aw] Failure Investigator (6h)

Question ID Question YES NO YES Rate
failures_investigated Did the agent investigate agentic workflow failures from the last 6 hours and produce findings? 1 0 100.0%
issues_created_or_closed Were fix sub-issues created for unresolved failures, or were resolved tracking issues closed? 0 1 0.0%

Quality Signals

  • action-taken: 0.0% YES across 4 answers in Design Decision Gate 🏗️.
  • adr-check-performed: 0.0% YES across 4 answers in Design Decision Gate 🏗️.
  • decision-justified: 0.0% YES across 4 answers in Design Decision Gate 🏗️.

Recommendations

  • Tighten the Design Decision Gate prompt so it produces an explicit action and justification, not just an implicit decision.
  • Revisit the PR Code Quality Reviewer rubric and evidence expectations so scope checks line up with the diff being reviewed.
  • Enforce binary evaluator outputs or update the parser, because UNKNOWN answers are materially depressing the reported YES rate.

References

Generated by 🧪 Daily Evals Feature Report · gpt54 · 16.6 AIC · ⌖ 3.15 AIC · ⊞ 11.1K ·

  • expires on Aug 5, 2026, 12:13 AM UTC-08:00

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions