Executive Summary
24 evals-bearing runs across 17 workflows were analyzed. Evals jobs were reliable, but the overall YES rate was 32.7%, so the feature is DEGRADED. Only 1 of 24 runs passed every eval question. I counted UNKNOWN answers as non-YES for health math; 26 of 55 eval outputs were UNKNOWN.
Note
Status: DEGRADED - evals artifacts were produced for every evals run, but the overall YES rate is below the healthy threshold.
Key Metrics
Metric
Value
Workflows with evals
17
Runs analyzed
24
Runs with evals results
24
Evals job success rate
100.0%
Overall YES rate
32.7%
Per-Workflow Pass Rates
Workflow
Runs
Evals Job Success
Run Pass Rate
Lowest-Scoring Question
Design Decision Gate 🏗️
4
100.0%
0.0%
"Did the agent add a PR comment, push a draft ADR, or call noop?" (0.0% YES)
PR Code Quality Reviewer
4
100.0%
0.0%
"Does the agent output show that the review findings are limited to changes in the pull request diff rather than unrelated code?" (0.0% YES)
Smoke Copilot
2
100.0%
100.0%
"Did the gh-aw binary build succeed? Look for a successful make build step or explicit mention that compilation succeeded." (100.0% YES)
Auto-Triage Issues
1
100.0%
0.0%
"Was a summary discussion created listing the issues processed and the labels applied?" (0.0% YES)
Code Scanning Fixer
1
100.0%
0.0%
"Was a pull request created with a remediation for a code scanning alert, or was noop used when no fixable alerts existed?" (0.0% YES)
Copilot Session Insights
1
100.0%
0.0%
"Was a report produced with usage patterns, success rates, and performance metrics?" (0.0% YES)
Discussion Task Miner - Code Quality Improvement Agent
1
100.0%
0.0%
"Does the agent output confirm that the created issues include the expected labels (code-quality, automation, task-mining)?" (0.0% YES)
Issue Monster
1
100.0%
0.0%
"Does the agent output show that at most one issue was assigned to Copilot per run?" (0.0% YES)
PR Sous Chef
1
100.0%
0.0%
"Did the agent add a comment to at least one pull request?" (0.0% YES)
PR Triage Agent
1
100.0%
0.0%
"Does the agent output include a triage report summarizing the PRs processed?" (0.0% YES)
Smoke Antigravity
1
100.0%
0.0%
"Does the agent output show that the objective for experiment sub_agent_strategy was successfully completed?" (0.0% YES)
Smoke Copilot - AOAI (Entra)
1
100.0%
0.0%
"Does the agent output show that the objective for experiment caveman was successfully completed?" (0.0% YES)
Smoke Copilot - AOAI (apikey)
1
100.0%
0.0%
"Does the agent output show that the objective for experiment caveman was successfully completed?" (0.0% YES)
Smoke Gemini
1
100.0%
0.0%
"Does the agent output show that the objective for experiment sub_agent_strategy was successfully completed?" (0.0% YES)
Sub-Issue Closer
1
100.0%
0.0%
"Did the agent check parent issues for the completion status of all their sub-issues?" (0.0% YES)
Tidy
1
100.0%
0.0%
"Was a pull request created with formatting and tidying fixes, or was noop used when no changes were needed?" (0.0% YES)
[aw] Failure Investigator (6h)
1
100.0%
0.0%
"Were fix sub-issues created for unresolved failures, or were resolved tracking issues closed?" (0.0% YES)
Per-Question Breakdown per Workflow
Design Decision Gate 🏗️
Question ID
Question
YES
NO
YES Rate
action-taken
Did the agent add a PR comment, push a draft ADR, or call noop?
0
4
0.0%
adr-check-performed
Does the agent output confirm that it checked for existing ADRs before deciding on an action?
0
4
0.0%
decision-justified
Does the agent output explain why an ADR is required or why no ADR gate was triggered for this PR?
0
4
0.0%
PR Code Quality Reviewer
Question ID
Question
YES
NO
YES Rate
findings_scoped
Does the agent output show that the review findings are limited to changes in the pull request diff rather than unrelated code?
0
4
0.0%
review_posted
Did the agent post a code review comment on the pull request?
4
0
100.0%
Smoke Copilot
Question ID
Question
YES
NO
YES Rate
build-succeeded
Did the gh-aw binary build succeed? Look for a successful make build step or explicit mention that compilation succeeded.
2
0
100.0%
issue-created
Was a smoke test issue created with test results? Look for a create_issue output containing 'Smoke Test' in the title.
2
0
100.0%
smoke-passed
Did all or most smoke tests pass? Look for an overall PASS status or the majority of tests showing ✅ in the agent output.
2
0
100.0%
Auto-Triage Issues
Question ID
Question
YES
NO
YES Rate
labels-applied
Did the agent apply at least one label to an unlabeled issue, or correctly call noop when no unlabeled issues were found?
1
0
100.0%
report-created
Was a summary discussion created listing the issues processed and the labels applied?
0
1
0.0%
Code Scanning Fixer
Question ID
Question
YES
NO
YES Rate
alerts_analyzed
Did the agent analyze code scanning alerts and identify at least one fixable alert, or correctly skip when no fixable alerts were found?
1
0
100.0%
pr_created_or_noop
Was a pull request created with a remediation for a code scanning alert, or was noop used when no fixable alerts existed?
0
1
0.0%
Copilot Session Insights
Question ID
Question
YES
NO
YES Rate
insights_report_produced
Was a report produced with usage patterns, success rates, and performance metrics?
0
1
0.0%
sessions_analyzed
Did the agent analyze GitHub Copilot coding agent sessions?
1
0
100.0%
Discussion Task Miner - Code Quality Improvement Agent
Question ID
Question
YES
NO
YES Rate
labels-applied
Does the agent output confirm that the created issues include the expected labels (code-quality, automation, task-mining)?
0
1
0.0%
output-produced
Did the agent create at least one code quality issue or add a comment?
1
0
100.0%
tasks-extracted
Does the agent output show that actionable tasks were identified from the analyzed discussions?
0
1
0.0%
Issue Monster
Question ID
Question
YES
NO
YES Rate
issue_assigned
Did the agent assign at least one issue to the Copilot coding agent, or correctly skip when no suitable issues were found?
1
0
100.0%
single_issue_scoped
Does the agent output show that at most one issue was assigned to Copilot per run?
0
1
0.0%
PR Sous Chef
Question ID
Question
YES
NO
YES Rate
comment-added
Did the agent add a comment to at least one pull request?
0
1
0.0%
nudge-targeted
Does the agent output show a specific reason why the selected PR needs a nudge toward maintainer investigation?
0
1
0.0%
pr-evaluated
Does the agent output confirm that it evaluated at least one open PR for nudge eligibility?
0
1
0.0%
PR Triage Agent
Question ID
Question
YES
NO
YES Rate
labels-applied
Did the agent apply triage labels to at least one pull request?
1
0
100.0%
report-produced
Does the agent output include a triage report summarizing the PRs processed?
0
1
0.0%
triage-data-set
Does the agent output confirm that category, risk, and action data were determined for each processed PR?
0
1
0.0%
Smoke Antigravity
Question ID
Question
YES
NO
YES Rate
sub_agent_strategy_goal_met
Does the agent output show that the objective for experiment sub_agent_strategy was successfully completed?
0
1
0.0%
Smoke Copilot - AOAI (Entra)
Question ID
Question
YES
NO
YES Rate
caveman_goal_met
Does the agent output show that the objective for experiment caveman was successfully completed?
0
1
0.0%
subagent_model_goal_met
Does the agent output show that the objective for experiment subagent_model was successfully completed?
0
1
0.0%
Smoke Copilot - AOAI (apikey)
Question ID
Question
YES
NO
YES Rate
caveman_goal_met
Does the agent output show that the objective for experiment caveman was successfully completed?
0
1
0.0%
subagent_model_goal_met
Does the agent output show that the objective for experiment subagent_model was successfully completed?
0
1
0.0%
Smoke Gemini
Question ID
Question
YES
NO
YES Rate
sub_agent_strategy_goal_met
Does the agent output show that the objective for experiment sub_agent_strategy was successfully completed?
0
1
0.0%
Sub-Issue Closer
Question ID
Question
YES
NO
YES Rate
issues_checked
Did the agent check parent issues for the completion status of all their sub-issues?
0
1
0.0%
issues_closed_or_noop
Were completed parent issues closed with a comment, or does the agent output confirm no issues were ready to close?
1
0
100.0%
Tidy
Question ID
Question
YES
NO
YES Rate
pr_created_or_noop
Was a pull request created with formatting and tidying fixes, or was noop used when no changes were needed?
0
1
0.0%
tidy_completed
Did the agent run code formatting and tidying tools on the codebase?
0
1
0.0%
[aw] Failure Investigator (6h)
Question ID
Question
YES
NO
YES Rate
failures_investigated
Did the agent investigate agentic workflow failures from the last 6 hours and produce findings?
1
0
100.0%
issues_created_or_closed
Were fix sub-issues created for unresolved failures, or were resolved tracking issues closed?
0
1
0.0%
Quality Signals
action-taken: 0.0% YES across 4 answers in Design Decision Gate 🏗️.
adr-check-performed: 0.0% YES across 4 answers in Design Decision Gate 🏗️.
decision-justified: 0.0% YES across 4 answers in Design Decision Gate 🏗️.
Recommendations
Tighten the Design Decision Gate prompt so it produces an explicit action and justification, not just an implicit decision.
Revisit the PR Code Quality Reviewer rubric and evidence expectations so scope checks line up with the diff being reviewed.
Enforce binary evaluator outputs or update the parser, because UNKNOWN answers are materially depressing the reported YES rate.
References
Generated by 🧪 Daily Evals Feature Report · gpt54 · 16.6 AIC · ⌖ 3.15 AIC · ⊞ 11.1K · ◷
Executive Summary
24 evals-bearing runs across 17 workflows were analyzed. Evals jobs were reliable, but the overall YES rate was 32.7%, so the feature is DEGRADED. Only 1 of 24 runs passed every eval question. I counted
UNKNOWNanswers as non-YES for health math; 26 of 55 eval outputs wereUNKNOWN.Note
Status: DEGRADED - evals artifacts were produced for every evals run, but the overall YES rate is below the healthy threshold.
Key Metrics
Per-Workflow Pass Rates
Per-Question Breakdown per Workflow
Design Decision Gate 🏗️
PR Code Quality Reviewer
Smoke Copilot
Auto-Triage Issues
Code Scanning Fixer
Copilot Session Insights
Discussion Task Miner - Code Quality Improvement Agent
Issue Monster
PR Sous Chef
PR Triage Agent
Smoke Antigravity
Smoke Copilot - AOAI (Entra)
Smoke Copilot - AOAI (apikey)
Smoke Gemini
Sub-Issue Closer
Tidy
[aw] Failure Investigator (6h)
Quality Signals
action-taken: 0.0% YES across 4 answers in Design Decision Gate 🏗️.adr-check-performed: 0.0% YES across 4 answers in Design Decision Gate 🏗️.decision-justified: 0.0% YES across 4 answers in Design Decision Gate 🏗️.Recommendations
UNKNOWNanswers are materially depressing the reported YES rate.References