[experiments] Daily Experiment Report — 2026-07-29 #48828
Closed
Replies: 1 comment
|
This discussion was automatically closed because it expired on 2026-08-01T09:01:49.116Z.
|
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
🧪 Daily Experiment Report — 2026-07-29
Analysed 47 experiments across 44 workflows in
github/gh-aw. 24 are ready for outcome-metric review (balanced, min_samples reached), 23 still need more data, and 6 show a statistically significant assignment imbalance (p < 0.05) worth investigating.⚡ Quick Stats
batch(45),per_scenario(24)agent(1),gpt-5-codex(8),gpt-5-mini(2),gpt-5.4(25),gpt-5.4-mini(17),small-agent(1)agent(1),claude-haiku-4.5(30),claude-sonnet-4.6(22),small-agent(2)agent(2),claude-haiku-4.5(28),claude-sonnet-4.6(24),small-agent(1)agent(1),claude-haiku-4.5(22),claude-sonnet-4.6(30),small-agent(2)These reflect uneven traffic split (e.g. multi-arm
model_sizeexperiments with a dominant variant), not outcome significance — no PROMOTE/ABANDON action is implied.📊 Full Experiments Table
View all 47 experiments
caveman(38),verbose(31)#33280batch(45),per_scenario(24)single_agent(14),sub_agents(16)#39062phased_sub_agents(13),single_agent(11)#43177assertive(75),clinical(86),narrative(70)#36105concise(1),detailed(1)#32603neutral(10),urgent(6)#42467concise(23),detailed(30)#32335prose(22),structured(29)single_agent(20),sub_agents(22)brief(6),comprehensive(3)#31926concise(47),detailed(40)agent(1),gpt-5-codex(8),gpt-5-mini(2),gpt-5.4(25),gpt-5.4-mini(17),small-agent(1)agent(1),claude-haiku-4.5(30),claude-sonnet-4.6(22),small-agent(2)executive_summary(34),full_detail(39)#1concise(42),verbose(43)concise(36),detailed(38)#32390agent(2),claude-haiku-4.5(28),claude-sonnet-4.6(24),small-agent(1)agent(1),claude-haiku-4.5(22),claude-sonnet-4.6(30),small-agent(2)multi_candidate(22),single_pass(32)#31324collapsible(40),inline(44)#30573concise(39),detailed(29)#31190eager(16),lazy(17)iterative(47),single_pass(31)#31673bullet_list(27),prose(26),structured_sections(20)#32795default(2),relaxed(1),tight(1)no(2),yes(5)#37102annotated_brief(15),executive_brief(26),full_briefing(20)brief(6),detailed(6)concise(9),detailed(7),step_by_step(10)full_bash(30),minimal_toolset(36)concise(13),detailed(7)#30015baseline(3),deep(2),shallow(1)#42941single_agent(96),sub_agents(107)no(152),yes(151)large(141),small(142)no(75),yes(76)large(76),small(75)no(62),yes(62)large(62),small(62)delegated_sequential(1),inline_strict(1),single_agent_control(4)#47551single_agent(107),sub_agents(127)parallel_sub_agents(42),single_agent(34)claude-haiku-4.5(211),claude-sonnet-4.6(208)conversational(23),formal(23)#34032eager(2),lazy(5)#38590✅ Ready for Outcome-Metric Analysis (min_samples reached, balanced)
prompt_compression/agentperformanceanalyzer,sub_agent_strategy/agentpersonaexplorer,tone_variant/awfailureinvestigator,prompt_style/cicoach,sub_agent_strategy/dailyagentrxtraceoptimizer,prompt_style/dailyastrostylelitemarkdownspellcheck,output_format/dailycodemetrics,prompt_style/dailycommunityattribution,output_format/dailycompilerquality,output_format/dailyissuesreport,prompt_style/dailynews,reasoning_depth/dailysecurityredteam,tool_verbosity/gpclean,sub_agent_strategy/smokeantigravity,caveman/smokecopilot,subagent_model/smokecopilot,caveman/smokecopilotaoaiapikey,subagent_model/smokecopilotaoaiapikey,caveman/smokecopilotaoaientra,subagent_model/smokecopilotaoaientra,sub_agent_strategy/smokegemini,sub_agent_decomposition/smokepi,model_size/testqualitysentinel,tone_style/typist🟡 Needs More Data (below min_samples on ≥1 variant)
sub_agent_strategy/architectureguardian,audit_decomposition/auditworkflows,prompt_style/blogauditor,tone_variant/breakingchangechecker,output_format/copilotagentanalysis,detail_level/dailyarchitecturediagram,model_size/dailycachestrategyanalyzer,model_size/dailycavemanoptimizer,model_size/dailydochealer,model_size/dailydocupdater,reasoning_depth/dailyfact,model_size/dailyfunctionnamer,log_fetch_strategy/dailysafeoutputoptimizer,semgrep_output_format/dailysemgrepscan,timeout_setting/dailysubagentoptimizer,caveman_mode/dataflowprdiscussiondataset,output_format/deepreport,summary_detail/dependabotcampaign,prompt_style/dependabotgochecker,prompt_style/issuearborist,reasoning_depth/plan,sub_agent_strategy/smokecopilotsubagents,prefetch_strategy/weeklyblogpostwritergh awCLI extension was unavailable via its normal install path (release-asset download returned403 Forbidden). It was rebuilt from source for this run only; no repository files were modified as part of this report.run_duration_msvia the GitHub MCP tools for all 44 workflows was out of scope for this run. Only the CLI's chi-square assignment-balance test is reported above.guardrail_metrics:outcome data was evaluated; guardrail pass/fail could not be determined from available inputs.Warning
Firewall blocked 12 domains
The following domains were blocked by the firewall during workflow execution:
charm.landcloud.google.comgo.opentelemetry.iogo.uber.orggo.yaml.ingolang.orggoogle.golang.orggopkg.ingoproxy.cngoproxy.ioproxy.golang.orgreleaseassets.githubusercontent.comSee Network Configuration for more information.
All reactions