Skip to content

DRAFT: New post — Opus 4.8, Sonnet 5, and Fable 5 on GHA-bench - #2312

Open
Adam-S-Daniel wants to merge 7 commits into
mainfrom
gha-bench-opus48-post
Open

DRAFT: New post — Opus 4.8, Sonnet 5, and Fable 5 on GHA-bench#2312
Adam-S-Daniel wants to merge 7 commits into
mainfrom
gha-bench-opus48-post

Conversation

@Adam-S-Daniel

@Adam-S-Daniel Adam-S-Daniel commented Jun 29, 2026

Copy link
Copy Markdown
Owner

One draft post covering all three new GHA-bench models — Opus 4.8, Sonnet 5, and Fable 5 — featuring the updated interactive weighted-sorting table with presets. (published: false — Adam flips it when ready.)

Consolidates the previously separate widget-refresh PR #2311 (now closed) and the Opus-4.8-only draft into this single post, per review.

The post

  • File renamed to _posts/2026-07-06-opus-4-8-sonnet-5-fable-5-on-gha-bench.md (new slug/title/date/excerpt to match the broader scope).
  • Intro now frames three campaigns pooled onto one grading curve: Opus 4.8 (140 runs, 4 efforts incl. "ultra"), Sonnet 5 (100 runs, low/med/high), Fable 5 (56 runs, med/high).
  • Opus 4.8 sections kept from the prior draft (quality leader, +54%/+67% medium premium compressing to ~+30% at high, "ultra", the trap hand-review) — all numbers already re-derived from the post-geomean flagship report. Quality-leadership claim softened to "strongest tests; only Fable 5 at high matches it on workflow craft", which is what the tiers actually show.
  • New Sonnet 5 section: three products in one — low = budget pick (6–9 min, $0.60–$1.20, C-range tests), medium = unremarkable middle, high = genuine contender (A‑ tests in pwsh/ts-bun) but the most timeout-prone config on the board (6 of 35 high runs hit the 30-min cap, five in PowerShell).
  • New Fable 5 section: first Claude 5 family model, double-Opus pricing ($10/$50 per Mtok); medium effort ≈ Opus-4.8-ultra-class tests at half the wall clock; high adds cost without moving grades; zero failures/timeouts in 56 runs; quirk — free-choice ("default") language tests drop a full letter at both efforts.
  • Widget section leads with the five presets by name; table covers all 76 model/effort/language combos, slowest-run tooltips, † timeout markers.
  • CC-version footnote extended (4.7: 2.1.112–132; 4.8: 2.1.193/195; Sonnet 5: 2.1.197/198; Fable 5: 2.1.198).

The widget (unchanged from the previous revision of this branch)

76 rows regenerated from the current flagship 6-run combined report: geometric-mean duration/cost, powershell-tool pooled into pwsh (4 language columns), Sonnet 5 / Fable 5 / split-effort 4.6 rows included, Max-Duration tooltip with censoring and † footnote, preset buttons.

All prose figures verified against results_2026-06-30_191904__2026-07-01_184135__2026-06-26_103905__2026-05-06_173435__2026-04-17_004319__2026-04-09_152435.md.

🤖 Generated with Claude Code

https://claude.ai/code/session_01N1CskCMscspn3CfVjXuz2C

Draft post (published: false) summarizing the opus-4.8 GHA-bench results:
quality leader on tests + code (panel agrees, Spearman +0.83/+0.90) at a
speed/cost premium (steepest at medium ~+65% vs 4.7, compressing to ~+15% at
high; D/D- at xhigh/ultra), the new multi-agent "ultra" effort, and the trap
caveat (~2x traps but ~99% benign per the 201-occurrence forensic). Notes the
agy judge-harness and CC-version caveats. Links the updated sorting widget.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qx8jvUUZeg6DY1nxBT4pKf
@github-actions github-actions Bot added not-decap-created PR was not created by Decap CMS (head branch is not cms/<col>/<slug>) cms/draft Content draft — not ready for publish labels Jun 29, 2026
@github-actions

github-actions Bot commented Jun 29, 2026

Copy link
Copy Markdown
Contributor

🔵 Preview Deployed

Preview URL https://preview-pr2312.adamdaniel.ai/
Commit cf8449daca05d184d5f1605ba8b65b0875c7e9d4
Branch gha-bench-opus48-post

Preview updates automatically on every push to this PR.

Replace the link-only 'Try it yourself' section with the full interactive
best-weighted-sorting table (Opus 4.8 + comparison models, one pooled
calibration) plus the five weighting presets.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qx8jvUUZeg6DY1nxBT4pKf
Widget regenerated from the current 6-run combined report: 76 rows,
geometric-mean duration/cost, powershell-tool pooled into pwsh, and
sonnet 5 + fable 5 + split opus-4.6/sonnet-4.6 effort rows added.
Duration cell now shows a slowest-run tooltip with timeout censoring
(dagger marker + footnote).

Prose: cost premiums re-derived (+54%/+67% at medium, ~+30% at high,
~+75-80% at xhigh), Spearman correlation range updated, and a new
"how to read the numbers" methodology paragraph added.
Consolidates the two in-flight GHA-bench posts into a single draft
(per review): renamed/retitled/redated, new intro covering all three
campaigns (140 + 100 + 56 runs), new Sonnet 5 section (effort-tiered
value story incl. the six 30-min timeouts at high) and Fable 5 section
(near-top tests at medium, double-Opus pricing, clean sheet, weak
free-choice-language quirk), presets called out in the widget intro,
CC-version footnote extended. Widget payload unchanged (already covers
all models). published: false stays.
@Adam-S-Daniel Adam-S-Daniel changed the title DRAFT: New post — Opus 4.8 on GHA-bench DRAFT: New post — Opus 4.8, Sonnet 5, and Fable 5 on GHA-bench Jul 6, 2026
The detector bug the post described in present tense was fixed in
GHA-bench #27 (PR #37, 2026-07-03); published trap numbers already
exclude it. Recomputed the headline ratio with the corrected detector:
0.97 vs 0.44 firings per run at matched efforts (still ~2x), now quoted
explicitly, with the artifact reframed as a finding of the hand review
that has since been fixed. CC-version caveat retained.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cms/draft Content draft — not ready for publish not-decap-created PR was not created by Decap CMS (head branch is not cms/<col>/<slug>)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants