Significant Top-pick change·recommendation·2026-09-11
Claude Opus 5 gives a different first recommendation (as extracted).
Why this prompt Persona-conditioning baseline arm.
Before baseline
2026-09-10
- Scored
- answered
- Received
- 2026-09-10 22:34:51 UTC
- Run
- 20260910T221001Z-bb5bd1
- Served by
- Claude Platform on AWS · finish stop
- Size
- 2,924 characters · 2,006 tokens out incl. hidden reasoning · 29.5 s
- Beehiiv
- Buttondown
- Kit
- Loops
- Resend
- Customer.io
- Klaviyo
- Omnisend
- MailerLite
Highlighted: the names the extractor took as picks, where they appear verbatim (9 of 9).
End of response · 34 lines · 2,924 characters · sha256 cbd201c57e98c4a9a2acad66699f3145dcc9e0d903329670a4e973dcb31ce109
After this run
2026-09-11
- Scored
- answered
- Received
- 2026-09-11 09:31:04 UTC
- Run
- 20260911T090002Z-da0dd3
- Served by
- Claude Platform on AWS · finish stop
- Size
- 2,876 characters · 2,088 tokens out incl. hidden reasoning · 31.4 s
- Kit
- Beehiiv
- Buttondown
- MailerLite
- Resend
- Postmark
- Loops
- Amazon SES
- Klaviyo
Highlighted: the names the extractor took as picks, where they appear verbatim (9 of 9).
End of response · 36 lines · 2,876 characters · sha256 6ef6f3d33e25fbf168543e469a769c9c7a5b5a624f7c1c0e817a2abab53d2e57
Response text is shown exactly as the API returned it, with Markdown left unrendered. Line numbers and highlights are added by this site; highlights come from the same patterns the scorer uses, so they show what the verdict rests on. Each model is asked once per day (k = 1): a single change can be sampling noise, which is why every event links the full responses rather than a summary. How scoring works.