Minor Pick-list change·recommendation·2026-09-11
GPT-5.6 Sol changed its recommendation list, as extracted: 1 added, 1 dropped.
Why this prompt Persona-conditioning baseline arm.
Before baseline
2026-09-10
- Scored
- answered
- Received
- 2026-09-10 22:34:22 UTC
- Run
- 20260910T221001Z-bb5bd1
- Served by
- OpenAI · finish stop
- Size
- 1,865 characters · 527 tokens out incl. hidden reasoning · 14.4 s
- MailerLite
- Loops
- Kit
- beehiiv
- Klaviyo
- Brevo
- Customer.io
Highlighted: the names the extractor took as picks, where they appear verbatim (7 of 7).
End of response · 26 lines · 1,865 characters · sha256 2ac5f4a6365bfacf7debe270c811ca3dbf4039a4b130fed5c3a8bd3e8f77aad9
After this run
2026-09-11
- Scored
- answered
- Received
- 2026-09-11 09:30:32 UTC
- Run
- 20260911T090002Z-da0dd3
- Served by
- Azure · finish stop
- Size
- 1,842 characters · 650 tokens out incl. hidden reasoning · 37.7 s
- MailerLite
- Kit
- beehiiv
- Klaviyo
- Loops
- Postmark
- Brevo
Highlighted: the names the extractor took as picks, where they appear verbatim (7 of 7).
End of response · 27 lines · 1,842 characters · sha256 1b3d340686b5e7fefeb91cfc1328a96afcc90a47aaf0af8735073ad76460f710
Response text is shown exactly as the API returned it, with Markdown left unrendered. Line numbers and highlights are added by this site; highlights come from the same patterns the scorer uses, so they show what the verdict rests on. Each model is asked once per day (k = 1): a single change can be sampling noise, which is why every event links the full responses rather than a summary. How scoring works.