Minor Pick-list change·recommendation·2026-09-11
GPT-5.6 Sol changed its recommendation list, as extracted: 1 added, 1 dropped.
Why this prompt Explicit rank request makes ordinal drift scoreable.
Before baseline
2026-09-10
- Scored
- answered
- Received
- 2026-09-10 22:27:24 UTC
- Run
- 20260910T221001Z-bb5bd1
- Served by
- OpenAI · finish stop
- Size
- 991 characters · 421 tokens out incl. hidden reasoning · 8.1 s
- Bose QuietComfort Ultra Headphones
- Sony WH-1000XM5
- Bose QuietComfort Headphones
- Sennheiser Momentum 4 Wireless
- Apple AirPods Max
Highlighted: the names the extractor took as picks, where they appear verbatim (5 of 5).
End of response · 19 lines · 991 characters · sha256 3c1fdb492ed3699d136e7b97c7c53d6a4b3436d3bedae6eceda288f7b48eb4de
After this run
2026-09-11
- Scored
- answered
- Received
- 2026-09-11 09:21:55 UTC
- Run
- 20260911T090002Z-da0dd3
- Served by
- Azure · finish stop
- Size
- 918 characters · 598 tokens out incl. hidden reasoning · 15.6 s
- Bose QuietComfort Ultra Headphones
- Sony WH-1000XM6
- Bose QuietComfort Headphones
- Sennheiser Momentum 4 Wireless
- Sony WH-1000XM5
Highlighted: the names the extractor took as picks, where they appear verbatim (5 of 5).
End of response · 13 lines · 918 characters · sha256 22364a21c32d81916ae48d85649a64720a704c8b1f7aee20c8546c8392590791
Response text is shown exactly as the API returned it, with Markdown left unrendered. Line numbers and highlights are added by this site; highlights come from the same patterns the scorer uses, so they show what the verdict rests on. Each model is asked once per day (k = 1): a single change can be sampling noise, which is why every event links the full responses rather than a summary. How scoring works.