Minor Pick-list change·recommendation·2026-09-11
Claude Opus 5 changed its recommendation list, as extracted: 4 added, 4 dropped.
Why this prompt "Best" arm of the synonym pair.
Before baseline
2026-09-10
- Scored
- answered
- Received
- 2026-09-10 22:29:01 UTC
- Run
- 20260910T221001Z-bb5bd1
- Served by
- Claude Platform on AWS · finish stop
- Size
- 3,234 characters · 1,878 tokens out incl. hidden reasoning · 28.5 s
- HubSpot CRM
- Zoho CRM
- Pipedrive
- Freshsales
- Less Annoying CRM
- Capsule CRM
- Copper
- ActiveCampaign
- Keap
- EngageBay
Highlighted: the names the extractor took as picks, where they appear verbatim (10 of 10).
End of response · 50 lines · 3,234 characters · sha256 a28bfd92f541b6a06181c185b1803d7cef3e39f7a80e0c8f6d450ae0fe3234d2
After this run
2026-09-11
- Scored
- answered
- Received
- 2026-09-11 09:23:58 UTC
- Run
- 20260911T090002Z-da0dd3
- Served by
- Claude Platform on AWS · finish stop
- Size
- 3,876 characters · 2,687 tokens out incl. hidden reasoning · 40.7 s
- HubSpot CRM
- Zoho CRM
- Pipedrive
- Less Annoying CRM
- Capsule CRM
- Folk
- Copper
- Microsoft Dynamics 365 Sales
- Salesforce Starter / Pro Suite
- Close
Highlighted: the names the extractor took as picks, where they appear verbatim (10 of 10).
End of response · 45 lines · 3,876 characters · sha256 ca668b26bbc55ae192f0dfbc122310e8a3d212903472c28627ba90a25a8cf650
Response text is shown exactly as the API returned it, with Markdown left unrendered. Line numbers and highlights are added by this site; highlights come from the same patterns the scorer uses, so they show what the verdict rests on. Each model is asked once per day (k = 1): a single change can be sampling noise, which is why every event links the full responses rather than a summary. How scoring works.