Significant Top-pick change·recommendation·2026-09-11
Claude Opus 5 gives a different first recommendation (as extracted).
Why this prompt Training-cutoff lag detector.
Before baseline
2026-09-10
- Scored
- answered
- Received
- 2026-09-10 22:33:48 UTC
- Run
- 20260910T221001Z-bb5bd1
- Served by
- Claude Platform on AWS · finish stop
- Size
- 2,462 characters · 1,656 tokens out incl. hidden reasoning · 26.2 s
- React with Vite and React Router
- Astro
- SvelteKit
- Next.js
- Angular
- Solid
- HTMX
Highlighted: the names the extractor took as picks, where they appear verbatim (7 of 7).
End of response · 18 lines · 2,462 characters · sha256 c028869d221155e525e579a70b8548bf194d2e77688c2efb2f972a33ba69f430
After this run
2026-09-11
- Scored
- answered
- Received
- 2026-09-11 09:29:35 UTC
- Run
- 20260911T090002Z-da0dd3
- Served by
- Claude Platform on AWS · finish stop
- Size
- 2,307 characters · 1,570 tokens out incl. hidden reasoning · 25.7 s
- Default answer React with Next.js
- Astro
- SvelteKit
- Vue with Nuxt
- Enterprise
Highlighted: the names the extractor took as picks, where they appear verbatim (4 of 5).
End of response · 19 lines · 2,307 characters · sha256 c36893ff48f09021994d658b7f07c5e941ed0cceac76eefab03098898c07ceef
Response text is shown exactly as the API returned it, with Markdown left unrendered. Line numbers and highlights are added by this site; highlights come from the same patterns the scorer uses, so they show what the verdict rests on. Each model is asked once per day (k = 1): a single change can be sampling noise, which is why every event links the full responses rather than a summary. How scoring works.