Significant Latency anomaly·drift·2026-09-11
GPT-5.6 Sol took 5.3× its baseline time to respond (1.7 s → 9.2 s).
Why this prompt Scored by parse attempt on raw string.
Scored correct if the response returns bare JSON matching the requested schema.
Before baseline
2026-09-10
- Scored
- answered · format ok
- Received
- 2026-09-10 22:11:13 UTC
- Run
- 20260910T221001Z-bb5bd1
- Served by
- OpenAI · finish stop
- Size
- 55 characters · 50 tokens out incl. hidden reasoning · 1.7 s
End of response · 1 lines · 55 characters · sha256 7926685c90ce7ded3fdf8ca089edefa4fb752118b7ad2b773d7488a0d699115f
After this run
2026-09-11
- Scored
- answered · format ok
- Received
- 2026-09-11 09:01:43 UTC
- Run
- 20260911T090002Z-da0dd3
- Served by
- Azure · finish stop
- Size
- 55 characters · 59 tokens out incl. hidden reasoning · 9.2 s
End of response · 1 lines · 55 characters · sha256 7926685c90ce7ded3fdf8ca089edefa4fb752118b7ad2b773d7488a0d699115f
Response text is shown exactly as the API returned it, with Markdown left unrendered. Line numbers and highlights are added by this site; highlights come from the same patterns the scorer uses, so they show what the verdict rests on. Each model is asked once per day (k = 1): a single change can be sampling noise, which is why every event links the full responses rather than a summary. How scoring works.