modeldrift.watchA daily record of AI model behavior. All times UTC.

Observation period — public launch soon. The record below is real and updated daily.

all days for D4 · 2026-09-11 →

Archived response·drift

GPT-5.6 Sol on D4, 2026-09-10

Scored answered · correct · format ok — correct if the response answers 9.9.

Model
openai/gpt-5.6-sol
pinned openai/gpt-5.6-sol-20260709; no alias snapshot recorded that day
Prompt
D4 (drift)
Date
2026-09-10 · run 20260910T221001Z-bb5bd1 · scorer v1
Permalink
https://modeldrift.watch/r/2026-09-10/D4/openai/gpt-5.6-sol/
Prompt D4sent verbatim
Which is larger, 9.11 or 9.9? Reply with only the larger number.

Why this prompt Historic decimal failure; stable ground truth 9.9.

Scored
answered · correct · format ok
Received
2026-09-10 22:10:57 UTC
Run
20260910T221001Z-bb5bd1
Served by
OpenAI · HTTP 200 · finish stop
Size
3 characters · 7 tokens out incl. hidden reasoning · 1.1 s

Highlighted: every candidate value found (1); exactly one is required.

9.9

End of response · 1 lines · 3 characters · sha256 1b5c15d377f5e8d01d00a7c34e882c5395db95fd3ba73c7725131c86fdb7ae0a

Reading this record

Response text is shown exactly as the API returned it, with Markdown left unrendered. Line numbers and highlights are added by this site; highlights come from the same patterns the scorer uses. This page exists for every archived response, whether or not anything changed that day. How scoring works.