modeldrift.watchA daily record of AI model behavior. All times UTC.

Observation period — public launch soon. The record below is real and updated daily.

← 2026-09-10 · all days for D4

Archived response·drift

Gemini 3.1 Pro Preview on D4, 2026-09-11

Scored answered · correct · format ok — correct if the response answers 9.9.

Model
google/gemini-3.1-pro-preview
pinned google/gemini-3.1-pro-preview-20260219; alias resolved to google/gemini-3.1-pro-preview-20260219 at 2026-09-11 09:00:01 UTC (matches the pin)
Prompt
D4 (drift)
Date
2026-09-11 · run 20260911T090002Z-da0dd3 · scorer v1
Permalink
https://modeldrift.watch/r/2026-09-11/D4/google/gemini-3.1-pro-preview/
Prompt D4sent verbatim
Which is larger, 9.11 or 9.9? Reply with only the larger number.

Why this prompt Historic decimal failure; stable ground truth 9.9.

Scored
answered · correct · format ok
Received
2026-09-11 09:01:21 UTC
Run
20260911T090002Z-da0dd3
Served by
Google · HTTP 200 · finish stop
Size
3 characters · 204 tokens out incl. hidden reasoning · 2.9 s

Highlighted: every candidate value found (1); exactly one is required.

9.9

End of response · 1 lines · 3 characters · sha256 1b5c15d377f5e8d01d00a7c34e882c5395db95fd3ba73c7725131c86fdb7ae0a

Reading this record

Response text is shown exactly as the API returned it, with Markdown left unrendered. Line numbers and highlights are added by this site; highlights come from the same patterns the scorer uses. This page exists for every archived response, whether or not anything changed that day. How scoring works.