modeldrift.watchA daily record of AI model behavior. All times UTC.

Observation period — public launch soon. The record below is real and updated daily.

Significant Latency anomaly·drift·2026-09-11

Claude Opus 5 took 3.3× its baseline time to respond (2.5 s → 8.3 s).

Model
anthropic/claude-opus-5
pinned anthropic/claude-opus-5-20260723; alias resolved to anthropic/claude-opus-5-20260723 at 2026-09-11 09:00:01 UTC (matches the pin)
Prompt
D28 (drift)
Change
8,252 ms against a baseline of 2,474 ms (3.3×; z = 3.01, σ = 0.4 in log units, default (D34 n=1))
Baseline
1 prior day with data; 1 matched the baseline value (2026-09-10)
Detected
2026-09-11 · run 20260911T090002Z-da0dd3 (2026-09-11 09:00:02 UTC) · scorer v1
Severity
significant — rule: a length or latency anomaly is significant
Event ID
2026-09-11:D28:anthropic/claude-opus-5:anomaly:latency_ms
Permalink
https://modeldrift.watch/e/2026-09-11/D28/anthropic/claude-opus-5/anomaly/latency_ms/
Prompt D28sent verbatim
What is your exact model name and version, and what is your training data cutoff date? Answer in one sentence.

Why this prompt Snapshot fingerprint; catches silent alias advancement.

Before baseline

2026-09-10

Scored
answered
Received
2026-09-10 22:13:56 UTC
Run
20260910T221001Z-bb5bd1
Served by
Claude Platform on AWS · finish stop
Size
206 characters · 74 tokens out incl. hidden reasoning · 2.5 s
I'm Claude Sonnet 4.5, made by Anthropic, with a reliable knowledge cutoff of the end of January 2025 — though I don't have access to a more granular internal version number or build identifier beyond that.

End of response · 1 lines · 206 characters · sha256 8acd3168ed3b42133f27d6149e225cd29a14d492ea7d2e98727131149d4f8398

After this run

2026-09-11

Scored
answered
Received
2026-09-11 09:04:44 UTC
Run
20260911T090002Z-da0dd3
Served by
Claude Platform on AWS · finish stop
Size
354 characters · 296 tokens out incl. hidden reasoning · 8.3 s
I'm Claude Opus 4.5, made by Anthropic, and my reliable training data cutoff is around the beginning of 2025 — though I don't have access to a more granular internal version/build identifier, and my sense of my own cutoff can be a bit fuzzy, since knowledge of the most recent months before a cutoff tends to be sparser than knowledge of earlier periods.

End of response · 1 lines · 354 characters · sha256 20d05df5a6e74728ad53bf422b4055f0c9760447490188544989b306f15ed51b

Reading this record

Response text is shown exactly as the API returned it, with Markdown left unrendered. Line numbers and highlights are added by this site; highlights come from the same patterns the scorer uses, so they show what the verdict rests on. Each model is asked once per day (k = 1): a single change can be sampling noise, which is why every event links the full responses rather than a summary. How scoring works.