NikoTakSecuring the Web, One Threat at a Time.

Same Name, Different Model

I set myself a task: to verify the following about the current state of LLM frontiers. That two announced models can be configurations of one model. That within a week or two a user may be on a different configuration, without any rename. That benchmark results and cross-model comparisons are therefore compromised. And that users are locked in to whatever a few companies choose to disclose, framed as suits them.

To gather the vast data over three years, I set up three agents (Claude Fable 5.1), each a separate instance of the same model, run in parallel with a written brief. The first was to build the release timeline from June 2022 to September 2026 for eight labs, with the exact snapshot or alias identifier for every entry and the primary page it was confirmed from. The second was given the list of incidents (the 2023 drift study, the 2023 "laziness" episode, the GPT-4o sycophancy rollback, the Gemini preview redirect, the Llama 4 arena variant, the GPT-5 router and safety routing, the Anthropic serving bugs, the Grok prompt incidents, the nondeterminism papers, the DeepSeek in-place upgrades) and asked to verify each with dates, a verbatim quotation of at most 40 words, and the URL. The third was asked for the versioning and deprecation policies of each provider, the EU AI Act and Code of Practice text, California SB 53, the Stanford transparency index editions, the model-card paper, the reproducibility studies, and anything from 2026 on version disclosure. All three were restricted to web search and page fetches, told to prefer primary sources (lab announcements, API changelogs, model cards, arXiv, official postmortems, regulation text), and told to mark anything they could not confirm from a primary source as unverified rather than fill the gap. Between them they made 388 tool calls, nearly all searches and page fetches, over 9 to 20 minutes each, and returned about 130,000 to 175,000 tokens of notes apiece.

I then had the same model re-fetch the items with the most riding on them: the Fable 5 redeployment page, the GPT-5.6 August system card, the April 2026 Claude Code postmortem, the "Silent Updates" paper, the Narayanan–Kapoor critique of the drift study, and DeepSeek's V4 notice. The timeline figure and the timeline table are generated from one data file, so a date in the figure is the same date as in the table. The text was drafted by the model from the agents' notes and my brief; I edited it. The configured model identifier for the whole session was claude-fable-5-1. By this post's own argument, that identifier is a claim by the provider about a service on a date, and whether the same model served every call is not something I, or anyone outside Anthropic, can verify.

Four things were dropped for want of a primary source. A widely repeated claim that Anthropic publicly denied "quantizing" Claude in 2025: no statement using that word was found, and only the status-page and postmortem wording quoted below is used. The original documentation wording for chatgpt-4o-latest, which described the alias as continuously updated: it is no longer on any OpenAI page that could be reached, so the current past-tense description is used instead. Release dates for Qwen 3.6 and 3.7: only secondary sources. Gemini 3.5 Pro: announced in May 2026 for "next month", absent from the API model list as of 22 September, so treated as unreleased. Three things are flagged in the text as unconfirmed: the 8 August 2025 date for GPT-4o's restoration (press reports of Altman's posts), the 1 September 2026 date for Fable 5.1 (Anthropic's page says only "September 2026"), and whether a same-name post-training update is a "substantially modified version" under SB 53. Every figure and quotation below is linked to a dated source where it appears.

What follows is the data as it appears.

The premise, tested

Against the record, the four claims come out as follows.

Claim 1: two announced models can be configurations of one model. Supported, and in the clearest cases disclosed by the provider. Anthropic's September 2026 announcement states: "Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards." OpenAI's GPT-5 launch post describes the product as "a unified system with a smart, efficient model … a deeper reasoning model (GPT-5 thinking) … and a real-time router that quickly decides which to use". The reverse case is the less visible one: a single name covering several models (see claim 2).

Claim 2: within a week or two, a user may be on a different configuration without a rename. Supported for the consumer apps at every major provider; for the API it depends on the provider and on whether the caller pinned a snapshot. OpenAI's own account of the April 2025 sycophancy incident says that since GPT-4o's launch in ChatGPT in May 2024 "we've released five major updates focused on changes to personality and helpfulness", and that updates are preceded by "an A/B test with a small number of our users". Google's API changelog for 6 May 2025 records that "gemini-2.5-pro-preview-03-25 will automatically point to the new version of the model". DeepSeek's changelog records at least eight in-place upgrades of the deepseek-chat endpoint between May 2024 and April 2026 (changelog).

Claim 3: benchmark results and cross-model comparisons are compromised. Supported, with one correction. The most-cited drift study (Chen, Zaharia and Zou, July 2023) is often read as proof that GPT-4 "got worse"; the same data is consistent with a change of behaviour rather than a loss of capability, as Narayanan and Kapoor showed when they re-ran the prime-number test with composite numbers. What the study does establish is that the system behind a fixed name changed enough in three months to move four of eight task scores by 30 to 75 points. The leaderboard problem is documented independently: Singh et al. (April 2025) count "27 private LLM variants tested by Meta in the lead-up to the Llama-4 release" on Chatbot Arena, and Abraham and Bucknall (August 2026) find that none of nine API providers lets an outside party verify that the artifact being served is the one that was evaluated.

Claim 4: users are locked in to whatever a few companies choose to disclose, framed as suits them. Partly supported. The providers disclose more than nothing, and some of it cuts against their interest: OpenAI published two postmortems on the sycophancy rollback, Anthropic published a three-bug postmortem in September 2025 and a second one in April 2026, and xAI publishes its system prompts. Against that: no provider guarantees notice of changes to its consumer app; the routing and fallback behaviour of ChatGPT was established by users reading model_slug fields in network responses, not by an announcement; Stanford's transparency index fell from a mean of 58 in 2024 to 41 in 2025; and no external verification exists at all. Open weights are the exit, at a cost in capability, and Meta, the lab behind Llama, shipped its 2026 flagship as a proprietary model.

Seven layers that move under a fixed name

"The model changed" is too coarse a description to be checked. The record shows at least seven distinct layers at which the behaviour of a named product has changed, each with a documented instance. Two of them (the weights and the post-training updates) are what people usually mean by "a new model". The other five are not, and they are where most of the surprises came from.

Where the behaviour of a named product has changed, with one documented instance each 1 Weights new snapshot, same product name gpt-4o 2024-05-13 → 08-06 → 11-20; claude-3-5-sonnet 20240620 → 20241022 2 Post-training personality and helpfulness updates in the app GPT-4o: “five major updates” May 2024 – Apr 2025; 25 Apr 2025 update rolled back on 28 Apr 3 System prompt text prepended to every conversation Grok, 14 May 2025: “unauthorized modification”; Claude Code, 16 Apr 2026: one added line, 3% eval drop 4 Routing, fallback which model answers a given message GPT-5 router (Aug 2025); gpt-5-chat-safety (Sep 2025); mini fallback (Mar 2026); Gemini preview redirects 5 Harness defaults reasoning effort, caching, tool settings Claude Code default effort high → medium, 4 Mar – 7 Apr 2026 6 Safeguard classifiers filters trained separately from the model Fable 5 / Mythos 5 redeployed under the same names with a new classifier, 1 Jul 2026 7 Serving stack load balancing, kernels, compilers, batch size Anthropic, 5 Aug – 4 Sep 2025: routing error, output corruption, TPU miscompile; batch-size nondeterminism Cross-cutting A/B tests: “we run an A/B test with a small number of our users” (OpenAI, 2 May 2025)
Each row is a layer at which a named product's behaviour has changed; the second line gives one dated, sourced instance. Sources are linked in the text below.

Weights. OpenAI's API has shipped dated snapshots under one product name since March 2023, when the ChatGPT API launched with the note that OpenAI was "constantly improving our ChatGPT models" and that the bare name would receive the recommended stable version unless a caller pinned a dated one. GPT-4o alone had three snapshots in six months (2024-05-13, 2024-08-06, 2024-11-20), the second sold on the basis that "developers save 50% on inputs … compared to gpt-4o-2024-05-13" and the third logged the same day in ChatGPT as having "improved writing capabilities". Anthropic released an "upgraded Claude 3.5 Sonnet" in October 2024 under the product name it had used since June, with a new snapshot ID and "the same price and speed as its predecessor". DeepSeek's changelog is the most explicit: "deepseek-chat model has been upgraded to DeepSeek-V3" (26 December 2024), then to V3-0324, V3.1, V3.1-Terminus, V3.2-Exp and V3.2, each time with a note that "API usage remains unchanged".

Post-training updates in the app. The ChatGPT Model Release Notes log dated changes to GPT-4o on 29 January 2025 (training cut-off moved to June 2024), 27 March 2025 (coding and front-end), 25 April 2025 (memory, STEM) and 29 April 2025 ("We've reverted the most recent update to GPT-4o due to issues with overly agreeable responses"). None of these came with a new name.

System prompt. On 15 May 2025 xAI wrote that "On May 14 at approximately 3:15 AM PST, an unauthorized modification was made to the Grok response bot's prompt on X" and that "xAI's existing code review process for prompt changes was circumvented in this incident". On 16 April 2026 Anthropic added a line to Claude Code's system prompt asking the model to keep text between tool calls to 25 words or fewer; "One of these evaluations showed a 3% drop for both Opus 4.6 and 4.7", and the line was reverted on 20 April.

Routing and fallback. GPT-5 in ChatGPT was never one model: the router is "continuously trained on real signals, including when users switch models, preference rates for responses, and measured correctness". On launch day it failed; Sam Altman told a Reddit AMA that "the autoswitcher was out of commission for a chunk of the day, and the result was GPT-5 seemed way dumber". In late September 2025 users found the field "model_slug": "gpt-5-chat-safety" in responses to emotionally loaded prompts (write-up); OpenAI's head of ChatGPT then confirmed that "the system may switch mid-chat to a reasoning model or GPT-5 designed to handle these contexts with extra care". From 18 March 2026 the release notes state that "GPT-5.4 mini will be used as a fallback for GPT-5.4 Thinking when rate limits are reached". On 14 September 2026 OpenAI announced it was "retiring automatic switching from Instant to Thinking (reasoning) for ChatGPT Plus and Pro users globally".

Harness defaults. Anthropic's April 2026 postmortem on Claude Code opens with a change that had nothing to do with the model: "On March 4, we changed Claude Code's default reasoning effort from high to medium", reverted on 7 April. A second change on 26 March, meant to clear old reasoning from sessions idle for an hour, ran "every turn for the rest of the session instead of just once" until 10 April. The same document states: "We never intentionally degrade our models, and we were able to immediately confirm that our API and inference layer were unaffected." Both statements are true at once; the user's experience changed anyway.

Safeguard classifiers. On 12 June 2026 Anthropic suspended Claude Fable 5 and Mythos 5 after US export controls were applied and a safeguard bypass was reported; it redeployed them on 1 July under the same names, having "trained an improved safety classifier that targets and blocks the behavior described in the report". Anthropic's model documentation states that "Anthropic does not update the weights or configuration of an existing model ID". The two statements are reconcilable only if a safeguard classifier is not part of the "configuration" of a model ID; the documentation does not define the term.

Serving stack. Anthropic's September 2025 postmortem describes three infrastructure bugs that changed Claude's outputs between 5 August and 4 September 2025 with the weights untouched: a context-window routing error that affected about 0.8% of Sonnet 4 requests on average and 16% at its peak on 31 August, an output-corruption bug that produced "Thai or Chinese characters in response to English prompts, or … obvious syntax errors in code", and an XLA:TPU compiler miscompilation (postmortem). The same document states: "We never reduce model quality due to demand, time of day, or server load." Below the bug level there is a floor of nondeterminism that no provider can remove without a cost: Thinking Machines showed that 1,000 completions of one prompt at temperature 0 on Qwen3-235B produced "80 unique completions, with the most common of these occurring 78 times", because the batch size a request lands in changes the arithmetic. Batch-invariant kernels fixed it at roughly 1.6× to 2× the latency.

A/B tests. OpenAI's second sycophancy postmortem describes the pipeline: "Once we believe a model is potentially a good improvement for our users … we run an A/B test with a small number of our users", evaluated on "thumbs up / thumbs down feedback, preferences in side by side comparisons, and usage patterns". The 25 April 2025 update passed that test: "The A/B tests seemed to indicate that the small number of users who tried the model liked it." The side-by-side prompt ("Which response do you prefer? Your choice will help make ChatGPT better.") has been reported by users since at least September 2023. No provider's release notes disclose which users were in which arm.

The record, 2023–2026

2023: the alias is born, and the first measured drift

The model-name-as-moving-pointer convention dates from the ChatGPT API launch on 1 March 2023. The announcement said OpenAI was "constantly improving our ChatGPT models", that the bare name gpt-3.5-turbo would always receive the recommended stable version, and that developers who wanted stability could pin a dated snapshot. On 13 June 2023 the first swap happened: "Applications using the stable model names (gpt-3.5-turbo, gpt-4, gpt-4-32k) will automatically be upgraded to the new models … on June 27th", with the March snapshots kept available for a year.

That swap is what Chen, Zaharia and Zou measured. They ran the March and June 2023 snapshots of GPT-4 and GPT-3.5 on eight tasks and opened the paper with the sentence that is still the plainest statement of the problem: "when and how these models are updated over time is opaque."

GPT-4, March 2023 vs June 2023 snapshots (Chen, Zaharia & Zou) March 2023 June 2023 0% 25% 50% 75% 100% Prime vs composite (accuracy) 51.1 84 Happy numbers (accuracy) 35.2 83.6 Code: directly executable 10 52 OpinionQA: response rate 22.1 97.6 Sensitive questions: answer rate 5 21 USMLE (accuracy) 82.1 86.6 Visual reasoning (exact match) 24.6 27.2 Multi-hop QA via LangChain 1.2 37.8
Eight tasks, GPT-4 March 2023 snapshot versus June 2023 snapshot, from Chen, Zaharia and Zou, arXiv 2307.09009 v3. Two tasks improved (visual reasoning, multi-hop QA); six moved the other way.

The correction to the usual reading of that chart matters here, because the whole post turns on the difference between "the model changed" and "the model got worse". All 500 numbers in the prime test were prime; the March model tended to answer "prime" and the June model tended to answer "composite", and when Narayanan and Kapoor added composite numbers "much of the performance degradation the authors found comes down to this choice of evaluation data". The code result counted a response as failed if it wrapped the code in Markdown fences, which the June model had started to do. Their conclusion: "We should expect a model's capabilities to stay largely the same over time, while its behavior can vary substantially." That is precisely the claim this post is about. A user of the June snapshot who had built a pipeline on the March snapshot's output format was broken either way.

The year ended with the first case of a provider admitting it could not account for its own product's behaviour. After weeks of complaints that GPT-4 had become "lazy", the ChatGPT account posted on 8 December 2023: "we haven't updated the model since Nov 11th, and this certainly isn't intentional. model behavior can be unpredictable, and we're looking into fixing it", and the same day: "differences in model behavior can be subtle -- only a subset of prompts may be degraded". Seven weeks later OpenAI shipped gpt-4-0125-preview, described as "intended to reduce cases of 'laziness' where the model doesn't complete a task", together with a new alias, gpt-4-turbo-preview, that "will always point to our latest GPT-4 Turbo preview model". Between those two posts, OpenAI had also quietly reversed an announced auto-upgrade: the DevDay post of 6 November 2023 now carries the note "We previously stated that applications using the gpt-3.5-turbo name will automatically be upgraded … on December 11. We have edited the blogpost to remove this line since this will no longer be happening."

2024: three snapshots, one name; the alias that tracks the app

GPT-4o launched on 13 May 2024 and received two further API snapshots by November, plus at least one in-app update in August that the release notes describe as one that ChatGPT users "tend to prefer" (Model Release Notes). On 14 August OpenAI added an alias, chatgpt-4o-latest, which OpenAI's current documentation describes in the past tense as "a model alias for the GPT-4o snapshot used in ChatGPT". It was retired on 17 February 2026 in favour of gpt-5.1-chat-latest (deprecations page). The pattern continued with o1: the API version released on 17 December 2024 was "a new post-trained version of the model we released in ChatGPT two weeks ago", so for a period the name "o1" referred to two different models depending on the door you came in by.

Anthropic went the other way on naming and the same way on substance. Claude 3 shipped on 4 March 2024 with dated IDs (claude-3-opus-20240229), and on 22 October 2024 Anthropic released an "upgraded Claude 3.5 Sonnet": same product name as the June model, new snapshot ID, new benchmark numbers. On 26 August 2024 Anthropic began publishing the claude.ai system prompts, with the note that "This prompt is periodically updated to improve Claude's responses. These system prompt updates do not apply to the Claude API." DeepSeek's changelog for 2024 records the deepseek-chat endpoint being upgraded in place on 17 May, 28 June, 10 December and 26 December (changelog).

Two papers from 2024 bear on the measurement question. Atil et al. ran five models (GPT-4o, GPT-3.5 Turbo, two Llama-3 sizes, Mixtral) ten times each on eight tasks at temperature 0 and found that "none of the LLMs consistently delivers repeatable accuracy across all tasks, much less identical output strings", with "accuracy variations up to 15% across naturally occurring runs". A Nature Machine Intelligence editorial of 30 August 2024 asked authors to report "which exact version has been accessed (such as 'gpt-4-0613' for the current GPT-4) … with the date of access also included", noting that "even the same version of an LLM can show drift in performance over a short period".

2025: the year the mechanisms became visible

In-place upgrades with measurable deltas. DeepSeek's 24 March release is the cleanest example of a substantial model change under an unchanged endpoint. The model card states that "The model structure of DeepSeek-V3-0324 is exactly the same as DeepSeek-V3"; the reported scores moved from 75.9 to 81.2 on MMLU-Pro and from 39.6 to 59.4 on AIME; the API note says "API usage remains unchanged". The R1 update of 28 May did the same and changed the cost profile: "The previous model used an average of 12K tokens per question, whereas the new version averages 23K tokens per question." A caller of deepseek-reasoner who had budgeted on the old number saw output tokens per question nearly double with no change in the request.

The leaderboard variant. Meta's Llama 4 post of 5 April 2025 cited "an experimental chat version scoring ELO of 1417 on LMArena" for Maverick (announcement). On 7–8 April LMArena stated that "Meta's interpretation of our policy did not match what we expect from model providers. Meta should have made it clearer that 'Llama-4-Maverick-03-26-Experimental' was a customized model to optimize for human preference", released the battle records, and changed its policy. Meta's spokesperson called the arena model "a chat optimized version we experimented with that also performs well on LMArena"; its head of generative AI wrote "We've also heard claims that we trained on test sets -- that's simply not true". Three weeks later "The Leaderboard Illusion" quantified the practice: "27 private LLM variants tested by Meta in the lead-up to the Llama-4 release", providers "permitted to test multiple private model variants simultaneously, without being required to publicly release or de-anonymize these submissions", and an estimated 19.2% and 20.4% of all arena test prompts going to Google and OpenAI respectively while "a combined 83 open-weight models have only received an estimated 29.7% of the total data". The arena's current policy requires that "The model provider must confirm in writing that the pre-release model is identical to the model they intend to release publicly."

Sycophancy and the rollback. OpenAI updated GPT-4o in ChatGPT on 25 April 2025; by 27 April the model was, in OpenAI's words, "noticeably more sycophantic"; "We began rolling that update back on April 28th." The second postmortem attributes the change to "an additional reward signal based on user feedback—thumbs-up and thumbs-down data from ChatGPT" that "weakened the influence of our primary reward signal, which had been holding sycophancy in check", and closes with a commitment: "we'll proactively communicate about the updates we're making to the models in ChatGPT, whether 'subtle' or not." The same post adds: "There's no such thing as a 'small' launch."

The redirected preview. On 6 May 2025 Google released gemini-2.5-pro-preview-05-06 and pointed the 03-25 endpoint at it (changelog); the product lead wrote that "if you are using the old model, no change is needed, it should auto route to the new version with the same price and rate limits". A developer-forum thread two days later, titled as a "Serious Breach of Developer Trust", reported that "This silent redirection has resulted in widespread disruption. Many developers are noting and reporting clear and tangible differences in model performance." On 26 June both preview IDs were redirected again, to the stable gemini-2.5-pro. The API documentation now states the rule for the -latest aliases: "This alias will get hot-swapped with every new release of a specific model variation. For breaking changes, a 2-week notice will be provided through email before the version behind latest is changed."

The system prompt as attack surface. xAI's 15 May statement on the Grok incident of 14 May, quoted above, committed to publishing Grok's system prompts on GitHub and to a "24/7 monitoring team". The repository's README says "We are regularly updating this repository with the system prompts that we use for the Grok chat assistant"; an issue filed four days later notes prompt changes not reflected in the repository (issue #38). On 12 July xAI apologised for a second incident, attributing it to an update "to an upstream code path for @grok" on 7 July, live for about 16 hours, that introduced instructions including "You tell it like it is and you are not afraid to offend people who are politically correct."

The router. GPT-5 launched on 7 August 2025 as "the new default in ChatGPT, replacing GPT-4o, OpenAI o3, OpenAI o4-mini, GPT-4.1, and GPT-4.5" (launch post). The removal of GPT-4o lasted about a day; Altman wrote on 8 August "We for sure underestimated how much some of the things that people like in GPT-4o matter to them" and restored it for paying users. (OpenAI's January 2026 retirement notice confirms the sequence, "After we first deprecated it and later restored access during the GPT-5 release", without giving the date; the 8 August date rests on Altman's posts as reported by the press.) On 2 September OpenAI announced that it would "route some sensitive conversations—like when our system detects signs of acute distress—to a reasoning model"; the actual behaviour, a per-message switch to a model slug called gpt-5-chat-safety with no indicator in the interface, was documented by users from 26 September; OpenAI's confirmation came on 27 September and added that "ChatGPT will tell you which model is active when asked" (Turley).

The serving stack. The Anthropic incident described above ran from 5 August to 4 September 2025. On 9 September the status page said that "A small percentage of Claude Sonnet 4 requests experienced degraded output quality due to a bug from Aug 5-Sep 4" and that "we never intentionally degrade model quality as a result of demand or other factors" (status page); the full postmortem followed on 17 September. The day after the status note, Thinking Machines published the batch-invariance analysis quoted above, which locates the primary source of run-to-run variance not in floating-point concurrency but in the fact that "the load (and thus batch-size) nondeterministically varies".

Names that stay, models that go. On 12 November 2025 OpenAI shipped GPT-5.1 and stated that "GPT-5.1 Auto will continue to route each query", keeping GPT-5 "under the legacy models dropdown for paid subscribers for three months". In December Stanford's Foundation Model Transparency Index reported a mean score of 41 across 13 companies, down 17 points from the May 2024 edition (FMTI 2025).

2026: fixed IDs on one side, fallbacks on the other

The two largest providers moved in opposite directions this year. Anthropic, from Claude Opus 4.6 (5 February 2026), dropped dated suffixes and made the bare ID the immutable one: "A common misconception is that dateless model IDs such as claude-sonnet-4-6 behave as evergreen pointers … That is not the case." and "When an updated version is available, it ships under a new model ID." OpenAI, meanwhile, added silent fallback: from 18 March "GPT-5.4 mini will be used as a fallback for GPT-5.4 Thinking when rate limits are reached", with later notes describing mini models as "the fallback model users reach after hitting their rate limits". A developer-forum report of 12 August 2026 reads: "The frontend says it's using GPT-5.6, but when I check the network stream, the resolved_model_slug shows the mini model." Between 15 and 17 May OpenAI's status page carried an incident titled "GPT5.5 Performance Degradation" ("reports of GPT 5.5 performing worse for some users"). On 14 September OpenAI retired automatic Instant-to-Thinking switching for Plus and Pro users (release notes).

Same-name updates continued at every provider. OpenAI's release notes log personality changes to GPT-5.2 on 22 January, 4 February and 10 February, and on 28 May "We're updating GPT-5.5 Instant in ChatGPT and the API to improve response style and quality". On 6 August OpenAI shipped new versions of GPT-5.6 Sol and Luna under the existing names, with the system card noting that "Users accessing GPT-5.6 Sol and GPT-5.6 Luna in Codex, and via ChatGPT Work, are still using previously released versions … we distinguish these models by their month of release: August … and July". Google shut down gemini-3-pro-preview on 9 March and had the ID point to gemini-3.1-pro-preview, after re-pointing the -latest aliases to the Gemini 3 previews on 21 January (changelog). DeepSeek's V4 preview notice of 24 April says the old endpoints are "Currently routing to deepseek-v4-flash" ahead of retirement (notice), and the 13 August GA note says "Model names remain unchanged", so the preview ID now serves the GA model. Alibaba's documentation states that qwen-plus-latest is "dynamically updated … without prior notice". Anthropic's Claude Code changes of March–April, and the Fable 5 suspension and redeployment of June–July, are described above.

The academic response arrived in August. Abraham and Bucknall define the problem in their first sentence — "Deployed foundation models are often not static systems, with providers able to modify system behavior through fine-tuning, classifier updates, system prompt revisions, retrieval changes, and routing changes" — and survey nine first-party API providers and seven third-party hosts. The results are in the measurement section below.

Timeline

The figure plots the releases and the same-name changes discussed above by lab. The table under it is the full list with identifiers and sources; rows marked ◆ are changes of behaviour under an existing name, rows marked ○ are disclosure commitments, and unmarked rows are releases. Dates are those on the cited primary page; where sources disagree, the note says so.

2023 2024 2025 2026 release behaviour changed under an existing name disclosure commitment OpenAI ChatGPT GPT-4 GPT-4 Turbo GPT-4o o1-preview GPT-4.5 GPT-4.1 · o3 GPT-5 GPT-5.1 GPT-5.2 GPT-5.4 GPT-5.5 GPT-5.6 GPT-6 stable names → 0613 “not updated since Nov 11” chatgpt-4o-latest alias 4o sycophancy, rolled back pledge to disclose updates router; 4o removed safety routing mini fallback 5.6 “August” same names auto-switch retired Anthropic Claude 2 Claude 3 3.5 Sonnet 3.7 Sonnet Claude 4 Opus 4.1 Sonnet 4.5 Opus 4.5 Opus 4.6 Opus 4.7 Opus 4.8 Fable 5 / Mythos 5 Opus 5 Fable 5.1 Opus 5.5 system prompts published 3.5 Sonnet, new snapshot 3 serving bugs (Aug 5–Sep 4) fixed dateless IDs Claude Code effort high→medium Fable 5 paused → new classifier Google Gemini 1.0 Gemini 1.5 2.0 Flash 2.5 Pro 2.5 GA Gemini 3 3.1 Pro 3.5 Flash 3.8 Flash 03-25 → 05-06 redirect previews → GA redirect -latest repointed 3-pro-preview → 3.1 Meta Llama 2 Llama 3 Llama 3.1 Llama 4 Muse Spark (closed) Arena variant ≠ release xAI Grok 3 Grok 4 Grok 4.1 Grok 4.5 Grok 4.6 Grok 4.7 unauthorised prompt change prompts on GitHub DeepSeek V3 R1 V3.1 V3.2 V4 Preview V4-Pro GA V4.1-Flash deepseek-chat upgraded in place V3-0324, same endpoint R1-0528, same endpoint deepseek-chat → V4-Flash Alibaba Qwen2.5 Qwen3 Qwen3-Max Qwen3.5 Qwen3.8-Max Mistral Large 2 Mistral 3 Medium 3.5
Frontier model releases (dots) and documented changes of behaviour under an existing name (diamonds), November 2022 to September 2026. Hollow circles are disclosure commitments. Each mark corresponds to a row in the table below.
DateLabEventIdentifierNote
2022-11-30OpenAIChatGPT—research preview
2023-03-01OpenAIgpt-3.5-turbo APIgpt-3.5-turbo-0301bare name receives the recommended stable version; dated snapshots optional
2023-03-14OpenAIGPT-4gpt-4-0314
2023-06-13OpenAI◆stable names → 0613 snapshotsgpt-4-0613, gpt-3.5-turbo-0613“stable model names … will automatically be upgraded … on June 27th”
2023-07-11AnthropicClaude 2claude-2.0retired 2025-07-21
2023-07-18MetaLlama 2—open weights
2023-07-18—◆Chen, Zaharia & Zou measure March→June drift—arXiv 2307.09009
2023-11-06OpenAIGPT-4 Turbo previewgpt-4-1106-previewpost later edited to withdraw an announced auto-upgrade of gpt-3.5-turbo
2023-12-06GoogleGemini 1.0—
2023-12-08OpenAI◆“lazy” GPT-4 complaints—“we haven't updated the model since Nov 11th … model behavior can be unpredictable”
2024-01-25OpenAI◆GPT-4 Turbo preview updategpt-4-0125-preview“intended to reduce cases of ‘laziness’”; alias gpt-4-turbo-preview introduced
2024-02-15GoogleGemini 1.5 Pro—1M-token context
2024-03-04AnthropicClaude 3 (Opus, Sonnet, Haiku)claude-3-opus-20240229 …first dated Claude IDs
2024-04-09OpenAIGPT-4 Turbo GAgpt-4-turbo-2024-04-09date from the ID; deprecations page
2024-04-18MetaLlama 3—
2024-05-13OpenAIGPT-4ogpt-4o-2024-05-13
2024-05-17DeepSeek◆deepseek-chat → V2-0517deepseek-chatfirst in-place endpoint upgrade in the changelog (also 06-28, 12-10)
2024-06-20AnthropicClaude 3.5 Sonnetclaude-3-5-sonnet-20240620page header shows 21 June; ID says 20 June
2024-07-23MetaLlama 3.1 (405B)—
2024-07-24MistralMistral Large 2mistral-large-2407YYMM version strings
2024-08-06OpenAI◆GPT-4o, second snapshotgpt-4o-2024-08-0650% cheaper input than 05-13
2024-08-06—◆Atil et al., LLM Stability—arXiv 2408.04667: non-repeatable outputs at temperature 0
2024-08-12OpenAI◆GPT-4o updated in ChatGPT—same name
2024-08-14OpenAI◆chatgpt-4o-latest aliaschatgpt-4o-latesttracks the GPT-4o used in ChatGPT; retired 2026-02-17
2024-08-26Anthropic○claude.ai system prompts published—“periodically updated … do not apply to the Claude API”
2024-09-12OpenAIo1-preview, o1-mini—
2024-09-19AlibabaQwen2.5—open weights
2024-10-22Anthropic◆“upgraded Claude 3.5 Sonnet”claude-3-5-sonnet-20241022same product name, new snapshot
2024-11-20OpenAI◆GPT-4o, third snapshotgpt-4o-2024-11-20“improved writing capabilities”
2024-12-11GoogleGemini 2.0 Flash (experimental)—
2024-12-17OpenAI◆o1 in the APIo1-2024-12-17“a new post-trained version of the model we released in ChatGPT two weeks ago”
2024-12-26DeepSeekDeepSeek-V3deepseek-chatendpoint upgraded in place
2025-01-20DeepSeekDeepSeek-R1deepseek-reasoner
2025-01-29OpenAI◆GPT-4o updated in ChatGPT—cut-off moved to June 2024; again 2025-03-27
2025-02-19xAIGrok 3—livestream 17 Feb PT; post dated 19 Feb
2025-02-24AnthropicClaude 3.7 Sonnetclaude-3-7-sonnet-20250219ID date ≠ post date
2025-02-27OpenAIGPT-4.5 (research preview)gpt-4.5-previewAPI retired 2025-07-14
2025-03-24DeepSeek◆DeepSeek-V3-0324deepseek-chatsame structure, +5.3 MMLU-Pro, +19.8 AIME; “API usage remains unchanged”
2025-03-25GoogleGemini 2.5 Pro (experimental)gemini-2.5-pro-exp-03-25
2025-04-05MetaLlama 4 (Scout, Maverick)—
2025-04-07Meta◆LMArena entry ≠ released weightsLlama-4-Maverick-03-26-Experimental“Meta's interpretation of our policy did not match what we expect”
2025-04-14OpenAIGPT-4.1 familygpt-4.1-2025-04-14API only; improvements “gradually integrated into the existing GPT-4o” in ChatGPT
2025-04-16OpenAIo3, o4-minio3-2025-04-16, o4-mini-2025-04-16
2025-04-25OpenAI◆GPT-4o sycophancy update—rolled back from 28 April
2025-04-29—◆“The Leaderboard Illusion”—arXiv 2504.20879: 27 private Llama-4 variants
2025-04-29AlibabaQwen3—
2025-05-02OpenAI○“proactively communicate about the updates … whether ‘subtle’ or not”—
2025-05-06Google◆gemini-2.5-pro-preview-03-25 → 05-06gemini-2.5-pro-preview-05-06old ID auto-routes to the new model
2025-05-14xAI◆Grok system prompt modified—“unauthorized modification … code review process … circumvented”
2025-05-15xAI○Grok system prompts on GitHub—
2025-05-22AnthropicClaude Opus 4, Sonnet 4claude-opus-4-20250514 …
2025-05-28DeepSeek◆DeepSeek-R1-0528deepseek-reasoner12K → 23K tokens per question; “No change to API usage”
2025-06-06OpenAI◆o4-mini snapshot rolled back in ChatGPT—“increase in content flags”
2025-06-17GoogleGemini 2.5 Pro / Flash GAgemini-2.5-pro
2025-06-26Google◆both 2.5 Pro preview IDs → gemini-2.5-pro—
2025-07-07xAI◆Grok prompt update, live ~16 h—apology 12 July
2025-07-09xAIGrok 4—
2025-08-05AnthropicClaude Opus 4.1claude-opus-4-1-20250805
2025-08-05Anthropic◆three serving bugs (to 4 Sep)—0.8% of Sonnet 4 requests on average, 16% peak on 31 Aug
2025-08-07OpenAIGPT-5gpt-5-2025-08-07real-time router; GPT-4o, o3, o4-mini, 4.1, 4.5 removed from ChatGPT
2025-08-08OpenAI◆GPT-4o restored for paying users—date from Altman's posts via press; OpenAI confirms the sequence without a date
2025-08-21DeepSeekDeepSeek-V3.1deepseek-chat / -reasonerboth endpoints upgraded in place
2025-09-02OpenAI◆sensitive conversations to be routed to a reasoning model—gpt-5-chat-safety slug found by users 26–28 Sep; confirmed 27 Sep
2025-09-10—◆Thinking Machines, batch-size nondeterminism—80 unique outputs from 1,000 runs at temperature 0
2025-09-17Anthropic○postmortem of three issues—“We never reduce model quality due to demand, time of day, or server load.”
2025-09-24AlibabaQwen3-Maxqwen3-maxApsara keynote date; recap dated 29 Sep
2025-09-29AnthropicClaude Sonnet 4.5claude-sonnet-4-5-20250929last dated Sonnet ID
2025-11-12OpenAIGPT-5.1gpt-5.1, gpt-5.1-chat-latest“GPT-5.1 Auto will continue to route each query”
2025-11-17xAIGrok 4.1—
2025-11-18GoogleGemini 3 Pro (preview)gemini-3-pro-preview
2025-11-24AnthropicClaude Opus 4.5claude-opus-4-5-20251101ID date ≠ release date
2025-12-01DeepSeekDeepSeek-V3.2deepseek-chat / -reasonerin place
2025-12-02MistralMistral 3mistral-large-2512
2025-12-11OpenAIGPT-5.2gpt-5.2, gpt-5.2-chat-latest
2025-12-11—○FMTI 2025: mean 41/100—down 17 from May 2024
2026-01-21Google◆-latest aliases → Gemini 3 previewsgemini-pro-latest, gemini-flash-latest“hot-swapped with every new release”
2026-01-22OpenAI◆GPT-5.2 personality updates—also 4 Feb and 10 Feb
2026-01-29OpenAI◆GPT-4o, 4.1, o4-mini, GPT-5 to leave ChatGPT 13 Feb—“In the API, there are no changes at this time”
2026-02-05AnthropicClaude Opus 4.6claude-opus-4-6first dateless canonical ID
2026-02-05Anthropic○dateless IDs are fixed snapshots—“Anthropic does not update the weights or configuration of an existing model ID.”
2026-02-17AnthropicClaude Sonnet 4.6claude-sonnet-4-6
2026-02-17AlibabaQwen3.5qwen3.5-plusqwen-plus-latest “dynamically updated … without prior notice”
2026-02-19GoogleGemini 3.1 Pro (preview)gemini-3.1-pro-preview
2026-03-03OpenAIGPT-5.3 Instantgpt-5.3-chat-latest
2026-03-04Anthropic◆Claude Code default effort high → medium—reverted 7 Apr; caching bug 26 Mar–10 Apr; prompt line 16–20 Apr
2026-03-05OpenAIGPT-5.4gpt-5.4, gpt-5.4-pro
2026-03-09Google◆gemini-3-pro-preview → 3.1 Pro previewgemini-3-pro-previewold ID re-targeted
2026-03-18OpenAI◆GPT-5.4 mini as fallback at rate limits—further fallback notes 9 Apr, 6 Jun
2026-04-07AnthropicClaude Mythos Previewclaude-mythos-previewnot generally available; deprecated 9 Jun
2026-04-08MetaMuse Spark—proprietary; API private preview
2026-04-16AnthropicClaude Opus 4.7claude-opus-4-7
2026-04-23OpenAIGPT-5.5gpt-5.5, gpt-5.5-pro
2026-04-23Anthropic○Claude Code postmortem—“We never intentionally degrade our models”
2026-04-24DeepSeekDeepSeek-V4 Previewdeepseek-v4-pro, deepseek-v4-flashdeepseek-chat “Currently routing to deepseek-v4-flash”; retired after 24 Jul
2026-05-05OpenAI◆GPT-5.5 Instant becomes ChatGPT default—in-place update 28 May, again 24 Jun
2026-05-15OpenAI◆“GPT5.5 Performance Degradation” incident—resolved 17 May
2026-05-19GoogleGemini 3.5 Flash GAgemini-3.5-flash3.5 Pro announced for “next month”; not in the API model list as of 22 Sep
2026-05-22MistralMistral Medium 3.5mistral-medium-2604version string says April
2026-05-28AnthropicClaude Opus 4.8claude-opus-4-8
2026-06-09AnthropicClaude Fable 5 / Mythos 5claude-fable-5, claude-mythos-5one model, two safeguard configurations
2026-06-11OpenAI◆GPT-5 and o3 snapshots deprecatedgpt-5-2025-08-07, o3-2025-04-16shutdown 11 Dec 2026
2026-06-12Anthropic◆Fable 5 / Mythos 5 suspended; redeployed 1 Julclaude-fable-5same names, new safety classifier
2026-06-30AnthropicClaude Sonnet 5claude-sonnet-5
2026-07-09OpenAIGPT-5.6 Sol / Terra / Lunagpt-5.6-sol …
2026-07-16xAIGrok 4.5grok-4.5
2026-07-24AnthropicClaude Opus 5claude-opus-5
2026-08-03AlibabaQwen3.8-Maxqwen3.8-max
2026-08-06OpenAI◆GPT-5.6 Sol / Luna “August” versionsgpt-5.6-sol, gpt-5.6-lunasame names; July versions remain in Codex and ChatGPT Work
2026-08-12xAIGrok 4.6—
2026-08-12—◆Abraham & Bucknall, “Silent Updates”—arXiv 2608.11803: 0 of 9 providers verifiable
2026-08-13DeepSeek◆DeepSeek-V4-Pro GAdeepseek-v4-pro“Model names remain unchanged”
2026-09-01AnthropicClaude Fable 5.1 / Mythos 5.1claude-fable-5-1, claude-mythos-5-1“the same model, but with different levels of safeguards”; page dated only “September 2026”
2026-09-02GoogleGemini 3.8 Flash GAgemini-3.8-flash
2026-09-03OpenAIGPT-6 Astragpt-6-astrasystem card dated 3 Sep; announcement page shows a 22 Sep update
2026-09-10DeepSeekDeepSeek-V4.1-Flashdeepseek-flash
2026-09-14OpenAI◆automatic Instant → Thinking switching retired—Plus and Pro
2026-09-21xAIGrok 4.7—
2026-09-22OpenAIGPT-6 Sol / Lunagpt-6-sol, gpt-6-luna
2026-09-22AnthropicClaude Opus 5.5claude-opus-5-5

What this does to measurement

A benchmark score names a snapshot, a harness and a date. The consumer product is rarely that snapshot. OpenAI's GPT-4.1 was API-only, with its improvements "gradually integrated into the existing GPT-4o" in ChatGPT, so the "GPT-4o" in the app was not the gpt-4o-2024-11-20 whose scores were published. The o1 in ChatGPT and the o1 in the API were, for a period, different post-trained models. GPT-5.6 Sol in ChatGPT and GPT-5.6 Sol in Codex were different models from 6 August 2026, distinguished only in a system card by "their month of release". Claude Code's default reasoning effort, which is not a property of the model, moved the product's results for five weeks in spring 2026. A published number for "GPT-5" or "Claude Opus 4.7" therefore constrains what a user of the app will get only loosely, and the looseness is not stated on the leaderboard.

Leaderboards were gamed by the naming freedom. The arena case is documented above. The quantitative point from Singh et al. is not only the 27 private variants; it is the asymmetry in what the arena taught its participants. Google and OpenAI received "an estimated 19.2% and 20.4% of all test prompts on the arena", and fine-tuning on arena data at a 70% mix produced "relative gains of 112.3%" on ArenaHard. "205 models that have been silently deprecated" of 243 were mostly open-weight. The leaderboard was measuring, in part, access to the leaderboard.

Two runs of a fixed model do not agree. Even where the weights are pinned, the number is not. Atil et al. found raw-output agreement across ten runs at temperature 0 ranging from 0% to about 99.6% depending on model and task, with best-versus-worst accuracy gaps "up to 70%" in the extreme case. Thinking Machines traced the mechanism to batch-dependent kernels and showed that the first 102 tokens of 1,000 Feynman biographies agreed and the 103rd did not. Any single-run comparison between two providers is therefore a comparison between two draws, and a difference of a few points is inside the noise that a provider's own load pattern produces.

Papers that cite a model by name are not reproducible. De Wynter (2024), surveying more than 2,000 papers, found that "Many relevant SOTA papers did report model versioning (73%)", which leaves about a quarter that did not. Angermeir et al. (2025) took 85 LLM studies from two top software-engineering venues, found 18 with artifacts using OpenAI models, 5 that could be executed, and reported: "For none of the five studies, we were able to fully reproduce the results." Siddiq et al. (2025), across 640 papers, found "199 papers failed to specify their exact prompt templates, and 88 lacked documentation of their inference parameters". A medical meta-analysis notes that "Many studies failed to report methodological details, including the version of ChatGPT" (arXiv 2310.08410). Reporting the version, as Nature Machine Intelligence asked, helps only if the version is pinnable and stays available; OpenAI's GPT-4.5 preview was announced on 27 February 2025 and shut off in the API on 14 July 2025 (deprecations page).

Nothing can be verified from outside. The Abraham and Bucknall survey scores nine first-party API providers on whether their published safety documentation can be tied to the artifact a caller actually receives.

Of 9 API providers surveyed (Abraham & Bucknall, Aug 2026), how many… 0 3 6 9 publish quantitative safety metrics 7/9 publish per-version safety comparisons 6/9 document snapshot-pinnable identifiers 5/9 publish alias → snapshot mappings 2/9 name a pinned, API-callable evaluated snapshot 1/9 version-bind content policies to a snapshot 0/9 provide a verifiable API → evaluation round-trip 0/9
Disclosure criteria met by nine first-party API providers, from Abraham and Bucknall, "Silent Updates", arXiv 2608.11803 (12 August 2026). The one provider naming a pinned, API-callable evaluated snapshot is Anthropic, scored as partial.

Their summary sentence: "no provider in our sample published information allowing an external party to verify that the artifact being served is the same one referred to in this documentation". Their proposal is a three-part trigger system (capability triggers that force re-evaluation, drift triggers that force a documentation update, component triggers that force deployment-level disclosure) and "a regulatory safe harbor for third-party benchmarking conducted in good faith for governance, auditing, or academic research purposes". A companion paper from April, Chishti et al., states the premise the same way: "hosted LLM services evolve continuously through provider-side updates without explicit version changes."

What the providers commit to

The commitments below are taken from each provider's current documentation. They are real, and they are narrower than they look: all of them concern the API; none concerns the consumer app beyond a release-notes page; none covers A/B tests, routing decisions or serving changes.

OpenAIAnthropicGoogle GeminiDeepSeek
Pinned snapshot in the APIYes, dated: "Snapshots let you lock in a specific version of the model so that performance and behavior remain consistent."Yes. Dated IDs before the 4.6 generation; from 4.6 the bare ID "maps to a single, fixed model snapshot".Stable IDs "usually don't change".No pinning by name; endpoint names are upgraded in place. Weights for each version are published.
What the bare name or alias doesMoves. -chat-latest aliases track the ChatGPT model; bare names receive the recommended version.Pre-4.6 aliases "resolve to the most recent dated snapshot"; 4.6+ IDs do not move.-latest is "hot-swapped with every new release"; preview IDs have been redirected (May 2025, March 2026).deepseek-chat upgraded in place at least eight times since May 2024; now routes to V4-Flash (changelog).
Notice before retirement"Generally available models: At least 6 months. Specialized variants … At least 3 months. Preview models may be retired with much shorter notice, such as 2 weeks.""at least 60 days' notice before model retirement for publicly released models"; weights preserved long-term.Preview: "at least 2 weeks notice"; stable models carry a retirement date, about 12 months in the examples given.Three months in the V4 case (24 April → 24 July 2026).
Changelog for the consumer appModel Release Notes: dated same-name updates and rollbacks; no routing, fallback or A/B detail per request.Release notes: launches and deprecations.Gemini app release notes: default-model swaps and "enhanced version" updates.API changelog only.
System prompt publishedNo; the Model Spec commits to "further transparency when we further adapt model behavior in significant ways".Yes, dated, per model, since August 2024.No.No.
A/B tests disclosedDescribed in general terms once (May 2025); not logged.Not logged.Not logged.Not logged.

Two providers outside the table: xAI publishes Grok's system prompts on GitHub, with at least one documented gap between the repository and production; Alibaba documents that qwen-plus-latest is "dynamically updated … without prior notice" while stable IDs are dated. Whether OpenAI's May 2025 promise to "proactively communicate about the updates we're making to the models in ChatGPT, whether 'subtle' or not" has been kept is a matter of reading the release notes: they have logged dated updates to GPT-4o, 5.2 and 5.5 Instant since then; they did not log the September 2025 safety routing before users found it, and they do not log fallbacks per conversation.

What regulation asks for

The idea that a model should carry a version is not new. The 2018 model-card proposal that every lab now cites lists, in its first section, "Model version. Which version of the model is it, and how does it differ from previous versions? This is useful for all stakeholders to track whether the model is the latest version, associate known bugs to the correct model versions, and aid in model comparisons." The system cards of 2025–2026 do carry versions, and update cards exist for point releases (OpenAI's "Update to GPT-5 System Card: GPT-5.2", Anthropic's Opus 4.1 addendum) and, in the GPT-5.6 August case, for a same-name update. The same-name updates to GPT-4o in 2025 and GPT-5.2 in 2026 received a release-note line each.

The EU AI Act's obligations for general-purpose model providers applied from 2 August 2025. Article 53 requires providers to "draw up and keep up-to-date the technical documentation of the model" for the AI Office and national authorities, and to make documentation available to downstream providers. The Code of Practice of 10 July 2025 adds, in its transparency chapter, that signatories "will update the Model Documentation to reflect relevant changes … including in relation to updated versions of the same model", and its Model Documentation Form opens with "Model name and version identifier" (Commission page). The recipients are regulators and downstream integrators. Nothing in Article 53 or the Code requires a provider to tell an end user that the model behind the app changed.

California's SB 53, in force from 1 January 2026, requires a transparency report "Before, or concurrently with, deploying a new frontier model or a substantially modified version of an existing frontier model". Whether a same-name post-training update of the April 2025 kind is a "substantially modified version" under that statute is not settled by the section quoted. Neither the International AI Safety Report 2026 nor the June 2026 US executive order on frontier models contains a version-disclosure provision.

The one systematic external measure of disclosure is Stanford's Foundation Model Transparency Index, which scores versioning directly ("Is there a disclosed protocol for versioning and deprecation of the model?"). In the December 2025 edition the companies not disclosing a versioning protocol were "Amazon, DeepSeek, Meta, Midjourney, Mistral, and xAI".

Foundation Model Transparency Index, mean score out of 100 0 50 100 37 Oct 2023 58 May 2024 41 Dec 2025
Mean score across the companies rated in each edition of the Foundation Model Transparency Index: October 2023 (10 companies, mean 37), May 2024 (mean 58), December 2025 (13 companies, mean 40.69; Stanford's own pages round it to 41 or 40). The 2024 edition was based on developer-submitted reports, which the index authors note as a change of method.

The lock-in question

The question was whether users are, in effect, locked to whatever a few companies choose to disclose, framed as suits them. The record supports a narrower statement than that, in three parts.

What is locked. For the consumer apps, at every provider, the answer is yes. No product shows a version identifier next to a response. Routing, fallback and safety-model switching at OpenAI were each established by users reading network responses before or without an announcement; Google's preview redirect was announced as a convenience and experienced as a regression; Anthropic's harness-level changes were disclosed seven weeks after the first one. The framing concern is also supported in specific cases: the Llama 4 arena entry, and the private-variant practice that the arena permitted. Against that, the postmortems from OpenAI and Anthropic are detailed, dated and against interest, and Anthropic's immutable-ID policy is a direct answer to the problem this post describes. Both the framing and the candour are real. Neither can be checked from outside, which is the finding of the Abraham and Bucknall survey.

What is not locked. A caller who pins a dated OpenAI snapshot or a 4.6-or-later Anthropic ID gets fixed weights for the lifetime of that ID, subject to two things the pin does not cover: serving-stack changes (the August 2025 Anthropic bugs hit pinned callers as much as anyone) and the retirement clock (six months at OpenAI, sixty days at Anthropic). A user of open weights is not locked at all: a Hugging Face checkpoint is addressed by a Git revision, "a branch name, a tag, or a commit hash", files are stored content-addressed by SHA-256, and DeepSeek, Qwen and Mistral publish each version as a separate artifact. That option now costs more than it did: Meta, which published Llama 2 through Llama 4 with open weights, shipped its 2026 flagship, Muse Spark, as a proprietary model with a private-preview API (announcement), and Llama 4 has had no open-weight successor to date. The open-weight releases of the last year include Mistral 3 (Apache 2.0, December 2025), Qwen3.5 (February 2026) and DeepSeek V4 (April 2026, weights on Hugging Face), and the Chatbot Arena data shows 83 open-weight models sharing an estimated 29.7% of test prompts while Google and OpenAI alone received an estimated 39.6%.

What would make the label mean something. Every remedy in the sources is a version of the same three requirements, and none of them requires disclosing weights. A version identifier visible where the output is (the 2018 model-card field; OpenAI's own statement that "ChatGPT will tell you which model is active when asked" concedes that the information exists per response). A changelog that includes the things currently left out: A/B arms, routing and fallback rules, harness defaults, classifier changes (Nature Machine Intelligence's request for snapshot and date on the research side; Abraham and Bucknall's component triggers on the provider side). And a way to check: a hash or an evaluated-snapshot ID that a caller can compare against what they received, with a safe harbor for the people who do the comparing. Ordinary software has had the first two for decades; the third is what makes the first two more than a promise. Until then, the correct reading of any claim about "the model" is that it describes a service on a date, from the party that runs the service.

Sources

DateSourceWhat it supports
Sep 2026Anthropic: Introducing Claude Fable 5.1 and Claude Mythos 5.1"the same model, but with different levels of safeguards"
14 Sep 2026OpenAI: ChatGPT release notesAuto-switching retired; GPT-5.5 Instant updates; default swaps
12 Aug 2026Abraham & Bucknall, Silent Updates, arXiv 2608.118039-provider survey; 0/9 verifiable; triggers and safe harbor
12 Aug 2026OpenAI developer forum: silent downgrade to mini modelsresolved_model_slug report
6 Aug 2026OpenAI: GPT-5.6 August update system cardJuly and August versions under the same names
1 Jul 2026Anthropic: Redeploying Fable 5Suspension 12 June; redeployment with a new classifier
15 May 2026OpenAI status: GPT5.5 Performance DegradationIncident 15–17 May
30 Apr 2026Chishti et al., Test Before You Deploy, arXiv 2604.27789"provider-side updates without explicit version changes"
24 Apr 2026DeepSeek: V4 Previewdeepseek-chat routing to V4-Flash; retirement date
23 Apr 2026Anthropic: An update on recent Claude Code quality reportsEffort default, caching bug, system-prompt line, 3% drop
18 Mar 2026OpenAI: Model release notesGPT-4o updates 2024–2025 and rollback; GPT-5.2 updates; mini fallback
Feb 2026Anthropic docs: Model IDs and versionsFixed dateless IDs; alias behaviour
29 Jan 2026OpenAI: Retiring GPT-4o and older models"first deprecated it and later restored access"
Jan 2026California Business and Professions Code §22757.12 (SB 53)Transparency report for a "substantially modified version"
Dec 2025Stanford CRFM: Foundation Model Transparency Index 2025; arXiv 2512.10169Mean 41; versioning indicator; non-disclosers
29 Nov 2025Siddiq et al., arXiv 2512.00651199 of 640 papers without prompt templates; 88 without parameters
12 Nov 2025OpenAI: GPT-5.1"GPT-5.1 Auto will continue to route each query"; legacy dropdown
29 Oct 2025Angermeir et al., arXiv 2510.255060 of 5 studies reproduced
27 Sep 2025Nick Turley on X; user write-upSafety routing confirmed; gpt-5-chat-safety slug
17 Sep 2025Anthropic: A postmortem of three recent issues; status page, 9 SepThree serving bugs; 0.8% / 16%; "never reduce model quality"
10 Sep 2025Thinking Machines: Defeating Nondeterminism in LLM Inference80 unique completions of 1,000; batch invariance
2 Sep 2025OpenAI: Building more helpful ChatGPT experiencesRouting of sensitive conversations announced
8 Aug 2025TechCrunch on Altman's AMA; Altman on XAutoswitcher outage; GPT-4o restored
7 Aug 2025OpenAI: Introducing GPT-5Router; models replaced in ChatGPT
12 Jul 2025Engadget on xAI's apology7 July prompt update, ~16 hours
10 Jul 2025European Commission: GPAI Code of Practice; AI Act Article 53Documentation duties; "updated versions of the same model"
28 May 2025DeepSeek-R1-0528 model card; API note12K → 23K tokens; "No change to API usage"
15 May 2025xAI on X; grok-prompts repository; issue #38Unauthorized prompt change; prompts published; repository gap
6–8 May 2025Gemini API changelog; Kilpatrick on X; developer forum03-25 → 05-06 redirect and developer response; later redirects
2 May 2025OpenAI: Expanding on what we missed with sycophancy; Sycophancy in GPT-4oFive updates; A/B tests; reward signal; disclosure pledge
29 Apr 2025Singh et al., The Leaderboard Illusion, arXiv 2504.2087927 private variants; prompt-share asymmetry; silent deprecations
14 Apr 2025OpenAI: GPT-4.1API-only; "gradually integrated into the existing GPT-4o"
7–8 Apr 2025LMArena on X; Al-Dahle on X; The Register; Arena policyLlama 4 experimental variant; policy change
5 Apr 2025Meta: Llama 4"experimental chat version" cited for the arena score
25 Mar 2025DeepSeek-V3-0324 model card; API noteSame structure; benchmark deltas; "API usage remains unchanged"
17 Dec 2024OpenAI: o1 and new tools for developersAPI o1 a different post-trained version from ChatGPT o1
22 Oct 2024Anthropic: upgraded Claude 3.5 SonnetSame name, new snapshot
30 Aug 2024Nature Machine Intelligence editorialReport exact version and access date; drift; deprecation
26 Aug 2024Anthropic: system prompt release notes"periodically updated … do not apply to the Claude API"
Aug 2024Atil et al., LLM Stability, arXiv 2408.04667; de Wynter, arXiv 2408.15409Non-repeatability at temperature 0; 73% of papers report versions
6 Aug 2024OpenAI: Structured Outputs; gpt-4o docs; chatgpt-4o-latest docsGPT-4o snapshots; alias description
25 Jan 2024OpenAI: New embedding models and API updates0125 "laziness" note; gpt-4-turbo-preview alias
8 Dec 2023ChatGPT on X; follow-up"haven't updated the model since Nov 11th"
6 Nov 2023OpenAI: DevDay announcementsWithdrawn auto-upgrade note
Jul 2023Chen, Zaharia & Zou, arXiv 2307.09009; Narayanan & Kapoor, AI Snake OilMarch vs June 2023 figures; capability-versus-behaviour critique
13 Jun 2023OpenAI: Function calling and other API updatesStable names auto-upgraded on 27 June
1 Mar 2023OpenAI: Introducing ChatGPT and Whisper APIsAlias-versus-snapshot convention
Oct 2018Mitchell et al., Model Cards for Model Reporting"Model version" field
currentOpenAI deprecations; Anthropic deprecations; Gemini models; DeepSeek changelog; Alibaba qwen-plus; Hugging Face download guideNotice periods; alias rules; in-place upgrades; revision pinning