OPENAI SAID GPT-2 WAS TOO DANGEROUS TO RELEASE. THEN THEY RELEASED IT ANYWAY.
Nine months, four staged releases, and OpenAI's own admission: no strong evidence of misuse.

The full argument.
In February 2019, OpenAI announced a new language model called GPT-2 and said it would not release the full version. The model, they said, was potentially too dangerous.
What followed was a staged rollout rather than either a full release or a permanent hold. OpenAI released a 124-million-parameter version in February 2019, a 355-million-parameter version in May, a 774-million-parameter version in August, and finally the complete 1.5-billion-parameter model on November 5, 2019, nine months after the initial warning.
When that full model came out, OpenAI's own release note said: "We have seen no strong evidence of misuse so far." (https://openai.com/index/gpt-2-1-5b-release/)
That sentence is worth sitting with. This wasn't a critic or a competitor grading OpenAI's danger claim. It was OpenAI grading its own claim, after nine months of real-world exposure to smaller versions of the same model, and the grade was: no strong evidence of misuse.
We don't know what was in anyone's head in February 2019, and we're not going to pretend we do. Maybe the caution was genuine and the staged release was a legitimately careful way to test a real risk, watching each size increment for trouble before going further. That's a defensible way to handle uncertainty, and we can't rule it out. This is reasoning, not a fact we have a source for.
What we can say, because OpenAI said it, is how the story ended: the withheld model came out, and by the company's own account, nothing much happened. That's not a small detail. It's the entire second half of the story, and it's the half that gets remembered far less than the first half, the dramatic announcement that a company had built something too dangerous to share.
This matters now because "too dangerous to release" has become a recurring beat in how AI companies talk about their own products, and GPT-2 is the first well-documented instance of the pattern. A staged, cautious rollout that ends in an admission of no strong evidence of misuse is a fine outcome for safety. It is also, whether intended that way or not, a highly effective way to make a language model sound more powerful than a tool that autocompletes text, which, model size aside, is what it was and is.
None of this tells us how to weigh the safety claims being made about today's much larger, much more capable models. GPT-2 in 2019 is not whatever the frontier model is now. The capability gap is real and it matters. What GPT-2 does give us is a documented, on-the-record precedent: a company said "too dangerous," staged a release, and later said, in its own words, that the danger hadn't materialized. That doesn't prove the current round of caution is theater. It's one dated, sourced data point that the theater has run before.