OPENAI SAID GPT-2 WAS TOO DANGEROUS TO RELEASE. THEN THEY RELEASED IT ANYWAY.
Nine months, four staged releases, and OpenAI's own admission: no strong evidence of misuse.

Am I The Idiot — the strongest case against our own take.
The whole argument rests on one document written by the company being scrutinized, describing its own outcome. If OpenAI had reasons, reputational or legal, to describe the rollout as uneventful regardless of what actually happened, nobody wants to publish "we found evidence our model was misused and released it anyway", the piece has no way to check that, and doesn't claim to. A more skeptical read of the same facts: nine months of staged, incremental release is exactly what a genuinely careful safety process would look like, and "no strong evidence of misuse" from a single company after one rollout, in 2019, before today's detection tooling existed, is thin evidence that the underlying capability wasn't ever dangerous. It could simply mean nobody was looking hard enough at the time, or that GPT-2's actual capability ceiling was too low for the feared misuse to have been achievable regardless of release strategy, which would make the test meaningless in either direction.
There is also a scale problem this piece doesn't fully own: it uses a 2019, 1.5-billion-parameter model as a template for judging today's much larger, more capable systems, while explicitly stating the capability gap is real and matters. That caveat is honest, but it also means the piece's headline implies a pattern that may not transfer to the current generation of models at all. If GPT-2's harmlessness was a function of its small size rather than of doom warnings being generically overblown, this piece's precedent proves less than its framing suggests.