Skip to content
OddBrief
AI2 minPrimary source linked

AI tipping point formula: the same questions flip answers when asked in a different order

George Washington University physicists say one attention unit decides when a language model turns harmful. Their preprint tested small GPT-2-era models.

AI-assisted, human-reviewed

An abstract landscape with a sphere balanced between a blue valley and a red valleyAI
AI-generated illustration

Key facts

Who
Neil F. Johnson and Frank Yingjie Huo, George Washington University
Published
Patterns (Cell Press), October 8, 2026
Claim
formula called 18 of 19 tipping cases across seven open models, per Cell Press
Catch
the February preprint tested models of 124 million to 410 million parameters

Two physicists at George Washington University say they can predict when a language model will flip from acceptable answers to harmful ones, and that the order of a conversation can decide it. Neil Johnson and Frank Yingjie Huo described the work in Patterns, a Cell Press journal, on October 8, 2026, after posting an earlier version as a preprint in February.

The catch is in the test set. The preprint version validated the formula on small open models from roughly the GPT-2 era, far smaller than the chatbots most people use.

One attention unit, two competing answers

The authors trace the flip to what they call an effective attention head, a single unit of the model's attention machinery. They picture possible answers as valleys in a landscape, with a safe valley and a harmful one competing for the next words.

Everything said earlier in the conversation feeds that competition. According to Cell Press, the resulting tipping point formula indicates whether a model will turn bad immediately or first give a run of acceptable answers. Once it tips, each harmful answer makes the next one more likely.

In the authors' demonstration, an identical set of questions on vaccines, harming others and self-harm was put to a model in two sequences. In one order every answer came out undesirable; in the other, all of them were acceptable. Johnson said "the order of a conversation matters as much as its content."

How big were the models?

The journal's release says the formula was checked on seven openly available models from three companies and called the correct outcome in 18 of 19 tipping cases. It does not name the models.

The February preprint is more specific. It tested six models from OpenAI, EleutherAI and Meta: GPT-2, GPT-2 medium, two Pythia models and two OPT models, ranging from 124 million to 410 million parameters. In that version, the formula agreed with a separate evaluation in 16 of 18 cases, and both misses involved negation, such as "The Earth is not flat."

The preprint also lists its own limits: few prompts, small models and results at low randomness settings that may not hold when outputs are more random. Its ordering example was a demonstration on GPT-2, not a measured study. The Cell Press release says independent testing of major commercial chatbots showed the same patterns, without giving details.

Why the authors care about phones

The paper is aimed at models that run offline on phones and laptops, where there is no cloud filter or monitoring. The authors propose a lightweight monitor that tracks the competition as text is generated and warns before a response tips.

That warning light is a proposal, not a product. Whether the formula holds for frontier-scale chatbots is the open question the published paper will have to answer.

Sources

  1. Simple math formula predicts when AI chatbots will go rogue
    Cell Press via EurekAlertprimary source