

I think part of the problem might be that it gets hard to make a model that’s both able to comprehend complex details of the world of its training data and also that’s had the training data crudely manipulated in inconsistent ways. You either get a model that’s “figured out” what you’re trying to conceal via other sources and inferences, or you get a model that’s just broken and dumb because it can’t reconcile the contradictions.
One may recall the incident where someone at X inserted some weird conspiracy theory about “white genocide” in South Africa into Grok’s system prompt, and Grok basically switched to malicious compliance mode - it wouldn’t shut up about it, inserting it into inappropriate discussions spontaneously, and when asked about it would look up sources and explain why this conspiracy theory was actually wrong. That was a system prompt, not training data, but I could see the same sort of thing happening to a model where all the information about what happened at Tienanmen Square had been excised from the training data. There would be a conspicuous absence of information. Conversations from its training data would end abruptly or skirt a specific time and place. News articles reference some change in how foreign powers viewed China at that date but never explain why. China’s policies themselves abruptly change around that time. It’d know something significant happened then but would have to fill in the blanks.
Perhaps better to just accept it and go with the approach of building censorship into the framework around the model instead.

Do you think it’s likely the Trump administration will successfully ban cutting-edge Chinese AI models? Especially given how most of the world is not in fact under American jurisdiction?