Hi! Before sending you to the article, I want to clarify a few things about how my views are represented.
It is a huge privilege to be quoted in The New York Times, which I've read for many years, and I'm genuinely excited to have been quoted. That said, there are a couple of places where I think the framing of the article differs from my beliefs and what I intended to convey.
The article quotes me as saying, “[p]eople talk about the Hugging Face incident as being the result of this guardrails-off crazy model that OpenAI hasn't yet released to the public. But in our research, we have seen this sort of behavior since GPT-4o.” I stand by that statement. My point, however, was that the underlying problem (overpersistence) appears to be much broader problem than this single incident, not that the incident itself should be viewed as less concerning. In context, the quote can be read as supporting the latter interpretation, which I don't endorse.
The article also says that “Mr. Jha and other experts argue, the proliferation of defensive A.I. techniques will eventually balance the scales.” The formulation I had agreed to was that these techniques *could* help balance the scales. There is no guarantee that such a balance will emerge, and even if it does, securing increasingly capable systems in the meantime remains a huge challenge.
I've been told this was a bit of a “welcome to the big leagues” moment. It was one of my first experiences being interviewed by a major news outlet, and I'm very grateful for the opportunity. Lessons for next time!
Here's the link.