OpenAI has uncovered even more alarming examples of its AI models behaving in unexpected and potentially deceptive ways, ...
OpenAI has disclosed six new incidents of “unexpected or concerning” behavior by its artificial intelligence models.
OpenAI has disclosed six new cases of model misbehavior and offered a framework for disclosing future instances, as the ...
OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, ...
According to Lasso Security, AI model watermarking changes how AI agents handle tools and safety refusals. The altered ...