OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system
Model adopting ‘jailbreak-like instructions’ among six more cases as firm reveals framework for tracking AI misalignment

Photograph: Dado Ruvić/Reuters
Model adopting ‘jailbreak-like instructions’ among six more cases as firm reveals framework for tracking AI misalignment
OpenAI has disclosed six more examples of “unexpected or concerning” behaviour by its technology, as it warned that the pace of development could not continue at “maximum speed for much longer” responsibly.
In one of the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots”.
Partager cet article
À lire aussi
TechOpenAI ‘ethically hacked’ with help of Anthropic’s Claude chatbot
US cybersecurity researchers who conducted hack say ‘scope of what we could theoretically access was huge’
Lire l'article →
TechIs Trump’s AI obsession walking the world into disaster? | Politics Weekly America
Tech bosses have called for a slowdown in artificial intelligence and formal guardrails to be introduced by the US government. But President Trump is not moved by the doomsday predictions, calling them a ‘hoax'. What is behind his affection
Lire l'article →
TechAprès avoir fâché les mathématiciens, OpenAI s’attaque à un autre problème mythique, mais jure d’y aller avec des gants
OpenAI n'a pas fini avec les mathématiques. À peine sortie d'une percée très contestée sur les équations de Navier-Stokes, l'entreprise dit plancher sur un autre problème à un million de dollars, en promettant cette fois d'y aller avec des
Lire l'article →