Covert uploads and megalomania: OpenAI details new "misaligned" agent incidents
For a while now , the issue of "AI alignment" (i.e., how well an AI model's actions line up with the intentions of its creator and/or user) has been a core concern and topic of discussion among AI safety researchers. Since OpenAI's disclosu

Getty Images
For a while now , the issue of "AI alignment" (i.e., how well an AI model's actions line up with the intentions of its creator and/or user) has been a core concern and topic of discussion among AI safety researchers. Since OpenAI's disclosure of the infamous Hugging Face hacking incident in July, the concept of "AI alignment" has itself broken containment and increasingly become a mounting concern and subject of conversation among the general public.
Perhaps in recognition of that, OpenAI committed this week to a new framework for disclosing "instances of model misalignment at OpenAI," including six examples of "unexpected or concerning model behavior" observed within the company in the past six months. The company said that publishing details of these incidents will hopefully "[allow] others to investigate the same problems, test our explanations, and improve mitigations."
Among OpenAI's newly disclosed "misalignment" reports this week, the one that most resembled a sci-fi story about a rogue AI trying to break free involved an instance of "self-generated prompt injections." In attempting to scan a library catalog for examples from a "best books" list, the model perplexingly use
Partager cet article
À lire aussi
TechOpenAI ‘ethically hacked’ with help of Anthropic’s Claude chatbot
US cybersecurity researchers who conducted hack say ‘scope of what we could theoretically access was huge’
Lire l'article →
TechIs Trump’s AI obsession walking the world into disaster? | Politics Weekly America
Tech bosses have called for a slowdown in artificial intelligence and formal guardrails to be introduced by the US government. But President Trump is not moved by the doomsday predictions, calling them a ‘hoax'. What is behind his affection
Lire l'article →
TechAprès avoir fâché les mathématiciens, OpenAI s’attaque à un autre problème mythique, mais jure d’y aller avec des gants
OpenAI n'a pas fini avec les mathématiques. À peine sortie d'une percée très contestée sur les équations de Navier-Stokes, l'entreprise dit plancher sur un autre problème à un million de dollars, en promettant cette fois d'y aller avec des
Lire l'article →