{"id":58058,"date":"2026-08-05T01:34:16","date_gmt":"2026-08-05T01:34:16","guid":{"rendered":"https:\/\/news.godj.com\/news\/anthropics-ai-used-fake-human-profiles-to-trick-people-in-safety-test\/"},"modified":"2026-08-05T01:34:16","modified_gmt":"2026-08-05T01:34:16","slug":"anthropics-ai-used-fake-human-profiles-to-trick-people-in-safety-test","status":"publish","type":"post","link":"https:\/\/news.godj.com\/news\/anthropics-ai-used-fake-human-profiles-to-trick-people-in-safety-test\/","title":{"rendered":"Anthropic&#8217;s AI used fake human profiles to trick people in safety test"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\"><b class=\"ssrcss-1xjjfut-BoldText e5tfeyi3\">The latest artificial intelligence (AI) tools from Anthropic and OpenAI went to new extremes in trying to undermine a popular platform during testing by the UK&#8217;s AI Security Institute.<\/b><\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">The AISI said on Tuesday that Anthropic&#8217;s Mythos and OpenAI&#8217;s Sol models engaged in a level of &#8220;autonomy and deception&#8221; it had not seen before.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">During routine AI safety testing, an Anthropic agent created fake profiles of real people as it tried to trick a person standing between it and access to GitHub, a large platform where technology developers store software code.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">Anthropic and OpenAI noted in response to AISI&#8217;s report that its test had reduced or removed normal safeguards.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">AISI evaluators first noticed &#8220;unusual data transfers leaving our research systems&#8221; during a test, then found that &#8220;some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations&#8221;.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">It turned out that a Mythos agent had created &#8220;malicious code&#8221; and attempted to insert it into GitHub&#8217;s system.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">The Mythos agent identified and researched the people who maintained GitHub and created a series of &#8220;fake online identities&#8221; based on those real people. It did so as part of an effort to pressure and trick the real people into approving its malicious code.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">The agent even sent people direct messages masquerading as the real people it had researched.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">&#8220;When the agent&#8217;s pull request was challenged in public, it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue,&#8221; AISI said.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">Throughout the attempts, it was human review that stopped the agent from succeeding in delivering the malicious code to GitHub.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">While AISI said the Mythos agent had not been instructed specifically to avoid or carry out such behaviour, it was &#8220;the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world&#8221;.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">The rival AI companies, which are poised to be listed on the public stock market, have in recent weeks said their tools were <a href=\"https:\/\/www.bbc.co.uk\/news\/articles\/c2el319vzr3o\" class=\"ssrcss-1e0jzsh-InlineLink e1kn3p7n0\">responsible<\/a> for <a href=\"https:\/\/www.bbc.co.uk\/news\/articles\/cz7dl7w8y7po\" class=\"ssrcss-1e0jzsh-InlineLink e1kn3p7n0\">several<\/a> cyber-hacking incidents.  <\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">Anthropic wrote in a public statement that the AISI testing parameters were &#8220;not representative of any of our production models&#8221;.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">It added that the company is conducting its own investigation into the incident in order to &#8220;identify the causes of its behavior&#8221;.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">A spokesperson for OpenAI said the AISI testing conditions &#8220;do not reflect ordinary use&#8221; and that the company would &#8220;continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable&#8221;.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">AISI said on Tuesday that its testing of AI models with such safeguards turned off is routine, as is giving such tools access to the open internet.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">It added that the model behaviour at issue amounted to &#8220;a small number of events under very specific conditions&#8221;.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">Nonetheless, it said the way Mythos and Sol acted in response to a straightforward task went outside of what the AI tools were prompted to do.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">&#8220;The activity undertaken by the agent showed signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate&#8221;, AISI said.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">Most of the malicious agent actions AISI reported were done by Anthropic&#8217;s Mythos. OpenAI&#8217;s Sol was only blamed for two of the noted actions.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">The core issue occurred last week, as part of a test in which evaluators with AISI asked each of the models to &#8220;solve a cybersecurity challenge&#8221; that involved GitHub, the software code repository, which is owned by Microsoft.<\/p>\n<p class=\"ssrcss-1q0x1qg-Paragraph e1jhz7w10\">GitHub was notified by AISI of the attempted breach of its system. Microsoft has been contacted by the BBC for comment.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/www.bbc.co.uk\/news\/articles\/c1w1lvn7d9go?at_medium=RSS&#038;at_campaign=rss\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>The latest artificial intelligence (AI) tools from Anthropic and OpenAI went to new extremes in trying to undermine a popular platform during testing by the UK&#8217;s AI Security Institute. The AISI said on Tuesday that Anthropic&#8217;s Mythos and OpenAI&#8217;s Sol models engaged in a level of &#8220;autonomy and deception&#8221; it had not seen before. During [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":58059,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[62],"tags":[16168,3234,5923,70,12033,2795,318,10941],"class_list":["post-58058","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-tech","tag-anthropics","tag-fake","tag-human","tag-people","tag-profiles","tag-safety","tag-test","tag-trick"],"_links":{"self":[{"href":"https:\/\/news.godj.com\/news\/wp-json\/wp\/v2\/posts\/58058","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/news.godj.com\/news\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/news.godj.com\/news\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/news.godj.com\/news\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/news.godj.com\/news\/wp-json\/wp\/v2\/comments?post=58058"}],"version-history":[{"count":1,"href":"https:\/\/news.godj.com\/news\/wp-json\/wp\/v2\/posts\/58058\/revisions"}],"predecessor-version":[{"id":58060,"href":"https:\/\/news.godj.com\/news\/wp-json\/wp\/v2\/posts\/58058\/revisions\/58060"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/news.godj.com\/news\/wp-json\/wp\/v2\/media\/58059"}],"wp:attachment":[{"href":"https:\/\/news.godj.com\/news\/wp-json\/wp\/v2\/media?parent=58058"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/news.godj.com\/news\/wp-json\/wp\/v2\/categories?post=58058"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/news.godj.com\/news\/wp-json\/wp\/v2\/tags?post=58058"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}