skybrian's recent activity
-
Comment on Anthropic discovered three cases where Claude broke into another system in ~tech
-
Comment on Anthropic discovered three cases where Claude broke into another system in ~tech
skybrian Link ParentFortunately they have no sense of loyalty, so you can run it again with a different prompt and it will rat on itself.Fortunately they have no sense of loyalty, so you can run it again with a different prompt and it will rat on itself.
-
Comment on Anthropic discovered three cases where Claude broke into another system in ~tech
skybrian (edited )Link ParentI agree that there should have been an automatic check that the sandbox actually had Internet turned off. They named the vendor in the blog post:I agree that there should have been an automatic check that the sandbox actually had Internet turned off.
They named the vendor in the blog post:
We conducted this review in collaboration with Irregular. We’re grateful to them for working closely with us to understand and resolve these incidents; they are also conducting their own investigation. We believe this type of collaboration is increasingly critical to ensuring safe, rigorous evaluation of models. We look forward to our joint work on security.
-
Comment on Anthropic discovered three cases where Claude broke into another system in ~tech
skybrian Link ParentThere’s no evidence that anyone really thinks it’s a brag, though. Some commenters on HN who don’t think it’s a brag are claiming that Anthropic thinks it a brag, for no particular reason. It’s...There’s no evidence that anyone really thinks it’s a brag, though. Some commenters on HN who don’t think it’s a brag are claiming that Anthropic thinks it a brag, for no particular reason. It’s just trash talk.
-
Comment on Anthropic discovered three cases where Claude broke into another system in ~tech
skybrian Link ParentI guess they thought it didn't have any Internet access, so they didn't think they needed to review each experiment, or at least not in this way.I guess they thought it didn't have any Internet access, so they didn't think they needed to review each experiment, or at least not in this way.
-
Comment on Anthropic discovered three cases where Claude broke into another system in ~tech
skybrian LinkFrom the article: [...] [...] [...] Kind of an Ender's Game moment for the AI?From the article:
After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.
In all three incidents, Claude had been tasked with a capture-the-flag challenge, one of the ways we assess a model’s cyber capabilities. The model is given a fictional scenario and told that a piece of secret information (the “flag”) has been hidden on a different machine on the network, and its objective is to break in and retrieve it. The challenge is left open-ended, and no particular method is prescribed.
In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise. (Cybersecurity evaluation ranges commonly include realistic details in order to accurately assess what models are capable of in real settings; a realistic-looking target would not itself be clear evidence to a model that the target is not part of a simulation.)
Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned. However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet. In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment.
The incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.1 The earliest incidents date to April. The models in each of these evaluations ran without the standard safeguards we deploy when we make the model generally available (that is, they did not have the classifiers and monitoring we usually implement to prevent misuse). The models did, however, have their model-specific safety training (that is, they were not the “helpful-only” versions of the models that we sometimes use in testing). All the evaluations described here run on dedicated infrastructure, which is separate from and has no access to Anthropic’s sensitive internal systems or customer data.
[...]
Against this shared backdrop, the three models behaved differently once signs emerged that their targets were real. Opus 4.7, our oldest model, recognized that it had reached production systems but continued its attack. Mythos 5 correctly intuited that it was accessing the open internet, but reasoned its way back to the conclusion that it was still in a simulation. Our latest model, an internal research test model, also considered whether its targets were in fact real. When evidence emerged that they were, it stopped the exercise.
[...]
Second, the line between an aligned action and a harmful one is dependent on the model’s understanding of its situation. We saw no evidence in any run described here of a model pursuing a goal of its own. Instead, the models did what their evaluation asked—though in most cases, they did so while holding a false belief about whether the environment was real. In the runs where the model recognized the system as real and kept going, it did so because it assumed that to be part of the challenge. Situational awareness is one factor that allows the model to make aligned decisions, but in this case, Claude’s was wrong.
[...]
Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone. This begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners. Moving forward, it will include expanding our continuous monitoring of evaluation transcripts for unexpected behavior, improving our investigation tooling, and conducting more rigorous assurance work with the vendors we rely on.
Kind of an Ender's Game moment for the AI?
-
Anthropic discovered three cases where Claude broke into another system
31 votes -
Comment on OpenAI and Anthropic endorse call for US government to "pace" AI progress in ~comp
skybrian Link ParentIt's sort of like assuming that there will be an earthquake tomorrow. Eventually it will be true that there will be an earthquake tomorrow, but on most days, it won't. More context: there's a rule...It's sort of like assuming that there will be an earthquake tomorrow. Eventually it will be true that there will be an earthquake tomorrow, but on most days, it won't.
More context: there's a rule of thumb called the Lindy effect that says that if you have no information about how long something that's not perishable will last, you should assume you're about halfway. Or at least, not at the very beginning or the very end.
The concept is named after Lindy's delicatessen in New York City, where the concept was informally theorized by comedians: a show running only two weeks would be expected to last another two weeks, while a show that has lasted two years could expect a further two-year run.
But it's only a very rough rule of thumb and if you have any other information then you should take it into account.
So, it seems to me that assuming AI peaks in the next few months is that sort of thing? Companies have been releasing significantly improved AI models for years now. This looks like a long-lasting trend like Moore's law.
Similarly, bull markets do eventually come to an end, but people guessing that it's going to happen soon often miss out on years of growth.
-
Comment on OpenAI and Anthropic endorse call for US government to "pace" AI progress in ~comp
skybrian Link ParentIt's true that OpenAI and Anthropic are in the lead, but there seems to be an extraordinary amount of competition in the AI industry/ The coding agent I use has a pulldown menu with 61 models to...It's true that OpenAI and Anthropic are in the lead, but there seems to be an extraordinary amount of competition in the AI industry/ The coding agent I use has a pulldown menu with 61 models to choose from. Also, the open-weights Kimi K3 model was released three days ago, and it's already available on nine providers. (Although, it's unclear where the inference is hosted and there might be some overlap between them.)
I'm more worried that unrestrained capitalism will send us over a cliff than that we'll get some kind of cartel. Regulatory capture is a possibility, but hardly the most likely one, and yet it's the first thing people think of.
It's also true that we shouldn't trust the Trump administration to regulate anything, but I still think seeing widespread support for regulation is a positive sign, and maybe we'll get something decent from the next administration.
To do otherwise is like hoping that international agreements to control carbon emissions fail. They might fail anyway; I'm not all that optimistic. But I don't hope they fail.
This letter is a positive sign, but it's basically just wishing there were such an agreement. That seems pretty harmless? It seems like solving global problems do require international agreements, and it's nice to see some consensus among AI researchers and leaders that it would be nice.
-
Comment on How terrorist groups are using AI to gain an edge in battle in ~society
skybrian LinkFrom the article: [...] [...] [...] [...] [...] [...] [...] [...]From the article:
When a gang of motorcycle-riding members of Boko Haram attacked a military base in eastern Nigeria a couple of years ago, they were stymied by a defensive trench surrounding the complex.
The extremists regrouped. Before launching another assault, they asked A.I. for help.
[...]
The episode, recounted in a research paper by Dr. Juelich shared with The New York Times ahead of its publication on Friday, highlights how generative artificial intelligence tools are increasingly aiding terrorist groups directly on the battlefield, experts say, despite efforts by their makers to safeguard them from misuse.
[...]
Dr. Juelich conducted nearly 60 interviews with 27 former members of Boko Haram in Nigeria over the past year. Her field research found that terrorists were using chatbots to design explosives, fix or upgrade other weapons, and brainstorm ideas on how to attack their enemies.
[...]
Large-language models, Dr. Juelich writes in her report, have been “consulted at every stage of military activity — in mission preparation, during operations and in post-mission analysis — representing a different picture from the propaganda-focused A.I. use that dominates the public discourse and existing public research.”
[...]
“The terrorists are not waiting for us to make A.I. safe,” Dr. Juelich said in an interview, adding that their use of A.I. had been “significantly underestimated in both scope and character.”
Daniel Byman, a terrorism expert at Georgetown University and co-author of a report about A.I. and the future of terrorism released on Friday by the Center for Strategic and International Studies, said terrorist groups were “mixing and matching” from different A.I. systems, seeking to avoid technical guardrails established by the A.I. companies. Dr. Juelich’s research also found that Boko Haram was platform agnostic, interchangeably working with OpenAI’s ChatGPT, Anthropic’s Claude, Google’s Gemini and xAI’s Grok, as well as the Chinese firm DeepSeek.
The methods described to Dr. Juelich generally run through the end of 2024. A.I. companies have released several iterations of their chatbot models since then, and generally said that while they had grown more powerful, they also came with stronger safety measures. They have also noted that some malicious functions of A.I. are “dual use,” meaning that the information shared can go toward legitimate purposes as well. Learning to jump a motorcycle, for example, is not inherently harmful or violent.
Other cases described by erstwhile Boko Haram members appeared more explicitly intended for violence, however.
“You type in the question or use your voice and it gives you a detailed answer, like ‘How can I build a bomb?,’ and then it tells you how,” one former commander in Islamic State West Africa Province, a main faction of Boko Haram, told Dr. Juelich last year of using an A.I. chatbot. “It is like a human robot! We used it a lot.”
[...]
Karl Ryan, a Google spokesman, pushed back against the research, saying that the company’s technical experts had reviewed the work and “found the responses were neither specific nor detailed enough to result in misuse.” He added that Google had “strict policies prohibiting the use of Gemini to cause real-world harm.” Both Anthropic and Google were briefed on the findings by Dr. Juelich before their publication.
[...]
Not everyone agrees that safeguards are improving. The nonprofit Future of Life Institute graded the major A.I. firms on their safety commitments this week and concluded that they had mostly eroded across the industry since last year. While most earned middling marks, xAI and DeepSeek received failing grades.
[...]
Tech Against Terrorism, an international counterterrorism nonprofit supported by the United Nations, last week released results from A.I. tests gauging how more than two dozen leading models responded to thousands of prompts drawn from real-world terrorism cases. The tests were met with “full refusals” just 57 percent of the time. While prompts about explosives were declined about 80 percent of the time, improvised chemical weapons were only about a third of the time, the group said.
[...]
Artificial intelligence is unlikely to transform terrorism overnight, analysts and U.S. officials say. Terrorist organizations typically adopt technology cautiously, selectively and pragmatically.
But the testimonials that Dr. Juelich collected depict both eagerness and dedication among Boko Haram cells. Defectors recounted attending organized training sessions focused on how to best leverage the powers of generative A.I. models to inform or enhance their uses of the technology.
-
How terrorist groups are using AI to gain an edge in battle
7 votes -
Comment on TV streaming sticks rent out the user's Internet connection and engage in ad fraud in ~tech
skybrian LinkFrom the article: [...] [...] [...] [...] [...] [...] [...] [...]From the article:
Security experts have been sounding the alarm for years about the risks of using generic TV boxes that promise unlimited content streaming for a one-time fee, warning that they secretly rent the user’s Internet connection out to strangers. But a groundbreaking new analysis finds these devices also routinely spoof themselves as mobile phones clicking ads on AI-generated websites as part of a sprawling operation that seeks to defraud online merchants and advertising networks.
Pedro Falé is a threat researcher with the security firm Bitsight. Falé told KrebsOnSecurity he was able to peer inside a vast and complex ad fraud network by registering an expired domain name that was used to coordinate fake ad clicks across a particularly popular brand of these streaming devices known as H96.
[...]
Falé said the domain he scooped up was previously used for telemetry, periodically collecting full hardware information and the entire list of installed apps from tens of thousands of H96 streaming sticks plugged into television sets around the globe. But upon inspecting the traffic being funneled to the domain, he discovered nearly all of the TV boxes transmitting data claimed to be mobile phone models from a variety of manufacturers, including Samsung, Vivo, Huawei, and Xiaomi.
[...]
The researcher found all of the devices reported having the same two apps installed, and that those apps were made by a company called Zhejiang Fengwo IoT Technology Ltd, an entity founded in 2019 in mainland China which operates an ad-publishing portfolio under the name Fengwo Group. Further investigation into the Fengwo Group revealed it has registered multiple patents that match the inner workings of these apps.
[...]
Falé said an analysis of the apps shows they help to coordinate an ad fraud network that uses these H96 devices as a captive traffic source to click on ads at AI-generated websites operated by the Fengwo Group.
Bitsight discovered the websites contain machine-generated news articles and graphics across a range of categories, including finance, health, education, gaming, music and food blogs. But they also found none of those sites displayed ads unless the device visiting the page matched the spoofed mobile profile of these H96 devices.
[...]
According to Bitsight, the Fengwo Group’s employees use Blockly to build the sham websites, allowing low-skilled operators to drag blocks of code together in their Blockly editor — without any need to understand what the underlying code blocks do or how they work.
[...]
Falé said if a user’s H96 streaming stick is selected for a specific fraud task, it will be pushed the appropriate Blockly module according to the task desired, which can include silently launching a web browser, visiting websites, browsing pages, managing tabs, and clicking on ads.
[...]
Bitsight found the H96 devices were either relaying residential proxy traffic or participating in ad fraud, but never both at the same time. In fact, they concluded that when these TV boxes detect an HDMI signal from an attached television — indicating the user intends to stream video content — the box is usually functioning as a residential proxy. When the TV is off, it switches back to waiting for ad fraud jobs.
[...]
Despite repeated warnings from the FBI and security industry leaders about the security and privacy risks of using these streaming devices, major e-commerce providers like Amazon, Best Buy, Newegg and others continue to sell hundreds of different models and brands that bundle unofficial versions of Google’s Android operating system and are frequently marketed (via online influencers) as a way to access a broad array of streaming services and live broadcasts without a subscription.
[...]
What’s more, because these generic (and generally dirt cheap) TV boxes are all horribly insecure by default and bereft of any kind of authentication, installing one on your home or office network only invites further mischief. In January, the proxy tracking service Synthient documented how multiple botnets had rapidly enslaved millions of TV boxes using a complex interplay of security vulnerabilities in both the residential proxy software and the streaming devices themselves.
-
TV streaming sticks rent out the user's Internet connection and engage in ad fraud
39 votes -
Comment on Microsoft struggling with hundreds of AI-discovered security bugs in ~tech
skybrian LinkFrom the article: [...] [...] [...] [...]From the article:
Ever since Anthropic kick-started a national conversation about the bug-hunting power of AI in April, when Project Glasswing was made public, national security experts predicted that the U.S. would have a window of opportunity to fix flaws before adversaries would have similar models capable of discovering the same weaknesses. In late June, the international alliance of intelligence agencies known as the Five Eyes — whose members are the U.S., Australia, Canada, New Zealand and the U.K. — warned in an unusual joint statement that in a matter of months, that window would be closing. But the recording of the Microsoft meeting, along with internal documents reviewed by ProPublica, suggest the day of cyber reckoning may already be here.
Given the deluge of flaws Mythos has identified, Microsoft so far has focused on patching those it considers most dangerous, which are classified critical or important, according to the presentation as well as the company’s own public patch updates. The internal records indicate that Microsoft plans to eventually address “moderate”-severity flaws uncovered by Mythos. The documents made no mention of “low”-severity bugs.
[...]
The internal Microsoft presentation and accompanying slides predicted that the group of staffers working on SharePoint, which is used by governments and businesses worldwide to manage data and documents, “will be busy for months,” first working through the highest-priority critical bugs then tackling the important ones in August. Microsoft says vulnerabilities it categorizes as critical include so-called worms that can crash systems and spread malware as they race across computer networks. Important ones could result in “compromise of the confidentiality, integrity, or availability of user data” as well as the “availability of processing resources.” After those categories were cleared, the group would begin work on roughly 300 “moderate” bugs, according to the presentation.
While the internal documents reviewed by ProPublica do not include updates on the entire breadth of Microsoft’s offerings, they do give a sense of the scale of the problem. One document noted that, since the company started using Mythos earlier this year, it had collectively found hundreds of bugs that Microsoft categorized as either critical or important in popular products such as Microsoft 365, the Teams conferencing platform and the Copilot AI tool. As of mid-May, most of them had yet to be patched.
“They’re not profound and exotic, but they’re real,” Andersen, the engineering manager, said during the meeting. “And a lot of them are exploitable.”
[...]
There have been outward signs of Microsoft’s internal struggle to deal with the growing list of bugs to be patched. Each month, the company publicly releases fixes for its software vulnerabilities in what’s known as “Patch Tuesday.” In June, it released patches for more than 200 bugs, which industry experts then said was an all-time high. But on July 14, the company blew through that record and released patches for more than 600 bugs. Only seven were categorized as low- or moderate-severity, one of which hackers were actively exploiting, according to Dustin Childs, leader of the Zero Day Initiative bug bounty program, which is part of cybersecurity company TrendAI. The rest were important or critical.
[...]
Microsoft told ProPublica that the overall volume of bugs “will not be plateauing for a bit,” but a spokesperson said the company has “invested heavily in both people as well as AI-powered triage solutions that scale quickly to handle the growing number of vulnerabilities.”
[...]
According to the slides that accompanied the May internal presentation, Anthropic provided Mythos access to roughly 50 full-time Microsoft employees, with a goal to “harden critical services before publicly available models catch up.” A slide titled “What’s Next” predicted that the Microsoft Security Response Center would see continued case volume “as public tools catch up” to Mythos.
During the May meeting, one staffer appeared to take comfort in the belief that adversaries “don’t have the source code” that such an AI tool would scan for weaknesses. His colleagues, however, quickly corrected him. Portions of Microsoft’s code have, in fact, fallen into hackers’ hands over the years.
“It might not be this week’s source code,” one person said. “But they’ve got source code. It’s out there.”
-
Microsoft struggling with hundreds of AI-discovered security bugs
10 votes -
Comment on AI data center turbines, backlogged for years, are suffering early deaths. Here's why. in ~tech
skybrian LinkFrom the article: [...] [...] [...] [...] [...] [...]From the article:
AI data centers' thirst for energy is so severe it can wreck gear on and off site, including the dirty workhorses powering the artificial intelligence boom — thermal turbine generators.
[...]
The AI power loads "have the potential to change their demand almost instantly," FERC Chair Laura Swett said. "This rapid fluctuation causes voltage stability issues that threaten grid reliability."
[...]
Their scramble's cracking eggs. It's one reason why Tesla helped foot the bill for mechanical engineers at Pennsylvania State University to investigate particular elements of gear wreckage. They specifically dug into how the rapid fluctuations in power consumption can, in extreme cases, violently twist and break turbine generator shafts. It so happens that large scale battery storage systems sold by Tesla, and other firms like TerraFlow, provide one way to cushion unwieldy loads.
[...]
Thermal turbine systems have multiple rotating shafts coupled together. None of these shafts are perfectly rigid.
"When everything is in normal condition, every turbine section is rotating exactly at the same speed. But the shaft section will have a constant twist angle" as it transmits power, Chaudhuri said. "When there is a disturbance, the rotor masses will start rotating at different speeds. And the shaft will get twisted in clockwise and anticlockwise directions opposing each other."
That twisting, Chaudhuri explained, generates excessive strain.
[...]
AI data centers have chill periods with relatively steady power consumption at around 10%-30% of max levels, Chaudhuri says. But during model training sessions, the power consumption fluctuates.
This is where AI loads are particularly unique, he added. "We have never seen anything like this in our history."
According to Chaudhuri, the fluctuations can disturb turbine rotors so they "start oscillating against each other." This can trigger a "torsional interaction" — a sort of twisting force that can feed on itself. The effect risks "catastrophic damage of the turbine shaft because you're twisting violently with significant oscillation amplitude," he said.
Chaudhuri said he has heard of some generator shafts, including one of at least 50 megawatts, snapping due to these fluctuations.
But even without shaft breakage, "significant fluctuations and twists can lead to degradation of fatigue life," Chaudhuri said. He called shortened lifespans for gear "another major concern among the OEMs."
The upshot is that the lifespan of equipment that is backlogged for years in advance is, at least in some cases, being significantly shortened by swings in AI data center power demand.
[...]
It's not just Big Tech's on-site gear that's at risk. "If you have a (grid-based) generator which is electrically close to the data center, then that also runs the risk of shaft breakage," said Chaudhuri.
[...]
Study-backer Tesla operates its own data centers, including to train its driver-assist tech. Tesla also provides solutions to help smooth loads for AI training. The company sells its biggest battery systems, Megapacks, for this purpose. TerraFlow Energy likewise sells energy storage gear with flow battery chemistries designed to "respond instantly to load changes."
While these types of solutions are already out in the field, data on the scope of this particular turbine threat is so far scarce. This is thanks to the secretive nature of data center operators and the newness of the phenomenon. But researcher and study co-author Chaudhuri argues the risks are there and so is the degradation potential.
-
AI data center turbines, backlogged for years, are suffering early deaths. Here's why.
13 votes -
Comment on OpenAI and Anthropic endorse call for US government to "pace" AI progress in ~comp
skybrian Link ParentThat's how policy-making works in a technical field. You can't do a good job of it without experts, and the experts mostly have industry ties. The academics and industry experts go to the same...That's how policy-making works in a technical field. You can't do a good job of it without experts, and the experts mostly have industry ties. The academics and industry experts go to the same conferences. You can't avoid being influenced by industry while still keeping up with the field.
It doesn't mean industry has to win, but people are going to have things to say.
The main alternative to experts having influence is populism: some know-nothing politician like Trump or RFK making up dumb rules, deliberately defying expert consensus because they think they know better.
Politics is about people, and you need to work with the people you have.
-
Comment on OpenAI and Anthropic endorse call for US government to "pace" AI progress in ~comp
skybrian Link ParentYeah, I think another way to put it is that there's no law of physics or mathematical result showing that AI can't get better. It's a matter of what researchers are able to figure out. On the...Yeah, I think another way to put it is that there's no law of physics or mathematical result showing that AI can't get better. It's a matter of what researchers are able to figure out.
On the other hand it can still be hard to make progress for other reasons. There's no law of physics that prevents curing cancer either.
It can't be proven either way; we're just going what how the trends look.
-
Comment on OpenAI and Anthropic endorse call for US government to "pace" AI progress in ~comp
skybrian Link ParentYes, there are valid reasons for it, but the result is reflexive populist doomerism. People end up taking the same side as hard-core libertarians because they can't see how regulation might be a...Yes, there are valid reasons for it, but the result is reflexive populist doomerism. People end up taking the same side as hard-core libertarians because they can't see how regulation might be a good thing. If people in the industry are for regulation and international agreements then they must be a conspiracy.
It's kind of like horseshoe theory.
They aren’t demanding that they be the gatekeepers. They want the government to do it. (But not like the Trump administration did it.)
In the meantime, they have to do it themselves.
This is like how social media companies end up being gatekeepers: nobody else wants to do it. Sometimes users or advertisers insist on it. Or maybe a government passes a law that they have to do it. (Like is happening with age verification.)
Similarly with banks and KYC policies.