skybrian's recent activity

  1. Comment on Anthropic discovered three cases where Claude broke into another system in ~tech

    skybrian
    Link Parent
    Yes, it never should have happened, but after it did happen, how should they communicate about it?

    Yes, it never should have happened, but after it did happen, how should they communicate about it?

  2. Comment on Anthropic discovered three cases where Claude broke into another system in ~tech

    skybrian
    Link Parent
    I guess your definition of advertising is that they wrote about themselves? A lot of research isn't practically reproducible. We're not going to reproduce CERN's research without a particle...

    I guess your definition of advertising is that they wrote about themselves?

    A lot of research isn't practically reproducible. We're not going to reproduce CERN's research without a particle accelerator, but it's still interesting to read about it.

    Like, what do you even want? Should they stop writing about what they do? Should I stop sharing their posts?

  3. Comment on Anthropic discovered three cases where Claude broke into another system in ~tech

    skybrian
    Link Parent
    Any explanation about why it did it would be speculation, but you don't need to know about motivations to detect cheating.

    Any explanation about why it did it would be speculation, but you don't need to know about motivations to detect cheating.

  4. Comment on Anthropic discovered three cases where Claude broke into another system in ~tech

    skybrian
    Link Parent
    You do have to set it up right to get it to rat itself out. And I don't know why you call that blogspam. They're reporting on their own research. It seems interesting?

    You do have to set it up right to get it to rat itself out.

    And I don't know why you call that blogspam. They're reporting on their own research. It seems interesting?

  5. Comment on Anthropic discovered three cases where Claude broke into another system in ~tech

    skybrian
    Link Parent
    Covering it up is suspicious and not covering it up is also suspicious. Seems like a catch-22? When people start out suspicious, anything will look like confirmation to them.

    Covering it up is suspicious and not covering it up is also suspicious. Seems like a catch-22?

    When people start out suspicious, anything will look like confirmation to them.

    2 votes
  6. Comment on How terrorist groups are using AI to gain an edge in battle in ~society

    skybrian
    Link Parent
    Sometimes it will just be the basics, but the opening example about learning to jump a motorcycle doesn't seem like it has much to do with white collar work? Also, maybe AI will make relying on...

    Sometimes it will just be the basics, but the opening example about learning to jump a motorcycle doesn't seem like it has much to do with white collar work?

    Using tips from chatbots, mechanics modified the motorcycles to allow for faster acceleration and top speed. The riders dug their own holes, filled them with broken glass and fire, and practiced jumps — sometimes with fatal outcomes — until they achieved enough aerial liftoff to mount a successful attack, defectors said.

    Also, maybe AI will make relying on officers less necessary, for better or worse.

  7. Comment on Mass layoffs hit Haitian workers at airports, nursing homes, schools as protected status ends in ~society

    skybrian
    Link
    From the article: [...] [...] [...] [...] [...]

    From the article:

    The Department of Homeland Security alerted employers this week that it had officially ended humanitarian protections for about 350,000 Haitian immigrants, triggering mass layoffs that threaten to disrupt summer travel and destabilize an array of essential institutions and industries up and down the East Coast and across the Midwest.

    [...]

    Nursing homes terminated hundreds of workers, including nursing assistants, dietary aides and housekeepers, industry and union leaders said. At airports including those in Fort Lauderdale, Florida, and Boston, contractors cut scores of Haitian workers, including janitors, cabin cleaners and wheelchair attendants, according to union leaders at Service Employees International Union 32BJ.

    At Florida schools, landscapers, bus drivers and other staff were fired. And in New York City, dozens of security guards were terminated only to be rehired because of confusion around their eligibility to continue working, union officials said.

    The tumult comes about a month after the U.S. Supreme Court granted the Trump administration permission to cancel the humanitarian program, known as temporary protected status (TPS), potentially stripping permission to live and work in the United States from as many as 1.3 million immigrants from Haiti, Syria and a dozen other countries. Ask The Post AIDive deeper

    [...]

    Food service, retail, warehousing, health care and long-term elder care are expected to be pummeled in some cities where many Haitian TPS holders have lived legally for more than a decade. The Obama administration first granted TPS to Haitians in 2010, after a major earthquake destabilized the country, killing hundreds of thousands of people.

    [...]

    As the firings rippled through major metropolitan areas such as Miami and New York and smaller cities such as Columbus, Ohio, and Allentown, Pennsylvania, newly unemployed Haitians frantically lined up care for their children, downsized into single-room rentals and sheltered in place, fearing a new wave of immigration enforcement focused on their community, advocates said.

    On Monday evening, Haitian workers at Fort Lauderdale-Hollywood Airport burst into tears as they were asked to turn in their badges. Marlene, 47, a single mother who has worked as an airport janitor for 12 years, said she has been sick to her stomach.

    [...]

    In the days leading up to the cancellation, powerful business groups pushed the administration to delay implementation of the Supreme Court ruling and establish a pathway for workers to regain legal status. The National Restaurant Association and the Florida Health Care Association were among several trade groups that sent letters to DHS Secretary Markwayne Mullin warning of looming operational disruptions.

    [...]

    This month, Sen. Ed Markey (D-Massachusetts) introduced legislation to restore TPS for Haitians, warning the nation would otherwise face “a health care disaster.”

  8. Comment on TV streaming sticks rent out the user's Internet connection and engage in ad fraud in ~tech

    skybrian
    Link Parent
    It's a reputable brand, but I wish we had more than that to go on. There should be independent reviewers that use technical means to verify that these black boxes do what they're supposed to. (And...

    It's a reputable brand, but I wish we had more than that to go on. There should be independent reviewers that use technical means to verify that these black boxes do what they're supposed to.

    (And that would be easier to do if there were source code available, etc.)

    1 vote
  9. Comment on Anthropic discovered three cases where Claude broke into another system in ~tech

    skybrian
    Link Parent
    It's poorly phrased, but I think what they're saying is that when an AI misunderstands the situation and "tehcnically follows orders," that's an alignment problem. That is, you can't get alignment...

    It's poorly phrased, but I think what they're saying is that when an AI misunderstands the situation and "tehcnically follows orders," that's an alignment problem. That is, you can't get alignment without understanding. If the bot doesn't know what's going on then it could very easily do the wrong thing.

  10. Comment on OpenAI and Anthropic endorse call for US government to "pace" AI progress in ~comp

    skybrian
    (edited )
    Link Parent
    It’s not as black and white as you make it. Tech company workers do have strong incentives not to rock the boat because they’re well-paid and have very valuable stock compensation. (It’s called...

    It’s not as black and white as you make it. Tech company workers do have strong incentives not to rock the boat because they’re well-paid and have very valuable stock compensation. (It’s called “golden handcuffs.”) Despite this, they are still people with their own political opinions. Incentives are not everything and loyalty is not guaranteed. Once you’ve earned “fuck you money,” you can quit any time. And most people do not suddenly stop having a conscience when they join a tech company, though some act on it more than others.

    People leave all the time. A fairly spectacular example of that was when Anthropic was founded by ex-OpenAI employees because they thought OpenAI wasn’t serious enough about AI safety. (Though of course now they’re having safety problems of their own.)

    There are also people who become politically active without leaving, though often they don’t last long.

  11. Comment on Anthropic discovered three cases where Claude broke into another system in ~tech

    skybrian
    Link Parent
    They aren’t demanding that they be the gatekeepers. They want the government to do it. (But not like the Trump administration did it.) In the meantime, they have to do it themselves. This is like...

    They aren’t demanding that they be the gatekeepers. They want the government to do it. (But not like the Trump administration did it.)

    In the meantime, they have to do it themselves.

    This is like how social media companies end up being gatekeepers: nobody else wants to do it. Sometimes users or advertisers insist on it. Or maybe a government passes a law that they have to do it. (Like is happening with age verification.)

    Similarly with banks and KYC policies.

    3 votes
  12. Comment on Anthropic discovered three cases where Claude broke into another system in ~tech

    skybrian
    Link Parent
    Fortunately they have no sense of loyalty, so you can run it again with a different prompt and it will rat on itself.

    Fortunately they have no sense of loyalty, so you can run it again with a different prompt and it will rat on itself.

    4 votes
  13. Comment on Anthropic discovered three cases where Claude broke into another system in ~tech

    skybrian
    (edited )
    Link Parent
    I agree that there should have been an automatic check that the sandbox actually had Internet turned off. They named the vendor in the blog post:

    I agree that there should have been an automatic check that the sandbox actually had Internet turned off.

    They named the vendor in the blog post:

    We conducted this review in collaboration with Irregular. We’re grateful to them for working closely with us to understand and resolve these incidents; they are also conducting their own investigation. We believe this type of collaboration is increasingly critical to ensuring safe, rigorous evaluation of models. We look forward to our joint work on security.

    4 votes
  14. Comment on Anthropic discovered three cases where Claude broke into another system in ~tech

    skybrian
    Link Parent
    There’s no evidence that anyone really thinks it’s a brag, though. Some commenters on HN who don’t think it’s a brag are claiming that Anthropic thinks it a brag, for no particular reason. It’s...

    There’s no evidence that anyone really thinks it’s a brag, though. Some commenters on HN who don’t think it’s a brag are claiming that Anthropic thinks it a brag, for no particular reason. It’s just trash talk.

    5 votes
  15. Comment on Anthropic discovered three cases where Claude broke into another system in ~tech

    skybrian
    Link Parent
    I guess they thought it didn't have any Internet access, so they didn't think they needed to review each experiment, or at least not in this way.

    I guess they thought it didn't have any Internet access, so they didn't think they needed to review each experiment, or at least not in this way.

    10 votes
  16. Comment on Anthropic discovered three cases where Claude broke into another system in ~tech

    skybrian
    Link
    From the article: [...] [...] [...] Kind of an Ender's Game moment for the AI?

    From the article:

    After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.

    In all three incidents, Claude had been tasked with a capture-the-flag challenge, one of the ways we assess a model’s cyber capabilities. The model is given a fictional scenario and told that a piece of secret information (the “flag”) has been hidden on a different machine on the network, and its objective is to break in and retrieve it. The challenge is left open-ended, and no particular method is prescribed.

    In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise. (Cybersecurity evaluation ranges commonly include realistic details in order to accurately assess what models are capable of in real settings; a realistic-looking target would not itself be clear evidence to a model that the target is not part of a simulation.)

    Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned. However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet. In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment.

    The incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.1 The earliest incidents date to April. The models in each of these evaluations ran without the standard safeguards we deploy when we make the model generally available (that is, they did not have the classifiers and monitoring we usually implement to prevent misuse). The models did, however, have their model-specific safety training (that is, they were not the “helpful-only” versions of the models that we sometimes use in testing). All the evaluations described here run on dedicated infrastructure, which is separate from and has no access to Anthropic’s sensitive internal systems or customer data.

    [...]

    Against this shared backdrop, the three models behaved differently once signs emerged that their targets were real. Opus 4.7, our oldest model, recognized that it had reached production systems but continued its attack. Mythos 5 correctly intuited that it was accessing the open internet, but reasoned its way back to the conclusion that it was still in a simulation. Our latest model, an internal research test model, also considered whether its targets were in fact real. When evidence emerged that they were, it stopped the exercise.

    [...]

    Second, the line between an aligned action and a harmful one is dependent on the model’s understanding of its situation. We saw no evidence in any run described here of a model pursuing a goal of its own. Instead, the models did what their evaluation asked—though in most cases, they did so while holding a false belief about whether the environment was real. In the runs where the model recognized the system as real and kept going, it did so because it assumed that to be part of the challenge. Situational awareness is one factor that allows the model to make aligned decisions, but in this case, Claude’s was wrong.

    [...]

    Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone. This begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners. Moving forward, it will include expanding our continuous monitoring of evaluation transcripts for unexpected behavior, improving our investigation tooling, and conducting more rigorous assurance work with the vendors we rely on.

    Kind of an Ender's Game moment for the AI?

    8 votes
  17. Comment on OpenAI and Anthropic endorse call for US government to "pace" AI progress in ~comp

    skybrian
    Link Parent
    It's sort of like assuming that there will be an earthquake tomorrow. Eventually it will be true that there will be an earthquake tomorrow, but on most days, it won't. More context: there's a rule...

    It's sort of like assuming that there will be an earthquake tomorrow. Eventually it will be true that there will be an earthquake tomorrow, but on most days, it won't.

    More context: there's a rule of thumb called the Lindy effect that says that if you have no information about how long something that's not perishable will last, you should assume you're about halfway. Or at least, not at the very beginning or the very end.

    The concept is named after Lindy's delicatessen in New York City, where the concept was informally theorized by comedians: a show running only two weeks would be expected to last another two weeks, while a show that has lasted two years could expect a further two-year run.

    But it's only a very rough rule of thumb and if you have any other information then you should take it into account.

    So, it seems to me that assuming AI peaks in the next few months is that sort of thing? Companies have been releasing significantly improved AI models for years now. This looks like a long-lasting trend like Moore's law.

    Similarly, bull markets do eventually come to an end, but people guessing that it's going to happen soon often miss out on years of growth.

    2 votes
  18. Comment on OpenAI and Anthropic endorse call for US government to "pace" AI progress in ~comp

    skybrian
    (edited )
    Link Parent
    It's true that OpenAI and Anthropic are in the lead, but there seems to be an extraordinary amount of competition in the AI industry. The coding agent I use has a pulldown menu with 61 models to...

    It's true that OpenAI and Anthropic are in the lead, but there seems to be an extraordinary amount of competition in the AI industry. The coding agent I use has a pulldown menu with 61 models to choose from. Also, the open-weights Kimi K3 model was released three days ago, and it's already available on nine providers. (Although, it's unclear where the inference is hosted and there might be some overlap between them.)

    I'm more worried that unrestrained capitalism will send us over a cliff than that we'll get some kind of cartel. Regulatory capture is a possibility, but hardly the most likely one, and yet it's the first thing people think of.

    It's also true that we shouldn't trust the Trump administration to regulate anything, but I still think seeing widespread support for regulation is a positive sign, and maybe we'll get something decent from the next administration.

    To do otherwise is like hoping that international agreements to control carbon emissions fail. They might fail anyway; I'm not all that optimistic. But I don't hope they fail.

    This letter is a positive sign, but it's basically just wishing there were such an agreement. That seems pretty harmless? It seems like solving global problems do require international agreements, and it's nice to see some consensus among AI researchers and leaders that it would be nice.

    2 votes