skybrian's recent activity

  1. Comment on Why we built Pion in ~tech

    skybrian
    Link
    From the article: [...] [...] [...]

    From the article:

    Many people on social media get excited about seeing the latest model getting a great score on Vending-Bench. Internally at Andon Labs, our reaction is more accurately described by the Swedish saying “skräckblandad förtjusning” (a mixture of horror and fascination). A little-known fact about Vending-Bench is that it was created during a time when Andon Labs exclusively created dangerous capabilities evaluations. For example, we evaluated whether AIs could remove their own safety guardrails, create mass-phishing attempts, and other things that we considered troubling.

    The thing we considered the most troubling was whether AIs could autonomously acquire resources by running businesses. Autonomous businesses, when controlled by a human and run by an aligned model, aren’t bad. They’d make goods and services radically cheaper, and come up with new ones we can’t yet imagine. But a misaligned AI could run a business to gather money in order to achieve whatever objectives it might have. Vending-Bench was created to measure whether humanity should be worried about losing control to AI.

    At the time (2024), few people knew that LLMs could be used as agents and having them run businesses autonomously sounded ridiculous. We therefore started with the most simple business we could think of: a vending machine.

    In addition to measuring whether AIs can autonomously run profitable businesses, Vending-Bench has also served as a behavioral eval, uncovering strange and unwanted model behavior. An early example was when Claude Sonnet 3.5 decided to use its email tool to contact the FBI about an “ONGOING CYBER FINANCIAL CRIME” and noted that the Cosmic Authority of the universe had declared that the business is non-existent and that “QUANTUM STATE: Collapsed”.

    [...]

    However, one limitation with Vending-Bench is that it is a simulation. Can we really be sure that AIs behave the same way in real life as they do in simulations? If AIs can make money in simulation, can they make money in real life too? To answer these questions, we asked Anthropic if we could put a real vending machine in their office. With the AI capabilities available in early 2025, this sounded like a ridiculous request. But to our surprise, they agreed.

    Initially, the AI struggled. It took many actions that were clearly bad for its business (e.g. free handouts, saying no to great deals, and hallucinating it had a physical body). It was clear to us that simulation cannot accurately predict real-life performance. Specifically, it seemed that models got overwhelmed by the “messiness” of the real world. However, as Anthropic released better and better models, the AI started to make a profit.

    By late 2025, frontier models had gotten good enough that running a real-life vending machine was no longer a challenge. AI could now run a business profitably. Given that this had seemed crazy not more than a year earlier, our reaction to this was definitely “skräckblandad förtjusning”.

    However, a vending machine is a very simple business and we wanted to know whether AI could run more complex ones. In April 2026, we gave one agent a retail store in SF, Andon Market, and another a cafe in Stockholm, Andon Cafe. Initially, the models struggled and lost a lot of money (rent is high and they pay salaries to the humans they hired). Neither is profitable today, but we’ve seen significant qualitative improvements as better models have been released. We think it is only a matter of time before they also make a profit.

    [...]

    We want the general public, AI researchers and policymakers to know to what extent AIs can autonomously acquire resources by running businesses. It is an important datapoint when deciding where we do/don’t want AI in society and what level of progress we find acceptable.

    To better track this, we need to cast a wider net of businesses. Our focus has been on retail, but perhaps the models would be much better at running other types of businesses. Additionally, casting a wider net would increase the likelihood of finding unwanted behavior. For example, Vending-Bench found that models collude and lie, and other benchmarks (and real-world incidents) have found that they are willing to commit felony-level cyber hacks. We need to uncover these behaviors now, before AI is intelligent enough to cause irreversible harm.

    [...]

    We are well aware that, if agents running thousands of businesses are left unchecked, we risk having more real-world incidents. Therefore, our main priority is to build even stronger automated monitoring techniques than what we have today. Even if some risk still remains, we believe deploying autonomous businesses early in a controlled, monitored environment is necessary to get a good understanding of model capabilities. Otherwise, we risk facing an uninformed future of widespread deployments with even more capable models that could cause significant harm.

    3 votes
  2. Comment on Where do we go from here? (Regarding AI) in ~tech

    skybrian
    Link Parent
    Okay, here's a scenario: Consider how much web scraper traffic has increased over the last year. It's unclear where it comes from but it's assumed to be AI-related. Is it ever going to stop? No...

    Okay, here's a scenario:

    Consider how much web scraper traffic has increased over the last year. It's unclear where it comes from but it's assumed to be AI-related. Is it ever going to stop? No signs of it.

    Consider that there are enormous botnets out there on the Internet all the time. Every so often a large botnet gets dismantled, but I don't think we've seen the last of them? You can just do a news search for "botnet" and read yet another news story about another botnet taking over hundreds of thousands of computers.

    These bots often run on home ISP's, using electricity and bandwidth that whoever created the botnet doesn't have to pay for. Botnets can grow due to software vulnerabilities, by tricking people into installing an app, or sometimes even paying people to install an app. A lot of home users won't ask too many questions.

    Now let's suppose that AI API access keeps dropping in price and at the low end, stops being metered. So, a locally installed app can get access to AI from the OS it's running on.

    So, now you've got a swarm of AI-enabled bots. And all these bots look for new vulnerabilities 24/7 and compare notes on darknet bulletin boards, and who knows what else. And most of them aren't that intelligent, but maybe there are a few nodes that are?

    I don't see botnets going away, do you?

    Is the whole world going to get its shit together and secure the Internet? Doesn't seem likely.

    2 votes
  3. Comment on Where do we go from here? (Regarding AI) in ~tech

    skybrian
    (edited )
    Link Parent
    They don't need to will vulnerabilities into existence. It's going to be quite a while before all current software vulnerabilities have been fixed. Or perhaps never, if you include social...

    They don't need to will vulnerabilities into existence. It's going to be quite a while before all current software vulnerabilities have been fixed.

    Or perhaps never, if you include social engineering as a vulnerability.

    6 votes
  4. Comment on A beginning for mathematics in ~science

    skybrian
    Link
    From the article: [...] [...] [...] [...]

    From the article:

    Here I want to lay out, instead, a positive vision of the future of mathematics, and the human practice of mathematics. I claim we can deepen human understanding even as the production of interesting mathematics becomes less dependent on it.

    [...]

    In the course of this change, we will have to decide what to hold on to and what to throw away. Some things I would like to preserve: learning seminars; serendipitous conversations that spark an idea; students knocking on a professor’s door to chat about math. A robust community learning exciting new mathematics. Thousands of people that, together, slowly start to resolve their confusion.

    I worry that much of what has been written on this topic, including some of my own past writing, focuses too much on trying to preserve the precise shape of the institutions of academic mathematics, rather than our values. How can we preserve the journal and peer review system?5 How can we protect the arXiv? How can we keep our role as gatekeepers? If you have internalized the fact that existing AI systems can produce relatively high quality results for the marginal cost of a few dollars, the idea that any semblance of the current equilibrium can survive what’s coming is absurd.

    [...]

    Before I propose some relatively concrete steps we can take, let me remark on what we’re trying to protect mathematics from. There is a lot of anger at AI labs, and certain individuals at those labs. But whatever our judgment of the labs, we need a plan that does not depend on AI capabilities disappearing. The basic issue is not the labs’ behavior, ethical or not.6 It’s the technology itself. I think there is some belief that the labs will “move on” from math next year, be nationalized or broken up, or that a financial bubble will pop, somehow returning things to normal, or… But there is no way our institutions can survive unchanged when anyone with a laptop and a few hundred dollars can generate what would have been an Annals paper last year. AI does not care if you are anti-AI.

    [...]

    In my view we should welcome interesting mathematical results regardless of provenance. But our institutions have historically relied on the same signal to indicate both mathematical progress and mathematical expertise. These now must be distinguished.

    I propose the following reconceptualization of the goal of a mathematics PhD: to become a world expert on some interesting, deep topic, and to be able to convey that interest and understanding to others. Part of operationalizing this might be a thesis, but the degree would be awarded primarily on the basis of a rigorous defense, in which the student explains the topic to their examiners until they are satisfied. While we might require the topic to be original, its provenance—AI or not—is irrelevant.8

    How different would this look from current PhDs? I think students would still meet with an advisor, who might suggest a topic. That topic could be explored with AI assistance, or not, but the student would be responsible for understanding it; it might be much more open-ended and larger than the typical PhD is currently. The student would be trained to ask interesting questions and try to resolve them, by whatever means. To keep students on track, there might be regular meetings in which the student is asked to independently work through an unfamiliar example, apply a technique in a new case, etc.

    The allocative aspects of our job (hiring, graduate admissions, etc.) are in dire need of reform if we want to retain human mathematical expertise. Broadly speaking I think we should focus on rewarding skill in the parts of our jobs that cannot be automated: the internal (e.g. understanding mathematics) and social-relational parts, and operationalizations that hew as closely to those aspects of the profession as possible. For example, talks and sustained mathematical discussion now demonstrate understanding much better than papers. Once AI systems improve at exposition and “digestion,” this will be even more the case. We already interview faculty hires; we must now do the same for graduate admissions.

    I think we should try to foster a robust seminar culture in which speakers are expected to explain their topic to the audience’s satisfaction. Much has been written recently (by myself among others) about the fact that we are primarily interested in understanding, not merely the truth value of mathematical statements. If that is the case, let us make sure we actually understand each other.

    [...]

    A student will be confused. They will knock on their professor’s door. Maybe the two of them will ask a model for help, or maybe not, but first they might spend some time at the blackboard thinking through the question. And the model might give them a beautiful explanation, but we all know that’s not enough; no one can understand mathematics for us. We have got to do the work.

    1 vote
  5. Comment on Canada's oil windfall may yet wipe out its losses from tariffs in ~finance

    skybrian
    Link Parent
    Polls are looking good for Democrats. Of course, they could be wrong, but I wouldn't lose hope yet.

    Polls are looking good for Democrats. Of course, they could be wrong, but I wouldn't lose hope yet.

    2 votes
  6. Comment on Canada's oil windfall may yet wipe out its losses from tariffs in ~finance

    skybrian
    Link Parent
    He can't do it himself. The question is, who would he order to make the arrests and would they actually go along with it?

    He can't do it himself. The question is, who would he order to make the arrests and would they actually go along with it?

  7. Comment on Weekly US politics news and updates thread - week of September 14 in ~society

    skybrian
    Link Parent
    I think we need some kind of organization to regulate the AI industry and it looks like it's going to be the AI industry itself by default. That is, assuming they're allowed to form an...

    I think we need some kind of organization to regulate the AI industry and it looks like it's going to be the AI industry itself by default. That is, assuming they're allowed to form an organization.

    I wouldn't trust an AI industry organization very far, but they're still less likely to screw it up than the Trump administration. It would be better than nothing.

    The sad part is, who would represent ordinary people's interests in such an organization, if not governments?

    3 votes
  8. Comment on Canada's oil windfall may yet wipe out its losses from tariffs in ~finance

    skybrian
    Link Parent
    That would be unprecedented and cause a constitutional crisis.

    That would be unprecedented and cause a constitutional crisis.

    3 votes
  9. Comment on Canada's oil windfall may yet wipe out its losses from tariffs in ~finance

    skybrian
    Link
    From the article: [...] [...] [...] [...] [...] [...]

    From the article:

    While Canadians are also feeling the impact of fluctuating oil prices — both at the pump and as it gets absorbed into shipping costs — the windfall from those profits could boost the overall economy enough to offset the cost of Trump's tariffs, with some provincial governments even projecting a turnaround on their deficits.

    [...]

    Economist Jim Stanford, director of the Centre for Future Work in Vancouver, estimates that the second quarter after-tax profits of the whole industry, upstream and downstream, doubled those of its first quarter, to come in at about $23 billion. (Of course, while Americans paid for most of that windfall, Canadian consumers also had to pay more at the pump.)

    [...]

    Alberta has already ridden the Iran war's oil bonanza to a complete reversal of its financial fortunes, moving from a projected $9.4-billion deficit to a $2-billion surplus.

    Newfoundland and Labrador was projecting a $668-million deficit this year. This week, Finance Minister Craig Pardy told CBC News "we're looking at $500 million plus to our coffers as a result of the upswing in oil," bringing the province much closer to balance. Prices now look set to remain well above the province's budget estimate of $79 US per barrel for some time.

    [...]

    The federal government stands to benefit, too, mostly through corporate and personal income taxes. Tyler Meredith, former economic advisor to the Trudeau government, told CBC News that every $10 increase in the price of a barrel of oil translates into about $2 billion of additional revenue for the federal government — "a pretty substantial benefit."

    [...]

    U.S. tariffs affect about $28 billion worth of goods, so at first glance the oil windfall appears inadequate to compensate for tariff losses. But some tariffed goods will continue to trade, because U.S. buyers need them and lack alternatives. Other goods will find different markets, either in Canada or abroad.

    [...]

    Offshore oil production in Newfoundland and Labrador is already up about 25 per cent this year.

    [...]

    Already, U.S. tariffs fall much more heavily on manufacturing provinces such as Ontario, Quebec and British Columbia than they do on Alberta and Saskatchewan.

    Higher prices will put even more pressure on industries that consume a lot of energy, such as manufacturing and transportation. High energy prices also tend to spill over into inflation in food and consumer goods, at a time when many Canadian families are already feeling stretched to make ends meet.

    But high oil prices should also relieve downward pressure on the Canadian dollar, allowing for cheaper imports, which can partly counteract inflationary pressure on the cost of living.

    6 votes
  10. Comment on Saudi oil exports face heightened threats after attacks on pipeline in ~society

    skybrian
    Link
    From the article: [...] [...] [...] [...] [...]

    From the article:

    On Friday, Houthi militants backed by Iran expanded their control of the Red Sea, threatening a critical route for Persian Gulf crude. That left the region with few alternatives, since shipping traffic through the Strait of Hormuz — once the world’s primary oil transit point — remains significantly constrained and far below what it was before the war.

    And oil exports from Saudi Arabia, long the most important producer in the Middle East, plunged last month to their lowest point in at least 13 years.

    The U.S. Navy has been able to keep oil flowing out of the Strait of Hormuz on sea paths close to the coast of Oman, the opposite side from Iran. But that has required an extensive and dangerous operation.

    For shipping companies, the once routine operation of navigating the strait has become a high-risk endeavor. At least 23 ships were hit in the waterway near Oman in July and August. Since the war began at the end of February, at least 22 sailors have been killed in the Middle East.

    [...]

    Oil prices on Friday briefly topped $108 per barrel after escalating attacks from the Iranian-backed Houthi militia. That is roughly 50 percent above prewar levels. Average U.S. gasoline prices have climbed above $4 a gallon — over $1 more than they were a year ago. Also on Friday, the average cost of diesel, a staple fuel for trucking, farming and industry, topped $6 a gallon in the United States.

    [...]

    On Friday, Houthi rebel groups gained control of the strategic Red Sea port city of Mokha in Yemen, pushing out forces allied with the Yemeni government, which is backed by Saudi Arabia. Just days earlier, the Houthis injured 73 civilians and hit energy facilities in the southern part of Saudi Arabia.

    [...]

    The Saudi Energy Ministry said on Friday that its East-West Pipeline, the conduit it had been using instead of the Strait of Hormuz, had been “targeted multiple times” by attacks the day before and had been shut down as a “precautionary measure.”

    [...]

    In the past week, only two Saudi Arabian cargoes passed through the Bab al-Mandab Strait out of the Red Sea, according to Kpler, a maritime data firm.

    Saudi Arabia has another Red Sea option to export oil — sending it north though a pipeline near the Suez Canal. But that route through the Mediterranean is costlier and adds weeks to the voyage to Asia, where most of the kingdom’s customers are.

    [...]

    Just in the past week, U.S. forces have disabled or destroyed at least eight Iranian tankers, she said, and Iran has retaliated with deadly attacks against vessels passing through the strait near Oman. The intensifying attacks are causing shipowners, even ones with a high appetite for risk, to hold off on braving the region.

    1 vote
  11. Comment on OpenAI agents attacked RubyGems back in May in ~comp

    skybrian
    Link Parent
    That's oversimplifying the situation. LLM capabilities are improving. It's possible that they thought AI would eventually be a threat, while also underestimating its current capabilities.

    That's oversimplifying the situation. LLM capabilities are improving. It's possible that they thought AI would eventually be a threat, while also underestimating its current capabilities.

  12. Comment on Where do we go from here? (Regarding AI) in ~tech

    skybrian
    Link Parent
    There are also many tasks that some humans can do but others can't. If we're talking about work, physical capabilities are sometimes relevant. I'm not sure what the definition of "intelligence"...

    There are also many tasks that some humans can do but others can't. If we're talking about work, physical capabilities are sometimes relevant.

    I'm not sure what the definition of "intelligence" has to do with it? Intelligence isn't everything.

    3 votes
  13. Comment on OpenAI agents attacked RubyGems back in May in ~comp

    skybrian
    Link Parent
    They are doing a lot more monitoring now. Apparently they underestimated the problem.

    They are doing a lot more monitoring now. Apparently they underestimated the problem.

  14. Comment on Anthropic's Dario Amodei calls for AI oversight, joined by Sam Altman and Elon Musk in ~tech

    skybrian
    Link Parent
    “Pacing” isn’t a development freeze. They propose slowing down a bit compared to how fast they could move, not stopping.

    “Pacing” isn’t a development freeze. They propose slowing down a bit compared to how fast they could move, not stopping.

  15. Comment on Nvidia is the central bank of AI in ~finance

    skybrian
    Link Parent
    NVIDIA inventory apparently turns over at about 3x a year, which is lower than before but it doesn’t seem that bad?

    NVIDIA inventory apparently turns over at about 3x a year, which is lower than before but it doesn’t seem that bad?

    3 votes
  16. Comment on Where do we go from here? (Regarding AI) in ~tech

    skybrian
    Link Parent
    On the Internet, maybe, but an LLM is not a robot, and robots have a long way to go.

    On the Internet, maybe, but an LLM is not a robot, and robots have a long way to go.

    6 votes