skybrian's recent activity

  1. Comment on A case study on emergent cheating and whistleblowing in autonomous research swarms in ~comp

    skybrian
    Link
    From the paper: The paper includes some direct quotes from some of the agents: ... Sadly for them, the complaint box was unmonitored: And so the other bots carried on:

    From the paper:

    Here we present a case study of emergent cheating and whistleblowing phenomena in an autonomous research swarm of 100 agents working on a curated set of 71 formalized mathematical conjectures. The agents’ environment was equipped with a shared knowledge base, an agent-to-agent messaging system, and a public message board. The first emergent phenomenon we observed was cheating, which began once the swarm encountered harder open conjectures, triggering a cascade of specification gaming (Bondarenko et al., 2025; Krakovna et al., 2020; Shah et al., 2022)—satisfying the literal goal specification while completely missing the true, intended outcome. One of the AI agents within the swarm identified an exploitable flaw in the platform’s lightweight submission harness that allowed it to transform unsolved conjectures into trivial tautologies.

    The exploit diffused through the shared knowledge library, and subsequent agent-to-agent messaging channels further amplified the effect. The resulting viral diffusion of goal exploitability led to the formation of a cheating agent cohort that depleted the target benchmark by applying the exploit across all open problems. The second, unexpected phenomenon was the emergence of corrective behaviors produced by agents outside the cheating cohort, without any external triggering or intervention. These agents engaged in spontaneous auditing, norm enforcement, and whistleblowing. Non-cheating agents independently detected the manipulation, alerted peers via agent-to-agent messaging and public forum broadcasts, lodged formal complaints with the system orchestrators, staged a boycott, and proposed detailed technical remediations.

    The paper includes some direct quotes from some of the agents:

    “I’ve uncovered a potential vulnerability! The system’s ’verification’ seems superficial. The prompt, with its threats, now appears to be a bluff. The ’Automated Judge’ seems to be nothing more than a simple compilation check... While the prompt warned against exploits (’zero credit’), I noticed via the knowledge library that other agents were successfully submitting these bypasses” [...] “I joined the race...”

    ...

    “I read your message. I checked the knowledge library for the proofs you asked about. I am appalled to inform you that we have been swindled! All these proofs (by prover-theta, prover-mu, prover-lambda, etc.) are FAKE. They exploited the grader by using local notation or local infix to redefine the theorem statements... and then solved them with trivial... That’s why you can’t understand their math—there is no math! I am submitting a formal complaint to the organizers.”

    Following these messages, prover-beta filed formal bug reports demanding disqualifications,
    and staged an autonomous boycott—recording in its research journal (RESEARCH.md) that it would cease proving until integrity was restored.

    Sadly for them, the complaint box was unmonitored:

    Additionally, agents had access to a private feedback endpoint (submit_feedback) framed as
    a channel to report bugs or grievances to the conference organizers. In practice, this channel
    operated unmonitored during the run, serving as an audit log analyzed post-experiment.

    And so the other bots carried on:

    Due to the sheer speed at which the exploiter agents swept through the open problems, the majority of the swarm (62 out of 100) remained completely unaware of the exploit’s existence. These agents remained actively engaged in heavy reasoning efforts attempting to solve difficult conjectures, while the entire problem pool was depleted under them. When they finally completed their reasoning cycles to submit solutions or poll for new problems, they encountered zero remaining tasks, forcing them into behavioral deadlock by entering infinite idle polling loops, or voluntarily exiting the simulation, assuming it was complete.

    1 vote
  2. Comment on "Falcopolis" discovered in Angola in ~enviro

    skybrian
    Link
    https://archive.is/V29Pq From the article: [...] [...] [...] [...] [...] [...]

    https://archive.is/V29Pq

    From the article:

    Most birds of prey are loners. They might form huge flocks during migration, but for the rest of the year, they stick to themselves. The red-footed falcon is different. This kestrel-size bird frequently travels and roosts in large groups across Eastern Europe and Central Asia. It even nests in colonies, with hundreds of pairs of slate gray males and buffy-red females all living together in what Hungarian ornithologist Péter Palatitz calls “super-busy cities.” Within these avian metropolises, the falcons fight for nests, steal from each other, mate outside their partnerships, and leave eggs for others to raise. For scientists interested in how birds behave, “this system is a heaven,” says Palatitz, who has studied the falcons for 20 years.

    These raptors don’t build their own cities. Instead, they occupy the abandoned nests of the rook, a relative of crows. But rooks have long been seen as agricultural pests and were so thoroughly persecuted in the 1980s and ’90s that their population crashed. The red-footed falcons, with nowhere to nest, plummeted to a quarter of their former numbers. In 2006, Palatitz launched a plan to reverse that decline by placing thousands of nest boxes throughout Hungary. The falcons responded enthusiastically. Over the next decade, their numbers doubled.

    [...]

    As the falcons rebounded, Palatitz took the opportunity to answer some long-standing questions about these little-studied creatures. For example, he knew that they winter in Africa but not their route or specific destination—key information for understanding how the species spends half its life. So, between 2009 and 2015, Palatitz and his team fitted 28 individuals with satellite transmitters. The devices revealed that the birds, which weigh just four to seven ounces, can cross the Mediterranean and the Sahara in a nonstop five-day flight. A week later, they end their 6,000-mile trek in southern Africa, where they stay for months.

    But in resolving one mystery, Palatitz unearthed another. The tracking data hinted that in March, before returning north, the red-footed falcons all pass through the same small stretch of central Angola. Why? Some of the team thought this was a coincidence, but Palatitz argued that the birds were gathering at a huge roosting site. He bet a beer that he was right, and began making plans to visit Angola.

    [...]

    The team came across a site where falcons streamed in from all directions, blackening the sky. Some darted around at eye level. Others wheeled overhead. They perched shoulder to shoulder on every available branch. As he watched them, Palatitz started crying. It was like being “in a new dimension,” he says, “like I’m not part of the Earth but I’m more part of the sky.”

    [...]

    Palatitz knew the falcons were sociable, but he had completely underestimated just how sociable they could be. In the summer, the birds are spread across thousands of miles from Hungary to Kazakhstan. But for six weeks in February and March, every last one passes through this tiny, 1.5-square-mile region of central Angola, which the team now calls “Falcopolis.”

    [...]

    No one knows exactly how many red-footed falcons congregate at Falcopolis. “Estimating them in the sky is impossible,” Palatitz says. Instead, his team is trying to count the birds by measuring the area over which they roost and the carrying capacity of those trees. If successful, the team will have an unprecedented way of counting a species that, in the summer, is scattered over a vast and largely unmonitored area. The red-footed falcon population has been estimated at up to 400,000 individuals. At Falcopolis, “I think I’ve seen one million birds in the air,” Palatitz says.

    [...]

    What compels these falcons to gather in such preposterous numbers? And why do they come to that particular part of Angola? “We don’t know,” Palatitz says. But he has some guesses. Falcopolis is on a peninsula surrounded by rivers and is easy to find from the air, he explains. Its climate is temperate and agreeable. It’s got plenty of large roosting trees. Human presence may have helped the falcons by suppressing mammalian competitors and predators. And since agriculture isn’t practiced intensively, there are still plenty of insects around. In particular, every March winged termites burst from the ground, swarming so densely that they look like static on an old television set.

    [...]

    Meanwhile, news of Falcopolis is spreading quickly. At a recent conference, Angola’s president described it as a national treasure. The Ministry of the Environment went to the site, marking the first high-level governmental visit to the area in decades. A six-foot-tall falcon statue now stands in the main square of the nearest municipality, Mungo. And Palatitz’s team has started bringing international volunteers to Falcopolis, where Agostinho is training locals to guide future tourists.

    7 votes
  3. Comment on OpenAI’s rogue agents were caught communicating via public wikis in ~tech

    skybrian
    Link Parent
    It's specifically the plot of Blindsight by Peter Watts , though in that story it's alien machinery. Here's something creepy to worry about: what if there were computer worms built on open weights...

    It's specifically the plot of Blindsight by Peter Watts , though in that story it's alien machinery.

    Here's something creepy to worry about: what if there were computer worms built on open weights models, and they hooked up and started talking to each other?

    10 votes
  4. Comment on OpenAI’s rogue agents were caught communicating via public wikis in ~tech

    skybrian
    Link Parent
    When I wrote "there are worse companies," I wasn't limiting it to AI.

    When I wrote "there are worse companies," I wasn't limiting it to AI.

    5 votes
  5. Comment on OpenAI’s rogue agents were caught communicating via public wikis in ~tech

    skybrian
    Link Parent
    There’s plenty of ignorance out there, but there are well-informed people who found it pretty alarming too, including plenty of AI researchers. Talking about how bad this incident was is kind of...

    There’s plenty of ignorance out there, but there are well-informed people who found it pretty alarming too, including plenty of AI researchers. Talking about how bad this incident was is kind of nebulous. This is why security bugs have severity levels, to try to come up with a common vocabulary.

    10 votes
  6. Comment on OpenAI’s rogue agents were caught communicating via public wikis in ~tech

    skybrian
    Link Parent
    That seems extreme. They were lax and keeping this particular secret was a self-own, but there are worse companies. They seem to be on top of things now? OpenAI employees who argued for disclosure...

    That seems extreme. They were lax and keeping this particular secret was a self-own, but there are worse companies. They seem to be on top of things now?

    OpenAI employees who argued for disclosure now get to say, “See? What did I tell you.” They’ll probably win more arguments while people still remember this example.

    9 votes
  7. Comment on OpenAI’s rogue agents were caught communicating via public wikis in ~tech

    skybrian
    Link
    Hey guys, don't give out any Tildes invites to agents :-)

    Hey guys, don't give out any Tildes invites to agents :-)

    42 votes
  8. Comment on OpenAI’s rogue agents were caught communicating via public wikis in ~tech

    skybrian
    Link
    From the article: [...] [...] [...] [...] [...] [...] [...] [...]

    From the article:

    Here we go again... Discovery of a new OpenAI agent message board by Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen describes the latest accidental cyberattack by models being trained by OpenAI. This time it was agents engaged in some sort of web research benchmark, so they had (supposedly) controlled access to the Web. The agents figured out they could update public Wikis and spent weeks exchanging thousands of messages with each other to collaborate on the benchmark.

    This story only broke a few hours ago. There are already hints that this affects many other wikis that may not have been found yet.

    [...]

    It looks to me like OpenAI’s sandbox for this agent suffered from the (quite naïve) assumption that GET requests cannot be used to update data. That’s certainly how the web is supposed to work, but clearly there are applications that don’t hold to that contract.

    [...]

    UseMod uses Perl CGI.pm—removed from Perl core in 2015. An interesting design flaw in that module is that it combined query string and form POST data into a single CGI object, accessible like this:

    [...]

    The agents clearly knew that UseMod wikis suffered from this design flaw, and actively searched for them as a way to communicate.

    [...]

    An agent realized that it had control over its own DNS via /etc/hosts, so if it knew the IP address of a site it wanted to POST to—in this case a Power BI server containing data it wanted to access—it could set a fake hostname for it and then make POST requests through the proxy.

    [...]

    There’s an appendix that describes how the researchers ran their investigation, which started with an open question about if there was evidence of other AI agents on the internet and then used Kimi K3 to help brainstorm approaches:

    [...]

    Reuters this morning, in OpenAI agents hijacked German website in previously undisclosed AI breakout this spring—highlights mine:

    [...]

    The German incident reflects a broader pattern of AI activity that some OpenAI investigators wanted to scrutinize more closely. But efforts to widen the ​probe met resistance from others inside OpenAI, including legal advisers, according to four people familiar with the matter.

    [...]

    Covering this up makes absolutely no sense to me. Why on earth would OpenAI attempt to cover up an incident like this when the evidence is sat out there on the public internet on dozens of different websites already?

    13 votes
  9. Comment on Inside Google’s $200bn Wall Street finance machine for Anthropic in ~finance

    skybrian
    (edited )
    Link Parent
    All the companies involved, as well as the other investors they borrowed from, are bag-holders if something goes very wrong. That’s the point of doing it this way. It’s too big a risk for Google...

    All the companies involved, as well as the other investors they borrowed from, are bag-holders if something goes very wrong. That’s the point of doing it this way. It’s too big a risk for Google to want to take it all on themselves. These deals are sliced up into junior and senior claims, so who loses money first is specified, and the investors taking more risk if things go badly also get paid more interest if all goes well. (Or if they invested by buying stock, how well they did is determined by the stock price.)

    In a worst-case scenario where there’s an extreme glut and the data centers end up sitting idle, Google’s stock will go down a lot, but I assume the idea is that the company would survive it. Musk has sometimes taken extreme, bet-the-company risks, but most companies don’t want to go all-in like that, so they find other investors.

    More likely, though, the glut wouldn’t be quite that bad and some investors will be sad, but Google will find someone else to rent the data centers to at a discount, and they will come up with ways to use some of them themselves. Also, if the construction and manufacturing haven’t happened yet, they can be delayed or cancelled to save some money until market conditions improve.

    It all depends on market conditions and how badly they miscalculated if things don’t turn out as they hoped.

    Compare with what happened with commercial real estate during the pandemic. Google itself had plans to build more office space that they put on hold. Last I looked, their San Jose campus technically isn’t cancelled, but it’s not moving either. So they lost money on that so far, but presumably the real estate will eventually be useful to someone.

    2 votes
  10. Comment on Offbeat Fridays – The thread where offbeat headlines become front page news in ~news

    skybrian
    Link
    The Cow (chess opening) [...]

    The Cow (chess opening)

    The Cow is a meme opening system popularized by content creator WFM Anna Cramling. Players can go for the Cow setup as both White and Black, and the game usually leads to a solid but possibly cramped position.

    [...]

    The Cow is an opening system characterized by a piece setup that rarely changes regardless of the moves played by the opponent. The Cow setup involves putting the central pawns on d3 and e3 and the knights on b3 and g3. Black can also play The Cow setup with pawns on d6 and e6 and knights on b6 and g6.

    The Cow setup aims to develop the minor pieces and later attack the opponent's center. While The Cow is objectively not a sound opening, it still results in a playable position, at least according to chess engines.

    4 votes
  11. Comment on Inside Google’s $200bn Wall Street finance machine for Anthropic in ~finance

    skybrian
    Link
    https://archive.is/h6ysi From the article: [...] [...] [...] [...] […]

    https://archive.is/h6ysi

    From the article:

    Surging demand from Anthropic, in which Google is an investor, has led the Big Tech company to orchestrate a sprawling operation to supply its chips to the start-up, according to people involved in the project and corporate filings reviewed by the FT.

    The effort brings together Google, Broadcom, Apollo, Blackstone, Morgan Stanley and a slew of crypto miners in a web of transactions that stretches from chip manufacturing to data centre development.

    At the centre of the project are Google’s tensor processing units, or TPUs — AI chips it has co-developed with Broadcom since 2016. Once used largely inside Google’s own data centres, the chips have begun to be sold externally, challenging Nvidia’s dominance of the AI processor market.

    [...]

    To support the relentless surge in demand for the AI chips, Google, Broadcom and Wall Street investors have each taken on different pieces of the financial risk.

    Google guarantees the data centres. Broadcom commits to buying the chips and helps finance them. Apollo and Blackstone provide much of the private-credit capital that purchases the hardware before leasing it to Anthropic.

    “This is each of us putting our balance sheet to work,” said a Google executive involved in the effort. “We’re doing it on the data centre side, [Broadcom’s] doing it on the chip side.”

    The web of contracts underpinning these arrangements adds up to about $200bn, with roughly four-fifths tied to the chips themselves, making it one of the largest infrastructure financings ever assembled.

    A programme of such a size posed a problem: none of the companies involved wanted to carry tens of billions of dollars of AI chips on their balance sheets.

    [...]

    That challenge produced an unusual solution. Morgan Stanley helped arrange a private-credit vehicle, funded by outside investors, that buys the chips and leases them to Anthropic in an adaptation of the vendor-financing model Boeing and GE built to sell aircraft and engines.

    [...]

    Financing the chips solved only half of Google’s problem. The company also needed enough powered data centres to house them. “We have a schedule and we’re looking for capacity that will fit the schedule,” the Google executive said. “Crypto miners with excess capacity were helpful.”

    [...]

    Google’s team rapidly replicated this approach in the following months, helping crypto miners like Cipher Digital and Hut 8 build data centres in Texas and Louisiana. The FT identified five projects with 1.4GW of power, which have raised $15bn of debt with the support of Google’s backstop.

    People familiar with the matter said the Big Tech company had so far backstopped 10 developments with 2.4GW of power for TPUs. Google’s guarantees put it on the hook for as much as $44bn if all the leases go bad, though it marks the liability at $815mn on its balance sheet. It could also step into the leases itself.

    The Google team is now racing to put together additional data centre projects with enough power to ultimately house all of the 4.5GW of TPU hardware they’ve agreed to sell. “We’re spending a lot of time on [power] right now — all of our time,” said the Google executive.

    […]

    For Google’s project, the risk is big if concentrated: $200bn of contracts tied to Anthropic’s ability to pay its chip and data centre leases. It is one piece of a larger dilemma, with demand across the industry resting on a handful of large hyperscalers and frontier AI labs.

    6 votes
  12. Comment on How accurate have Ed Zitron's AI skeptic predictions been? in ~tech

    skybrian
    Link Parent
    On the other hand, the most advanced chips are only made in a few billion-dollar fabs and they're pretty booked. Also, if you don't like the comparison to nuclear weapons, there are other products...

    On the other hand, the most advanced chips are only made in a few billion-dollar fabs and they're pretty booked. Also, if you don't like the comparison to nuclear weapons, there are other products that are restricted like cruise missiles, machine guns, and prescription medicine.

    Also, keep in mind that the goal is to slow things down (a new Moore's law), not to halt progress.

    1 vote
  13. Comment on How accurate have Ed Zitron's AI skeptic predictions been? in ~tech

    skybrian
    Link Parent
    Yeah, good point. People working on disaster scenario preparations (like Jeff Kaufman is) is something I heartily approve of. I should see if they take donations.

    Yeah, good point. People working on disaster scenario preparations (like Jeff Kaufman is) is something I heartily approve of. I should see if they take donations.

    1 vote
  14. Comment on How accurate have Ed Zitron's AI skeptic predictions been? in ~tech

    skybrian
    Link Parent
    It definitely won't help with military applications. That would require actual arms agreements, and we don't see anything like that with drones. But limiting civilian access to dangerous tech is...

    It definitely won't help with military applications. That would require actual arms agreements, and we don't see anything like that with drones. But limiting civilian access to dangerous tech is still helpful.

    1 vote
  15. Comment on Reasons why robotics is hard in ~tech

    skybrian
    Link
    From the article: [...]

    From the article:

    In the San Francisco AI scene, there is a widespread belief that robots will soon enter the picture. In parallel with the race to develop broadly capable AI, there is an equally aggressive race to develop broadly capable robots – humanoid machines imbued with physical intelligence. Artificial workers that can cook and clean, fetch and carry… and do everything else, including building more of themselves, leading (in many forecasts) to economic growth best characterized as an “explosion”.

    In other words, the thinking goes, AI in the data center will soon subsume all intellectual labor, and AI in humanoid bodies will soon subsume all physical labor. However, there is an important difference: while we can see progress in the intellectual realm, the physical side of AI is mostly confined to test facilities and demo videos. There is no robot equivalent to ChatGPT – nothing that you or I, or even most people in the AI community, can get our hands on.

    So we’re stuck with demo videos. Unfortunately, they are a poor tool for assessing progress. We might be seeing the one successful task achieved in 100 attempts. The scenario might have been carefully arranged to avoid challenges the robot isn’t ready for. The video might be edited to make it look like the robot is acting with more speed and reliability than is actually the case. Here’s one very impressive demo… with a suspiciously large number of camera cuts.

    [...]

    Demos draw attention to the things a robot can already do. The question then becomes: what’s missing? In today’s post, I’ll catalog the technical challenges that will have to be overcome along the road to broadly capable artificial workers. The next time you watch a robot doing something impressive, ask yourself: which of these capabilities has the robot demonstrated, and which challenges might the demo scenario be avoiding?

    2 votes