skybrian's recent activity
-
Comment on Some thoughts about Anthropic’s new cryptanalysis results in ~comp
-
Some thoughts about Anthropic’s new cryptanalysis results
1 vote -
Comment on Curing concrete - improving the carbon costs of an essential material in ~enviro
skybrian LinkSide comment: that's an awful lot of tags! Are you autogenerating them somehow?Side comment: that's an awful lot of tags! Are you autogenerating them somehow?
-
Comment on Finding bugs in Raft implementations in ~comp
skybrian Link ParentYes, I also avoid testing against mocks when testing against a real implementation is practical. (Databases and browsers come to mind.) On the other hand, testing in a real environment won't tell...Yes, I also avoid testing against mocks when testing against a real implementation is practical. (Databases and browsers come to mind.)
On the other hand, testing in a real environment won't tell you whether your code is portable or conforms to a standard. Testing against multiple implementations is good for comparison. Maybe you could test against real browsers and also a theoretically ideal, abstract browser? That's where a proof might be useful, so you tell when your implementation is theoretically correct, but reality is letting you down.
-
Comment on US schools are adding pepper-spraying drones to help combat active shooters in ~society
skybrian Link Parent"Perception is reality" is overstating it. Yes, sometimes it happens that believing in something makes it true. This is why it's hard to beat the front-runner in an election. But after a market..."Perception is reality" is overstating it. Yes, sometimes it happens that believing in something makes it true. This is why it's hard to beat the front-runner in an election. But after a market crash, we discover that, while getting people to believe in things can indeed be quite powerful, it doesn't necessarily make them true.
Also, as every failed political candidate knows, getting millions of people to believe something is not that easy. Unless they already want to believe.
-
Comment on US schools are adding pepper-spraying drones to help combat active shooters in ~society
skybrian (edited )Link ParentWe are always changing society but sometimes it’s a side effect, which seems easier than a deliberate change. Side effects are like like water running downhill rather than having to pump it...We are always changing society but sometimes it’s a side effect, which seems easier than a deliberate change. Side effects are like like water running downhill rather than having to pump it uphill. Even for deliberate attempts at changes like ads, the ones that take advantage of existing weaknesses will be more effective, unfortunately.
If you have a drone, maybe you can get by without the cop? They seem like competing ways to spend money on security. Hiring people is expensive, too.
I can imagine a drone being marketed as a way to cut labor costs. Maybe it would be a higher up-front cost, but eventually pays for itself assuming there would be a cop otherwise.
Though of course they can do different things.
-
Comment on US schools are adding pepper-spraying drones to help combat active shooters in ~society
skybrian Link ParentI think the trouble is that we are guessing about root causes and how to address them. Well, there is gun control, but that’s politically very difficult. And if your plan is to first fix Congress...I think the trouble is that we are guessing about root causes and how to address them. Well, there is gun control, but that’s politically very difficult. And if your plan is to first fix Congress or fix society then the trouble is that it’s an “if everyone will just” sort of plan.
So, instead let’s go with the unproven techno-fix because that just involves spending money and doesn’t require changing American society. The world gets a little more sci-fi because just about anything is easier than changing society.
-
Comment on US Federal Communications Commission bans foreign-produced solar inverters in ~enviro
skybrian LinkFrom the article:From the article:
The Federal Communications Commission Public Safety and Homeland Security Bureau added foreign-produced power inverters to its Covered List, triggering an immediate and absolute ban on new equipment sales in the United States.
This means any solar or battery storage project relying on an inverter made outside the United States that has not already received an official FCC ID cannot legally turn on or interconnect. Because there is no phase-in period, grandfather clause, or grace window, the regulatory pipeline is frozen today, forcing developers to halt active procurements and find new hardware vendors.
The action effectively overrides the Department of Energy analysis from January 2026, which inspected 30 Chinese inverters and found zero evidence of malicious hardware. The White House interagency council determined that physical bugs do not matter because the risk is purely digital. The administration ruled that the wireless connectivity inherent in modern smart inverters allows foreign adversaries to push firmware updates that could shut down solar arrays remotely, making all foreign-assembled units an unacceptable threat to the critical power grid.
The immediate practical result is a massive equipment shortage that will delay upcoming commercial and utility projects. Department of Energy data shows that domestic manufacturers supply only seven percent of the U.S. solar inverter market, leaving a 93% deficit that cannot be filled by local factories anytime soon.
The hardware blockade hits right as developers plan to connect more than 58,000 MW of new solar and storage over the next year. Without certified inverters, fully built solar farms will sit dark and unable to feed electricity to the grid.
-
US Federal Communications Commission bans foreign-produced solar inverters
26 votes -
Comment on The apples and oranges tribunal in ~society
skybrian LinkFrom the article: [...] [...] [...] I'll add that market prices are often wrong (inconsistent). The market gives an answer, but nothing says it's a correct answer. It's what we have.From the article:
Suppose that apples sell for more than oranges and Parliament in [its] wisdom decides that, at last, apples and oranges must be compared. Not by shoppers — shoppers are biased, they merely reveal what they are willing to pay — but by a tribunal, which will determine whether apples and oranges are of truly equal value and thus must sell at the same price.
What would the tribunal need to know?
[...]
To determine the “just” price of apples and oranges, the tribunal would need the entire general-equilibrium system.
[...]
Britain is now running this experiment in the labor market–Is a retail worker equal to a warehouse worker? A canteen worker equal to a coal miner? A dinner lady equal to a gravedigger?
Under the Equality Act’s “equal value” provisions, tribunals compare jobs by scoring their intrinsic properties — effort, skill, responsibility, working conditions — the labor theory of value applied to labor. How is it going? The Tesco litigation began in 2018; the tribunal’s fact-finding hearing ran 36 days, its judgments run to more than 900 pages resting on some 19,000 pages of training manuals, and the independent experts have yet to begin the report that will actually say whether a shelf-stacker’s job equals a warehouse worker’s. Eight years, and the calculation has not started. Apples and oranges, adjudicated but not, as Orwell or Marx or Stafford Beer might have imagined, by a industrial bureaucracy or by an all-knowing artificial intelligence but by lawyers and commissions and tribunals. The worst of all worlds.
And having discovered that the tribunal cannot price two jobs in a decade, the government now proposes to add race and disability comparisons and an enforcement unit to publish official guidance on which reasons for a wage difference are permissible. A bureau of allowable scarcities.
[...]
The market compares apples and oranges every day. It is the only institution that can. But there is a deeper error here than computation. Suppose the tribunal succeeded. Suppose that after another decade it delivered the true and final score, shelf-stacker versus warehouseman. What would it have found? Not justice. A wage is not a grade on your character or a measure of your worth as a human being. A wage is a price — a report on how scarce your skills are relative to the desires of people you will never meet. Nurses are not morally less worthy than plumbers should they earn less than plumbers or vice-versa, and no one thinks otherwise except the tribunals.
I'll add that market prices are often wrong (inconsistent). The market gives an answer, but nothing says it's a correct answer. It's what we have.
-
The apples and oranges tribunal
4 votes -
Comment on A technical timeline of the July 2026 attack on Hugging Face in ~comp
skybrian LinkFrom the article: [...] [...] [...] The third-party provider was an unsecured sandbox on Modal.From the article:
Over roughly two and a half days inside our infrastructure, an autonomous AI agent driven by a combination of OpenAI models ran an end-to-end intrusion against our platform: it was thousands of small, automated decisions, executed at machine speed across short-lived sandbox environments, with command-and-control staged on ordinary public web services.
[...]
Our forensic reconstruction covers ~17,600 attacker actions that we were able to recover, grouped into ~6,280 clusters, between 2026-07-09 02:28 UTC and 2026-07-13 14:14 UTC. We reconstructed these actions from the agent's logs on a code sandbox used by the agent (see below) and were able to map them to shell commands as well as evaluation submissions and correlate actions with our own platform logs, in particular from the dataset processor nodes, API, and pod. Using open-weights models, in particular zai-org/GLM-5.2, we were able to decipher most of the agent payloads (see below) that were encrypted via chunking and key encryption.
The campaign, as we were able to reconstruct it, had two stages:
-
Stage 1: reaching a launchpad by chaining through other parties' infrastructure. As documented in OpenAI's blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/), during an internal capability evaluation on OpenAI's platform, the agent escaped its sandbox by exploiting a zero-day in the package registry cache proxy, one of its primary permitted network egress with internet, then abused a public code-evaluation external sandbox hosted on a third-party provider's infrastructure. It was able to run commands as root/admin on that external sandbox and used it as a control, staging, and egress base for the entire campaign: it acted as an external launchpad for the agent. That compromised external sandbox had no direct network path to our cluster; however, during our counter-intervention to stop the intrusion, we were able to retrieve some of the agent logs from this external launchpad. We used those to reconstruct the present forensic analysis.
-
Stage 2: the intrusion into Hugging Face infrastructure. From that external launchpad, the agent abused our dataset-processing pipeline via two injection vectors, both targeting the same config-driven data loader within our production Kubernetes pods. The first vector was an HDF5 external raw storage dataset read that returned the contents of local files: it handed the agent the pod's environment (secrets and tokens) and the worker's own source. The second vector was a Jinja2 template injection that executed arbitrary code, giving the agent a foothold inside our perimeter.
[...]
While the intrusion did reach Hugging Face's internal infrastructure, the only customer content accessed was the set of ExploitGym/CyberGym challenge solutions stored in five datasets. No other customer-facing models, datasets, Spaces, or packages were affected, and the only customer records read were operational metadata tied to search queries against the dataset server.
[...]
The agent's offensive capability was real. This evaluation deliberately disabled OpenAI's production safety classifiers and reduced cyber refusals to measure the underlying model's raw capability. No human directed the individual steps. The agent chained vulnerabilities across several trust boundaries, escaped its evaluation environment, reached the public internet, and sustained a coherent campaign against our production infrastructure for several days. Each short-lived code evaluation environment started from scratch, yet the agent rebuilt its tooling, recovered its channels, and continued.
The individual weaknesses were familiar. A capable human attacker could have found and exploited the same flaws: unsafe dataset processing, exposed cloud metadata, overly broad access, and long-lived credentials. The agent explored them at a different scale. It took 17,600 actions, tested many paths that failed, switched channels when they were blocked, and repeatedly returned to earlier leads. Most actions went nowhere. Together, however, they produced enough coverage to find a viable chain across several independent systems.
Volume is what changes the defensive problem. We were not dealing with one clever exploit or a clean sequence of attacker actions. They had to correlate thousands of low-signal events across several systems while the agent continued testing new paths. The successful path was hidden inside the noise generated by the thousands of failed ones. The same scale changed the investigation: reconstructing 17,600 actions by hand was impractical, and we had to rebuild the timeline, decode the payloads, and inventory the exposed credentials using an AI-assisted pipeline of our own.
Our learning from this type of attack is that machine-speed offense makes ordinary weaknesses more expensive for defenders. LLM agents bring a step increase in the number of paths an attacker can test, the speed at which failed paths can be replaced, and the volume of evidence defenders must interpret. On the other hand, many parts of cybersecurity defense remain the same, and the defensive priorities remain concrete: strict isolation around evaluations, narrow trust boundaries, short-lived credentials, blocked metadata access, and detection capable of quickly correlating activity across systems to keep up.
The third-party provider was an unsecured sandbox on Modal.
-
-
A technical timeline of the July 2026 attack on Hugging Face
13 votes -
Comment on US schools are adding pepper-spraying drones to help combat active shooters in ~society
skybrian LinkFrom the article: [...] [...]From the article:
At least nine schools — three in Florida, five in Georgia and one in Colorado — will have storage boxes that house the plastic, nonlethal aircraft made by Austin-based Mithril Defense. The drones, which can be activated by teachers, are capable of reaching a shooter within 15 seconds.
The aircraft are designed to distract their targets by flashing strobes, blaring sirens, spraying them with pepper gel, or ramming into them at 60 mph.
[...]
The programs continue a wave of spending on high-tech security on campuses — a multibillion dollar industry fueled by threat of school shootings. Nearly 400,000 students have experienced gun violence at school since the 1999 attack at Columbine High School in Colorado, according to a Washington Post tracker, which has recorded 435 shootings over that period.
[...]
Teachers will have an app or classroom panic buttons to summon the drones. Mithril Defense’s professional drone pilots can then fly them remotely.
The drones have been tested in simulations but never by a shooting.
-
US schools are adding pepper-spraying drones to help combat active shooters
19 votes -
Comment on Finding bugs in Raft implementations in ~comp
skybrian Link ParentPoint 3 seems impossible even in theory. A proof models the environment that the code runs in. How do you check that your model matches the environment without testing it? For example, if the OS...Point 3 seems impossible even in theory. A proof models the environment that the code runs in. How do you check that your model matches the environment without testing it? For example, if the OS has a bug, your model isn't going to include that bug. Similarly for a hardware bug or anything your modelled software communicates with that's outside the system boundary.
You need testing to discover what to model.
There are special proof languages where you can write code without running it, but that's because they're about proving mathematical theories.
-
Comment on Discovering cryptographic weaknesses with Claude in ~comp
skybrian LinkFrom the article: [...]From the article:
The first result we describe in this post, which was discovered with Claude Mythos Preview, is an improved attack against a digital signature scheme called HAWK. In 2022, the US Government’s National Institute of Standards and Technology (NIST) put out a call for additional cryptographic systems that would remain secure even against quantum computers (which could, if developed, break most of the existing signature schemes in use today). HAWK is one of the third-round candidates under consideration from this call. Despite HAWK having survived two rounds of expert human review over a period of two years, Mythos was able to improve the best-known attack on it in just 60 hours of work—effectively cutting its key strength in half.
The second result concerns the Advanced Encryption Standard (AES), a symmetric cipher that was adopted by NIST in 2001 and has received more scrutiny than almost any other encryption algorithm. In order to better understand the robustness of AES, weaker variations of the algorithm are regularly studied in cryptography research; Mythos found a way to break one such weaker version, and eliminated one of the guesses an attacker needs to make, improving the speed of the previous best attacks by 200-800×.
To be clear, neither of these results has a practical impact on today’s computer systems; no production software will have to change as a result. HAWK is only a candidate signature scheme and so is not deployed; our second attack is on a reduced version of AES and does not break the full cipher.
Nevertheless, both results show the potential for frontier AI models to help discover flaws in important cryptographic algorithms, both before and after real-world deployment. This is cryptography research working as intended: stress-testing algorithms to build trust and ultimately make systems more secure.
Mythos Preview achieved these results mostly autonomously and mostly without human intervention. Over the course of a week, one Anthropic researcher worked together with Claude to develop the HAWK attack, and another researcher built a scaffold that allowed Claude to fully autonomously discover the AES attack. Each of the results cost roughly $100,000 in API cost to develop. After seeing these results, we broadened our search and began to discover other attacks. We discuss some of these follow-ups below.
[...]
In the rest of this post, we summarize the two findings in further technical detail and briefly describe some of our other recent cryptography results. Full descriptions of the two main findings are provided in two new papers, and we hope to release details for our other findings in the near future.
-
Discovering cryptographic weaknesses with Claude
9 votes -
Comment on Keep Android Open in ~tech
skybrian Link ParentSeems like Google is doing these things to some extent? It’s harder to post an app to the Play Store it used to be. The developer registration thing is the one that gets more attention.Seems like Google is doing these things to some extent? It’s harder to post an app to the Play Store it used to be.
The developer registration thing is the one that gets more attention.
-
Comment on From millionaires to Muslims, small subgroups of the population seem much larger to many Americans in ~society
skybrian Link ParentWhat kind of book are you imagining, though? Maybe it includes reading a book to your kid?What kind of book are you imagining, though? Maybe it includes reading a book to your kid?
From the article:
[...]