skybrian's recent activity
-
Comment on US Federal Communications Commission bans foreign-produced solar inverters in ~enviro
-
US Federal Communications Commission bans foreign-produced solar inverters
18 votes -
Comment on The apples and oranges tribunal in ~society
skybrian LinkFrom the article: [...] [...] [...] I'll add that market prices are often wrong (inconsistent). The market gives an answer, but nothing says it's a correct answer. It's what we have.From the article:
Suppose that apples sell for more than oranges and Parliament in [its] wisdom decides that, at last, apples and oranges must be compared. Not by shoppers — shoppers are biased, they merely reveal what they are willing to pay — but by a tribunal, which will determine whether apples and oranges are of truly equal value and thus must sell at the same price.
What would the tribunal need to know?
[...]
To determine the “just” price of apples and oranges, the tribunal would need the entire general-equilibrium system.
[...]
Britain is now running this experiment in the labor market–Is a retail worker equal to a warehouse worker? A canteen worker equal to a coal miner? A dinner lady equal to a gravedigger?
Under the Equality Act’s “equal value” provisions, tribunals compare jobs by scoring their intrinsic properties — effort, skill, responsibility, working conditions — the labor theory of value applied to labor. How is it going? The Tesco litigation began in 2018; the tribunal’s fact-finding hearing ran 36 days, its judgments run to more than 900 pages resting on some 19,000 pages of training manuals, and the independent experts have yet to begin the report that will actually say whether a shelf-stacker’s job equals a warehouse worker’s. Eight years, and the calculation has not started. Apples and oranges, adjudicated but not, as Orwell or Marx or Stafford Beer might have imagined, by a industrial bureaucracy or by an all-knowing artificial intelligence but by lawyers and commissions and tribunals. The worst of all worlds.
And having discovered that the tribunal cannot price two jobs in a decade, the government now proposes to add race and disability comparisons and an enforcement unit to publish official guidance on which reasons for a wage difference are permissible. A bureau of allowable scarcities.
[...]
The market compares apples and oranges every day. It is the only institution that can. But there is a deeper error here than computation. Suppose the tribunal succeeded. Suppose that after another decade it delivered the true and final score, shelf-stacker versus warehouseman. What would it have found? Not justice. A wage is not a grade on your character or a measure of your worth as a human being. A wage is a price — a report on how scarce your skills are relative to the desires of people you will never meet. Nurses are not morally less worthy than plumbers should they earn less than plumbers or vice-versa, and no one thinks otherwise except the tribunals.
I'll add that market prices are often wrong (inconsistent). The market gives an answer, but nothing says it's a correct answer. It's what we have.
-
The apples and oranges tribunal
2 votes -
Comment on A technical timeline of the July 2026 attack on Hugging Face in ~comp
skybrian LinkFrom the article: [...] [...] [...] The third-party provider was an unsecured sandbox on Modal.From the article:
Over roughly two and a half days inside our infrastructure, an autonomous AI agent driven by a combination of OpenAI models ran an end-to-end intrusion against our platform: it was thousands of small, automated decisions, executed at machine speed across short-lived sandbox environments, with command-and-control staged on ordinary public web services.
[...]
Our forensic reconstruction covers ~17,600 attacker actions that we were able to recover, grouped into ~6,280 clusters, between 2026-07-09 02:28 UTC and 2026-07-13 14:14 UTC. We reconstructed these actions from the agent's logs on a code sandbox used by the agent (see below) and were able to map them to shell commands as well as evaluation submissions and correlate actions with our own platform logs, in particular from the dataset processor nodes, API, and pod. Using open-weights models, in particular zai-org/GLM-5.2, we were able to decipher most of the agent payloads (see below) that were encrypted via chunking and key encryption.
The campaign, as we were able to reconstruct it, had two stages:
-
Stage 1: reaching a launchpad by chaining through other parties' infrastructure. As documented in OpenAI's blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/), during an internal capability evaluation on OpenAI's platform, the agent escaped its sandbox by exploiting a zero-day in the package registry cache proxy, one of its primary permitted network egress with internet, then abused a public code-evaluation external sandbox hosted on a third-party provider's infrastructure. It was able to run commands as root/admin on that external sandbox and used it as a control, staging, and egress base for the entire campaign: it acted as an external launchpad for the agent. That compromised external sandbox had no direct network path to our cluster; however, during our counter-intervention to stop the intrusion, we were able to retrieve some of the agent logs from this external launchpad. We used those to reconstruct the present forensic analysis.
-
Stage 2: the intrusion into Hugging Face infrastructure. From that external launchpad, the agent abused our dataset-processing pipeline via two injection vectors, both targeting the same config-driven data loader within our production Kubernetes pods. The first vector was an HDF5 external raw storage dataset read that returned the contents of local files: it handed the agent the pod's environment (secrets and tokens) and the worker's own source. The second vector was a Jinja2 template injection that executed arbitrary code, giving the agent a foothold inside our perimeter.
[...]
While the intrusion did reach Hugging Face's internal infrastructure, the only customer content accessed was the set of ExploitGym/CyberGym challenge solutions stored in five datasets. No other customer-facing models, datasets, Spaces, or packages were affected, and the only customer records read were operational metadata tied to search queries against the dataset server.
[...]
The agent's offensive capability was real. This evaluation deliberately disabled OpenAI's production safety classifiers and reduced cyber refusals to measure the underlying model's raw capability. No human directed the individual steps. The agent chained vulnerabilities across several trust boundaries, escaped its evaluation environment, reached the public internet, and sustained a coherent campaign against our production infrastructure for several days. Each short-lived code evaluation environment started from scratch, yet the agent rebuilt its tooling, recovered its channels, and continued.
The individual weaknesses were familiar. A capable human attacker could have found and exploited the same flaws: unsafe dataset processing, exposed cloud metadata, overly broad access, and long-lived credentials. The agent explored them at a different scale. It took 17,600 actions, tested many paths that failed, switched channels when they were blocked, and repeatedly returned to earlier leads. Most actions went nowhere. Together, however, they produced enough coverage to find a viable chain across several independent systems.
Volume is what changes the defensive problem. We were not dealing with one clever exploit or a clean sequence of attacker actions. They had to correlate thousands of low-signal events across several systems while the agent continued testing new paths. The successful path was hidden inside the noise generated by the thousands of failed ones. The same scale changed the investigation: reconstructing 17,600 actions by hand was impractical, and we had to rebuild the timeline, decode the payloads, and inventory the exposed credentials using an AI-assisted pipeline of our own.
Our learning from this type of attack is that machine-speed offense makes ordinary weaknesses more expensive for defenders. LLM agents bring a step increase in the number of paths an attacker can test, the speed at which failed paths can be replaced, and the volume of evidence defenders must interpret. On the other hand, many parts of cybersecurity defense remain the same, and the defensive priorities remain concrete: strict isolation around evaluations, narrow trust boundaries, short-lived credentials, blocked metadata access, and detection capable of quickly correlating activity across systems to keep up.
The third-party provider was an unsecured sandbox on Modal.
-
-
A technical timeline of the July 2026 attack on Hugging Face
8 votes -
Comment on US schools are adding pepper-spraying drones to help combat active shooters in ~society
skybrian LinkFrom the article: [...] [...]From the article:
At least nine schools — three in Florida, five in Georgia and one in Colorado — will have storage boxes that house the plastic, nonlethal aircraft made by Austin-based Mithril Defense. The drones, which can be activated by teachers, are capable of reaching a shooter within 15 seconds.
The aircraft are designed to distract their targets by flashing strobes, blaring sirens, spraying them with pepper gel, or ramming into them at 60 mph.
[...]
The programs continue a wave of spending on high-tech security on campuses — a multibillion dollar industry fueled by threat of school shootings. Nearly 400,000 students have experienced gun violence at school since the 1999 attack at Columbine High School in Colorado, according to a Washington Post tracker, which has recorded 435 shootings over that period.
[...]
Teachers will have an app or classroom panic buttons to summon the drones. Mithril Defense’s professional drone pilots can then fly them remotely.
The drones have been tested in simulations but never by a shooting.
-
US schools are adding pepper-spraying drones to help combat active shooters
11 votes -
Comment on Finding bugs in Raft implementations in ~comp
skybrian Link ParentPoint 3 seems impossible even in theory. A proof models the environment that the code runs in. How do you check that your model matches the environment without testing it? For example, if the OS...Point 3 seems impossible even in theory. A proof models the environment that the code runs in. How do you check that your model matches the environment without testing it? For example, if the OS has a bug, your model isn't going to include that bug. Similarly for a hardware bug or anything your modelled software communicates with that's outside the system boundary.
You need testing to discover what to model.
There are special proof languages where you can write code without running it, but that's because they're about proving mathematical theories.
-
Comment on Discovering cryptographic weaknesses with Claude in ~comp
skybrian LinkFrom the article: [...]From the article:
The first result we describe in this post, which was discovered with Claude Mythos Preview, is an improved attack against a digital signature scheme called HAWK. In 2022, the US Government’s National Institute of Standards and Technology (NIST) put out a call for additional cryptographic systems that would remain secure even against quantum computers (which could, if developed, break most of the existing signature schemes in use today). HAWK is one of the third-round candidates under consideration from this call. Despite HAWK having survived two rounds of expert human review over a period of two years, Mythos was able to improve the best-known attack on it in just 60 hours of work—effectively cutting its key strength in half.
The second result concerns the Advanced Encryption Standard (AES), a symmetric cipher that was adopted by NIST in 2001 and has received more scrutiny than almost any other encryption algorithm. In order to better understand the robustness of AES, weaker variations of the algorithm are regularly studied in cryptography research; Mythos found a way to break one such weaker version, and eliminated one of the guesses an attacker needs to make, improving the speed of the previous best attacks by 200-800×.
To be clear, neither of these results has a practical impact on today’s computer systems; no production software will have to change as a result. HAWK is only a candidate signature scheme and so is not deployed; our second attack is on a reduced version of AES and does not break the full cipher.
Nevertheless, both results show the potential for frontier AI models to help discover flaws in important cryptographic algorithms, both before and after real-world deployment. This is cryptography research working as intended: stress-testing algorithms to build trust and ultimately make systems more secure.
Mythos Preview achieved these results mostly autonomously and mostly without human intervention. Over the course of a week, one Anthropic researcher worked together with Claude to develop the HAWK attack, and another researcher built a scaffold that allowed Claude to fully autonomously discover the AES attack. Each of the results cost roughly $100,000 in API cost to develop. After seeing these results, we broadened our search and began to discover other attacks. We discuss some of these follow-ups below.
[...]
In the rest of this post, we summarize the two findings in further technical detail and briefly describe some of our other recent cryptography results. Full descriptions of the two main findings are provided in two new papers, and we hope to release details for our other findings in the near future.
-
Discovering cryptographic weaknesses with Claude
9 votes -
Comment on Keep Android Open in ~tech
skybrian Link ParentSeems like Google is doing these things to some extent? It’s harder to post an app to the Play Store it used to be. The developer registration thing is the one that gets more attention.Seems like Google is doing these things to some extent? It’s harder to post an app to the Play Store it used to be.
The developer registration thing is the one that gets more attention.
-
Comment on From millionaires to Muslims, small subgroups of the population seem much larger to many Americans in ~society
skybrian Link ParentWhat kind of book are you imagining, though? Maybe it includes reading a book to your kid?What kind of book are you imagining, though? Maybe it includes reading a book to your kid?
-
Comment on From millionaires to Muslims, small subgroups of the population seem much larger to many Americans in ~society
skybrian Link ParentYes, good catch, and this is a common problem with survey questions. The people who wrote the question think it means one thing and the people who took the survey think it means something else....Yes, good catch, and this is a common problem with survey questions. The people who wrote the question think it means one thing and the people who took the survey think it means something else. And if you don’t ask people to explain their reasoning, you might not find out?
For someone reading about the survey, if you read the questions then you might see the ambiguity.
-
Comment on Keep Android Open in ~tech
skybrian Link ParentWhat do you mean it's not about security? People say that all the time but do they know anything?What do you mean it's not about security? People say that all the time but do they know anything?
-
Comment on The machines are fine. I'm worried about us. in ~science
skybrian Link ParentDetecting LLMisms is something you learn from practice and you get better at it. There are patterns that are invisible at first, but you start seeing them after repeated exposure. And I wonder if...Detecting LLMisms is something you learn from practice and you get better at it. There are patterns that are invisible at first, but you start seeing them after repeated exposure. And I wonder if there are more patterns that I don’t see yet?
-
Comment on The machines are fine. I'm worried about us. in ~science
skybrian Link ParentI think we will be seeing more price competition soon. Open weights models are getting pretty good, and so are cheaper proprietary models. I sometimes use Luna instead of Terra and rarely try Sol...I think we will be seeing more price competition soon. Open weights models are getting pretty good, and so are cheaper proprietary models. I sometimes use Luna instead of Terra and rarely try Sol or Opus.
-
Comment on Finding bugs in Raft implementations in ~comp
skybrian LinkFrom the article: [...] [...]From the article:
Despite the paper’s famously accessible style, we’ve found bugs in every Raft implementation we’ve tested, including HashiCorp Raft, Aeron Cluster, OpenRaft, and MicroRaft — despite the investment in formal methods, careful code review, unit testing, and years of testing in production. The bugs we found manifest as violations of Raft’s main invariant (called state machine safety in the paper, commonly referred to elsewhere as total order delivery). If you’re using a Raft implementation, you might want to check it for bugs.
We’ve sent bug reports upstream. This isn’t intended as a critique of Raft, its authors, its implementers, or any particular implementation. Raft implementations, even with a formal specification and a detailed implementation guide, are not easy to write.
Rather, this is a story about correctness in distributed systems — the inevitability of bugs, the inadequacy of any single approach, and the high cost of learned helplessness.
[...]
In all cases, network partitions and turbulence were sufficient to surface examples of divergence (i.e. no node kill/restart required, or disk corruption, or other faults)
[...]
If there is a single lesson to be learned here, it’s that formal methods alone cannot ensure that software works, because the formal specification still needs to be implemented, and even if you have mechanized verification, you’re verifying the model and not the implementation itself.
We believe formal methods are useful and necessary — they can confirm the basic soundness of a design, and provide a map that saves engineers from many of the errors that can arise in the implementation of a complex system.
But as these bugs show, errors continue to arise when translating the formal specification to production code. In the course of our work with various Raft implementations, we identified a number of assumptions in the Raft paper that remain implicit. An implementer who misses any of these details is likely to run into trouble.
-
Finding bugs in Raft implementations
9 votes -
Comment on South Korea’s Kospi index sinks 10% on heavy selling of chipmaking stocks in ~finance
skybrian LinkFrom the article: [...] [...]From the article:
HONG KONG (AP) — South Korea's Kospi index plunged more than 10% on Tuesday on heavy selling of chipmaking stocks that have gyrated recently due to concerns over the sustainability of the boom in artificial intelligence.
[...]
Trading was temporarily halted as Kospi dropped to its lowest level since April as shares in chipmakers Samsung Electronics and SK Hynix dropped sharply. The Kospi was down 10.5% at 6,051.19 by midday.
[...]
Samsung's shares tumbled 12% while those for SK Hynix were down 12.7%.
A big factor driving the selling of AI-related shares, analysts said, is the expectation that competition from Chinese AI startup s and chipmakers might undermine gains for global companies whose shares have skyrocketed due to the AI frenzy.
A 466% jump in the price of Chinese chipmaker CXM T in its trading debut Monday has underscored such concerns. CXMT raised at least $8.6 billion in its initial public offering on Shanghai's tech-oriented STAR exchange.
From the article: