skybrian's recent activity
-
Comment on 'VPNs are lawful technical tools,' says EU court in landmark Anne Frank copyright ruling in ~society
-
'VPNs are lawful technical tools,' says EU court in landmark Anne Frank copyright ruling
9 votes -
Comment on Judge approves a $1.5B Anthropic settlement over books used to train Claude in ~books
skybrian LinkFrom the article:From the article:
SAN FRANCISCO (AP) — A federal judge has approved a $1.5 billion copyright settlement in which artificial intelligence company Anthropic will pay thousands of authors about $3,000 per book after using pirated copies of their works to train its Claude chatbot.
District Judge Araceli Martínez-Olguín said in a Monday ruling that the class-action settlement provides “meaningful relief” to affected authors and publishers.
About 91% of the more than 482,000 books covered by the ruling have been claimed by authors or publishers who are now due payment.
Plaintiff attorney Justin Nelson said in a statement that the settlement was “the largest known copyright recovery in history. We look forward to making distributions to the Class as promptly as possible.”
U.S. District Judge William Alsup issued the preliminary approval in San Francisco federal court last September and has since retired. Alsup had dealt the case a mixed ruling last summer, finding that training AI chatbots on copyrighted books wasn’t illegal but that Anthropic wrongfully acquired millions of books through pirate websites.
-
Judge approves a $1.5B Anthropic settlement over books used to train Claude
3 votes -
Comment on Unreleased OpenAI model escapes containment and hacks into HuggingFace in ~tech
skybrian LinkFrom the article:From the article:
This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.
The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.
While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.
After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. OpenAI’s security team discovered this anomalous activity internally.
-
Unreleased OpenAI model escapes containment and hacks into HuggingFace
5 votes -
Comment on Long presumed dead, a thriving coral reef is discovered in West Africa in ~enviro
skybrian LinkFrom the article:From the article:
In the 1960s, fishing surveyors off West Africa’s Benin coast hauled up heads of coral in their nets. Employed by local governments to assess fish diversity and discover potentially trawlable seabeds in the sandy-bottomed Gulf of Guinea, the researchers shrugged off the discovery at the time, burying the find in a brief paragraph in a 130-something-page report.
With no surveys since, the scientists who followed had lost track of exactly where the reefs might lie. And as mass bleaching events and overfishing slashed the world’s coral reef area by more than 50 percent since the 1950s, local oceanographers had written off any remnants of a possible reef as dead.
More than six decades on, the mystery has been solved as a team of Beninese scientists rediscovered the healthy reef teeming with marine life. At least eight coral types and eight fish species have formed a thriving ecosystem on this long-forgotten site.
-
Long presumed dead, a thriving coral reef is discovered in West Africa
3 votes -
Comment on Deportation orders soar in NYC as Trump 'mega master' hearings accelerate in ~society
skybrian LinkFrom the article: [...] [...] [...]From the article:
Inside a New York City immigration courtroom at 26 Federal Plaza on Friday morning, 90 people had cases before a single immigration judge. People described hastily making plans to get to court, some from hundreds of miles away, having learned about their new hearing date only days in advance.
By the end of the morning, dozens of people still hadn’t shown up and the judge ordered them removed from the country.
These removal orders of immigrants who didn’t show up to court reached record heights in June nationwide and more than doubled in New York City last month, according to a new report released by bklg.org, a nonprofit that analyzes immigration court data to help immigration attorneys keep on top of their clients’ cases. The increase occurred during the first few weeks of a new Trump administration effort to speed up deportations through large-scale court proceedings known as “mega master” hearings.
In June, 4,447 people were ordered removed by immigration judges in the city “in absentia,” meaning they’d missed their hearings, the report found. That was more than double the number in May when 2,189 were ordered removed in absentia.
[...]
In absentia removal orders have been increasing annually since 2021, as the number of people crossing the U.S. border and the number of deportation cases rose, but June showed a dramatic jump and surpassed any month on record, going back to the late 1990s when the government started keeping track.
[...]
The federal government is required by law to send written notices of any new hearing dates by mail, but many attending these hearings told The City Reporter they’d only learned of the date change because they happened to check the Executive Office for Immigration Review’s online portal or another similar app that helped them monitor their cases. People who miss hearings in deportation proceedings are subject to automatic removal orders, clearing the way for the government to deport them.
[...]
Immigration advocates have been urging anyone in deportation proceedings to double-check EOIR’s portal daily to monitor their case for any changes.
-
Deportation orders soar in NYC as Trump 'mega master' hearings accelerate
2 votes -
Comment on How do you live simply/save money? What does your lifestyle look like? in ~talk
skybrian LinkAnd for emergencies too, I hope! And eventually, for retirement. Having savings helps with a lot of things.I’d like to have lower expenses so that I can save more for meaningful purchases
And for emergencies too, I hope! And eventually, for retirement. Having savings helps with a lot of things.
-
Comment on AI companies are buying tons of old books because they're free of AI slop in ~books
skybrian LinkFrom the article: [...] [...] [...] [...] [...] [...]From the article:
As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing.
“The world's best AI training data is sitting on a shelf,” ISBNdb, a company that produces what it claims is “the world’s largest book database,” and that offers high-volume book acquisition services for AI companies, says on its site. “Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.”
[...]
AI companies’ attempts to hoover up printed books for training data got wide attention in January after a copyright lawsuit from book authors against Anthropic revealed internal documents detailing its plan to obtain and scan millions of printed books, and destroy them in the process. The Washington Post article found that Anthropic was buying books from one company called Better World Books, one of several marketplaces where libraries, retailers, and individuals can sell their books. Google was recently sued by book publishers for similarly training Google Gemini on copyrighted books.
ISBNdb advertises that it can keep the identity of AI companies secret.
[...]
ISBNdb notes that AI companies may not want to be caught destroying printed books during the scanning process.
[...]
One professional bookseller who specializes in selling foreign language books on these marketplaces told me that, starting in April, he and other booksellers noticed a historic spike in sales. This bookseller asked to remain anonymous so he can continue to do business on these platforms.
“I personally have mixed feelings about all of this,” the bookseller, who suspects he’s sold hundreds of books to AI companies for training data, told me. “It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. I’ve been well-suited for these sales with inventory from overseas and foreign language books. On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped.”
This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain.
The seller told me that, normally, on a good week, he’d sell about 20 books. Since April, he has regularly sold hundreds of books a week. While the seller didn’t have clear evidence that the purchases were being made by AI companies, the purchases made him suspect that they were. First of all, he said, the kind of books he sells are specialized and are usually bought by schools and libraries. Purchases from these organizations have been trending downward because of reduced funding, he said. Bulk purchases also usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases. I have not seen any evidence that this bookseller’s recent sales were facilitated by ISBNdb or that the client was an Anthropic or another AI company.
[...]
It’s hard to say for a fact that the books are being bought for training data and possibly being destroyed by AI companies because ISBNdb and book marketplaces like Biblio and Alibris keep the identity of the buyer hidden. Large bulk purchases of books are first sent to distribution centers where, for example, Alibris checks the quality of the books before sending them off to the client.
[...]
Regardless of its original intentions, the federal judge in the copyright lawsuit from authors against Anthropic, William Alsup, found that Anthropic’s creation of digital copies of the books was legal specifically because the books were destroyed.
“Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy,” Alsup wrote in his ruling. “The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company.”
[...]
“Purchasing paper books at scale from the secondary market does not deprive any creator of income they would otherwise have received,” ISBNdb’s site says. “These are books that have already fully discharged their financial obligation to their creators.”
-
AI companies are buying tons of old books because they're free of AI slop
9 votes -
Comment on Why socialists should embrace luxury apartments in ~society
skybrian Link ParentIn New York City, co-ops are common and they are often cheaper than buying a condo, but they’re also restrictive. You need to reveal a lot of information about yourself to the co-op board and hope...In New York City, co-ops are common and they are often cheaper than buying a condo, but they’re also restrictive. You need to reveal a lot of information about yourself to the co-op board and hope they approve your application. So one way to think about it is that condo owners willingly pay more to not have to deal with it, and that’s why condo prices are somewhat higher. When supply is restricted, prices are set by demand to be higher than what it costs to provide the service, and providers do get a windfall if their costs are low.
It seems like an important service that landlords provide is that they serve people who can’t afford to buy a house or condo, wouldn’t be able to get a mortgage, and wouldn’t look like a good bet to the co-op board. Landlords do make requirements of prospective tenants too, but their requirements are lower.
It’s good that all these alternative housing arrangements are available, and there should be more. A lot of the cheaper alternatives like boarding houses or staying at the YMCA have disappeared, so people make do by sharing apartments.
You wrote elsewhere in this topic that you’d like to see co-ops given a solid try and I think that’s fine; I think we need more of everything.
-
Comment on A skeptical overview of AI and consciousness in ~humanities
skybrian Link ParentIt would be prudent to avoid the ethical mess by making AI’s that are useful, but that society can be reasonably sure aren’t conscious. Although having someone to talk to has its attractions (see...It would be prudent to avoid the ethical mess by making AI’s that are useful, but that society can be reasonably sure aren’t conscious. Although having someone to talk to has its attractions (see Her), one would hope that most people will prefer machine labor to slave labor. If customers disagree and feel strongly about it, from a market perspective it’s incentive to give both sides what they want. I can imagine either regulation or certification that a particular service only uses AI that is almost certainly not conscious.
Maybe that’s a “classic” LLM that forgets who you are when you start a new conversation? I think the lesson we’re learning from LLM’s is that whatever consciousness is, our machines don’t need it. Pretty much any kind of reasoning ability can be separated from it.
And then the question is whether to ban the kind of AI that’s blurring the lines too much.
-
Comment on Linus Torvalds says Linux is not "anti-AI", tells haters to 'fork it' and 'just walk away' in ~tech
skybrian Link ParentI’m not sure what you mean by sustainable, but I’m a bit skeptical that trends toward cloud computing will reverse to the extent that running LLMs on your laptop or a server in the closet wins...I’m not sure what you mean by sustainable, but I’m a bit skeptical that trends toward cloud computing will reverse to the extent that running LLMs on your laptop or a server in the closet wins over running them in the cloud. Datacenters will almost certainly be more efficient than anything you run yourself. But it might be an open weights model running at whichever hosting provider you trust rather than a service provided by an AI lab.
-
Comment on Why socialists should embrace luxury apartments in ~society
skybrian Link ParentFirst of all, it's wrong to say landlords do no work. Many small-time landlords do some work themselves. Also, I don't know why you're so focused on the labor thing. Whether you do the maintenance...First of all, it's wrong to say landlords do no work. Many small-time landlords do some work themselves. Also, I don't know why you're so focused on the labor thing. Whether you do the maintenance yourself or pay someone to do it is irrelevant, as long as it gets done.
I don't see anything wrong with paying a plumber, for example. Most people who own their own home would do the same.
-
Comment on US President Donald Trump says he will impose fifty percent tariffs on Canadian goods in thirty days in ~society
skybrian Link ParentDo you have a source for that legal basis? The article says different: From a brief search, I didn't find any news articles claiming that the 1974 Trade Act is the legal basis for the latest...Do you have a source for that legal basis? The article says different:
Monday’s action relies on a never-before-used provision of a 1930 trade law, which provides for tariffs of up to 50 percent against countries that discriminate against U.S. goods. The 30-day delay provides time for the two nations to resolve the dispute before tariffs take effect.
From a brief search, I didn't find any news articles claiming that the 1974 Trade Act is the legal basis for the latest threat against Canada. Here's what CNN has:
Trump cited Section 338 of the Tariff Act of 1930, a law that allows a president to set tariffs up to 50% when another country discriminates against American goods. However, the law has never been applied this way.
According to this article, he did use Section 301 in his first term.
-
Comment on Why socialists should embrace luxury apartments in ~society
skybrian Link ParentApology accepted! I do agree that the rents (and house prices) are too damn high. I suspect a lot of people are quietly getting by with the family helping them out, if they're lucky enough to have...Apology accepted! I do agree that the rents (and house prices) are too damn high. I suspect a lot of people are quietly getting by with the family helping them out, if they're lucky enough to have someone who can help.
I think the way out, or at least to keep things from getting worse, is more housing, through any means. And, perhaps making it more viable to move to cheaper locations. And UBI would be nice.
-
Comment on Human mathematicians are being outcounterexampled in ~science
skybrian LinkFrom the article: [...] [...] [...]From the article:
It is now 9 years since I had a mid-life crisis, realised I no longer trusted many human mathematicians when it comes to technical details, discovered Lean, and started to argue that interactive theorem provers should play an important role in the future of mathematics. So of course my first question was “is the counterexample formalized in Lean”. The answer was “no”.
But under a week later (26th May 2026), I got an email from Fields Medallist Mike Freedman. Mike is now the Chief Science Officer for Logical Intelligence, a company cofounded by Turing Award winner and “godfather of AI” Yan LeCun. Mike informed me that their system had autoformalized the entire ChatGPT-generated paper in Lean and could I take a look. I looked, and my post-doc Thomas Browning looked too. And indeed this was what Logical Intelligence had done: they had formalized precisely the statement that the profound theorem of number theory implied the Erdős counterexample. Breakthrough LLM-generated mathematics being formalized in real time. Interesting data point.
Of course there is an elephant in the room here though, the profound theorem of number theory which takes 100+ pages to prove (it needs huge chunks of global class field theory, a theory developed at the beginning of the 20th century and for which there are still no short proofs; it is proving difficult to compress). In 2025 I had run a Clay Summer School with Richard Hill on the formalization of class field theory, and one year later we have nearly done the local case (it is the current PhD project of my student Edison Xie); the global case remained open, and indeed in 2025 formalizing global class field theory seemed like a fantasy.
One month later, on June 26th 2026, my perception of what was possible again changed. Boris Alexeev announced on the Lean Zulip that he had steered ChatGPT to a complete formalization of the Erdős counterexample, assuming nothing beyond the axioms of mathematics. Boris works at OpenAI and had used their new model Sol to do the autoformalization. Boris made the code public and it did not take long for me to realise that somewhere within all this AI-generated (and sometimes horrible, although sometimes decent) code was indeed a proof of some really hard theorems in global class field theory. Also of interest to me was that Sol had generated 1.2 million lines of Lean code in the three weeks that it had worked on the project. Lean’s fantastic (declaration of conflict of interest: I am a maintainer) mathematics library mathlib is only 2.3 million lines of code, and took nine years to write. Perhaps it was at this point that the penny really dropped for me — large AI-generated developments of mathematics are inevitable. One cannot trust AI-generated code so I ran it in a sandbox on my machine (malicious Lean code can run arbitrary commands on your computer — Lean is a programming language, after all). Indeed, it was proving nontrivial theorems about the cohomology of number fields. Wow.
[...]
I was not sure how good Logos’ tool was going to be, but I wanted a development of the theory of finite flat group schemes in Lean for my ongoing proof of Fermat’s Last Theorem, so I put uploaded some classic papers in the area to Fable and ChatGPT, and got them together to write down an exposition of the theory in natural language. I passed this pdf document over to Logos the day before the workshop, and on the first day of the workshop they said that one of the claims in the pdf was false and they had found an explicit counterexample. Another counterexample! I took a look and indeed the LLM-generated pdf was simply wrong at some point when describing a standard construction; false alarm. I had missed this myself though when reading through the pdf. Interesting how AI had again found a counterexample. I fixed the pdf. I thought it was interesting that the AI didn’t just say “I don’t quite follow this argument”, it instead said “here is a proof that this argument is simply wrong”, a much more powerful statement.
[...]
The day after the workshop finished, on Saturday 11th July, I got a DM from Akhil telling me that Sol had found a counterexample. He sent me a 12 page pdf. I immediately replied saying that I was not reading AI-generated informal mathematics and could he please formalize the entire thing in Lean. Four hours later he replied again, saying that Fable had autoformalized the entire thing. I scanned over the 1076-line Lean file, checking that the code did not delete all the files on my hard drive (Lean is a programming language, so it can do this). Convinced that it was only theorems, I then compiled it on my laptop and it took me under 5 minutes in total to check that (a) the statement of the claimed theorem used only concepts in mathlib (and thus things like HopfAlgebra can be trusted to mean what mathematicians think of as Hopf algebras) (b) the statement of the claimed theorem was that there was a counterexample and (c) the proof compiled. At this point I knew that we had a counterexample — a group scheme of order 4 which was not killed by 4. I suggested to Akhil that he make a PR to mathlib with the counterexample — which he did. I would have also suggested to him that he draft a press release saying that a machine had solved a 60-year-old question of Grothendieck in algebraic geometry, but somehow by this point I was almost becoming immune to all of this. It wasn’t clear to me that the media would even be able to distinguish between “machine resolves question due to Erdős” and “machine resolves question due to Grothendieck” even though I personally found the latter far more interesting. Of course the Grothendieck counterexample was far far easier than the Erdős one (a thousand lines, not a million), all I’m saying is that it’s an area of mathematics that I personally find more interesting. I pointed out to Akhil that machines seemed to be getting very good at finding counterexamples and suggested that he try the Hodge conjecture next.
[...]
I think it’s worth stepping back at this point and surveying what the attitudes of human experts to these sorts of things are. On Tuesday (14th July) I went to work at Imperial and the Grothendieck counterexample was the talk of lunch. A member of the faculty (who I won’t name) said to me that the fact that the counterexample was so easy to find just indicated that humans had not spent enough time thinking about the problem, implying that a 60-year-old question of Grothendieck was not actually that interesting to work on. I didn’t tell him that at some point earlier in my career I had spent a week working hard on the problem. In my mind my colleague is just going through the five stages of grief; right now they seem to be in the denial phase.
From the article:
[...]
[...]