skybrian's recent activity
-
Comment on What is the oldest file you still have? in ~talk
-
Comment on What are we in a golden age of? in ~talk
skybrian Link ParentMy guess is that it's going to get cheaper, but current prices aren't bad. I'm using Luna with a $20/month ChatGPT subscription (also used for asking questions). It's not free, but it's not that...My guess is that it's going to get cheaper, but current prices aren't bad. I'm using Luna with a $20/month ChatGPT subscription (also used for asking questions). It's not free, but it's not that expensive as hobbies go.
At today's RAM prices, I'd put off buying an expensive machine for local inference.
-
Comment on What are we in a golden age of? in ~talk
skybrian Link ParentThis reminds me of how I randomly saw one of Hiromi's concert videos on YouTube a few years ago. Then I searched for all her other videos and made a playlist, started listening to her albums on...This reminds me of how I randomly saw one of Hiromi's concert videos on YouTube a few years ago. Then I searched for all her other videos and made a playlist, started listening to her albums on Spotify, went to see her in concert, and recently bought a book of sheet music from a store in Japan and started learning to play one of her songs.
Seems like I should listen to other jazz musicians more.
-
Comment on What are we in a golden age of? in ~talk
skybrian Link ParentGoogle does seem worse than it used to be. Also, finding news articles is somewhat worse. It's harder to find articles that aren't paywalled in some way, and there are more bot challenges, though...Google does seem worse than it used to be. Also, finding news articles is somewhat worse. It's harder to find articles that aren't paywalled in some way, and there are more bot challenges, though I don't block cookies so maybe it affects me less. (I also have NYT and Washington Post subscriptions, which helps.)
I've had better luck with ChatGPT than with Google. It might be due to having a subscription, resulting in using a better LLM by default.
I can ask for a chart and if it doesn't find anything suitable, it might download the data and make one.
-
Comment on What are we in a golden age of? in ~talk
skybrian LinkThis is a golden age for asking questions. Often I’m wondering how to do something and then I think, wait, I could ask ChatGPT about that! Of course, this happened before with search engines....This is a golden age for asking questions. Often I’m wondering how to do something and then I think, wait, I could ask ChatGPT about that!
Of course, this happened before with search engines. Those of us who are old enough can remember how cool it was to be able to search the web instead of going to the library.
But there are many questions that I wouldn’t bother to search on because I expected the results to be crappy, or I would try a different search that’s a bit more indirect but might dredge something up. That’s a mental block that’s no longer helpful. You can often ask questions in the most direct, naive way, and get good results.
-
Comment on What are we in a golden age of? in ~talk
skybrian LinkIf you’ve ever wanted to write your own personal software, this is a fantastic time to do it. You might make a mess, but with AI, you will make messes much faster and learn from them. And if you...If you’ve ever wanted to write your own personal software, this is a fantastic time to do it. You might make a mess, but with AI, you will make messes much faster and learn from them. And if you manage to avoid getting too ambitious, you might even build something you actually use.
-
Comment on An AI researcher writes about his crisis of faith in ~tech
skybrian Link ParentI’m not entirely sure how to define an “act of will,” but winning elections doesn’t seem like an “act of will” either, for the same reason: people will oppose you! Convincing people to vote for...I’m not entirely sure how to define an “act of will,” but winning elections doesn’t seem like an “act of will” either, for the same reason: people will oppose you! Convincing people to vote for your candidate instead of one of the others isn’t easy.
-
Comment on An AI researcher writes about his crisis of faith in ~tech
skybrian (edited )Link ParentPutin thought that conquering Ukraine was an act of will. Israel convinced Trump that regime change in Iran was an act of will. Turns out that other people will oppose acts of will.Putin thought that conquering Ukraine was an act of will. Israel convinced Trump that regime change in Iran was an act of will.
Turns out that other people will oppose acts of will.
-
Comment on An AI researcher writes about his crisis of faith in ~tech
skybrian (edited )Link ParentIf you want to see what that looks like today, talk to people who have retired, early or otherwise. There are plenty of retirees to talk to! In my case, I do spend time helping relatives and I'm...If you want to see what that looks like today, talk to people who have retired, early or otherwise. There are plenty of retirees to talk to!
In my case, I do spend time helping relatives and I'm glad I've been able to drop everything and go when people needed me. I'm not helping random strangers, though, except with my donations. That's real work and I don't think it's going to automated any time soon?
Assuming there is a decline in other kinds of work (I'm uncertain about that), I imagine more work is going to be care-giving of various sorts. Healthcare is already a growing industry.
But even "everything except care-giving is automated" seems rather utopian? There are too many fields that AI has had no effect on yet.
Even for something like teaching where the effect of AI is important, I see it having an effect like "kids learn some more things from computers," but not "therefore nobody needs to take care of them." I imagine that the aspects of the job that are more about taking care of kids and motivating them get even more emphasis?
-
Comment on AI text watermarking is free and good in ~comp
skybrian Link ParentYes, I imagine the watermark detection algorithm won't work as well when the text is hightly constrained by context. It will always need a minimum amount of text to work with, and for...Yes, I imagine the watermark detection algorithm won't work as well when the text is hightly constrained by context. It will always need a minimum amount of text to work with, and for highly-constrained text, it will need more.
But if it doesn't have enough text to pick up a signal, wouldn't it answer "not detected" most of the time? That is, the input will affect the score randomly and the score will be close to what it would be for random noise.
-
Comment on US judge to baby: file an asylum application in ~society
skybrian Link ParentThat scenario seems pretty plausible too.That scenario seems pretty plausible too.
-
Comment on An AI researcher writes about his crisis of faith in ~tech
skybrian LinkFrom the blog post: [...] [...]From the blog post:
Even if the nightmare scenario of full human redundance is not imminent, the fact that a considerable number of people claim that this is their goal is unsettling. This view is perhaps most succinctly stated by the startup Mechanize, which aims to “fully automate the economy”. As a participant in the economy, I find this goal deeply objectionable. I think that trying to fully replace all human labor is not a morally acceptable goal to have. It’s frankly horrifying. To be clear, “fully automate the economy” means “destroy society as we know it”. Unfortunately, similar goals are articulated by quite a few inside and outside Silicon Valley. I think the callousness of such objectives should be called out whenever they are encountered.
The great irony of this is that those who claim to want to automate all human labor are typically the kind of competitive, smart, high-agency persons who absolutely need to have a purpose and something to build. They would hate to be redundant. Yet, here they fly, like so many moths to a flame.
[...]
Why didn't I just quit AI research? My predicament would seem like that of a vegan butcher, or a monk who makes money on Onlyfans. But me quitting and becoming an Uber driver would not make the world better. The pace of AI progress would clearly not slow down noticeably. And I assure you, I still love AI research, even if I sometimes hate what AI does to the world. I'm not even very good at anything else. AI is what I do. And I think that I can do more good by trying to steer my field in a good direction than if I became an Uber driver. So I’d rather think of myself as akin to a hypochondriac doctor, or a pilot with a fear of heights.
[...]
From this perspective, the history of AI is a history of attempts to mimic the specific combination of behaviors and capabilities that humans have; most of them successful in some way, but all of them quite different to humans. The onslaught of LLMs becomes a push in a particular capability direction. Understanding that direction becomes crucial to figuring out which types of human intellectual patterns and capabilities will become more important in the future. Where the new domains of human excellence will appear. But this understanding can also help us develop different types of AI that are more complementary to what humans can do and like to do. Seeing intelligence as a scalar, where machines can overtake humans, is a recipe for paralysis; dissolving this faulty notion gives us the freedom to act. More of this argument in the article I linked above; what's important here is that there are things that can be done. Indeed, there's a lot to do. Such as building mechanisms for meaningfully incorporating humans in creative search processes, and open-ended learning and discovery processes that are not based on imitating what humans do.
But not everything has a technical fix. Norms, structures, and laws are probably more important. I still think that trying to automate humans out of the processes that give them, and our civilization, meaning is immoral. We need to build counter-narratives, and be vocal that "fully automating the economy" is not an acceptable goal to work towards. We also need to push hard to counter the centralization of power that so easily comes with lavishly funded tech companies trying to achieve monopolies on some layers of the AI stack. Our best bet for a future where we all matter is one with a myriad different forms of intelligence, open and accessible for all to use as tools for our natural intelligences. Let's get to work.
-
An AI researcher writes about his crisis of faith
24 votes -
Comment on Your executable is a SQLite database in ~comp
skybrian LinkFrom the article: [...] [...] [...] [...]From the article:
I never let the idea go and with the recent improvements with LLMs, I find it compelling to revisit these ideas to explore further. Specifically, can we replace ELF with SQLite as an executable format? 🤔
Not “a database that describes an executable”, but the actual file you
chmod +xand run.[...]
I developed a pretty fleshed out prototype. It is called SELF, the Structured Executable & Linkable Format, because I am unoriginal. It is on GitHub if you are interested. I’m surprised about all the interesting things that fall out of this idea.
[...]
There is a fixed ~5 ms to open SQLite and start the interpreter, plus a copy proportional to the image. That copy is worse than it looks, because the b-tree pages are not mapped into memory. Two processes running the same SELF binary do not share text pages the way a normally-mmap‘d ELF does, because the bytes are copied out of the b-tree rather than mapped.33You might notice that
curl(274 KiB, 27 libraries) starts slower than ELFgit(4.6 MiB, 5 libraries). That isld.sodoing work proportional to the number of objects rather than the number of bytes, which I have complained about before.[...]
611.9 MiB of database against 644.4 MiB of ELF files. The whole userland, as one queryable file, is smaller than the files it came from. The b-tree cost that doubled a single
helloamortises to nearly nothing across 1,123 objects and is roughly 6% over the actual program bytes.The libraries and closure are shared across the executables very similar to how Nix might share them across multiple closures, if the store-path was the same. If every root shipped its own private closure (i.e. the AppImage model), the same 723 programs would come to 5.53 GiB but the deduplication of libraries and symbols falls out naturally from the database schema.
[...]
The format is done and round-trips between ELF and SELF losslessly. The tooling is done and can query, modify, and pack closures. Lookup through SQL works on unmodified glibc programs perfectly and the native-SQL loader works enough to explore it as a possibility for ideas.
The whole thing is at fzakaria/selfdb.
nix run .#self-vmboots a NixOS VM wherehellois a SQLite database. 🙌 -
Your executable is a SQLite database
21 votes -
Comment on AI text watermarking is free and good in ~comp
skybrian Link ParentIt's going to arbitrarily choose among identifiers that the LLM thinks are equally good. That will depend on context. For example, the LLM might prefer to maintain consistency with surrounding...It's going to arbitrarily choose among identifiers that the LLM thinks are equally good. That will depend on context. For example, the LLM might prefer to maintain consistency with surrounding code, which means that in a particular context, a consistent choice is better. Or, maybe there's a style guide in context?
It will also change as models get smarter and/or more opinionated about good coding style.
-
Comment on US judge to baby: file an asylum application in ~society
skybrian Link ParentClearly, this result is absurd and unjust. But maybe the judge knew that? Maybe it's their way of calling attention to the injustice of the system? That's pure speculation about the motives of a...Clearly, this result is absurd and unjust. But maybe the judge knew that? Maybe it's their way of calling attention to the injustice of the system?
That's pure speculation about the motives of a perfect stranger and I don't consider it particularly likely that I guessed right. But you're doing the same. We all speculate sometimes, but we shouldn't be confident that our speculations are correct. "Must certainly be" seems a bit much?
We shouldn't judge people without an investigation, and it's unreasonable and no fun to expect readers to investigate most of what we read in the news. I think that usually means saying "that sounds terrible" and moving on.
-
Comment on AI text watermarking is free and good in ~comp
skybrian Link ParentLet me try to explain it a slightly different way: Yes, some randomly-chosen word choices are better and others are worse. However, the LLM doesn't "know" that. The AI uses the random number...Let me try to explain it a slightly different way:
Yes, some randomly-chosen word choices are better and others are worse. However, the LLM doesn't "know" that. The AI uses the random number generator to choose among what the LLM considers to be "equivalent" paraphrases. If they don't seem equivalent to you, it's because you know better than the model.
Sometimes the LLM does "know" that one word is better than other in a given circumstance. For example one word is the right answer and another word is wrong. But if it knew that, it wouldn't defer its choice to the random number generator. The probability distribution would be so skewed that it forces the right answer.
How can an AI lab optimize the probability distribution to serve you better? They could train a new model or continue to train one that they already have. Messing with the random number generator isn't going to do it.
-
Comment on AI text watermarking is free and good in ~comp
skybrian Link ParentAI chat is often using a random number generator to decide what to write. If you're concerned about LLM's not serving you wholeheartedly, maybe you should be concerned about that, too? They're...AI chat is often using a random number generator to decide what to write. If you're concerned about LLM's not serving you wholeheartedly, maybe you should be concerned about that, too? They're rolling the dice to decide what to tell you! How does that serve you?
Swapping one random number generator for another isn't going to change that.
Though of course it's not just random. The weights bias the results, making some answers much more likely than others.
Overall, the answers being chosen from are in some sense equivalent. Even though they might not seem at all equivalent to you, the AI has no preference between them.
But perhaps a better model would have a preference? That's pretty much what happens when switching to an improved model.
Running the same query multiple times can be a good way of seeing what a model considers to be equivalent. Though, maybe this watermarking scheme would reduce the variety since it's using a biased generator? It's not going to make it better or worse on average, but it will reduce the number of answers that it's choosing from.
-
Comment on AI text watermarking is free and good in ~comp
skybrian Link ParentAnthropic links to this paper. Apparently the tokens are chosen to score highly according to a random scoring function. This is not very clear to me, but here’s the key bit: Google’s AI summary...Anthropic links to this paper. Apparently the tokens are chosen to score highly according to a random scoring function. This is not very clear to me, but here’s the key bit:
The key idea of Tournament sampling is to use a tournament-like process to choose an output token that scores highly with respect to some random watermarking functions. An illustration is given in Fig. 2 (top). First, we take the random seed rt provided by the random seed generator. This seed is passed to m (in this case, m = 3) watermarking functions g1, g2, g3, …, gm—these are independent pseudorandom number functions that assign a score gℓ(xt, rt) (in this case, a 0 or 1) to any candidate token xt ∈ V.
In the second stage (Fig. 2, bottom), we start by sampling M = 2m candidate tokens from the LLM distribution pLM(⋅∣x<t) (some tokens may appear multiple times): these are the initial participants of the m-layer tournament. We randomly divide these candidates into M/2 pairs, and, in the first tournament layer, in each pair the token with the higher score under g1(⋅, rt) is selected, and the other discarded (any ties are broken randomly). The remaining M/2 tokens are regrouped randomly into M/4 pairs, and the function g2(⋅, rt) determines the winners for this second tournament layer. This iterative process continues until one token emerges as the final winner, which becomes the output token xt. A formal description of Tournament sampling is given in Algorithm 2 in Methods.
By design, Tournament sampling selects a token from the LLM distribution that is likely to score higher under the random watermarking functions g1(⋅, rt), …, gm(⋅, rt). To detect whether a piece of text x = x1, …, xT is watermarked, we measure how highly x scores with respect to these functions.
Google’s AI summary seems a bit clearer, so maybe the thing to do is to ask your friendly AI what this all means?
It sounds like they start with the scoring functions and produce the text specifically so it scores high. So, they only need the scoring functions to verify the text. But they do need to keep them secret to prevent forgery, so we will have to rely on their API to tell us whether the text was from their LLM or not. Yeah, that’s not ideal.
Do scanned photos and documents count? If they do, the scans that my wife and I have from going through our parents' stuff are the oldest. That's what I see when I scroll all the way back in Google Photos.
Although, technically I've taken photos of much older things at museums. I didn't try to file them under their original dates, though :-)
Actual files, but maybe not readable: perhaps some of the floppy disks from the Commodore 64 at my Mom's house are still readable?