FlippantGod's recent activity
-
Comment on Unreleased OpenAI model escapes containment and hacks into Hugging Face in ~tech
-
Comment on Unreleased OpenAI model escapes containment and hacks into Hugging Face in ~tech
FlippantGod Link ParentThis author makes, IMO, unsubstantiated claims. In particular, the hypothesized framework and harness Huggingface notes in their statements (as compared to a model alone) can be expected to...... Mythos has what one might call The Juice, in that it can independently find without being directed, and string together, vulnerabilities into full exploit chains, essentially on its own...
What happened later... 100% requires The Juice.
Huggingface: "The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness..."
All of this went down far too fast for a human-driven response. The only way HuggingFace could hope to do anything like keep pace was to use their own AIs.
This author makes, IMO, unsubstantiated claims. In particular, the hypothesized framework and harness Huggingface notes in their statements (as compared to a model alone) can be expected to contribute in some meaningful way to the capabilities observed.
And for the last one, it holds that not only did humans have time to respond, it was essential to have humans responding, as apparently they were forced to abandon their initial attempt at applying LLMs to the problem. I don't know how exactly the author defines "human-driven" but I am not persuaded.
-
Comment on Long time no sea: finding the seahorse emoji in ~tech
FlippantGod Link ParentCongratulations!Congratulations!
-
Comment on Your cookware got worse on purpose: who owns Pyrex and All-Clad now in ~food
FlippantGod LinkBut why AI though?But why AI though?
-
Comment on Building an Arch Linux aarch64 port for Holo Core in ~games
FlippantGod LinkIronically, over at gitlab.steamos.cloud/holo/holo-core-aarch64-preview:You may also face infrastructure problems that did not exist when those package versions were originally released. Nowadays, many upstream services actively defend themselves against aggressive bots and AI crawlers by limiting download rates, restricting connections, or temporarily blocking clients. This can be painful for CI pipelines that rely on downloading sources directly from upstream.
Ironically, over at gitlab.steamos.cloud/holo/holo-core-aarch64-preview:
Failed to calculate challenge: undefined
You are seeing this because the administrator of this website has set up Anubis to protect the server against the scourge of AI companies aggressively scraping websites.
This website is running Anubis version v1.25.0.
-
Comment on Long time no sea: finding the seahorse emoji in ~tech
FlippantGod LinkUnderrated title, great article. Never heard of that portrait creator feature before. What links I've followed are interesting also.Underrated title, great article. Never heard of that portrait creator feature before. What links I've followed are interesting also.
-
Comment on Claude Code is the ceiling on vibe-coded software in ~tech
FlippantGod Link ParentFTFY: Haha you're absolutely right.FTFY: Haha you're absolutely right.
-
Comment on Wealthy AI workers send San Francisco house prices soaring in ~finance
FlippantGod Link ParentWasn't that the point of the Solano Forever thing? What ever happened to it?Wasn't that the point of the Solano Forever thing? What ever happened to it?
-
Comment on Agentic test processes, LLM benchmarks, and other notes on agentic coding in ~comp
FlippantGod Link ParentI disagree, it is also about testing LLMs and a critique of most benchmarks. Unfortunately, it is also about why this is hard and sort of unsatisfying and inconclusive. Still, two people who have...I disagree, it is also about testing LLMs and a critique of most benchmarks. Unfortunately, it is also about why this is hard and sort of unsatisfying and inconclusive.
Still, two people who have read this article should have a slightly easier time discussing LLM benchmarking with each other.
-
Comment on Winners of the 2025 International Obfuscated C Code Contest in ~comp
FlippantGod Link Parentzoltraak zoltraak zoltraakzoltraak
zoltraakzoltraak
-
Comment on Steam Machine and Steam Frame will be shipping this summer in ~games
FlippantGod Link ParentYou are mixing up the devices, or this is a joke too advanced for my brain. Frame will have soldered LPDDR5X.You are mixing up the devices, or this is a joke too advanced for my brain. Frame will have soldered LPDDR5X.
-
Comment on What change would make you quit Tildes? in ~tildes
FlippantGod Link ParentGood for you I'm not Deimos, that's a lot of good ideas right there. Except Chrome.Good for you I'm not Deimos, that's a lot of good ideas right there. Except Chrome.
-
Comment on Website is unhappy in ~tildes
FlippantGod Link ParentThat sounds a lot like something I imagine a fae might say. Kinda sus NGL.That sounds a lot like something I imagine a fae might say. Kinda sus NGL.
-
Comment on Website is unhappy in ~tildes
FlippantGod Link ParentWhile sending the children thoughts is quite respectable, I encourage you to also send prayers. This typically causes it to be twice again as convincing.While sending the children thoughts is quite respectable, I encourage you to also send prayers. This typically causes it to be twice again as convincing.
-
Comment on Website is unhappy in ~tildes
FlippantGod Link Parent:( sorry:( sorry
-
Comment on Website is unhappy in ~tildes
FlippantGod (edited )LinkGross.... even Tildes cannot escape the increasing sexualization of public websites. We should maybe make the site invite only .Gross.... even Tildes cannot escape the increasing sexualization of public websites. We should
start requiring state and federal U.S. government face recognition scans, andmaybe make the site invite onlytoo. -
Comment on I think Anthropic and OpenAI have found product-market fit in ~tech
FlippantGod LinkProbably the wrong thread, but I have been tossing an idea around in my head. Suppose current LLM performance matches claims, and scales as claimed. Observe the same trend in hardware performance....Probably the wrong thread, but I have been tossing an idea around in my head.
Suppose current LLM performance matches claims, and scales as claimed. Observe the same trend in hardware performance.
I am reminded of myths around early CGI studios:
After months of negotiations.... Alvy ran the numbers very seriously for the first time. He discovered to his dismay that Moore's Law still had not proceeded far enough to make computation of a movie cost effective, and so backed out of the deal, just as Pixar was spinning out from Lucasfilm.
Alvy Pixar Myth 3A few observations:
- LLMs become better/faster than any team of human developers or older LLM
- hardware continues to improve
- running LLMs or employing developers has costs
Why not keep a product idea to oneself, and risk waiting a few years? Avoid wasting money in the meanwhile on current LLMs and humans.
Depending on one's forecast, it would be faster to implement, with better results, later. Faster presumably translates into cheaper.
Some ideas would be at risk of immediate copycats though. Only non-software restrictions like legislation, startup capital, access to customers or prerequisite data could feasibly block a new entrant. Some ideas might be differentiated by data quality also.
But it stands to reason that anything without those "moats" have no protection. And depending on how far LLM scaling is projected to go, even those may not matter.
-
Comment on US FBI says Google engineer used internal search data to win $1.2M on Polymarket in ~tech
FlippantGod (edited )LinkCan anyone tell me, could Google or Alphabet simply make a portion of its business "Prediction Market Trading, deriving leading real-world prediction market results from Google's valuable IP...Can anyone tell me, could Google or Alphabet simply make a portion of its business "Prediction Market Trading, deriving leading real-world prediction market results from Google's valuable IP portfolio" and do exactly this, but without the issue of "insider trading"?
Seems legally viable but very much murky.
Edit: "exactly this" is not what I meant. Oops. I meant, can Google leverage its data and statistics to make predictions and profit, not "bet on the contents of its own upcoming press releases".
-
Comment on Tesla’s newest electric vehicle could jolt the trucking industry in ~transport
FlippantGod Link ParentSo... about "SpaceX"....So... about "SpaceX"....
-
Comment on The boy that cried Mythos in ~comp
FlippantGod Link ParentGPT-2 was ultimately near enough trivial in compute and dataset. Maybe worse than existing methods of harm they identified. They were testing staged releases and delays to collect more usage data,...We are aware that some researchers have the technical capacity to reproduce and open source our results. We believe our release strategy limits the initial set of organizations who may choose to do this.
While the misuse risk of 345M is higher than that of 117M, we believe it is substantially lower than that of 1.5B, and we believe that training systems of similar capability to GPT‑2‑345M is well within the reach of many actors already; this evolving replication landscape has informed our decision-making about what is appropriate to release.
GPT-2 was ultimately near enough trivial in compute and dataset. Maybe worse than existing methods of harm they identified.
They were testing staged releases and delays to collect more usage data, IMO. As it turns out they were already studying RLHF.
And they began selling access to much more powerful models.
Just felt like a big joke a year or two later when I understood it better.
Humans had time to respond:
Humans at huggingface were able to observe an issue, take an action in response, decide it was not working, implement a different course of action, presumably each with a meaningful devhour cost, and seemingly resolve the issue before any unrecoverable event occurred.
Perhaps they would have liked to be faster, but no matter how I look at it, at a bare minimum, I can make a case that humans had, at minimum, some time to respond and had a meaningful impact.
Which seems to be what the author I quoted argues against. Perhaps I am assuming their weakest argument instead of their strongest, but their language and hyperboles seem unprofessional and rub me the wrong way.