FlippantGod's recent activity

  1. Comment on Unreleased OpenAI model escapes containment and hacks into Hugging Face in ~tech

    FlippantGod
    Link Parent
    Humans had time to respond: Humans at huggingface were able to observe an issue, take an action in response, decide it was not working, implement a different course of action, presumably each with...

    Humans had time to respond:

    Humans at huggingface were able to observe an issue, take an action in response, decide it was not working, implement a different course of action, presumably each with a meaningful devhour cost, and seemingly resolve the issue before any unrecoverable event occurred.

    Perhaps they would have liked to be faster, but no matter how I look at it, at a bare minimum, I can make a case that humans had, at minimum, some time to respond and had a meaningful impact.

    Which seems to be what the author I quoted argues against. Perhaps I am assuming their weakest argument instead of their strongest, but their language and hyperboles seem unprofessional and rub me the wrong way.

    3 votes
  2. Comment on Unreleased OpenAI model escapes containment and hacks into Hugging Face in ~tech

    FlippantGod
    Link Parent
    This author makes, IMO, unsubstantiated claims. In particular, the hypothesized framework and harness Huggingface notes in their statements (as compared to a model alone) can be expected to...

    ... Mythos has what one might call The Juice, in that it can independently find without being directed, and string together, vulnerabilities into full exploit chains, essentially on its own...

    What happened later... 100% requires The Juice.

    Huggingface: "The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness..."

    All of this went down far too fast for a human-driven response. The only way HuggingFace could hope to do anything like keep pace was to use their own AIs.

    This author makes, IMO, unsubstantiated claims. In particular, the hypothesized framework and harness Huggingface notes in their statements (as compared to a model alone) can be expected to contribute in some meaningful way to the capabilities observed.

    And for the last one, it holds that not only did humans have time to respond, it was essential to have humans responding, as apparently they were forced to abandon their initial attempt at applying LLMs to the problem. I don't know how exactly the author defines "human-driven" but I am not persuaded.

    2 votes
  3. Comment on Long time no sea: finding the seahorse emoji in ~tech

  4. Comment on Your cookware got worse on purpose: who owns Pyrex and All-Clad now in ~food

  5. Comment on Building an Arch Linux aarch64 port for Holo Core in ~games

    FlippantGod
    Link
    Ironically, over at gitlab.steamos.cloud/holo/holo-core-aarch64-preview:

    You may also face infrastructure problems that did not exist when those package versions were originally released. Nowadays, many upstream services actively defend themselves against aggressive bots and AI crawlers by limiting download rates, restricting connections, or temporarily blocking clients. This can be painful for CI pipelines that rely on downloading sources directly from upstream.

    Ironically, over at gitlab.steamos.cloud/holo/holo-core-aarch64-preview:

    Failed to calculate challenge: undefined

    You are seeing this because the administrator of this website has set up Anubis to protect the server against the scourge of AI companies aggressively scraping websites.

    This website is running Anubis version v1.25.0.

    3 votes
  6. Comment on Long time no sea: finding the seahorse emoji in ~tech

    FlippantGod
    Link
    Underrated title, great article. Never heard of that portrait creator feature before. What links I've followed are interesting also.

    Underrated title, great article. Never heard of that portrait creator feature before. What links I've followed are interesting also.

  7. Comment on Claude Code is the ceiling on vibe-coded software in ~tech

    FlippantGod
    Link Parent
    FTFY: Haha you're absolutely right.

    FTFY: Haha you're absolutely right.

    7 votes
  8. Comment on Wealthy AI workers send San Francisco house prices soaring in ~finance

    FlippantGod
    Link Parent
    Wasn't that the point of the Solano Forever thing? What ever happened to it?

    Wasn't that the point of the Solano Forever thing? What ever happened to it?

    1 vote
  9. Comment on Agentic test processes, LLM benchmarks, and other notes on agentic coding in ~comp

    FlippantGod
    Link Parent
    I disagree, it is also about testing LLMs and a critique of most benchmarks. Unfortunately, it is also about why this is hard and sort of unsatisfying and inconclusive. Still, two people who have...

    I disagree, it is also about testing LLMs and a critique of most benchmarks. Unfortunately, it is also about why this is hard and sort of unsatisfying and inconclusive.

    Still, two people who have read this article should have a slightly easier time discussing LLM benchmarking with each other.

  10. Comment on Winners of the 2025 International Obfuscated C Code Contest in ~comp

  11. Comment on Steam Machine and Steam Frame will be shipping this summer in ~games

    FlippantGod
    Link Parent
    You are mixing up the devices, or this is a joke too advanced for my brain. Frame will have soldered LPDDR5X.

    You are mixing up the devices, or this is a joke too advanced for my brain. Frame will have soldered LPDDR5X.

    6 votes
  12. Comment on What change would make you quit Tildes? in ~tildes

    FlippantGod
    Link Parent
    Good for you I'm not Deimos, that's a lot of good ideas right there. Except Chrome.

    Good for you I'm not Deimos, that's a lot of good ideas right there. Except Chrome.

    10 votes
  13. Comment on Website is unhappy in ~tildes

    FlippantGod
    Link Parent
    That sounds a lot like something I imagine a fae might say. Kinda sus NGL.

    That sounds a lot like something I imagine a fae might say. Kinda sus NGL.

    9 votes
  14. Comment on Website is unhappy in ~tildes

    FlippantGod
    Link Parent
    While sending the children thoughts is quite respectable, I encourage you to also send prayers. This typically causes it to be twice again as convincing.

    While sending the children thoughts is quite respectable, I encourage you to also send prayers. This typically causes it to be twice again as convincing.

    12 votes
  15. Comment on Website is unhappy in ~tildes

  16. Comment on Website is unhappy in ~tildes

    FlippantGod
    (edited )
    Link
    Gross.... even Tildes cannot escape the increasing sexualization of public websites. We should maybe make the site invite only .

    Gross.... even Tildes cannot escape the increasing sexualization of public websites. We should start requiring state and federal U.S. government face recognition scans, and maybe make the site invite only too.

    12 votes
  17. Comment on I think Anthropic and OpenAI have found product-market fit in ~tech

    FlippantGod
    Link
    Probably the wrong thread, but I have been tossing an idea around in my head. Suppose current LLM performance matches claims, and scales as claimed. Observe the same trend in hardware performance....

    Probably the wrong thread, but I have been tossing an idea around in my head.

    Suppose current LLM performance matches claims, and scales as claimed. Observe the same trend in hardware performance.

    I am reminded of myths around early CGI studios:

    After months of negotiations.... Alvy ran the numbers very seriously for the first time. He discovered to his dismay that Moore's Law still had not proceeded far enough to make computation of a movie cost effective, and so backed out of the deal, just as Pixar was spinning out from Lucasfilm.
    Alvy Pixar Myth 3

    A few observations:

    1. LLMs become better/faster than any team of human developers or older LLM
    2. hardware continues to improve
    3. running LLMs or employing developers has costs

    Why not keep a product idea to oneself, and risk waiting a few years? Avoid wasting money in the meanwhile on current LLMs and humans.

    Depending on one's forecast, it would be faster to implement, with better results, later. Faster presumably translates into cheaper.

    Some ideas would be at risk of immediate copycats though. Only non-software restrictions like legislation, startup capital, access to customers or prerequisite data could feasibly block a new entrant. Some ideas might be differentiated by data quality also.

    But it stands to reason that anything without those "moats" have no protection. And depending on how far LLM scaling is projected to go, even those may not matter.

    3 votes
  18. Comment on US FBI says Google engineer used internal search data to win $1.2M on Polymarket in ~tech

    FlippantGod
    (edited )
    Link
    Can anyone tell me, could Google or Alphabet simply make a portion of its business "Prediction Market Trading, deriving leading real-world prediction market results from Google's valuable IP...

    Can anyone tell me, could Google or Alphabet simply make a portion of its business "Prediction Market Trading, deriving leading real-world prediction market results from Google's valuable IP portfolio" and do exactly this, but without the issue of "insider trading"?

    Seems legally viable but very much murky.

    Edit: "exactly this" is not what I meant. Oops. I meant, can Google leverage its data and statistics to make predictions and profit, not "bet on the contents of its own upcoming press releases".

    3 votes
  19. Comment on Tesla’s newest electric vehicle could jolt the trucking industry in ~transport

  20. Comment on The boy that cried Mythos in ~comp

    FlippantGod
    Link Parent
    GPT-2 was ultimately near enough trivial in compute and dataset. Maybe worse than existing methods of harm they identified. They were testing staged releases and delays to collect more usage data,...

    We are aware that some researchers have the technical capacity to reproduce and open source our results. We believe our release strategy limits the initial set of organizations who may choose to do this.

    While the misuse risk of 345M is higher than that of 117M, we believe it is substantially lower than that of 1.5B, and we believe that training systems of similar capability to GPT‑2‑345M is well within the reach of many actors already; this evolving replication landscape has informed our decision-making about what is appropriate to release.

    GPT-2 was ultimately near enough trivial in compute and dataset. Maybe worse than existing methods of harm they identified.

    They were testing staged releases and delays to collect more usage data, IMO. As it turns out they were already studying RLHF.

    And they began selling access to much more powerful models.

    Just felt like a big joke a year or two later when I understood it better.

    3 votes