5 votes

Astra is hard to monitor

1 comment

  1. streblo
    Link
    An article breaking down the loss of monitor-ability in OpenAI's latest model, some possible reasons why, and what could be next. At the end of the article, the author is making the case that what...

    An article breaking down the loss of monitor-ability in OpenAI's latest model, some possible reasons why, and what could be next.

    All AIs think both in ways that we can monitor, since we can at least monitor the output and this includes some of the thinking. They also all think in at least some ways we can’t monitor, because that’s what happens when computers do math.

    All signs point to monitorable Chain of Thought going away with some combination of larger and more capable models and the current training techniques, even if everyone otherwise behaves responsibly. The smarter you are, the more you can hold in your head and System-1-style thoughts, the less you need to put your thoughts into your System-2-style CoT in order to accomplish things.

    This is not a reason to stop fighting as hard as we can, to preserve as much monitorability as we can in as many ways as we can. Taboos around breaking down such techniques exist for a reason and should not be messed with lightly.

    At the end of the article, the author is making the case that what we need right now is more monitor-ability of the AI labs, something I strongly agree with. Here is something related I saw on Bluesky a few weeks ago that really resonated with me:

    Repulsed by the idea that we should sit idly by while frontier labs determine the trajectory of the AI transition. They don’t know how to make systems “safe” or “aligned“ nor do any such notions make sense without democratic input. There’s so much research to do here.

    I think, without ever acknowledging it explicitly, I had sort of resigned myself to this. But now I feel weirdly re-inspired to do some useful work here

    5 votes