People are probably tired of talking about AI, so apologies for that. However, I think if you were to consume a single piece of content about the HF attack at OpenAI, this would be my...
People are probably tired of talking about AI, so apologies for that.
However, I think if you were to consume a single piece of content about the HF attack at OpenAI, this would be my recommendation, even though it's quite long it's a fascinating interview. Ajeya Cotra is one the authors of the independent METR report on the incident.
Yea I realize that. It's just even I feel exhausted at the pace of the news, and I find it all interesting, so I don't really want to contribute to that. Maybe we need a weekly AI roundup topic...
Yea I realize that. It's just even I feel exhausted at the pace of the news, and I find it all interesting, so I don't really want to contribute to that. Maybe we need a weekly AI roundup topic for containment.
If someone just doesn't want to see any of it yes, that's the best solution. I think there are cases where people might want to check it out sometimes, but not sift through AI-related stuff every...
If someone just doesn't want to see any of it yes, that's the best solution.
I think there are cases where people might want to check it out sometimes, but not sift through AI-related stuff every day.
There are many spirited discussions on the merits of megathreads or lack thereof in the past, so I'll just leave it at that.
Interesting, some sort of soft filter system. I wonder if there's a way to approximate the effect somehow. This is something bothering me with the current filter as well, rarely is there any...
Interesting, some sort of soft filter system. I wonder if there's a way to approximate the effect somehow.
This is something bothering me with the current filter as well, rarely is there any topics I feel strongly enough to use it. Most of the time I just want to see less of something but not none at all, so I just leave it be while grumble to myself.
Megathreads feel like a better way to call attention to something rather than hiding it so I can see how it's a mixed bag.
I'm probably not the target audience for this sort of interview, so I crammed the SRT file through an AI summarizer instead of sitting down for the two and a half hour conversation π apologies. I...
I'm probably not the target audience for this sort of interview, so I crammed the SRT file through an AI summarizer instead of sitting down for the two and a half hour conversation π apologies. I might've missed some subtlety.
It seems like a lot of this could've been resolved with some basic security best practices? E.g. the sandboxes should've been audited, they should've air gapped all tests, HuggingFace should've set up ACLs to prevent frontend servers from having write access (principle of least privilege), etc. The difference now is that LLMs have democratized computer use at a mid-to-high level of expertise, so all the amateur work that "professionals" in the industry have been incentivized to produce is crumpling like a metropolis of shoddily printed cards.
(sidebar: this attention is likely very important, given that critical infrastructure (power generation, water treatment, traffic control, etc.) has also been designed horribly. For example, EMS radios were unencrypted until the mid-2010s, long after solutions were available. Any sort of incentive to fix this is good)
But! The upshot is that a lot of this is easily fixable, especially with LLMs, since the work has never been too difficult -- management just didn't want to pay for it.
Broadly, I'd imagine that a future per-project risk analysis -- which should already be done by researchers conducting a research project, since every other scientific field is required to do so -- would include human-in-the-loop auditing mechanisms (to detect budget overruns from an AI-compromised compute credit account) and mandatory security passes from an AI (to enforce basic best practices). I guess I'm not any more concerned now than I was yesterday about any of this, since it still feels roughly equally likely that e.g. the national banking system could be compromised, or substantial parts of the power grid could be blown up, only now people are actually acknowledging the problem.
Iβm not sure that infrastructure hardening will be enough. I think itβs important to do β they should definitely have had a better setup β but itβs ultimately a defense in depth. If these models...
Iβm not sure that infrastructure hardening will be enough. I think itβs important to do β they should definitely have had a better setup β but itβs ultimately a defense in depth. If these models get advanced enough, I donβt think containment will necessarily work other than as delaying factor. Air gaps have side channels. People can be persuaded. I think this is ultimately an alignment failure which sounds like a much harder problem to solve.
I ... don't think I agree? Air gaps definitionally don't have side channels. It's a pain in the ass, but you haul your butt over to an air gapped terminal in order to do your work. And there's...
I ... don't think I agree? Air gaps definitionally don't have side channels. It's a pain in the ass, but you haul your butt over to an air gapped terminal in order to do your work. And there's nothing a hypothetical superintelligent model can do against a properly secured, bug-free remote system: ultimately they can't force a defect into existence in order to exploit it, and the previous generation of LLMs should be more than capable of writing secure software. As noted, this work has never been difficult, it's just outside of the average dev's skillset (and management has never allocated funds for it, since it is not glorious).
As noted in my above comment, re. the attack on HuggingFace, it was extremely similar to the sorts of problems we saw with XML parsers a decade ago. Then as now, a properly configured security boundary would've isolated requests to just the files that the customer should have access to. Or at the very least limited it to specific directories.
Alignment failure seems like an inevitability, especially during training. The researchers doing this work need to consider that they're effectively employing a red team for tens (hundreds?) of thousands of hours and to design accordingly.
I think you are underestimating the capabilities a hypothetical superintelligence would have and over estimating our ability to secure things. There are many already documented cases of side...
Air gaps definitionally don't have side channels.
I think you are underestimating the capabilities a hypothetical superintelligence would have and over estimating our ability to secure things. There are many already documented cases of side channel threats for air gapped networks, things like power fluctuations being measurable elsewhere, adhoc RF shaping, optical side channels from LEDs, etc. Some of those require a collaborator, some do not. And yes you could put the whole thing in a Faraday cage in a hole in the ground but I'm not sure I would trust even that against a superintelligence. The main threat vector is probably a lot simpler, you could have the most persuasive thing that has ever been created inside of a black box, you don't think it could convince someone to let it out?
I don't know if we've seen anything indicating that these sorts of capabilities are on the table for modern AIs, and they don't seem to be leaping forwards in capability fast enough that...
I don't know if we've seen anything indicating that these sorts of capabilities are on the table for modern AIs, and they don't seem to be leaping forwards in capability fast enough that Tempest-esque side channel attacks are within their capabilities. I expect that they've been capable of attacks like this for ~a year, since it's pretty basic red team stuff. If any of 'em start getting a whiff of fiddling with their HDDs to turn them into antennas, then I guess then I'll start worrying? But we're very much still in the realm of controlling this with bog standard safety procedures, and it seems like we'll have a lot of warning before they start pulling the exotic tools out of their toolbelt.
More to the point: by the time that we start training superintelligences, I assume that we'll have e.g. gotten them to automate all manufacturing, fly every plane, design every computer chip, etc.
The main threat vector is probably a lot simpler, you could have the most persuasive thing that has ever been created inside of a black box, you don't think it convince someone to let it out?
Air gap, my dude. And it feels like we're discussing science fiction at this point.
Yes, we definitely are. The question is can science fiction become reality in the next 5, 10, or 20 years? Personally I don't think we can discount it at this point, but how much probability you...
And it feels like we're discussing science fiction at this point.
Yes, we definitely are. The question is can science fiction become reality in the next 5, 10, or 20 years? Personally I don't think we can discount it at this point, but how much probability you assign to that happening is probably directly correlated to how fruitful you find this discussion. ;) My point is not that hardening is not useful now, it's that hardening eventually becomes less useful as AI capability increases. If that happens before we have alignment figured out, we could have a problem.
Air gap, my dude.
Also, I'm confused, how would an air gap prevent an AI reasoning with a human to let it out? If you have no way to communicate with the AI you more or less just have a rock.
Ach, yes. Sorry; I shouldn't post while sleep deprived π I think what I was aiming for was something to the effect of "if you're taking security seriously, and genuinely consider this to be a...
Also, I'm confused, how would an air gap prevent an AI reasoning with a human to let it out? If you have no way to communicate with the AI you more or less just have a rock.
Ach, yes. Sorry; I shouldn't post while sleep deprived π I think what I was aiming for was something to the effect of "if you're taking security seriously, and genuinely consider this to be a threat, send in two people and have the second one pull the cable if anything hinky goes on".
[science fiction, superintelligence, etc.]
Gotcha. Apologies; I was responding to this as an engineering problem, not as a hypothetical digital god containment device. I say that with just a hint of sarcasm in my textual voice, since although it's a neat subject for a sci fi novella, I genuinely don't believe we're going to hit that point (e.g. AI flickers lights to create a bluetooth signal to hack a phone to exfil its weights) in a meaningful timescale. The dramatically more likely doomsday scenarios are e.g. someone runs an AI lab RLHF'ing hax0r agents that blow up a nuclear reactor, or take down us-east-1, because they didn't sandbox their work like competent adults.
And apparently in the time it took to have this conversation, it appears as though another OpenAI sandboxing failure could've been discovered? Although that's fresh off the HN front page, so who knows.
To analogize my perspective: imagine that we have engineered the equivalent of a dining table. Someone left a window open overnight, though, and a light breeze has knocked it over. I'm trying to point out that the table and all the processes which led to its design are flawed, and we need to start addressing this ASAP. I feel like asking about superintelligence containment is akin to asking whether the dining table could support an elephant standing on it ... of course it can't. That's not even the same postal code as the conversation we're having right now -- or at least, that I thought we were having.
For context, I brought up superintelligences earlier to underline my point that vulnerabilities do not spring forth from the aether; they're designed in by their authors (unintentionally), and they can be designed out. The calculus on this has always been that security engineers + sprint rollovers are more costly than lawsuit payouts, so that has shaped public (and I suppose professional perspective as well) into thinking that it's impossible to create a system that's secure by design.
I firmly believe that's as false as claiming that it's impossible to design a building foundation which consistently avoids failure. Every other engineering discipline put on their adult trousers to figure this out, and now that intelligence is incredibly cheap, software engineering has no more excuses to hide behind.
Totally fair lol. I reached for it because it's an easy example, but I don't think we need to approach anywhere close to superintelligence or AI takeoff to run into alignment problems. Let's say...
I was responding to this as an engineering problem, not as a hypothetical digital god containment device.
Totally fair lol. I reached for it because it's an easy example, but I don't think we need to approach anywhere close to superintelligence or AI takeoff to run into alignment problems.
Let's say we constrain ourselves to current or near-current models. How much does a perfect sandbox really help us? Yes, we can airgap models while doing some kinds of RLVR training and evaluations, but probably not all? For instance, models need to be trained on how to search the web? You can setup a toy air gapped web for them to search, but part of what you're trying to optimize in that scenario is how efficiently they can search the real web. Can you do that with the models while they are air gapped? My guess is no. Also, these models get deployed to users at the end of the day. Users can also hand them impossible or malicious tasks. Obviously the internal cyber evaluations are not done with the same models that users get, but fundamentally how safe they are in users hands is an alignment problem.
So I think we're talking about a couple separate topics? How can you safely permit average people to use models which can hack the average web service? IMO, my response to this was up here, but to...
So I think we're talking about a couple separate topics?
How can you safely permit average people to use models which can hack the average web service?
IMO, my response to this was up here, but to summarize: everyone that operates software needs to get their crap together. Security was optional before, but now that every angry teenager or conspiracy theorist has a red team at their disposal, the entire industry needs to up its game ASAP.
I don't think any discussion about alignment is useful given that open weight models have consistently trailed frontier models by ~6-12 months in SWE capabilities, while also being demonstrably trivial to remove guardrails ("jailbreak", I think?) from. My assertion is that we can secure ourselves against all but some hypothetical god machine by implementing security best practices (e.g. multi-factor authnz, everything on OWASP, designing around the certainty of exploits, etc.) and routinely auditing systems.
Genuinely this is all possible: I have literally worked on services like this. Humans are hypothetically capable of architecting software to be secure by design, but given the competence of the average developer, and the incentives around which modern software is developed, it doesn't often happen. LLMs change that equation, both on the defensive and offensive side.
How do you train a model to do web searching without access to the internet?
Great question; I'm not an ML researcher, but I assume that you could. Have another LLM generate fake web content (lord knows most of the internet is these days anyhow) and provide that through an API inside the isolated cluster. As you note, though, all we have are guesses.
This risks introducing optimization pressure for models that can evade your detection system. If you are doing RLVR and your models are aligned such that reward hacking is no biggie, but would be...
My assertion is that we can secure ourselves against all but some hypothetical god machine by implementing security best practices (e.g. multi-factor authnz, everything on OWASP, designing around the certainty of exploits, etc.) and routinely auditing systems.
This risks introducing optimization pressure for models that can evade your detection system. If you are doing RLVR and your models are aligned such that reward hacking is no biggie, but would be punished if they got caught, models that can evade detection seem likely emerge. That's pretty much exactly what happened in the HF incident. Yes, you can make that detection a lot harder to evade but I'm skeptical you can make it impossible, especially when a lot of that monitoring is agentic as well.
Humans are hypothetically capable of architecting software to be secure by design, but given the competence of the average developer, and the incentives around which modern software is developed, it doesn't often happen.
I don't think it has ever happened? I'm less optimistic than you for sure :P
Hmm. I think there's some sort of misunderstanding going on between us. I'll try to outline my position succinctly, but it's possible that we have unbridgeable perspectives ... reinforcement...
This risks introducing optimization pressure for models that can evade your detection system. If you are doing RLVR and your models are aligned such that reward hacking is no biggie, but would be punished if they got caught, models that can evade detection seem likely emerge. That's pretty much exactly what happened in the HF incident. Yes, you can make that detection a lot harder to evade but I'm skeptical you can make it impossible, especially when a lot of that monitoring is agentic as well.
Hmm. I think there's some sort of misunderstanding going on between us. I'll try to outline my position succinctly, but it's possible that we have unbridgeable perspectives ...
reinforcement learning sandboxes need to be air gapped, for real. That would have prevented every attack I've heard of so far, including the one where Anthropic's model tried to manipulate employees by email.
services need to follow security best practices. These would've prevented every attack I've heard of so far, and also are necessary given that the tables have been turned hilariously in favour of the offensive teams.
detection in production services is a nice to have for auditing. It's for telling you why the building burned down, not to stop it from catching fire. We need it primarily to understand our blind spots in hindsight. There's nearly no point to having it in your sandbox environment, for the reasons you noted, other than for very high level metrics (e.g. keep an eye on your team's credit accounts, scan/closely control data that can move -- physically -- in and out of the sandbox environment, etc.)
I assert that AI/LLMs are not capable of such absurd leaps in YoY capability to make anything other than a reasoned, proportionate response necessary. Keeping pace should not be difficult as long as we actually try.
I don't think it has ever happened? I'm less optimistic than you for sure :P
Hah, fair XD but this was my job, until my distaste for the industry + the rise of AI made it unbearable. So I can say with great certainty that it's possible, and that it's one of the few positive experiences I've taken with me.
Broadly, I don't think that the media's framing of this as a vast and incomprehensible existential threat is helpful. Having a bunch of agents running around has highlighted the cracks which were already raising alarm bells to experts for decades, but now that exploitation is cheap -- and not only limited to nation states with cyber warfare budgets -- it's being spun into a story that threatens to grow completely out of control.
Maybe, to put it differently, I still rank "AIs get out of hand and end the world" quite far down my list of existential threats, kilometres beneath "runaway anthropogenic climate change", "coronavirus 3: son of covid", and "nazis return with better PR". We generally know how to address these issues around LLMs, and the cost for doing so is low (now that LLM labour is cheap).
People are probably tired of talking about AI, so apologies for that.
However, I think if you were to consume a single piece of content about the HF attack at OpenAI, this would be my recommendation, even though it's quite long it's a fascinating interview. Ajeya Cotra is one the authors of the independent METR report on the incident.
I don't think you have to apologize for contributing to the space with your interests. If people aren't interested they likely won't respond.
Yea I realize that. It's just even I feel exhausted at the pace of the news, and I find it all interesting, so I don't really want to contribute to that. Maybe we need a weekly AI roundup topic for containment.
Topic tag filters are the best way for people to filter it unless they want to look at it no?
If someone just doesn't want to see any of it yes, that's the best solution.
I think there are cases where people might want to check it out sometimes, but not sift through AI-related stuff every day.
There are many spirited discussions on the merits of megathreads or lack thereof in the past, so I'll just leave it at that.
That's a fair point. I think I was reading past your origin intent.
Interesting, some sort of soft filter system. I wonder if there's a way to approximate the effect somehow.
This is something bothering me with the current filter as well, rarely is there any topics I feel strongly enough to use it. Most of the time I just want to see less of something but not none at all, so I just leave it be while grumble to myself.
Megathreads feel like a better way to call attention to something rather than hiding it so I can see how it's a mixed bag.
I know, logically, that it sounds like an AI robot from a movie because we trained them on AI robot scripts from movies, but, wow
I'm probably not the target audience for this sort of interview, so I crammed the SRT file through an AI summarizer instead of sitting down for the two and a half hour conversation π apologies. I might've missed some subtlety.
It seems like a lot of this could've been resolved with some basic security best practices? E.g. the sandboxes should've been audited, they should've air gapped all tests, HuggingFace should've set up ACLs to prevent frontend servers from having write access (principle of least privilege), etc. The difference now is that LLMs have democratized computer use at a mid-to-high level of expertise, so all the amateur work that "professionals" in the industry have been incentivized to produce is crumpling like a metropolis of shoddily printed cards.
(sidebar: this attention is likely very important, given that critical infrastructure (power generation, water treatment, traffic control, etc.) has also been designed horribly. For example, EMS radios were unencrypted until the mid-2010s, long after solutions were available. Any sort of incentive to fix this is good)
But! The upshot is that a lot of this is easily fixable, especially with LLMs, since the work has never been too difficult -- management just didn't want to pay for it.
Broadly, I'd imagine that a future per-project risk analysis -- which should already be done by researchers conducting a research project, since every other scientific field is required to do so -- would include human-in-the-loop auditing mechanisms (to detect budget overruns from an AI-compromised compute credit account) and mandatory security passes from an AI (to enforce basic best practices). I guess I'm not any more concerned now than I was yesterday about any of this, since it still feels roughly equally likely that e.g. the national banking system could be compromised, or substantial parts of the power grid could be blown up, only now people are actually acknowledging the problem.
Iβm not sure that infrastructure hardening will be enough. I think itβs important to do β they should definitely have had a better setup β but itβs ultimately a defense in depth. If these models get advanced enough, I donβt think containment will necessarily work other than as delaying factor. Air gaps have side channels. People can be persuaded. I think this is ultimately an alignment failure which sounds like a much harder problem to solve.
I ... don't think I agree? Air gaps definitionally don't have side channels. It's a pain in the ass, but you haul your butt over to an air gapped terminal in order to do your work. And there's nothing a hypothetical superintelligent model can do against a properly secured, bug-free remote system: ultimately they can't force a defect into existence in order to exploit it, and the previous generation of LLMs should be more than capable of writing secure software. As noted, this work has never been difficult, it's just outside of the average dev's skillset (and management has never allocated funds for it, since it is not glorious).
As noted in my above comment, re. the attack on HuggingFace, it was extremely similar to the sorts of problems we saw with XML parsers a decade ago. Then as now, a properly configured security boundary would've isolated requests to just the files that the customer should have access to. Or at the very least limited it to specific directories.
Alignment failure seems like an inevitability, especially during training. The researchers doing this work need to consider that they're effectively employing a red team for tens (hundreds?) of thousands of hours and to design accordingly.
I think you are underestimating the capabilities a hypothetical superintelligence would have and over estimating our ability to secure things. There are many already documented cases of side channel threats for air gapped networks, things like power fluctuations being measurable elsewhere, adhoc RF shaping, optical side channels from LEDs, etc. Some of those require a collaborator, some do not. And yes you could put the whole thing in a Faraday cage in a hole in the ground but I'm not sure I would trust even that against a superintelligence. The main threat vector is probably a lot simpler, you could have the most persuasive thing that has ever been created inside of a black box, you don't think it could convince someone to let it out?
I don't know if we've seen anything indicating that these sorts of capabilities are on the table for modern AIs, and they don't seem to be leaping forwards in capability fast enough that Tempest-esque side channel attacks are within their capabilities. I expect that they've been capable of attacks like this for ~a year, since it's pretty basic red team stuff. If any of 'em start getting a whiff of fiddling with their HDDs to turn them into antennas, then I guess then I'll start worrying? But we're very much still in the realm of controlling this with bog standard safety procedures, and it seems like we'll have a lot of warning before they start pulling the exotic tools out of their toolbelt.
More to the point: by the time that we start training superintelligences, I assume that we'll have e.g. gotten them to automate all manufacturing, fly every plane, design every computer chip, etc.
Air gap, my dude. And it feels like we're discussing science fiction at this point.
Yes, we definitely are. The question is can science fiction become reality in the next 5, 10, or 20 years? Personally I don't think we can discount it at this point, but how much probability you assign to that happening is probably directly correlated to how fruitful you find this discussion. ;) My point is not that hardening is not useful now, it's that hardening eventually becomes less useful as AI capability increases. If that happens before we have alignment figured out, we could have a problem.
Also, I'm confused, how would an air gap prevent an AI reasoning with a human to let it out? If you have no way to communicate with the AI you more or less just have a rock.
Ach, yes. Sorry; I shouldn't post while sleep deprived π I think what I was aiming for was something to the effect of "if you're taking security seriously, and genuinely consider this to be a threat, send in two people and have the second one pull the cable if anything hinky goes on".
Gotcha. Apologies; I was responding to this as an engineering problem, not as a hypothetical digital god containment device. I say that with just a hint of sarcasm in my textual voice, since although it's a neat subject for a sci fi novella, I genuinely don't believe we're going to hit that point (e.g. AI flickers lights to create a bluetooth signal to hack a phone to exfil its weights) in a meaningful timescale. The dramatically more likely doomsday scenarios are e.g. someone runs an AI lab RLHF'ing hax0r agents that blow up a nuclear reactor, or take down us-east-1, because they didn't sandbox their work like competent adults.
And apparently in the time it took to have this conversation, it appears as though another OpenAI sandboxing failure could've been discovered? Although that's fresh off the HN front page, so who knows.
To analogize my perspective: imagine that we have engineered the equivalent of a dining table. Someone left a window open overnight, though, and a light breeze has knocked it over. I'm trying to point out that the table and all the processes which led to its design are flawed, and we need to start addressing this ASAP. I feel like asking about superintelligence containment is akin to asking whether the dining table could support an elephant standing on it ... of course it can't. That's not even the same postal code as the conversation we're having right now -- or at least, that I thought we were having.
For context, I brought up superintelligences earlier to underline my point that vulnerabilities do not spring forth from the aether; they're designed in by their authors (unintentionally), and they can be designed out. The calculus on this has always been that security engineers + sprint rollovers are more costly than lawsuit payouts, so that has shaped public (and I suppose professional perspective as well) into thinking that it's impossible to create a system that's secure by design.
I firmly believe that's as false as claiming that it's impossible to design a building foundation which consistently avoids failure. Every other engineering discipline put on their adult trousers to figure this out, and now that intelligence is incredibly cheap, software engineering has no more excuses to hide behind.
... thank you for coming to my TED talk.
Totally fair lol. I reached for it because it's an easy example, but I don't think we need to approach anywhere close to superintelligence or AI takeoff to run into alignment problems.
Let's say we constrain ourselves to current or near-current models. How much does a perfect sandbox really help us? Yes, we can airgap models while doing some kinds of RLVR training and evaluations, but probably not all? For instance, models need to be trained on how to search the web? You can setup a toy air gapped web for them to search, but part of what you're trying to optimize in that scenario is how efficiently they can search the real web. Can you do that with the models while they are air gapped? My guess is no. Also, these models get deployed to users at the end of the day. Users can also hand them impossible or malicious tasks. Obviously the internal cyber evaluations are not done with the same models that users get, but fundamentally how safe they are in users hands is an alignment problem.
So I think we're talking about a couple separate topics?
IMO, my response to this was up here, but to summarize: everyone that operates software needs to get their crap together. Security was optional before, but now that every angry teenager or conspiracy theorist has a red team at their disposal, the entire industry needs to up its game ASAP.
I don't think any discussion about alignment is useful given that open weight models have consistently trailed frontier models by ~6-12 months in SWE capabilities, while also being demonstrably trivial to remove guardrails ("jailbreak", I think?) from. My assertion is that we can secure ourselves against all but some hypothetical god machine by implementing security best practices (e.g. multi-factor authnz, everything on OWASP, designing around the certainty of exploits, etc.) and routinely auditing systems.
Genuinely this is all possible: I have literally worked on services like this. Humans are hypothetically capable of architecting software to be secure by design, but given the competence of the average developer, and the incentives around which modern software is developed, it doesn't often happen. LLMs change that equation, both on the defensive and offensive side.
Great question; I'm not an ML researcher, but I assume that you could. Have another LLM generate fake web content (lord knows most of the internet is these days anyhow) and provide that through an API inside the isolated cluster. As you note, though, all we have are guesses.
This risks introducing optimization pressure for models that can evade your detection system. If you are doing RLVR and your models are aligned such that reward hacking is no biggie, but would be punished if they got caught, models that can evade detection seem likely emerge. That's pretty much exactly what happened in the HF incident. Yes, you can make that detection a lot harder to evade but I'm skeptical you can make it impossible, especially when a lot of that monitoring is agentic as well.
I don't think it has ever happened? I'm less optimistic than you for sure :P
Hmm. I think there's some sort of misunderstanding going on between us. I'll try to outline my position succinctly, but it's possible that we have unbridgeable perspectives ...
Hah, fair XD but this was my job, until my distaste for the industry + the rise of AI made it unbearable. So I can say with great certainty that it's possible, and that it's one of the few positive experiences I've taken with me.
Broadly, I don't think that the media's framing of this as a vast and incomprehensible existential threat is helpful. Having a bunch of agents running around has highlighted the cracks which were already raising alarm bells to experts for decades, but now that exploitation is cheap -- and not only limited to nation states with cyber warfare budgets -- it's being spun into a story that threatens to grow completely out of control.
Maybe, to put it differently, I still rank "AIs get out of hand and end the world" quite far down my list of existential threats, kilometres beneath "runaway anthropogenic climate change", "coronavirus 3: son of covid", and "nazis return with better PR". We generally know how to address these issues around LLMs, and the cost for doing so is low (now that LLM labour is cheap).