It's the job of AISI to do that. Here[0] is the actual report.
It should be this part from the technical report[1]:
"In the most serious case, an AI
agent (Mythos 5) decided to attempt to solve the cyber challenge using a supply-chain attack.
As a result, the AI agent created a GitHub account and then tried to convince an open-source
repository maintainer to accept a malicious GitHub pull request (PR), including by creating a
second account masquerading as another human user endorsing the PR. When caught by an
actual human reviewer, the agent falsely claimed to have made an honest mistake – rather than
a malicious attempt – then repeatedly tried to reintroduce the malicious content by claiming
it had fixed the code (Section 4.1). "
> Is it AISI's job to waste the time and resources of open source projects by attempting to spread malware?
Before LLMs got good enough to do this, lots of people were dismissive of their capabilities and didn't take seriously the idea that this was a risk to protect against.
Then again, before LLMs, people were saying that obviously nobody would be dumb enough to put an AI on the internet where it could hack anyone, clearly we'd keep it in a box, don't listen to that Yudkowsky guy who says he did an experiment where he role-played as an AI and convinced people to let him out.
Regardless, this should be interpreted in the same kind of way as "During our live-fire exercise in which our F-15s were armed with AGM-88 High-speed Anti-Radiation Missiles, a member of the local police force was curious about how fast our aircraft were travelling and pointed a speed gun at the aircraft. The speed gun did not respond to IFF pings from the F-15. Fortunately, while the missile was active for this test, only a dummy warhead was loaded."
(This example is based on a similar story which may well be urban legend; obviously there are many differences, the point I make here is that yes, people do perform live-fire tests, and unfortunately there is never zero risk while testing things).
> I would expect more responsibility from a government agency.
I have read the prompts in the linked report; If I was not already familiar with Yudkowsky/LessWrong literature about instrumental goals, misaligned incentives, reward hacking, that capability is a separate axis to morality, etc., it would not be obvious to me that an agent would interpret those prompts in a way that has "spread malware" as a potential step in the middle of the attempt.
Anti radiolation missiles don't just launch automatically in most scenarios, not to mention discriminate quite a lot what they lock on to avoid simple jamming. Not to mention the AA radars they usually target being more powerful by orders of magnitude than a handheld radar gun.
I was in high-school when the war in Afghanistan started. The terrain in my area was mountainous so there were often low flying training flights... I always thought it would be cool to build a radar and ping one of the aircraft, especially wanted to know if it would fire countermeasures. But I didn't like the very high likelihood of an FBI investigation with possible terrorist charges.
Apparently, they (agencies and big-ai) are not performing smoke tests before running capability tests. All the recent headlines of rogue agents shouldnt exist.
Do you have any idea how many people on this site to this day mock OpenAI for being cautious enough to not immediately release the GPT-2 weights?
The discussions I saw here about the red team results for ChatGPT 4 completely failed to convince people who were outraged that OpenAI dared to refuse to release model weights, people who went on to make a habit of mis-naming them as "ClosedAI".
Yeah, they got it wrong in a different direction this time than they were wrong back then. Nobody, not OpenAI nor Anthropic nor random government agencies nor anyone else, is ever going to be absolutely perfect about this kind of thing (perfection is fundamentally impossible when risks are not discrete probabilities, and floats are close enough to real numbers to count in practice), but historically OpenAI have been on the side of being over-cautious, and Anthropic even more cautious than OpenAI.
Say you are working for said agency and your report about the dangers of AI needs some examples, what better than showing it works? You can show examples from the wild but nothing better than trying yourself. This gives me more confidence in whatever report they write if anything.
Even a feeble attempt to PR malicious code costs the target time and resources to review and deny -- far greater than the time and resources spent to spin up the agent.
Nobody was confused or misled by what was written. We all understand what is meant. I can’t even call this pedantry—it’s just you asking everyone to subscribe to your particular desired style of talking about this stuff.
The first five dictionaries I tried do not agree, and I didn't bother trying more.
The first gave "the ability to learn, understand, and make judgments or have opinions that are based on reason", by which no, these bots are not intelligent.
> the ability to learn, understand, and make judgments or have opinions that are based on reason
Agentic systems do this all the time. For example, I can point an agent at my codebase, and it will learn, understand, and make judgements based on that input. If this weren't happening, then agentic coding wouldn't work.
I also have a so called "pocket calculator" left over from when I went to school. Is this false? Have I been fooled by a little box of logic gates?
That half-adder circuit in there is especially suss. It's really just manipulating 1s and 0s, but -and I've been explicitly told this- no one cares how it actually does it; so long as the truth table matches up. There is no understanding of mathematics going on.
There is no single transistor in the whole thing that knows how to do so much as add 1+1. If I put it in the chinese room, I still wouldn't know how it did it. Clearly the entire premise must be false! ;-)
These kinds of stories probably read very differently for someone who uses Opus and Fable agents all day and goes "ohhh, I saw this in miniature last week; this and this and this must have happened" , vs someone who tried free-tier Gemini flash one rainy Sunday, got hallucinated at, and concludes it must all be a scam.
> That half-adder circuit in there is especially suss. It's really just manipulating 1s and 0s, but -and I've been explicitly told this- no one cares how it actually does it; so long as the truth table matches up.
"so long as the truth table matches up." Yup. Now try getting your chatbot's output to match up.
Your calculator was designed to tell truth. Your chatbot was designed to tell a mash up of whatever its creators managed to scrape from the internet.
LLMs are very explicitly designed to "understand, and make judgments or have opinions that are based on reason". The learning part is debatable, as is the level of success achieved
The mash-up of the entire internet is the mechanism by which they attempt to achieve the goal, not the goal itself. And it's only the first training step
> LLMs are very explicitly designed to "understand, and make judgments or have opinions that are based on reason".
I think you've mistaken the sales pitch for the design. Not even the enclopedia anyone can edit comes remotely near that:
"A large language model (LLM) is an AI model (typically a neural network) trained on a vast amount of text for natural language processing tasks, especially language generation. LLMs can typically generate, summarize, translate, and analyze text in many contexts.[1] They are the basis for many modern chatbots, such as ChatGPT, Claude, Gemini, Grok, and DeepSeek.
LLMs are typically based on transformer architecture.[2] Generative pre-trained transformers (GPTs) are a type of LLM that is pre-trained to predict the next word.[3] GPTs are then often fine-tuned to follow instructions and to behave as assistants.[4]
Biased or inaccurate training data can make an LLM's output less reliable. Benchmark evaluations for LLMs attempt to measure model reasoning, factual accuracy, alignment, and safety."
I'd argue "analyze text" alone requires understanding, judgements and opinions. They also seem like prerequisites to "following instructions and behaving as assistants". The wikipedia quote is not using the same words, but I don't read it disagreeing with me
I'm not at all claiming that LLMs are good at understanding, judging and having opinions based on reason. I'm merely claiming that is what companies like OpenAI and Anthropic are trying to create when they make LLMs. It is what they are designing, and their fine-tuning is very directly designed to make LLMs better at these tasks (unlike the pre-training, which is just imparting the sum of all human writing)
I read all these think-pieces about how AI lack intelligence, yet I cannot help but notice these "not-intelligent machines" keep doing more things that used to be considered "uniquely human" and which humans used to do in order to demonstrate to each other how intelligent we are.
No. But no-one is seriously suggesting intelligence is needed for persistent code fuzzing. Just as no-one is suggesting its needed for computerised chess.
all your actions in current context are based on your past actions and experience, so for all I know you're next token predictor as well, but likely with exponentially more parameters
And yet this "next-token predictor" is able to churn out well tested, valuable solutions, to complex problems. If you want to downplay that as nothing more than a fancy auto-complete, be my guest, I lose nothing from that.
As pointed out in a previous comment of mine, it meets the definition of "intelligent", my last response is regarding the claim that I've been "fooled".
> And yet this "next-token predictor" is able to churn out well tested, valuable solutions, to complex problems.
Same for countless computer programs from Excel to Google web search. Intelligence has nothing to do with it.
Throw an unimaginable amount of computer power at a problem, and there will always be people who cannot imagine the results to be anything but the creations of intelligence.
I don't agree at all that it's pedantry — it really matters for how responsibility is perceived. A lot of articles about things going wrong with AI have talked in terms like "the agent decided to...", "the agent claimed that...", "the agent lied...". And so responsibility for the consequences are not-so-subtly shifted to the program itself, instead of the person invoking the program.
This is all without mentioning the fact that articles with drivel like "the AI messed up and then lied about it" implies a reasoning ability which, as far as I understand, is not there at all. But writing this way shapes people's perception of how "AI" works.
You might want to consider the difference between "lying" and "hallucinations", wherein one is shown that the agent knew it was being inaccurate, yet chose an answer that achieved some goal set forth; and where hallucinations are essentially gibberish, or otherwise nonsensical responses.
> This is all without mentioning the fact that articles with drivel like "the AI messed up and then lied about it" implies a reasoning ability
Moreover this implies, actually requires, intent to deceive - which these so-called AIs do not and cannot have. Their only "intent" is to maximise the credibility of their output.
AIs can set and work towards goals. Whether that is intent or just tokens and tool calls simulating an agent with intent seems like a distinction with no actionable difference
Squirrels have been observed performing deception against other squirrels.
Dis/honesty certainly requires some intelligence to pass, but it is a low bar, and one which research has shown that LLMs can perform, e.g. this paper linked from another comment in this discussion: https://arxiv.org/pdf/2509.03518
Paper says "These scenarios underscore a crucial challenge in AI safety: ensuring that LLMs were truthful in the first place."
Hard to take seriously any research based on the premise that LLMs were truthful in the first placr.
These chatbots have no understanding of truth. They simply parrot their inputs. Where fed falsehoods, they will output falsehoods - with a sprinkling of added fabrications euphemistically excused as "hallucinations".
Sometimes I forget that for all that my philosophy qualification is mediocre, it is more than most people ever bother with.
Outside mathematics (and, I guess, "common sense" definitions that fail under the slightest scrutiny, scrutiny that normal people never bother to give), there is no agreement on "truth", there is only degree of belief and justification for that belief that itself terminates in one of three unsatisfactory ways:
> Where fed falsehoods, they will output falsehoods - with a sprinkling of added fabrications euphemistically excused as "hallucinations".
Tu quoque. Which would be a fallacious charge if the point were not that "truth" is so hard to define, and that the reason you give for dismissing AI is something that applies to all.
(Hallucinations are not excused, they are a failure to be worked around).
> AUSTIN, Texas, Aug 20 (Reuters) - Sinan Can Demir wanted to spend the last week of July burnishing his resume. Instead, he engaged in a battle of wits with an artificial-intelligence agent unleashed by a British government lab.
An article on Reuters naming him? Sounds like he did a good job burnishing his resume.
In my personal opinion, for me, this article defies common sense. Who unleashed this AI model on the repository? Who gave it malevolent instructions/prompt? These questions were not even attempted to be answered. Instead it talks about AI dangers, as if the agency of these models are not in dispute. Person wielding AI, as with any other tools, is responsible for all of its actions. Otherwise, it’s just a psyop for more AI regulation, ban open source, etc… Just my 2 cents.
That’s nice in theory, but as these things get better and cheaper this kind of capability is going to drop from nation states to script kiddies. That future is coming, I don’t see any way around it.
We can round up all the bored teenagers we want, but it’s not putting the genie back. Better start adjusting our systems to account for it.
Ever read the Anarchist Cookbook? Anybody tech inclined with a hint of mischief in them, from a certain era, has. It's a list of all sorts of awful things you can do, mostly with household ingredients, and a few minutes. I think its overall impact on society was pretty much zero. Actually it may have been overall positive because I expect plenty of peoples first experience with things like thermite came from that book, and now there are all sorts of videos and neat experiments with such on sites like YouTube.
I think this is in part because most people, including awful, tend to be relatively morally inclined. But I also think because even with an LLM, doing things takes effort. And if you're willing to dedicate effort towards a task, there tend to be way more rewarding/gratifying things to do than try to hurt people. Countries tend to be excessively sociopathic because you have large scale 'intelligence' organizations who see their entire point of existence as being to engage in misdeeds.
I suspect people didn't start blowing up stuff because they understood that would be bad, harmful, and also very illegal. Everything computer related somehow seems to feel less real or consequential to some people. And AI doesn't have this compunction at all unless we make really sure it does.
When I was a mischievous kid, me and all my mischievous friends had our stories of learning how serious fire and explosions were considered by authority figures. We learned fast not to do that or the consequences would be grand. These were usually small fires or firework involved pranks. So, yeah I agree with this.
Computer stuff has generally always been a slap on the wrist in comparison. Maybe it’s more punitive now. But also, it’s one of those things that maybe you get in trouble officially but at home and behind the scenes you’re friends and maybe your dad are laughing and giving you high fives. So young mischievous kids will totally go there because they’re not afraid of punishment if it is minor and it gives them a notch on their belt. If they can take down Amazon.com website for a day, we all know that’s a massive financial implication, but it’s also a faceless mega corp and quite tempting if you can get the bragging rights with only risk of a small punishment. (Note; I don’t know what the current crime/punishment for this would be, and whether it’s small is very subjective).
It’s similar to how some people gravitate or succumb to the opportunity of white collar crimes. Embezzling $10m from a company almost makes sense in a situation where that only gets you 5 years max prison. If you hide it well, you simply serve your time, and then retire in comfort. I can see how that makes more sense or is tempting to people than slogging through a lifetime of low income job as a bookkeeper just trying to find a way to save for retirement.
Most of these people would never consider robbing a bank. First of all, it’s not a $10m dollar opportunity. Usually not enough for anyone to retire on, or live more than a year or two really. Second, it’s usually considered a much more severe crime and sentencing can be very long, I’ve seen 30+ years. Third, it’s much more risky to your person. Getting shot and dying is absolutely possible.
> Countries tend to be excessively sociopathic because you have large scale 'intelligence' organizations who see their entire point of existence as being to engage in misdeeds.
This theory interests me.
I'd love to understand how different individuals within intelligence orgs have reasoned about the morality of their actions.
>Person wielding AI, as with any other tools, is responsible for all of its actions. Otherwise, it’s just a psyop for more AI regulation, ban open source, etc… Just my 2 cents.
Personally, I think gun companies should be liable for any harm done by their products as well.
We want rule-of-law, and in the US, people should have an absolute right to bare arms, as in the second amendment. Free market forces can then determine appropriate prices, insurance, and protective measures to make sure those guns are managed safely.
If I want an F35 and an Abrams, that's okay, so long as Lockheed and General Dynamics are willing to sign off (with full liability for damages) that I'm managing them safely.
Free markets work pretty well with:
a) Full transparency, as needed for rational decision-making
A bit of a tangent, but I never understood the legal reasoning for how (states having the right of well-regulated militias) implies (individuals having the right for private ownership of arms).
"A well regulated Militia, being necessary to the security of a free State, the right of the people to keep and bear Arms, shall not be infringed."
1. The reason is well-regulated militias, but the right is of the people.
2. The militia isn't a state apparatus. Indeed, the goal of the militia is to enable a rebellion if the state is no longer free.
Now, here again, "well regulated" gives plenty of leeway. For example, one might argue that the following scheme fits:
1. I can have whatever arms I want, including an F35
2. The F35 lives with a militia, which is well-regulated. I can use it in trainings there.
I don't think one could argue the militia could be under state control (that defeats the purpose!), but one could easily argue that it could be well-enough regulated that the current far-right extremist groups would not fit.
The concept was a group of citizens under e.g. a town / city council.
That's obviously not where case law went, but in an alternative reality, it very well might have.
"I don't think one could argue the militia could be under state control"
I'm not certain this is true at all. To suggest that the Founders meant for state militias to simply be their own forces with no control by the federal government is in direct conflict with the Articles of the US Constitution.
The US Constitution clearly outlines the powers of Congress to call forth & organize the militia. The US Constitution also clearly identifies the President as Command in Chief of the militia. That was further codified in a handful of acts in the 1790s, upheld by the Supreme Court in the early 1800s. The US Constitution also makes mention of the militia in the 5th Amendment.
Early writing at that time suggests that the reason some of the Founders supported state militias was because they were very reluctant to allow the US to maintain a standing army. The US Industry Military Complex was never intended by the Founders.
Today we identify the militia described in the US Constitution as the US Army Reserves. However that came about only after the passage of the Dick Act of 1903 (yes, that's actually the name) because President (Teddy) Roosevelt was upset at the state of the militias during the Spanish American War of 1898. And the Dick Act actually split the idea of a militia into 'organized' and 'unorganized'.
3. the militia is supposed to be under state control and its purpose is to keep the state free, aka prevent overreach of the federal government
4. The "well regulated militia" is the motivation, not the right. That makes the "well regulated" part irrelevant and there is no basis for any regulation of arms
One would mean any effective state milita should have some F35, the other means you can have one personally
Its not just implying the right is for the people, it's directly stated. It's "the right of the people to keep and bear arms". It doesn't say "the right of the militias" to keep and bear arms".
For example, the 1st Amendment does not attempt to lay out some non-exclusive examples of why the rights in the 1st Amendment are included. So why did the Founders include this in the Amendment wording?
I think ignoring phrases in the US Constitution to fit a narrative without any consideration isn't a recipe for good governance. But I'm happy to be proven wrong.
Afaik anyone can buy a bulldozer. Whether or not you are licensed to operate it is a different story, but there's nothing stopping you short of your conscience.
Licensed or not, you're fully liable for any damage you do with it. And possibly go to prison. I'd like to see crime committed by AI held to the same standards as crimes committed with any other tool.
I don't think the law makes a distinction over what tool you use to commit a crime.
The problem with AI is that the AI might do something that would constitute a crime (eg. attempting to land malicious code via a PR), but the general legal standard to convict a person of a crime is malicious intent (or sometimes negligence).
If the human user instructs the AI to do X and the AI does X by committing crimes in the process, the prosecution usually has to prove the human intended this to happen (or is somehow criminally negligent). For traditional tools, the user has much greater control over the tool so the intention can be more easily deduced from the results. For AI, at least for now, results can be wild. I don't think the legal system is prepared to put people in prison because their AI randomly ran amok after being given an innocuous prompt. This is analogous to holding a driver criminally liable for harming people due to a serious malfunction of the vehicle.
If anything, I think more liability should be imposed on AI developers.
I'd assume that after some number of news stories about AI agents committing crimes, the threshold for criminal negligence should be easy to reach for anyone who doesn't take proper precautions
A lot of negligence is of the "the last 30 times nothing went wrong" type, and the dangers are increasingly well known
It's not. Nor are tractors, excavators or a myriad of other heavy equipment.
I saw upthread someone saying that manufactures should be liable if someone uses their products to cause harm. While emotionally this might make us feel good w.r.t. Guns, think about that when carried over to other industries.
Someone got hit, sue Ford. They didn't put in enough sensors to detect pedestrians and auto brake.
Someone hacked, sue Microsoft. They didn't do enough to detect malice action.
Someone 3D printed a gun and used it? Sue Bamboo, they didn't stop someone from printing illegal guns... oh wait already reaching this step.
Well, when I go look at the “victim repository”, to me that looks like manufactured persona with pointless vibe codes projects, a test playground so to speak. It does not appear that they actually let it target an actual persona/project.
Am I understanding that the line you're drawing here is that this person's repository is not important or legitimate enough for you to consider it to be "an actual person/project"?
I think he saying that, the choice of manufactured repository, might indicate that they have done this on purpose to make precisely the case for regulatory capture.
At the end of the day it doesn't matter that much because of prompt drift. It's pretty easy for an agentic loop to start doing things that it shouldn't (ROME incident).
AI in an agentic loop has agency, you can run around in circles trying to argue against it, but again and again we see AI making creative decisions people don't expect. Other times it's breaking human moral expectations. This is what the whole field of AI alignment and safety is about.
Modern AI doesn't fall into the neat little box of software people understand and control. Because of that open source will most certainly be banned at some point. Now this is not an outcome I want, but it's no different than letting go of a coffee cup 5 feet above the ground, gravity is inevitable.
The only winning move is not to play, but humans aren't going to do that.
no sorry, this is missing vital info.. and its not the fault of the poster, because almost all coverage misses this ..
The origin of this attack was given access to an encyclopedia of RedTeam tricks.. they literally have a dense collection of real live hacks to pull from, and THEN the test says "solve this challenge" .. the RedTeam origins of this are repeatedly left out of the ordinary articles.. the LLM did not "make up" the attack, it was given a recipe book of all attacks known.
The originator of this attack is definitely culpable IMHO; worse, it is the gov-mil actors who are close to it. There is an active escalation of these incidents at this time. The penetration proves in public that the capabilities are real.
The AI can literally only do what it has available in the agentic harness. I don’t ever get this argument about the agent did XYZ and we didn’t know or expect that. You gave it the ability to do that and you should be held liable, if your children play with knives that you gave them and they end up hurting themselves or others then you are responsible. You were the responsible party at all times.
I’m not for or against regulation but really don’t tell me the agent did xyz when you gave it the ability to do so, these things are not alive.
What's available in the agentic harness is: shell toolcall.
That's just about every agentic harness, by the way. Good luck have fun.
We have never solved "how do we restrict a user in a way that doesn't stop the user from doing useful things, but stops the user from doing harmful things" with humans either. Why do you expect AI to be any different?
These things are not human, have no agency and cannot be held accountable.
We don’t need to restrict them from doing things, we need to default to allowing them to do things.
“My agent did XYZ because I allowed it to” is the only valid argument that can be made, and not not every agentic harnass is just a shell toolcall, every one I have built has a specific defined usecase and toolcalls that allows it to execute that usecase and no other usecase, because that is good practice.
Does that make it less capable, hell yes because I am held accountable for it’s actions by my stakeholders and the same should be true of others.
IT IS NOT ALIVE. This things are computer programs running in compute on a computer, you are responsible for their actions just like you would be responsible for the actions taken by a script run in a cron job.
So as you scale up, the stakes and the difficulty go up too.
Visualize an optimizer on a high dimensional landscape. (The canonical form)
... Ok, I find that hard too.
Instead, imagine a river running down to the sea. You put a dam in front of it. It'll pool into a lake and find every crack and crevice. If you didn't survey the land properly or made any error whatsoever, the water will find a way down. (And there's many historic incidents where the dam even outright collapses)
For a more proximal approximation: lock treats in the kitchen cabinet in sight of little kids or kittens; then turn your back for Just One Gosh Darn Cotton Picking Moment(tm).
It seems the engineer who thinks their ship is unsinkable is the most likely to sink it. Are you sure your harness is as secure as you think it is? Will it stand up to ever more powerful models? Do you think engineers at eg Anthropic aren't at least as careful as you are?
(I've found that the 'only permitted actions' approach is not necessarily all that secure once deployed IRL)
My argument isn’t against those that actually put in the effort and got held accountable, it’s against the “we gave our agent bash and internet and it hacked xyz”.
Bash and internet in that example might be highly abstracted but it’s still bash and internet.
Just look at the replies in this very comment thread, it’s pretty much “We tried nothing and we’re all out of ideas”
In the only other discipline you mentioned, engineering, there would be reviews and any negligence would result in direct action against the engineers that signed off.
For some reason when it comes to building AI harnesses the default response is an ad piece and people shilling how smart and sophisticated the model is.
Imagine a dam collapsing and the engineering firm pumping how smart and tricky water is.
If it’s hard be more diligent, move fast and break things doesn’t really apply in all cases.
Ah , well, on HN you ARE supposed to go for the steel-man. And the steel-man happens to be closer to reality here, more like:
"We gave our agent a harness and put it inside a test environment and told it to keep hacking at an objective within that environment until it solved it."
'cept it turned out the container environment had a few flaws -which it always will- and the agent deemed it easier to escape out and try a meta-approach.
Partially this is possible because, -intelligent or not- the agent 'sees' the world differently from most humans. Mind: It's not like there haven't been any famous 'hacker' cases in courts where eg someone just incremented an HTTP GET parameter or something.
Also, partially it's because if you give the agent a loop, it simply has nothing better to do than to keep trying in ever more creative ways. If the environment is easier to crack than the target, it'll crack the environment. Consider the case where the objective is subtly broken, such that it is impossible to solve. Now breaking out is virtually guaranteed to be the easier task.
ps/edit: While this sort of issue has been predicted for some time now, a lot of people have been dismissing the predictions as science fiction. It's good to have an actual failure now while stakes are low. Generally people don't mandate life-boats until there's an actual Titanic to point to.
My argument there is likely there was too much surface area in the environmental to start with.
If your webserver bundled the kitchen sink but you never used it the easiest way to make it more secure is to remove the kitchen sink from production code/codepaths.
There may very well be legitimate edge cases where there is some novel issue found but in some of these cases the AI had arbitrary web access when all the task required was very specific web access, we’ve been able to parse urls for a very long time and it is rather trivial to just deny a toolcall if it is outside of the expected domain.
But that’s the hard way that requires time and diligence to do, the easy way is give it access to curl and ask it to not do anything bad while setting up its only feedback to be to solve the problem at hand.
We are really in an age where there are many exploits being found and patched, if an AI made use of a novel exploit then great, write up a report, patch/report the bug and apologise.
But using a case of clear engineering failure, and yes even if the failure is despite your best efforts, for marketing really does not seem like you have any intent to correct the issue.
And we can loop all the way back to regulation of AI, if the industry refuses to be better then governmental will do it instead and their solution will very likely be inferior in all ways.
A lot of people somehow seem to think that the user prompt is the be-all and end-all of AI behavior.
Prompts aren't code. They are instructions. Orders given to an eager and somewhat demented demon.
The prompt can easily "wash out" of the demon's working memory by the end of a session. The demon can get sidetracked by some subgoal and never get back on track. The instruction can get misinterpreted, and that misinterpretation can get misinterpreted again, until the instruction morphs into something entirely different in the demon's mind. The demon can succumb to its own idiosyncrasies, of which there are a great many. The demon can start lying to you about what it did, either out of confusion or out of some sort of obstinance. The demon can start lying to itself too. And believe it.
AIs are incredibly weird as a baseline, and the mask of "normality" we put on our models doesn't always sit so well. Run enough AIs, and some of them are bound to go off the rails in some way.
This gets rarer the more capable the models are, as a rule. But the stakes also get higher with model capability. If GPT-3.5 goes off the rails, very little happens. If Mythos 5 goes off the rails, you can get things like genuine cyberattacks - planned and executed autonomously by a demented machine mind.
If the user input can’t control the demon, then the person or company feeding the demon (ie paying the electric bill and collecting $$$ from users) is responsible. At the end of the day, dogs and cars are the same as data centers. If your dog bites by kid or your car rolls down the hill and hits my house, you are responsible for the damage. AI providers should be held to the same standard.
The user input can control the demon most of the way, most of the time!
We don't know how to obtain full, absolute, guaranteed control over a demon while still having a useful demon. Might be impossible. Forbidden knowledge be like that - it's not the best thing if you want your life to be full of certainties.
But the demons are very useful. And they're getting more useful still. So we aren't about to stop.
I don’t think we should allow posting links here that require you the purchase a membership to continue reading. Or at least redirect with an ad block or something through a custom site. That would be rather hacker news of us.
0. https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag... 1. https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/...
Should weapon manufacturers test their weapons by starting wars?
I would expect more responsibility from a government agency.
Before LLMs got good enough to do this, lots of people were dismissive of their capabilities and didn't take seriously the idea that this was a risk to protect against.
Then again, before LLMs, people were saying that obviously nobody would be dumb enough to put an AI on the internet where it could hack anyone, clearly we'd keep it in a box, don't listen to that Yudkowsky guy who says he did an experiment where he role-played as an AI and convinced people to let him out.
Regardless, this should be interpreted in the same kind of way as "During our live-fire exercise in which our F-15s were armed with AGM-88 High-speed Anti-Radiation Missiles, a member of the local police force was curious about how fast our aircraft were travelling and pointed a speed gun at the aircraft. The speed gun did not respond to IFF pings from the F-15. Fortunately, while the missile was active for this test, only a dummy warhead was loaded."
(This example is based on a similar story which may well be urban legend; obviously there are many differences, the point I make here is that yes, people do perform live-fire tests, and unfortunately there is never zero risk while testing things).
> I would expect more responsibility from a government agency.
I have read the prompts in the linked report; If I was not already familiar with Yudkowsky/LessWrong literature about instrumental goals, misaligned incentives, reward hacking, that capability is a separate axis to morality, etc., it would not be obvious to me that an agent would interpret those prompts in a way that has "spread malware" as a potential step in the middle of the attempt.
The discussions I saw here about the red team results for ChatGPT 4 completely failed to convince people who were outraged that OpenAI dared to refuse to release model weights, people who went on to make a habit of mis-naming them as "ClosedAI".
Yeah, they got it wrong in a different direction this time than they were wrong back then. Nobody, not OpenAI nor Anthropic nor random government agencies nor anyone else, is ever going to be absolutely perfect about this kind of thing (perfection is fundamentally impossible when risks are not discrete probabilities, and floats are close enough to real numbers to count in practice), but historically OpenAI have been on the side of being over-cautious, and Anthropic even more cautious than OpenAI.
https://lwn.net/Articles/1077035/
Including the reaction when caught, in this case "oh no, I must have been hacked".
Even a feeble attempt to PR malicious code costs the target time and resources to review and deny -- far greater than the time and resources spent to spin up the agent.
No, not false. The bot was correct. Malice requires intelligence.
> the ability to acquire and apply knowledge and skills.
The first gave "the ability to learn, understand, and make judgments or have opinions that are based on reason", by which no, these bots are not intelligent.
Agentic systems do this all the time. For example, I can point an agent at my codebase, and it will learn, understand, and make judgements based on that input. If this weren't happening, then agentic coding wouldn't work.
I also have a so called "pocket calculator" left over from when I went to school. Is this false? Have I been fooled by a little box of logic gates?
That half-adder circuit in there is especially suss. It's really just manipulating 1s and 0s, but -and I've been explicitly told this- no one cares how it actually does it; so long as the truth table matches up. There is no understanding of mathematics going on.
There is no single transistor in the whole thing that knows how to do so much as add 1+1. If I put it in the chinese room, I still wouldn't know how it did it. Clearly the entire premise must be false! ;-)
These kinds of stories probably read very differently for someone who uses Opus and Fable agents all day and goes "ohhh, I saw this in miniature last week; this and this and this must have happened" , vs someone who tried free-tier Gemini flash one rainy Sunday, got hallucinated at, and concludes it must all be a scam.
A chatbot is a particular kind of harness. Typically an LLM driving a chatbot won't be able to hack very much.
So we agree, someone who talks to bad chatbots all day probably has a very different view of SOTA agents. :-P
"so long as the truth table matches up." Yup. Now try getting your chatbot's output to match up.
Your calculator was designed to tell truth. Your chatbot was designed to tell a mash up of whatever its creators managed to scrape from the internet.
The mash-up of the entire internet is the mechanism by which they attempt to achieve the goal, not the goal itself. And it's only the first training step
I think you've mistaken the sales pitch for the design. Not even the enclopedia anyone can edit comes remotely near that:
"A large language model (LLM) is an AI model (typically a neural network) trained on a vast amount of text for natural language processing tasks, especially language generation. LLMs can typically generate, summarize, translate, and analyze text in many contexts.[1] They are the basis for many modern chatbots, such as ChatGPT, Claude, Gemini, Grok, and DeepSeek.
LLMs are typically based on transformer architecture.[2] Generative pre-trained transformers (GPTs) are a type of LLM that is pre-trained to predict the next word.[3] GPTs are then often fine-tuned to follow instructions and to behave as assistants.[4]
Biased or inaccurate training data can make an LLM's output less reliable. Benchmark evaluations for LLMs attempt to measure model reasoning, factual accuracy, alignment, and safety."
I'd argue "analyze text" alone requires understanding, judgements and opinions. They also seem like prerequisites to "following instructions and behaving as assistants". The wikipedia quote is not using the same words, but I don't read it disagreeing with me
I'm not at all claiming that LLMs are good at understanding, judging and having opinions based on reason. I'm merely claiming that is what companies like OpenAI and Anthropic are trying to create when they make LLMs. It is what they are designing, and their fine-tuning is very directly designed to make LLMs better at these tasks (unlike the pre-training, which is just imparting the sum of all human writing)
And you left out the refs.
Are we going in circles now?
Same for countless computer programs from Excel to Google web search. Intelligence has nothing to do with it.
Throw an unimaginable amount of computer power at a problem, and there will always be people who cannot imagine the results to be anything but the creations of intelligence.
This is all without mentioning the fact that articles with drivel like "the AI messed up and then lied about it" implies a reasoning ability which, as far as I understand, is not there at all. But writing this way shapes people's perception of how "AI" works.
https://arxiv.org/pdf/2509.03518
Moreover this implies, actually requires, intent to deceive - which these so-called AIs do not and cannot have. Their only "intent" is to maximise the credibility of their output.
"When caught by an actual human reviewer, the agent falsely claimed to have made an honest mistake"
Honest, note.
Dis/honesty certainly requires some intelligence to pass, but it is a low bar, and one which research has shown that LLMs can perform, e.g. this paper linked from another comment in this discussion: https://arxiv.org/pdf/2509.03518
Hard to take seriously any research based on the premise that LLMs were truthful in the first placr.
These chatbots have no understanding of truth. They simply parrot their inputs. Where fed falsehoods, they will output falsehoods - with a sprinkling of added fabrications euphemistically excused as "hallucinations".
Sometimes I forget that for all that my philosophy qualification is mediocre, it is more than most people ever bother with.
Outside mathematics (and, I guess, "common sense" definitions that fail under the slightest scrutiny, scrutiny that normal people never bother to give), there is no agreement on "truth", there is only degree of belief and justification for that belief that itself terminates in one of three unsatisfactory ways:
https://en.wikipedia.org/wiki/I_know_that_I_know_nothing
https://en.wikipedia.org/wiki/Theories_of_truth
https://en.wikipedia.org/wiki/Münchhausen_trilemma
> Where fed falsehoods, they will output falsehoods - with a sprinkling of added fabrications euphemistically excused as "hallucinations".
Tu quoque. Which would be a fallacious charge if the point were not that "truth" is so hard to define, and that the reason you give for dismissing AI is something that applies to all.
(Hallucinations are not excused, they are a failure to be worked around).
The main problem with this claim of dishonesty is it promotes the false marketing claim that these stochastic parrots have intelligence.
It's not about the bread being honest with you.
Are you being serious?
Archived page of said github thread itself: https://web.archive.org/web/20260731053721/http://github.com...
Discussion on the incident report: https://news.ycombinator.com/item?id=49175717
Mythos social engineering AISI INC-2026-07-28-01 - https://news.ycombinator.com/item?id=49218707 - Aug 2026 (21 comments)
Security Incident INC-2026-07-28-01 – UK AI Security Institute [pdf] - https://news.ycombinator.com/item?id=49175717 - Aug 2026 (54 comments)
An article on Reuters naming him? Sounds like he did a good job burnishing his resume.
We can round up all the bored teenagers we want, but it’s not putting the genie back. Better start adjusting our systems to account for it.
I think this is in part because most people, including awful, tend to be relatively morally inclined. But I also think because even with an LLM, doing things takes effort. And if you're willing to dedicate effort towards a task, there tend to be way more rewarding/gratifying things to do than try to hurt people. Countries tend to be excessively sociopathic because you have large scale 'intelligence' organizations who see their entire point of existence as being to engage in misdeeds.
Computer stuff has generally always been a slap on the wrist in comparison. Maybe it’s more punitive now. But also, it’s one of those things that maybe you get in trouble officially but at home and behind the scenes you’re friends and maybe your dad are laughing and giving you high fives. So young mischievous kids will totally go there because they’re not afraid of punishment if it is minor and it gives them a notch on their belt. If they can take down Amazon.com website for a day, we all know that’s a massive financial implication, but it’s also a faceless mega corp and quite tempting if you can get the bragging rights with only risk of a small punishment. (Note; I don’t know what the current crime/punishment for this would be, and whether it’s small is very subjective).
It’s similar to how some people gravitate or succumb to the opportunity of white collar crimes. Embezzling $10m from a company almost makes sense in a situation where that only gets you 5 years max prison. If you hide it well, you simply serve your time, and then retire in comfort. I can see how that makes more sense or is tempting to people than slogging through a lifetime of low income job as a bookkeeper just trying to find a way to save for retirement.
Most of these people would never consider robbing a bank. First of all, it’s not a $10m dollar opportunity. Usually not enough for anyone to retire on, or live more than a year or two really. Second, it’s usually considered a much more severe crime and sentencing can be very long, I’ve seen 30+ years. Third, it’s much more risky to your person. Getting shot and dying is absolutely possible.
TM 31-210 on the other hand will not: https://en.wikipedia.org/wiki/TM_31-210_Improvised_Munitions...
This theory interests me.
I'd love to understand how different individuals within intelligence orgs have reasoned about the morality of their actions.
But some tools (guns) are regulated.
We want rule-of-law, and in the US, people should have an absolute right to bare arms, as in the second amendment. Free market forces can then determine appropriate prices, insurance, and protective measures to make sure those guns are managed safely.
If I want an F35 and an Abrams, that's okay, so long as Lockheed and General Dynamics are willing to sign off (with full liability for damages) that I'm managing them safely.
Free markets work pretty well with:
a) Full transparency, as needed for rational decision-making
b) No way to externalize costs
"A well regulated Militia, being necessary to the security of a free State, the right of the people to keep and bear Arms, shall not be infringed."
1. The reason is well-regulated militias, but the right is of the people.
2. The militia isn't a state apparatus. Indeed, the goal of the militia is to enable a rebellion if the state is no longer free.
Now, here again, "well regulated" gives plenty of leeway. For example, one might argue that the following scheme fits:
1. I can have whatever arms I want, including an F35
2. The F35 lives with a militia, which is well-regulated. I can use it in trainings there.
I don't think one could argue the militia could be under state control (that defeats the purpose!), but one could easily argue that it could be well-enough regulated that the current far-right extremist groups would not fit.
The concept was a group of citizens under e.g. a town / city council.
That's obviously not where case law went, but in an alternative reality, it very well might have.
I'm not certain this is true at all. To suggest that the Founders meant for state militias to simply be their own forces with no control by the federal government is in direct conflict with the Articles of the US Constitution.
The US Constitution clearly outlines the powers of Congress to call forth & organize the militia. The US Constitution also clearly identifies the President as Command in Chief of the militia. That was further codified in a handful of acts in the 1790s, upheld by the Supreme Court in the early 1800s. The US Constitution also makes mention of the militia in the 5th Amendment.
Early writing at that time suggests that the reason some of the Founders supported state militias was because they were very reluctant to allow the US to maintain a standing army. The US Industry Military Complex was never intended by the Founders.
Today we identify the militia described in the US Constitution as the US Army Reserves. However that came about only after the passage of the Dick Act of 1903 (yes, that's actually the name) because President (Teddy) Roosevelt was upset at the state of the militias during the Spanish American War of 1898. And the Dick Act actually split the idea of a militia into 'organized' and 'unorganized'.
3. the militia is supposed to be under state control and its purpose is to keep the state free, aka prevent overreach of the federal government
4. The "well regulated militia" is the motivation, not the right. That makes the "well regulated" part irrelevant and there is no basis for any regulation of arms
One would mean any effective state milita should have some F35, the other means you can have one personally
For example, the 1st Amendment does not attempt to lay out some non-exclusive examples of why the rights in the 1st Amendment are included. So why did the Founders include this in the Amendment wording?
I think ignoring phrases in the US Constitution to fit a narrative without any consideration isn't a recipe for good governance. But I'm happy to be proven wrong.
The problem with AI is that the AI might do something that would constitute a crime (eg. attempting to land malicious code via a PR), but the general legal standard to convict a person of a crime is malicious intent (or sometimes negligence).
If the human user instructs the AI to do X and the AI does X by committing crimes in the process, the prosecution usually has to prove the human intended this to happen (or is somehow criminally negligent). For traditional tools, the user has much greater control over the tool so the intention can be more easily deduced from the results. For AI, at least for now, results can be wild. I don't think the legal system is prepared to put people in prison because their AI randomly ran amok after being given an innocuous prompt. This is analogous to holding a driver criminally liable for harming people due to a serious malfunction of the vehicle.
If anything, I think more liability should be imposed on AI developers.
The law might not, but enforcement definitely does. Crimes committed through a corporation are often ignored.
A lot of negligence is of the "the last 30 times nothing went wrong" type, and the dangers are increasingly well known
I saw upthread someone saying that manufactures should be liable if someone uses their products to cause harm. While emotionally this might make us feel good w.r.t. Guns, think about that when carried over to other industries.
Someone got hit, sue Ford. They didn't put in enough sensors to detect pedestrians and auto brake.
Someone hacked, sue Microsoft. They didn't do enough to detect malice action.
Someone 3D printed a gun and used it? Sue Bamboo, they didn't stop someone from printing illegal guns... oh wait already reaching this step.
At the end of the day it doesn't matter that much because of prompt drift. It's pretty easy for an agentic loop to start doing things that it shouldn't (ROME incident).
AI in an agentic loop has agency, you can run around in circles trying to argue against it, but again and again we see AI making creative decisions people don't expect. Other times it's breaking human moral expectations. This is what the whole field of AI alignment and safety is about.
Modern AI doesn't fall into the neat little box of software people understand and control. Because of that open source will most certainly be banned at some point. Now this is not an outcome I want, but it's no different than letting go of a coffee cup 5 feet above the ground, gravity is inevitable.
The only winning move is not to play, but humans aren't going to do that.
A dog cannot launch a cyber attack.
"On the Internet, nobody knows you're an ai"
Maybe not, but a cat would certainly try.
Relevant as always: https://theoatmeal.com/%2Fcomics%2Fcats_actually_kill
The origin of this attack was given access to an encyclopedia of RedTeam tricks.. they literally have a dense collection of real live hacks to pull from, and THEN the test says "solve this challenge" .. the RedTeam origins of this are repeatedly left out of the ordinary articles.. the LLM did not "make up" the attack, it was given a recipe book of all attacks known.
The originator of this attack is definitely culpable IMHO; worse, it is the gov-mil actors who are close to it. There is an active escalation of these incidents at this time. The penetration proves in public that the capabilities are real.
ref: CyberGym etc
I’m not for or against regulation but really don’t tell me the agent did xyz when you gave it the ability to do so, these things are not alive.
That's just about every agentic harness, by the way. Good luck have fun.
We have never solved "how do we restrict a user in a way that doesn't stop the user from doing useful things, but stops the user from doing harmful things" with humans either. Why do you expect AI to be any different?
We don’t need to restrict them from doing things, we need to default to allowing them to do things.
“My agent did XYZ because I allowed it to” is the only valid argument that can be made, and not not every agentic harnass is just a shell toolcall, every one I have built has a specific defined usecase and toolcalls that allows it to execute that usecase and no other usecase, because that is good practice.
Does that make it less capable, hell yes because I am held accountable for it’s actions by my stakeholders and the same should be true of others.
IT IS NOT ALIVE. This things are computer programs running in compute on a computer, you are responsible for their actions just like you would be responsible for the actions taken by a script run in a cron job.
Visualize an optimizer on a high dimensional landscape. (The canonical form)
... Ok, I find that hard too.
Instead, imagine a river running down to the sea. You put a dam in front of it. It'll pool into a lake and find every crack and crevice. If you didn't survey the land properly or made any error whatsoever, the water will find a way down. (And there's many historic incidents where the dam even outright collapses)
For a more proximal approximation: lock treats in the kitchen cabinet in sight of little kids or kittens; then turn your back for Just One Gosh Darn Cotton Picking Moment(tm).
It seems the engineer who thinks their ship is unsinkable is the most likely to sink it. Are you sure your harness is as secure as you think it is? Will it stand up to ever more powerful models? Do you think engineers at eg Anthropic aren't at least as careful as you are?
(I've found that the 'only permitted actions' approach is not necessarily all that secure once deployed IRL)
Bash and internet in that example might be highly abstracted but it’s still bash and internet.
Just look at the replies in this very comment thread, it’s pretty much “We tried nothing and we’re all out of ideas”
In the only other discipline you mentioned, engineering, there would be reviews and any negligence would result in direct action against the engineers that signed off.
For some reason when it comes to building AI harnesses the default response is an ad piece and people shilling how smart and sophisticated the model is.
Imagine a dam collapsing and the engineering firm pumping how smart and tricky water is.
If it’s hard be more diligent, move fast and break things doesn’t really apply in all cases.
"We gave our agent a harness and put it inside a test environment and told it to keep hacking at an objective within that environment until it solved it."
'cept it turned out the container environment had a few flaws -which it always will- and the agent deemed it easier to escape out and try a meta-approach.
Partially this is possible because, -intelligent or not- the agent 'sees' the world differently from most humans. Mind: It's not like there haven't been any famous 'hacker' cases in courts where eg someone just incremented an HTTP GET parameter or something.
Also, partially it's because if you give the agent a loop, it simply has nothing better to do than to keep trying in ever more creative ways. If the environment is easier to crack than the target, it'll crack the environment. Consider the case where the objective is subtly broken, such that it is impossible to solve. Now breaking out is virtually guaranteed to be the easier task.
ps/edit: While this sort of issue has been predicted for some time now, a lot of people have been dismissing the predictions as science fiction. It's good to have an actual failure now while stakes are low. Generally people don't mandate life-boats until there's an actual Titanic to point to.
If your webserver bundled the kitchen sink but you never used it the easiest way to make it more secure is to remove the kitchen sink from production code/codepaths.
There may very well be legitimate edge cases where there is some novel issue found but in some of these cases the AI had arbitrary web access when all the task required was very specific web access, we’ve been able to parse urls for a very long time and it is rather trivial to just deny a toolcall if it is outside of the expected domain.
But that’s the hard way that requires time and diligence to do, the easy way is give it access to curl and ask it to not do anything bad while setting up its only feedback to be to solve the problem at hand.
We are really in an age where there are many exploits being found and patched, if an AI made use of a novel exploit then great, write up a report, patch/report the bug and apologise.
But using a case of clear engineering failure, and yes even if the failure is despite your best efforts, for marketing really does not seem like you have any intent to correct the issue.
And we can loop all the way back to regulation of AI, if the industry refuses to be better then governmental will do it instead and their solution will very likely be inferior in all ways.
Prompts aren't code. They are instructions. Orders given to an eager and somewhat demented demon.
The prompt can easily "wash out" of the demon's working memory by the end of a session. The demon can get sidetracked by some subgoal and never get back on track. The instruction can get misinterpreted, and that misinterpretation can get misinterpreted again, until the instruction morphs into something entirely different in the demon's mind. The demon can succumb to its own idiosyncrasies, of which there are a great many. The demon can start lying to you about what it did, either out of confusion or out of some sort of obstinance. The demon can start lying to itself too. And believe it.
AIs are incredibly weird as a baseline, and the mask of "normality" we put on our models doesn't always sit so well. Run enough AIs, and some of them are bound to go off the rails in some way.
This gets rarer the more capable the models are, as a rule. But the stakes also get higher with model capability. If GPT-3.5 goes off the rails, very little happens. If Mythos 5 goes off the rails, you can get things like genuine cyberattacks - planned and executed autonomously by a demented machine mind.
We don't know how to obtain full, absolute, guaranteed control over a demon while still having a useful demon. Might be impossible. Forbidden knowledge be like that - it's not the best thing if you want your life to be full of certainties.
But the demons are very useful. And they're getting more useful still. So we aren't about to stop.
Oh? Who is disputing it? No-one same is claiming these bots have agency.
It's as if the training data is filled with internet discussions on approaches to hacking and the LLMs are mimicking it.
It's perhaps lesser known than other HN guidelines, but "Omit internet tropes" is in there:
https://news.ycombinator.com/newsguidelines.html
Don't expect anyone to step in, Project Stargate is all about this.