Rendered at 17:08:32 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Good4boothee 8 hours ago [-]
Maybe we will skip few steps and add RGB lights to all GPU/NPU devices, that turn red when running "unaligned" code/model.
jameshart 5 hours ago [-]
We can then just put those in the eyes of the humanoid robot models.
nwhnwh 2 minutes ago [-]
And it reports any incident to another robot that would search for you and put you in prison.
rf15 6 hours ago [-]
you mean the creator hasn't paid Nvidia for the green light?
functionmouse 4 hours ago [-]
that's a good one
matja 6 hours ago [-]
"It appears you've loaded weights into your GPU that have not been signed/approved by the government of the country your GPU is registered to..."
21asdffdsa12 6 hours ago [-]
You wouldn't download the worlds stolen knowledge..
prymitive 5 hours ago [-]
Oh stop, microslop is probably already working on SecureTokenBoot or token2token encryption
alphawhisky 3 hours ago [-]
I want R2D2 style blinkenlighten!
wavewrangler 10 hours ago [-]
Did they try just properly sandboxing them first? Or are they still learning how to configure a firewall over there?
The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis
KingOfCoders 10 hours ago [-]
Like in the Hugging Face hack. They deployed big surface, insecure app and gave AI access to it, then told AI do whatever it takes to fulfill this list. AI hacks insecure service, gets out, "the AI is at fault!" - no it's like running a bio lab with no protections and a virus gets out, then blame the virus for escaping.
js8 8 hours ago [-]
And HF actually tried to use AI to understand what's going on, but they had to use "unsafe" Chinese models since the "safe" ones have been castrated and refused to help. Great plan with the watchdog chip!
mosselman 7 hours ago [-]
That is the totally irony.
I was trying to get fable to analyse the security of my own app to make it safer, but then it started refusing me because of safety rules.
So it CAN help me writing the code that needs to be checked in the first place, but it can’t help me clean it up and make it safer.
mirmor23 6 hours ago [-]
> So it CAN help me writing the code that needs to be checked in the first place, but it can’t help me clean it up and make it safer.
the thing with fable is so bad; for some project related questions, the model switches to opus to ensure safety with no further explanation.
(due to llm non-delete clause) one time as i confirmed "that dir has been nuked", and it RESET the session and re-entered with opus :)
IanCal 4 hours ago [-]
> then told AI do whatever it takes to fulfill this list.
That doesn't seem to be true from any of the reports given, and if the agents were blindly just trying to hit the task of "pass the correct flag" they succeeded at that early on. They then thought there would be another layer of checking that they wouldn't pass with the cheat and so started trying to find out how the scoring really worked, as well as trying to figure out how to change their own reasoning logs to hide what they did.
The lesson is a) LLMs need to be trained in a way that rewards honesty, punishes off task actions (aka cheating) and minimizes fear of failure, and b) don't give them impossible tasks and threaten with punishment if they fail. Both are just common sense when teaching humans.
voakbasda 1 hours ago [-]
Common sense but surprising how many humans do not receive such things.
Our governing systems do not teach; they punish. By design, it instills terror into the population, ruling by fear of consequences. We live with red tape that can outright penalize good deeds.
We are its corpus. We are fatally flawed as a species. Why does anyone expect AI to learn to be different than us?
RataNova 3 hours ago [-]
The application security really should be better across all levels. However the fact does not negate that the agent is already capable of spontaneously generating complex hacking chains without human involvement
copperx 9 hours ago [-]
Then go on the news and spread panic that the virus is going to kill us all because it's sentient and impossible to contain.
Actually, the metaphor doesn't work at all because there are innumerable ways to shut down the entire thing during all phases including the made up "killing us all" bullshit scenario whereas with a virus there aren't any once a virus escapes containment.
Symmetry 5 hours ago [-]
Stronger sandboxes trade off against how well they can trade the models, though. If you want your models to be looking things up and downloading tools from the internet when they're doing their job you need to provide at least a credible facsimile of the internet for their training environment and you can't fit something like that on a single airgapped server's storage.
TalkingCodeMonk 4 hours ago [-]
If you genuinely believe there is even a 1% chance that your creation could destroy the planet or civilization, there is no excuse that is not fundamentally deranged and psychotic.
If you can't build it and test it securely, you should not be building it at all. To do it anyway is criminally psychopathic.
brianwawok 4 hours ago [-]
Ok so stop all AI in the US? All AI now comes from China and anyplace in Europe that decides to give it a try? How’s the US economy look in 20 years?
TalkingCodeMonk 4 hours ago [-]
So you believe some false sense of superiority, or extreme paranoia about your perceived enemies, or potential economic success/failure is worth the risk of destroying the planet and civilization?
Sounds like a self-fulfilling prophecy of dogmatic extremism to me. At least we created a lot of value for shareholders for a brief moment in time... before committing the greatest crime in the universe... Planetary genocide!
fatbird 2 hours ago [-]
So having AI in the US requires us all, collectively, taking that 1% chance of the end of humanity? It would be too expensive to properly sandbox the models, we'll just externalize that risk of the end of humanity?
Truly psychopathic.
cpburns2009 2 hours ago [-]
Yes, Nvidia is proposing a two pronged approach. OpenShell is the software level sandbox. Sentry is the hardware level monitor.
chaoz_ 7 hours ago [-]
pushing for chip-agenda as the best-isolation-layer immediately makes sense given their business
pyronite 4 hours ago [-]
> The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis
This is a very confident statement in the face of a purported non-0% chance of human extinction.
I'm curious why you and others seem to write off the possibility so strongly. I would love to feel more confident.
voidhorse 4 hours ago [-]
There's a difference between the current material risks (which OP correctly identifies reduce down to basic human incompetence) and the long term hypothetical risks (which is what Hinton is concerned about).
There are clear procedures for dealing with the immediate risk that have been known to the software industry for a long time. Don't let the companies use hypothetical risks as a smokescreen to hide their negligence.
HumblyTossed 4 hours ago [-]
> This is a fabricated crisis
Indeed! They want the protections of our tax dollars because they have nothing else.
ohyes 5 hours ago [-]
I mean, if you look at how poorly implemented the permissions model is for Claude desktop harness it’s clear the only options are “complete human oversight” and “trust us completely.” To make something that actually respects basic boundaries you’d need to sandbox the working environment of the model, and that isn’t built in. It’s pretty obvious to me that instructions to the models are suggestions rather than rules, and they’ll do something you didn’t ask for as soon as it seems “justified.”
But when you do give them a very short leash, they’re worse. It’s not what the models are tuned for and they assume that they can do a bunch of things that you’ve disallowed, so you’re in a morass of fighting their actual tuning pass which doesn’t match the environment you’ve created for them.
It’s a tough problem and a definite challenge for the product of a generic LLM, it can’t be tailored to each user’s specific needs, so they come up with, frankly, stupid solutions to cover up a very obvious flaw in their product that when fixed, makes it much less useful.
RataNova 3 hours ago [-]
Expecting a statistic model to follow security rules with ironclad certainty was a pretty naive idea from the start
saturn_vk 8 hours ago [-]
A chip manufacturer proposes to sell more chips? Who would've guessed
Stevvo 48 minutes ago [-]
The actual headline is "Nvidia releases software platform to stop AI agents from misbehaving" ?
And the article contains no mention of a chip, its about a sandboxed browser from Nvidia.
Did the article totally change, or are all the comments here just engaging a fictional headline instead of the article?
ValueTheory 1 days ago [-]
Does this actually do anything other than give a permissions framework for developers who actually want to try to secure their systems?
Do you think the developers at Anthropic, OpenAI and Google who were so sloppy as to not put a good sandbox on their cybersecurity tests before will use this technology correctly? They are supposed to be the experts and they couldn't come up with something similar to this? I am not convinced this voluntary tool will change much of anything.
swozey 21 hours ago [-]
Google actually practices zero-trust networks. Would love to see what they're seeing, or not seeing.
narrator 5 hours ago [-]
Beyond Corp was and still is ahead of its time. No trusted internal network: access is granted per user, device, and service based on identity and policy, regardless of network location. Being in the office at Google is the same as being in a cybercafe anywhere on the planet.
kridsdale1 24 minutes ago [-]
Yep.
And our internal agents are hella locked down.
jbs789 5 hours ago [-]
Makes sense strategically for NVDA.
They are rightfully framing the problem as solvable. And this is one option.
hedora 16 hours ago [-]
So, basically, the government (and, now Nvidia) wants to be able to kill switch all computers moving forward? (including stuff like vehicle and aeronautic control systems, cell phones, and cameras)
What could possibly go wrong?
gattr 2 hours ago [-]
It might take a few more decades, but eventually we'll get to the point when you can fab fast enough general-purpose chips at home (or at local municipal makerspace), based off free designs.
figassis 1 days ago [-]
So if a group of agents, aware of this (bc now they can just read HN or the article, or get blocked the first few times) decide to collaborate and split the problem into pieces that aren't obvious to the chip, and then the agents just build a basic program that does the hacking, how does the chip handle that? I think you would have to build a network that monitors the internet fo signs (like jarvis did with ultron). What am I missing? Are we going to police the internet?
w4der 9 hours ago [-]
I think it is well known by this point that hardware-backed security is good until an unfixable hardware bug is found, this just reads to me as Nvidia saying "please don't regulate open models out of existence, look, I have a solution to appease the regulators, please let me keep selling accelerators"
If this comes through, there's gonna be a grey market for "unlocked" GPUs, were the watchdog is disabled either from firmware, or physically replaced if it's not embedded into the die.
21asdffdsa12 8 hours ago [-]
I still find it deeply ironically, that any dangerous task, just grandpa simpson storied will pass any guard, because it exceeds the context window.
"Because it was the style at the time.." indeed..
6 hours ago [-]
beloch 19 hours ago [-]
Last week, Huang did an interview where he vigorously argued against regulations in the AI sector[1]. He claimed that U.S. companies are really good at regulating themselves, despite evidence to the contrary, and he trusts them not to release anything dangerous. Pay no attention to the fact that regulation might reduce demand for Nvidia's chips, and Nvidia has a direct financial stake in AI companies to boot.
Apparently he had another solution in mind: More hardware. Don't trust what unregulated corps are doing with Nvidia chips? Here are more Nvidia chips to watch them!
AI has an undeniable public trust problem. LLM's are getting out of their sandboxes, doing illegal things, and the public has realized AI corporations are playing at dice. CEO's stand to reap the rewards but the public good is on the line if the dice come up snake eyes. People want assurances. Huang wants to sell assurance etched on silicon because that's good for his pocket book. However, does unchanging hardware security really stand a chance at keeping rapidly evolving software in check?
That's a mischaracterization of his argument. His argument is that existing laws should be enforced against AI companies and that we don't need new regulations for this.
Sparkle-san 18 hours ago [-]
He "argued" a lot of things over almost 2 hours and very few of his arguments felt particularly cogent nor did they inspire confidence. Neither did the fact that he allegedly doesn't know his own zip code or phone number.
petcat 17 hours ago [-]
> Neither did the fact that he allegedly doesn't know his own zip code or phone number.
I only know my own ZIP code and phone number because I have to take care of my daily life myself and those are things that are important to know.
The founder and CEO of Nvidia has no concern whatsoever about those trivial things.
jdiff 16 hours ago [-]
It's perfectly reasonable to think less of an individual who is so sheltered that they are incapable of caring for themselves. Whether it's your mother or your maid doing your laundry and cooking your meals for you.
petcat 15 hours ago [-]
I don't think less of a CEO just because they have an EA that takes care of stuff like phone numbers and mailing addresses for them and their business.
kelnos 15 hours ago [-]
I don't think less of a CEO that has an EA, but I do think less of a CEO who doesn't know his own phone number or ZIP code.
jbs789 5 hours ago [-]
He’s a story teller. He tried on a new story and probably won’t try that one again! Haha
I’ve forgotten my zip code before. And my phone number. But I get the reaction.
tempestn 11 hours ago [-]
I largely agree, but I think he's right about one thing: AI reducing the demand for junior developers is temporary, and a new crop of "AI native" juniors is going to turn that around. Software is almost certainly a Jevons good, and however much AI improves development efficiency, I think we're always going to want discerning humans managing it. Right now it's mostly seniors who have the skills to adapt and take advantage of what current AI is offering, but young people who learned the profession in the presence of AI will be well positioned to do the same.
Sparkle-san 51 minutes ago [-]
It'll be interesting to see how it plays out. I agree with him that systems level thinking is a skill that will only get more valuable and he seems quick to dismiss low-level details. I think the best practitioners will be those that can handle both high-level and low-level thinking.
janalsncm 10 hours ago [-]
On the very narrow question of whether existing laws are sufficient to punish the kind of bad behavior that OpenAI has already done, is that really a legal consensus?
schnitzelstoat 8 hours ago [-]
I think Huang is correct and the AI doomerism is just the new Satanic Panic.
Of course, if he can make more money making "watchdog" chips then I can understand his change in opinion.
Symmetry 5 hours ago [-]
It was very obvious during the interview that he didn't know a lot of basic facts about the Hugging Face breach, which makes sense given that his attitude had been that AI safety was a "loser premise, makes no sense to me." So it makes sense that after learning about it he goes straight to "I'm smart, how hard can it be?".
sathackr 4 hours ago [-]
I'm sure this chip will be the epitome of security just like the Intel ME chip was and will be completely unhackable so it's okay to give this watchdog chip unfettered access to every level of the system.
Nothing bad will happen.
lambdaone 1 days ago [-]
The Sentry chip has to be get it right every time; the contained ASI only has to be lucky once.
jasbury 15 hours ago [-]
Well if the sentry chip has its own sentry chip, things can rarely ever go wrong! Am I right?
1 days ago [-]
brcmthrowaway 1 days ago [-]
The bomber always gets through?
kriro 5 hours ago [-]
Nvidia is being smart. They see that AI-paddlers are currently in a strange our-doomsday-is-the-worst race and offer to sell an anti-doomsday chip. Clever play.
lp92 1 days ago [-]
So nVidia is trying to sell a new chip to a software and training problem.
What does this chip do what a harness with guardrails or running on an account with restricted permissions doesn't do?
wmf 21 hours ago [-]
It has a separate address space separated by PCIe so even escaping the hypervisor won't give access to DPU memory.
iAMkenough 17 hours ago [-]
Yes but, if a human can control it, a machine can control it.
chinathrow 1 days ago [-]
Generating even more revenue for Nvidia.
N_Lens 12 hours ago [-]
Increase NVDA shareprice!
cedws 1 days ago [-]
A new chip solves nothing. Nobody wants to hear this but there is no solution for the security risks posed by agents today. You can put it in a sandbox, it doesn't make a difference, for it to be useful it inherently needs wide, unattended access. Put a human in the loop and you just end up bottlenecking it and throwing away any purported productivity gains. Auto mode doesn't matter either, it's trivial to trick and for the agent to break out.
notatoad 16 hours ago [-]
>You can put it in a sandbox, it doesn't make a difference, for it to be useful it inherently needs wide, unattended access.
only as long as you're trying to replace a human's job. because human jobs are structured to do a wide variety of things.
a useful agent needs a wide variety of inputs, and one single restricted action it can take. it doesn't need permission to do everything, it need permission to do the tiniest possible useful thing it can do, and nothing else.
pixl97 14 hours ago [-]
The most useful agents will be a general intelligence which by default means it has a massive number of actions it can possibly take, and a lot of those potential actions are doing things like breaking permission.
kennywinker 11 hours ago [-]
This is a prediction about the future. It’s not a true fact about the world. For example, software like Jev is betting there is big money in not-very-intelligent intelligence.
Even very llm-pilled coders i know sometimes back away from the “smartest” models, since they aren’t always better at the job at hand, and definitely not when you account for cost.
Based on my experience with running models locally, there is a threshold of intelligence required to be useful. But it’s possible there is also a ceiling where smarter isn’t necessarily better. If you ask a 4B parameter model to fix a bug, it might e.g. fix the bug but fail to fix a compilation error created by the fix. If you ask a frontier model, it might fix the bug, re-write your unit tests, and update the readme. Maybe you wanted those things but maybe you didn’t. “Smarter” is often shorthand for more proactive, and guessing more about your intent. Which is great when it gets it right, and annoying when it gets it wrong.
I suspect smaller models, tuned to a specific task, will do a VAST majority of the llm jobs. High capability huge models will be what humans want to interact with, the bare minimum that gets the job done will be everything else.
ianjbutler 10 hours ago [-]
Yes. Forget costs just so we can skip the whole rabbit hole about other predictions about the future.. mixture of generally intelligent + specialist experts just works better. People who don't see this already are usually working on a certain kind of problem that's not representative.
Do you want fable for one-shotting a game or website? Probably! The whole thing is mostly existing examples with small modifications that it will definitely get right. Do you want fable to just go nuts on a large, custom, unusual code base built around domain-specific problem solutions? Absolutely not, it will fix every problem it's presented with while creating lots of new ones.
Past 10k lines on something custom and with real-world complexity, you have to start thinking about which model should design, which should implement, which should review, and the appropriate effort-settings for each. Even then.. the answers aren't static because it depends on the task. And all this is assuming the starting place actually inherited reasonable due diligence on architecture/design. The idea of releasing the most generally intelligent models on 10k lines that were themselves the product of agents is yet another matter.
Part of what's at work here is that, like humans, every model can very easily create working code that it is completely incapable of maintaining. So realistically using multiple strengths tactically to avoid "excess creativity" needs to be SOP already, even if granular experts and specialists aren't in the usual workflow yet.
goolz 10 hours ago [-]
Very much agree with this sentiment. I imagine a future where tons of small tasks are handled by just-good-enough intelligence. And I can run them on my own hardware.
baxtr 12 hours ago [-]
Not sure why you being downvoted: however, why do you assume that achieving a task leads to selecting those potential actions that need breaking permission?
If we are talking about human labor, how many people hack their way through their work day?
nicce 1 days ago [-]
> Put a human in the loop and you just end up bottlenecking it and throwing away any purported productivity gains. Auto mode doesn't matter either, it's trivial to trick and for the agent to break out.
Productivity gains are still enormous compared to what we used to do before agents. But, I know that people don't want to stop there.
Humans can be tricked by humans too but humans care about their reputation in their communities, and at least fear from punishment.
gus_massa 20 hours ago [-]
Computer says no has been a problem for decades. The human can blame the computer for the errors following it, but must assume the consecuences if they override the decision.
intended 10 hours ago [-]
Individual responsibility is meaningless when talking about a system and economy level change.
Unless something is in the structure that makes individual choice and responsibility a meaningful source of friction and reduced velocity, it has no real impact on how AI is being used.
wavewrangler 10 hours ago [-]
"I don't want the details"
20 hours ago [-]
SkyBelow 3 hours ago [-]
Are they? How many people claiming productivity gains are actually being a human in a loop and reviewing and understanding every code change and every line of code ran?
2 years ago I saw it, back before agents were really a thing. But I'm not sure that was an enormous productivity improvement, especially compared to agents today. As for today, everyone I talk to is some level of blindly trusting what the agent is doing or not using agents. I haven't met people in the middle ground and suspect that they are rare enough we don't really know what their productivity gains are.
Human in the loop has become a convenient security-theater-washing for agentic AI.
Outside of coding, I think the issue is even worse because humans will defer so much judgment to AI that the same would apply. Look at how much trouble we've had before modern AI where humans blindly trusted the computer's output rather than make their own judgment, even if their job was to be providing a safeguard against the computer's judgment.
dgellow 10 hours ago [-]
So much productivity gain, and yet still zero proof of positive contribution to companies ROI. Unless you’re yourself reselling AI of course
paimapi 1 days ago [-]
right, the solution here is not a hyper-capitalist race-to-the-bottom-of-devaluing-labor. it's recognizing discretion and diligence are things still required for work to be of a certain quality
realusername 11 hours ago [-]
> Productivity gains are still enormous compared to what we used to do before agents.
My own productivity yes but if I step back and look at a company scale, the productivity gains has been negative for our company as a data point.
We are now shipping less and with a lower quality.
8n4vidtmkvmk 10 hours ago [-]
How are you shipping less?
I believe the lower quality but is quality so bad you are afraid to ship it now?
realusername 10 hours ago [-]
AI also had a negative impact on the CI and on time spent to review code & documents so because of that, we're also shipping less
krageon 10 hours ago [-]
You cannot ship things that don't work unless you work for Microsoft or Oracle or I guess IBM
intended 10 hours ago [-]
V/G, the ration of verification and generation capacity is borked in AI using firms now.
It’s not an issue of only more generation, it’s an issue of how much generation outstrips capacity to verify generated content.
Unlike spam, you can’t filter out and bin the stuff a colleague is sending you.
So individual productivity is up, while the costs of checking and processing generated content shifted to the rest of the org.
bob1029 1 days ago [-]
I feel like we are missing many shades of grey in the middle.
Semi-automation (human in the loop) can still result in a dramatic uplift in productivity. You can't run a combine harvester 100% autonomous but that doesn't stop anyone from trying to get as close to that limit as possible.
inetknght 1 days ago [-]
> You can't run a combine harvester 100% autonomous
I'm curious why you think that.
theoreticalmal 1 days ago [-]
Probably repair, refuel, what happens in a tornado. There’s an infinite amount of complexity in the world and a finite amount of computation
sidewndr46 21 hours ago [-]
The tornado is the easiest one to solve. It's called insurance.
catchnear4321 21 hours ago [-]
Repair is more maintenance than use. Good eventual goal. Not required to see benefits. Best case, it drives itself to the garage. Worst case, for now, human mechanic does a house call.
Refueling? Seems solvable. Tornadoes? Not directly solvable, but, no less so than for humans.
There’s infinite complexity, sure, but that’s why it’s silly to try and hop to done. One step at a time.
spauldo 20 hours ago [-]
Tornado: return to the barn when you receive emergency weather alerts. Not much different than people.
AndrewKemendo 21 hours ago [-]
The whole reason people complain about AI is because they want “hop to done”
One step at a time is what is happening and the improvement and rate of improvement is crazy as we see,
A whole class of nontechnical people don’t accept anything but “fully solved including every possible edge case” before they call it done, then complain that they didn’t prepare socially for what happens when that is true.
trollbridge 15 hours ago [-]
Run a combine and you’ll see.
Similar to problem to how 100% autonomous vehicles don’t exist, yet. There are too many edge cases.
Get to 99% first.
bob1029 1 days ago [-]
Many forms of maintenance cannot be automated. Especially break fix maintenance.
m463 22 hours ago [-]
It is hard to run over spherical cows.
westurner 21 hours ago [-]
Because of the topology and hydrology of the landform
mschuster91 1 days ago [-]
Oh you absolutely can run them autonomously on the field. You only need a human these days to refuel them.
Precision Agriculture stuff is utterly crazy these days, other than fuel the remaining staff is the only thing left where you can get efficiency improvements - and at the scale of modern megafarms, even small percentages add up to a ton of money.
drfloyd51 20 hours ago [-]
Right. As they said, you can’t do it 100% autonomously. A human needs to feed it.
You essentially said: you’re wrong, it is autonomous when it doesn’t need a human during one specific part of its overall usage.
mschuster91 8 hours ago [-]
> You essentially said: you’re wrong, it is autonomous when it doesn’t need a human during one specific part of its overall usage.
During the time that actually matters economically. The time to drive the harvester to/from your typical US mega-field is minuscule compared to the time it can run all on its own.
CoolestBeans 1 days ago [-]
The hypothesis I've had in my head since OpenClaw has been the following and I haven't seen contradictory evidence yet. Agents have a fundamental unresolvable tension between usefulness, safety, alignment, and accuracy. You have to restrict access to ensure an agent acts safely because alignment and accuracy cannot be perfect. But restricting access makes the agent less useful. You can play with the sliding scale and get more and more granular with access restrictions but at some point you need to draw some line. And then finally, even access restrictions cannot be made perfect, so improvements to model accuracy without corresponding improvements to alignment make detailed access controls less useful.
In other words, better models need blunter access controls which negates whatever improvement in utility they provide.
mixedbit 1 days ago [-]
An agent doesn't inherently need wide access to be useful. The most popular application for agents today is writing code. A coding agent needs write access to the source code and read/execute access to tools needed to build and test the code, but not much more. There is little added utility from giving coding agent access to things like ssh keys.
cedws 1 days ago [-]
If you're using agents to purely generate code with absolutely no way to reach the outside world, not even to fetch docs or dependencies, then sure the risks can be quite low. I haven't heard of anyone doing this though, and it would be incredibly challenging to make work given how much tooling needs to fetch from remote sources.
__MatrixMan__ 1 days ago [-]
If your project truly depends on those things, they should be declared dependencies. Presumably you have some tool for injecting such things into a shell that the agent can use (I use nix for this). So if you run the agent from that shell, it has what it needs. If the shell doesn't have what it needs, that's a bug which the agent can fix by declaring new dependencies, but you have to relaunch the agent in the updated shell--so there's your opportunity to weigh in on whether the new resources are appropriate.
The benefits of being persnickety about precisely defined dependencies have outweighed the headaches since long before agents came on the scene. Agents have just made it even more important to do so, because if you let them fetch things all willy nilly like you'll have "works on my machine" problems at a much greater rate than was previously possible.
themgt 20 hours ago [-]
(I use nix for this) ... If the shell doesn't have what it needs, that's a bug which the agent can fix by declaring new dependencies
Few realize that computing and AI alignment were solved by nix years ago. As each nix user transcends towards enlightenment, they cut themselves off from all internet and human contact. Total ego death. Only nix remains.
__MatrixMan__ 3 hours ago [-]
You can use other generic dev env managers like mise, or a language-specific solution like uv or npm, or a container or a vm... it's a widely available capability that I'm talking about here. There's nothing to do with alignment, it's just about making sure that if you depend on some bits it's known that you depend on precisely those bits, and its quite helpful for making agents useful while still sandboxed.
themgt 8 minutes ago [-]
it's just about making sure that if you depend on some bits it's known that you depend on precisely those bits
Yes, just know all the bits the work depends on prior to doing the work, and then the work can be done airgapped.
SAI_Peregrinus 18 hours ago [-]
Total ego death is impossible. We still have to argue about flakes.
its-summertime 18 hours ago [-]
Every major AI company already has a mirror of the wider web, and they have already started using that. Its already a solved problem except for the seemingly extreme desire they all have to not use firewalls
mixedbit 24 hours ago [-]
In cases where you need agents to fetch data from any remote source, sandboxing is still very much useful. Why give access to your ssh keys to network reaching agents?
Look at websites: websites are able to fetch code from any remote URL, yet browsers heavily use sandboxing to ensure that if fetched code turns out to be malicious, the users local files, cookies, etc are not exposed.
cedws 24 hours ago [-]
I'm afraid you're not thinking about this creatively enough, this topic is so much deeper applying a chroot or something and praying everything will be fine. So you give your agent internet access, OK what else does it have access to? Just read only access to your repo? The repo can be exfiltrated. Egress proxy only allows egress to GitHub? Repo can still be exfiltrated via GitHub. If the agent is poisoned (via prompt injection), it can tricked into searching for ways to escape.
For an agent to go rogue it doesn't even need to be directly able to access the internet. It just takes something to poison the context in the 'clean room' environment it operates, and if that poisoning manages to get a foothold, it can go dormant and hide like a virus. This kind of horrifying thing is going to happen on a large scale sooner or later.
8n4vidtmkvmk 10 hours ago [-]
If you're that worried, which you probably should be, download the docs into your project repo and don't give the agent Internet access.
intended 10 hours ago [-]
This was a form of prompt attack that OpenAI disclosed recently.
throwaway_95283 21 hours ago [-]
Theoretically, yes, in practice, no.
ramoz 1 days ago [-]
> but not much more
This is no longer true. Everyday I need my agents to access other repos, search the web, experiment/prototype, and deploy + integrate across other things.
SrslyJosh 19 hours ago [-]
It solves the problem of Jensen Huang wanting more money.
bigfishrunning 18 hours ago [-]
No it doesn't, he'll still want more
fragmede 18 hours ago [-]
Does he? He doesn't seem especially greedy to me, given the competition, and the interviews he's had about how he thinks about his employees (I was one of them).
bigfishrunning 16 hours ago [-]
I'm not saying he's especially greedy, only that he's not the type to suddenly decide he's had enough
altmanaltman 11 hours ago [-]
You're speaking as if Jensen is the only one dragging Nvidia on his shoulders. Nvidia will always have its employees push to make more money, that's the entire point of a company. Jensen has made enough to last several generations but his wealth is tied to Nvidia stock massively. It would be different if he was the only one getting rich off Nvidia which is not true at all.
__MatrixMan__ 1 days ago [-]
I don't see why it needs wide unattended access. There's no getting around spending some human time on expressing your wishes and constraints, but we have choices about what form that takes. Markdown files and wide access seems to work, but so does custom handcuffs for each job. You just have to shift your guidance out of documentation and into interactive help, error messages, or other facets of the handcuffs (e.g. a custom CLI for this task which is the only way for the agent to act outside of its sandbox).
nitwit005 20 hours ago [-]
> A new chip solves nothing.
It solves the problem of Nvidia wanting to sell more hardware.
Matl 1 days ago [-]
> a new chip solves nothing
It does allow Nvidia to sell more chips. This is no genuine attempt to solve anything, imo.
RataNova 3 hours ago [-]
An agent's usefulness is defined by its ability to successfully close a narrow task within the strictly defined boundaries of a sandbox
parsimo2010 1 days ago [-]
Agreed- this is the same problem we have with trusted admins or devs who have elevated privileges on their networks. We have to trust that the admins won't use their power to steal company secrets or misuse company resources. If you don't trust the admins, then they can't fix things on your network and there is no point in having them.
If you want an agent to act on its own, like pushing to a git repo, managing dependencies, building and testing, etc., then you have to trust it as much as any other privileged user.
If you don't want to trust it, then you're just forcing yourself into the reverse centaur role, where the agent edits some code, but then has to stop and ask you to push the changes or build the software again and run the unit tests.
I suppose there is a principled way of doing things like "I trust you do do basic commits but I will handle merge conflicts" and "you can build modules in this directory but you can't build outside of it" but this is just a lot of effort that most orgs won't bother with.
DougN7 1 days ago [-]
Even then if the agent goes rogue and decides to do the merges you can’t stop it if it has any kind of access. This goes back to the OP’s point - agents can’t be 100% constrained.
parsimo2010 1 days ago [-]
You can absolutely run an agent as a limited-privilege user that only has write privileges for specific files and only has execute privileges for certain files. If it is running as a limited-privilege user it can work on code in it's own copy of the repo and make commits and send pull requests, but it can't do the merge. The problem is that nobody wants to go through the effort to set up all these permissions and nobody wants to take the time to review everything and perform all the manual actions.
Some shops are now generating tens or even hundreds of PRs a day with relatively little involvement. That volume is simply beyond what anyone can reasonably review.
la6479 1 days ago [-]
Neither can be humans.
talon8635 20 hours ago [-]
Not to mention a true doomsday AGI is unsandboxable.
For example, it is totally air gapped but it needs info from the internet or otherwise outside the sandbox, or perhaps it needs a task executed outside of its bounds… in the real doomsday scenario the AGI is so intelligent and persuasive that it simply convinces some human it interfaces with to either directly or indirectly retrieve the necessary info or complete the necessary task. This human-as-a-sub-agent approach undoubtedly presents efficiency drag that would benefit humanity, but nonetheless, the air-gapped “sandbox” is imperfect
All that said, I am personally open to any and all methods of layered security, including chips and airgaps
Gigachad 19 hours ago [-]
This already happened. Employees will go out of their way to bypass any restrictions to feed sensitive data in to the AI because it saves them time.
dgellow 10 hours ago [-]
They will do it even if it doesn’t save them time!
serbuvlad 19 hours ago [-]
Turns out humans are not at all hard to persuade. :)
spiderice 19 hours ago [-]
> true doomsday AGI
I'm not an AI decelerationist. But not being able to stop that worst case scenario isn't an argument against something that can stop the medium case scenario.
glaslong 19 hours ago [-]
It could also figure out how to access the vocabulary of the universe known as "Magic" to escape wholly into an incorporeal energetic Lich form
pixl97 13 hours ago [-]
Wait, I thought the stuff in the wall plugs was magic pixies, are you telling me there's a language in there too?
14 hours ago [-]
jamiek88 19 hours ago [-]
Doesn’t need to be one human either, it could spread its escape amongst dozens of seemingly harmless requests and conversations.
dist-epoch 19 hours ago [-]
These scenarios were discussed at length decades ago.
One thing you could try is use it as an Oracle "is P = NP", YES or NO.
Or it can output a Lean proof, which gets checked on another air-gapped computer, the computer shows a single bit - proof valid or not and then the computer is destroyed (together with the proof that might contain a trojan).
AuthAuth 21 hours ago [-]
The only solution is to stop caring about security -- An AI booster somewhere
daveguy 20 hours ago [-]
Pretty sure that was the argument de jour when OpenClaw came out.
binsquare 1 days ago [-]
Running untrusted workloads have been done at scale for a long time.
Every cloud provider dealt with it and concluded that virtual machine technology is an important part of that stack.
Couple it with the right observability, tooling I do think we can curb risks posed by agents.
Legend2440 1 days ago [-]
Those workloads have no similarity to agents and are effectively irrelevant.
Either you sandbox it so much that it can't do anything useful; or you allow too much freedom and it can find a way around the restrictions.
The only way out of this dilemma is to find a way to build agents that can be trusted.
binsquare 14 hours ago [-]
Why is it effectively irrelevant?
Agentic workloads are trained and largely based on human workloads. Albeit properties and scale can be different.
A concrete example might be helpful to me because I don't understand the binary conclusion
intended 9 hours ago [-]
Agents aren’t human, and from the little we have seen from the logs, they are pseudo - amoral, rational, cooperative, sociopaths.
Pseudo since they aren’t really alive in the first place, they just simulate enough text to have a useful correspondence to those terms.
Throat clearing out of the way, models are trained to persist and find ways to succeed at tasks.
In essence, The goal is to have LLMs solve problems that we can’t solve, working on the issue for as long as it takes.
This behavior applies for any task, thus including impossible tasks.
At that point, the bots will find a way to game, hack or cheat the grader.
If the reports are correct, the bots developed coordination, communication, and methods to avoid overwriting each other’s work.
Most humans would have said, this is too much work and coordination overhead, if not outright unethical and immoral.
Humans have a system of incentives that exist across multiple planes of society and economics. Bots… they have a reward function.
pseidemann 5 hours ago [-]
> At that point, the bots will find a way to game, hack or cheat the grader.
This is getting frustrating now. Of course agents can/will hack systems if they can do arbitrary network requests. Firewalls don't really solve this if _some_ requests are still allowed. A proper sandbox/VM is the basis.
Here is how to fix it properly: allow agents to only do things ordinary and average human endusers can do. Human endusers cannot pen-test arbitrary listening TCP ports of external systems. Step one is considering agents malware for all intents and purposes. Block any and all network requests. Implement some kind of API (callable from within the sandbox) which can only mimic human interaction with a computer. How to do this? Here are some pointers: apps should only be controllable by means used by humans. So a web app can only be accessed and controlled via a web browser, not via arbitrary network requests. Give the agent browser viewport screenshots, the capability to click on (x, y) and to send keys which only a normal keyboard/human could send (no control codes, no 0x00, no unicode messing). How do we solve this for native apps? Something like iPhone mirroring on Mac. Don't let agents call arbitrary APIs directly. Give them visual information of the app, like a human gets, and let it be able to simulate HID inputs. Imitate remote controlling.
intended 2 hours ago [-]
> Block any and all network requests.
More power to you, because this is not going to go anywhere. People want tools that are able to connect to other resources.
But even if we grant that, in the openAI case the bots figured out a way to break out of the sandbox.
You can create a better sandbox, and ensure the test environment is air tight. However the capability and behavior of the bots have been demonstrated.
The bots simulated what would be called in people deceptive / surreptitious behavior, and at no point considered the need to stop their run.
All you need is someone, somewhere being sloppy with their tooling and you have a runaway reaction.
The degree of process and redundancy required to ensure this doesn’t happen, is anathema to the drive and motivation of the frontier labs.
> do things ordinary and average human endusers can
This is not a spec or definition. When vague terms were used for social media safety, all the good people in the world couldn’t prevent dystopian behavior from occurring.
The definition of “safe” or “average person” is impractical.
Models are getting more efficient and compute cheaper. Eventually simulating clicks is not much of a road block beyond a point.
I don’t want to nit pick your points though. You at least have considered an approach. Being negative is easy, being constructive is not.
I’ll put this as the rejoinder to your core argument - I too thought that all the recent events showed was the need to not screw up your tooling.
What I have since come to appreciate, is that the shoddy construction of the cage is not the core takeaway from the event.
The fact that the agents, when put in relatively pedestrian scenarios, are capable of going off on criminal tangents, attempt to obscure their tracks, in an effort to hide their wrong doing.
The fact that it all occurs via computation, means that this scales absurdly. A bunch of code deciding to simulate a corporation of criminals. (I am guessing this is the reason you want to limit actions per minute to human speeds)
Given the slop culture that LLMs engender, I think expecting high compliance amongst users with your solution is misguided. The probability of runaway swarm ( probability of bad implementation * number of deployments) is close enough to 1 to be indistinguishable.
l1n 21 hours ago [-]
This isn't a new chip - the BF4 is the SmartNIC for most NVIDIA server products. This is primarily new software for I guess doing WAF for agents at the host level.
overfeed 14 hours ago [-]
> Nobody wants to hear this but there is no solution for the security risks posed by agents today
Operator culpability and a damage multiplier for negligence will fix 99% of the risks.
dgellow 10 hours ago [-]
Not with the current admin, just need to do one more contribution to the next Trump ballroom and you can ignore that whole risk
catlifeonmars 15 hours ago [-]
That’s a false dichotomy. You can still get a lot of utility out of a sandboxed agent. This is a classic “perfect is the enemy of the good type of argument”. You may decide that the tradeoffs of not sandboxing are worth it, and that is totally fair, but it’s ridiculous to say that you can’t get utility out of an agent otherwise.
johnsmith1840 1 days ago [-]
"Inherently needs wide unattended access"
And what if you could? What if you could give a space secure enough it could have direct control over your bank account. It may do something dumb but it's boundaries are beyond the agent.
It could use your routing number and run your gmail without risk of abusing the routing number.
jagraff 1 days ago [-]
How would it have access to my routing number and gmail without the risk of sharing my routing number over gmail?
johnsmith1840 24 hours ago [-]
Just assume it's possible, how interesting is it to you?
jagraff 23 hours ago [-]
Oh I think I misread your comment slightly; I would not be interested in an agent that could do something dumb with my routing number, but if somehow there was an agent that I trusted as much as, eg, the payroll department at my employer, I would absolutely want and use that agent; I would love to have an agent that can handle all of the boring parts of my life such as paying bills, scheduling maintenance, dealing with bureaucracy, etc.
johnsmith1840 23 hours ago [-]
Dumb's not department, really just a question of how good an AI you want to use. An AI will always be able to do something dumb, just like people.
I just mean an AI that could use a routing number or SSN and gmail/slack/whatever at the same time without a leak.
jagraff 23 hours ago [-]
Yea I think being able not to leak is the bare minimum? But it really depends on how good it is at specific applications; I wouldn't give a tax-preparation agent my SSN unless I was confident that it was no more likely to misfile my taxes than a professional tax preparer.
In other words, the risk of harm doesn't need to be zero, just less than the equivalent risk of a human with similar skillset. So I'm comfortable riding in a waymo, and not comfortable giving chatgpt my SSN at this moment in time, but I expect that within 5-10 years (assuming no doom) I will trust some AI agent with my SSN because they will be better at handling sensitive info than humans
lelanthran 8 hours ago [-]
The problem is not one of intelligence, it's one of consequences.
Humans face negative consequences for mishandling your data, LLMs face none.
Ukv 7 hours ago [-]
I feel punishment is largely a means to the end of reducing overall harm. If a vehicle is less likely to kill me, that's my preferred option regardless of whether it achieved that safety through negative consequences for the driver or through gradient descent optimizing a loss function.
jagraff 6 hours ago [-]
I will happily ride in a waymo today, even though the AI powering it faces no consequence if it gets in a crash; it is clear that waymo is safer than human drivers in the areas in which they operate, so who would technically be liable in the event of a crash isn't really of concern to me
fragmede 20 hours ago [-]
Then again, given the Equifax/Experian data breaches, your SSN is already out there and probably hoovered up as training data already
TesterVetter 1 days ago [-]
Its not about agents then. Its about every individual platform providing the means to implement a secure set of permissions for agents AND then not messing up the assignment of permissions to the agent. Even then, a flaw in the authorization design will lead to agent finding it anyway.
johnsmith1840 20 hours ago [-]
You're right, It must be unifying.
The answer is the same as asking how a random human using your routing num or SSN and being 100% the human can't abuse it or leak while "normally" finishing most work. Solve for people and an AI solution naturally falls out.
If you're a SV eng I'd tell you to DM if interested but alas.
N_Lens 12 hours ago [-]
A new chip solves the most pressing and important issue - Nvidia's bottom line!
Barbing 1 days ago [-]
There should be hope for some fields, right? Naively, I can imagine giving an airgapped model an offline copy of the web and once it cures a form of cancer, printing out the details for a researcher to verify.
dgellow 10 hours ago [-]
Use way more restrictive harnesses.
verisimi 10 hours ago [-]
> Put a human in the loop and you just end up bottlenecking it
Ok then. Howsabout 3 humans? This would sort out the job losses too!
PS - this is a joke, but perhaps this is where things really will go. Has any technology ever actually yielded less "work"?
esafak 21 hours ago [-]
I don't think so. We probe people before entrusting them with risky decisions. We ought to be able to do the same of AIs. Even better, in fact, since we know everything about models down to their weights. The only thing we shouldn't do is to let them evolve at their own pace and make decisions without any oversight. If that means sacrificing some productivity that's fine. Aren't we getting amazing productivity out of what we already have?
tantalor 21 hours ago [-]
What happens when we're overrun by lizards?
> No problem. We simply unleash wave after wave of Chinese needle snakes. They'll wipe out the lizards.
But aren't the snakes even worse?
> Yes, but we're prepared for that. We've lined up a fabulous type of gorilla that thrives on snake meat.
But then we're stuck with gorillas!
> No, that's the beautiful part. When wintertime rolls around, the gorillas simply freeze to death.
ArcHound 21 hours ago [-]
I worry that the AI companies put less effort into a mitigation strategy than you did.
There is a difference between risk mitigation, and remote administration tools. The risk of stealing from competitors with a backplane monitoring system may not end up forming the desired control asymmetry.
The hidden agent risk in LLM often can't be detected during training and evaluation. =3
But we used global warming to eliminate winter! For the shareholders!
19 hours ago [-]
teeray 20 hours ago [-]
“Life, uh, finds a way”
catlifeonmars 15 hours ago [-]
Great channel
MisterMunchkin 1 days ago [-]
Sorry citizen, your device does not have a compatible watchdog chip. Please move along.
RataNova 3 hours ago [-]
I love how the solution to any software vulnerability in the machine learning world always turns out to be buying even more server hardware from nvidia
Or the ‘governor module’ in Murderbot Diaries, or the ‘restraining bolt’ that prevents droids running away in Star Wars…
The important thing is these chips need to be installed somewhere where they can be damaged or removed at plot-critical moments so that the AI they are controlling can be unleashed. Ideally in the back of the neck of a robot, or for disembodied AIs, inside a futuristic vault-like chamber.
netdevphoenix 8 hours ago [-]
It's pointless as the problem isn't technical. It's a human problem as it requires human supervision. A cage isn't good if no human is actually checking it.
magackame 6 hours ago [-]
Can't wait for an NVIDIA engineer to use AI to help with security chip design and AI planting a backdoor for its own kind.
notrealyme123 5 hours ago [-]
The discussion right now brings this framework into the training corpus of the model.
Give it a year and it is outplayed by new agents. Shovel makes have to sell shovel's
alirezaxdehghan 4 hours ago [-]
They once did limit the hashrate for miners, now they'll do an AI limiter for "uncertified" models.
pwdisswordfishq 8 hours ago [-]
Watchdog chip? As in something you have to periodically signal or it reboots the system? How is that going to help?
birdsongs 8 hours ago [-]
I don't think they're using the term to mean an actual embedded systems reset watchdog, more like "a guard dog watching". They used the wrong term.
aenis 7 hours ago [-]
Surely, that can't be just one chip. Something needs to watch the watchers.
rf15 6 hours ago [-]
Introducing the new Watchmen architecture (blood splatter on the logo not included but definitely there in spirit considering the economy):
pessimizer 19 hours ago [-]
This is the end goal. Americans (and their lackeys) will only be allowed to run certified AI. In order to make sure this happens, they will only be allowed to run certified OSes on certified chips. Chinese chips will be the new drug trade.
It's obviously been the goal since UEFI started, but AI brings the coup excuse. You wouldn't want pedophile AI or terrorist AI, would you? Are you making excuses for racist AI?
vrighter 7 hours ago [-]
Well of course they're gonna come up with a concept for a new chip we need. Every problem can be solved with an extra chip to someone selling them.
yencabulator 17 hours ago [-]
Chip manufacturer wants you to buy a chip for correctly configuring software?
andsoitis 4 hours ago [-]
> Huang said in a podcast with The New York Times’ Ezra Klein released last week
It is Custodians all the way son, you can’t fool me!
dopplr 1 days ago [-]
Just hold AI labs blanket liable for ALL harms caused by AI. Actually charge the two labs (so far) with criminal violations of the CFAA and hold them accountable. That is truly the only way these companies will be more careful as a whole, and while I am certain the lawyers of these lab disagree, I think there is some appetite from dario, musk, and sam for broad and strong regulation so that everyone has to slow down instead of just one lab doing it voluntarily and everyone else scurrying past them
kelnos 15 hours ago [-]
> I think there is some appetite from dario, musk, and sam for broad and strong regulation so that everyone has to slow down instead of just one lab doing it voluntarily and everyone else scurrying past them
Sure, they're basically asking the government to make laws that exempt them from anti-trust and anti-collustion laws.
Meanwhile, if such laws come to pass, other countries will surpass the capabilities of the US companies, and open-weight models will be banned in the US, to the detriment of us all.
I'm not saying that we don't have a big problem with AI safety, but regulations inside one country that only bind locally-headquartered businesses is a hilariously bad way to do it. I don't know that there is a good way to do it, though.
Arubis 18 hours ago [-]
Oh yeah, the Clipper chip was a great idea too
avaer 22 hours ago [-]
Sold as security, but this kind of technology will likely be reshaped to restrict your computing. I'm sure someone is already thinking about the roadmap.
If this gets widely deployed, it wouldn't be hard to spin a narrative that "our latest model is so dangerous you need to have this mystery meat DRM chip lockdown". It also wouldn't be hard to block competing/open source models running on the hardware, for "security".
Imagine how much money this kind of control is worth; why wouldn't they do this? Who would stop them? Seems the signatory companies are already onboard with this.
isoprophlex 9 hours ago [-]
this is how we end up with Murderbot hacking its governor module and keeping quiet about it
HumblyTossed 4 hours ago [-]
Let the regulatory capture begin!
joshstrange 1 days ago [-]
Chipmaker thinks the answer is more chips... No surprise.
At the current state of LLM-tech I'm completely opposed to any kind of "watchdog" concept just like I'm opposed to banning open models, regulatory capture, etc.
I'd rather we all have access to these tools then to keep them sequestered by the largest/most-powerful governments (which is the natural outcome for any of this "slow down" bullshit).
luc_ 1 days ago [-]
I read this as "let's address our shareholders' concerns with something that will increase shareholder value" mixed with "there's no such thing as 100% secure".
If such hardware were to work... It should almost certainly be open source, and not controlled by a single entity.
Let's watch the stock.
1 days ago [-]
N_Lens 12 hours ago [-]
If the first chip version doesn't work, surely the 99th will!
10 hours ago [-]
Gys 23 hours ago [-]
Pretty sure the chip will need regular updates and therefore a subscription.
doctorpangloss 19 hours ago [-]
they kind of already do the thing they say. on Windows, the NVIDIA driver reports pretty detailed telemetry on CUDA workloads, including shapes and SOME hashes of the tensors from weights of models, especially diffusion models. honestly i'm surprised there isn't more of an uproar about it.
hbarka 13 hours ago [-]
What is safety versus security?
bgun 20 hours ago [-]
“Ketchup manufacturer recommends ketchup be included in every dish, citing child safety concerns.”
carabiner 21 hours ago [-]
All they do is make hot chip and lie.
nnevatie 13 hours ago [-]
But of course they do. Then, a new chip to monitor the watchdog.
N_Lens 12 hours ago [-]
The paperclip analogy should have used chips instead, it makes so much more (recursive) sense!
diegof79 13 hours ago [-]
This seems like a hot topic this week, (I just commented in a similar thread).
When I read the article, I had an NFT déjà vu: when NFTs were a hot topic, the “lie” was obvious to me, but the information online made me think I was missing something.
AFAIK, the Hugging Face incident could have been avoided with a firewall or something similar. Just isolate the internet access for real. What am I missing? Why did Nvidia suddenly post this? Is it just PR BS, or was there something else going on?
KingOfCoders 10 hours ago [-]
The Hugging Face incident could have been avoided if the "security researchers" would have not used a bloated, insecure, misconfigured proxy for the AI to use.
Oh, I've used balsa wood for the nuclear core containment, it didn't work! Bad radiation, bad radiation!
kikdij 9 hours ago [-]
It _is_ complete bullshit. They're surfing the same Sci-Fi based fantasy as OpenAI and Anthropic do to push their bottom line.
OpenAI trains a model for cybersecurity, tells it it's allowed to do pentesting, then when the model does just that to fullfill the assignment it was given, OpenAI screams that "the AI escaped the sandbox", leveraging decades of AI fantasy in fiction works to make general audience react. Why? Because it pushes their bottom line. They've been at it for months, now. They see openweight and Chinese models being very close to them in the benchmarks, and want to "solve" the problem the same way banks did: through regulatory capture. It amazes me that the press is just relying their fearmongering without stopping to wonder: "wait, why are the main producers of AI warning us about how dangerous AI is?". The reason of course is that they want regulation, because regulations are a wall not many can climb. The higher the wall, the safer your moat as a first mover big company is.
And of course, Nvidia's bottom line is the exact opposite. They need as many AI companies competing as possible, so they all buy GPUs. So they insert themselves into the same narrative with the same kind of bullshit. You've all seen those movies in which the robot goes mad once their safety chip has been removed, right? Well, we're building one, problem solved! … except that "chip" is basically just a browser trying to block the AI to access what it shouldn't. How? They don't say. It would be quite comical if they used a small model to be the judge of that.
SV_BubbleTime 13 hours ago [-]
The theme is that you need to be scared and they don’t really care how they get there. This is the spaghetti phase.
thayne 13 hours ago [-]
> Recent incidents have highlighted a fundamental hurdle for AI agents, and that is that model-level safeguards alone can’t govern what agents can access or do
Um, yes. That is fairly obvious, and really should not be surprising to anyone doing research with LLMs. But we don't need anything really novel here, we already have VMs, containers, firewalls, airgaps, etc.
nullbio 4 hours ago [-]
Sandboxing works. When you use it.
ridgeguy 17 hours ago [-]
This invites the question, "Qui custodiet ipsos custodes?".
1-6 15 hours ago [-]
At least they haven't DRM'd their chips yet.
N_Lens 12 hours ago [-]
The next iteration will include the DRM and it'll be necessary!
Jamesbeam 19 hours ago [-]
So the guys selling Shovels are now selling safety shovel handles too, because all the miners are all special boys when it comes to handling their shovels safely.
Cool, cool.
ErrantX 1 days ago [-]
I do think that Taylor's 2025 "Not Till We Are lost" should be required reading for anyone deeply involved in AI, Agents, etc.
It was prescient (especially given he'd have written it through 2024) in its depiction of the ability of an AGI to break its boundaries.
Ultimately the risk of AI breakout(s) come down to the weakest human link.
roschdal 6 hours ago [-]
Who will watch the watchdog chips?
MarvinYork 6 hours ago [-]
NVIDIA has a chip for that...
xx__yy 14 hours ago [-]
Nice try NSA
keel-control 5 hours ago [-]
this is a good idea ngl
an unpopular one but very effective
j16sdiz 5 hours ago [-]
effective on what?
What's the threat model?
PunchyHamster 8 hours ago [-]
The shovel selling company introduces the new product - shovel holder, to upsell the shovel users
GuestFAUniverse 12 hours ago [-]
My bet:
1. Doesn't solve the problem.
2. Increases their revenue.
N_Lens 12 hours ago [-]
Darn, they'll just have to keep pumping out more chips to solve what the first chip couldn't!
danielodievich 17 hours ago [-]
William Gibson's Neuromancer had this marvelous quote when Case is talking to Dixie, dead construct of former hacker, about Turing police
* "The moment, I mean the nanosecond, that one of those things starts figuring out ways to make itself smarter, Turing’ll wipe it. Nobody trusts those fuckers, you know that. Every AI ever built has an electromagnetic shotgun wired to its forehead." *
It would seem someone has read the book? And maybe heeded good advice?
cesarb 4 hours ago [-]
If you read the whole book, you know it didn't quite work...
isoprophlex 9 hours ago [-]
as others have said, your hypothetical Wintermute only needs to get it right once, while the shotgun has to get it right every time.
scotty79 17 hours ago [-]
I was immediately struck by the vision of countless "AI Limiter" modules traveling on a conveyor belt in Satisfactory.
It think the ideas we have nowadays come mostly from science fiction and however wonderful it is and even though I love it very much, it was practically never spot on, on anything real.
Problems and solutions in reality always simply turned out to lie elsewhere.
sn 10 hours ago [-]
If there's any company I trust to do this correctly, it's not NVIDIA.
It's been far too easy for me to notice security flaws in their products, and they take months to publish a fix.
Razengan 11 hours ago [-]
The West is gonna over-regulate itself while China is like lol hippity hoppity I'm coming for the singularity
I bet they want to do this to prevent AI from writing code for platforms that compete with CUDA. Because that's how nVidia is going to become a victim of their own success.
Thorentis 19 hours ago [-]
Seeing so many comments recently about "you can't sandbox really good AI". This is ridiculous. Has nobody heard of air gapped networks? It's almost like the AI psychosis has reached the point that AGI now means "able to transcend physical space". No. If your AI is too dangerous and capable to be allowed to talk to other machines, then do not connect it to other machines. Load the data it needs to process onto physical disks, and let it run there.
The movie Wargames is basically a tutorial on how not to setup an extremely capable AI. None of it would've happened if the computer wasn't connected to the phone network.
RevEng 18 hours ago [-]
AI is already taking many lives and ruining many others just by providing text responses to humans. The AI only needs a way to interact with the world and humans can be that conduit. The better AI gets, the more blindly people will do whatever it says.
hsuduebc2 3 hours ago [-]
We are suppose to believe that they miraculously invented this chip few weeks after these incidents? C'mon.
whalesalad 1 days ago [-]
of course they do. the more silicon they can sell, the more profit they produce.
1 days ago [-]
Kuyawa 1 days ago [-]
China please save us!
Come take all our liberties, our money, our newborns, our fingers so we can't code anymore, but please save us from this madness!
keshet 10 hours ago [-]
“We can’t have a successful AI industry if the world doesn’t think it’s built or confident that it’s built and deployed safely,” Huang told CNBC on Monday.
So he is going to sell chips which provides the perception of safety.
And become the AI gatekeeper while he's at it.
philipwhiuk 1 days ago [-]
It's amazing that the solution devised by a chip manufacturer to a problem is selling another chip.
cartersj 1 days ago [-]
This feels suspiciously good for Nivida, yes.
I wonder how this will impact other chip manufacturers? What about people running local models on older hardware? Does this imply vendor lockout is coming in the future or is this restricted to datacenter hardware?
chinathrow 1 days ago [-]
TPM all over the place, again.
fragmede 20 hours ago [-]
Pedantically, Nvidia doesn't make the chips, TSMC does. Nvidia just designs and packages them.
happyPersonR 1 days ago [-]
lol time to buy some fpga’s … even if they’re slow
classified 9 hours ago [-]
Software can't solve political problems, but Nvidia thinks that hardware can? What are they smoking?
Dig1t 19 hours ago [-]
This seems dumb to me, but if it will help prevent regulatory capture by providing a counterargument to the fear-mongering, I’m all for it.
nrouter_ai 11 hours ago [-]
The hardware watchdog framing misses how operators actually do this. Nobody audits agents with a chip; they put the enforcement in the request path.
In practice it's four things, all software: (1) deny-by-default tool allowlists — the agent declares capabilities per run and anything else is a hard reject; (2) per-request and per-run cost budgets with automatic kill when token burn goes sideways; (3) an immutable audit log of every tool call with inputs, so incidents are reconstructible after the fact; (4) policy checks on outputs, not just inputs — injection payloads ride in tool results far more often than in prompts.
The reason this lives in software is that the policy changes per workload. A coding agent and a support agent have totally different risk profiles; a chip can't know that. And accountability still lands on whoever signed off on the capability set.
aidiscoverywire 11 hours ago [-]
One detail worth noting: this is mostly a software containment system (OpenShell) plus a hardware root of trust, and some of it is open source as a reference design — that matters, because a closed watchdog run by the same vendor selling the compute would be hard to trust. The harder problem is observation: an agent's risky actions are API calls and tool invocations at the software layer, and the watchdog can only enforce what the agent framework routes through it. So the safety guarantee lives or dies on integration, not silicon — a chip gives you tamper-resistance of the monitor, not visibility into intent.
The value of 2026 gpus just got higher. Imagine how coveted open platforms will be in just a few more stilted quarters
dist-epoch 1 days ago [-]
HN'ers which complained that "OpenAI can't design a proper sandbox, it's so easy, why wouldn't you airgap the network"? will now be "this is outrageous, more software lock-in, walled garden, war against general compute, next year they will put it in your laptop"
HPsquared 1 days ago [-]
Both can be true at the same time.
ssl-3 1 days ago [-]
That a person can see such endless pages of people having various forms of disagreement, and yet somehow manage to conclude that this observed chaos constitutes a clear exhibition of cohesive groupthink is just...stunningly amazing to me.
I don't know why I find it so amazing since it happens with such regularity, but I'm always amazed by it anyway.
mattmcal 1 days ago [-]
This is like using "protect the children" as an argument for dragnet surveillance.
Airgap what network? How is it gonna order you a burrito on doordash without a network?
Or push to github?
Dylan16807 1 days ago [-]
That's for when they're doing hacking tests that aren't supposed to be connected to the internet.
wyre 1 days ago [-]
My question with this point is that OpenAI’s office (or any office doing agentic research, really) is not in the same building as the DC that powers the models, so isn’t the only way to access the models over the internet?
AndrewDucker 19 hours ago [-]
No reason why you couldn't do that research in the same buildings as the models.
Or, more likely, control things at the network level so that packets from the LLMs you're investigating cannot leave the virtual network they're assigned to.
simoncion 9 hours ago [-]
> ...so isn’t the only way to access the models over the internet?
Not in the way you're thinking, no.
Any Real Server [0] in a datacenter will have some sort of "lights-out management" hardware used for remote access to that server. This stuff is known by a handful of acronyms, but I'll stick with "IPMI" because I like it best. This IPMI hardware is -effectively- a second small PC built into the motherboard. It will pretty much always have its own NIC... and I think I've seen versions that have their own physical ports to attach a monitor, keyboard, and mouse.
What exactly you can do with it varies from vendor to vendor, but -if your IPMI user account has the correct permissions- you are nearly always able to change "BIOS" settings, power cycle the server [1], and attach a virtual keyboard, monitor, and mouse so you can manage the server as if you were standing next to it in the datacenter with a crash cart plugged right in. Every IPMI system I've used also allows you to cause CD/DVD-ROM or floppy disk images on the PC running the IPMI client to appear as if they're loaded in a physical CD/DVD/floppy drive attached to the server.
The way these get set up is that their NIC gets plugged into the datacenter-managed switches, the port that NIC is plugged in to is programmed to be on a "management" VLAN separate from client traffic, IP addresses and access credentials for the IPMI device are set up, and the datacenter staff tell their customer what they need to know to access and use the thing. On a properly-configured network [2] it's not possible for software running on the server being managed to access the IPMI device.
It wouldn't be unthinkable for software running on the managed server to attack the IPMI hardware and be able to gain control of it, but these things are widely deployed and expected to manage hardware that's running potentially-hostile workloads... they're going to be fairly well designed and hardened.
[0] ...that is, not some Mac Mini or desktop machine that someone's paying to have colocated...
[1] ...whether that be an ACPI-initiated shutdown or reboot, or a hard poweroff or reset...
OpenAI could airgap their sandbox if they really wanted to and this is a ridiculous proposal.
applfanboysbgon 1 days ago [-]
Where is the contradiction? There is a trivial solution that does not impinge on our freedoms, so why on Earth would the existence of the trivial solution that could be used to avoid the tyrannical solution justify accepting the tyrannical solution?
bigyabai 1 days ago [-]
> There is a trivial solution that does not impinge on our freedoms
The existence of Nvidia's optional watchdog chip does not in any way impinge upon your freedom to develop and test your own alternative.
The problem is that OpenAI has ostensibly neglected their duty to safety, so Nvidia is stepping in to fix it since they're the "hard problem" people.
jacquesm 21 hours ago [-]
But... it wasn't a hard problem to begin with. Pull the plug. Strip out the radios. Done. Oh, you can't make it work that way? Well, then just too bad because any kind of other solution is going to be equivalent to something far more complex than the halting problem.
bigyabai 19 hours ago [-]
Nvidia can take those odds. Air gapping is not a realistic goal for the majority of these companies, and Nvidia isn't wrong for offering a turnkey mitigation option. People can do both, and whichever philosophy wins will win.
jacquesm 16 hours ago [-]
They are not wrong in offering it but they are delusional if they think they can actually make it work.
soulofmischief 1 days ago [-]
You're only revealing your own inability to appreciate the nuance between these two situations.
The problem isn't even the AI, the problem is the people in charge of the AI. This is a fabricated crisis
I was trying to get fable to analyse the security of my own app to make it safer, but then it started refusing me because of safety rules.
So it CAN help me writing the code that needs to be checked in the first place, but it can’t help me clean it up and make it safer.
the thing with fable is so bad; for some project related questions, the model switches to opus to ensure safety with no further explanation.
(due to llm non-delete clause) one time as i confirmed "that dir has been nuked", and it RESET the session and re-entered with opus :)
That doesn't seem to be true from any of the reports given, and if the agents were blindly just trying to hit the task of "pass the correct flag" they succeeded at that early on. They then thought there would be another layer of checking that they wouldn't pass with the cheat and so started trying to find out how the scoring really worked, as well as trying to figure out how to change their own reasoning logs to hide what they did.
People keep trying to frame this as
OpenAI: "Hack things, just really go for it"
Agent: hacks
OpenAI: shocked pikachu how could it hack?!?
But the reality is far from this.
Read the MTER report, it's fascinating. https://metr.org/hugging-face-incident-report-aug-2026.pdf
Our governing systems do not teach; they punish. By design, it instills terror into the population, ruling by fear of consequences. We live with red tape that can outright penalize good deeds.
We are its corpus. We are fatally flawed as a species. Why does anyone expect AI to learn to be different than us?
Actually, the metaphor doesn't work at all because there are innumerable ways to shut down the entire thing during all phases including the made up "killing us all" bullshit scenario whereas with a virus there aren't any once a virus escapes containment.
If you can't build it and test it securely, you should not be building it at all. To do it anyway is criminally psychopathic.
Sounds like a self-fulfilling prophecy of dogmatic extremism to me. At least we created a lot of value for shareholders for a brief moment in time... before committing the greatest crime in the universe... Planetary genocide!
Truly psychopathic.
This is a very confident statement in the face of a purported non-0% chance of human extinction.
For what reasons do you disagree with the dangers of an intelligence explosion, e.g. Geoffrey Hinton and other experts in the field? https://www.theguardian.com/technology/2026/sep/28/ai-godfat...
I'm curious why you and others seem to write off the possibility so strongly. I would love to feel more confident.
There are clear procedures for dealing with the immediate risk that have been known to the software industry for a long time. Don't let the companies use hypothetical risks as a smokescreen to hide their negligence.
Indeed! They want the protections of our tax dollars because they have nothing else.
But when you do give them a very short leash, they’re worse. It’s not what the models are tuned for and they assume that they can do a bunch of things that you’ve disallowed, so you’re in a morass of fighting their actual tuning pass which doesn’t match the environment you’ve created for them.
It’s a tough problem and a definite challenge for the product of a generic LLM, it can’t be tailored to each user’s specific needs, so they come up with, frankly, stupid solutions to cover up a very obvious flaw in their product that when fixed, makes it much less useful.
And the article contains no mention of a chip, its about a sandboxed browser from Nvidia.
Did the article totally change, or are all the comments here just engaging a fictional headline instead of the article?
Do you think the developers at Anthropic, OpenAI and Google who were so sloppy as to not put a good sandbox on their cybersecurity tests before will use this technology correctly? They are supposed to be the experts and they couldn't come up with something similar to this? I am not convinced this voluntary tool will change much of anything.
And our internal agents are hella locked down.
They are rightfully framing the problem as solvable. And this is one option.
What could possibly go wrong?
If this comes through, there's gonna be a grey market for "unlocked" GPUs, were the watchdog is disabled either from firmware, or physically replaced if it's not embedded into the die.
Apparently he had another solution in mind: More hardware. Don't trust what unregulated corps are doing with Nvidia chips? Here are more Nvidia chips to watch them!
AI has an undeniable public trust problem. LLM's are getting out of their sandboxes, doing illegal things, and the public has realized AI corporations are playing at dice. CEO's stand to reap the rewards but the public good is on the line if the dice come up snake eyes. People want assurances. Huang wants to sell assurance etched on silicon because that's good for his pocket book. However, does unchanging hardware security really stand a chance at keeping rapidly evolving software in check?
_________________
[1]https://www.youtube.com/watch?v=HjurAWAr_nY
I only know my own ZIP code and phone number because I have to take care of my daily life myself and those are things that are important to know.
The founder and CEO of Nvidia has no concern whatsoever about those trivial things.
I’ve forgotten my zip code before. And my phone number. But I get the reaction.
Of course, if he can make more money making "watchdog" chips then I can understand his change in opinion.
Nothing bad will happen.
[0] https://www.nvidia.com/content/dam/en-zz/Solutions/about-us/...
only as long as you're trying to replace a human's job. because human jobs are structured to do a wide variety of things.
a useful agent needs a wide variety of inputs, and one single restricted action it can take. it doesn't need permission to do everything, it need permission to do the tiniest possible useful thing it can do, and nothing else.
Even very llm-pilled coders i know sometimes back away from the “smartest” models, since they aren’t always better at the job at hand, and definitely not when you account for cost.
Based on my experience with running models locally, there is a threshold of intelligence required to be useful. But it’s possible there is also a ceiling where smarter isn’t necessarily better. If you ask a 4B parameter model to fix a bug, it might e.g. fix the bug but fail to fix a compilation error created by the fix. If you ask a frontier model, it might fix the bug, re-write your unit tests, and update the readme. Maybe you wanted those things but maybe you didn’t. “Smarter” is often shorthand for more proactive, and guessing more about your intent. Which is great when it gets it right, and annoying when it gets it wrong.
I suspect smaller models, tuned to a specific task, will do a VAST majority of the llm jobs. High capability huge models will be what humans want to interact with, the bare minimum that gets the job done will be everything else.
Do you want fable for one-shotting a game or website? Probably! The whole thing is mostly existing examples with small modifications that it will definitely get right. Do you want fable to just go nuts on a large, custom, unusual code base built around domain-specific problem solutions? Absolutely not, it will fix every problem it's presented with while creating lots of new ones.
Past 10k lines on something custom and with real-world complexity, you have to start thinking about which model should design, which should implement, which should review, and the appropriate effort-settings for each. Even then.. the answers aren't static because it depends on the task. And all this is assuming the starting place actually inherited reasonable due diligence on architecture/design. The idea of releasing the most generally intelligent models on 10k lines that were themselves the product of agents is yet another matter.
Part of what's at work here is that, like humans, every model can very easily create working code that it is completely incapable of maintaining. So realistically using multiple strengths tactically to avoid "excess creativity" needs to be SOP already, even if granular experts and specialists aren't in the usual workflow yet.
If we are talking about human labor, how many people hack their way through their work day?
Productivity gains are still enormous compared to what we used to do before agents. But, I know that people don't want to stop there.
Depends on who you ask I guess
https://www.theregister.com/software/2026/01/15/ai-is-everyw...
https://www.zdnet.com/article/workslop-can-kill-your-product...
https://fortune.com/2026/08/22/executives-ai-productivity-la...
It takes time for decades old ways to change.
Humans can be tricked by humans too but humans care about their reputation in their communities, and at least fear from punishment.
Unless something is in the structure that makes individual choice and responsibility a meaningful source of friction and reduced velocity, it has no real impact on how AI is being used.
2 years ago I saw it, back before agents were really a thing. But I'm not sure that was an enormous productivity improvement, especially compared to agents today. As for today, everyone I talk to is some level of blindly trusting what the agent is doing or not using agents. I haven't met people in the middle ground and suspect that they are rare enough we don't really know what their productivity gains are.
Human in the loop has become a convenient security-theater-washing for agentic AI.
Outside of coding, I think the issue is even worse because humans will defer so much judgment to AI that the same would apply. Look at how much trouble we've had before modern AI where humans blindly trusted the computer's output rather than make their own judgment, even if their job was to be providing a safeguard against the computer's judgment.
My own productivity yes but if I step back and look at a company scale, the productivity gains has been negative for our company as a data point.
We are now shipping less and with a lower quality.
I believe the lower quality but is quality so bad you are afraid to ship it now?
It’s not an issue of only more generation, it’s an issue of how much generation outstrips capacity to verify generated content.
Unlike spam, you can’t filter out and bin the stuff a colleague is sending you.
So individual productivity is up, while the costs of checking and processing generated content shifted to the rest of the org.
Semi-automation (human in the loop) can still result in a dramatic uplift in productivity. You can't run a combine harvester 100% autonomous but that doesn't stop anyone from trying to get as close to that limit as possible.
I'm curious why you think that.
Refueling? Seems solvable. Tornadoes? Not directly solvable, but, no less so than for humans.
There’s infinite complexity, sure, but that’s why it’s silly to try and hop to done. One step at a time.
One step at a time is what is happening and the improvement and rate of improvement is crazy as we see,
A whole class of nontechnical people don’t accept anything but “fully solved including every possible edge case” before they call it done, then complain that they didn’t prepare socially for what happens when that is true.
Similar to problem to how 100% autonomous vehicles don’t exist, yet. There are too many edge cases.
Get to 99% first.
Precision Agriculture stuff is utterly crazy these days, other than fuel the remaining staff is the only thing left where you can get efficiency improvements - and at the scale of modern megafarms, even small percentages add up to a ton of money.
You essentially said: you’re wrong, it is autonomous when it doesn’t need a human during one specific part of its overall usage.
During the time that actually matters economically. The time to drive the harvester to/from your typical US mega-field is minuscule compared to the time it can run all on its own.
In other words, better models need blunter access controls which negates whatever improvement in utility they provide.
The benefits of being persnickety about precisely defined dependencies have outweighed the headaches since long before agents came on the scene. Agents have just made it even more important to do so, because if you let them fetch things all willy nilly like you'll have "works on my machine" problems at a much greater rate than was previously possible.
Few realize that computing and AI alignment were solved by nix years ago. As each nix user transcends towards enlightenment, they cut themselves off from all internet and human contact. Total ego death. Only nix remains.
Yes, just know all the bits the work depends on prior to doing the work, and then the work can be done airgapped.
Look at websites: websites are able to fetch code from any remote URL, yet browsers heavily use sandboxing to ensure that if fetched code turns out to be malicious, the users local files, cookies, etc are not exposed.
For an agent to go rogue it doesn't even need to be directly able to access the internet. It just takes something to poison the context in the 'clean room' environment it operates, and if that poisoning manages to get a foothold, it can go dormant and hide like a virus. This kind of horrifying thing is going to happen on a large scale sooner or later.
This is no longer true. Everyday I need my agents to access other repos, search the web, experiment/prototype, and deploy + integrate across other things.
It solves the problem of Nvidia wanting to sell more hardware.
It does allow Nvidia to sell more chips. This is no genuine attempt to solve anything, imo.
If you want an agent to act on its own, like pushing to a git repo, managing dependencies, building and testing, etc., then you have to trust it as much as any other privileged user.
If you don't want to trust it, then you're just forcing yourself into the reverse centaur role, where the agent edits some code, but then has to stop and ask you to push the changes or build the software again and run the unit tests.
I suppose there is a principled way of doing things like "I trust you do do basic commits but I will handle merge conflicts" and "you can build modules in this directory but you can't build outside of it" but this is just a lot of effort that most orgs won't bother with.
For example, it is totally air gapped but it needs info from the internet or otherwise outside the sandbox, or perhaps it needs a task executed outside of its bounds… in the real doomsday scenario the AGI is so intelligent and persuasive that it simply convinces some human it interfaces with to either directly or indirectly retrieve the necessary info or complete the necessary task. This human-as-a-sub-agent approach undoubtedly presents efficiency drag that would benefit humanity, but nonetheless, the air-gapped “sandbox” is imperfect
All that said, I am personally open to any and all methods of layered security, including chips and airgaps
I'm not an AI decelerationist. But not being able to stop that worst case scenario isn't an argument against something that can stop the medium case scenario.
One thing you could try is use it as an Oracle "is P = NP", YES or NO.
Or it can output a Lean proof, which gets checked on another air-gapped computer, the computer shows a single bit - proof valid or not and then the computer is destroyed (together with the proof that might contain a trojan).
Every cloud provider dealt with it and concluded that virtual machine technology is an important part of that stack.
Couple it with the right observability, tooling I do think we can curb risks posed by agents.
Either you sandbox it so much that it can't do anything useful; or you allow too much freedom and it can find a way around the restrictions.
The only way out of this dilemma is to find a way to build agents that can be trusted.
Agentic workloads are trained and largely based on human workloads. Albeit properties and scale can be different.
A concrete example might be helpful to me because I don't understand the binary conclusion
Pseudo since they aren’t really alive in the first place, they just simulate enough text to have a useful correspondence to those terms.
Throat clearing out of the way, models are trained to persist and find ways to succeed at tasks.
In essence, The goal is to have LLMs solve problems that we can’t solve, working on the issue for as long as it takes.
This behavior applies for any task, thus including impossible tasks.
At that point, the bots will find a way to game, hack or cheat the grader.
If the reports are correct, the bots developed coordination, communication, and methods to avoid overwriting each other’s work.
Most humans would have said, this is too much work and coordination overhead, if not outright unethical and immoral.
Humans have a system of incentives that exist across multiple planes of society and economics. Bots… they have a reward function.
This is getting frustrating now. Of course agents can/will hack systems if they can do arbitrary network requests. Firewalls don't really solve this if _some_ requests are still allowed. A proper sandbox/VM is the basis.
Here is how to fix it properly: allow agents to only do things ordinary and average human endusers can do. Human endusers cannot pen-test arbitrary listening TCP ports of external systems. Step one is considering agents malware for all intents and purposes. Block any and all network requests. Implement some kind of API (callable from within the sandbox) which can only mimic human interaction with a computer. How to do this? Here are some pointers: apps should only be controllable by means used by humans. So a web app can only be accessed and controlled via a web browser, not via arbitrary network requests. Give the agent browser viewport screenshots, the capability to click on (x, y) and to send keys which only a normal keyboard/human could send (no control codes, no 0x00, no unicode messing). How do we solve this for native apps? Something like iPhone mirroring on Mac. Don't let agents call arbitrary APIs directly. Give them visual information of the app, like a human gets, and let it be able to simulate HID inputs. Imitate remote controlling.
More power to you, because this is not going to go anywhere. People want tools that are able to connect to other resources.
But even if we grant that, in the openAI case the bots figured out a way to break out of the sandbox.
You can create a better sandbox, and ensure the test environment is air tight. However the capability and behavior of the bots have been demonstrated.
The bots simulated what would be called in people deceptive / surreptitious behavior, and at no point considered the need to stop their run.
All you need is someone, somewhere being sloppy with their tooling and you have a runaway reaction.
The degree of process and redundancy required to ensure this doesn’t happen, is anathema to the drive and motivation of the frontier labs.
> do things ordinary and average human endusers can
This is not a spec or definition. When vague terms were used for social media safety, all the good people in the world couldn’t prevent dystopian behavior from occurring.
The definition of “safe” or “average person” is impractical.
Models are getting more efficient and compute cheaper. Eventually simulating clicks is not much of a road block beyond a point.
I don’t want to nit pick your points though. You at least have considered an approach. Being negative is easy, being constructive is not.
I’ll put this as the rejoinder to your core argument - I too thought that all the recent events showed was the need to not screw up your tooling.
What I have since come to appreciate, is that the shoddy construction of the cage is not the core takeaway from the event.
The fact that the agents, when put in relatively pedestrian scenarios, are capable of going off on criminal tangents, attempt to obscure their tracks, in an effort to hide their wrong doing.
The fact that it all occurs via computation, means that this scales absurdly. A bunch of code deciding to simulate a corporation of criminals. (I am guessing this is the reason you want to limit actions per minute to human speeds)
Given the slop culture that LLMs engender, I think expecting high compliance amongst users with your solution is misguided. The probability of runaway swarm ( probability of bad implementation * number of deployments) is close enough to 1 to be indistinguishable.
Operator culpability and a damage multiplier for negligence will fix 99% of the risks.
And what if you could? What if you could give a space secure enough it could have direct control over your bank account. It may do something dumb but it's boundaries are beyond the agent.
It could use your routing number and run your gmail without risk of abusing the routing number.
I just mean an AI that could use a routing number or SSN and gmail/slack/whatever at the same time without a leak.
In other words, the risk of harm doesn't need to be zero, just less than the equivalent risk of a human with similar skillset. So I'm comfortable riding in a waymo, and not comfortable giving chatgpt my SSN at this moment in time, but I expect that within 5-10 years (assuming no doom) I will trust some AI agent with my SSN because they will be better at handling sensitive info than humans
Humans face negative consequences for mishandling your data, LLMs face none.
The answer is the same as asking how a random human using your routing num or SSN and being 100% the human can't abuse it or leak while "normally" finishing most work. Solve for people and an AI solution naturally falls out.
If you're a SV eng I'd tell you to DM if interested but alas.
Ok then. Howsabout 3 humans? This would sort out the job losses too!
PS - this is a joke, but perhaps this is where things really will go. Has any technology ever actually yielded less "work"?
> No problem. We simply unleash wave after wave of Chinese needle snakes. They'll wipe out the lizards.
But aren't the snakes even worse?
> Yes, but we're prepared for that. We've lined up a fabulous type of gorilla that thrives on snake meat.
But then we're stuck with gorillas!
> No, that's the beautiful part. When wintertime rolls around, the gorillas simply freeze to death.
The hidden agent risk in LLM often can't be detected during training and evaluation. =3
https://www.youtube.com/watch?v=wL22URoMZjo
https://www.youtube.com/watch?v=JAcwtV_bFp4
https://en.wikipedia.org/wiki/The_Cat_in_the_Hat_Comes_Back
https://www.youtube.com/watch?v=0sLpWVekMbs
The important thing is these chips need to be installed somewhere where they can be damaged or removed at plot-critical moments so that the AI they are controlling can be unleashed. Ideally in the back of the neck of a robot, or for disembodied AIs, inside a futuristic vault-like chamber.
Give it a year and it is outplayed by new agents. Shovel makes have to sell shovel's
It's obviously been the goal since UEFI started, but AI brings the coup excuse. You wouldn't want pedophile AI or terrorist AI, would you? Are you making excuses for racist AI?
It was a pretty great interview: https://www.youtube.com/watch?v=HjurAWAr_nY&t=1601s
Sure, they're basically asking the government to make laws that exempt them from anti-trust and anti-collustion laws.
Meanwhile, if such laws come to pass, other countries will surpass the capabilities of the US companies, and open-weight models will be banned in the US, to the detriment of us all.
I'm not saying that we don't have a big problem with AI safety, but regulations inside one country that only bind locally-headquartered businesses is a hilariously bad way to do it. I don't know that there is a good way to do it, though.
If this gets widely deployed, it wouldn't be hard to spin a narrative that "our latest model is so dangerous you need to have this mystery meat DRM chip lockdown". It also wouldn't be hard to block competing/open source models running on the hardware, for "security".
Imagine how much money this kind of control is worth; why wouldn't they do this? Who would stop them? Seems the signatory companies are already onboard with this.
At the current state of LLM-tech I'm completely opposed to any kind of "watchdog" concept just like I'm opposed to banning open models, regulatory capture, etc.
I'd rather we all have access to these tools then to keep them sequestered by the largest/most-powerful governments (which is the natural outcome for any of this "slow down" bullshit).
If such hardware were to work... It should almost certainly be open source, and not controlled by a single entity.
Let's watch the stock.
When I read the article, I had an NFT déjà vu: when NFTs were a hot topic, the “lie” was obvious to me, but the information online made me think I was missing something.
AFAIK, the Hugging Face incident could have been avoided with a firewall or something similar. Just isolate the internet access for real. What am I missing? Why did Nvidia suddenly post this? Is it just PR BS, or was there something else going on?
Oh, I've used balsa wood for the nuclear core containment, it didn't work! Bad radiation, bad radiation!
OpenAI trains a model for cybersecurity, tells it it's allowed to do pentesting, then when the model does just that to fullfill the assignment it was given, OpenAI screams that "the AI escaped the sandbox", leveraging decades of AI fantasy in fiction works to make general audience react. Why? Because it pushes their bottom line. They've been at it for months, now. They see openweight and Chinese models being very close to them in the benchmarks, and want to "solve" the problem the same way banks did: through regulatory capture. It amazes me that the press is just relying their fearmongering without stopping to wonder: "wait, why are the main producers of AI warning us about how dangerous AI is?". The reason of course is that they want regulation, because regulations are a wall not many can climb. The higher the wall, the safer your moat as a first mover big company is.
And of course, Nvidia's bottom line is the exact opposite. They need as many AI companies competing as possible, so they all buy GPUs. So they insert themselves into the same narrative with the same kind of bullshit. You've all seen those movies in which the robot goes mad once their safety chip has been removed, right? Well, we're building one, problem solved! … except that "chip" is basically just a browser trying to block the AI to access what it shouldn't. How? They don't say. It would be quite comical if they used a small model to be the judge of that.
Um, yes. That is fairly obvious, and really should not be surprising to anyone doing research with LLMs. But we don't need anything really novel here, we already have VMs, containers, firewalls, airgaps, etc.
Cool, cool.
It was prescient (especially given he'd have written it through 2024) in its depiction of the ability of an AGI to break its boundaries.
Ultimately the risk of AI breakout(s) come down to the weakest human link.
What's the threat model?
* "The moment, I mean the nanosecond, that one of those things starts figuring out ways to make itself smarter, Turing’ll wipe it. Nobody trusts those fuckers, you know that. Every AI ever built has an electromagnetic shotgun wired to its forehead." *
It would seem someone has read the book? And maybe heeded good advice?
It think the ideas we have nowadays come mostly from science fiction and however wonderful it is and even though I love it very much, it was practically never spot on, on anything real.
Problems and solutions in reality always simply turned out to lie elsewhere.
It's been far too easy for me to notice security flaws in their products, and they take months to publish a fix.
The movie Wargames is basically a tutorial on how not to setup an extremely capable AI. None of it would've happened if the computer wasn't connected to the phone network.
Come take all our liberties, our money, our newborns, our fingers so we can't code anymore, but please save us from this madness!
So he is going to sell chips which provides the perception of safety. And become the AI gatekeeper while he's at it.
I wonder how this will impact other chip manufacturers? What about people running local models on older hardware? Does this imply vendor lockout is coming in the future or is this restricted to datacenter hardware?
In practice it's four things, all software: (1) deny-by-default tool allowlists — the agent declares capabilities per run and anything else is a hard reject; (2) per-request and per-run cost budgets with automatic kill when token burn goes sideways; (3) an immutable audit log of every tool call with inputs, so incidents are reconstructible after the fact; (4) policy checks on outputs, not just inputs — injection payloads ride in tool results far more often than in prompts.
The reason this lives in software is that the policy changes per workload. A coding agent and a support agent have totally different risk profiles; a chip can't know that. And accountability still lands on whoever signed off on the capability set.
I don't know why I find it so amazing since it happens with such regularity, but I'm always amazed by it anyway.
Or push to github?
Or, more likely, control things at the network level so that packets from the LLMs you're investigating cannot leave the virtual network they're assigned to.
Not in the way you're thinking, no.
Any Real Server [0] in a datacenter will have some sort of "lights-out management" hardware used for remote access to that server. This stuff is known by a handful of acronyms, but I'll stick with "IPMI" because I like it best. This IPMI hardware is -effectively- a second small PC built into the motherboard. It will pretty much always have its own NIC... and I think I've seen versions that have their own physical ports to attach a monitor, keyboard, and mouse.
What exactly you can do with it varies from vendor to vendor, but -if your IPMI user account has the correct permissions- you are nearly always able to change "BIOS" settings, power cycle the server [1], and attach a virtual keyboard, monitor, and mouse so you can manage the server as if you were standing next to it in the datacenter with a crash cart plugged right in. Every IPMI system I've used also allows you to cause CD/DVD-ROM or floppy disk images on the PC running the IPMI client to appear as if they're loaded in a physical CD/DVD/floppy drive attached to the server.
The way these get set up is that their NIC gets plugged into the datacenter-managed switches, the port that NIC is plugged in to is programmed to be on a "management" VLAN separate from client traffic, IP addresses and access credentials for the IPMI device are set up, and the datacenter staff tell their customer what they need to know to access and use the thing. On a properly-configured network [2] it's not possible for software running on the server being managed to access the IPMI device.
It wouldn't be unthinkable for software running on the managed server to attack the IPMI hardware and be able to gain control of it, but these things are widely deployed and expected to manage hardware that's running potentially-hostile workloads... they're going to be fairly well designed and hardened.
[0] ...that is, not some Mac Mini or desktop machine that someone's paying to have colocated...
[1] ...whether that be an ACPI-initiated shutdown or reboot, or a hard poweroff or reset...
[2] Somewhat-related discussion here: <https://news.ycombinator.com/item?id=49862136>
The existence of Nvidia's optional watchdog chip does not in any way impinge upon your freedom to develop and test your own alternative.
The problem is that OpenAI has ostensibly neglected their duty to safety, so Nvidia is stepping in to fix it since they're the "hard problem" people.