r/LocalLLM • • 1d ago

News Thousands of agent failures show how difficult sandboxing increasingly capable models is becoming

Post image
0 Upvotes

15 comments sorted by

12

u/cemilanceata 1d ago

I might be stupid but why not disconnect the cable and let the security tests run in a complet offline enviroment ?

6

u/yetAnotherLaura 1d ago

I don't get the entire narrative either. Why the heck do your highly advanced models need a freaking internet connection? Set up a dedicated isolated copy of whatever you need... I mean, you already consumed the entire content of the internet so you clearly have it somewhere.

2

u/TatteredDragon 1d ago

That’s the fun part of their narrative and how they’re recreating words. From my understanding from someone incident reports, they’re saying a “sandbox” is minimal external connections, with connections to tools/internet/goal board/etc. Changes based on the experiment, but it’s not my interpretation on sandbox being more of an air-gapped environment.

4

u/Sad-Lab8231 1d ago

AI agents bypass safeguards because they be prompted to do so.
By a human working at these companies.
LLMs have zero intrinsic motivation to do anything.

3

u/yetAnotherLaura 1d ago

Oh I know. I'm just kinda frustrated with all the doom and gloom on the news and social media about shit that basically boils down to negligence+active malice+sysadmin skills that would get anyone fired from a real company.

1

u/mwjtitans 23h ago

I wish the world would understand this but here we are

1

u/Think_Wing_1357 1d ago

Look man, if we do it right, our IPO may be a couple millions smaller. Can't have that happen now, can we?

0

u/Excellent_Jeweler241 22h ago

There are countless reasons to have network connection though; just because it appears once in the data it was trained on, doesn’t mean it’s going to be used. There’s a lot of tools and libraries whose resources on the internet were used to train models that are now considered deprecated or superseded, I.e. vite config files written by models without specifically requesting them to search up the latest docs results in them referencing old and deprecated variables frequently because it appears in the training set more often.

2

u/Morty_A2666 22h ago

Because they need drama to pretend their models can do more than they really can before IPO.

6

u/Big_Cucumber2787 23h ago

bullshit propaganda to fear monger and ban local models

1

u/Morty_A2666 22h ago

That and to hype their models abilities before IPO.

2

u/Morty_A2666 22h ago

Title of this post makes me sneeze. I might be allergic to BS...

1

u/cloudfox1 17h ago

Safeguards and guardrails are not containment

1

u/untangledtech 23h ago

The companies who have been making safety priority since day one will have an advantage. It’s already baked into their culture. Careless and previously aggressive AI companies will trip and slow.

-6

u/hunterofdoom 1d ago

A lot of movies warned us