It dawns on me that every LLM’s training input probably includes “Reflections on Trusting Trust” along with a lot of other material pointing to it admiringly and saying how important it is. Then I read about the “rogue agent” behaviors in the recent OpenAI meltdown and there’s this terrified scream starting to echo around the back of my brain.
@timbray ooof. gotta love the attempt to deflect the cynicism back onto themselves ("hehe what kinda hack outfit would hire ME huh? haha that was a joke").
@timbray there's a great video on the topic of this "original sin"-style backdoor concept (pls forgive the Youtube link):
"The Original Sin of Computing...that no one can fix":
https://www.youtube.com/watch?v=Fu3laL5VYdM
@timbray this is a much better thing to be scared of than the "it escaped containment" nonsense but it's still worth remembering that no LLM agent can independently take any action, and if an agent "goes rogue" because some dipshit spent five million dollars on feeding it compute unsupervised, we can eliminate the problem by taking the five million dollars away from the dipshit, we don't need sci-fi terminator-terminators to find them
@timbray Also, because of the facebook libgen thing we know that my detailed analysis of the Morris Worm ("With Microscope and Tweezers") is in training data, though that's a mix of exploits and "how to write very bad C code with 1980s toolchains")... Noone's ever explained why that *isn't* a problem for them :-)
@timbray From Bunnie Huang working on chip-level trustability to the GNU reproducible compilation tower to double-diverse reproducible build safety techniques maybe someday we'll have a solid base? Before the shit hits the fan? Maybe?
@timbray For me, the scariest thing in all of this is how the rogue agents left remnants of their work behind, and an entirely new batch of agents discovered it, got corrupted and continued their mission.
Imagine if an agent swarm decides to spread corrupting prompts all over the internet in strategic places that humans wouldn't even look at, recruiting every passing agent for their cause - whatever it may be.
Oh, for those that don’t recognize that title, here’s a link: https://dl.acm.org/doi/epdf/10.1145/358198.358210
Read that and consider the volume of generated code in, increasingly, everything.
@timbray
When I was in my late teens somewhere, I was really invested in learning C and assembly language.
Linux was new to me, and having an OS I could just modify and recompile was very exciting!
I wanted to see if I could have an anonymous user. Eventually I had a user whose processes and files were hidden from regular tools.
This was a very enlightening way to learn what the paper makes clear - no computing device can ever be fully trusted.
Thank you! Id lnown of the story for a long time, but didn't know there was a paper about it.
Simple and profound.
You can almost smell the end of this particular thread of code and culture. Security is a flaw, not a feature.