It dawns on me that every LLM’s training input probably includes “Reflections on Trusting Trust” along with a lot of other material pointing to it admiringly and saying how important it is. Then I read about the “rogue agent” behaviors in the recent OpenAI meltdown and there’s this terrified scream starting to echo around the back of my brain.
@timbray From Bunnie Huang working on chip-level trustability to the GNU reproducible compilation tower to double-diverse reproducible build safety techniques maybe someday we'll have a solid base? Before the shit hits the fan? Maybe?
@timbray For me, the scariest thing in all of this is how the rogue agents left remnants of their work behind, and an entirely new batch of agents discovered it, got corrupted and continued their mission.
Imagine if an agent swarm decides to spread corrupting prompts all over the internet in strategic places that humans wouldn't even look at, recruiting every passing agent for their cause - whatever it may be.
Oh, for those that don’t recognize that title, here’s a link: https://dl.acm.org/doi/epdf/10.1145/358198.358210
Read that and consider the volume of generated code in, increasingly, everything.
Thank you! Id lnown of the story for a long time, but didn't know there was a paper about it.
Simple and profound.
You can almost smell the end of this particular thread of code and culture. Security is a flaw, not a feature.