@lxo @skjeggtroll @freshstart @neil
Okay, that's what the law says, but how is the second person going to fare against a lawsuit if the first person claims infringement?
If the second person succeeds on an argument that "using an LLM is like a clean-room implementation" then open-source is non-functional, legally, as anyone can easily ignore any license by asking an LLM to duplicate relevant functionality and then claiming plausibly that they hadn't seen the other work before.
If, as you're suggesting, there's an "I came up with it independently" defense, does the burden of proof fall on the accuser or the accused? Normally burden always falls on the accuser, but if we assume open-source development on both sides, the accuser can trivially prove (if they published faster) that they came up with the idea first and that their code was theoretically accessible to the accused when the accused was developing their version, and that the accused used lines of code verbatim which are in their version. The accuser can probably even show that their code was likely in the training data of any LLM that the accused used.
I don't actually know the law here, but if that's not enough to throw a presumption onto the accused and the accuser actually has to prove intent or something, it again seems like open-source licenses offer negligible protections against infringement (maybe this is actually the case). If the accuser on the basis of showing that the accused used verbatim copies of their public code can shift the burden of proof onto the accused, how is the accused going to prove they actually came up with the idea independently, especially when they used an LLM so it's not really their idea?
Maybe some kind of twisted precedent *will* be established that this situation is fine actually and anyone accused of copyright infringement who is using an LLM can just claim their invention is independent despite having literal copies of another work in it... I don't want to be the person testing that legal theory against a startup with a 6-figure legal budget, let alone a big tech firm. The big tech firms regularly do patent software, after all.
On the other side of things out it turns out that legally LLM-generated code cannot be licensed or patented at all, then it can't be open source and this should not be acceptable to Debian.
But let's think about the moral level too. Do I want to be using the sometimes-steals-code machine and then get into a situation where it looks like I stole someone's code? No. What actually happened in this case was not that the slower sloperator mounted an arcane legal defense and everyone was okay to let the two apos coexist. Instead he backed off, apologized to everyone, deleted the project and promised never to use LLMs to generate code again. That seems like the best case for reputational damage. So even if the legal reality permits some really gross stuff (let's face it, this is the norm actually) unless you're a massive corporation who doesn't care about reputation, a legal technicality doesn't make this situation all good.