Book shredding, statistical watermarking and keylogging are all the same story: a desperate effort to avoid model collapse.
Large-language generative models are a dead end. Open weight models are bullshit.
Small models with carefully curated datasets pointed at specific problems are the future. Stop scraping the world and start hiring librarians.
https://www.anthropic.com/news/claude-text-watermark
https://futurism.com/artificial-intelligence/new-chatgpt-feature-collects-every-keystroke-you-make