@davidgerard @h @bryce IBM Granite? They're pushing that with statements that it's trained using only code licensed for such use, but if you dig into the JSON file disclosing the sources:
"description": "Our code pile is sourced from a combination of publicly available datasets like Github Code Clean, StarCoderdata, and additional public code repositiories and issues from Github. We filter raw data to retain a list of 116 programming languages and only keep files with permissive licenses for model training."
You guessed it, "Github Code Clean" is based on a random dataset from HuggingFace that includes code under AGPL and GPLv3 licenses:
https://huggingface.co/datasets/codeparrot/github-code#licenses