2 articles about training data
A Munich court ruled that Suno's AI models memorized copyrighted musical compositions during training, finding that songs including 'Rasputin' and 'Daddy Cool' are 'reproducibly contained' in the models. The ruling grants German collecting society GEMA injunctive relief and damages.
NVIDIA's Nemotron open-data initiative releases over 10 trillion pre-training tokens and millions of post-training samples for building AI agents, plus region-specific personas and a Prompt Atlas.