正文 Markdown
For those that are deeply passionate about model distillation as a problem, can you explain why you think it’s both more unfair than the original training runs of models on web data and/or different from us all collectively publishing our outputs of models on github and training on that? It’s clear that AI labs should do everything they can to prevent distillation, because of course they’re going to try and protect their assets as much as possible. But why do we think it’s a definitely past threshold legally or philosophically than training? Happy for any answers to this.