• 1 Post
  • 55 Comments
Joined 3 years ago
cake
Cake day: March 22nd, 2024

help-circle






  • Admittedly, no, but this is slowly improving. Stepfun published their SFT dataset, smaller labs are publishing their datasets for task specific tunes. I believe there was another Chinese lab that published bulk pretraining data, but I can’t find it in my history at the moment.

    And, notably, these comparatively tiny labs generally aren’t scraping the internet so abusively like OpenAI/Meta. They don’t need as much. Going by statements in their papers, they tend to use existing archives of web data, buy commercial data, or (more recently) generate a lot synthetic data.