https://openreview.net/forum?id=G-7UQeQAVD&referrer=%5Bthe%20profile%20of%20Leon%20Weber-Genzel%5D(%2Fprofile%3Fid%3D~Leon_Weber-Genzel1)
The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset | OpenReview
As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The...
rootscorpus
https://huggingface.co/bigscience-historical-texts
bigscience-historical-texts (BigScience: LMs for Historical Texts)
historical texts, named-entity recognition, big science
historical textslms
https://wandb.ai/telidavies/ml-news/reports/BLOOMZ-MT0-BigScience-Releases-New-Finetuned-Large-Language-Models--VmlldzoyOTEzMzU3
BLOOMZ & MT0: BigScience Releases New Finetuned Large Language Models
Nov 4, 2022 - BigScience has released a new lineup of finetuned models based on their own BLOOM model as well as Google's MT5 model: BLOOMZ and MT0 respectively. These...
releases newbloomzlargelanguagemodels
https://cohere.com/research/papers/bigscience-a-case-study-in-the-social-construction-of-a-multilingual-large-language-model-2023-05-05
BigScience: A Case Study in the Social Construction of a Multilingual Large Language Model
The BigScience Workshop was a value-driven initiative that spanned one and half years of interdisciplinary research and culminated in the creation of ROOTS
a case study
https://github.com/bigscience-workshop/t-zero
GitHub - bigscience-workshop/t-zero: Reproduce results and replicate training fo T0 (Multitask...
Reproduce results and replicate training fo T0 (Multitask Prompted Training Enables Zero-Shot Task Generalization) - bigscience-workshop/t-zero