DataLearner logo
BigScience

BigScience

2 models tracked · latest release 2023-03-16

UL

Product line release timeline

Generational evolution by product line · dot = one model release · dashed line connects successive generations · click a dot to open the model page

Models
2
Time span
248days
Avg. gap
248days

Published models

2 models

Models published by BigScience, grouped into 1 categories.

About this organization

'

The acceleration of artificial intelligence (AI) and natural language processing (NLP) will have a fundamental impact on society because these technologies are at the core of the tools we use every day. Currently, a considerable part of the work in NLP is training larger and larger language models on more and more texts.

Unfortunately, the resources needed to create the best-performing models are largely in the hands of large tech companies. Control of this transformative technology poses problems from a research advancement, environmental, ethical and social perspective.

For example, while recent models such as OpenAI/Microsoft’s GPT3 exhibit interesting behavior from a research perspective, these models are private and inaccessible to many academic institutions. Moreover, even if accessible, these tools were not designed as research artifacts, such as lack of training datasets or checkpoints, which makes it impossible to answer many important research questions about these models (capabilities, limitations, potential improvements, biases, ethics, environmental impacts, general AI/cognitive research landscape). The current situation also promotes duplication of energy requirements and environmental costs due to the repeated training of large models in private settings. Finally, these models are often English-centric, and the text corpora they are trained on have flaws, ranging from non-representative populations to potentially harmful stereotypes or containing personally identifiable information.

The BigScience project aims to demonstrate an alternative approach to creating, studying, and sharing large language models and large research artifacts within the AI/NLP research community.

The project draws inspiration from scientific creation programs in other fields of science, such as CERN and the Large Hadron Collider (LHC) in particle physics, where open scientific collaboration helps create large artifacts useful to the entire research community.

Gathering the larger research community around the creation of these artifacts allows us to consider in advance many of the research issues surrounding large language models (capabilities, limitations, potential improvements, biases, ethics, environmental impacts, overall picture of the intelligence/cognition research field). It is then interesting to use the artifacts, discussions, and tools created to answer as many of these questions as possible and facilitate conversations around key aspects of the research area.

The BigScience open science project is seen as a proposal for an international and inclusive way of conducting collaborative research. In addition to the research artifacts created and shared, the project aims to bring together all skills, conditions, and lessons learned for future experiments in such large-scale scientific collaborations.

Ultimately, therefore, the founding members are convinced that the project's success will ultimately measure its long-term impact on the fields of NLP and AI by proposing "an alternative way of conducting large-scale scientific projects."

a research working group

The collaboration is a one-year large language model research workshop: "Summer of Language Models 21" 🌸.

This workshop will:

Taking place online during the year: from May 2021 to May 2022

Include real-time events scattered throughout the year (first online, subsequent possibly offline), including at least the opening and closing ceremonies.

Conduct a series of collaborative tasks aimed at creating, sharing, and evaluating a large multilingual dataset and large language models as research tools.

This workshop will promote discussion and reflection on the research issues of large language models (power, limitations, potential improvements, biases, ethics, environmental impacts, and role in the field of AI/cognition research), as well as the challenges of creating and sharing such models and datasets for research and among the research community.

These collaborative tasks are quite large, requiring millions of GPU hours of computation on supercomputers.

If successful, this workshop could also be run in the future, including an updated or different set of collaborative tasks.

'