The Data is Better Together community has released an open preference dataset for text-to-image generation, addressing a long-standing challenge in the field of AI research.

What Happened

The new dataset, called "open-image-preferences-v1," was created by the Data is Better Together community to address the lack of open preference datasets for text-to-image generation. The dataset consists of over 10,000 preference pairs across common image generation categories, with varying model families and prompt complexities.

To create the dataset, the community used a multi-model approach that combined two text-based and two image-based classifiers as filters to remove NSFW prompts and images. They also employed synthetic data generation using distilabel to enhance the dataset's diversity and quality.

Background and Context

The ability to collect large datasets of human preferences from text-to-image users is often limited to companies, making such datasets inaccessible to the public. This has hindered research in the field of text-to-image generation, where understanding user preferences is crucial for developing effective models.

Previous efforts to create open image preference datasets have been limited by their scope and complexity. For example, the Pick-a-Pic dataset, created by Yuval Kirstain et al., consists of over half-a-million examples of human preferences over model-generated images but has some limitations in terms of its size and diversity.

Why It Matters to the Industry

The release of the open-image-preferences-v1 dataset is significant for several reasons. Firstly, it provides a large-scale, open dataset that can be used by researchers and developers to train and evaluate text-to-image generation models. This can lead to improved model performance and more accurate predictions of user preferences.

Secondly, the dataset's focus on varying model families and prompt complexities makes it an ideal resource for fine-tuning existing models or developing new ones that can handle diverse input prompts. This is particularly relevant in the adult industry, where text-to-image generation models are used to create high-quality images for various applications.

Thirdly, the dataset's openness and availability can facilitate collaboration among researchers and developers, leading to more innovative solutions and advancements in the field of text-to-image generation.

What Comes Next

The Data is Better Together community plans to continue organizing community sprints on the Hugging Face Hub. They encourage community members to participate in future sprints, propose their own sprints or requests for high-quality datasets, and contribute to the development of new models and fine-tuning existing ones.

Additionally, the community invites researchers and developers to use the open-image-preferences-v1 dataset to train and evaluate their text-to-image generation models. They also encourage collaboration and knowledge-sharing among community members to advance the field of text-to-image generation.

Key Facts

  • The Data is Better Together community has released an open preference dataset for text-to-image generation, called "open-image-preferences-v1."
  • The dataset consists of over 10,000 preference pairs across common image generation categories.
  • The dataset was created using a multi-model approach that combined two text-based and two image-based classifiers as filters to remove NSFW prompts and images.
  • The dataset's focus on varying model families and prompt complexities makes it an ideal resource for fine-tuning existing models or developing new ones.
  • The Data is Better Together community plans to continue organizing community sprints on the Hugging Face Hub.