A growing pushback against big tech companies, and greater awareness of the value of data, is spurring interest in data collectives and cooperatives, which give communities control over the collection, management, and distribution of their data. This alternative allows creators to benefit from data sets that may otherwise be ignored or misused.

A handful of tech companies dominate the generative artificial intelligence industry, with American and Chinese frontier models controlling the lion’s share of the market. Some countries are building their own large language models because their language and culture are not adequately represented in GPT, Gemini, Claude, or Qwen.

The big companies have built themselves up on the backs of all these people creating data, who think it’s time to set their own terms now.”Raffi Krikorian, chief technology officer, Mozilla Foundation

Communities that possess smaller or unusual data sets can gain from having control over them, Raffi Krikorian, chief technology officer at Mozilla Foundation, told Rest of World. The nonprofit last year set up Mozilla Data Collective to provide a platform for such data sets from communities, organizations, and individuals around the world.