Who Gets to Own the Data That Builds AI?
Mozilla Data Collective aims to make AI data fair, diverse, and community-controlled.
Opinions expressed by Entrepreneur contributors are their own.
You're reading Entrepreneur United Kingdom, an international franchise of Entrepreneur Media.
Artificial Intelligence needs more than scraped internet data. Mozilla Data Collective, a fairer marketplace for AI data, is building a model for consent, fair value exchange, and community control – while helping AI companies access richer, more diverse, and trustworthy datasets. Entrepreneur UK talks to founder EM Lewis-Jong to find out more…
What did you see happening with AI that made you want to build Mozilla Data Collective?
I was watching AI being built on a diet of indiscriminately scraped web content, which meant models often inherited the internet’s worst habits: its prejudices, its language coverage, its homogenisation of culture. I spent half a decade on Common Voice, and it taught me that, when given the chance, people will show up to help build technology that really sees them, and works for the people and communities they care about. There were also a lot of organisations, and people, who were asking tough governance questions about how to make sure that dataset owners and stewards – be they museums or podcasters or tech start ups – were benefitting from tech companies using their datasets. What they wanted that benefit to look like, was varied! Sometimes compensation (e.g. large western labs should pay for a commercial license), sometimes transparency (we want you to tell us what you’re going to do with that data), or sometimes in-kind exchanges (we’ll give you the data for free if you let us use the model for free). What I saw was a structural problem that needed a structural fix: human agency and fair value exchange, built right throughout the AI data supply chain. A platform that could make CHOICE really easy for people without big tech teams or legal teams. Not an open repository – downloaders need quality, curation, and compliance. And not a closed marketplace – uploaders want transparency, collective bargaining and self-service flows. That is the gap Mozilla Data Collective exists to fill.
What is broken about the way AI companies source and use training data?
To start, not all AI companies have behaved the same in this regard, but often, I think it’s fair to say that the default has been extractive, and skips the important questions, like where did this dataset come from, and did the people in it actually agree? How are they benefitting from their dataset being used in commercial products? Is it a fair value exchange? A suspiciously cheap dataset is usually cheap because somebody upstream went unpaid or underpaid. Take speech synthesis: a student might sell a perpetual, unlimited clone of their voice for a few hundred quid, not knowing that a professional voice actor has unionised to charge thousands for the same thing, or that platforms will then rent that voice out to enterprises, on six- and seven-figure contracts. That model of either scraping, or farming out precarious gig work as the way to power AI is getting increasingly brittle: legally, reputationally, and commercially. The regulatory floor beneath it is rising quickly – just look at Article 53 of the EU AI Act – companies now have to account for how they got their data.
Can communities genuinely own and control the data used to build AI?
Yes, but only if the platform makes it easy to state and enforce their choices. Real sovereignty means the people who create the datasets should decide what happens to it: they can share openly if that is right for them, or attach conditions such as compensation, attribution, collaboration, restriction to education or research, or limits by geography or type of organisation. We make that real through legal licences, authentication and access checks, automations that flag particular requests under the hood, bespoke request-to-download workflows for sensitive material, and community moderation. Some people call these types of safeguards “developer friction.” We don’t think that respecting consent is a bug, it’s a feature, and we’re committed to making it work smoothly for all sides.
What has been hardest about building a business around that idea?
Saying no! As a start-up CEO, I know budgets are not infinite, targets are ambitious, and doing this work properly is expensive. So the hardest moments are the ones where money is on the table, but taking it would mean violating our own rules on how data workers get compensated, or how we deal with the stewards we partner with. We have turned that work down, and we will keep turning it down, because the entire proposition collapses the moment we start behaving like everyone else. The other hard part is that there is no silver bullet or magical tech solution for making human agency and value exchange practical at scale. It takes a triple stack of technical features, legal safeguards, and real community engagement, all reinforcing one another across thousands of datasets and multiple jurisdictions. We have an R&D team building automated checks off the back of months of very human review work, counsel in different countries keeping everyone in binding contracts, take-down processes, and a user community that flags things they spot out in the wild. None of that gets us to a zero chance of bad actors; but that is not a reason to give up! It is a reason we keep building better tools.
What do you think the AI industry is getting wrong right now?
The Mozilla Data Collective launch video is like three minutes long. But we filmed for three hours. I have a couple of articles posted online, and maybe 30 drafts sitting in my files. The industry needs to get beyond the internet because it is a tiny tiny tiny fraction of the data that is out there. It is also the source being legally squeezed from every direction. For thousands of languages, accents, variants, culturally specific images: they are not sitting on the open web waiting to be scraped. The only way to get that data is to establish ongoing, trustworthy partnerships with the people who have it. Not through obfuscating, or race-to-the-bottom pricing but through working in good faith with trusted intermediaries and showing up as an entity that is actually going to build things with and for communities, and enable them to build for themselves. The top 10 languages of the internet account for 80% of the content and anything beyond that is going to require new approaches.
Where do you see the biggest opportunity in AI that others are missing?
The scarce resource is not really compute or clever architecture; it is clean, abundant, contextualised, consentful data, and that is still the real bottleneck for building multimodal models that people outside of the US and UK can actually trust. We have had Hazaragi literature from Afghanistan, oral histories in Mada from Cameroon, Romansh newspapers from Switzerland, and Elders sharing Ekpeye folktales in Nigeria. AI tech company founders keep telling me their product is brilliant… as long as you speak English… like you’re from Surrey. That is fine for Beta! But when they go-to-market in new countries, their product needs to step up its multilingual, multicultural performance. But the deeper opportunity is a business insight that people sometimes mistake as a ‘soft’ or abstract ethics thing: aka, that our model deliberately involves not being greedy. I hold that this is actually a competitive advantage. A community that trusts you, that can see how its work is being used, and who gets real benefit back, that is a business asset. We are building a self-service alternative, one where communities keep 100% of what they charge and we take a 5% platform fee that just covers our costs. It’s so rare in tech companies that you get a perfectly aligned model, where the product, the growth, the social impact and the revenue strategies all sit together in harmony. You’re not having to make strategy compromises because your poker platform only stays in business by helping third parties sell jackets. If we get this moment right – unlock data abundance in a way that also benefits communities – we can all have AI that is actually worth using.
Artificial Intelligence needs more than scraped internet data. Mozilla Data Collective, a fairer marketplace for AI data, is building a model for consent, fair value exchange, and community control – while helping AI companies access richer, more diverse, and trustworthy datasets. Entrepreneur UK talks to founder EM Lewis-Jong to find out more…
What did you see happening with AI that made you want to build Mozilla Data Collective?
I was watching AI being built on a diet of indiscriminately scraped web content, which meant models often inherited the internet’s worst habits: its prejudices, its language coverage, its homogenisation of culture. I spent half a decade on Common Voice, and it taught me that, when given the chance, people will show up to help build technology that really sees them, and works for the people and communities they care about. There were also a lot of organisations, and people, who were asking tough governance questions about how to make sure that dataset owners and stewards – be they museums or podcasters or tech start ups – were benefitting from tech companies using their datasets. What they wanted that benefit to look like, was varied! Sometimes compensation (e.g. large western labs should pay for a commercial license), sometimes transparency (we want you to tell us what you’re going to do with that data), or sometimes in-kind exchanges (we’ll give you the data for free if you let us use the model for free). What I saw was a structural problem that needed a structural fix: human agency and fair value exchange, built right throughout the AI data supply chain. A platform that could make CHOICE really easy for people without big tech teams or legal teams. Not an open repository – downloaders need quality, curation, and compliance. And not a closed marketplace – uploaders want transparency, collective bargaining and self-service flows. That is the gap Mozilla Data Collective exists to fill.
What is broken about the way AI companies source and use training data?
To start, not all AI companies have behaved the same in this regard, but often, I think it’s fair to say that the default has been extractive, and skips the important questions, like where did this dataset come from, and did the people in it actually agree? How are they benefitting from their dataset being used in commercial products? Is it a fair value exchange? A suspiciously cheap dataset is usually cheap because somebody upstream went unpaid or underpaid. Take speech synthesis: a student might sell a perpetual, unlimited clone of their voice for a few hundred quid, not knowing that a professional voice actor has unionised to charge thousands for the same thing, or that platforms will then rent that voice out to enterprises, on six- and seven-figure contracts. That model of either scraping, or farming out precarious gig work as the way to power AI is getting increasingly brittle: legally, reputationally, and commercially. The regulatory floor beneath it is rising quickly – just look at Article 53 of the EU AI Act – companies now have to account for how they got their data.