TheVortiq
Inteligencia Artificial

The shadow of copyright: the ethical dilemma behind AI training

Revelations regarding the use of unauthorized databases at Meta and their impact on the credibility of the European AI ecosystem

October 9, 2026 · 4 min read

A fine white mesh fabric draped over a dark surface under blue light

TL;DR: The use of millions of pirated books to train AI models challenges the ethics of big tech. The industry must move toward licensed sources to avoid legal risks and maintain its legitimacy.

A disturbing precedent for the industry

The artificial intelligence industry is going through a moment of accountability. Mediapart's revelations about Meta are not just a case of large-scale piracy; they represent a fracture in the 'open source' AI narrative that has dominated the tech discourse since 2022. The contradiction is palpable: while tech companies demand clear regulatory frameworks to protect their models, they have built the foundations of these models on the systematic plunder of others' intellectual property. This phenomenon is reminiscent of the software piracy era in the 90s or the digital media crisis against Google News, where the value generated by aggregation consistently outweighed the rights of the original creators.

The heart of the scandal: LibGen and LLaMA models

The core of the conflict lies in Library Genesis (LibGen), a repository that operates outside of commercial legality and has served as a quarry to fuel Meta's LLaMA models. The Mediapart investigation, supported by internal emails and court documents, confirms that the use of LibGen was not an anomaly, but a deliberate strategy initiated in October 2022. By processing nearly 70 terabytes of data, Meta was not only seeking volume, but linguistic quality: books, being edited and structured, offer a higher density of knowledge than the open web (Common Crawl). The technical implication is clear: the performance of current models depends on a data architecture that, under a strict interpretation of the law, could be considered 'fruit of the poisonous tree'. If the models are legally invalidated, the industry faces an unprecedented scenario of 'forced data cleansing'.

The scale of the challenge: 70 terabytes of uncertainty

The figure of 70 terabytes is revealing. To put it in context, it is an amount of data that far exceeds what any physical library could house. Historically, the 'fair use' doctrine in the U.S. has allowed data analysis for transformative purposes, but the large-scale commercialization of tools based on this content is forcing courts to re-evaluate whether a language model is a 'reader' or a 'competitor'. The industry is at a technical crossroads: the scarcity of high-quality data (tokens) is forcing companies to seek alternative sources, and the temptation to resort to repositories like LibGen is, in essence, an arms race where ethics have been sacrificed in favor of model latency and precision.

The responsibility of Meta's leadership

What differentiates this case from previous copyright disputes is the evidence of corporate knowledge. The documents suggest that Meta's leadership, under the direct supervision of Mark Zuckerberg, endorsed these practices, aware that obtaining legal licenses for millions of books would have been economically prohibitive, as well as logistically unfeasible. This revelation dismantles the shield of 'plausible deniability' often used by Big Tech. It is a reminder that, in the race for technological supremacy, legal risk is calculated simply as another operating cost within the model training budget.

The rebound effect on the European ecosystem

The case takes on particular gravity as it involves key figures who now lead AI development in Europe. The transfer of talent from Silicon Valley to the European ecosystem has brought with it not only intellectual capital, but also the 'bad practices' of the 'move fast and break things' culture. If European AI champions are building their models on the same questionable foundations as their American counterparts, the differential value of 'sovereign' and ethical AI vanishes. This poses an existential question for the European market: can Europe claim leadership in human-centric AI if its base models replicate the vices of opacity and copyright infringement that they criticize externally?

What should companies and users expect?

  • Data audits: Transparency will be the new currency. Companies will need to implement 'model cards' that detail the provenance of their datasets to mitigate reputational and legal risks.
  • Legal certainty and litigation: Companies that integrate Meta or third-party models into their workflows could be held vicariously liable. Due diligence in the selection of AI providers will be mandatory.
  • Paradigm shift towards 'Data Curation': The era of indiscriminate training is coming to an end. We will see a transition towards models trained exclusively with licensed data (such as OpenAI's agreements with Axel Springer or Reddit), which will increase operating costs but protect intellectual property.

In conclusion, AI is not a magical black box; it is a mirror of the information it consumes. If that information is the result of digital plunder, the legitimacy of the results will always be in question. The industry must transition towards a sustainable compensation model if it wants to avoid a judicial paralysis that could halt technological progress in the coming years.

Keep reading