AskVantage

Security and privacy

Major AI research breakthrough helps AI’s forget copyrighted content

Previously trying to get AI to forget anything was impossible because it retained "memories" of what it had learned. But now it might actually be able to forget ...

Key takeaways

  • When people learn things they should not know, getting them to forget that information can be tough.
  • Underscoring this issue, The New York Times recently sued OpenAI, and so did hundreds of artists, maker of ChatGPT, arguing that the AI company illegally used its articles as training data to help its chatbots generate content.
  • Image-to-image models are the primary focus of this research.
Cite or link to this article

Griffin, M. (2024) 'Major AI research breakthrough helps AI’s forget copyrighted content', 311 Institute, 31 March. Available at: https://www.311institute.com/major-ai-research-breakthrough-helps-ais-forget-copyrighted-content/ (Accessed: 1 October 2026).

One of the big problems that companies have been having when it comes to realising that their Artificial Intelligences (AI) have infringed copyright, or are being asked to forget something, perhaps as the result of a GDPR request, is that they can’t find an effective way to actually get their AIs to forget what they’ve learned … either not at all or not very effectively because “old memories” remain no matter how much you try to get them to forget or retrain them.

When people learn things they should not know, getting them to forget that information can be tough. This is also true of rapidly growing AI programs that are trained to think as we do, and it has become a problem as they run into challenges based on the use of copyright protected material and privacy issues.

To respond to this challenge, researchers at the University of Texas at Austin have developed what they believe is the first "Machine Unlearning" method applied to image-based Generative AI. This method offers the ability to look under the hood and actively block and remove any violent images or copyrighted works without losing the rest of the information in the model. The study is published on the arXiv preprint server.

This new tool becomes especially valuable when you realise that under new EU AI laws European law makers can request that AI’s that aren’t compliant to the letter of the law, whether that’s from a copyright, ethics, or even safety perspective, can be deleted which in some cases, bearing in mind companies are spending billions training their AI’s, is understandably causing some companies to freak out.

"When you train these models on such massive data sets, you're bound to include some data that is undesirable," said Radu Marculescu, a professor in the Cockrell School of Engineering's Chandra Family Department of Electrical and Computer Engineering and one of the leaders on the project.

"Previously, the only way to remove problematic content was to scrap everything, start anew, manually take out all that data and retrain the model. Our approach offers the opportunity to do this without having to retrain the model from scratch."

Generative AI models are trained primarily with data on the internet because of the unrivalled amount of information it contains. But it also contains massive amounts of data that is protected by copyright, in addition to personal information and inappropriate content.

Underscoring this issue, The New York Times recently sued OpenAI, and so did hundreds of artists, maker of ChatGPT, arguing that the AI company illegally used its articles as training data to help its chatbots generate content.

"If we want to make generative AI models useful for commercial purposes, this is a step we need to build in, the ability to ensure that we're not breaking copyright laws or abusing personal information or using harmful content," said Guihong Li, a graduate research assistant in Marculescu's lab who worked on the project as an intern at JPMorgan Chase and finalized it at UT.

Image-to-image models are the primary focus of this research. They take an input image and transform it - such as creating a sketch, changing a particular scene and more - based on a given context or instruction.

This new machine unlearning algorithm provides the ability of a machine learning model to "forget" or remove content if it is flagged for any reason without the need for retraining the model from scratch. Human teams handle the moderation and removal of content, providing an extra check on the model and ability to respond to user feedback.

Machine unlearning is an evolving branch of the field that has been primarily applied to classification models. Those models are trained to sort data into different categories, such as whether an image shows a dog or a cat.

Applying machine unlearning to generative models is "relatively unexplored," the researchers write in the paper, especially when it comes to images.

FAQ

Why does this matter?

Previously trying to get AI to forget anything was impossible because it retained "memories" of what it had learned. But now it might actually be able to forget ...

Matthew Griffin

About the author

Matthew Griffin Founder, 311 Institute

Matthew Griffin is a multi-award winning Futurist and expert in Disruption and Innovation, Geopolitics, Leadership, and Technology, who NASA have described as a "walking encyclopaedia of the future" and a "futurist Polymath."

Read full bio

Matthew Griffin is a multi-award winning Futurist and expert in Disruption and Innovation, Geopolitics, Leadership, and Technology, who NASA have described as a "walking encyclopaedia of the future" and a "futurist Polymath." 15-time best selling author of the "Codex of the Future" series, Matthew is the Founder and Futurist in Chief of the 311 Institute, a global Futures and Deep Futures advisory firm working with royal households, world leaders, G7, G20, and G77 governments, NGOs, and multi-national mid and mega cap firms to help them explore, shape, and lead the next 50 years of business and society.

An award-winning YouTube creator with over a million followers, with an unrivalled global reach and impact, Matthew is a highly sought-after international keynote speaker, lecturer, and mentor who collaborates with global leaders through the United Nations Alliance of Civilizations (UNAOC) and United Nations General Assembly (UNGA) to shape pivotal initiatives such as the UN’s AI for Humanity program, the United Nations Conference of the Parties (UN COP), and the World Economic Forum in Davos.

As the former Global Head of Cloud, National Security, and Enterprise Sales for companies including Atos, Dell-EMC, and IBM, Matthew has a proven track record of building multi-billion dollar business units and turning failing divisions into market leaders. His ability to identify, analyse, and communicate the implications of hundreds of emerging technologies and trends is unparalleled, and his insights are trusted by many of the world’s most respected organisations, including ABB, Accenture, Adidas, AON, ARM, BCG, Centrica, Citi, Coca-Cola, Dentons, Deloitte, Dow Jones, EY, Google, KPMG, Lego, Legal & General, LinkedIn, Microsoft, PepsiCo, Qualcomm, RWE, Samsung, Siemens AG and Siemens Energy, T-Mobile, UBS, VISA, Walmart, Workday, Worldpay and many others.

Regularly featured in the global media including the AP, BBC, Bloomberg, CNBC, Discovery, Forbes, Khaleej Times, Telegraph, TIME, ViacomCBS, WIRED, and the WSJ, Matthews mission is to help organisations create a fair and sustainable future whose benefits are shared by everyone irrespective of their ability, background, or circumstances.

What future do you need to see?

Choose one to get started on security and privacy and the future of your organisation.

Where should Matthew reply?

Takes 30 seconds. No obligation. Matthew replies quickly. Privacy

Tag Cloud

Starburst opens that technology on the interactive 311 Starburst.

Sources and further reading

  1. 2402.00351 arxiv.org
  2. New york times open ai microsoft lawsuit nytimes.com

Source: first published by the 311 Institute on 31 March 2024. Cite as: Griffin, M. (2024). Major AI research breakthrough helps AI’s forget copyrighted content. 311 Institute. https://www.311institute.com/major-ai-research-breakthrough-helps-ais-forget-copyrighted-content/

You are welcome to quote this article with credit and a link to the original.

Book a Keynote