AskVantage

AI and intelligence

Black box AI’s learn to express themselves so researchers can read their minds

Some AI's are black boxes which means people don't know how they do what they do, or how they reach their decisions, and that's a major issue for different industry use cases.

Key takeaways

  • The Artificial Intelligence (AI) behind self driving cars, medical image analysis and other computer vision applications rely on what's called Deep Neural Networks, or DNN’s for short.
  • By modifying the reasoning process behind the AI’s predictions the team were able to troubleshoot the networks and understand whether they were trustworthy.
  • The team's work appeared Dec. 7 in the journal Nature Machine Intelligence.
Cite or link to this article

Griffin, M. (2020) 'Black box AI’s learn to express themselves so researchers can read their minds', 311 Institute, 22 December. Available at: https://www.311institute.com/black-box-ais-learn-to-express-themselves-so-researchers-can-read-their-minds/ (Accessed: 1 October 2026).

The Artificial Intelligence (AI) behind self driving cars, medical image analysis and other computer vision applications rely on what's called Deep Neural Networks, or DNN’s for short. Loosely modelled on the brain DNN’s consist of layers of interconnected "neurons" - mathematical functions that send and receive information like the neurons in the human brain that "fire" like neurons in response to features of the input data.

The first layer in a DNN processes a raw data input, such as pixels in an image, then passes that information to the next layer above, triggering some of those neurons, which then pass a signal to even higher layers until eventually the AI figures out what it’s looking at.

But here's the problem, says Duke University computer science professor Cynthia Rudin: "We can input, say, a medical image, and observe what comes out the other end, for example 'this is a picture of a malignant lesion', but it's hard to know what happened in between."

It's what's known as AI’s black box problem and if you want to know how an AI made the decision it did, for whatever reason, then it’s one of, if not the, fields biggest problem.

Furthermore, as AI is increasingly used to diagnose everything from disease to take out hostile targets for the military, whether it’s in real life or in the cyber world, not knowing why an AI did what it did so you can query it or validate it is a pressing problem that has to be solved sooner rather than later.

"The problem with deep learning models is they're so complex that we don't actually know what they're learning or how they’re applying it," said Zhi Chen, a Ph.D. student in Rudin's lab at Duke. "They can often leverage and use information we don't want them to, and their reasoning processes can be completely wrong."

Now though, as the field of so called Explainable AI, where people try read the minds of these AI’s, by using anything from cell biology, oddly, to different visualisation techniques, takes root, Rudin, Chen and Duke undergraduate Yijie Bei have come up with a way to address this black box issue. By modifying the reasoning process behind the AI’s predictions the team were able to troubleshoot the networks and understand whether they were trustworthy.

Most methods in this field often attempt to uncover what led a computer vision system to the right answer after the fact, by pointing to the key features or pixels that identified an image: "The growth in this chest X-ray was classified as malignant because, to the model, these areas are critical in the classification of lung cancer." But such approaches don't reveal the network's reasoning, just where it was looking.

The Duke team tried a different tack. Instead of attempting to account for a network's decision-making on a post hoc basis, their method trains the network to show its work by expressing its understanding about concepts along the way. Their method works by revealing how much the network calls to mind different concepts to help decipher what it sees.

"It disentangles how different concepts are represented within the layers of the network," Rudin said.

Given an image of a library, for example, the approach makes it possible to determine whether and how much the different layers of the neural network rely on their mental representation of "books" to identify the scene.

The researchers found that, with a small adjustment to a neural network, it is possible to identify objects and scenes in images just as accurately as the original network, and yet gain substantial interpretability in the network's reasoning process.

"The technique is very simple to apply," Rudin said.

The method controls the way information flows through the network. It involves replacing one standard part of a neural network with a new part. The new part constrains only a single neuron in the network to fire in response to a particular concept that humans understand. The concepts could be categories of everyday objects, such as "book" or "bike." But they could also be general characteristics, such as such as "metal," "wood," "cold" or "warm." By having only one neuron control the information about one concept at a time, it is much easier to understand how the network "thinks."

The researchers tried their approach on a neural network trained by millions of labelled images to recognize various kinds of indoor and outdoor scenes, from classrooms and food courts to playgrounds and patios. Then they turned it on images it hadn't seen before. They also looked to see which concepts the network layers drew on the most as they processed the data.

Chen pulls up a plot showing what happened when they fed a picture of an orange sunset into the network. Their trained neural network says that warm colors in the sunset image, like orange, tend to be associated with the concept "bed" in earlier layers of the network. In short, the network activates the "bed neuron" highly in early layers. As the image travels through successive layers, the network gradually relies on a more sophisticated mental representation of each concept, and the "airplane" concept becomes more activated than the notion of beds, perhaps because "airplanes" are more often associated with skies and clouds.

It's only a small part of what's going on, to be sure. But from this trajectory the researchers are able to capture important aspects of the network's train of thought.

The researchers say their module can be wired into any neural network that recognizes images. In one experiment, they connected it to a neural network trained to detect skin cancer in photos.

Before an AI can learn to spot melanoma, it must learn what makes melanomas look different from normal moles and other benign spots on your skin, by sifting through thousands of training images labelled and marked up by skin cancer experts.

But the network appeared to be summoning up a concept of "irregular border" that it formed on its own, without help from the training labels. The people annotating the images for use in artificial intelligence applications hadn't made note of that feature, but the machine did.

"Our method revealed a shortcoming in the dataset," Rudin said. Perhaps if they had included this information in the data, it would have made it clearer whether the model was reasoning correctly. "This example just illustrates why we shouldn't put blind faith in black box models with no clue of what goes on inside them, especially for tricky medical diagnoses," Rudin said.

The team's work appeared Dec. 7 in the journal Nature Machine Intelligence.

FAQ

Why does this matter?

Some AI's are black boxes which means people don't know how they do what they do, or how they reach their decisions, and that's a major issue for different industry use cases.

Matthew Griffin

About the author

Matthew Griffin Founder, 311 Institute

Matthew Griffin is a multi-award winning Futurist and expert in Disruption and Innovation, Geopolitics, Leadership, and Technology, who NASA have described as a "walking encyclopaedia of the future" and a "futurist Polymath."

Read full bio

Matthew Griffin is a multi-award winning Futurist and expert in Disruption and Innovation, Geopolitics, Leadership, and Technology, who NASA have described as a "walking encyclopaedia of the future" and a "futurist Polymath." 15-time best selling author of the "Codex of the Future" series, Matthew is the Founder and Futurist in Chief of the 311 Institute, a global Futures and Deep Futures advisory firm working with royal households, world leaders, G7, G20, and G77 governments, NGOs, and multi-national mid and mega cap firms to help them explore, shape, and lead the next 50 years of business and society.

An award-winning YouTube creator with over a million followers, with an unrivalled global reach and impact, Matthew is a highly sought-after international keynote speaker, lecturer, and mentor who collaborates with global leaders through the United Nations Alliance of Civilizations (UNAOC) and United Nations General Assembly (UNGA) to shape pivotal initiatives such as the UN’s AI for Humanity program, the United Nations Conference of the Parties (UN COP), and the World Economic Forum in Davos.

As the former Global Head of Cloud, National Security, and Enterprise Sales for companies including Atos, Dell-EMC, and IBM, Matthew has a proven track record of building multi-billion dollar business units and turning failing divisions into market leaders. His ability to identify, analyse, and communicate the implications of hundreds of emerging technologies and trends is unparalleled, and his insights are trusted by many of the world’s most respected organisations, including ABB, Accenture, Adidas, AON, ARM, BCG, Centrica, Citi, Coca-Cola, Dentons, Deloitte, Dow Jones, EY, Google, KPMG, Lego, Legal & General, LinkedIn, Microsoft, PepsiCo, Qualcomm, RWE, Samsung, Siemens AG and Siemens Energy, T-Mobile, UBS, VISA, Walmart, Workday, Worldpay and many others.

Regularly featured in the global media including the AP, BBC, Bloomberg, CNBC, Discovery, Forbes, Khaleej Times, Telegraph, TIME, ViacomCBS, WIRED, and the WSJ, Matthews mission is to help organisations create a fair and sustainable future whose benefits are shared by everyone irrespective of their ability, background, or circumstances.

What future do you need to see?

Choose one to get started on AI and intelligence and the future of your organisation.

Where should Matthew reply?

Takes 30 seconds. No obligation. Matthew replies quickly. Privacy

Tag Cloud

Starburst opens that technology on the interactive 311 Starburst.

Sources and further reading

  1. Deep Neural Networks en.wikipedia.org

Source: first published by the 311 Institute on 22 December 2020. Cite as: Griffin, M. (2020). Black box AI’s learn to express themselves so researchers can read their minds. 311 Institute. https://www.311institute.com/black-box-ais-learn-to-express-themselves-so-researchers-can-read-their-minds/

You are welcome to quote this article with credit and a link to the original.

Book a Keynote