Deepminds newest Embodied AI Agents learn with significantly less data
AI is running out of public data to train from, and private data is expensive, so companies want AI to learn without needing huge volumes of data - or even any at all.
Key takeaways
- One is the use of so called Embodied AI agents that can interact with the physical world and companies like Meta and Google DeepMind believe that these hold immense potential for various applications.
- The real world presents several challenges to data collection in embodied AI.
- First, physical environments are much more complex and unpredictable than the digital world.
Cite or link to this article
Griffin, M. (2024) 'Deepminds newest Embodied AI Agents learn with significantly less data', 311 Institute, 31 December. Available at: https://www.311institute.com/deepminds-newest-embodied-agents-learn-with-significantly-less-data/ (Accessed: 1 October 2026).
Today Artificial Intelligence (AI) learns by consuming huge volumes of different kinds of data, but in the future we can see a time where AI learns either using only small data sets or maybe even no data at all much the same way that baby animals “intuitively” learn how to do things such as walk, or flee from predators, using a technology called Few Shot or Zero Shot Learning which we’ve seen be successful a few times already at creating intuitive machines and general purpose robots. And there are a few different approaches that companies are embracing to achieve few or zero shot learning.
One is the use of so called Embodied AI agents that can interact with the physical world and companies like Meta and Google DeepMind believe that these hold immense potential for various applications. But the scarcity of training data remains one of their main hurdles.
To explore this researchers from Imperial College London and Google DeepMind have introduced Diffusion Augmented Agents (DAAG), a novel framework that leverages the power of Large Language Models (LLMs), Large Vision Models (VLMs) and diffusion models to enhance the learning efficiency and transfer learning capabilities of embodied agents.
The Future of Artificial Intelligence and Generative AI, by Keynote Futurist Matthew Griffin
The impressive progress in LLMs and VLMs in recent years has fuelled hopes for their application to robotics and embodied AI. However, while LLMs and LVMs can be trained on massive text and image datasets scraped from the internet, embodied AI systems on the other hand need to learn by interacting with the physical world.
The real world presents several challenges to data collection in embodied AI. First, physical environments are much more complex and unpredictable than the digital world. Second, robots and other embodied AI systems rely on physical sensors and actuators, which can be slow, noisy, and prone to failure.
The researchers believe that overcoming this hurdle will depend on making better use of the agent’s existing data and experience.
“We hypothesize that embodied agents can achieve greater data efficiency by leveraging past experience to explore effectively and transfer knowledge across tasks,” the researchers write.
Diffusion Augmented Agent (DAAG), the framework proposed by the Imperial College and DeepMind team, is designed to enable agents to learn tasks more efficiently by using past experiences and generating synthetic data.
“We are interested in enabling agents to autonomously set and score subgoals, even in the absence of external rewards, and to repurpose their experience from previous tasks to accelerate learning of new tasks,” the researchers write.
The researchers designed DAAG as a lifelong learning system, where the agent continuously learns and adapts to new tasks.
DAAG works in the context of a Markov Decision Process (MDP). The agent receives instructions for a task at the beginning of each episode. It observes the state of its environment, takes actions and tries to reach a state that aligns with the described task.
It has two memory buffers: a task-specific buffer that stores experiences for the current task and an “offline lifelong buffer” that stores all past experiences, regardless of the tasks they were collected for or their outcomes.
DAAG combines the strengths of LLMs, LVMs and diffusion models to create agents that can reason about tasks, analyze their environment, and repurpose their past experiences to learn new objectives more efficiently.
The LLM acts as the agent’s central controller. When the agent receives a new task, the LLM interprets instructions, breaks them into smaller subgoals, and coordinates with the VLM and diffusion model to obtain reference frames for achieving its goals.
To make the best use of its past experience, DAAG uses a process called Hindsight Experience Augmentation (HEA), which uses the VLM and the diffusion model to augment the agent’s memory.
First, the VLM processes visual observations in the experience buffer and compares them to the desired subgoals. It adds the relevant observations to the agent’s new buffer to help guide its actions.
If the experience buffer does not have relevant observations, the diffusion model comes into play. It generates synthetic data to help the agent “imagine” what the desired state would look like. This enables the agent to explore different possibilities without physically interacting with the environment.
“Through HEA, we can synthetically increase the number of successful episodes the agent can store in its buffers and learn from,” the researchers write. “This allows to effectively reuse as much data gathered by the agent as possible, substantially improving efficiency especially when learning multiple tasks in succession.”
The researchers describe DAAG and HEA as the first method “to propose an entire autonomous pipeline, independent from human supervision, and that leverages geometrical and temporal consistency to generate consistent augmented observations.”
The researchers evaluated DAAG on several benchmarks and across three different simulated environments, measuring its performance on tasks such as navigation and object manipulation. They found that the framework delivered significant improvements over baseline reinforcement learning systems.
For example, DAAG-powered agents were able to successfully learn to achieve goals even when they were not provided with explicit rewards. They were also able to reach their goals more quickly and with less interaction with the environment compared to agents that did not use the framework. And DAAG is better suited to effectively reuse data from previous tasks to accelerate the learning process for new objectives.
The ability to transfer knowledge between tasks is crucial for developing agents that can learn continuously and adapt to new situations. DAAG’s success in enabling efficient transfer learning in embodied agents has the potential to pave the way for more robust and adaptable robots and other embodied AI systems.
“This work suggests promising directions for overcoming data scarcity in robot learning and developing more generally capable AI agents,” the researchers write.
FAQ
Why does this matter?
AI is running out of public data to train from, and private data is expensive, so companies want AI to learn without needing huge volumes of data - or even any at all.

About the author
Matthew Griffin Founder, 311 Institute
Matthew Griffin is a multi-award winning Futurist and expert in Disruption and Innovation, Geopolitics, Leadership, and Technology, who NASA have described as a "walking encyclopaedia of the future" and a "futurist Polymath."
Read full bio
Matthew Griffin is a multi-award winning Futurist and expert in Disruption and Innovation, Geopolitics, Leadership, and Technology, who NASA have described as a "walking encyclopaedia of the future" and a "futurist Polymath." 15-time best selling author of the "Codex of the Future" series, Matthew is the Founder and Futurist in Chief of the 311 Institute, a global Futures and Deep Futures advisory firm working with royal households, world leaders, G7, G20, and G77 governments, NGOs, and multi-national mid and mega cap firms to help them explore, shape, and lead the next 50 years of business and society.
An award-winning YouTube creator with over a million followers, with an unrivalled global reach and impact, Matthew is a highly sought-after international keynote speaker, lecturer, and mentor who collaborates with global leaders through the United Nations Alliance of Civilizations (UNAOC) and United Nations General Assembly (UNGA) to shape pivotal initiatives such as the UN’s AI for Humanity program, the United Nations Conference of the Parties (UN COP), and the World Economic Forum in Davos.
As the former Global Head of Cloud, National Security, and Enterprise Sales for companies including Atos, Dell-EMC, and IBM, Matthew has a proven track record of building multi-billion dollar business units and turning failing divisions into market leaders. His ability to identify, analyse, and communicate the implications of hundreds of emerging technologies and trends is unparalleled, and his insights are trusted by many of the world’s most respected organisations, including ABB, Accenture, Adidas, AON, ARM, BCG, Centrica, Citi, Coca-Cola, Dentons, Deloitte, Dow Jones, EY, Google, KPMG, Lego, Legal & General, LinkedIn, Microsoft, PepsiCo, Qualcomm, RWE, Samsung, Siemens AG and Siemens Energy, T-Mobile, UBS, VISA, Walmart, Workday, Worldpay and many others.
Regularly featured in the global media including the AP, BBC, Bloomberg, CNBC, Discovery, Forbes, Khaleej Times, Telegraph, TIME, ViacomCBS, WIRED, and the WSJ, Matthews mission is to help organisations create a fair and sustainable future whose benefits are shared by everyone irrespective of their ability, background, or circumstances.
What future do you need to see?
Choose one to get started on AI and intelligence and the future of your organisation.
Sources and further reading
- Diffusion Augmented Agents arxiv.org
- Markov Decision Process en.wikipedia.org
Source: first published by the 311 Institute on 31 December 2024. Cite as: Griffin, M. (2024). Deepminds newest Embodied AI Agents learn with significantly less data. 311 Institute. https://www.311institute.com/deepminds-newest-embodied-agents-learn-with-significantly-less-data/
You are welcome to quote this article with credit and a link to the original.