Modular manufacturing systems optimized faster with inverse models

Distributed Optimization of Modular Production Systems using Model-based Reinforcement Learning with Inverse Models

Artificial IntelligenceMachine Learning

Summary

Modular factories can be hard to control because they have many separate parts working together in different ways. The authors introduced a new method that helps robots and machines learn better how to run these modular systems by teaching the control software how actions translate into changes. They use a special kind of learning called reinforcement learning but improve it by separating how machine actions link to results. This makes training faster and the systems perform better in practice.

What this means in practice

  • For manufacturing engineers: Improve control and efficiency of flexible production lines with faster training of automated controllers using inverse model integration.
  • For industrial robotics teams: Enhance modular robot setups by applying model-based reinforcement learning that better separates actuation and state dynamics.
  • For factory automation vendors: Develop control software products for modular production systems that train more efficiently and optimize performance.$Commercial implications: Enables building and selling advanced reinforcement learning control software tailored for modular flexible manufacturing.

Authors

Andreas Schwung, Steve Yuwono, Sofiene Lassoued, Dorothea Schwung

Abstract

This paper presents a novel approach for data-driven self-learning control of highly flexible, modular manufacturing systems. Specifically, we employ a novel framework for model-based reinforcement learning which introduces approximate inverse process models within the training of reinforcement policies. This approach disentangles the learning of actuation dynamics and the dynamics in state space, resulting in RL-based training solely within the task space. We propose a lightweight feedforward architecture for approximate inverse models and integrate them within the policy network of standard RL algorithms. We apply the approach to a laboratory modular production testbed with heterogeneous production modules. The results underline the efficiency improvements for modular manufacturing units in terms of both performance and training speed, particularly for off-policy algorithms.