2022/09/25 by Vittorio Giammarino, Giammarino, Vittorio, Meyer, Andrew J +1
Computer Science · Engineering · #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Mobile Crowdsensing and Crowdsourcing #Reinforcement Learning in Robotics #Robot Manipulation and Learning #Robotics (cs.RO) #Systems and Control (eess.SY) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2209.12350
openalex publication_date 2022/09/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We focus on an unloading problem, typical of the logistics sector, modeled as a sequential pick-and-place task. In this type of task, modern machine learning techniques have shown to work better than classic systems since they are more adaptable to stochasticity and better able to cope with large uncertainties. More specifically, supervised and imitation learning have achieved outstanding results in this regard, with the shortcoming of requiring some form of supervision which is not always obtainable for all settings. On the other hand, reinforcement learning (RL) requires much milder form of supervision but still remains impracticable due to its inefficiency. In this paper, we propose and theoretically motivate a novel Unsupervised Reward Shaping algorithm from expert's observations which relaxes the level of supervision required by the agent and works on improving RL performance in our task.