2018/05/12 by Juan Cruz Barsce, Barsce, Juan Cruz, Jorge A. Palombarini +3
Computer Science · #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG)
paper · pdf · doi:10.48550/arxiv.1805.04748
openalex publication_date 2018/05/12 · openalex created_date 2023/02/18 · openalex updated_date 2026/07/28
With the increase of machine learning usage by industries and scientific\ncommunities in a variety of tasks such as text mining, image recognition and\nself-driving cars, automatic setting of hyper-parameter in learning algorithms\nis a key factor for achieving satisfactory performance regardless of user\nexpertise in the inner workings of the techniques and methodologies. In\nparticular, for a reinforcement learning algorithm, the efficiency of an agent\nlearning a control policy in an uncertain environment is heavily dependent on\nthe hyper-parameters used to balance exploration with exploitation. In this\nwork, an autonomous learning framework that integrates Bayesian optimization\nwith Gaussian process regression to optimize the hyper-parameters of a\nreinforcement learning algorithm, is proposed. Also, a bandits-based approach\nto achieve a balance between computational costs and decreasing uncertainty\nabout the Q-values, is presented. A gridworld example is used to highlight how\nhyper-parameter configurations of a learning algorithm (SARSA) are iteratively\nimproved based on two performance functions.\n