2019/11/22 by Zhiliang Chen, Chen, Zhiliang
Computer Science · Decision Sciences · #Auction Theory and Applications #Blockchain Technology Applications and Security #Computer Science and Game Theory (cs.GT) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multiagent Systems (cs.MA) #Privacy-Preserving Technologies in Data #cs.GT #cs.LG #cs.MA
paper · pdf · doi:10.48550/arxiv.1911.11555
Undergraduate Thesis at National University of Singapore, School of Computer Science; in preparation for publication in 2020
arxiv created 2019/11/22 · openalex publication_date 2019/11/22 · arxiv updated 2019/11/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
High performance machine learning models have become highly dependent on the availability of large quantity and quality of training data. To achieve this, various central agencies such as the government have suggested for different data providers to pool their data together to learn a unified predictive model, which performs better. However, these providers are usually profit-driven and would only agree to participate inthe data sharing process if the process is deemed both profitable and fair for themselves. Due to the lack of existing literature, it is unclear whether a fair and stable outcome is possible in such data sharing processes. Hence, we wish to investigate the outcomes surrounding these scenarios and study if data providers would even agree to collaborate in the first place. Tapping on cooperative game concepts in Game Theory, we introduce the data sharing process between a group of agents as a new class of cooperative games with modified definition of stability and fairness. Using these new definitions, we then theoretically study the optimal and suboptimal outcomes of such data sharing processes and their sensitivity to perturbation.Through experiments, we present intuitive insights regarding theoretical results analysed in this paper and discuss various ways in which data can be valued reasonably.