2021/01/03 by Dom Huh, Huh, Dom
Business, Management and Accounting · Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Financial Distress and Bankruptcy Prediction #Imbalanced Data Classification Techniques #Machine Learning (cs.LG) #Machine Learning and Data Classification
paper · pdf · doi:10.48550/arxiv.2101.00728
openalex publication_date 2021/01/03 · openalex created_date 2021/02/15 · openalex updated_date 2026/07/28
Given the inherent class imbalance issue within student performance datasets, samples belonging to the edges of the target class distribution pose a challenge for predictive machine learning algorithms to learn. In this paper, we introduce a general framework for synthetic embedding-based data generation (SEDG), a search-based approach to generate new synthetic samples using embeddings to correct the detriment effects of class imbalances optimally. We compare the SEDG framework to past synthetic data generation methods, including deep generative models, and traditional sampling methods. In our results, we find SEDG to outperform the traditional re-sampling methods for deep neural networks and perform competitively for common machine learning classifiers on the student performance task in several standard performance metrics.