2021/04/12 by Rebecca West, Khalifeh Al Jadda, West, Rebecca +7
Computer Science · #Advanced Text Analysis Techniques #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Text and Document Classification Technologies #Web Data Mining and Analysis #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2104.05504
arxiv created 2021/04/12 · openalex publication_date 2021/04/12 · arxiv updated 2021/04/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
For e-commerce companies with large product selections, the organization and grouping of products in meaningful ways is important for creating great customer shopping experiences and cultivating an authoritative brand image. One important way of grouping products is to identify a family of product variants, where the variants are mostly the same with slight and yet distinct differences (e.g. color or pack size). In this paper, we introduce a novel approach to identifying product variants. It combines both constrained clustering and tailored NLP techniques (e.g. extraction of product family name from unstructured product title and identification of products with similar model numbers) to achieve superior performance compared with an existing baseline using a vanilla classification approach. In addition, we design the algorithm to meet certain business criteria, including meeting high accuracy requirements on a wide range of categories (e.g. appliances, decor, tools, and building materials, etc.) as well as prioritizing the interpretability of the model to make it accessible and understandable to all business partners.