vix.ing · top · new · best · stats · spec

Network Sampling Based on NN Representatives

2014/02/07 by Miloš Kudělka, Milos Kudelka, Šárka Zehnalová +6
Computer Science · Physics and Astronomy · #Advanced Clustering Algorithms Research #Complex Network Analysis Techniques #Data Management and Algorithms #FOS: Computer and information sciences #FOS: Physical sciences #Physics and Society (physics.soc-ph) #Social and Information Networks (cs.SI) #cs.SI #physics.soc-ph

paper · pdf · doi:10.48550/arxiv.1402.1661

arxiv created 2014/02/07 · openalex publication_date 2014/02/07 · arxiv updated 2014/02/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

The amount of large-scale real data around us increase in size very quickly and so does the necessity to reduce its size by obtaining a representative sample. Such sample allows us to use a great variety of analytical methods, whose direct application on original data would be infeasible. There are many methods used for different purposes and with different results. In this paper we outline a simple and straightforward approach based on analyzing the nearest neighbors (NN) that is generally applicable. This feature is illustrated on experiments with weighted networks and vector data. The properties of the representative sample show that the presented approach maintains very well internal data structures (e.g. clusters and density). Key technical parameters of the approach is low complexity and high scalability. This allows the application of this approach to the area of big data.

Related