Data Summarization with Social Contexts

Zhuang, Hao; Rahman, Rameez; Hu, Xia; Guo, Tian; Hui, Pan; Aberer, Karl

doi:10.1145/2983323.2983736

conference paper

Data Summarization with Social Contexts

Zhuang, Hao

•

Rahman, Rameez

•

Hu, Xia

2016

Cikm'16: Proceedings Of The 2016 Acm Conference On Information And Knowledge Management

25th ACM Conference on Information and Knowledge Management (CIKM)

While social data is being widely used in various applications such as sentiment analysis and trend prediction, its sheer size also presents great challenges for storing, sharing and processing such data. These challenges can be addressed by data summarization which transforms the original dataset into a smaller, yet still useful, subset. Existing methods find such subsets with objective functions based on data properties such as representativeness or informativeness but do not exploit social contexts, which are distinct characteristics of social data. Further, till date very little work has focused on topic preserving data summarization, despite the abundant work on topic modeling. This is a challenging task for two reasons. First, since topic model is based on latent variables, existing methods are not well-suited to capture latent topics. Second, it is difficult to find such social contexts that provide valuable information for building effective topic-preserving summarization model. To tackle these challenges, in this paper, we focus on exploiting social contexts to summarize social data while preserving topics in the original dataset. We take Twitter data as a case study. Through analyzing Twitter data, we discover two social contexts which are important for topic generation and dissemination, namely (i) CrowdExp topic score that captures the influence of both the crowd and the expert users in Twitter and (ii) Retweet topic score that captures the influence of Twitter users' actions. We conduct extensive experiments on two real-world Twitter datasets using two applications. The experimental results show that, by leveraging social contexts, our proposed solution can enhance topic-preserving data summarization and improve application performance by up to 18%.

Name

data summarization_Zhuang.pdf

Type

Preprint

Version

http://purl.org/coar/version/c_71e4c1898caa6e32

Access type

openaccess

Size

449.88 KB

Format

Adobe PDF

Checksum (MD5)

480845868a9759f406d3c9be6b3253b6