Implementation of Data Analysis and Document Summarization in Social Media Data Using R and Python

Deepika Raman, S. Jayalakshmi, K. ARUMUGAM, A.V. Athipathi Raj, D. Balaji, R. Brightsingh · 2022 4th International Conference on Inventive Research in Computing Applications (ICIRCA) · 2022

In today's internet environment, people mostly use smartphones, laptops, and other peripheral devices as communication devices to connect with people. The emerging different uses of social media platforms also necessitates a significant research attention in this field. Analysis of data from various social media resources, analysis of each user's activity, analysis of social media post content, and analysis of social media discussion message are recently gaining a significant research interest. These statistical analysis were performed based on the summary data and it can be classified as positive and negative. Python and R are two programming languages that may be used to perform this study. R is a programming language primarily created for performing data analysis. R is the key to unlock the door between the data challenges that needs to be solved and the answers required to reach the solution. A new strategy based on well-known hierarchical Bayesian title model is introduced in this study. Bayesians argue that the relevant information regarding decision-making and updating beliefs cannot be ignored and that hierarchical modeling has the potential to overrule classical methods in applications, where the respondents provide multiple observational data. The suggested contextual topic model will effectively detect the sentiments present in the articles that have been shared widely on social media. It might be helpful to identify any prohibited or violent content and then the data will be sent to the cloud for performing an additional review. The proposed model outperforms HLDA and LDA in document modelling, according to the quantitative assessment findings. It might be used to track out who sends how many messages and forwards a significant number of them. Furthermore, a practical example shows the implementation of the proposed model in a summary system and how it greatly increases the summarizing performance and bring it up to pace with current-generation summarizing systems. Finally, Exploratory Data Analysis (EDA) is a statistical method or approach for reviewing data sets to explain its important and major qualities utilizing visual aids to illustrate the categorised data from the dataset to be studied using R programming. Basic data pre-processing steps like null value imputation and removal of unwanted data.

Read the paper · More papers on PaperTik