Multilingual Cyberbullying Detection System

Pawar, Rohit S.

Multilingual Cyberbullying Detection System

Files

MULTILINGUAL CYBERBULLYING DETECTION SYSTEM.pdf (1.13 MB)

If you need an accessible version of this item, please email your request to digschol@iu.edu so that they may create one and provide it to you.

Date

2019-05

Authors

Pawar, Rohit S.

Language

American English

Committee Chair

Raje, Rajeev R.

Committee Members

Tuceryan, Mihran
Durresi, Arjan

Degree

M.S.

Degree Year

2019

Grantor

Purdue University

Abstract

Since the use of social media has evolved, the ability of its users to bully others has increased. One of the prevalent forms of bullying is Cyberbullying, which occurs on the social media sites such as Facebook©, WhatsApp©, and Twitter©. The past decade has witnessed a growth in cyberbullying – is a form of bullying that occurs virtually by the use of electronic devices, such as messaging, e-mail, online gaming, social media, or through images or mails sent to a mobile. This bullying is not only limited to English language and occurs in other languages. Hence, it is of the utmost importance to detect cyberbullying in multiple languages. Since current approaches to identify cyberbullying are mostly focused on English language texts, this thesis proposes a new approach (called Multilingual Cyberbullying Detection System) for the detection of cyberbullying in multiple languages (English, Hindi, and Marathi). It uses two techniques, namely, Machine Learning-based and Lexicon-based, to classify the input data as bullying or non-bullying. The aim of this research is to not only detect cyberbullying but also provide a distributed infrastructure to detect bullying. We have developed multiple prototypes (standalone, collaborative, and cloud-based) and carried out experiments with them to detect cyberbullying on different datasets from multiple languages. The outcomes of our experiments show that the machine-learning model outperforms the lexicon-based model in all the languages. In addition, the results of our experiments show that collaboration techniques can help to improve the accuracy of a poor-performing node in the system. Finally, we show that the cloud-based configurations performed better than the local configurations.

Description

Indiana University-Purdue University Indianapolis (IUPUI)

Keywords

Distributed Computing, Natural Language Processing, Machine Learning, Indian Languages, Cloud

Rights

Attribution 3.0 United States

Type

Thesis

Permanent Link

https://hdl.handle.net/1805/18942
http://dx.doi.org/10.7912/C2/2364

Collections

Computer & Information Science Department Theses and Dissertations

Full item page