Hi. We are Group 37.

And this is our project, An Analysis on the Supposed Rise in EDSA Revolution dis/misinformation during the 2022 election period. As data science students from the University of the Philippines, we aim to gather pertinent data about the allegations involving the EDSA revolution from tweets spanning from 2016-2022. Using these data and our knowledge in machine learning, we hope to make an analysis that can help elucidate the situation about the rampant spread of mis/disinformation on Twitter.

Data Science Team

  • Vincent Angelo Dispo, WFU
  • Kristina Cassandra Castañeda, WFX
  • Loridge Anne Gacho, WFX

Here's an overview of our project.

Motivation

In light of the recent change in the EDSA holiday's date, we are interested in knowing how people perceive the historical significance of the EDSA revolution and how the rise of mis/disinformation has affected the amount of skepticism about its legitimacy and validity.

Problem

As social media platforms become much accessible to the mass people, it has become a means for some people to distribute wrong information for their own benefit or sometimes they are just misinformed too. This has caused a rise in misinformation and disinformation regarding many fields including medicine, politics, and even history itself.

Problem Formulation

The research problem and hypothesis are as shown below.

Research Problem

Is there an observed rise in EDSA Revolution dis/misinformation during the 2022 election period?

Hypothesis

A rise in EDSA Revolution dis/misinformation during the 2022 election period is observed.

Null Hypothesis

A rise in EDSA Revolution dis/misinformation during the 2022 election period is not observed or not significant.

Action Plan

Examine the frequency and posting times of tweets that spread false information about the EDSA Revolution.

We mined the internet for fake news data.

After the formulation of our topic, we scoured through Twitter to look for tweets containing mis/disinformation

Mode of Collection

We collected our data manually by using Twitter's advanced search and browsing through all tweets looking for potential mis/disinformation. After collecting all these suspected tweets, we manually picked out those tweets that really give mis/disinformation, we then manually encoded the tweet data to our dataset one by one.

Dates of Collection

We collected tweets on two batches. The first batch or the first 100 tweets are collected from March 22 to March 30 of the year 2023. The second batch on the other hand, was all collected on April 18, 2023.

Data Selection Criteria

The data collected are fact-checkable and not mere opinions. They adhere to the following standards:
Is it a statement of fact?
Is it relevant?
Is it feasible to verify using reliable sources?
In order to gather tweets on the subject, we looked into the topic of "EDSA revolution legitimacy and the alleged downfall of the Philippines after it." To do this, we used keywords like "fake edsa," "EDSA revolution economy," "edsa fake revolution," and "EDSA, peke." To be considered relevant to the issue, tweets had to be posted between 2016 and 2022. Additionally, we ensured that each tweet is distinct, therefore we do not include retweets as unique tweets.

Sample Size Collected

We collected a total of 442 sample tweets.

Data Features

These are the following data features in our dataset:
ID
This column contains each unique ID for the collected tweets.
Timestamp
This column contains the date and time the tweet was collected.
Group
This column contains our group number from our CS 132 class.
Category
This column contains the general category of the tweets we are collecting. In our case this is “AQNO”.
Topic
This column contains the subject of the collected tweet with mis/disinformation.
Keywords
This column contains the keywords we used to explore the mis/disinformation found in Twitter.
Account handle
This column contains the unique username of each collected tweet. It always begins with the “@” symbol.
Account name
This column contains the user’s account name.
Account bio
This column contains the user’s bio, which is a small public summary about the user’s Twitter profile picture.
Account type
This column refers to the user’s identity. It contains the categories “Identified” and “Anonymous”.
Joined
This column contains the month and year the user has joined Twitter.
Following
This column contains the user’s following count.
Followers
This column contains the user’s follower count.
Location
This column contains the user’s location.
Tweet
This column contains the tweet posted by the user.
Tweet Type
This column may contain the categories “Text”, “Image”, “Video”, “URL”, “Reply”, and “Quote Tweet”. This column may contain more than one category.
Date posted
This contains the day, month, year and time the tweet was posted.
Content type
This column may contain the categories “Emotional”, “Rational” and “Transactional”. This column may contain more than one category.
Likes
This column contains the amount of likes the tweet has garnered.
Replies
This column contains the amount of replies the tweet has garnered.
Retweets
This column contains the amount of retweets the tweet has garnered.
Rating
This column may contain the categories “FAKE”, “FALSE”, “MISLEADING”, “UNPROVEN”, “INACCURATE”, “NEED CONTEXT” and “FLIPFLOP”.
Reasoning
This column contains the reason why the tweet is classified as mis/disinformation.

Data Exploration

Now let's take a look of how we processed our data from preprocessing to data visualization

Here's what we found out.

We applied statistical methods and machine learning methods applicable to time series analysis in order to help us explore the data. This includes peak detection, change point detection and linear regression.

Peak Detection

Time Series analysis examines a series of data points over a period time. In this study, we examine data from 2016-2022. The figure above highlights the peaks observed in our data.

Change Point Detection

According to our findings, there is a significant peak in number of tweets around 2021-02, 2022-02, and 2022-06. Furthermore, a change point is observed in the Change Points plot. We took note of the dates within each timeframe.

Based on our research and observations, these dates are relevant to significant events that relate to our research question. The observable events from the dates are the following:

  1. Bongbong Marcos declared his candidacy for president[1]: In October 2021, Bongbong Marcos filed his candidacy for the Philippine presidency in 2022. This sparked controversy among the public as can be seen on various social media platforms. This prompted mention of the EDSA Revolution among Twitter users.
  2. EDSA Revolution Anniversary: In commemoration of the historical event, the People Power Revolution is celebrated as a holiday every year. Therefore, it resurfaces every February in remembrance. Our findings show that it garnered the most engagement in 2021 and 2022.
  3. The 2022 campaign period[2]: The presidential election took place on May 9, 2022, but the campaign began three months prior, in February 2022. Social media was used to resurface issues, including the EDSA Revolution, during the campaign to support the people's preferred candidates. As a result, more tweets were posted during this time period.
  4. Premiere of 'Maid In Malacañang'[3]: In July 2022, Daryll Yap announced that his film 'Maid in Malacañang' would premiere in August 2022. The announcement received a generally negative response, with certain scenes spurring controversy over misrepresentations of history. Meanwhile, it was promoted in Twitter by a number of people who supported the film.

Linear Regression

The graph above shows the result of our segmented linear regression with respect to the number of tweets. The data is segmented for actual tweets regarding EDSA misinformation (1) before Bongbong Marcos declared his candidacy (blue scatter plots) and the (2) time period after announcing his candidacy up to the day of the election (green scatter plots). Segmented Regression 1 is the predicted linear regression for time period 1 and Segmented Regression 2 is the predicted linear regression for time period 2.
The statistical model used is t-test between the slope values of the linear predictors for the two time periods to test for a significant change in the number of tweets between the said period. T-test is also used on the mean number of tweets on the said periods to test whether there is a significant rise to the tweets. The t-test returned a p-value of 0.0138, and with a significance level of 0.05, we accept the hypothesis.

Conclusion

Based on the findings, we accept our hypothesis given that the 2022 election and campaign period is among the events with a significant rise in EDSA Revolution dis/misinformation. This implies that the EDSA Revolution resurfaced in social media, particularly Twitter, during the 2022 election and provoked the controversy as a result of people spreading false information about the historical event to favor their candidate.

Behind the study

Group Members


Kristina Castaneda

Hello! I'm Kristina. I am a 2nd-year student in the University of the Philippines, Diliman, actively working towards earning a Bachelor's degree in Computer Science. Through this Data Science course handled by Professor Paul Regonia, I hope to enhance my skills in machine learning methods while being able to contribute to society by showing the influence of the rampant spread of EDSA Revolution misinformation around Philippine Twitter.

Vincent Dispo

Hello! I'm Vincent. I am a 2nd-year standing BS Computer Science student in the University of the Philippines Diliman. I have a great interest in learning about software programming and also web development. This study made me learn many things about Data Science, especially in the process of data preprocessing, presentation, data modeling and more.I want to be able to help the community through means related to my field like making software to help people, and doing studies like this that expose the rampand spread of misinformation in our social media platforms.

Loridge Gacho

Hello! My name is LA, and I'm a second-year student at the University of the Philippines, Diliman. I'm a current undergraduate student in the BS Computer Science program. This study has contributed in improving my machine learning skills and I look forward to learning more in the future. In addition to working with data, I want to improve my web programming abilities. I aspire to be able to contribute to society by working on projects that fight disinformation utilizing the skills I am developing.