2022 · Public Health

ANTi-Vax: a novel Twitter dataset for COVID-19 vaccine misinformation detection

Hayawi, K.; Shahriar, S.; Serhani, M.A.; Taleb, I.; Mathew, S.S.

Verdict

UNSCIENTIFIC

Neither the paradigm nor the study inside it passes. The verdict is unscientific.

1·Study methodology · inside the paradigm

UNSCIENTIFIC

Does not follow the scientific method within Institutional Consensus.

2·Paradigm · Institutional Consensus

UNSCIENTIFIC

Fails the tests. The worse layer decides.

How the verdict is decided

The verdict grades both layers: the paradigm this study assumes, and how the study was carried out inside it. Both must be scientific for the whole thing to be scientific. A clean method inside an unscientific paradigm is unscientific. An unscientific method inside a better paradigm is unscientific. A split on either layer is mixed. The worse layer decides.

1·Study methodology

Did this study test its claim with methods that are independent, falsifiable, and non-circular?

UNSCIENTIFIC

This paper claims to build a labeled Twitter dataset for detecting COVID-19 vaccine misinformation and train machine learning models to classify it. All three tests fail.

Independently Verifiable

Fail

The tweets are visible text, but the misinformation label is assigned by matching them against an answer key from the CDC and Public Health.

Falsifiable

Fail

No result could disprove the claim because the labeling scheme presupposes that the institutional sources define what counts as misinformation.

Non-Circular

Fail

The definition of misinformation comes from the same institutional sources used to validate the labeling.

Why

They collected 15 million tweets about vaccines. That part is real. Then they took common statements like the vaccine can alter DNA or the vaccine causes infertility, got them from the CDC and Public Health, and called them myths. They read through tweets and labeled any tweet containing those statements as misinformation. Everything else was labeled general. The tweets are visible. The label is not. A person who does not accept the CDC position would not agree that those statements are misinformation. They would see someone making a claim about a vaccine. The study grades tweets against an answer key written by institutions, then trains a model to reproduce that grading. The model confirms the answer key. The answer key confirms the model. Nobody tested whether the statements labeled misinformation are actually false. They assumed it, built the labels from the assumption, and reported the labels as a dataset.

2·Paradigm · Institutional Consensus

Does the framework this study assumes pass the three tests?

UNSCIENTIFIC

Independently Verifiable

Fail

A tweet is visible. Calling it misinformation, conspiracy, or disordered speech requires the institutional consensus framework to interpret the text first.

Falsifiable

Fail

If a tweet disputes the label, the field calls it coded language or a dog whistle, so no classification can count against the paradigm.

Non-Circular

Fail

The institutions that define what counts as misinformation are the same institutions the paradigm uses to validate its classifications.

Why

Institutional Consensus treats text written by people as visible data, then applies its own labels to that text and calls the labels discovery. A person typing vaccine can alter DNA is visible. Calling that misinformation requires accepting that the CDC has already settled the question. The framework never tests whether the settled answer is correct. It assumes the answer, grades the population against it, and reports non-conformity as misinformation. If someone disputes the label, the field has absorbers ready. The statement is called coded language, a dog whistle, or conspiracy thinking. The paradigm cannot die because every possible response from a speaker is already classified as either compliant or disordered. There is no outcome that would make the framework conclude it was wrong about what counts as misinformation.

From the paper

Methods
To label the misinformation, some common myths regarding the COVID-19 vaccines were obtained from reliable sources including Public Health, Healthline, the Centers for Disease Control and Prevention (CDC), and the University of Missouri Health Care.
Methods
Some of the common myths and misinformation include 'The vaccine can alter DNA,' 'The vaccine can cause infertility,' 'The vaccine contains dangerous toxins,' and 'The vaccine contains tracking device.' In this process, tweets containing this common misinformation were manually read and labeled/flagged.
Methods
Tweets other than these common myths were considered not misinformation and included general opinions regarding the vaccine, official news, and appointment details of vaccination centers.