| Published Version Download ( PDF | 572kB) | License: Creative Commons Attribution 4.0 |
A Corpus of Memes from Reddit: Acquisition, Preparation and First Case Studies
Schmidt, Thomas
, Schiller, Fabian, Götz, Mathias and Wolff, Christian
(2023)
A Corpus of Memes from Reddit: Acquisition, Preparation and First Case Studies.
In: Klein, Maike and Krupka, Daniel and Winter, Cornelia and Wohlgemuth, Volker, (eds.)
INFORMATIK 2023. Designing Futures: Zukünfte gestalten.
Lecture Notes in Informatics (LNI), 337.
Gesellschaft für Informatik e.V. (GI), Bonn, pp. 795-804.
ISBN 978-3-88579-731-9.
Date of publication of this fulltext: 28 May 2024 11:04
Book section
Abstract
We present a corpus of memes and their textual components that were acquired from the popular meme platform r\memes, a subreddit of Reddit and one of the major outlets of online meme culture. The corpus consists of the most popular memes from 2013-2021 on the platform and we acquired 11,701 memes and 280,351 text tokens. We conduct several case studies focused on diachronic analysis to highlight ...
We present a corpus of memes and their textual components that were acquired from the popular meme platform r, a subreddit of Reddit and one of the major outlets of online meme culture. The corpus consists of the most popular memes from 2013-2021 on the platform and we acquired 11,701 memes and 280,351 text tokens. We conduct several case studies focused on diachronic analysis to highlight the possibilities of the corpus for research in internet studies and online culture. We examine the general activity on the platform throughout the years and identify a significant increase in meme production beginning 2017. Results of sentiment analysis show a tendency towards memes with positively classified texts. The analysis of most frequent words per half-year spotlights the importance of certain cultural events for meme culture (e.g. the 2016 US election). Using the LIWC to analyze swear and sexual words shows an overall decrease in the usage of these words pointing to an increased moderation of the platform. The corpus is publicly available for the research community for further studies.
Alternative links to fulltext
Involved Institutions
Details
| Item type | Book section | ||||
| ISBN | 978-3-88579-731-9 | ||||
| Title of Book: | INFORMATIK 2023. Designing Futures: Zukünfte gestalten | ||||
|---|---|---|---|---|---|
| Publisher: | Gesellschaft für Informatik e.V. (GI) | ||||
| Open Access Type: | CC-License | ||||
| Place of Publication: | Bonn | ||||
| Other Series: | Lecture Notes in Informatics (LNI) | ||||
| Volume: | 337 | ||||
| Page Range: | pp. 795-804 | ||||
| Date | September 2023 | ||||
| Institutions | Languages and Literatures > Institut für Information und Medien, Sprache und Kultur (I:IMSK) > Lehrstuhl für Medieninformatik (Prof. Dr. Christian Wolff) Informatics and Data Science > Department Human-Centered Computing > Lehrstuhl für Medieninformatik (Prof. Dr. Christian Wolff) | ||||
| Identification Number |
| ||||
| Related URLs |
| ||||
| Keywords | memes, internet studies, corpus, natural language processing, sentiment analysis, Reddit | ||||
| Dewey Decimal Classification | 000 Computer science, information & general works > 004 Computer science | ||||
| Status | Published | ||||
| Refereed | Yes, this version has been refereed | ||||
| Created at the University of Regensburg | Yes | ||||
| URN of the UB Regensburg | urn:nbn:de:bvb:355-epub-582593 | ||||
| Item ID | 58259 |
Download Statistics
Download Statistics