Data

  • Attitudinally-positioned European sample dataset

    Twitter users of 9 countries positioned along a Left-Right polarization dimension and an anti-elite/institutions dimension

    See also Deliveable D2.1.

    Data set
  • Reddit collective field experiment 2024

    5 waves survey (pre, post survey + 3 check-in surveys, 4th check-in survey integrated in post survey wave) + 4 week discussion data in 6 experimental subreddits, 520 participants (survey waves with varying completeness), ~ 5600 comments internally, ~12000 comments externally (additionally, pilot data for 1 week, two surveys, 1 subreddit and 50 participants)

     Corresponding publication: Lisa Oswald et al. Disentangling participation in online political discussions with a collective field experiment, Science Advances (2026).

    Data set
  • Anonymized retweet networks from German Twitter Trends 2021 - 2023

    They were analyzed in the following paper:

    Pournaki, A., Gaisbauer, F., Olbrich, E. (2025). How Influencers and Multipliers Drive Polarization and Issue Alignment on Twitter/X. To appear in Proceedings of the International AAAI Conference on Web and Social Media (ICWSM). Vol. 19. 2025.

    See also the corresponding github repository for more information. 

    Data set
  • Telegram channel embeddings

    This repository contains the dataset of Telegram channel embeddings accompanying the paper: Peeters, S. and Willaert, T. (2026), "Dark metrics" and the mainstreaming of political extremism on Dutch-speaking Telegram: A comparative reading of platform affordances". Platforms and Society.

    Data set
  • Anonymized message classification data and actantial analyses of public Telegram channels

    Anonymized message classification data and actantial analyses of public Telegram channels pertaining to the paper:
     

    Willaert, T. (2023). A computational analysis of Telegram’s narrative affordances. Plos one, 18(11), e0293508.

    Data set
  • An Inductive Analysis of the Kremlin's Weaponization of Digital Diplomacy on Telegram

    This dataset accompanies the paper: Willaert, T., & Tuters, M. (2025). From denazification to the Golden Billion: an inductive analysis of the Kremlin’s weaponisation of digital diplomacy on Telegram. Humanities and Social Sciences Communications, 12(1), 1-16.

    Data set
  • Europe Day & Berlin Wall Commemorations on X (Slovenia, Italy, Germany, France): Annotated dataset

    This data set contains the research materials for the article “Conflict, Antagonistic Tone, and Deliberative Quality in Online Memory Debates: Europe Day and the Fall of the Berlin Wall on Twitter/X” It includes the anonymised and annotated datasets used to analyse X posts around Europe Day and the fall of the Berlin Wall in the four countries, the LLM-assisted prompt and the topic-modelling outputs used to identify thematic hotspots of conflict.

    Horvat, M., & Koražija, J. (2026). Conflict, Antagonistic Tone, and Deliberative Quality in Online Memory Debates: Europe Day and the Fall of the Berlin Wall on Twitter/X. ANNALES, SERIES HISTORIA ET SOCIOLOGIA, 36(2), 247–266. doi.org/10.19233/ASHS.2026.14

    Data set
  • Slovenian Day of Resistance X & news corpus

    The dataset contains social media posts from X and traditional media articles from online news sources related to the Slovenian commemorations of the Day of Resistance.

    We used two types of data: For the social media analysis, we collected X posts covering the period from April 2023 to April 2024. This dataset was gathered by Sciences Po under the SoMe4Dem project. The collection focused on commemorative discussions in Slovenian and comprised 753 posts. The X dataset was compiled using the query terms “Dan upora proti okupatorju” and “Dan upora”, with special-character normalization to ensure broader retrieval of relevant posts.

    To analyze traditional media, we collected relevant news articles using Media Cloud (https://www.mediacloud.org/), an open-source platform developed by the Berkman Klein Center for Internet & Society at Harvard University, which compiles and organizes online news content to facilitate research on attention, representation, influence, and language in global media ecosystems. The Slovenian database was queried using the following 14 case-sensitive keywords: »dan upora«, »dnevu upora«, »dan OF«, »dneva OF«, »proti okupatorju«, »državna proslava«, »državne proslave«, »državni proslavi«, »dan spomina«, »dnevu spomina«, »osvobodilna fronta«, »osvobodilne fronte«, »protiimperialistična fronta« and »protiimperialistične fronte«. Additional news material was collected through links found in the X dataset and manually retrieved from three Slovenian weekly publications: Delo, Demokracija, and Mladina. We included all relevant news articles published on this topic for three consecutive years, from 2022 to 2024.

    After collecting traditional media news articles from Media Cloud and X links, 144 irrelevant or duplicated articles were identified, thus reducing the media part of our dataset from 308 to 164 articles.

    For publication and data-sharing purposes, version 1.1 transforms the original version of the X dataset into an anonymized, feature-based analytical dataset. The published version contains post-level entries with derived features such as Greimasian actantial coding, actant clusters, actor and character fields, author stance, antagonism score, discourse-function indicators (+ action coding), HDBSCAN-based cluster information, and average cluster scores.

    Horvat, M., Koražija, J., Babnik, J., Škvorc, T., Darovec, D., Oman, Žiga, Lampe, U., Ergaver, A., & Robnik-Šikonja, M. (2026). Mapping Contested Cultural Memory: An LLM-Supported Approach to Analysing Narrative Structures, Discursive Modes and Discourse Functions. ANNALES, SERIES HISTORIA ET SOCIOLOGIA, 36(2), 183–204. https://doi.org/10.19233/ASHS.2026.11

    Data set
  • Giorno del Ricordo (Twitter & News)

    This data set contains the research materials for the article “Agonistic Engagement in Memory Politics: Media Arenas, Normative Orientations, and Debates on Giorno del ricordo in Italy and Slovenia.” It includes anonymised and annotated datasets used to analyse debates across X and online news (2022–2024), the coding schemes, label definitions, the LLM prompts and the opinion poll questionnaire.

    Lampe, U., Horvat, M., Koražija, J., Ergaver, A., & Darovec, D. (2026). Agonistic Engagement in Memory Politics: Media Arenas, Normative Orientations, and Debates on the Giorno del Ricordo in Italy and Slovenia. ANNALES, SERIES HISTORIA ET SOCIOLOGIA, 36(2), 227–246. https://doi.org/10.19233/ASHS.2026.13

    Data set