GESIS Leibniz Institute for the Social Sciences: Go to homepage

Text and Data Mining

Text and data mining comprises the development and application of methods which are designed to extract knowledge that is relevant to the social sciences from unstructured texts or data streams.

Main research areas are:

  • Detection of statistical regularities in data and text and alignment of these regularities with variables of interest such as political leaning or gender
  • Combine digital behavioral data and survey data to create new types of user models
  • Semantic enrichment and analysis of collaboratively generated documents (e.g. wikipedia articles or scientific publications) and the social dynamics of the creation process (e.g. conflicts, productivity)
  • Statistical modelling of sequential human behavior (e.g., the decisions made when navigating on the web or individual movement in urban surroundings)
  • Detection, disambiguation and linking of entities which are of interest for the social sciences in academic publications (especially references to research data)
  • Extraction of key information from texts and (semi-)automatic indexing
Name Department Team Email Telephone
Assenmacher, Dennis
Computational Social Science
Data Science Methods
+49 (0221) 47694-484
Bensmann, Felix
Knowledge Technologies for the Social Sciences
Information Extraction and Linking
+49 (0221) 47694-524
Bleier, Dr. Arnim
Computational Social Science
Transparent Social Analytics
+49 (0221) 47694-514
Breuer, Dr. Johannes
Computational Social Science
Digital Society Observatory
+49 (0221) 47694-471
Dahou, Abdelhalim Hafedh
Knowledge Technologies for the Social Sciences
FAIR Data
+49 (0221) 47694-430
Froehling, Leon
Computational Social Science
Digital Society Observatory
+49 (0221) 47694-585
Kohne, Julian
Computational Social Science
Designed Digital Data
+49 (0221) 47694-222
Linzbach, Stephan
Knowledge Technologies for the Social Sciences
Big Data Analytics
+49 (0221) 47694-715
Mathiak, Dr. Brigitte
Knowledge Technologies for the Social Sciences
FAIR Data
+49 (0221) 47694-510
Mayr, Dr. Philipp
Knowledge Technologies for the Social Sciences
Information and Data Retrieval
+49 (0221) 47694-533
Otto, Wolfgang
Knowledge Technologies for the Social Sciences
Information Extraction and Linking
+49 (0221) 47694-543
Soldner, Felix
Computational Social Science
Digital Society Observatory
+49 (0221) 47694-234
Stier, Prof. Dr. Sebastian
Computational Social Science
+49 (0221) 47694-221
Strohmaier, Prof. Dr. Markus
Präsidialbereich
+49 (0221) 47694-225
Wagner, Prof. Dr. Claudia
Computational Social Science
+49 (0221) 47694-224
Ziaja, Dr. Sebastian
Survey Data Curation
Survey Data Augmentation
+49 (0221) 47694-462
Zloch, Dr. (rer. nat.) Matthäus
+49 (0221) 47694-534
  • Ben Aichaoui, Shaimaa, Nawel Hiri, Abdelhalim Hafedh Dahou, and Mohamed Amine Cheragui. 2022. "Automatic Building of a Large Arabic Spelling Error Corpus." SN Computer Science 2 (4): 108. doi: https://doi.org/10.1007/s42979-022-01499-x.
  • Sen, Indira, Mattia Samory, Claudia Wagner, and Isabelle Augenstein. 2022. "Counterfactually Augmented Data and Unintended Bias: The Case of Sexism and Hate Speech Detection." In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, edited by Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz, 4716–4726. Seattle: Association for Computational Linguistics. doi: https://doi.org/10.18653/v1/2022.naacl-main.347.
  • Soldner, Felix, Bennett Kleinberg, and Shane Johnson. 2022. Confounds and Overestimations in Fake Review Detection: Experimentally Controlling for Product-Ownership and Data-Origin. https://osf.io/29euc/?view_only=d382b6f03e1444ffa83da3ea04f1a04a.
  • Batzdorfer, Veronika. 2022. "Conspiracy theories on Twitter: Emerging motifs and temporal dynamics during the COVID-19 pandemic." ODISSEI Conference for Social Science in the Netherlands 2022, Open Data Infrastructure for Social Science and Economic Innovations, Utrecht, 2022-11-03.
  • Martins Rosa, Jorge, N. Gizem Bacaksizlar Turbic, Alda Magalhães Telles, Clara González Tosat, Cristian Jiménez Ruiz, Kalliopi Moraiti, Özgür Karadeniz, and Valentina Pallacci. 2022. "Exploring User Engagement with Portuguese Political Party Pages on Facebook: Data Sprint as Workflow." Dígitos. Revista de Comunicación Digital 8 127-154. doi: https://doi.org/10.7203/drdcd.v1i8.233.