Wikipedia Text Reuse

Within and Without

verfasst von: Milad Alshomary, Michael Völske, Tristan Licht, Henning Wachsmuth, Benno Stein, Matthias Hagen, Martin Potthast
Abstract: We study text reuse related to Wikipedia at scale by compiling the first corpus of text reuse cases within Wikipedia as well as without (i.e., reuse of Wikipedia text in a sample of the Common Crawl). To discover reuse beyond verbatim copy and paste, we employ state-of-the-art text reuse detection technology, scaling it for the first time to process the entire Wikipedia as part of a distributed retrieval pipeline. We further report on a pilot analysis of the 100 million reuse cases inside, and the 1.6 million reuse cases outside Wikipedia that we discovered. Text reuse inside Wikipedia gives rise to new tasks such as article template induction, fixing quality flaws, or complementing Wikipedia’s ontology. Text reuse outside Wikipedia yields a tangible metric for the emerging field of quantifying Wikipedia’s influence on the web. To foster future research into these tasks, and for reproducibility’s sake, the Wikipedia text reuse corpus and the retrieval pipeline are made freely available.
Externe Organisation(en): Universität Paderborn
Bauhaus-Universität Weimar
Martin-Luther-Universität Halle-Wittenberg
Universität Leipzig
Typ: Aufsatz in Konferenzband
Seiten: 747-754
Anzahl der Seiten: 8
Publikationsdatum: 2019
Publikationsstatus: Veröffentlicht
Peer-reviewed: Ja
ASJC Scopus Sachgebiete: Theoretische Informatik, Informatik (insg.)
Elektronische Version(en): https://doi.org/10.1007/978-3-030-15712-8_49 (Zugang: Geschlossen)

BibTeX

@inproceedings{95304f127da44d2d9424b2724841981c,
title = "Wikipedia Text Reuse: Within and Without",
abstract = "We study text reuse related to Wikipedia at scale by compiling the first corpus of text reuse cases within Wikipedia as well as without (i.e., reuse of Wikipedia text in a sample of the Common Crawl). To discover reuse beyond verbatim copy and paste, we employ state-of-the-art text reuse detection technology, scaling it for the first time to process the entire Wikipedia as part of a distributed retrieval pipeline. We further report on a pilot analysis of the 100 million reuse cases inside, and the 1.6 million reuse cases outside Wikipedia that we discovered. Text reuse inside Wikipedia gives rise to new tasks such as article template induction, fixing quality flaws, or complementing Wikipedia{\textquoteright}s ontology. Text reuse outside Wikipedia yields a tangible metric for the emerging field of quantifying Wikipedia{\textquoteright}s influence on the web. To foster future research into these tasks, and for reproducibility{\textquoteright}s sake, the Wikipedia text reuse corpus and the retrieval pipeline are made freely available.",
author = "Milad Alshomary and Michael V{\"o}lske and Tristan Licht and Henning Wachsmuth and Benno Stein and Matthias Hagen and Martin Potthast",
year = "2019",
doi = "10.1007/978-3-030-15712-8_49",
language = "English",
isbn = "9783030157111",
series = "Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)",
publisher = "Springer Verlag",
pages = "747--754",
editor = "Benno Stein and Norbert Fuhr and Leif Azzopardi and Philipp Mayr and Djoerd Hiemstra and Claudia Hauff",
booktitle = "Advances in Information Retrieval",
address = "Germany",
note = "41st European Conference on Information Retrieval, ECIR 2019 ; Conference date: 14-04-2019 Through 18-04-2019",
}

Details zu Publikationen

Wikipedia Text Reuse

Within and Without

Gefördert vom