Direkt zum Inhalt

Madge, Chris ; Yu, Juntao ; Chamberlain, Jon ; Kruschwitz, Udo ; Paun, Silviu ; Poesio, Massimo

Crowdsourcing and Aggregating Nested Markable Annotations

Conference or workshop item

Madge, Chris, Yu, Juntao, Chamberlain, Jon, Kruschwitz, Udo , Paun, Silviu and Poesio, Massimo (2019) Crowdsourcing and Aggregating Nested Markable Annotations. In: 57th Annual Meeting of the Association for Computational Linguistics, July, 2019, Florence, Italy.

DOI to cite this document: 10.5283/epub.43402


Abstract

One of the key steps in language resource creation is the identification of the text segments to be annotated, or markables, which depending on the task may vary from nominal chunks for named entity resolution to (potentially nested) noun phrases in coreference resolution (or mentions) to larger text segments in text segmentation. Markable identification is typically carried out ...

One of the key steps in language resource creation is the identification of the text segments to be annotated, or markables, which depending on the task may vary from nominal chunks for named entity resolution to (potentially nested) noun phrases in coreference resolution (or mentions) to larger text segments in text segmentation. Markable identification is typically carried out semi-automatically, by running a markable identifier and correcting its output by hand—which is increasingly done via annotators recruited through crowdsourcing and aggregating their responses. In this paper, we present a method for identifying markables for coreference annotation that combines high-performance automatic markable detectors with checking with a Game-With-A-Purpose (GWAP) and aggregation using a Bayesian annotation model. The method was evaluated both on news data and data from a variety of other genres and results in an improvement on F1 of mention boundaries of over seven percentage points when compared with a state-of-the-art, domain-independent automatic mention detector, and almost three points over an in-domain mention detector. One of the key contributions of our proposal is its applicability to the case in which markables are nested, as is the case with coreference markables; but the GWAP and several of the proposed markable detectors are task and language-independent and are thus applicable to a variety of other annotation scenarios.



Involved Institutions


Details

Item typeConference or workshop item (UNSPECIFIED)
Title of BookProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy
PublisherAssociation for Computational Linguistics
Page Rangepp. 797-807
DateJuly 2019
Date of publication29 Jun 2020 13:02
InstitutionsLanguages and Literatures > Institut für Information und Medien, Sprache und Kultur (I:IMSK) > Lehrstuhl für Informationswissenschaft (Prof. Dr. Udo Kruschwitz)
Informatics and Data Science > Department Human-Centered Computing > Lehrstuhl für Informationswissenschaft (Prof. Dr. Udo Kruschwitz)
Identification Number
ValueType
10.18653/v1/P19-1077DOI
Dewey Decimal Classification000 Computer science, information & general works > 020 Library & information sciences
StatusPublished
RefereedYes, this version has been refereed
Created at the University of RegensburgYes
URN of the UB Regensburgurn:nbn:de:bvb:355-epub-434024
Item ID43402

Export bibliographical data

Owner only: item control page

nach oben