AHA
Back to Research

GWC2025 — Pavia

Global WordNet Conference, June 2025

Conference Paper

A Semi-Automated Approach to the Annotation of Argument Structures in Turkish Datasets

Pavia, Italy · Global WordNet Conference 2025

Abstract

This paper presents a PropBank annotation project for Turkish, focusing on core arguments in matrix clauses. The dataset comprises five different corpora with 25,580 sentences labeled for verbal predicates and core arguments. Using a semi-automatic approach, a dependency layer was leveraged to pre-assign some ARG0s and ARG1s, followed by manual corrections. This work feeds into the development of Abstract Meaning Representations (AMRs), enhancing Turkish NLP resources for semantic parsing and higher-level language tasks.

Methodology

  • Predicate selection: identifying verbal predicates while excluding noun phrases and nominal clauses
  • Dependency-based pre-labeling: automatically assigning ARG0/ARG1 from NSUBJ and OBJ relations, with manual correction for unaccusative and passive predicates
  • Manual argument selection using the StarDust annotation interface, assigning frame-file roles from TRopBank v2.0

Datasets & Argument Distribution

ARG0 typically marks the agent or causer, while ARG1 denotes the patient or theme. ARG2 and ARG3 represent more specialized roles depending on the verb's semantics.

DatasetARG0ARG1ARG2ARG3
Atis99516,0906980
Tourism1,2796,2591420
FrameNet1,2113,0011598
Penn TreeBank18,28331,35975625
KeNet11,81124,7971,26742
Total32,54276,9832,85373

Conference Info

Pavia, Italy — Global WordNet Conference 2025

DOI: 10.18653/v1/2025.gwc-1.28

Authors

Neslihan Cesur, Sabri İnce, Ali Hakkı Aydın, Ece Su Eren, Deniz Gücükçavuş, Murat Papaker, Kaan Bayar, Deniz Baran Aslan, Yelda Fırat, Olcay Taner Yıldız