GWC2025 — Pavia
Global WordNet Conference, June 2025
A Semi-Automated Approach to the Annotation of Argument Structures in Turkish Datasets
Pavia, Italy · Global WordNet Conference 2025
Abstract
This paper presents a PropBank annotation project for Turkish, focusing on core arguments in matrix clauses. The dataset comprises five different corpora with 25,580 sentences labeled for verbal predicates and core arguments. Using a semi-automatic approach, a dependency layer was leveraged to pre-assign some ARG0s and ARG1s, followed by manual corrections. This work feeds into the development of Abstract Meaning Representations (AMRs), enhancing Turkish NLP resources for semantic parsing and higher-level language tasks.
Methodology
- Predicate selection: identifying verbal predicates while excluding noun phrases and nominal clauses
- Dependency-based pre-labeling: automatically assigning ARG0/ARG1 from NSUBJ and OBJ relations, with manual correction for unaccusative and passive predicates
- Manual argument selection using the StarDust annotation interface, assigning frame-file roles from TRopBank v2.0
Datasets & Argument Distribution
ARG0 typically marks the agent or causer, while ARG1 denotes the patient or theme. ARG2 and ARG3 represent more specialized roles depending on the verb's semantics.
| Dataset | ARG0 | ARG1 | ARG2 | ARG3 |
|---|---|---|---|---|
| Atis | 995 | 16,090 | 698 | 0 |
| Tourism | 1,279 | 6,259 | 142 | 0 |
| FrameNet | 1,211 | 3,001 | 159 | 8 |
| Penn TreeBank | 18,283 | 31,359 | 756 | 25 |
| KeNet | 11,811 | 24,797 | 1,267 | 42 |
| Total | 32,542 | 76,983 | 2,853 | 73 |
Conference Info
Pavia, Italy — Global WordNet Conference 2025
Authors
Neslihan Cesur, Sabri İnce, Ali Hakkı Aydın, Ece Su Eren, Deniz Gücükçavuş, Murat Papaker, Kaan Bayar, Deniz Baran Aslan, Yelda Fırat, Olcay Taner Yıldız