Turkish AMR & Question Answering System
Özyeğin University · TÜBİTAK 1001 (Project No. 123E027)
2023 – 2025 · 36 months
Learning Abstract Meaning Representation for Turkish and Building a Question Answering System
Summary
A multi-layer NLP project building a comprehensive Turkish FrameNet, integrating it with the existing Turkish PropBank (TRopBank) and dependency treebanks to produce semi-automatic Abstract Meaning Representations (AMR) for Turkish, with a graph neural network-based question answering system as the downstream application.
Original Aspects
- Expanding Turkish FrameNet from 2,561 to 20,000 verb synsets to cover all Turkish verbs
- Producing the first Abstract Meaning Representation for Turkish, built on PropBank and FrameNet frames
- Semi-automatic AMR annotation of 5 dependency treebanks (KeNet, Penn, Tourism, Atis, FrameNet) covering ~70,000 sentences
- First graph neural network-based AMR parser for Turkish
- A Turkish question answering system built on a Turkish translation of SQuAD
Work Packages
Turkish FrameNet frame creation
5 monthsReviewing all verbs in Turkish WordNet (KeNet) and creating frames aligned with English FrameNet for ~20,000 verb synsets.
Annotating 5 treebanks (~70,000 sentences)
11 monthsMarking verbal predicates, FrameNet frame elements, and PropBank arguments (ARG0–ARG3) across the 5 treebanks.
AMR generation & corpus annotation
8 monthsDesigning a rule-based AMR parser tailored to Turkish's agglutinative morphology, then semi-automatically annotating the corpus.
Graph neural network parser
8 monthsTraining fastText embeddings and a graph autoencoder, then a supervised GNN model predicting AMR relations between word pairs.
Question answering system
8 monthsBuilding a machine reading comprehension system on a Turkish-translated SQuAD dataset, combining AMR and dependency similarity to locate answers.
Team
Coordinated by the project lead with 4 undergraduate and 3 graduate linguistics students handling annotation, and 1 computer engineering PhD student leading the GNN and QA components.
Resources Used
- Turkish WordNet (KeNet) — ~80,000 word groups
- TRopBank v2.0 — Turkish PropBank, 17,691 verb frame files
- Turkish FrameNet (initial phase: 139 frames, 2,561 synsets)
- 5 Universal Dependencies treebanks: Penn, KeNet, Tourism, Atis, FrameNet (~70,000 sentences)
- Turkish translation of SQuAD for question answering