AHA
Back to Research

Turkish AMR & Question Answering System

Özyeğin University · TÜBİTAK 1001 (Project No. 123E027)

2023 – 2025 · 36 months

TÜBİTAK ProjectCompleted

Learning Abstract Meaning Representation for Turkish and Building a Question Answering System

Summary

A multi-layer NLP project building a comprehensive Turkish FrameNet, integrating it with the existing Turkish PropBank (TRopBank) and dependency treebanks to produce semi-automatic Abstract Meaning Representations (AMR) for Turkish, with a graph neural network-based question answering system as the downstream application.

Original Aspects

  • Expanding Turkish FrameNet from 2,561 to 20,000 verb synsets to cover all Turkish verbs
  • Producing the first Abstract Meaning Representation for Turkish, built on PropBank and FrameNet frames
  • Semi-automatic AMR annotation of 5 dependency treebanks (KeNet, Penn, Tourism, Atis, FrameNet) covering ~70,000 sentences
  • First graph neural network-based AMR parser for Turkish
  • A Turkish question answering system built on a Turkish translation of SQuAD

Work Packages

Turkish FrameNet frame creation

5 months

Reviewing all verbs in Turkish WordNet (KeNet) and creating frames aligned with English FrameNet for ~20,000 verb synsets.

Annotating 5 treebanks (~70,000 sentences)

11 months

Marking verbal predicates, FrameNet frame elements, and PropBank arguments (ARG0–ARG3) across the 5 treebanks.

AMR generation & corpus annotation

8 months

Designing a rule-based AMR parser tailored to Turkish's agglutinative morphology, then semi-automatically annotating the corpus.

Graph neural network parser

8 months

Training fastText embeddings and a graph autoencoder, then a supervised GNN model predicting AMR relations between word pairs.

Question answering system

8 months

Building a machine reading comprehension system on a Turkish-translated SQuAD dataset, combining AMR and dependency similarity to locate answers.

Team

Coordinated by the project lead with 4 undergraduate and 3 graduate linguistics students handling annotation, and 1 computer engineering PhD student leading the GNN and QA components.

Resources Used

  • Turkish WordNet (KeNet) — ~80,000 word groups
  • TRopBank v2.0 — Turkish PropBank, 17,691 verb frame files
  • Turkish FrameNet (initial phase: 139 frames, 2,561 synsets)
  • 5 Universal Dependencies treebanks: Penn, KeNet, Tourism, Atis, FrameNet (~70,000 sentences)
  • Turkish translation of SQuAD for question answering