Projects
    Open benchmark dataset · CC BY 4.0

    TurkMorfBench

    A Turkish morphology benchmark: suffix generation on real and nonce stems

    TurkMorfBench measures how well a language model knows the Turkish suffix system: vowel harmony, consonant assimilation, consonant softening and buffer consonants. 512 items, 8 suffix tasks, 64 stems; 40 real, 24 nonce (wug). Two evaluation modes: free generation and forced choice among phonological distractors.

    0

    Items

    8 tasks × 64 stems · 320 real + 192 nonce

    8
    Suffix tasks
    64
    Stems
    40
    Real stems
    24
    Nonce stems
    2
    Evaluation modes
    CC BY 4.0
    License
    Why nonce stems

    Was the rule learned, or the form memorised?

    Producing the right suffix on a real word is possible by memorisation alone: the model has seen the form kitabı tens of thousands of times in its corpus. A nonce stem never occurs in any corpus; the model produces the correct suffix only if it knows the rule. The gap between the two shows whether the model learned rules or forms.

    This is an adaptation to Turkish suffixes of the wug test used in language-acquisition research since 1958.

    Examples of nonce stems
    çölgapzelbikpürsatbrelakkuntaşglodekfönükdratok

    These stems appear in no corpus; the correct suffix can only come from rule knowledge. Nonce stems were chosen in patterns that do not trigger vowel drop, reproducibly with seed 1337.

    Coverage

    Eight suffix tasks

    Every task is applied to all 64 stems: 8 × 64 = 512 items.

    TaskKeySuffixExample (real)Example (nonce)
    Pluralcogul-ler/-larkapı → kapılarkuntaş → kuntaşlar
    Dativehal_yonelme-e/-a (+y)çocuk → çocuğabrelak → brelağa
    Locativehal_bulunma-de/-da/-te/-taçölgap → çölgaptayontuç → yontuçta
    Ablativehal_ayrilma-den/-dan/-ten/-tankuş → kuştankrint → krintten
    Accusativehal_belirtme-i/-ı/-u/-ü (+y)ağaç → ağacıdratok → dratoğu
    Possessive, 1st singulariyelik_1sg-im/-ım/-um/-üm (+m)masa → masamçölgap → çölgabım
    Possessive, 1st pluraliyelik_1pl-imiz/… (+miz)ev → evimizglodek → glodeğimiz
    Possessive, 3rd singulariyelik_3sg-i/… (+si)araba → arabasıfönük → fönüğü

    Coverage is limited to regular stems; monosyllabic t/k stems (süt, kat) and special cases such as nk→ng are deliberately out of scope.

    Rules tested

    Four phonological rules

    Vowel harmony

    Two-way harmony (-ler/-lar) and four-way harmony (-i/-ı/-u/-ü).

    Consonant assimilation

    After a voiceless consonant, -de → -te and -den → -ten.

    Consonant softening

    p/ç/t/k → b/c/d/ğ before a vowel-initial suffix.

    Buffer consonants

    -y-, -s-, -m- on vowel-final stems.

    Gold answers

    Verified three independent ways

    Softening is lexical in Turkish: kitap → kitabı but süt → sütü. Real stems were therefore restricted to those whose softening behaviour is uncontroversial; for nonce stems a single documented convention applies: a multi-syllabic nonce stem ending in p/ç/t/k softens before a vowel-initial suffix. The benchmark does not measure this, it assumes it.

    01

    Generator engine

    morfbench_v2.py: a phonological engine applying vowel harmony, assimilation, softening and buffer consonants; the stem pool is cleared of vowel-drop and exception words.

    02

    Independent rule table

    kural_denetimi.py: a second generator written separately from the engine; it re-derives and compares every answer: 512/512.

    03

    Morphological analyser

    crosscheck.py: real-word answers are parsed with the zeyrek Turkish morphological analyser; an item is flagged if the expected suffix tag is missing from the parse.

    Version note · v1 → v2

    The first version (329 items, 7 tasks) was never published; it contained 18 wrong gold answers on real stems (çocuka, südü, renği and the like) and an inconsistent rule for nonce stems. v2 fixes these, adds the accusative task and raises the stem count from 47 to 64. All reported numbers were measured with v2.

    Evaluation modes

    Two modes, two different questions

    01

    Generation

    The model sees the question (add the -e/-a dative suffix to the word çocuk; write only the word) and answers freely; punctuation and case are stripped and an exact match against gold is required. It measures format compliance together with knowledge; untuned base models score low here.

    02

    Forced choice

    The highest log-likelihood option is chosen among the gold answer and phonological distractors (çocuğa / çocuğe / çocuka …). Format compliance is taken out of the equation; only knowledge is measured. Items for which no distractor can be built are skipped.

    The gap between the two modes is informative: a continued-pretraining model we measured beat its base by +36 points in generation but by +3.7 in forced choice. Most of the difference was format compliance, not knowledge. Do not report the generation score on its own.
    Results (v2)

    Qwen3-32B and Erk-32B on the same items

    Every difference carries a paired-bootstrap 95% confidence interval from 10,000 resamples; an interval that includes zero is written as not demonstrated.

    Forced choice · 502 items (no distractor could be built for 10)

    ModelTotal (502)Real (312)Nonce (190)
    Qwen3-32B89.44%94.55%81.05%
    Erk-32B · LoRA 0,4089.84%+0.40 [−1.20, +2.19]not demonstrated95.19%+0.64 [−0.64, +1.92]not demonstrated81.05%0.00 [0.00, 0.00]not demonstrated
    Erk-32B · LoRA 0,75released91.43%+1.99 [−0.20, +4.18]not demonstrated96.47%+1.92 [−0.32, +4.17]not demonstrated83.16%+2.11 [−2.11, +6.32]not demonstrated
    Erk-32B · LoRA 0,8091.83%+2.39 [+0.20, +4.78]97.12%+2.56 [+0.64, +4.81]83.16%+2.11 [−2.63, +6.84]not demonstrated

    Generation · 512 items, exact match, Qwen3 thinking mode on

    ModelTotal (512)
    Qwen3-32B31.64%
    Erk-32B · LoRA 0,4055.86%+24.22 [+18.95, +29.49]
    Erk-32B · LoRA 0,7561.52%+29.88 [+24.61, +35.16]
    Erk-32B · LoRA 0,75 · identity system promptreleased59.38%+27.73 [+22.46, +33.01]
    Erk-32B · LoRA 0,40 · identity system prompt49.02%+17.38 [+11.91, +22.85]

    The roughly 28-point gap between the two modes is format compliance: the base model stays inside its thinking block and never writes the word; in forced choice the same model knows 89%. The system prompt's effect on the generation score depends on scale (−2.15 at 0.75, not demonstrated; −6.84 at 0.40, significant); measure with the configuration you deploy.

    The real–nonce gap

    What the benchmark is built to measure

    Qwen3-32B · forced choice · accuracy

    Real
    94.55%
    Nonce
    81.05%

    On the base model, real-word accuracy is 94.55% and nonce-stem accuracy 81.05%: the model knows forms far better than rules. On a continued-pretraining model (Erk-32B) the nonce-stem score never separated from the base at any scale; Turkish continued pretraining did not measurably change the suffix system.

    Usage

    Measure with two commands

    Output: total / real / nonce accuracy and a per-item prediction vector (JSON). To compare two models on the same items, bootstrap.py gives paired confidence intervals.

    bash
    $ pip install torch transformers$ python turkmorfbench_olc.py --model Qwen/Qwen3-32B --kip zorunlu$ python turkmorfbench_olc.py --model Qwen/Qwen3-32B --kip uretim
    turkmorfbench.json
    {"kok": "çocuk", "gorev": "hal_yonelme", "cevap": "çocuğa", "tip": "gercek",
     "soru": "'çocuk' kelimesine -e/-a yönelme eki ekle. Sadece kelimeyi yaz.",
     "kural": "yönelme + ünlü uyumu + yumuşama"}
    kok

    The stem (root) word.

    gorev

    Suffix task key (cogul, hal_yonelme, …).

    cevap

    Gold answer.

    tip

    gercek (real) or wug (nonce).

    soru

    The prompt shown to the model.

    kural

    The phonological rules tested by the item; used for subset analysis.

    Limits

    Known limitations

    01

    Coverage is limited to regular stems; lexical exceptions (vowel drop burun → burnu, non-softening monosyllables, nk→ng) are deliberately excluded.

    02

    The nonce-stem convention is an assumption; softening productivity may vary between speakers. The nonce subset score is relative to that assumption.

    03

    With 24 nonce stems the nonce subset is small (192 items, ±5 points); expect not demonstrated for small differences.

    04

    Limited to Turkish nominal inflection; no verb inflection, derivation or syntax.

    Citation

    If you use this benchmark

    citation
    eCloud Tech. (2026). TurkMorfBench: Gerçek ve uydurma gövdelerde Türkçe ek
    üretimi kıyası (v2). https://github.com/ecloudtechnology/turkmorfbench

    License: data and scripts CC BY 4.0. Contact: [email protected]

    Measure your model with TurkMorfBench

    Data and scripts are open under CC BY 4.0. Contact us to collaborate on evaluation for Turkish models.