TurkMorfBench
A Turkish morphology benchmark: suffix generation on real and nonce stems
TurkMorfBench measures how well a language model knows the Turkish suffix system: vowel harmony, consonant assimilation, consonant softening and buffer consonants. 512 items, 8 suffix tasks, 64 stems; 40 real, 24 nonce (wug). Two evaluation modes: free generation and forced choice among phonological distractors.
Items
8 tasks × 64 stems · 320 real + 192 nonce
Was the rule learned, or the form memorised?
Producing the right suffix on a real word is possible by memorisation alone: the model has seen the form kitabı tens of thousands of times in its corpus. A nonce stem never occurs in any corpus; the model produces the correct suffix only if it knows the rule. The gap between the two shows whether the model learned rules or forms.
This is an adaptation to Turkish suffixes of the wug test used in language-acquisition research since 1958.
These stems appear in no corpus; the correct suffix can only come from rule knowledge. Nonce stems were chosen in patterns that do not trigger vowel drop, reproducibly with seed 1337.
Eight suffix tasks
Every task is applied to all 64 stems: 8 × 64 = 512 items.
| Task | Key | Suffix | Example (real) | Example (nonce) |
|---|---|---|---|---|
| Plural | cogul | -ler/-lar | kapı → kapılar | kuntaş → kuntaşlar |
| Dative | hal_yonelme | -e/-a (+y) | çocuk → çocuğa | brelak → brelağa |
| Locative | hal_bulunma | -de/-da/-te/-ta | çölgap → çölgapta | yontuç → yontuçta |
| Ablative | hal_ayrilma | -den/-dan/-ten/-tan | kuş → kuştan | krint → krintten |
| Accusative | hal_belirtme | -i/-ı/-u/-ü (+y) | ağaç → ağacı | dratok → dratoğu |
| Possessive, 1st singular | iyelik_1sg | -im/-ım/-um/-üm (+m) | masa → masam | çölgap → çölgabım |
| Possessive, 1st plural | iyelik_1pl | -imiz/… (+miz) | ev → evimiz | glodek → glodeğimiz |
| Possessive, 3rd singular | iyelik_3sg | -i/… (+si) | araba → arabası | fönük → fönüğü |
Coverage is limited to regular stems; monosyllabic t/k stems (süt, kat) and special cases such as nk→ng are deliberately out of scope.
Four phonological rules
Vowel harmony
Two-way harmony (-ler/-lar) and four-way harmony (-i/-ı/-u/-ü).
Consonant assimilation
After a voiceless consonant, -de → -te and -den → -ten.
Consonant softening
p/ç/t/k → b/c/d/ğ before a vowel-initial suffix.
Buffer consonants
-y-, -s-, -m- on vowel-final stems.
Verified three independent ways
Softening is lexical in Turkish: kitap → kitabı but süt → sütü. Real stems were therefore restricted to those whose softening behaviour is uncontroversial; for nonce stems a single documented convention applies: a multi-syllabic nonce stem ending in p/ç/t/k softens before a vowel-initial suffix. The benchmark does not measure this, it assumes it.
Generator engine
morfbench_v2.py: a phonological engine applying vowel harmony, assimilation, softening and buffer consonants; the stem pool is cleared of vowel-drop and exception words.
Independent rule table
kural_denetimi.py: a second generator written separately from the engine; it re-derives and compares every answer: 512/512.
Morphological analyser
crosscheck.py: real-word answers are parsed with the zeyrek Turkish morphological analyser; an item is flagged if the expected suffix tag is missing from the parse.
The first version (329 items, 7 tasks) was never published; it contained 18 wrong gold answers on real stems (çocuka, südü, renği and the like) and an inconsistent rule for nonce stems. v2 fixes these, adds the accusative task and raises the stem count from 47 to 64. All reported numbers were measured with v2.
Two modes, two different questions
Generation
The model sees the question (add the -e/-a dative suffix to the word çocuk; write only the word) and answers freely; punctuation and case are stripped and an exact match against gold is required. It measures format compliance together with knowledge; untuned base models score low here.
Forced choice
The highest log-likelihood option is chosen among the gold answer and phonological distractors (çocuğa / çocuğe / çocuka …). Format compliance is taken out of the equation; only knowledge is measured. Items for which no distractor can be built are skipped.
Qwen3-32B and Erk-32B on the same items
Every difference carries a paired-bootstrap 95% confidence interval from 10,000 resamples; an interval that includes zero is written as not demonstrated.
Forced choice · 502 items (no distractor could be built for 10)
| Model | Total (502) | Real (312) | Nonce (190) |
|---|---|---|---|
| Qwen3-32B | 89.44% | 94.55% | 81.05% |
| Erk-32B · LoRA 0,40 | 89.84%+0.40 [−1.20, +2.19]not demonstrated | 95.19%+0.64 [−0.64, +1.92]not demonstrated | 81.05%0.00 [0.00, 0.00]not demonstrated |
| Erk-32B · LoRA 0,75released | 91.43%+1.99 [−0.20, +4.18]not demonstrated | 96.47%+1.92 [−0.32, +4.17]not demonstrated | 83.16%+2.11 [−2.11, +6.32]not demonstrated |
| Erk-32B · LoRA 0,80 | 91.83%+2.39 [+0.20, +4.78] | 97.12%+2.56 [+0.64, +4.81] | 83.16%+2.11 [−2.63, +6.84]not demonstrated |
Generation · 512 items, exact match, Qwen3 thinking mode on
| Model | Total (512) |
|---|---|
| Qwen3-32B | 31.64% |
| Erk-32B · LoRA 0,40 | 55.86%+24.22 [+18.95, +29.49] |
| Erk-32B · LoRA 0,75 | 61.52%+29.88 [+24.61, +35.16] |
| Erk-32B · LoRA 0,75 · identity system promptreleased | 59.38%+27.73 [+22.46, +33.01] |
| Erk-32B · LoRA 0,40 · identity system prompt | 49.02%+17.38 [+11.91, +22.85] |
The roughly 28-point gap between the two modes is format compliance: the base model stays inside its thinking block and never writes the word; in forced choice the same model knows 89%. The system prompt's effect on the generation score depends on scale (−2.15 at 0.75, not demonstrated; −6.84 at 0.40, significant); measure with the configuration you deploy.
What the benchmark is built to measure
Qwen3-32B · forced choice · accuracy
On the base model, real-word accuracy is 94.55% and nonce-stem accuracy 81.05%: the model knows forms far better than rules. On a continued-pretraining model (Erk-32B) the nonce-stem score never separated from the base at any scale; Turkish continued pretraining did not measurably change the suffix system.
Measure with two commands
Output: total / real / nonce accuracy and a per-item prediction vector (JSON). To compare two models on the same items, bootstrap.py gives paired confidence intervals.
$ pip install torch transformers$ python turkmorfbench_olc.py --model Qwen/Qwen3-32B --kip zorunlu$ python turkmorfbench_olc.py --model Qwen/Qwen3-32B --kip uretim
{"kok": "çocuk", "gorev": "hal_yonelme", "cevap": "çocuğa", "tip": "gercek",
"soru": "'çocuk' kelimesine -e/-a yönelme eki ekle. Sadece kelimeyi yaz.",
"kural": "yönelme + ünlü uyumu + yumuşama"}kokThe stem (root) word.
gorevSuffix task key (cogul, hal_yonelme, …).
cevapGold answer.
tipgercek (real) or wug (nonce).
soruThe prompt shown to the model.
kuralThe phonological rules tested by the item; used for subset analysis.
Known limitations
Coverage is limited to regular stems; lexical exceptions (vowel drop burun → burnu, non-softening monosyllables, nk→ng) are deliberately excluded.
The nonce-stem convention is an assumption; softening productivity may vary between speakers. The nonce subset score is relative to that assumption.
With 24 nonce stems the nonce subset is small (192 items, ±5 points); expect not demonstrated for small differences.
Limited to Turkish nominal inflection; no verb inflection, derivation or syntax.
If you use this benchmark
eCloud Tech. (2026). TurkMorfBench: Gerçek ve uydurma gövdelerde Türkçe ek üretimi kıyası (v2). https://github.com/ecloudtechnology/turkmorfbench
License: data and scripts CC BY 4.0. Contact: [email protected]
Measure your model with TurkMorfBench
Data and scripts are open under CC BY 4.0. Contact us to collaborate on evaluation for Turkish models.