Files
Leo Galambos b29699b763 Refresh multilingual benchmarks and fix overlapping gold evaluation
Recompute published benchmark results for all default language models,
exclude Polish Polimorf, add Hebrew documentation, and record the current
benchmark environment. Evaluate repeated surface forms as an overlapping
gold cover and publish only applicable metrics for candidate policies.
2026-07-23 17:06:41 +02:00

3.1 KiB

Language Benchmark Pages

This section splits Radixor stemmer benchmark results by language. Each of the 20 registered default models has one language page containing the refreshed corpus, patch-command distribution, exact-root accuracy, runtime performance, and pairwise stemming-quality tables for both dictionary-processing modes.

Reference Pages

Page Purpose
Methodology Workload design, normalization, speed metrics, and exact-root quality metrics. Pairwise quality definitions are also reproduced on every language page.
Corpora Dictionary sizes and changed-token timing workloads.
Environment and reports Hardware, JVM, JMH settings, report files, and badge policy.
English dictionary coverage Quality/speed operating curve for contracted Radixor tries built from 100% down to 10% of English dictionary rows.
Candidate evaluation Included and skipped stemmer candidates.

Languages

Language Resource Benchmark page
Czech CS_CZ Czech
Danish DA_DK Danish
Dutch NL_NL Dutch
English US_UK English
Finnish FI_FI Finnish
French FR_FR French
German DE_DE German
Hebrew HE_IL Hebrew
Hungarian HU_HU Hungarian
Italian IT_IT Italian
Norwegian Bokmal NB_NO Norwegian Bokmal
Norwegian Nynorsk NN_NO Norwegian Nynorsk
Persian FA_IR Persian
Polish PL_PL Polish
Portuguese PT_PT Portuguese
Russian RU_RU Russian
Spanish ES_ES Spanish
Swedish SV_SE Swedish
Ukrainian UK_UA Ukrainian
Yiddish YI Yiddish

Methodology Notes

  • Speed benchmarks process only changed dictionary tokens where the surface form differs from the expected root.
  • Accuracy benchmarks process the complete dictionary and report All exact, Changed exact, and Root preserved.
  • Radixor speed must be interpreted together with exact-root quality. A slower Radixor row must not be read as a simple performance weakness when Radixor is also the row with accuracy close to 100% and competing stemmers are much lower. Many fast light, minimal, possessive, or aggressive rule-based stemmers are fast because they do much less linguistic work. The measured Radixor cost buys dictionary-trained precision, and that precision is what improves search quality when queries and indexed text are reduced to the same intended roots. The EnglishRadixorDictionaryCoverageBenchmark table shows this contracted-trie operating curve explicitly.
  • Results are comparable only within the same language and benchmark family.
  • The historical Porter badge is retired; no JMH badge JSON is generated.