Files
Radixor/docs/benchmarks/languages/index.md
Leo Galambos 05f3855b99 feat(benchmarks): expand multilingual stemming quality evaluation
* cover all Radixor dictionary languages
* add PRIMARY_OUTPUT, ANY_CANDIDATE, and ALL_CANDIDATES policies
* measure pairwise over-stemming and under-stemming
* add balanced accuracy and complementary quality metrics
* compare single-output and multi-output stemmers fairly
* improve result validation, reporting, and documentation
* move stemming quality tests into the standard test source set
* preserve the existing JMH benchmark structure and badge output
2026-07-20 23:20:17 +02:00

3.0 KiB

Language Benchmark Pages

This section splits Radixor stemmer benchmark results by language. Each language page preserves the existing exact-root accuracy and runtime-performance results and adds pairwise stemming-quality tables for both dictionary-processing modes.

Reference Pages

Page Purpose
Methodology Workload design, normalization, speed metrics, and exact-root quality metrics. Pairwise quality definitions are also reproduced on every language page.
Corpora Dictionary sizes and changed-token timing workloads.
Environment and reports Hardware, JVM, JMH settings, report files, and badge policy.
English dictionary coverage Quality/speed operating curve for contracted Radixor tries built from 100% down to 10% of English dictionary rows.
Candidate evaluation Included and skipped stemmer candidates.

Languages

Language Resource Benchmark page
Czech CS_CZ Czech
Danish DA_DK Danish
Dutch NL_NL Dutch
English US_UK English
Finnish FI_FI Finnish
French FR_FR French
German DE_DE German
Hungarian HU_HU Hungarian
Italian IT_IT Italian
Norwegian Bokmal NB_NO Norwegian Bokmal
Norwegian Nynorsk NN_NO Norwegian Nynorsk
Persian FA_IR Persian
Polish PL_PL Polish
Portuguese PT_PT Portuguese
Russian RU_RU Russian
Spanish ES_ES Spanish
Swedish SV_SE Swedish
Ukrainian UK_UA Ukrainian
Yiddish YI Yiddish

Methodology Notes

  • Speed benchmarks process only changed dictionary tokens where the surface form differs from the expected root.
  • Accuracy benchmarks process the complete dictionary and report All exact, Changed exact, and Root preserved.
  • Radixor speed must be interpreted together with exact-root quality. A slower Radixor row must not be read as a simple performance weakness when Radixor is also the row with accuracy close to 100% and competing stemmers are much lower. Many fast light, minimal, possessive, or aggressive rule-based stemmers are fast because they do much less linguistic work. The measured Radixor cost buys dictionary-trained precision, and that precision is what improves search quality when queries and indexed text are reduced to the same intended roots. The EnglishRadixorDictionaryCoverageBenchmark table shows this contracted-trie operating curve explicitly.
  • Results are comparable only within the same language and benchmark family.
  • The historical Porter badge is retired; no JMH badge JSON is generated.