- add the Rust-backed Python API with PyStemmer compatibility - distribute standard compiled models as a separate Python package - generate model artifacts during builds instead of storing them in Git - add GitHub release and Pages-backed package index workflows - add Python tests, benchmarks, documentation, and Gradle integration - refresh the documentation site, branding, and language benchmarks
5.6 KiB
Radixor for Python
The radixor package is Radixor's native Python implementation. It is not
a wrapper around the Java library and does not require a JVM: it is a
compiled extension (Rust, via PyO3 and
maturin) that loads precompiled patch-command tries
derived from the same canonical UniMorph data as the Java models.
from radixor import Stemmer
s = Stemmer("en")
s.stem("running") # 'run'
s.stem_batch(["cats", "ran"]) # ['cat', 'run']
- Fast Track — install and produce the first stem.
- Quick Start — the complete application-oriented learning path.
- Installation and building — Linux, Windows, macOS.
- Usage and examples — batch API, caching, and custom models.
- Dictionary compilation — prepare a version 7 binary once and share it with Python or Java.
- Performance — fair, reproducible comparisons vs PyStemmer, snowballstemmer, NLTK Porter, and CISTEM.
The language and model-ID mapping is shared with Java and maintained on the
Built-in Languages page. Installing radixor
also resolves the separate pure radixor-models-standard distribution containing
the 20 default compiled models; Java applications select independently
versioned model JARs.
!!! note "Same models, same results, different runtime"
The standard Python models are compiled from the identical canonical
dictionaries with the identical production reduction configuration
(MERGE_SUBTREES_WITH_EQUIVALENT_DOMINANT_GET_RESULTS, 75 % / 3×,
uniform-subtree contraction, LOWERCASE_WITH_LOCALE_ROOT, AS_IS
diacritics, storeOriginal=true). For a word present in a model, both
implementations return the same dominant stem. The compiled binary format
is shared (see below), so a model compiled by one side loads in the other.
Java vs. Python: read this first
The two implementations solve the same problem but make different runtime trade-offs. Mixing their mental models causes confusion, so the differences are stated explicitly. Neither is “better” — they target different runtimes.
| Aspect | Java (org.egothor:radixor) |
Python (radixor) |
|---|---|---|
| Runtime | JVM library | Compiled extension (Rust/PyO3), no JVM |
| Distribution | Maven JAR + model JARs | abi3 wheel (one wheel per OS/arch, Python ≥ 3.9) |
| Hot-path data structure | CompiledNode graph; routines operate on caller-owned char[] with zero-copy normalized lookups and visitor sinks (EntrySink) |
Flat CSR arrays (no per-node objects); reused UTF‑16 scratch buffers |
| Result cache | None — get() is stateless and re-stems every call |
Bounded, 10,000 entries by default (matching PyStemmer); Stemmer(cache_size=0) disables it |
| Batch API | Not a batch call; you loop and reuse char[]/visitors to avoid allocation |
stem_batch() / stem_all_batch() — one Python↔Rust crossing amortized over the whole list |
| Reduction modes | All three modes selectable at compile time | Fixed to the production DOMINANT mode |
| Extending a compiled trie | Supported — add words/transformations to an already-compiled trie without recompiling | Not exposed — compile from a dictionary (or load a compiled binary) |
| Model resolution | ServiceLoader registry, descriptors, SHA‑256 integrity checks |
Separate standard data package; catalog/format/SHA‑256 validation before synchronous native load |
| Normalization control | Case and diacritic modes fully configurable | lowercase toggle; diacritics AS_IS (models are built this way) |
| Binary format | StemmerPatchTrieBinaryIO v7 read/write (versioned, fingerprinted) |
v7 read/write, inner stream byte-identical to Java; v7 only (no legacy v1–v6) |
| Multiple stems | getAll(...) |
stem_all() / stem_all_batch() |
Runtime capabilities that differ
To avoid surprises, these Java capabilities are not in the Python package:
- Extending / incrementally growing a compiled trie. Python compiles from a source dictionary (or loads a compiled binary); it does not add words to an existing compiled trie at runtime.
- Selectable reduction modes. Only the production
DOMINANTmode is used. - Pluggable provider discovery. Python currently resolves one known standard provider directly; entry-point plugins are not yet exposed.
- Legacy binary versions. Only stream version 7 is read/written.
- Diacritic-removal modes beyond
AS_IS(the bundled models areAS_IS).
Python-specific capabilities
- A batch API (
stem_batch) that amortizes the Python↔native boundary — the single most important call for throughput from Python. - A bounded result cache (
cache_size=10_000by default) for workloads with repeated tokens. It is shared bystem(),stemWord(),stem_batch(), andstemWords(); passcache_size=0to disable it. Thestem_all*()methods are not cached. - A
lowercase=Falsemode to skip per-lookup lowercasing when the caller guarantees already-lowercased input.
Interoperability
The compiled binary is Radixor's v7 trie stream, and the Python runtime writes
the inner stream byte-for-byte identically to the Java
StemmerPatchTrieBinaryIO. Consequently:
- a model compiled by Java (
org.egothor.stemmer.Compile/StemmerPatchTrieBinaryIO.write) loads in Python, and - a model compiled by Python (
radixor.compile(...)) loads in Java.
(The outer gzip wrapper bytes differ between the two gzip implementations; this is irrelevant — both sides decompress to the same v7 stream.)