Files
Radixor/python/README.md
Leo Galambos 5e3d3c7c7d feat(python): add native distribution and release infrastructure
- add the Rust-backed Python API with PyStemmer compatibility
- distribute standard compiled models as a separate Python package
- generate model artifacts during builds instead of storing them in Git
- add GitHub release and Pages-backed package index workflows
- add Python tests, benchmarks, documentation, and Gradle integration
- refresh the documentation site, branding, and language benchmarks
2026-08-10 22:34:32 +02:00

254 lines
9.8 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# radixor — Fastest Stemming for Python
**radixor** is a Python extension for the [Radixor](https://github.com/leogalambos/Radixor) stemmer library, built on a Rust core via [PyO3](https://pyo3.rs/). It provides sub-microsecond per-word stemming with a batch API that amortises the Python↔Rust bridge overhead across thousands of words at once.
## Why radixor?
| Library | Approach | Batch API |
|---|---|---|
| **radixor** | Compiled patch-command trie in Rust | ✅ `stem_batch()` |
| PyStemmer (Snowball) | C extension (`libstemmer`) | ✅ `stemWords()` |
| snowballstemmer | Pure-Python Snowball | ✅ (Python loop) |
| NLTK Porter / CISTEM | Pure Python | ❌ |
**Performance.** On the shared UniMorph gold-standard corpus, measuring runtime
stemming only (construction excluded) and with a fair, cache-disabled,
same-input methodology, radixor won all **18 / 18** direct comparisons with
PyStemmer 3.1.0 (Snowball's C `libstemmer`) in the published 2026-08-08 run.
At batch size 100, the geometric-mean speedup was **1.67×**. The complete
machine metadata and current results are in the [Python performance
documentation](../docs/python/performance.md); benchmark implementation and
fairness notes are in [`benchmarks/`](benchmarks/README.md).
## Installation
From PyPI, once publication is enabled:
```bash
python -m pip install --only-binary=:all: radixor
```
The GitHub Releases-backed index is the independent alternative:
```bash
python -m pip install --only-binary=:all: \
--index-url https://leogalambos.github.io/Radixor/python/simple/ radixor
```
The GitHub command becomes usable after the first model and native releases
populate that index. See the [installation guide](../docs/python/installation.md)
for current availability and source-checkout builds.
Wheels are provided for Linux, macOS, and Windows (Python 3.9+). The install
also resolves the mandatory pure `radixor-models-standard` dependency
with 20 precompiled standard models. Building the native source distribution
requires Rust ≥ 1.75 and [maturin](https://www.maturin.rs/).
## Quick start
```python
from radixor import Stemmer
s = Stemmer("en") # English (us-uk-default model)
s.stem("running") # → "run"
s.stem("cats") # → "cat"
s.stem("unknown_word") # → None
```
## Batch API — the fast path
```python
words = ["running", "cats", "stemming", "quickly"]
# Amortises the Python→Rust bridge cost across all words at once
stems = s.stem_batch(words)
# → ["run", "cat", "stem", "quick"]
```
For large corpora (tens of thousands of words) the batch call is the recommended interface. It avoids per-call Python frame overhead and keeps the hot loop entirely inside Rust.
## Migrating from PyStemmer
Radixor provides PyStemmer's `stemWord` and `stemWords` method names. These
compatibility methods also follow PyStemmer's fallback behavior: when the trie
has no patch command, they return the original word instead of `None`.
```python
# PyStemmer: import Stemmer
import radixor as Stemmer
stemmer = Stemmer.Stemmer("english")
stemmer.stemWord("running") # → "run"
stemmer.stemWord("unknown_word") # → "unknown_word"
stemmer.stemWords(["running", "unknown"]) # → ["run", "unknown"]
```
The original Radixor methods remain unchanged: `stem` and `stem_batch` return
`None` for words without a matching patch command. Radixor accepts PyStemmer's
full language names for the languages represented by its bundled models, as
well as its existing two-letter codes and model IDs.
## Supported languages
| Code | Language | Model ID |
|---|---|---|
| `cs` | Czech | `cs-cz-default` |
| `da` | Danish | `da-dk-default` |
| `de` | German | `de-de-default` |
| `en` | English | `us-uk-default` |
| `es` | Spanish | `es-es-default` |
| `fa` | Persian | `fa-ir-default` |
| `fi` | Finnish | `fi-fi-default` |
| `fr` | French | `fr-fr-default` |
| `he` | Hebrew | `he-il-default` |
| `hu` | Hungarian | `hu-hu-default` |
| `it` | Italian | `it-it-default` |
| `nb` | Norwegian Bokmål | `nb-no-default` |
| `nl` | Dutch | `nl-nl-default` |
| `nn` | Norwegian Nynorsk | `nn-no-default` |
| `pl` | Polish | `pl-pl-unimorph` |
| `pt` | Portuguese | `pt-pt-default` |
| `ru` | Russian | `ru-ru-default` |
| `sv` | Swedish | `sv-se-default` |
| `uk` | Ukrainian | `uk-ua-default` |
| `yi` | Yiddish | `yi-default` |
## API reference
### `Stemmer(language=None, *, path=None, compiled=None, backward=None, store_original=True, lowercase=True, cache_size=10_000)`
Create a stemmer for the given language code, model ID, custom textual
dictionary, or previously compiled version 7 trie. Textual dictionaries are
compiled in Rust; `compiled=` loads a prepared binary directly.
```python
s = Stemmer("de") # by language code
s = Stemmer("de-de-default") # by model ID
s = Stemmer(path="/data/custom.gz") # custom gzipped dictionary
s = Stemmer(compiled="/data/custom.rxc") # prepared v7 binary
```
`backward` selects the traversal direction; when left as `None` it is derived
from the language (BACKWARD, except right-to-left `fa`/`he`/`yi` which use
FORWARD). `store_original` (default `True`) maps each canonical stem to a no-op
patch so the stem itself is recognised. `lowercase=False` skips runtime
lowercasing for already-normalized input, and `cache_size` enables the bounded
result cache. The default holds up to 10,000 entries, matching PyStemmer;
`cache_size=0` disables it. One cache is shared by `stem()`, `stemWord()`,
`stem_batch()`, and `stemWords()`; the `stem_all*()` methods are not cached.
### `stem(word: str) → str | None`
Return the stem, or `None` when the compiled trie finds no applicable patch
command. This does not mean that lookup is restricted to exact training words.
### `stem_batch(words: list[str]) → list[str | None]`
Stem an entire list. Preferred for large inputs.
### `stemWord(word: str) → str`
PyStemmer-compatible scalar method. Return the original word if it cannot be
stemmed.
### `stemWords(words: list[str]) → list[str]`
PyStemmer-compatible batch method. Return each unrecognized word unchanged.
### `stem_all(word: str) → list[str]`
Return all stems ordered by descending corpus frequency. Useful when multiple valid stems exist.
### `stem_all_batch(words: list[str]) → list[list[str]]`
Return all stems for each word in a batch.
## Compiling a model (compile once, load instantly)
Compiling the trie from a textual dictionary takes time for large languages
(seconds). You can compile it **once** to Radixor's binary format and then load
it near-instantly — the same workflow Java users have:
```python
import radixor
radixor.compile("stemmer.gz", "en.rxc", language="en") # or backward=True/False
s = radixor.Stemmer(compiled="en.rxc") # instant load, no re-compile
```
The compiled file uses Radixor's **v7 trie format and is byte-compatible with
the Java `StemmerPatchTrieBinaryIO`** (the inner stream is identical), so a file
compiled by Java can be loaded by Python and vice versa. `Stemmer(path=...)`
auto-detects whether it was given a compiled trie or a textual dictionary.
## Using a custom model
Provide your own gzipped source dictionary (tab-separated
`stem<TAB>variant1<TAB>variant2…` per line, `#` / `//` line remarks allowed) and
load it directly:
```python
s = Stemmer(path="my_dictionary.gz") # BACKWARD by default
s = Stemmer(path="my_rtl_dictionary.gz", backward=False) # right-to-left
```
The dictionary is compiled to a patch-command trie in Rust at construction time.
## Building from source
```bash
cd Radixor/
pip install maturin build setuptools wheel pytest
./gradlew pythonBuildStandardModels
pip install --no-deps build/python/dist/standard/radixor_models_standard-0.0.0-py3-none-any.whl
cd python/
maturin develop --release # editable install with release optimisations
```
From the repository root, Gradle builds the native wheel/sdist and pure
standard-model wheel/sdist without installing them globally:
```bash
./gradlew pythonBuild
```
The convenience tasks `pythonBuildLinux`, `pythonBuildWindows`, and
`pythonBuildMacos` use the host build when the requested platform matches the
current system. Other platforms are cross-compiled with the corresponding Rust
target and therefore require that target and its linker/SDK to be installed.
Override a default target with, for example,
`-PpythonWindowsTarget=x86_64-pc-windows-gnu`. Build artifacts are written below
`build/python/dist/`.
The complete batch benchmark runs Radixor for all bundled languages and every
available comparison engine for the languages it supports, using batch sizes
10, 20, 50, and 100:
Comparison engines are auto-detected in the environment of `pythonExecutable`.
Install `python/benchmarks/requirements-bench.txt` there to enable the complete
comparison set.
```bash
./gradlew pythonBenchmarkAllLanguagesBatch
```
Use `pythonBenchmarkWords`, `pythonBenchmarkRepeats`, and
`pythonBenchmarkWarmup` Gradle properties to tune the run. CSV and JSON reports
are written below `build/reports/python-benchmarks/`.
Neither runtime distribution contains textual dictionaries. The standard data
sdist contains build-ready gzip v7 `.rxc` files, a checksummed provenance
manifest, and per-model CC BY-SA 3.0 notices. They are generated below `build/`
from canonical `models/*/src/modelInput/stemmer.gz` inputs and are never stored
in Git. `./gradlew regeneratePythonStandardModels` performs this deterministic
generation; repository topology selects the 20 defaults and excludes optional
`pl-pl-polimorf`.
`radixor` requires `radixor-models-standard>=1.0,<2.0`. The Python distribution
version is independent of its `2026.1` Java model-catalog identity and of the
individual model versions recorded in the manifest.
## License
The native/API package is BSD-3-Clause — see [LICENSE](LICENSE). Model data
is separately licensed under CC BY-SA 3.0 in its packaged notices.