feat(python): add native distribution and release infrastructure

- add the Rust-backed Python API with PyStemmer compatibility
- distribute standard compiled models as a separate Python package
- generate model artifacts during builds instead of storing them in Git
- add GitHub release and Pages-backed package index workflows
- add Python tests, benchmarks, documentation, and Gradle integration
- refresh the documentation site, branding, and language benchmarks
This commit is contained in:
2026-08-10 22:34:32 +02:00
parent b45e143c84
commit 5e3d3c7c7d
139 changed files with 11420 additions and 747 deletions

96
docs/python/index.md Normal file
View File

@@ -0,0 +1,96 @@
# Radixor for Python
The **`radixor`** package is Radixor's native Python implementation. It is not
a wrapper around the Java library and does not require a JVM: it is a
compiled extension (Rust, via [PyO3](https://pyo3.rs/) and
[maturin](https://www.maturin.rs/)) that loads precompiled patch-command tries
derived from the same canonical UniMorph data as the Java models.
```python
from radixor import Stemmer
s = Stemmer("en")
s.stem("running") # 'run'
s.stem_batch(["cats", "ran"]) # ['cat', 'run']
```
- [Fast Track](fast-track.md) — install and produce the first stem.
- [Quick Start](quick-start.md) — the complete application-oriented learning path.
- [Installation and building](installation.md) — Linux, Windows, macOS.
- [Usage and examples](usage.md) — batch API, caching, and custom models.
- [Dictionary compilation](model-compilation.md) — prepare a version 7 binary
once and share it with Python or Java.
- [Performance](performance.md) — fair, reproducible comparisons vs PyStemmer,
snowballstemmer, NLTK Porter, and CISTEM.
The language and model-ID mapping is shared with Java and maintained on the
[Built-in Languages](../built-in-languages.md) page. Installing `radixor`
also resolves the separate pure `radixor-models-standard` distribution containing
the 20 default compiled models; Java applications select independently
versioned model JARs.
!!! note "Same models, same results, different runtime"
The standard Python models are compiled from the identical canonical
dictionaries with the identical production reduction configuration
(`MERGE_SUBTREES_WITH_EQUIVALENT_DOMINANT_GET_RESULTS`, 75 % / 3×,
uniform-subtree contraction, `LOWERCASE_WITH_LOCALE_ROOT`, `AS_IS`
diacritics, `storeOriginal=true`). For a word present in a model, both
implementations return the same dominant stem. The compiled **binary format
is shared** (see below), so a model compiled by one side loads in the other.
## Java vs. Python: read this first
The two implementations solve the same problem but make different runtime
trade-offs. Mixing their mental models causes confusion, so the differences are
stated explicitly. **Neither is “better”** — they target different runtimes.
| Aspect | Java (`org.egothor:radixor`) | Python (`radixor`) |
|---|---|---|
| Runtime | JVM library | Compiled extension (Rust/PyO3), no JVM |
| Distribution | Maven JAR + model JARs | `abi3` wheel (one wheel per OS/arch, Python ≥ 3.9) |
| Hot-path data structure | `CompiledNode` graph; routines operate on caller-owned **`char[]`** with zero-copy normalized lookups and visitor sinks (`EntrySink`) | Flat **CSR arrays** (no per-node objects); reused UTF16 scratch buffers |
| Result cache | **None**`get()` is stateless and re-stems every call | **Bounded**, 10,000 entries by default (matching PyStemmer); `Stemmer(cache_size=0)` disables it |
| Batch API | Not a batch call; you loop and reuse `char[]`/visitors to avoid allocation | **`stem_batch()` / `stem_all_batch()`** — one Python↔Rust crossing amortized over the whole list |
| Reduction modes | All three modes selectable at compile time | Fixed to the production `DOMINANT` mode |
| Extending a compiled trie | **Supported** — add words/transformations to an already-compiled trie without recompiling | **Not exposed** — compile from a dictionary (or load a compiled binary) |
| Model resolution | `ServiceLoader` registry, descriptors, SHA256 integrity checks | Separate standard data package; catalog/format/SHA256 validation before synchronous native load |
| Normalization control | Case and diacritic modes fully configurable | `lowercase` toggle; diacritics `AS_IS` (models are built this way) |
| Binary format | `StemmerPatchTrieBinaryIO` v7 read/write (versioned, fingerprinted) | v7 read/write, **inner stream byte-identical** to Java; **v7 only** (no legacy v1v6) |
| Multiple stems | `getAll(...)` | `stem_all()` / `stem_all_batch()` |
### Runtime capabilities that differ
To avoid surprises, these Java capabilities are **not** in the Python package:
- **Extending / incrementally growing a compiled trie.** Python compiles from a
source dictionary (or loads a compiled binary); it does not add words to an
existing compiled trie at runtime.
- **Selectable reduction modes.** Only the production `DOMINANT` mode is used.
- **Pluggable provider discovery.** Python currently resolves one known
standard provider directly; entry-point plugins are not yet exposed.
- **Legacy binary versions.** Only stream version 7 is read/written.
- **Diacritic-removal modes** beyond `AS_IS` (the bundled models are `AS_IS`).
### Python-specific capabilities
- A **batch API** (`stem_batch`) that amortizes the Python↔native boundary — the
single most important call for throughput from Python.
- A **bounded result cache** (`cache_size=10_000` by default) for workloads with
repeated tokens. It is shared by `stem()`, `stemWord()`, `stem_batch()`, and
`stemWords()`; pass `cache_size=0` to disable it. The `stem_all*()` methods are
not cached.
- A `lowercase=False` mode to skip per-lookup lowercasing when the caller
guarantees already-lowercased input.
## Interoperability
The compiled binary is Radixor's **v7 trie stream**, and the Python runtime writes
the *inner stream byte-for-byte identically to the Java*
`StemmerPatchTrieBinaryIO`. Consequently:
- a model compiled by **Java** (`org.egothor.stemmer.Compile` /
`StemmerPatchTrieBinaryIO.write`) loads in **Python**, and
- a model compiled by **Python** (`radixor.compile(...)`) loads in **Java**.
(The outer gzip wrapper bytes differ between the two gzip implementations; this
is irrelevant — both sides decompress to the same v7 stream.)