The dataset viewer is not available because its heuristics could not detect any supported data files. You can try uploading some data files, or configuring the data files location manually.
Synthelion worddata
Per-language linguistic data used by Synthelion's rule-based prompt compressor: function words (stopwords), lemma maps, proper-noun lists, part-of-speech tags, and IDF (inverse document frequency) scores, for 56+ languages.
This data is loaded at runtime by every compress() call — it is not an ML model, just precomputed per-language tables. It is downloaded once into ~/.synthelion/worddata/ on first use (or via synthelion worddata install) rather than bundled in the PyPI wheel, to keep the published package under PyPI's size limits.
Files
Brotli-compressed (.br) YAML/JSON per language (ISO 639-3 code):
| Suffix | Contents |
|---|---|
<lang>.yaml.br |
Function words / stopwords |
<lang>.pos.yaml.br |
Part-of-speech tag map |
<lang>.idf.br |
Inverse document frequency scores |
<lang>.generic.yaml.br |
Generic/high-frequency word list |
<lang>.excl.yaml.br |
Exclusion list overrides |
_index.br |
ISO 639-1 <-> ISO 639-3 language index |
License
MIT License. See LICENSE for details.
- Downloads last month
- 265