Dataset Viewer

The dataset viewer is not available because its heuristics could not detect any supported data files. You can try uploading some data files, or configuring the data files location manually.

Synthelion worddata

Per-language linguistic data used by Synthelion's rule-based prompt compressor: function words (stopwords), lemma maps, proper-noun lists, part-of-speech tags, and IDF (inverse document frequency) scores, for 56+ languages.

This data is loaded at runtime by every compress() call — it is not an ML model, just precomputed per-language tables. It is downloaded once into ~/.synthelion/worddata/ on first use (or via synthelion worddata install) rather than bundled in the PyPI wheel, to keep the published package under PyPI's size limits.

Files

Brotli-compressed (.br) YAML/JSON per language (ISO 639-3 code):

Suffix Contents
<lang>.yaml.br Function words / stopwords
<lang>.pos.yaml.br Part-of-speech tag map
<lang>.idf.br Inverse document frequency scores
<lang>.generic.yaml.br Generic/high-frequency word list
<lang>.excl.yaml.br Exclusion list overrides
_index.br ISO 639-1 <-> ISO 639-3 language index

License

MIT License. See LICENSE for details.

Downloads last month
265