-
How multilingual is Multilingual BERT?
Paper • 1906.01502 • Published -
Unsupervised Cross-lingual Representation Learning at Scale
Paper • 1911.02116 • Published • 4 -
The State and Fate of Linguistic Diversity and Inclusion in the NLP World
Paper • 2004.09095 • Published -
No Language Left Behind: Scaling Human-Centered Machine Translation
Paper • 2207.04672 • Published • 4
Collections
Discover the best community collections!
Collections including paper arxiv:1911.02116
-
Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on Hugging Face
Paper • 2401.13822 • Published • 1 -
Attention Is All You Need
Paper • 1706.03762 • Published • 144 -
HuggingFace's Transformers: State-of-the-art Natural Language Processing
Paper • 1910.03771 • Published • 33 -
Model Cards for Model Reporting
Paper • 1810.03993 • Published • 7
-
Attention Is All You Need
Paper • 1706.03762 • Published • 144 -
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Paper • 1810.04805 • Published • 33 -
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Paper • 1907.11692 • Published • 10 -
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Paper • 1910.01108 • Published • 24
-
How multilingual is Multilingual BERT?
Paper • 1906.01502 • Published -
Unsupervised Cross-lingual Representation Learning at Scale
Paper • 1911.02116 • Published • 4 -
The State and Fate of Linguistic Diversity and Inclusion in the NLP World
Paper • 2004.09095 • Published -
No Language Left Behind: Scaling Human-Centered Machine Translation
Paper • 2207.04672 • Published • 4
-
Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on Hugging Face
Paper • 2401.13822 • Published • 1 -
Attention Is All You Need
Paper • 1706.03762 • Published • 144 -
HuggingFace's Transformers: State-of-the-art Natural Language Processing
Paper • 1910.03771 • Published • 33 -
Model Cards for Model Reporting
Paper • 1810.03993 • Published • 7
-
Attention Is All You Need
Paper • 1706.03762 • Published • 144 -
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Paper • 1810.04805 • Published • 33 -
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Paper • 1907.11692 • Published • 10 -
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Paper • 1910.01108 • Published • 24