Which language is this?
Character n-grams and perplexity turn language identification into a question of which model is least surprised.
Try the method
Interactive explainerHow it works.
Build the character models
Count character n-grams for English, Spanish, French, Indonesian, Italian, Dutch, Portuguese, and Tagalog.
From the original project.
Saved artifacts · click to inspect
Original artifact ↗
Continue in the source.
Open the source notebook in Jupyter, Colab, or the environment described in the README. Data and model downloads may be required.
git clone https://github.com/eforus-overseer/Language-Detection.gitRead the setup and requirements ↗Project artifacts.
Source links point to the original public repository. Credit belongs to the project authors and the dependencies credited there.