Data Sources & Attribution
LangDex aggregates linguistic data from multiple open-source projects. We gratefully acknowledge the following sources and their contributors.
Language Metadata
| Source | Description | License | Link |
|---|---|---|---|
| Glottolog | Language family trees, geographic data, and metadata for 8,000+ languages | CC BY 4.0 | glottolog.org |
| CLDR | Localized language names, script associations, and locale data | Unicode License | cldr.unicode.org |
| ISO 639-3 | Language code standards | Public Domain | iso639-3.sil.org |
Cross-Lingual Semantics
| Source | Description | License | Link |
|---|---|---|---|
| PanLex | Cross-lingual lexical database linking meanings across 5,700+ languages | CC0 1.0 | panlex.org |
| Concepticon | Standardized concept lists for cross-linguistic comparison | CC BY 4.0 | concepticon.clld.org |
Dictionaries & Lexical Data
| Source | Description | License | Link |
|---|---|---|---|
| Kaikki (Wiktionary) | Multilingual dictionary extracts from Wiktionary | CC BY-SA 3.0 | kaikki.org |
| JMDict | Japanese-Multilingual Dictionary | CC BY-SA 4.0 | edrdg.org |
| KanjiDic2 | Japanese kanji dictionary | CC BY-SA 4.0 | edrdg.org |
Morphology & Phonology
| Source | Description | License | Link |
|---|---|---|---|
| UniMorph | Morphological paradigms for 100+ languages | Varies by language | unimorph.github.io |
| PHOIBLE | Phonological inventories | CC BY-SA 3.0 | phoible.org |
| IPA-Dict | IPA pronunciations for multiple languages | MIT | github.com/open-dict-data/ipa-dict |
Example Sentences
| Source | Description | License | Link |
|---|---|---|---|
| Tatoeba | Multilingual sentence corpus with translations | CC BY 2.0 FR | tatoeba.org |
Frequency Data
| Source | Description | License | Link |
|---|---|---|---|
| OpenSubtitles | Word frequency lists derived from subtitle corpora | Open Data | opus.nlpl.eu |
WordNet & Semantic Networks
| Source | Description | License | Link |
|---|---|---|---|
| Open Multilingual Wordnet | Linked wordnets for 100+ languages | Various (mostly CC BY) | omwn.org |
| Princeton WordNet | English WordNet | WordNet License | wordnet.princeton.edu |
License Compatibility
LangDex combines data under various open licenses. Due to the ShareAlike requirements of CC BY-SA sources (Kaikki, JMDict, KanjiDic2, PHOIBLE), derivative works incorporating this data are released under CC BY-SA 4.0.
Your Obligations When Using LangDex Data
- Attribution - Credit LangDex and the original data sources
- ShareAlike - If you modify and redistribute, use a compatible license
- No additional restrictions - You may not apply DRM or legal terms that restrict others
How to Cite
Academic Citation
LangDex: Universal Linguistic Database Data aggregated from: Glottolog, PanLex, Wiktionary (via Kaikki), JMDict, UniMorph, PHOIBLE, Tatoeba, and others. Available at: https://langdex.co
Attribution for Applications
Linguistic data provided by LangDex (langdex.co), aggregating data from Glottolog, PanLex, Wiktionary, JMDict, Tatoeba, and other open sources. Licensed under CC BY-SA 4.0.
Data Updates
This database represents a static snapshot of the source data. We periodically update from upstream sources but do not guarantee real-time synchronization.
Contact
For questions about data licensing or attribution, contact:
- Email: legal@langdex.io
Last updated: 20 December 2025
See also: Terms of Service | Privacy Policy