Source transparency
Data & licences
Semantide combines independent open-language projects into a searchable learning experience. Those projects are not affiliated with, and do not endorse, Semantide or lyricalearn.
EDRDG dictionary files
Word, name, kanji, and radical records are adapted from JMdict, JMnedict, KANJIDIC2, KRADFILE, and RADKFILE. Copyright belongs to James William Breen and the Electronic Dictionary Research and Development Group (EDRDG); RADKFILE2 and KRADFILE2 are copyright Jim Rose. These files are used under the EDRDG licence and applicable CC BY-SA 4.0 terms. Semantide consumes JSON conversions published by jmdict-simplified, indexes and normalizes the records, and makes the adapted dictionary data available through search and the JSON endpoint. Those adaptations remain subject to the source terms.
Tatoeba examples
Example sentence text is sourced from Tatoeba. Tatoeba’s downloadable sentence data is generally available under CC BY 2.0 FR, with CC0 and public-domain subsets identified by Tatoeba. Individual examples link back to their source where an identifier is available. Semantide does not redistribute Tatoeba audio.
OJAD pitch accent
Tokyo Japanese pitch accent is provided by OJAD, an educational dictionary from Minematsu Lab at the University of Tokyo. Semantide opens OJAD in a new tab and does not copy, cache, or rehost its data.
KanjiVG diagrams
Stroke-order diagrams are derived from KanjiVG, copyright Ulrich Apel and contributors, under CC BY-SA 3.0.
JLPT reference data
Word-level classifications use Bluskyo’s JSON conversion of Jonathan Waller’s JLPT Resources. Waller’s lists are offered under Creative Commons Attribution; the conversion code and formatting are MIT-licensed. Kanji-level classifications also use David Luzgouveia’s MIT-licensed kanji-data, which identifies Waller’s lists as its JLPT source. These classifications are unofficial study aids—the JLPT does not publish an official post-2010 vocabulary or kanji list.
CC-CEDICT
Chinese word entries are adapted from CC-CEDICT, copyright MDBG and contributors, used under CC BY-SA 4.0. Traditional and simplified writings, pinyin, and English glosses remain subject to those terms.
Wiktionary / Wiktextract
French and English entries are adapted from the English Wiktionary, extracted with Wiktextract and published via kaikki.org. Wiktionary content is available under CC BY-SA and the GNU Free Documentation License.
Frequency lists
Common-word ranking uses Hermit Dave’s FrequencyWords lists derived from OpenSubtitles. Please see that project for attribution of the subtitle corpus.
Unicode Unihan
Han character stroke counts are taken from the Unicode Unihan database where a total-stroke value is published.
Questions or corrections
If a credit is incomplete or a record needs correction, please contact the relevant source project. Source projects keep all rights not granted by their respective licences.