← Umlaut

Sources & Licenses

Umlaut is free and has no budget, so every piece of content is built from sources that are genuinely free to reuse in a product, not just free to browse. This page lists what's actually used in the live app today and how.

Used in the shipped product

FreeDict (deu-eng / eng-deu)

License: GPL v2+

Backbone bilingual dictionary for the vocabulary word list. The dictionary data is kept as its own clearly-licensed file, not merged into Umlaut's own source code.

hermitdave/FrequencyWords (de_50k.txt)

License: CC BY-SA 4.0 (data), MIT (code)

Word-frequency ranking used to sort vocabulary into A1/A2/B1 difficulty tiers.

UniMorph (unimorph/deu)

License: CC BY-SA 3.0

Cross-referenced against inflected German word forms to catch and fix a data-quality bug where capitalized noun homographs of common words (e.g. „Aber“ the rare noun vs. „aber“ the conjunction) were being selected incorrectly.

Grimm Grammar (COERLL, UT Austin)

License: CC BY 4.0

Informed the scope and sequencing of Umlaut's grammar curriculum. Explanations themselves are original writing, not copied or adapted text.

Thorsten-Voice

License: CC0 (public domain dedication)

Source of most listening clips and speaking-shadowing reference clips — real recorded human speech, not synthesized audio. Some newer listening clips are TTS-generated instead (see OpenAI TTS entry below); both are labeled honestly, never presented as the other.

OpenAI text-to-speech

License: Commercial API, paid per use

Generates audio for some listening clips and for AI Coach-generated listening content (Pro-gated). Synthesized, not recorded human speech — never presented as the latter.

LanguageTool

License: LGPL 2.1 (the open-source core)

Self-hosted grammar/spelling/punctuation checker powering the Writing module's feedback. Umlaut runs its own instance; no text is sent to a third-party service.

ts-fsrs

License: MIT

Implements the FSRS spaced-repetition scheduling algorithm used for vocabulary, grammar, reading, and listening review.

OpenAI language models

License: Commercial API, paid per use

Powers the AI Coach layer (Pro-gated): AI-generated reading passages and listening scripts, structured feedback on writing and speaking submissions, AI conversation practice, and quick translate.

Reference only — nothing copied

Goethe-Institut exam format

All rights reserved on their materials

Umlaut's writing-task shapes and reading/listening length targets are modeled on the real Goethe A1–B1 exam's structure (task types, word-count anchors). No Goethe-Institut text, wordlist, or exam content is copied or stored.

MERLIN corpus

Research/reference use

Consulted only as a calibration reference for what genuine B1-level learner German looks like, while authoring original reading passages. No MERLIN text is published in the app.

This page grows as new content sources are added. See the project's own docs/design.md for the full research trail, including sources that were investigated and rejected or not yet integrated.