Benchmark · Measured 2026-08-25 · 6 min read

German speech to text, with the benchmark that picked the engine

Drop in a German recording and get a timestamped transcript with speaker labels — then translate it into any of 63 languages, keeping the German original beside it.

Most pages that sell German transcription either quote someone else’s benchmark or give no number at all. We ran our own: every clip of the German validation split of FLEURS through both of our engines, scored against the published reference. Parakeet-TDT-0.6b-v3 won at 3.82% word error rate and German is routed to it. The table, the method and the sample size are all below — including the part where twenty clips is a small sample.

Free

3 files a day free · 5-minute preview each · no credit card

How accurate is German speech to text?

We measured it rather than estimating it. Every clip in the German validation split of FLEURS — Google’s public read-speech corpus — went through both of our engines, scored against the published reference with punctuation stripped and case folded. German is routed to the one that won.

Engine Word error rate Wrong words Clips with no error
Parakeet-TDT-0.6b-v3the engine we route German to 3.82% 15 of 393 10 of 20
Whisper large-v3-turboour general-purpose engine 4.33% 17 of 393 10 of 20

Parakeet won on accuracy and was not close on speed: it decoded all 238 seconds of audio in 2 seconds on an idle GPU. 10 of the 20 clips came back with no errors at all, and the whole error count sits in five recordings.

Reading those five word by word changes what the number means. Several “errors” are spelling conventions rather than mishearings: the reference writes kruger where we write krüger, simbabwe where we write zimbabwe, and m where we expand meter. Intelligibility is better than 3.82% suggests.

For outside context on the same corpus, ElevenLabs publishes 2.7% for Scribe, 3.9% for Gemini Flash 2, 4.5% for Whisper large-v3 and 9.1% for Deepgram Nova 2. Our own Whisper run landing within 0.2 points of their Whisper row is the reason to trust the method rather than the marketing.

Method and limits. FLEURS German (de_de), validation split; 20 clips, 393 words, 238 seconds of audio; single run, measured 2026-08-25. The Whisper rows went through the production API with language auto-detection, Parakeet was called directly with German pinned — same audio, same reference, same scoring, but not an identical harness. Twenty clips fix the order of magnitude, not the third decimal. If you want the general picture rather than the German one, we explain how transcription accuracy is measured and where AI still trails a human transcriber.

Which AI model transcribes German best?

Parakeet-TDT-0.6b-v3, on our clips, by half a point of word error rate. That is 15 wrong words against 17 — a small margin, which is exactly why it is worth measuring instead of guessing. Choosing the engine per language is the point of running the test at all: the model that wins German is not automatically the model that wins English. The language itself is detected automatically, and you can pin it by hand when a recording switches between German and English.

Can I translate a German transcript into another language?

Yes, and the order matters. The recording is recognised as German first; the finished transcript is then translated into any of 63 languages — English, Spanish, French, Chinese, Japanese and 58 more. Translation runs as a separate step on completed text, so the German original stays beside the translation and you can translate again into a second language without re-uploading the file.

This is the part most competitors skip. A page that transcribes German but cannot hand you the English is only half the job when the recording is an interview you need to quote in another language.

Does it identify different German speakers?

Speaker diarization labels up to 10 speakers automatically and marks every segment with who said it — which is what turns a German interview or Zoom call into something readable instead of one unbroken block. Rename a speaker once and the name changes everywhere in the document, including the exported file and any translation you generate afterwards. There is more on how speaker diarization decides who is speaking, and on recording interviews so it has a chance.

Editing and exporting a German transcript

The transcript opens in an editor with the audio beside it: click a word, hear it, fix it. Timestamps stay attached while you edit, so corrections never break the alignment between text and audio.

Exports come as TXT, DOCX, PDF, SRT and VTT — a document, a PDF, or subtitles that go straight onto a German video — the SRT export is the one that carries the timings.

What German audio and video formats can I upload?

13 container formats: MP3, WAV, M4A, AAC, FLAC, OGG and WMA for audio; MP4, MOV, AVI, MKV, WebM and MPEG for video. A German video file is transcribed directly — there is no need to extract the audio track first. The full list of supported formats covers what each one is good for.

Is there a free way to transcribe German audio?

The free plan covers 3 files a day with a 5-minute preview of each — enough to check German accuracy on your own recording rather than on our benchmark. No credit card is required for the first file, and the preview uses the same engine as a paid transcription, so what you read in the preview is what the full transcript will look like. Longer German files need a paid plan — the plans are here.