Writing & text
Make a local transcript with Whisper and check it against the audio
Prepare a short recording, run a CPU transcription, review uncertain passages and save corrected text and captions without a transcription API.
Choose a small recording and a clear permission boundary
Start with a 30–60 second recording of your own voice, or material you have permission to process. Write the expected language, approximate duration and intended use beside the file. Keep an untouched original. A short test catches installation, language and audio-stream mistakes before you commit a long interview to a slow machine.
This workflow uses the open-source Whisper package, not a paid transcription API. Its code and weights use the MIT license. Software and model downloads need an internet connection initially; later local processing still uses your computer's memory, storage and electricity. No account, API key or subscription is part of this path. The commands below are a recipe checked against official documentation; they were not executed as an audio benchmark for this guide.
Set up once and inspect the input
Use the official Whisper installation instructions for a compatible Python environment and FFmpeg. In a separate virtual environment, install openai-whisper with pip; record the installed package and Python versions. Do not paste an arbitrary installer from a search advertisement. FFmpeg's binary license depends on the build: its base license is LGPL, while optional GPL components change obligations. This guide distributes no program binaries or model weights.
After installation, run the help and version commands below. If the recording has several audio tracks, listen to them and choose the intended one before transcription. The example extraction selects the first audio stream and makes a mono 16 kHz WAV working copy. It does not remove background noise or recover lost speech. Using a new output filename and -n avoids overwriting an existing file.
python -m pip show openai-whisper
ffmpeg -version
whisper --help
ffmpeg -n -i "source.mp4" -map 0:a:0 -vn -ac 1 -ar 16000 -c:a pcm_s16le "speech.wav"OpenAI Whisper: installation, models and license ↗FFmpeg: stream selection and command reference ↗FFmpeg: license and legal considerations ↗
Run a bounded CPU transcription
For a Turkish sample, use a multilingual model and specify Turkish. The command requests transcription in the spoken language, CPU processing, two threads and text/timing outputs in a new directory. Start with base as a practical small trial; it is not an accuracy guarantee. First use may download weights. A device that cannot finish the short trial should use a smaller model or manual transcription before attempting a large file.
Whisper's translate task targets English; it is not a general Turkish dubbing command. The turbo model is not trained for translation. Transcription produces text and timing estimates, not a new spoken voice, reliable speaker identification or proof that a person said a disputed phrase. Keep those tasks separate.
whisper "speech.wav" --model base --language Turkish --task transcribe --device cpu --fp16 False --threads 2 --output_format all --output_dir "transcript-run-01"OpenAI Whisper: installation, models and license ↗Whisper command-line options in the official source ↗
Correct the transcript while listening
Open the text and replay each segment. Mark uncertain passages with timestamps rather than filling them from context. Prioritise names, numbers, negatives and domain terms. A confident-looking word can be wrong. The official model card documents hallucinated or repeated text and uneven performance across languages and speakers; a plausible paragraph during silence deserves particular scrutiny.
For your first pass, keep an error log: timestamp, machine text, heard text, confidence of the human decision and action. If the phrase is inaudible, retain an uncertainty marker. A second reviewer can compare the marked passages with the recording. A writing model may suggest punctuation after this pass, but should not reconstruct missing speech or invent a quotation.
00:12.4–00:14.0 | draft: “fifteen” | heard: “fifty” | replay and confirm
00:24.0–00:27.2 | draft: a full sentence | audio: silence | remove hallucinated text
00:31.5–00:33.1 | name unclear | keep [unclear name] | ask speaker if permittedMake captions readable and faithful
Correct timing as well as words. A cue that appears before a reveal can spoil the explanation; a cue that stays after the speaker finishes can confuse a change of voice. Divide text at natural phrase boundaries and preview on the intended screen. For meaningful sounds, add concise descriptions when needed to understand the scene. Do not label a speaker by identity unless you know it.
W3C distinguishes same-language captions from translated subtitles and explains that accessible captions include necessary non-speech audio. Machine output therefore needs an editorial pass. Preserve the corrected transcript separately from SRT or VTT so later timing changes do not lose your verified wording. The cue below is an authored format example, not measured speech timing.
1
00:00:01,000 --> 00:00:04,000
Keep the original recording.
2
00:00:04,500 --> 00:00:07,000
Check each number while listening.Package evidence and limits with the files
Save the original, working audio, raw transcript, corrected transcript, corrected caption file and a short run note. Include model name, package version, language, command, reviewer and unresolved timestamps. A hash can identify which file was reviewed; it cannot prove that the transcript is accurate. If you need a fully offline run, fetch the required model first and test with networking disabled under your own device policy.
If FFmpeg reports no audio stream, inspect the source rather than adding random flags. If output is empty or repetitive, listen for silence and confirm the language. If processing is too slow, shorten the sample before changing many settings. Delete temporary copies according to your agreed retention plan; do not assume an editor, synced folder or backup service has done this for you.
KITFORMA
Put it into practice
Reading guides is free and needs no account. Linked tools explain any account or Pro requirements.
Further reading
Sources last reviewed:
KitForma prepared this guide with AI assistance. Examples illustrate a workflow; they are not measurements of real sales, search volume or success. Check the linked sources and the result with your own file.
How to use this guide
Published and maintained by KitForma. The linked references explain the relevant formats and definitions. Examples use sample inputs; they do not establish a speed, quality or compatibility guarantee for your files. Check the tool’s stated limits and inspect your downloaded result.