PDF guides
Extract text from a PDF: text layers, scans and reading order
Save existing PDF text as UTF-8 TXT and learn how to recognize a scanned document that needs OCR instead.
Check whether the PDF already contains text
Two PDFs can look identical while storing their words differently. One may contain digital text; the other may be a photograph of a printed page. Open the document and try selecting a sentence or searching for a distinctive word. Successful selection is a useful initial sign that a text layer exists, although the extracted characters still need checking.
KitForma’s Extract PDF text tool reads that existing text layer. It does not perform optical character recognition, or OCR. An image-only scan cannot produce usable text here merely because the letters look clear on screen. If the selected pages contain no text, the tool reports that situation instead of inventing a transcription.
Extract a focused page range
Use a short representative range first, especially when the PDF has columns, equations or unusual fonts. Page numbers follow the actual PDF sequence. Blank range input selects all pages; 1, 3-5 selects four pages while skipping page 2.
- Choose the PDF in Extract PDF text and enter the pages you need.
- Start extraction. If the document requests a password, provide the legitimate password in the tool’s password field.
- Download the UTF-8 TXT output and open it in a text editor.
- Compare a complete paragraph with the PDF, paying attention to names, accented letters, numbers and line endings.
Selected pages: 1, 3-5
Output page marker: --- 3 ---
The marker identifies the original PDF page, not a new document page.Expect a text file rather than a recreated layout
The result includes page markers and extracted text, not the original fonts, pictures, tables or page design. A two-column article may produce an unexpected reading order. Headers and footers can repeat, and a word split across a line may need to be joined manually.
Check a table row against the source before using its values in a spreadsheet. Plain text does not guarantee that columns still align or that a number remains attached to the correct heading. For quotations or data entry, keep the PDF available as the reference and correct the extraction only after comparing it with the source.
Clean the output without losing meaning
A whitespace cleaner can help with extra spaces, but run it on a copy. Removing every line break can join separate table rows or erase paragraph boundaries. Decide which formatting is noise and which carries meaning before applying a broad transformation.
KitForma accepts a source PDF up to 40 MB and 300 pages, with an extraction limit of two million characters. Work in smaller ranges if the document exceeds the output limit or the browser struggles. The selected file is processed in the browser, and the text download must be saved before leaving the page.
For a scanned document without text, obtain a text-bearing original or use an OCR workflow, then proofread its recognition results. Extracting an OCR-generated layer is possible once it exists, but extraction does not correct OCR mistakes. Always review ambiguous characters such as zero and the letter O before relying on the final text.
KITFORMA
Put it into practice
No account needed. Open an article, then try the matching tool.