How to extract text from a PDF (and what to do when you can't copy it)
Copy or download the text of a PDF, why scanned PDFs have nothing to copy, why pasted text sometimes comes out garbled, and how OCR helps.
You need a paragraph from a PDF, or the whole thing as plain text so you can edit it, search it or paste it into an email. Sometimes selecting and copying works perfectly. Other times nothing highlights, or what you paste is gibberish. Understanding why tells you which tool to reach for.
Two kinds of PDF
PDFs look alike on screen but are built in two very different ways:
- Text PDFs are made from word processors, spreadsheets, websites and most modern software. The words are stored as text, so they can be selected, copied and searched.
- Scanned PDFs are pictures of pages from a scanner or phone camera. They look like text but contain none, only the image of it. Nothing can be selected.
A quick test: open the file and try to select a word, or press Ctrl+F and search for a word you can see. If the word is found or highlights, the PDF has text. If not, it is a scan.
Extract text from a text PDF
- Open Extract Text from PDF and add your file.
- If you need only part of it, enter pages such as
1-3, 5. Leave it blank for all pages. - Choose Keep line breaks to match the layout of the page, or Join lines into paragraphs so the text flows when pasted into a document. The second option also re-joins words that were split with a hyphen at the end of a line.
- Click Extract text. The text appears in a box where you can edit it. Copy it, or download it as a .txt file.
Tick “Mark where each page starts” if you want to see where one page ends and the next begins, which helps when you need to cite page numbers.
Why the text sometimes looks wrong
- Columns and tables. PDFs store text in the order it was drawn, which may not be reading order. Two-column pages and tables can come out interleaved. Extract one page at a time, or one column’s worth, and reorder by hand.
- Garbled characters. Some PDFs, particularly older ones and many Hindi documents made with non-Unicode fonts, map letters to the wrong characters internally. Copying them gives nonsense even though the page looks right. This comes from how the PDF was made, not from the tool, and the usual fix is to read the page as an image with OCR.
- Missing spaces or extra breaks. Layout-heavy PDFs sometimes need a quick clean-up after extraction.
When the PDF is a scan: use OCR
If the tool reports that no text was found, you have a scanned PDF. The OCR Wizard in Nyay Sahayak reads text from scanned images and pages. You select the region you want, and it returns editable text, which you can then correct. OCR is never perfect, so proofread anything important, especially numbers, names and dates.
For better OCR results:
- Scan at 200 to 300 DPI. Lower resolutions lose detail, and much higher ones add file size without much gain.
- Keep pages straight and well lit. If a scan is sideways, fix it first with Rotate PDF or Organize Pages.
- Trim dark borders with Crop PDF Pages.
Other uses of extracted text
- Count words with the figure shown under the box, which is handy for assignments and translation quotes.
- Search a long document in a plain text editor, where search is fast and shows every match.
- Reuse paragraphs in a letter or a reply, with the layout cleaned up.
The PDF is read on your own device and not uploaded, which matters for contracts, case papers and personal documents. See how in-browser tools keep files private.