PDF to text
Reads the text layer and does not OCR. It runs in this tab. The file is not uploaded.
A $6 day pass is 24 hours, charged once, and it does not renew. Pricing
No ad script loads from this choice. It stays in this browser.
What this text keeps
PDF to text reads the text layer already in the PDF and downloads document.txt. Each page keeps items that have a string str and a transform array. x is transform index 4 and y is transform index 5. An item whose str trims to empty is ignored. Lines are baselines within 2 units: items whose y values are within 2 share a line. Lines are ordered by y descending, so the top of the page is first. Within a line, items are ordered by x ascending, left to right, and their str values are joined with a single space. Lines of one page are joined with a newline. Pages are joined with a blank line. A second column that shares a baseline lands on that same line. This page does not reconstruct columns, fonts, or pictures.
The reading runs in this tab with the open-source workers. The file is not uploaded. There is no GridFS copy. Closing the tab drops the copy. PDF to Word is /pdf-to-word and that page uploads. OCR is /ocr-pdf. A picture of each page is /pdf-to-jpg.
There is no published page cap. The tab has to hold the file. If the browser cannot keep the PDF, the read cannot finish. There is no account on this page. The day pass is $6, charged once, for 24 hours, and it does not renew. The year is $60 for 365 days. Neither charge is required to read a text layer here. The free server file is one PDF to Word conversion, up to 20 pages, and it is not this page and not OCR. OCR is not free.
An empty file, and any file that cannot be opened, fails with "This PDF could not be read." A password-protected PDF fails with "This PDF is protected. Unlock it in this tab first." This page does not guess a password. A PDF with no pages fails with "This PDF has no pages." A scan with no text layer fails with "This PDF has no text layer. A scan needs OCR, which this page does not do." The worker returns the string. It does not return a PDF and it does not call OCR.
Steps
- Drop the file. It stays in this tab.
- Set the one option on the page, if it has one.
- Download the result. Closing the tab drops the copy.
Questions
What does document.txt contain?
The text layer. Lines are baselines within 2 units, top of the page first, left to right, joined by a single space. Pages are separated by a blank line. The file is named document.txt. It is not a Word file and not a PDF. Columns, fonts, and pictures are not reconstructed. A second column that shares a baseline lands on that same line.
Does PDF to text upload the file?
No. It runs in this tab with the open-source workers. The file is not uploaded. There is no GridFS copy. Closing the tab drops the copy. PDF to Word at /pdf-to-word is the page that uploads.
Will two columns stay as two columns?
No. Lines are baselines within 2 units, left to right. A second column that shares a baseline lands on that same line. Columns, fonts, and pictures are not reconstructed. The download is document.txt, not a Word file and not a PDF.
What if the PDF is a scan?
A scan with no text layer is refused with "This PDF has no text layer. A scan needs OCR, which this page does not do." OCR is a different page, /ocr-pdf, and it is a paid upload. The day pass is $6, charged once, for 24 hours, and it does not renew. The year is $60 for 365 days. OCR is not free. The free server file is one PDF to Word conversion, up to 20 pages, and it is not this page and not OCR. Pictures of a page are /pdf-to-jpg, which stays in this tab and does not read the letters.