I have a list of PDF files from which I would like to extract a specific piece of information. It is a number that appears somewhere on the page and always starts with “4423…”.
Since I am not able to process or extract data directly from PDF files, my current approach is as follows:
For Each File in Folder
Open the PDF in Word scope
Take a screenshot
Use OCR with UiPath Screen OCR
However, I am facing an issue:
With OCR, I can only interact with clearly defined UI elements, as the selector requires a specific reference to a particular document. This does not work in my case—when opening the second document, the selector can no longer find a match.
If there is a better or more efficient solution, I would greatly appreciate any suggestions.
May I ask why you can’t extract text from the PDF files directly? This way you could just use a “Read PDF” activity and then use Regular Expressions to extract the number you’re looking for.
Alternatively, I would suggest using AI. Use Generative AI activities. You can take a screenshot of the PDF file and then pass it into the GenAI activity. You need to define a prompt of what you’re looking for. But it will easily extract the data you need.
hey @Klang
Best Solution: Use UiPath PDF Text Extraction
For Each file in folder → Extract Text activity (from PDF plugin) → Use Matches activity with regex: “4423\d+” → Get your number you know why
No OCR needed
No selector issues
Works on all PDFs automatically
Fast & reliable
The problem is that some parts of the PDF are not recognized as text, and therefore the PDF text extraction does not always work.
Is there also a solution for this?
Yes @Klang There Are Multiple Solutions
Problem Analysis
Your PDFs contain scanned content or images mixed with text The text extraction only works on actual text layers, not on embedded images or scanned pages.
Solution Use OCR on PDFs (BEST for Mixed Content)
For Each File in Folder
Use “Read PDF with OCR” Activity (instead of Extract Text)
This converts scanned areas to readable text
Apply Regex: “4423\d+”
Get your number
Why this works OCR reads both text AND scanned images Package needed UiPath.PDF.Activities (with OCR enabled)