Business and implementations20:04
Prepare PDF scans for AI: OCR parameters and file naming for Monday
How to set scanning to 300–600 DPI, enable the text layer, and name files as Type_Date-Version_Status before connecting documents to the AI index.

Listen to the article summary
Synthetic voice.Treat this text as practical notes shared for free: we show how it usually works out, not how it must look in your specific case. Open an OCR tool: Azure AI Document Intelligence, Google Document AI, or local OCRmyPDF. In the end, you get a folder of PDF scans equipped with a text layer and names ready to connect to an AI index. Before you begin, check access to the scans folder and permissions to create an OCR processor: in the cloud this requires a paid plan, while in a local installation an account with disk write permissions is sufficient.
Create OCR processor and set scan parameters (1–2 hours)
In the Azure console, open the Azure AI Document Intelligence section and select the option to create an OCR processor. In the processing settings, find the color mode field and switch it to grayscale/black-and-white mode: color does not improve recognition quality, it only increases processing time. Next to the resolution field, set a value between 300–600 DPI. The system defaults to 300 DPI; keep this value if scans are sharp, and raise it to 600 DPI when letters blur on stamps or blueprints. Field details are available in the Document Intelligence documentation.

Process sample of 5,000 files and measure missing text (2–4 weeks)
Run batch processing on a sample of 5,000 files. When complete, open the quality report. Count documents without a text layer: the acceptable threshold is less than 5%. If the result is higher, return to processor settings and raise the resolution to 400 or 600 DPI. In Google Cloud, perform the equivalent operation by selecting an OCR processor from the list and checking the output text field after processing: see the Document AI documentation.
Rename files and flag final versions (1 week)
For every file with a valid text layer, set the name according to the pattern [type]_[ISO-date]_[version]_[status].pdf, for example Invoice_2025-01-15_v2_final.pdf. Mark working versions as draft and exclude them from the index. Detect duplicates by calculating a SHA-256 hash: if two files share the same hash, remove the newer one. Move draft versions older than 12 months from finalization to an archive or delete them.

Set final filter and build index (1 day)
In the indexing tool, select the folder containing the prepared files. In the indexing rule, set the status filter to final, and mark draft files as excluded. Run the initial indexing. Once complete, test three documents from different departments to verify that search returns the correct invoice or contract when querying an excerpt of the text.
One month after implementation, finding an invoice stops taking 40 minutes, as the system returns the document in 20 seconds using the text layer. A pilot on 5,000 files run by 2–3 people saves about 60 hours per year compared to searching network drives manually. Read more about such implementations on the HEXART services page.
Source materials
Sources consulted during research. The text above is our own; these third parties are not responsible for its content and have not authorised it.
