The scanned PDF problem: finding sensitive data in Jira’s hardest files

There are two kinds of PDF, and the difference matters enormously for data protection. A digital PDF has a text layer — the words are real text you can select, copy, and search. A scanned PDF is a picture of a document: someone ran paper through a scanner, and the result looks like text but is actually an image. Contracts, invoices, signed forms, ID documents, and medical records overwhelmingly arrive as the second kind. They’re also where some of the most sensitive data in Jira hides, and where ordinary scanning quietly fails. If you need to find sensitive data in scanned PDFs in Jira, you have to treat them as images, not documents.

Why scanned PDFs slip through

The danger of a scanned PDF is that it looks searchable when it isn’t. A tool that extracts text from PDFs will happily process a digital PDF and return nothing for a scanned one — not an error, just an empty result — and it’s easy to mistake that silence for “no sensitive data here.” Jira’s search indexes fields, not file contents, so it never touches either kind. Malware scanning sees a clean file. The net effect is false confidence: you believe you’ve scanned your PDFs, when in fact you’ve only scanned the ones that happened to have a text layer, and missed exactly the contracts and forms most likely to carry regulated data.

OCR is the only way in

To read a scanned PDF you have to recognise the characters in the image — OCR. Attachment Scanner for Jira runs OCR on scanned PDFs (and images) using an AI vision model, converting the picture of a document back into searchable text, then matches your patterns against it. A national ID on a scanned passport, an IBAN on a scanned invoice, a diagnosis on a scanned medical form — all become detectable. Digital PDFs with a real text layer are read directly, without OCR, so the app handles both kinds; the crucial point is that it doesn’t go blind on the scanned ones the way text-only tools do.

Full scan vs document-only — an important distinction

This is the one place where the app’s two scan modes need care. A full scan reads everything, including images and all PDF types, using OCR where needed — this is the mode that opens scanned PDFs. A document-only scan covers Office and text files and deliberately skips all PDFs, even ones with a readable text layer, because PDF extraction routes through the OCR-capable service; skipping them is what lets document-only run with zero credits and never contact the OCR service. The practical takeaway: if scanned PDFs are your concern, you must run a full scan. Document-only is the right choice when you explicitly want a free, GPU-free pass over Office and text files — but it will not look inside any PDF.

Scoping, credits, and size limits

Because OCR consumes one credit per OCR’d PDF page, scanned PDFs are the most credit-intensive files you’ll process, so scoping matters. Use JQL to target the projects and date ranges where scanned documents accumulate — legal, finance, HR, customer verification — rather than scanning the whole instance at once. Keep the per-file cap in mind too: PDFs up to 50 MB are supported. A focused scope keeps a scanned-PDF scan both affordable and reviewable.

Reviewing with confidence — including what was skipped

Each match shows the issue key, file name, extraction type (OCR for scanned PDFs), matched text, and context, so you can confirm findings without opening every document. Just as importantly, the results surface files that were skipped or threw warnings rather than hiding them — so if a PDF was too large or failed to process, you know it wasn’t covered. That visibility is what prevents the false confidence scanned PDFs are so good at creating: you can state precisely what was read and what wasn’t.

Privacy, limits, and the takeaway

Scanned documents are often the most sensitive files you hold, so the processing model matters: OCR on dedicated EU/EEA GPU hardware, no public AI service, attachments processed in memory and discarded, and only matched snippets stored in Atlassian’s Forge storage, isolated per site. OCR accuracy depends on scan quality, scanning is on-demand rather than continuous, and the app is Jira Cloud only for now. The core message is simple: scanned PDFs are where regulated data hides and where text-based tools fail silently. Reading them takes OCR and a full scan — and once you have both, the files you were most worried about become the files you can actually account for. You can start a free 30-day trial from the Atlassian Marketplace.

Want
to know more?

Contact us to talk to our experts and have all your questions answered.

Request
free offer