Skip to content

PDF work for accountants and bookkeepers

Finance PDF work is mostly extraction and submission: getting numbers out of documents that were designed to be read rather than parsed, and getting bundles into portals with unhelpful limits. Both have specific failure modes worth knowing before a filing deadline.

Getting a table out of a statement, and checking it

A table in a PDF is not a table. It is text positioned to look like one, and reconstructing the rows and columns is inference from those positions. It works well on a cleanly ruled statement and badly on one with merged headers, wrapped descriptions or a running balance in a floating column.

So the step that matters is verification, and there is a fast one: total a numeric column in the spreadsheet and compare it against the closing balance printed on the statement. If they agree, the extraction held. If they do not, a row wrapped or a column shifted, and the difference will usually point straight at it.

Never file or reconcile from an extraction you have not checked against a figure printed on the original.

Statements that are pictures

A statement that arrived as a scan, or was printed and re-scanned, has no text in it at all — extraction will return nothing or nonsense. Try selecting a figure with the cursor: if it does not highlight, it is an image.

Recognition adds a text layer and makes extraction possible, but it is a reading of the scan rather than a transcript. Digits are exactly where recognition is least forgiving, so check totals against the printed figure before anything downstream depends on them.

Portal size limits

Tax and regulatory portals commonly cap uploads at a few megabytes per file and sometimes limit the number of files. A bundle of scanned invoices will exceed that immediately.

Compress before splitting — it is the step that costs nothing in content. If the documents are scans, most of the size is image data and the reduction can be substantial. If compression is not enough, splitting by document rather than by arbitrary page count keeps each part meaningful to whoever opens it.

Bundling a month of invoices

Merging invoices into a single document with page numbers makes the pack referenceable and stops files being lost in an email thread. Adding a header carrying the period or the client reference makes each page identifiable if it is printed and separated.

Where the recipient wants the individual files rather than a bundle, an archive keeps them distinct while still being one thing to send.

Client data

Bank statements, tax computations and payroll are among the most sensitive documents anybody handles, and the identifiers in them — account numbers, sort codes, tax references, PAN — are exactly what fraud needs.

Where a document has to be shared with a scope narrower than the whole file, remove what is not needed rather than covering it. Scanning a document for identifiers first will usually find things you had forgotten were in it.

Working papers and the audit trail

A schedule assembled from twenty source documents needs to be traceable back to each of them, and page numbers running continuously across the bundle are what makes a reference in a working paper resolvable a year later.

Add the numbering after merging, not to the parts. A header carrying the client, the period and the preparer makes a page identifiable if it is printed, photocopied or separated from the bundle — which in practice is what happens to working papers.

Keep the sources separable as well as merged. Extracting a single statement from a bundle is straightforward; reconstructing the original twenty files from one merged PDF is not.