Why a computer sometimes misreads a document

It isn't that the program is stupid, it's that a crooked scan is crooked for it too

A document lying crooked on the glass of an open scanner

Two files that look alike and aren't

When you hand a PDF to a program so it can pull the data out, the thing that decides almost everything is how that PDF was born, and there are two roads

A PDF born on a computer, for instance the one your supplier's business software prints by itself when it issues the invoice, has the text written inside it, and the program that opens it has nothing to interpret, it takes the characters and carries them off exactly as they are

A scanned PDF is another thing entirely, it's the photograph of a sheet of paper, inside, on its own, there is no text at all, there are light dots and dark dots, and to work out that a group of dots is an eight somebody has to look at them and decide, and the one deciding is the optical reading program

You can check the difference in ten seconds, open the file and try to select a line with the mouse. If it draws a rectangle over it there is no text inside to take, almost always because what you have is a photograph. If it highlights instead, the text is usually really there, but watch out, a scan that has already been through an optical reader lets you select it too, because it has a layer of characters laid over the image, and that layer is already a reading with its own errors inside it, as the guide on a PDF born on a computer or scanned explains

Why a photograph reads badly

A photographed document reads badly for reasons that are ordinary and very physical, the sheet went into the scanner crooked and the lines run uphill, the copy is a photocopy of a photocopy and the characters have blurred together, the blue stamp landed right on top of the taxable amount, the fold of the envelope runs through the middle of the number

Then there are the resemblances that fool a tired person too, the zero and the letter O, the one and the seven written without a bar, the thousands separator and the decimal comma that swap places depending on who printed the document

Whatever reads that sheet doesn't have your head, it doesn't know that an invoice from that particular supplier can't be for twelve thousand euros, it sees marks and tries to say what they are

Reading the characters and knowing where they sit are two different problems

Even when the optical reading gets every single character right, the harder problem is still there, knowing that a number is the total and not the taxable amount, and that a figure sits in the quantity column and not in the price one

In tables with no separating lines, and there are a great many of them, the columns exist only in the eye of whoever is looking, the program sees numbers scattered over a page and has to work out on its own which one belongs to what

And then there is the most annoying nuisance of the lot, a supplier who changes the invoice template, moves the total to the left or gives it another name, and what the system had learned about those documents no longer holds

Where the work is really decided

Anyone who promises you a perfect reading is selling you a word, because the perfect reading of a photographed sheet doesn't exist, and what counts is what happens right afterwards

A document that has been read goes through a series of checks before a single value touches the business software, the taxable amount plus the tax has to make the total written on the document, the VAT number has to be valid and has to match a supplier already in your records, the date has to be a sensible one and not in 2031, the amount has to fall inside the thresholds you decided, and the document number must not already be recorded

A check like that is a net, it doesn't understand the document, but it can tell when the sums don't add up

An example with made-up numbers

Say a scanned invoice comes in, taxable amount 1.240,00, tax 272,80, total 1.512,80

The optical reading catches a smudge on the four and reads the taxable amount as 1.248,00, and the sum of what it read, 1.248,00 plus 272,80, makes 1.520,80, while the total written on the document says 1.512,80

They don't match, and that is all it takes, the system doesn't need to work out where the error is, it's enough for it to know there is one, it stops and puts the document in a queue for a person, with the image and the value that doesn't add up right there in front of them

Without that check, that invoice would have gone into the accounts with eight euros too many and nobody would have noticed until the reconciliation, if ever

What you can do to make it read better

The thing that solves the problem at the root isn't technical, it's a phone call, ask the suppliers who send you photographed PDFs whether they can send you the native file or the XML, plenty of them have it and don't know, and that file reads exactly

For whatever is left to scan the usual dull rules apply, the sheet straight, the original instead of the photocopy of a photocopy, and the stamp put somewhere free instead of over the numbers

And when you show a process to whoever has to automate it, don't hand over the tidy documents, hand over the crooked ones too, because those are the ones that decide whether the work holds up

The question to ask whoever offers you automatic reading

Don't ask how well it reads, because the answer will be a percentage you can't verify, ask what it does when it isn't sure

If the answer is that it stops and puts the case in a queue with the reason written down, you're talking to someone who has already seen a crooked invoice, and if you want to see how these nets are made there's the guide on how you stop an automation from doing damage

A system that would rather stop than guess asks you for a few minutes of a person's time now and then, and saves you the errors that turn up too late, when they cost a great deal more

What people usually ask us

Try selecting a line with the mouse, if the text highlights the PDF was born on a computer and the data comes out exact, if the mouse draws a rectangle over an image you have a scan, that is a photograph that has to be interpreted
No, an automation built properly writes nothing before it has verified that the sums add up, and when a total doesn't match or a VAT number isn't in your records the system stops and puts that document in a queue for a person
Because it would be a made-up number, it depends on how your documents are made, on how straight the scan is and on how many stamps land on top of the numbers, that's why Muffin Suite would rather talk about the checks that sit around the reading
Yes, a stamp over an amount covers exactly the marks needed to recognize it, and it's one of the most common reasons a scan reads badly, asking whoever does the stamping to leave the area of the amounts free costs nothing and helps
Yes, when a supplier can send the native file or the XML the reading problem disappears at the root, because the fields are already defined and there is nothing to interpret, and it's the road Muffin Suite suggests trying first
It depends, if your scans come out straight and legible the scanner you have is perfectly fine, the problem is generally not the machine but the photocopies of photocopies and the sheets fed in crooked, and that gets sorted out without spending anything

Show us what you still do by hand

Describe the process or send a short screen recording and we'll tell you what can be automated

WA