But if you believe in your manual heuristics enough to ship them, you must alrea...

But if you believe in your manual heuristics enough to ship them, you must already have a body of tests that you're happy with, right?

Also seems like this is a case where generating synthetic data would be a big help. You don't have to use only real-world documents for training, just examples of the sorts of things real-world documents have in them. Make a vast corpus of semi-random documents in semi-random fonts and settings, printed from Word, Pandoc, LaTeX, etc.