What it is
Python libraries for PDF. pdfplumber carefully extracts text and tables; PyMuPDF can also change a page: erase and draw blocks without touching the rest.
How we use it
pdfplumber reads the PDF statements of four banks in the finance assistant. PyMuPDF stamps an accident commissioner's contacts onto an electronic car insurance policy: the form number, the data and the QR code stay untouched, and the file grows from 380 to 520 KB rather than to 6.5 MB.
Where it helps a business
- Staff copy numbers from PDFs by hand: statements, invoices, certificates.
- Your contacts or a stamp need to go onto a finished PDF without damaging the document.
How we use it
- Reading statements. In the finance assistant pdfplumber reads PDF statements from four Russian banks; the bank, account and card are detected from the file itself.
- Editing a policy. In the insurance policy bot PyMuPDF analyses the page (text blocks, images, the QR code) and places an accident commissioner's card without entering protected zones: the form number, the policyholder's data, VIN, QR code and signature block.
Common problems
- The bank changes its statement layout. Parsing is tied to the current PDF layout; when a bank changes it, the sums stop matching the totals and the parser has to be fixed.
- The file balloons. By default erasing text also redrew images, and a policy grew from 380 KB to 6.5 MB. Now only text is erased and fonts are trimmed to the glyphs used: the output is about 520 KB.
- The insurer's signature. Once the file is modified, the built-in electronic signature no longer verifies; authenticity is confirmed through the QR code, and the bot warns about this.
When you do not need it
If a bank or service provides data as CSV, XLSX or through an API, there is no need to parse PDFs: the finance assistant has a fallback CSV and XLSX parser for other banks.