Skip to content
Document extraction

BOL Data Extraction Software

Bills of lading turned into structured fields, including the handwriting

A bill of lading is the record of what actually moved. Pysar.AI reads it — printed or handwritten, clean scan or phone photo — and returns BOL and PRO numbers, shipper and consignee, pieces, weight, freight class, and the exception notes scribbled in the margin.

Overview

The BOL is the hardest document in the stack

Invoices are typed. Rate confirmations are generated. The bill of lading gets handled at a loading dock: stamped, signed, annotated with a corrected piece count, folded into a driver's pocket, and photographed under bad light hours later. It is the document most likely to be unreadable and the one most likely to matter in a dispute.

That is exactly why generic OCR struggles with it. Converting a skewed, low-contrast image into a wall of text produces something no downstream system can use, and the handwritten correction — the part that changes the invoice — is usually the first thing lost.

Pysar.AI treats the BOL as a structured document with known parts. It knows a bill of lading has a consignee block, a freight description table, and signature areas, so it looks for those regions and reports what it finds in each, including whether the handwriting agrees with the printed figures.

OCR gives you text; extraction gives you fields

The practical distinction is what your systems can consume. OCR output is a string; there is no way to ask it for the freight class without writing a parser, and that parser breaks on the next carrier's form.

Extraction returns labelled values: this is the PRO number, this is the consignee, this is 38,500 lbs of class 85 freight across 48 pallets. That structure is what a TMS import, an audit rule, or a cross-document comparison actually needs.

It also makes accuracy measurable. A confidence score per field tells you which values to trust and which to review, which is impossible when the output is one undifferentiated block of text.

Weight and piece count are where money leaks

Most billing disputes on a load trace back to two numbers on the BOL. If the rate was quoted on 38,500 lbs and the BOL records 42,000, there is a reweigh charge coming. If the rate con says 48 pallets and the BOL is annotated down to 46, the delivery is short and the invoice should reflect it.

Because these numbers are extracted rather than eyeballed, they can be compared automatically against the rate confirmation and the carrier invoice on every load — not on the sample a busy team has time to check.

Handwritten changes get special treatment: they are extracted, marked as handwriting, and given a lower confidence score, so a human confirms the correction rather than a machine silently accepting it.

Fields extracted from a BOL

  • BOL number and PRO number
  • Ship date and delivery date
  • Shipper name, address, and contact
  • Consignee name, address, and contact
  • Carrier name, MC number, and trailer
  • Piece count, packaging type, and weight
  • Freight class and NMFC code
  • Commodity description
  • Special instructions and accessorial requests
  • Signature presence for shipper, carrier, and receiver

Caught · Printed piece count reads 48 pallets; the delivery copy is annotated by hand to 46 with 'two short'. The handwritten correction is extracted and flagged so the invoice is adjusted rather than paid in full.

What the extractor handles

Handwriting and annotations

Corrected counts, damage notes, and delivery exceptions are read and marked as handwriting for review.

Poor-quality images

Skewed phone photos, faxes, and low-contrast scans are corrected before reading rather than rejected.

Multi-page and merged PDFs

A merged submission is split into its component documents, each classified before extraction.

Signature detection

Shipper, carrier, and receiver signature areas are checked for actual ink, not just the presence of a signature line.

How it works

  1. 1

    BOL arrives in any form

    Email attachment, scanner output, driver photo, or API upload — all are accepted.

  2. 2

    Image is prepared

    Deskew, contrast correction, and page splitting run before reading, which is what makes bad scans usable.

  3. 3

    Fields and signatures are extracted

    Printed and handwritten values are read, with signature areas checked and confidence scored.

  4. 4

    Compared against the load

    Weight, pieces, and parties are checked against the rate confirmation and carrier invoice for the same load.

Questions

BOL data extraction, answered

What is BOL data extraction?

Reading a bill of lading and converting it into structured fields — BOL and PRO numbers, shipper and consignee, carrier, pieces, weight, freight class, NMFC code, and special instructions — so the data can be validated or pushed into a TMS without typing.

Does BOL extraction work on scanned and photographed bills of lading?

Yes. Drivers photograph BOLs on phones and warehouses fax them; extraction handles skewed, low-contrast, and faxed images, returning a confidence score per field so questionable reads get reviewed.

Can it read handwritten notes and exceptions on a BOL?

Handwritten piece counts, damage notes, and delivery exceptions are detected and extracted, then flagged for review because handwriting carries lower confidence than printed text.

How is BOL OCR different from BOL data extraction?

OCR returns a wall of text with no meaning attached. Extraction returns labelled fields — it knows which value is the consignee and which is the freight class — which is what downstream matching and TMS imports actually need.

Ready to Stop Manual Document Processing?

See Pysar.AI in action with your own documents. Book a 30-minute personalized demo.

Join companies already processing 500K+ documents with Pysar.AI