Skip to main content Scroll Top

Unstructured Vendor Data Normalizer

Mid-Size Manufacturer · ERP & Procurement

PROBLEM

A mid-size manufacturer managing 400+ vendor relationships receives a steady volume of shipping notices, invoices, packing slips, and PO confirmations through email — much of it from smaller suppliers without EDI or standardized document formats. Procurement and receiving staff were spending 15+ hours per week opening attachments, identifying the related purchase order, and manually re-keying vendor data into their ERP. Inconsistent document layouts and repetitive entry also created transcription errors that surfaced later as receiving discrepancies, reconciliation issues, and additional accounts payable work.

The manufacturer’s ERP had no practical API for creating these records directly. It did, however, support structured bulk imports, creating an opportunity to automate document intake and normalization while keeping final ERP entry within the system’s existing controls.

OBSTACLES

  • Vendor documents arrived in multiple formats, including native PDFs, scanned documents, multi-page packing slips, spreadsheet attachments, and order details contained directly in email bodies.
  • Suppliers used different terminology and layouts for the same information. A reusable vendor mapping layer was developed to normalize variations such as “Qty,” “Units,” and “No. of Pieces,” as well as “Ship Date,” “Dispatch Date,” and similar vendor-specific fields.
  • Extracted records had to be reconciled against open purchase orders. Matching logic accounted for partial shipments, line-level quantity differences, duplicate documents, and cases where supplier SKUs or descriptions did not exactly match internal part numbers.
  • Extraction confidence alone was not enough to approve a record. A document could be read correctly but still reference an invalid PO, unexpected vendor, unmatched line item, or quantity outside the manufacturer’s accepted tolerance.
  • Because the ERP could not accept direct API writes, the workflow needed to produce validated, ERP-compatible import records rather than bypassing the manufacturer’s existing receiving and approval process.
  • Staff needed to be able to correct individual fields, resolve exceptions, and revalidate a record without restarting document processing from the beginning.

OUTCOME

An AI-assisted intake workflow monitors the shared procurement inbox and processes incoming vendor documents and message content. It identifies the vendor and document type, extracts required fields including PO number, vendor ID, line items, quantities, unit prices, shipment dates, carrier, and tracking information, then normalizes the results into a consistent internal schema.

The normalized data is validated against open purchase-order data before being staged for import. Confidence scoring is combined with rule-based checks for PO validity, vendor matching, line-item reconciliation, duplicate detection, and quantity tolerances. Records that pass validation are added to a Google Sheet review queue in an ERP-ready structure; questionable records are held with the specific field or validation rule requiring attention clearly identified.

Procurement staff can review the original document alongside the extracted record, correct individual values when necessary, and approve batches for export. Approved records are converted into the CSV format required by the ERP’s existing bulk-import process. Each staged record retains its source document reference and processing history, providing a traceable path from the original vendor communication through final ERP import.

Automation stops at the appropriate boundary: records without a reliable PO match, with unresolved line items, or below the configured confidence and validation thresholds are never silently passed into the import queue.

RESULTS

  • Approximately 75–80% of incoming vendor documents could be extracted, matched, and staged without manual data entry.
  • Weekly manual processing time fell from 15+ hours to roughly 3 hours, shifting staff effort from repetitive transcription to reviewing exceptions and approving completed batches.
  • Data-entry and reconciliation errors dropped substantially because PO mismatches, missing fields, and questionable quantities were identified before ERP import rather than during receiving or AP reconciliation.
  • Routine documents that previously waited for manual entry could be prepared for ERP import the same day they arrived, reducing a process that frequently stretched across 2–3 business days.
  • The normalized document structure also created a reusable data layer for later vendor-performance and procurement analytics without requiring changes to the legacy ERP.