A supplier redesigns its material certificate, and the extraction that worked last month starts leaving batch records blank or filling them with the wrong values. Your automation can keep working if you treat each layout as its own tested route: identify the form version, send it to the extractor built for it, map the fields to your agreed supplier, product and batch records, and prove the old layouts still work before going live.

Version-aware extraction is a document workflow that identifies which version of a form has arrived, then applies the extractor, field mapping and checks approved for that version. A layout the router does not recognise goes to a reviewer instead of the nearest known route.

What does version-aware extraction give your team?

AWS guidance on Amazon Textract adapters, published on 28 September 2026, separates document classification, adapter selection and extraction into distinct stages. Each separation carries a business benefit:

  • A new layout leaves working routes alone. Textract applies one adapter per page per feature type, so a routing step upstream chooses the right extractor for each form version.
  • Reviewers see the certificates that need them. The router reads markers such as form titles, revision dates and distinctive field labels, and results flow to databases, workflow engines or human review queues. An unknown marker becomes a visible review item rather than a silent blank field.
  • Changes go live without a code release. With the active adapter ID stored as configuration, promoting a new version means updating one parameter, with no application redeployment.
Explanatory illustration of Keep document automation useful when suppliers change their forms

Why does the business model matter more than the extractor?

A certificate becomes useful only when its extracted text reaches the right record. Functional AI Solutions builds a foundational ontology for this: one shared model of your suppliers, products, material grades, heats, batches and certificates, and the relationships between them.

The ontology gives every form version a stable target. Supplier layouts change; the link between a certificate, its heat number, the supplier, the material grade and your receiving batch does not. Each version maps its own labels, such as 'Heat No.' or 'Cast No.', to that agreed field. A new layout needs a new mapping, not a new data model. Your receiving, quality and procurement systems still need connectors to that model. FAIS then builds the routing, review workflow and automation on top of it.

A revised certificate, before and after

Consider a Taiwan manufacturer of machined parts that receives a material certificate with every steel delivery. One mill issues a revised certificate that moves the heat number into a table and renames the chemistry columns.

Before: the single extractor for that supplier keeps running. It returns a blank heat number or reads the wrong cell. Receiving finds the mismatch later, or the team goes back to keying in data.

After: the router finds a revision marker it does not recognise and sends the certificate to review with the original file attached. A reviewer enters the values once, and those corrections become samples for a new route. Old-form certificates keep flowing through the old route.

SphereGen's own case study of a Connecticut manufacturer shows the review side in production: staff approve or reject extracted certificate data before upload to ERP, quality and procurement systems, and certificates the automation cannot process go to an exception queue with the original file and attempted values.

How can you add a new supplier layout without disrupting working ones?

Add the new layout as a separate version with its own samples, extractor and mapping, and promote it only when it passes:

  1. Marker defined. A revision date, form code or distinctive label identifies the version.
  2. Samples separated. Textract needs at least five training and five test documents per adapter. Hold back some real certificates of the new layout for testing only.
  3. Critical fields scored. Check precision and recall on the fields that decide release: heat number, grade, chemistry and mechanical results.
  4. Mapping verified. Every field lands on the agreed supplier, product and batch record.
  5. Regression passed. Existing versions match their baseline results on their own test sets.
  6. Rollback and queries recorded. The previous version stays active, and query definitions live in your own store, because only trained model weights transfer when adapters are copied between accounts.
  7. Auto-update held. AWS suggests disabling auto-update in production until regression testing exists.
  8. Release authority named. The automation prepares the record; an authorised quality staff member releases the batch.

Adapter training takes 2–30 hours, depending on dataset size and AWS Region; the review queue carries the revised form meanwhile.

Which numbers show the workflow is healthy?

Track corrections per field, review minutes per certificate and the share sent to the unknown-layout queue, split by supplier and form version. A blended accuracy score can hide one revised form that is failing unnoticed.

Start with one certificate family

Choose the certificate family with the most suppliers or revisions. Collect old and new versions, agree the batch and product fields that matter, and name who approves release. FAIS models those records in your ontology and connects routing, extraction and review to them, so the next family reuses the same records and checklist.

When a supplier changes its form, you add a route instead of rebuilding the workflow.

Sources

  1. Automating Amazon Textract adapter lifecycle management across accounts | Artificial Intelligence
  2. Automating the Intake of Material Certificates - SphereGen
  3. Customizing your Queries Responses - Amazon Textract
  4. Evaluating and improving your adapters - Amazon Textract

Questions operators ask

Version-aware extraction is a document workflow that identifies which version of a supplier form has arrived, applies the extractor and field mapping approved for that version, and sends unrecognised layouts to a human reviewer.

How do you add a new supplier form layout without breaking the ones that already work?

Add it as a separate version with its own marker, samples, extractor and field mapping. Score the critical fields on held-back test certificates. Confirm existing versions still match their baselines, keep the previous version for rollback, then promote the new one.

What happens to revised certificates while a new extractor is being trained?

They go to a review queue with the original file attached. A reviewer enters the values once, and those corrections can become samples for the new route. Amazon Textract adapter training takes 2–30 hours, depending on dataset size and AWS Region.