News

September 7, 2026

Observability for Document Pipelines

Observability for Document Pipelines

Postponing the AI Act until December 2027 does not change the first actual deadline, which is July for eFTI and AMLR. This is just enough time to set up a document pipeline: trust by field, evaluations, drift, and lineage.

Regulation (EU) 2026/1744 postponed the obligations under the AI Act for high-risk systems until December 2, 2027, while for logistics and banking, the clock is ticking faster: eFTI and AMLR take effect on July 9 and 10 of that year. With the timeline clear until then, the key question is no longer a legal one but an engineering one: what needs to be implemented in a document pipeline to ensure that those deadlines do not require the entire process to be redone.

The uncomfortable reality is that day-one accuracy does not predict accuracy on day ninety. The document mix shifts: new supplier templates come in, along with lower-quality scans, rotated pages, and attachments taken with a cell phone. Templates are updated, and with them, outputs that no one had looked at again change. Deep Analysis, after reviewing more than 100 manufacturers, describes how accuracy in repetitive document processing tasks deteriorates with use. A pilot project that achieved 97% accuracy reflects the results from the day it was measured.

Implementing this involves four specific steps. Confidence calibrated by domain, used as a flow control: a value of 0.62 should trigger routing to human review, not an exception in the log. Evaluations of the system’s own documents, with versioned ground truth and regression tests that run with every change to the prompt, schema, or model—just like any other test in continuous integration. Monitoring of drift by document type rather than in aggregate, because the aggregate hides precisely the type that failed. And evidence lineage: which page and which text each value came from.

There is also a cost argument. In anyformat’s July 2026 public comparison, parsing 1,000 pages with a state-of-the-art general-purpose model costs about four times as much as using a specialized parser with equivalent performance. And an agent that re-reads the same PDF with every query pays state-of-the-art context costs for a task that is performed only once: parsing the structure, caching it, and retrieving it.

What’s striking is that these four components are almost exactly what Annex III will require in December 2027: event logging, version-specific technical documentation, effective human oversight, and data traceability back to its source. Anyone who implements them for engineering reasons will meet the deadline without having to launch a compliance project. At anyformat, we call this discipline “Document Operations,” by analogy with what DevOps did for deployment: what isn’t measured in production doesn’t hold up in production.

When Tech Show Madrid opens its doors on November 4, there will be eight months left until July 2027. A pipeline without evaluations can't be fixed in August.

SEE MORE NEWS
Loading

Partners

Institutional Support


 

Institutional Support


 

Institutional Support


 

Institutional Support


 

Event Partner


 

Event Partner


 

Strategic Partner


 

Event Partner


 

Event Partner


 

Event Partner


 

Event Partner


 

Event Partner


 

Event Partner


 

Event Partner


 

Event Partner


 

Media Partner


 

Media Partner


 

Media Partner


 

Media Partner


 

Media Partner


 

Media Partner


 

Media Partner


 

Media Partner


 

Media Partner


 

Media Partner


 

Media Partner


 

Media Partner


 

Media Partner


 

MediaPartner


 

Partner


 

Partner


 

UX Partner


 

CX Partner


 

3D Partner