Back to Blog Technology

What AI Can and Cannot Extract from EPC Documents Today

Jiyeon Park
Technical document parsing pipeline with highlighted data fields

Building SnapScale required us to be honest, in public and in our own design reviews, about where AI document extraction works reliably today and where it does not. EPC documents are a harder target than most practitioners realize when they first approach them, and a good number of the limitations are structural rather than a function of the current generation of models.

This article is a technical assessment, not a pitch. We are going to describe what extraction works well, what is unreliable, and where the hard limits are. If you are evaluating any AI-based document extraction tool for EPC use, you should be asking every vendor the same questions we ask ourselves.

The EPC Document Format Problem

EPC project documents are not a homogeneous category. They include PDFs of CAD drawings, spreadsheet exports converted to PDF, Word-derived datasheets, scanned legacy documents with varying OCR quality, vendor-formatted submittals in proprietary layouts, and multi-sheet drawing sets where the meaningful content is distributed across a sheet border legend and embedded symbol libraries.

General-purpose document extraction models are trained primarily on text-heavy documents: contracts, reports, technical articles. They handle those formats well. EPC documents are a different structural problem. A P&ID at full scale is a graphical document. The "text" in a P&ID is instrument tag labels, line numbers, valve symbols with alphanumeric identifiers, and process condition annotations, all spatially positioned relative to each other rather than flowing in a left-to-right, top-to-bottom reading order.

Reading a P&ID is not a text extraction problem. It is a spatial understanding problem: a tag label is meaningful only in relation to the instrument symbol it is connected to, and that relationship is defined by graphical proximity and connection lines, not by text proximity. A model that extracts the text content of a P&ID without understanding the spatial structure will collect a list of tag labels and annotations, but will not correctly associate each annotation with the element it describes.

What Works: Structured Tabular Documents

The extraction category that works most reliably for EPC documents is structured tabular data in PDFs that were generated from spreadsheets or structured templates. Instrument datasheets from most engineering firms follow a predictable format: a title block, a set of fixed fields (service, tag number, process conditions, range), and a vendor data section. If the PDF was generated from a template rather than scanned from paper, the text layer is clean and the cell structure is recoverable.

For these documents, a well-tuned extraction pipeline can reliably pull tag numbers, process conditions (normal and design temperature and pressure), engineering range, set point values, material specifications, and document metadata (drawing number, revision, revision date). Reliability here means above 95% field-level accuracy on a diverse sample of datasheets from different engineering firms, with errors concentrated in unusual formatting, multi-value cells, and non-standard field names.

Line lists and equipment lists in tabular PDF format behave similarly. The header row defines the column schema, the data rows fill it consistently, and extraction is primarily a table detection and column alignment problem. This category of document is where the first generation of general-purpose extraction tools performs adequately, because it is structurally close to the document types those tools were trained on.

What Is Harder: Semi-Structured Documents

Process data sheets, engineering specifications, and cause-and-effect matrices introduce layout variability that reduces extraction reliability significantly. These documents are often produced from Word templates that different engineers have modified over the years. The field names vary between firms and projects. Sections appear in different orders. Tables are sometimes replaced with free-form paragraphs. Footnotes carry critical exceptions to values stated in the main body.

For semi-structured documents, extraction accuracy drops to the 75 to 85% range on individual field values, depending heavily on the document source. Errors are not random; they cluster around fields where the document format is ambiguous, where values appear in context (e.g., a pressure stated as part of a sentence rather than in a dedicated cell), and where units are inconsistently specified (bar, bara, barg, kPa, psi all appearing in the same document set).

Unit normalization is a non-trivial problem that often goes unacknowledged. An EPC project document set will contain pressure values in at least three different unit conventions from different discipline sources. An extraction layer that does not normalize units before comparing values across documents will generate false positives constantly: a datasheet stating 14.5 psi will appear to conflict with a P&ID annotation of 1 bar, even though both express the same pressure. Getting unit normalization right requires a domain-specific conversion layer, not just text extraction.

What Is Genuinely Hard: P&ID Extraction

Automating P&ID reading at the level of accuracy needed for cross-document consistency checking is the hardest problem in EPC document extraction. Graphical P&IDs in PDF format contain structured information, but that structure is spatial, not textual. The industry has been working on P&ID digitization for many years, and the progress has been real but limited.

Current approaches to P&ID extraction tend to fall into two categories. The first is symbol-based recognition: train a model to recognize instrument symbols, valve types, and line annotations, then infer the process conditions associated with each element from the annotations that appear in its vicinity. This works tolerably for standardized instrumentation symbols but degrades quickly when projects deviate from standard symbol sets (which most do at some level) or when P&IDs are dense enough that annotation proximity becomes ambiguous.

The second approach is to treat the P&ID as a source of structured tag-attribute pairs and extract them by reading the revision history tables and text blocks rather than by parsing the graphical content. This extracts less information but with higher reliability. You can get revision blocks, title block data, and tagged annotation text without fully parsing the drawing. This is what SnapScale does today: we extract what is reliable (tag labels, drawing references, process condition text when annotated near tags) and flag what requires human verification rather than propagating uncertain extractions into the cross-reference.

We are not saying full automated P&ID parsing is impossible. We are saying that the current state of the technology does not deliver the accuracy needed for consistency-checking applications without a human review step. Tools that claim full automated P&ID parsing without qualification should be tested rigorously against a representative sample of your actual project drawings before procurement decisions are made.

The Verification Architecture That Makes Extraction Useful

The honest design response to extraction reliability limits is not to claim higher accuracy. It is to build a verification architecture that separates high-confidence extractions from low-confidence ones, routes low-confidence extractions for human review, and accumulates validated data as a training signal for project-specific tuning.

In practice this means: every extracted value carries a confidence score; values below a threshold are presented to the user as requiring confirmation before they enter the cross-reference; confirmed values are used to tune extraction behavior on similar future documents in that project's format; and the system's error rate on a given project decreases over time as the user validates the initial extraction batch.

The extraction layer is useful not because it gets everything right automatically, but because it handles the 80% of values that are high-confidence efficiently and focuses human attention on the 20% that require judgment. A process engineer doing a manual cross-check spends equal time on every value. An extraction-assisted cross-check concentrates the engineer's time where ambiguity actually exists. That is the real value proposition, and it is more honest than claiming extraction accuracy that does not hold up under real project conditions.