The most reliable production approach to PDF extraction for engineering combines layout segmentation, orientation-aware detection, and specialized parsing into a single hybrid pipeline that outputs schemaed JSON with full provenance. Generic OCR alone cannot resolve rotated dimensions, GD&T frames, or dense multi-view layouts, which is why preprocessing and human-in-the-loop verification remain non-negotiable, not optional extras bolted on afterward.
In short
01
A standard OCR engine was built to read paragraphs of upright, left-to-right text. An engineering drawing is almost the opposite of that assumption, and the mismatch explains most of the failures teams encounter when they first try to automate extraction.
Rotated dimension strings, often at 45 or 90 degrees, confuse character recognizers trained on horizontal text. Overlapping views, where a plan, section, and detail share the same sheet, create ambiguity about which annotation belongs to which geometry. Geometric dimensioning and tolerancing (GD&T) frames pack multiple symbols, datums, and tolerance values into a single compact box that generic OCR tends to read as garbled characters rather than structured data. Title blocks vary by company, discipline, and country, so a model trained on one template often fails silently on another.
A less obvious problem is that born-digital PDFs, the ones exported directly from CAD software, are frequently assumed to be “structured” when they are not. The text layer might exist, but critical fields such as dimension values or revision tags are often converted to curves or outlines during export, stripping out any machine-readable encoding. Treating that text layer as authoritative, rather than verifying it against a rasterized fallback, is a common and costly mistake.

Ambiguous legends compound the problem. A symbol that means one thing on a mechanical drawing can mean something else on a structural or electrical sheet, and a language model without project-specific grounding will often guess wrong with confidence rather than flag uncertainty.
Common failure points include:
One domain study reported numerical VLM F1 scores up to 0.963, with some annotation types reaching perfect scores, while alphabetical (text-heavy) fields scored lower, confirming that numeric and symbolic data behave very differently from free text in engineering drawings.
02
Three broad architectural patterns dominate production engineering extraction, and the right choice depends on input format, field type, and how much error tolerance a given use case allows.
The three-stage hybrid architecture is the most battle-tested pattern: layout detection identifies views, title blocks, and notes; orientation-aware annotation localization finds and corrects rotated dimension strings and GD&T frames; and a specialized parser, whether OCR or a vision-language model (VLM), extracts the actual values. A multi-stage hybrid framework applied this exact sequence to 2D multi-view engineering drawings and reported strong performance on numerical and symbolic fields, while flagging that free-form alphabetical text remained harder to parse reliably.
OCR-free VLMs change the calculus for numeric and symbolic fields specifically. Instead of detecting characters and reconstructing words, these models reason directly over image patches and can interpret a GD&T frame or a dimension callout as a semantic unit rather than a string of glyphs. Research on automated parsing of engineering drawings found that Donut-style, fine-tuned VLMs performed very strongly on numeric and symbolic fields, but noted that these same models showed higher hallucination rates on text-heavy title blocks, particularly in zero-shot settings without fine-tuning. The practical implication: use OCR-free VLMs where fields are numeric or symbolic, and fall back to OCR or rule-based parsing where fields are long-form text.
Vector-format extraction applies only to a narrower case: born-digital PDFs that retain their original vector paths rather than converting everything to outlines or raster images. When this is true, dimension lines, text runs, and layer metadata can sometimes be pulled directly from the PDF’s internal structure without running any visual model at all, which is both faster and more accurate. The catch is that this path only works for genuinely vector-preserving exports, and any rasterized or scanned sheet defaults back to the visual pipeline.
Choosing among these architectures comes down to a few practical trade-offs:
Most production systems end up as a blend: vector parsing where possible, OCR-free VLMs for numeric and symbolic fields, and traditional OCR or rule-based parsers as a fallback for long-form notes and specifications.
03
Building a pipeline that holds up under real project volume means treating extraction as a sequence of discrete, auditable stages rather than a single model call. Each stage has its own failure modes, and isolating them makes debugging and accuracy improvement tractable.
An open-source project demonstrating this exact structure is gost-ocr, which implements a three-stage preprocess-localize-extract pipeline for title-block metadata, outputting structured JSON with per-field bounding boxes and confidence scores for each extracted value.
Pro Tip: Log the original image patch alongside every extracted field, not just the text value, so a human reviewer can verify a low-confidence result in seconds rather than reopening the full drawing.
The stakes of getting this wrong are not abstract. A misread tolerance or an undetected rotated dimension can propagate into a bill of quantities, a fabrication order, or a compliance check, and the cost of catching that error after fabrication dwarfs the cost of a careful verification step during extraction.
04
No single library covers every stage of an engineering extraction pipeline, so most teams assemble a stack from several tool categories rather than adopting one end-to-end product.
For legend sheets specifically, where symbol meaning is ambiguous without project context, an interactive approach called In-Context Multimodal Annotation Prompting proved effective. The ICICLE system grounded a general-purpose VLM with a single annotated example and a referential text prompt, reporting extraction accuracy between 96 and 100 percent on tested legend sheets, without requiring full model retraining for each new project’s symbol conventions.
Preprocessing work of this kind is not a cosmetic step. Pipelines modeled on domain-specific reconstruction work, such as the chemoCR approach to chemical structure diagrams, organize processing into preprocessing, reconstruction, postprocessing, and validation stages, with vectorization and connected-component analysis materially improving what the recognition stage can achieve. The same principle holds for engineering drawings: a detector fed a deskewed, properly scaled page performs measurably better than one fed a raw scan.
05
Accuracy in engineering extraction has to be measured per field type, not as a single aggregate number, because numeric fields, symbolic fields, and free text behave very differently under the same model.
Field-level precision, recall, and F1 scores tell you where a pipeline is strong and where it needs a fallback. A hallucination rate, the proportion of extracted values that are confidently wrong rather than simply missing, matters just as much as recall for engineering data, since a wrong dimension is often more dangerous than a flagged gap. Research on automated drawing parsing found that OCR-free VLMs can show elevated hallucination on text-heavy fields in zero-shot settings, which is a direct argument for fine-tuning on labeled, domain-specific samples before trusting a model’s output.
Building a labeled dataset should prioritize the fields that drive downstream decisions: title blocks (project number, revision, scale), GD&T frames, numeric dimensions and tolerances, and legend or symbol sheets. A few hundred well-annotated sheets covering a company’s actual template variety typically outperform a much larger generic dataset.
| Validation pattern | What it catches | Typical trigger |
|---|---|---|
| Deterministic rule checks | Impossible values (negative dimensions, out-of-range tolerances) | Always runs on every field |
| Parallel model cross-check | Disagreement between OCR and VLM outputs | Fields below a confidence threshold |
| Human reviewer routing | Low-confidence or disputed fields | Confidence score below threshold or rule failure |
| Audit logging | Full provenance for compliance review | Every extracted field, regardless of confidence |
A domain-specific framework reported numerical field F1 scores as high as 0.963 using exactly this kind of staged validation, underscoring that strong numeric accuracy is achievable when detection, parsing, and verification are treated as separate, auditable steps rather than a single opaque model call.
Logging provenance, meaning the source page, bounding box coordinates, and confidence score for every field, turns an extraction pipeline from a black box into something a compliance team can actually audit after the fact.
06
Extraction only has value once the output can be consumed by the systems engineers already use, which means the schema design matters as much as the extraction accuracy itself.
A unified JSON schema, with each field carrying its page number, bounding box, extracted value, and confidence score, gives downstream systems both the data and the means to verify it. For teams working in spreadsheet-driven workflows, exporting the same data to Excel or CSV keeps it accessible without requiring a new tool. For BIM-centric workflows, mapping fields into an IFC-lite structure allows extracted quantities and attributes to align with existing model data rather than living in a separate silo, a pattern directly relevant to how quantities get derived from an IFC model.
Integration generally follows one of two patterns: direct API write-back into a PLM, CAD, or ERP system, or a staging export that a human reviews before it touches production data. The staging pattern is safer for early rollouts, since it adds a checkpoint before extracted data can corrupt a live system of record.
A few pitfalls recur across integration projects:
Structural estimating and bidding workflows illustrate the stakes well: steel estimating platforms depend on clean, deduplicated quantity data flowing in from drawings, and a schema mismatch upstream turns into a bidding error downstream.
07
General-purpose extraction models tend to degrade sharply as drawing complexity increases, because they lack the domain priors that let a specialized system distinguish a GD&T frame from a stray annotation or recognize a title block layout it has never seen before. Domain-tuned pipelines, built around the architecture described above, reduce hallucination specifically because they are validated against engineering-specific rules rather than generic language patterns, and they preserve traceability by design rather than as an afterthought.
Procurement teams evaluating a vendor or an internal build should request concrete evidence before committing:
During a demo, ask for an end-to-end sample export on a representative drawing set, with bounding boxes and confidence scores visible, so the traceability claims can be checked directly rather than taken on faith. Our own approach to domain-specific AI for civil engineering follows these same principles: specialization narrows the problem enough that accuracy and auditability both improve together.
08
High accuracy on engineering PDF extraction is achievable, but it rarely arrives as a single deployment event. Teams that succeed tend to stage the rollout: automate title block extraction first, since it is the most standardized field type, then move to numeric dimensions, and only later tackle GD&T and symbolic fields, which carry more error cost and need the most validation.
Expect to retrain or recalibrate detection and parsing models periodically as new drawing templates, disciplines, or standards enter the pipeline. Dataset drift, where a model’s accuracy quietly degrades as the input distribution shifts, is a maintenance cost that should be planned for from the outset, not discovered after accuracy complaints start arriving. Budgeting for ongoing labeling work, even a modest amount per quarter, tends to matter more for long-term reliability than any single architectural choice made at the start.
09
For construction and infrastructure teams that need this kind of pipeline running in production rather than prototyped in a notebook, we built Yesper as an AI civil engineer that works inside your actual projects. Rather than handing back raw extracted fields, our platform searches project documents, checks regulatory compliance, and automates multi-step workflows that turn extracted data into finished deliverables, reports, tenders, spreadsheets, and quantity takeoffs from BIM models and drawings.
Our customers report significant time savings on projects, and errors caught that manual review missed, because our platform flags inconsistencies a human pass might miss. Because we are built for construction workflows, our agents understand domain-specific regulations and can integrate with project management systems, performing work end-to-end rather than just answering questions about a drawing.
What this looks like in practice:
Teams ready to see how this fits an existing workflow can explore our enterprise platform for construction and infrastructure or review our services for integration and operational deployment.
FAQ
No single tool handles every field type well, which is why production systems combine a layout segmentation model, an orientation-aware detector, and a specialized parser rather than relying on one general tool. The strongest results come from matching the parser to the field type: OCR-free VLMs for numeric and symbolic data, OCR or rule-based parsing for free text, as shown in hybrid framework research.
PDFs remain the dominant exchange format for engineering drawings because they preserve layout fidelity across software and organizations, and there is no sign of that changing industry-wide. The more relevant shift is toward structured data layers and IFC-based model exchange running alongside PDFs, not replacing them outright.
A general-purpose model like ChatGPT can read simple, well-structured PDFs but tends to struggle with rotated dimensions, GD&T frames, and dense multi-view layouts without the orientation-aware detection and domain fine-tuning that specialized pipelines use. For reliable field-level accuracy on technical drawings, a hybrid pipeline with verification steps outperforms a general chatbot interface, as explored in our breakdown of engineers’ questions about AI.
Fine-tuned Donut-style vision-language models perform strongly on numeric and symbolic fields, while free-form text like specifications and notes is often handled more reliably by OCR or a general VLM paired with rule-based checks. Research comparing these approaches found that OCR-free VLMs can hallucinate more on text-heavy fields in zero-shot use, which supports pairing specialized models by field type rather than picking one model for an entire drawing.
Field-level accuracy depends heavily on field type, with numeric and symbolic fields reaching very high F1 scores in staged, validated pipelines, as reported in hybrid framework benchmarks. Free-form text fields remain harder, which is why human-in-the-loop review for low-confidence outputs stays part of any production deployment.
Sources
This post was written with AI assistance and published by Yesper. General information, not professional advice: requirements vary by project and jurisdiction, and the professional responsible for the project decides what applies. Spotted an error? Write to benjamin@yesper.ai.
Get news and articles in your inbox.
Yesper is the AI civil engineer for construction and infrastructure. AFRY, COWI, NRC Group and other Nordic firms use it to halve the time on a study, rerun calculations in minutes, and catch errors that would otherwise slip through. Get in touch if you'd like to see what it can do for you.
Book demo