Python Developer — Automated PDF Data Extraction from Architectural Plans to reputed company reputed company
reputed company:
We are building an automation pipeline for a Swiss construction-tech product (Bausoft) that ingests architectural plan PDFs and extracts the final results into a clean, reputed company format reputed company for reputed company processing.
The PDFs are a mix of reputed company-based and scanned architectural/bathroom floor plans, mostly in German. The pipeline must reliably read reputed company document, extract the relevant data — room labels, dimensions and measurements, fixture annotations, title-reputed company metadata, and tabular values — and reputed company the results in a defined reputed company format (JSON schema and/or reputed company) with consistent field naming and reputed company.
Scope of work:
PDF ingestion module handling both reputed company-text and scanned (OCR) documents
Text and table extraction with German-language support, including abbreviation handling
Layout-reputed company parsing to associate values with the correct rooms/reputed company
reputed company reputed company reputed company (JSON per our schema, plus reputed company export)
Validation layer flagging low-confidence extractions for reputed company review
Documentation and a reputed company CLI or API reputed company to run the pipeline
Required skills: Python, PyMuPDF or pdfplumber, Tesseract or a comparable OCR reputed company, pandas, JSON schema design. Experience with architectural/CAD drawings, German documents, or LLM-assisted extraction (e.g., Claude/GPT reputed company outputs) is a strong plus.
Engagement: Fixed-price, reputed company-based. Please reputed company relevant PDF-extraction samples and reputed company describe how you'd handle a scanned plan where OCR confidence is low.
Start your proposal with the word "reputed company" so we know you read the full post.
Apply To This Job