Files
Conflict_Checker/docs/superpowers/specs/2026-08-12-text-layer-grounding-design.md
T

8.0 KiB

Text-Layer Grounding — Design Spec

Date: 2026-08-12 · Branch: agent-mode · Status: approved by user (2026-08-12)

Problem

The pipeline is vision-only for extraction, but most CAD-produced drawing sets carry a real vector text layer. Two worst documented failure modes are text problems being solved with pixels:

  1. Wave-1 text misreads propagate immutably — e.g. job 959e16407573: vision read "(2) 2x6 STUD PACK" where the sheet says "(5)"; text-only downstream specialists treated the misread as ground truth → confident false-positive findings.
  2. Silent extraction loss — failed/under-extracted pages are invisible (job 475a6f184dd1: 42% extraction loss), producing false missing_expected_sheets warnings and missed conflicts.

Priority (user, 2026-08-12): reduce false positives and missed items; more accurate conflicts.

Approach

Extract the PDF text layer deterministically (PyMuPDF) once per job, and make it a first-class citizen at three points: extractor grounding, the grounding guard, and the wave-5b verifier (as text oracle + high-DPI evidence crops).

Inspired by hamzaabduljabbar/construction-drawing-analyzer (patterns only — its license is source-available/no-resale; all code here is original).

Components

1. New module backend/text_layer.py (deterministic, no LLM)

  • extract_text_layers(pdf_path) -> Dict[int, dict] — per 1-based page: {"text": str, "words": [{"text", "bbox": (x0,y0,x1,y1)}, ...], "has_text_layer": bool}. Pages with < TEXT_LAYER_MIN_CHARS of text are has_text_layer=False (scanned/raster sheets stay vision-only; logged).
  • find_evidence_bbox(words, needle) -> bbox | None — best-effort fuzzy substring match of an evidence source_text against word sequence; returns union rect of matched words.
  • render_crop(pdf_path, page_number, bbox, dpi, margin_pts) -> bytes — PyMuPDF page.get_pixmap(clip=rect, dpi=dpi) → JPEG bytes.

Both runners call extract_text_layers right after convert_pdf_to_images and attach page["text_layer"] = <text or None> to each page dict. Word positions stay in a separate page_words: Dict[int, list] runner-local map (not attached to page dicts — they get serialized).

2. Extractor grounding (both pipelines)

  • Static paragraph added to _EXTRACTOR_SYSTEM_TEMPLATE in backend/prompts.py (no new placeholder): when a TEXT LAYER block is present in the user message it is authoritative for alphanumeric content (counts, dimensions, member tags, notes); the image is for geometry, symbols, linework, and anything absent from the text layer.
  • Text-layer content is appended programmatically at each extractor call site (classic extractor.py::_extract_one, agent extractors.py::SheetExtractorAgent.run) — NOT a new {placeholder} in the shared template (two-render-path trap: render() silently leaves missing keys as literals). Block capped at TEXT_LAYER_MAX_CHARS. Format: \n\nTEXT LAYER (authoritative for alphanumeric content — trust it over the image for numbers, tags, and note text):\n<text>

3. Grounding guard rescue tier (pipeline/extractor.py::_normalize_sheet)

Current guard drops an object when its primary value's digit-runs aren't in its own source_text. New tier, only when a text layer exists for the page:

  • digits ⊆ source_text → keep (unchanged)
  • digits ⊆ page text layer but ⊄ source_text → keep, stamp grounding: "text_layer" on the assertion (recall rescue — vision quoted imperfectly but the value is real page text)
  • otherwise → drop (unchanged)

_is_grounded gains an optional page_text param; existing callers/tests unaffected. Dropped/ rescued counts logged per page.

4. Verifier: text oracle + high-DPI crops (wave 5b)

Wherever verify scopes are built (agent runner confirmed; classic runner to be checked — integrate at both if present):

  • Scope payload gains text_layer_excerpt: concatenated text of the finding's cited sheets, capped at VERIFY_TEXT_MAX_CHARS. VERIFY_USER_INSTRUCTION gains a {text_layer} placeholder with instructions to treat it as deterministic page text (verdicts may cite it as actual_text). Both render sites (agent verifier + any classic-path render) must substitute it — grep the template name across backend/agents/ and backend/pipeline/.
  • When VERIFY_HI_DPI_CROPS and the page has words: for each evidence item, find_evidence_bbox on the cited page's words; on hit, render_crop at VERIFY_CROP_DPI with margin → crop images replace full-page images (up to AGENT_CONFLICT_MAX_IMAGES). On any miss/failure → fall back to the current full-page image. Zero-resolved-images ⇒ scope skipped (I2 guard preserved).

5. Coverage signal (recall)

After extraction in both runners: for each page with has_text_layer=True whose extraction failed or returned 0 objects, log [TextLayer] Page N: text layer present (M chars) but no objects extracted — possible extraction gap and add the page to the existing gap-finding path (agent: orchestrator.stats.failed_scopes-style finding; classic: log only).

Config knobs (backend/config.py, env-overridable, documented in .env.example)

Key Default Effect
TEXT_LAYER_ENABLED true Master switch
TEXT_LAYER_MIN_CHARS 20 Below this per page → has_text_layer=False
TEXT_LAYER_MAX_CHARS 12000 Cap per sheet injected into extractor prompt
VERIFY_TEXT_MAX_CHARS 8000 Cap of text-layer excerpt in verify scope
VERIFY_HI_DPI_CROPS true Evidence-located crops in verifier
VERIFY_CROP_DPI 300 Crop render DPI
VERIFY_CROP_MARGIN_PTS 36 Padding around evidence bbox (PDF points)

Known traps (from project history — designed around)

  • Two render paths: no new {placeholder} in extractor templates; the one new placeholder ({text_layer} in VERIFY_USER_INSTRUCTION) substituted at every render site; a render test asserts no {...} literals remain.
  • ProjectMemory closed registry: no new memory keys. Text artifacts dump via plain file writes under outputs/<job>/text/ (agent: under agent/).
  • slim_clusters: no new cluster fields — unchanged.
  • I2 zero-image path: crops replace full-page images only on confident bbox match; never reduce image count to zero.
  • Base64 hygiene: page dicts already carry base64; text_layer strings must not leak into clusters.json dumps — reuse _without_base64 pattern if assertions ever carry page refs (they don't today).

Dependencies

PyMuPDF>=1.23 added to requirements.txt (Docker image rebuild picks it up; pdf2image/poppler unchanged).

Testing

  • tests/test_text_layer.py — build tiny PDFs with PyMuPDF in-test: extraction, has_text_layer thresholds, find_evidence_bbox hit/miss, render_crop dimensions.
  • Extractor guard: rescue-tier unit tests (keep-with-flag, still-drop, unchanged behavior without text layer).
  • Prompt render test: extractor + verify instructions fully substituted at every site (both pipelines).
  • Runner-level (pattern from tests/agents/test_wave5b_suppression.py): stubbed waves, assert text layer reaches extract scopes and verify scopes (excerpt present, crop fallback on no-match), full run_agent_pipeline.
  • Full pytest tests/ green before push.

Validation (post-deploy)

Re-run the Cypress set (source PDF persists at /app/backend/outputs/959e16407573/source.pdf on sits-docker) per the documented re-run workflow. Success criteria:

  1. The "(2) vs (5)"-class findings are not generated, or are verifier-refuted with text-layer evidence cited.
  2. Coverage-gap log lines appear for any page with text but no objects.
  3. No new finish_reason=length in waves 1/4; cost delta reported vs baseline job.

Out of scope (future PRs)

  • Legend/symbol-library wave injected into extractor + critic prompts.
  • Deterministic schedule-row recall pass (text-layer tables → assertions).
  • pdf-markup export for the review UI.
  • Takeoff/polygon geometry (belongs to AI_Takeoffs, not this product).