feat: text-layer grounding (extractor authority, guard rescue tier, verifier oracle + hi-DPI crops)
Docker Release / build-and-push (push) Successful in 1m25s
Docker Release / release (push) Skipped

- backend/text_layer.py: PyMuPDF text-layer extraction, fuzzy evidence
  bbox matching, 300-DPI crop rendering, coverage-gap signal
- extractor (classic + agent): TEXT LAYER block appended at call sites;
  grounding guard gains text-layer rescue tier (grounding=text_layer stamp)
- verifier: {text_layer} oracle excerpt + evidence-located hi-DPI crops
  replacing full-page images (fallback preserved, I2 guard intact)
- coverage gaps: text-bearing pages with zero extraction -> failed-scope
  gap findings (agent) / log-only (classic)
- config knobs: TEXT_LAYER_ENABLED/MIN_CHARS/MAX_CHARS, VERIFY_TEXT_MAX_CHARS,
  VERIFY_HI_DPI_CROPS, VERIFY_CROP_DPI, VERIFY_CROP_MARGIN_PTS
- tests: 22 new (text_layer unit, grounding/render, runner-level flow)
Spec: docs/superpowers/specs/2026-08-12-text-layer-grounding-design.md
This commit is contained in:
2026-08-12 14:27:00 -05:00
parent 349b357e5c
commit 570300324f
14 changed files with 1665 additions and 16 deletions
@@ -0,0 +1,936 @@
# Evidence Verification + Cross-Sheet Correlation Implementation Plan
> **For Hermes:** Use subagent-driven-development skill to implement this plan task-by-task.
**Goal:** Stop vision-extraction misreads (e.g. "(2) 2x6 STUD PACK" vs the actual "(5) 2x6") from becoming confident downstream findings, and correlate the same physical element across sheets (S101/S205/S401) so no stage reasons from one sheet's text in isolation.
**Architecture:** Three independently shippable phases on the agent pipeline (`backend/agents/runner.py`):
1. **Disputed-value detection** — deterministic post-link pass that flags contradictory extracted values inside a cluster and surfaces them to the critic/specialist prompts.
2. **Cross-sheet xref linking** — linker gains detail-reference/tag buckets that join assertions across levels (today `(level, family)` bucketing splits S101/S205/S401 apart).
3. **Evidence verification wave (5b)** — a bounded vision fact-check agent re-reads the cited sheet images for high-severity / disputed findings before the Brain merge, annotates or suppresses findings built on phantom text.
**Tech Stack:** Python 3.14, pytest (`tests/`), existing `call_json` LLM wrapper (supports `images_b64`, `reasoning_effort`, `reasoning_max_tokens`).
---
## Current context / root cause (from job 959e16407573)
Finding `validated_issues[3]` ("FRONT PERSPECTIVE detail on Sheet S401", HSS16x4 on
"(2) 2x6 STUD PACK", severity critical) is a **false positive built on a wave-1 vision
misread**. The sheet actually shows a (5) 2x6 stud pack (matching S205/S101). Chain of failure:
1. Wave 1 (`SheetExtractorAgent`) froze the misread into text. From then on it is "ground truth".
2. Wave 3 linker (`backend/agents/linker.py:33` `build_link_scopes`) buckets by
`(level, family)`. S101 (foundation), S205 (details), S401 (sections) get different
`level` values, so assertions about the same front-wall header never share a link scope
or cluster. No cross-sheet corroboration happened.
3. Wave 5 `ConstructabilityAgent` (`backend/agents/construct_agent.py:53`) calls
`call_json` **with no images** — in this job 120/120 constructability calls were `+0img`.
It reasoned arithmetically from the misread text ("2 x 1.5in = 3in < 4in -> unbuildable").
It even held the "(5) 2x6 STUD PACK" assertion in the same scope but labeled it
"Ambiguous column size specification" instead of arbitrating.
4. Nothing between wave 5 and the report ever looks at a sheet image again. Only the wave-4
conflict critic receives images, and only for its own cluster's pages.
Also confirmed in this log (separate known bug, fixed in Task 7 while we're here): wave-4
conflict critic truncates on Gemini thinking tokens because `conflict_critic.py:59` passes
`max_tokens=config.REASON_MAX_TOKENS` (4096) with no reasoning budget — 13/121 calls hit
`finish_reason=length`.
## Assumptions
- Assertions carry `id`, `attribute`, `value`, `source_text`, `location_key`
(`room`/`grid`/`detail_reference`/`tag`/`level`) — see `linker._payload` and
`_serialize.slim_assertion`.
- Extractor assertions already carry a `confidence` field (per test fixtures).
- `validate_issue` in `backend/pipeline/_stage.py` guarantees each finding an `issue_id`.
- Test convention: `unittest.mock.patch("backend.agents.<module>.call_json", ...)`
see `tests/agents/test_sheet_extractor_fallback.py`. Run tests with
`.venv/bin/python -m pytest tests/ -x -q`.
- `AgentResult.error` defaults to `""` (not None) in assertions.
---
## Phase 1 — Disputed-value detection + prompt hardening
### Task 1: `find_disputes` pure function (TDD)
**Objective:** Detect "same attribute, different values" inside one cluster's assertions.
**Files:**
- Create: `backend/agents/disputes.py`
- Test: `tests/agents/test_disputes.py`
**Step 1: Write failing test**
```python
# tests/agents/test_disputes.py
from backend.agents.disputes import annotate_clusters, find_disputes
def _a(id_, attribute, value):
return {"id": id_, "attribute": attribute, "value": value,
"source_text": value}
def test_find_disputes_flags_same_attribute_different_values():
assertions = [
_a("a1", "stud_pack_size", "(2) 2x6 STUD PACK"),
_a("a2", "stud_pack_size", "(5) 2x6 STUD PACK"),
_a("a3", "beam_size", "HSS16X4X5/8"),
]
disputes = find_disputes(assertions)
assert len(disputes) == 1
assert disputes[0]["attribute"] == "stud_pack_size"
assert disputes[0]["values"] == ["(2) 2x6 STUD PACK", "(5) 2x6 STUD PACK"]
assert disputes[0]["assertion_ids"] == ["a1", "a2"]
def test_find_disputes_ignores_agreeing_values_and_blanks():
assertions = [
_a("a1", "beam_size", "HSS16X4X5/8"),
_a("a2", "beam_size", " hss16x4x5/8 "), # same after normalize
_a("a3", "", "orphan"), # no attribute -> skipped
_a("a4", "beam_size", ""), # no value -> skipped
]
assert find_disputes(assertions) == []
def test_annotate_clusters_writes_disputed_attributes():
clusters = [
{"key": "c1", "assertions": [
_a("a1", "stud_pack_size", "(2) 2x6"),
_a("a2", "stud_pack_size", "(5) 2x6"),
]},
{"key": "c2", "assertions": [_a("a3", "x", "1"), _a("a4", "x", "1")]},
]
assert annotate_clusters(clusters) == 1
assert clusters[0]["disputed_attributes"][0]["attribute"] == "stud_pack_size"
assert "disputed_attributes" not in clusters[1]
```
**Step 2: Run test to verify failure**
Run: `.venv/bin/python -m pytest tests/agents/test_disputes.py -v`
Expected: FAIL — `ModuleNotFoundError: backend.agents.disputes`
**Step 3: Implement**
```python
# backend/agents/disputes.py
"""Deterministic detection of contradictory extracted values within a cluster.
Extraction is a vision pass: quantities and sizes can be misread ("(2) 2x6" vs
"(5) 2x6"). Cluster members are supposed to describe the same real-world
element, so two members asserting different values for the same attribute are
a probable misread. Flag these so downstream text-only stages treat the value
as unverified instead of reasoning from one reading.
"""
import re
from typing import Dict, List
def _norm(value) -> str:
return re.sub(r"\s+", " ", str(value or "").strip().lower())
def find_disputes(assertions: List[Dict]) -> List[Dict]:
"""Same attribute with >= 2 distinct normalized values = disputed."""
groups: Dict[str, Dict[str, set]] = {}
for assertion in assertions:
attribute = _norm(assertion.get("attribute"))
value = _norm(assertion.get("value"))
if not attribute or not value:
continue
groups.setdefault(attribute, {}).setdefault(value, set()).add(
assertion.get("id")
)
disputes = []
for attribute, values in sorted(groups.items()):
if len(values) < 2:
continue
disputes.append({
"attribute": attribute,
"values": sorted(values),
"assertion_ids": sorted(
aid for ids in values.values() for aid in ids if aid
),
})
return disputes
def annotate_clusters(clusters: List[Dict]) -> int:
"""Attach disputed_attributes to each cluster that has any. Returns count."""
annotated = 0
for cluster in clusters:
disputes = find_disputes(cluster.get("assertions") or [])
if disputes:
cluster["disputed_attributes"] = disputes
annotated += 1
return annotated
```
**Step 4: Run test to verify pass**
Run: `.venv/bin/python -m pytest tests/agents/test_disputes.py -v`
Expected: 3 passed
**Step 5: Commit**
```bash
git add backend/agents/disputes.py tests/agents/test_disputes.py
git commit -m "feat: deterministic disputed-value detection for cluster assertions"
```
---
### Task 2: Wire `annotate_clusters` into the runner + serialization
**Objective:** Disputes must be visible to the wave-4 critic and wave-5 constructability prompts.
**Files:**
- Modify: `backend/agents/runner.py` (after `memory.replace("clusters", clusters)`, ~line 116)
- Modify: `backend/pipeline/_serialize.py` (`slim_clusters`, line 46)
- Test: `tests/agents/test_disputes.py` (append)
**Step 1: Write failing test**
```python
def test_slim_clusters_preserves_disputed_attributes():
from backend.pipeline._serialize import slim_clusters
cluster = {"key": "c1", "assertions": [],
"disputed_attributes": [{"attribute": "a", "values": ["1", "2"],
"assertion_ids": ["x", "y"]}]}
slim = slim_clusters([cluster])[0]
assert slim["disputed_attributes"][0]["values"] == ["1", "2"]
```
**Step 2: Run test to verify failure**
Run: `.venv/bin/python -m pytest tests/agents/test_disputes.py::test_slim_clusters_preserves_disputed_attributes -v`
Expected: FAIL — `KeyError: 'disputed_attributes'`
**Step 3: Implement**
In `backend/pipeline/_serialize.py` `slim_clusters`, add the key:
```python
def slim_clusters(clusters: List[Dict]) -> List[Dict]:
return [
{
"key": c.get("key"),
"location": c.get("location"),
"disciplines": c.get("disciplines"),
"kind": c.get("kind"),
**({"disputed_attributes": c["disputed_attributes"]}
if c.get("disputed_attributes") else {}),
"assertions": [slim_assertion(a) for a in c.get("assertions", [])],
}
for c in clusters
]
```
In `backend/agents/runner.py`, right after `clusters = [...]` / `object_graph = build_object_graph(clusters)` (before `memory.replace("clusters", clusters)`):
```python
from backend.agents.disputes import annotate_clusters
...
object_graph = build_object_graph(clusters)
disputed_count = annotate_clusters(clusters)
if disputed_count:
orchestrator.log(
f"[Link] {disputed_count} clusters carry disputed extracted values"
)
```
(Check `Orchestrator` for the actual log method name — `orchestrator.stage(...)` exists;
if no `.log`, use the module's existing logging/print convention. Adjust to match.)
**Step 4: Run tests**
Run: `.venv/bin/python -m pytest tests/agents/ -v`
Expected: all pass (including existing `test_runner_review_gate.py`)
**Step 5: Commit**
```bash
git add backend/agents/runner.py backend/pipeline/_serialize.py tests/agents/test_disputes.py
git commit -m "feat: surface disputed extracted values to critic and specialist prompts"
```
---
### Task 3: Prompt hardening — extracted text is fallible
**Objective:** Tell text-only specialists how to handle disputed/unverified values so they stop asserting buildability conclusions from a single (possibly misread) number.
**Files:**
- Modify: `backend/prompts.py` `CONSTRUCTABILITY_SYSTEM_PROMPT` (line 533) and `CONSTRUCTABILITY_USER_INSTRUCTION` (line 551)
**Step 1: Edit prompts**
Append to `CONSTRUCTABILITY_SYSTEM_PROMPT` Rules list (after line 547, before "Use plain ASCII"):
```
- Assertions are machine-extracted from sheet images and may contain misread values,
especially quantities and member sizes (e.g. "(2) 2x6" vs "(5) 2x6").
- When the cluster lists disputed_attributes, or two evidence items disagree on a
numeric value, do NOT assert a buildability conclusion from one reading. Report the
ambiguity itself (category "detail_gap", confidence "low") and state that the value
needs verification against the sheet.
```
Append to `CONSTRUCTABILITY_USER_INSTRUCTION` after the `Cross-discipline conflicts already found: {conflicts}` line:
```
Disputed extracted values in this cluster (possible vision misreads - treat as unverified): {disputes}
```
**Step 2: Wire the `{disputes}` placeholder in `construct_agent.py`**
In `backend/agents/construct_agent.py` `run()`, extend the `substitutions` dict:
```python
substitutions = {
"assertions": dumps(cluster["assertions"]),
"clusters": dumps(slim_clusters([cluster])),
"conflicts": dumps(scope.payload.get("conflicts") or []),
"disputes": dumps(cluster.get("disputed_attributes") or []),
}
```
**Step 3: Run full test suite (prompt edits can break runner tests that snapshot prompts)**
Run: `.venv/bin/python -m pytest tests/ -q`
Expected: all pass
**Step 4: Commit**
```bash
git add backend/prompts.py backend/agents/construct_agent.py
git commit -m "feat: constructability prompt treats disputed extracted values as unverified"
```
---
## Phase 2 — Cross-sheet xref linking
### Task 4: detail-reference / tag xref buckets in the linker (TDD)
**Objective:** Assertions sharing a `detail_reference` or a member `tag` get linked across levels, so S101/S205/S401 details of the same physical element land in one scope.
**Files:**
- Modify: `backend/agents/linker.py` (`build_link_scopes`, line 33)
- Test: `tests/agents/test_linker_xref.py`
**Step 1: Write failing test**
```python
# tests/agents/test_linker_xref.py
from backend.agents.base import AgentScope
from backend.agents.linker import build_link_scopes
def _sheet(number, page, level, assertions):
return {"sheet_number": number, "page_number": page,
"discipline": "Structural", "level": level,
"assertions": assertions}
def _assertion(id_, ref=None, tag=None, level=None):
return {"id": id_, "attribute": "stud_pack_size", "value": "(5) 2x6",
"source_text": "(5) 2x6 STUD PACK",
"location_key": {"detail_reference": ref, "tag": tag,
"level": level}}
def test_xref_scope_joins_same_detail_reference_across_levels():
sheets = [
_sheet("S101", 10, "foundation", [_assertion("a1", ref="A/S205")]),
_sheet("S205", 20, "roof", [_assertion("a2", ref="A/S205")]),
_sheet("S401", 30, "roof", [_assertion("a3", ref="A/S205")]),
]
scopes = build_link_scopes(sheets)
xref = [s for s in scopes if s.scope_id.startswith("xref:")]
assert xref, "expected a cross-level detail-reference scope"
ids = {a["id"] for s in xref for a in s.payload["assertions"]}
assert ids == {"a1", "a2", "a3"}
def test_xref_scope_requires_two_distinct_sheets():
sheets = [
_sheet("S401", 30, "roof", [_assertion("a1", ref="A/S205"),
_assertion("a2", ref="A/S205")]),
]
scopes = build_link_scopes(sheets)
assert not [s for s in scopes if s.scope_id.startswith("xref:")]
def test_xref_scope_joins_shared_member_tag():
sheets = [
_sheet("S102", 5, "roof", [_assertion("a1", tag="HSS16X4X5/8")]),
_sheet("S401", 30, "unknown", [_assertion("a2", tag="HSS16X4X5/8")]),
]
scopes = build_link_scopes(sheets)
xref = [s for s in scopes if s.scope_id.startswith("xref:")]
assert xref
```
**Step 2: Run test to verify failure**
Run: `.venv/bin/python -m pytest tests/agents/test_linker_xref.py -v`
Expected: FAIL — no `xref:` scopes produced
**Step 3: Implement**
Rewrite `build_link_scopes` in `backend/agents/linker.py` (keep the existing
`(level, family)` bucketing, add the xref pass):
```python
def _xref_keys(assertion: Dict) -> List[str]:
"""Cross-level join keys: detail references and member tags."""
location = assertion.get("location_key") or {}
keys = []
ref = re.sub(r"\s+", "", str(location.get("detail_reference") or "")).upper()
if ref:
keys.append(f"detail:{ref}")
tag = re.sub(r"\s+", "", str(location.get("tag") or "")).upper()
if re.match(r"^[A-Z]{2,}\d", tag): # member marks: HSS16X4X5/8, W12X26, ...
keys.append(f"tag:{tag}")
return keys
def build_link_scopes(sheets: List[Dict]) -> List[AgentScope]:
"""Partition facts by level and object/tag family, then enforce a hard cap.
A second pass joins assertions that share a detail_reference or member tag
ACROSS levels, so plan/detail/section sheets describing the same physical
element are linked together even though their levels differ.
"""
buckets: Dict[Tuple[str, str], List[Dict]] = defaultdict(list)
xref: Dict[str, List[Dict]] = defaultdict(list)
for sheet in sheets:
for assertion in sheet.get("assertions", []):
enriched = {
**assertion,
"discipline": sheet.get("discipline") or "Unknown",
"sheet_number": sheet.get("sheet_number"),
"page_number": sheet.get("page_number"),
}
level = str((assertion.get("location_key") or {}).get("level")
or sheet.get("level") or "unknown").lower()
buckets[(level, _family(assertion))].append(enriched)
for key in _xref_keys(assertion):
xref[key].append(enriched)
scopes: List[AgentScope] = []
cap = max(2, config.AGENT_LINK_MAX_ASSERTIONS)
for (level, family), assertions in sorted(buckets.items()):
for offset in range(0, len(assertions), cap):
chunk = assertions[offset:offset + cap]
if len(chunk) < 2:
continue
scopes.append(AgentScope(
scope_id=f"{level}:{family}:{offset // cap + 1}",
payload={"assertions": chunk, "level": level, "family": family},
))
for key, assertions in sorted(xref.items()):
sheets_present = {a.get("sheet_number") for a in assertions}
if len(assertions) < 2 or len(sheets_present) < 2:
continue
scopes.append(AgentScope(
scope_id=f"xref:{key}",
payload={"assertions": assertions[:cap],
"level": "xref", "family": key},
))
return scopes
```
**Step 4: Run tests**
Run: `.venv/bin/python -m pytest tests/agents/test_linker_xref.py tests/agents/ -v`
Expected: all pass (watch existing runner tests for scope-count coupling)
**Step 5: Commit**
```bash
git add backend/agents/linker.py tests/agents/test_linker_xref.py
git commit -m "feat: cross-level xref link scopes via detail_reference and member tag"
```
---
## Phase 3 — Evidence verification wave (5b)
### Task 5: Config knobs + verify prompts
**Objective:** Add the tuning surface and prompts for the vision fact-check agent.
**Files:**
- Modify: `backend/config.py` (near line 44, with the other AGENT_* knobs)
- Modify: `backend/prompts.py` (append near the CONFLICT prompts, ~line 460)
- Modify: `backend/.env.example`
**Step 1: Add config knobs to `backend/config.py`**
```python
AGENT_VERIFY_MODEL = os.getenv("AGENT_VERIFY_MODEL", "") or MODEL
AGENT_VERIFY_CONCURRENCY = int(os.getenv("AGENT_VERIFY_CONCURRENCY", "4"))
AGENT_VERIFY_MAX_CHECKS = int(os.getenv("AGENT_VERIFY_MAX_CHECKS", "20"))
AGENT_VERIFY_SEVERITIES = {
s.strip().lower()
for s in os.getenv("AGENT_VERIFY_SEVERITIES", "critical,high").split(",")
if s.strip()
}
AGENT_VERIFY_REASONING_EFFORT = os.getenv("AGENT_VERIFY_REASONING_EFFORT", "low").strip()
VERIFY_MAX_TOKENS = int(os.getenv("VERIFY_MAX_TOKENS", "8192"))
```
Append to `backend/.env.example`:
```
# Wave 5b evidence verification (vision fact-check of cited sheet text)
AGENT_VERIFY_MAX_CHECKS=20
AGENT_VERIFY_SEVERITIES=critical,high
AGENT_VERIFY_REASONING_EFFORT=low
VERIFY_MAX_TOKENS=8192
```
**Step 2: Add prompts to `backend/prompts.py`**
```python
VERIFY_SYSTEM_PROMPT = """You are a meticulous construction document checker verifying machine-extracted evidence against the actual drawing sheet images.
For each evidence item you are given the sheet it was extracted from and the verbatim text the extractor claims appears there.
Judge each item against the images:
- confirmed: the text (or an obvious equivalent) appears on the cited sheet and means what the finding claims.
- corrected: the sheet shows a DIFFERENT value than the extracted text. Give the actual verbatim text.
- not_found: nothing like the extracted text appears on the cited sheet.
Be strict about numbers, quantities, and member sizes: "(2) 2x6" and "(5) 2x6" are different values. HSS16x4 and HSS16x16 are different values.
Use plain ASCII only.
Respond only with valid JSON."""
VERIFY_USER_INSTRUCTION = """Verify this finding's evidence against the attached sheet images.
Respond ONLY with a valid JSON object - no markdown fences, no explanation:
{ "verdicts": [ { "sheet": "string", "source_text": "the evidence text judged", "verdict": "confirmed | corrected | not_found", "actual_text": "verbatim sheet text when corrected, else null", "notes": "string or null" } ] }
Finding: {finding}"""
```
**Step 3: Sanity check**
Run: `.venv/bin/python -c "from backend import config, prompts; print(config.AGENT_VERIFY_MAX_CHECKS, config.AGENT_VERIFY_SEVERITIES); print(prompts.VERIFY_SYSTEM_PROMPT[:40])"`
Expected: `20 {'critical', 'high'}` and prompt text
**Step 4: Commit**
```bash
git add backend/config.py backend/prompts.py backend/.env.example
git commit -m "feat: config knobs and prompts for evidence verification wave"
```
---
### Task 6: `EvidenceVerifierAgent` + runner wave 5b (TDD)
**Objective:** Re-read cited sheet images for selected findings; annotate verified findings, suppress refuted ones before the Brain merge.
**Files:**
- Create: `backend/agents/verifier.py`
- Modify: `backend/agents/runner.py` (new wave between wave 5 and wave 6, ~line 167)
- Modify: `backend/agents/construct_agent.py` line 67 (stamp `cluster_key` for dispute-based selection)
- Test: `tests/agents/test_verifier.py`
**Step 1: Write failing test**
```python
# tests/agents/test_verifier.py
from unittest.mock import patch
from backend.agents.base import AgentScope, AgentUsage
from backend.agents.verifier import (
EvidenceVerifierAgent, apply_verdicts, select_findings,
)
def _finding(sev="critical", issue_id="i1", sheets=("S401",), cluster_key=None):
f = {"issue_id": issue_id, "severity": sev, "confidence": "high",
"source_stage": "constructability", "sheets": list(sheets),
"description": "HSS16x4 on (2) 2x6 STUD PACK is unbuildable",
"evidence": [{"sheet": "S401", "source_text": "(2) 2x6 STUD PACK",
"asserted_value": "3-inch width"}]}
if cluster_key:
f["cluster_key"] = cluster_key
return f
def test_select_findings_by_severity_and_dispute():
findings = [_finding("critical"), _finding("low", "i2"),
_finding("medium", "i3", cluster_key="c9")]
clusters = [{"key": "c9", "disputed_attributes": [{"attribute": "a"}]}]
selected = select_findings(findings, clusters, max_checks=20,
severities={"critical", "high"})
assert [f["issue_id"] for f in selected] == ["i1", "i3"]
def test_select_findings_respects_cap():
findings = [_finding("critical", f"i{n}") for n in range(30)]
selected = select_findings(findings, [], max_checks=5,
severities={"critical"})
assert len(selected) == 5
def test_run_attaches_verdicts_and_marks_refuted():
agent = EvidenceVerifierAgent(usage=AgentUsage())
scope = AgentScope(scope_id="verify:0", payload={
"finding_index": 0,
"finding": _finding(),
"images_b64": ["QUJD"],
})
verdicts = {"verdicts": [
{"sheet": "S401", "source_text": "(2) 2x6 STUD PACK",
"verdict": "corrected", "actual_text": "(5) 2x6 STUD PACK",
"notes": "callout reads (5)"},
]}
with patch("backend.agents.verifier.call_json", return_value=verdicts):
result = agent.run(scope)
assert not result.error
artifact = result.artifacts[0]
assert artifact["finding_index"] == 0
assert artifact["status"] == "refuted" # no evidence confirmed
assert artifact["verdicts"][0]["actual_text"] == "(5) 2x6 STUD PACK"
def test_apply_verdicts_annotates_and_suppresses():
findings = [_finding("critical", "i1"), _finding("high", "i2")]
from backend.agents.base import AgentResult
results = [AgentResult(scope_id="verify:0", artifacts=[
{"finding_index": 0, "status": "refuted", "verdicts": []},
{"finding_index": 1, "status": "confirmed", "verdicts": []},
])]
suppressed = apply_verdicts(findings, results)
assert suppressed == [findings[0]]
assert findings[0]["verification"]["status"] == "refuted"
assert findings[1]["verification"]["status"] == "confirmed"
```
**Step 2: Run test to verify failure**
Run: `.venv/bin/python -m pytest tests/agents/test_verifier.py -v`
Expected: FAIL — `ModuleNotFoundError: backend.agents.verifier`
**Step 3: Implement `backend/agents/verifier.py`**
```python
"""Wave 5b: vision fact-check of extracted evidence against cited sheet images.
Downstream specialists are text-only; a wave-1 vision misread ("(2) 2x6" vs
"(5) 2x6") otherwise becomes immutable ground truth. For high-severity or
dispute-linked findings, re-read the cited sheets and adjudicate each evidence
item: confirmed / corrected / not_found. Findings whose evidence is entirely
unconfirmed are suppressed before the Brain merge.
"""
from typing import Dict, List, Optional, Set
from backend import config
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
from backend.llm import call_json
from backend.pipeline._serialize import dumps
from backend.pipeline._stage import collect_list, render
from backend.prompts import VERIFY_SYSTEM_PROMPT, VERIFY_USER_INSTRUCTION
_SEVERITY_RANK = {"critical": 0, "high": 1, "medium": 2, "low": 3}
_VERDICTS = ("confirmed", "corrected", "not_found")
def select_findings(
findings: List[Dict],
clusters: List[Dict],
max_checks: int,
severities: Set[str],
) -> List[Dict]:
"""Severity-gated selection plus any finding tied to a disputed cluster."""
disputed_keys = {
cluster.get("key") for cluster in clusters
if cluster.get("disputed_attributes")
}
selected = [
finding for finding in findings
if str(finding.get("severity") or "").lower() in severities
or finding.get("cluster_key") in disputed_keys
]
selected.sort(key=lambda f: _SEVERITY_RANK.get(
str(f.get("severity") or "").lower(), 9))
return selected[:max_checks]
def _valid_verdict(item: Dict) -> Optional[Dict]:
if not isinstance(item, dict):
return None
verdict = str(item.get("verdict") or "").lower()
if verdict not in _VERDICTS:
return None
return {
"sheet": item.get("sheet") or "",
"source_text": item.get("source_text") or "",
"verdict": verdict,
"actual_text": item.get("actual_text"),
"notes": item.get("notes"),
}
def _status(verdicts: List[Dict]) -> str:
if not verdicts:
return "unverified"
confirmed = sum(1 for v in verdicts if v["verdict"] == "confirmed")
if confirmed == len(verdicts):
return "confirmed"
if confirmed == 0:
return "refuted"
return "mixed"
class EvidenceVerifierAgent:
name = "verify"
def __init__(self, usage: AgentUsage) -> None:
self.usage = usage
def run(self, scope: AgentScope) -> AgentResult:
try:
finding = scope.payload["finding"]
instruction = render(
VERIFY_USER_INSTRUCTION, {"finding": dumps(finding)}
)
parsed = call_json(
system_prompt=VERIFY_SYSTEM_PROMPT,
user_text=instruction,
images_b64=scope.payload.get("images_b64") or [],
max_tokens=config.VERIFY_MAX_TOKENS,
model=config.AGENT_VERIFY_MODEL,
reasoning_effort=config.AGENT_VERIFY_REASONING_EFFORT or None,
usage_tracker=self.usage,
usage_stage="agent.verify",
)
verdicts = collect_list(parsed, "verdicts", _valid_verdict)
return AgentResult(scope_id=scope.scope_id, artifacts=[{
"finding_index": scope.payload["finding_index"],
"status": _status(verdicts),
"verdicts": verdicts,
}])
except Exception as exc:
return failure(scope, exc)
def apply_verdicts(
findings: List[Dict], verify_results: List[AgentResult]
) -> List[Dict]:
"""Annotate findings with verification; return refuted ones to suppress."""
by_index: Dict[int, Dict] = {}
for result in verify_results:
for artifact in result.artifacts:
by_index[artifact["finding_index"]] = artifact
suppressed = []
for index, finding in enumerate(findings):
artifact = by_index.get(index)
if not artifact:
continue
finding["verification"] = {
"status": artifact["status"],
"verdicts": artifact["verdicts"],
}
if artifact["status"] == "refuted":
finding["confidence"] = "low"
suppressed.append(finding)
return suppressed
```
**Step 4: Stamp `cluster_key` on constructability findings**
In `backend/agents/construct_agent.py` line 66-67, change:
```python
for finding in findings:
finding.update(agent=self.name, scope_id=scope.scope_id)
```
to:
```python
for finding in findings:
finding.update(agent=self.name, scope_id=scope.scope_id,
cluster_key=cluster.get("key"))
```
**Step 5: Wire wave 5b into `backend/agents/runner.py`**
After `memory.extend("findings", specialist_findings)` (line 167) and before
`gap_findings` / wave 6:
```python
orchestrator.stage("Agent wave 5b: evidence verification")
sheet_to_page = {
sheet.get("sheet_number"): sheet.get("page_number") for sheet in sheets
}
verify_targets = select_findings(
specialist_findings, clusters,
max_checks=config.AGENT_VERIFY_MAX_CHECKS,
severities=config.AGENT_VERIFY_SEVERITIES,
)
target_indexes = {id(f): i for i, f in enumerate(specialist_findings)}
verify_scopes = [
AgentScope(
scope_id=f"verify:{target_indexes[id(finding)]}",
payload={
"finding_index": target_indexes[id(finding)],
"finding": finding,
"images_b64": [
page_to_b64[sheet_to_page[name]]
for name in (finding.get("sheets") or [])
[:config.AGENT_CONFLICT_MAX_IMAGES]
if sheet_to_page.get(name) in page_to_b64
],
},
)
for finding in verify_targets
]
verify_results = orchestrator.run_scopes(
EvidenceVerifierAgent(usage), verify_scopes,
config.AGENT_VERIFY_CONCURRENCY,
)
suppressed = apply_verdicts(specialist_findings, verify_results)
if suppressed:
suppressed_ids = {id(f) for f in suppressed}
specialist_findings = [
f for f in specialist_findings if id(f) not in suppressed_ids
]
memory.replace("suppressed", suppressed)
memory.extend("findings", specialist_findings) # see note below
```
NOTE for implementer: `memory.extend("findings", ...)` already ran with the
un-suppressed list. Adjust ordering so verification happens BEFORE
`memory.extend("findings", specialist_findings)` — i.e. move the extend to after
wave 5b — so the Brain never sees refuted findings. Keep `gap_findings` logic
unchanged. Also add imports at top of runner.py:
```python
from backend.agents.verifier import (
EvidenceVerifierAgent, apply_verdicts, select_findings,
)
```
And in the report dicts (both the `require_review` branch ~line 217-225 and the
wave-7 branch ~line 293-300), populate suppressed issues:
```python
"suppressed_issues": memory.snapshot().get("suppressed") or [],
```
Finally, `verify` results cost shows up as `agent.verify` in
`summary.cost_by_stage` automatically via `usage_stage="agent.verify"`.
**Step 6: Update existing runner tests**
`tests/agents/test_runner_review_gate.py` monkeypatches `BrainAgent` and
`convert_pdf_to_images` but lets waves 1-5 run against... check how LLM calls
are stubbed there (likely `call_json` returns None -> empty artifacts, which is
fine). The new wave must no-op cleanly when `select_findings` returns `[]`
(zero scopes -> `run_scopes` returns `[]` per orchestrator.py:60). Verify by
running the suite; if a runner test now fails because verification selects a
stubbed finding, monkeypatch `select_findings` to `lambda *a, **k: []` in that
test file's `_patch_brain` helper.
**Step 7: Run full suite**
Run: `.venv/bin/python -m pytest tests/ -q`
Expected: all pass
**Step 8: Commit**
```bash
git add backend/agents/verifier.py backend/agents/runner.py backend/agents/construct_agent.py tests/agents/test_verifier.py tests/agents/test_runner_review_gate.py
git commit -m "feat: wave 5b evidence verification - vision fact-check before Brain merge"
```
---
### Task 7 (small, related): reasoning budget for the conflict critic
**Objective:** Fix the wave-4 truncation found in this same job (13/121 calls hit `finish_reason=length` at the 4096 cap with ~3.7k thinking tokens).
**Files:**
- Modify: `backend/agents/conflict_critic.py:55-63`
- Modify: `backend/pipeline/conflict_checker.py:85` (same pattern, classic path)
**Step 1: Apply the extractor's reasoning-knob pattern**
```python
parsed = call_json(
system_prompt=CONFLICT_SYSTEM_PROMPT,
user_text=instruction,
images_b64=images,
max_tokens=config.REASON_MAX_TOKENS,
model=config.AGENT_CONFLICT_MODEL,
reasoning_effort=config.EXTRACT_REASONING_EFFORT or None,
reasoning_max_tokens=config.EXTRACT_REASONING_MAX_TOKENS or None,
usage_tracker=self.usage,
usage_stage="agent.conflict",
)
```
(Match the exact kwarg names `SheetExtractorAgent` uses — check
`backend/agents/extractors.py` for whether it passes `None` when the budget is
0, and mirror that guard.)
**Step 2: Run tests**
Run: `.venv/bin/python -m pytest tests/ -q`
Expected: all pass
**Step 3: Commit**
```bash
git add backend/agents/conflict_critic.py backend/pipeline/conflict_checker.py
git commit -m "fix: reasoning budget for conflict critic (wave-4 max_tokens truncation)"
```
---
## Tests / validation
1. `.venv/bin/python -m pytest tests/ -q` — full suite green.
2. **Targeted repro of the original failure:** pull the S401 page image from job
959e16407573's output dir on sits-docker (or re-render the PDF page), then run
one `EvidenceVerifierAgent` scope locally against the finding JSON from
`validated_issues[3]`. Expected: verdict `corrected`,
`actual_text: "(5) 2x6 STUD PACK"`, status `refuted`.
3. **End-to-end:** rerun the same Cypress TX PDF through the pipeline (local with
`LLM_CACHE`/`LLM_RAW_DUMP` per the conflict-checker skill). Expected:
- log shows `Agent wave 5b: evidence verification` with a bounded number of calls;
- the S401 stud-pack finding is either absent from `validated_issues` and present
in `suppressed_issues` with verification verdicts, or downgraded to low confidence;
- `agent.verify` appears in `summary.cost_by_stage`;
- zero `finish_reason=length` lines in wave 4 (Task 7).
4. Cost check: wave 5b adds at most `AGENT_VERIFY_MAX_CHECKS` (20) vision calls —
for this job's profile that is well under $1.
## Risks, tradeoffs, open questions
- **Dispute false positives:** cluster members with legitimately different values
(e.g. two doors in one door cluster) will produce `disputed_attributes`. Mitigation:
prompts treat disputes as "unverified", not "wrong"; only severity-gated findings
burn verification calls. Tune later by restricting `find_disputes` to numeric-ish
values if noise is high.
- **Verifier can also misread.** It is one model checking another with the same eyes.
Mitigation: verdict requires `actual_text` verbatim evidence for `corrected`, and
only fully-unconfirmed findings are suppressed (mixed keeps the finding with a note).
- **xref cost:** extra link scopes. Bounded by the >= 2 distinct sheets gate and the
existing assertion cap; expect a handful of extra scopes per set.
- **Suppression in review mode:** refuted findings land in `suppressed_issues` — the
review UI/finalizer must tolerate that list being non-empty (it is currently always
`[]` in agent mode). Open question: surface suppressed items in the human review
queue as informational, or keep them report-only?
- **Open question:** should wave-4 conflict findings (which already saw images) also be
verification-eligible? Plan says no (they had the pixels); revisit if critics show
the same misread pattern.
+15
View File
@@ -78,3 +78,18 @@ AGENT_VERIFY_MAX_CHECKS=20
AGENT_VERIFY_SEVERITIES=critical,high
AGENT_VERIFY_REASONING_EFFORT=low
VERIFY_MAX_TOKENS=8192
# Text-layer grounding (deterministic PDF text layer via PyMuPDF)
# TEXT_LAYER_ENABLED: master switch for text-layer extraction/grounding
# TEXT_LAYER_MIN_CHARS: below this per page the sheet stays vision-only
# TEXT_LAYER_MAX_CHARS: cap of text layer injected into the extractor prompt
# VERIFY_TEXT_MAX_CHARS: cap of the text-layer excerpt in verify scopes
# VERIFY_HI_DPI_CROPS: evidence-located high-DPI crops in the verifier
# VERIFY_CROP_DPI / VERIFY_CROP_MARGIN_PTS: crop render DPI / padding (PDF points)
TEXT_LAYER_ENABLED=true
TEXT_LAYER_MIN_CHARS=20
TEXT_LAYER_MAX_CHARS=12000
VERIFY_TEXT_MAX_CHARS=8000
VERIFY_HI_DPI_CROPS=true
VERIFY_CROP_DPI=300
VERIFY_CROP_MARGIN_PTS=36
+4 -3
View File
@@ -6,7 +6,7 @@ from typing import Dict
from backend import config
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
from backend.llm import call_json
from backend.pipeline.extractor import _normalize_sheet
from backend.pipeline.extractor import _normalize_sheet, _text_layer_block
from backend.pipeline.sheet_index import _index_input
from backend.prompts import (
EXTRACTOR_SYSTEM_PROMPT,
@@ -65,7 +65,7 @@ class SheetExtractorAgent:
page = scope.payload["page"]
instruction = EXTRACTOR_USER_INSTRUCTION.replace(
"{sheet_hint}", str(scope.payload.get("sheet_hint") or "")
)
) + _text_layer_block(page)
parsed = _wrap_bare_list(self._call(instruction, page),
page["page_number"])
if not isinstance(parsed, dict):
@@ -79,7 +79,8 @@ class SheetExtractorAgent:
)
if not isinstance(parsed, dict):
raise ValueError("no structured extraction returned")
sheet = _normalize_sheet(parsed, page["page_number"])
sheet = _normalize_sheet(parsed, page["page_number"],
page_text=page.get("text_layer"))
return AgentResult(scope_id=scope.scope_id, artifacts=[sheet])
except Exception as exc:
return failure(scope, exc)
+72 -3
View File
@@ -1,5 +1,6 @@
"""Public entry point for the scoped Agent-mode pipeline."""
import base64
import json
import os
from typing import Callable, Dict, Optional
@@ -30,6 +31,9 @@ from backend.pipeline.report import build_report, to_markdown
from backend.pipeline.sheet_index import derive_project_meta_from_cover
from backend.review.gate import build_review_queue
from backend.review.store import ReviewStore
from backend.text_layer import (
attach_text_layers, coverage_gaps, find_evidence_bbox, render_crop,
)
def run_agent_pipeline(
@@ -57,6 +61,9 @@ def run_agent_pipeline(
orchestrator.stage("Agent ingest: PDF -> images")
pages = convert_pdf_to_images(pdf_path)
page_to_b64 = {page["page_number"]: page["base64"] for page in pages}
text_dir = os.path.join(agent_dir, "text") if agent_dir else None
page_words = attach_text_layers(pdf_path, pages, text_dir=text_dir)
page_to_text = {page["page_number"]: page.get("text_layer") for page in pages}
orchestrator.stage("Agent wave 1: extract sheets")
extract_scopes = [
@@ -77,6 +84,13 @@ def run_agent_pipeline(
sheets.sort(key=lambda sheet: sheet.get("page_number") or 0)
memory.replace("sheets", sheets)
memory.dump("01-extract.json")
# Coverage signal: text layer present but extraction failed/empty reuses
# the failed-scopes gap-finding path (finding built below wave 6).
for gap_page in coverage_gaps(pages, sheets):
orchestrator.stats.failed_scopes.append(
f"sheet_extractor:sheet:{gap_page}: extraction gap "
f"(text layer present, no objects extracted)"
)
cover_meta = derive_project_meta_from_cover(
sheets, source_name or os.path.basename(pdf_path)
@@ -183,19 +197,34 @@ def run_agent_pipeline(
target_indexes = {id(f): i for i, f in enumerate(specialist_findings)}
verify_scopes = []
for finding in verify_targets:
images = [
page_to_b64[sheet_to_page[str(name)]]
for name in (finding.get("sheets") or [])[:config.AGENT_CONFLICT_MAX_IMAGES]
cited_pages = [
sheet_to_page[str(name)]
for name in (finding.get("sheets") or [])
if sheet_to_page.get(str(name)) in page_to_b64
]
images = [
page_to_b64[p]
for p in cited_pages[:config.AGENT_CONFLICT_MAX_IMAGES]
]
if not images:
continue # never judge evidence against images we could not load
# Text oracle: concatenated text layer of the cited sheets, capped.
excerpt = "\n\n".join(
f"--- Page {p} ---\n{page_to_text[p]}"
for p in cited_pages
if page_to_text.get(p)
)[:config.VERIFY_TEXT_MAX_CHARS]
if config.VERIFY_HI_DPI_CROPS:
images = _evidence_crops(finding, cited_pages, sheet_to_page,
page_words, page_to_b64, pdf_path,
fallback=images)
verify_scopes.append(AgentScope(
scope_id=f"verify:{target_indexes[id(finding)]}",
payload={
"finding_index": target_indexes[id(finding)],
"finding": finding,
"images_b64": images,
"text_layer_excerpt": excerpt,
},
))
verify_results = orchestrator.run_scopes(
@@ -394,6 +423,46 @@ def _dump(out_dir: str, name: str, value) -> None:
json.dump(value, f, indent=2)
def _evidence_crops(
finding: Dict,
cited_pages: list,
sheet_to_page: Dict,
page_words: Dict,
page_to_b64: Dict,
pdf_path: str,
fallback: list,
) -> list:
"""High-DPI crops around each evidence item's source_text, located via the
page text layer. Crops REPLACE full-page images when at least one evidence
location resolves confidently; otherwise the full-page fallback is kept.
Never returns an empty list when fallback is non-empty (I2 guard)."""
crops: list = []
for item in finding.get("evidence") or []:
if len(crops) >= config.AGENT_CONFLICT_MAX_IMAGES:
break
if not isinstance(item, dict):
continue
source_text = item.get("source_text") or ""
if not source_text:
continue
# Prefer the page named on the evidence item, then any cited page.
candidates = []
named_page = sheet_to_page.get(str(item.get("sheet") or ""))
if named_page in cited_pages:
candidates.append(named_page)
candidates.extend(p for p in cited_pages if p not in candidates)
for page in candidates:
bbox = find_evidence_bbox(page_words.get(page) or [], source_text)
if bbox is None:
continue
crop = render_crop(pdf_path, page, bbox)
if not crop:
continue
crops.append(base64.b64encode(crop).decode("utf-8"))
break
return crops or fallback
def _counts(items, key: str) -> Dict[str, int]:
counts: Dict[str, int] = {}
for item in items:
+5 -1
View File
@@ -55,7 +55,11 @@ class EvidenceVerifierAgent:
def run(self, scope: AgentScope) -> AgentResult:
try:
finding = scope.payload["finding"]
instruction = render(VERIFY_USER_INSTRUCTION, {"finding": dumps(finding)})
instruction = render(VERIFY_USER_INSTRUCTION, {
"finding": dumps(finding),
"text_layer": scope.payload.get("text_layer_excerpt")
or "(no text layer available for the cited sheets)",
})
parsed = call_json(
system_prompt=VERIFY_SYSTEM_PROMPT,
user_text=instruction,
+12
View File
@@ -59,6 +59,18 @@ AGENT_VERIFY_SEVERITIES = {
AGENT_VERIFY_REASONING_EFFORT = os.getenv("AGENT_VERIFY_REASONING_EFFORT", "low").strip()
VERIFY_MAX_TOKENS = int(os.getenv("VERIFY_MAX_TOKENS", "8192"))
# -- Text-layer grounding (deterministic PDF text layer via PyMuPDF) ----
# The vector text layer is extracted once per job and grounds the extractor,
# rescues misquoted-but-real values in the grounding guard, and serves the
# wave-5b verifier as a text oracle plus high-DPI evidence crops.
TEXT_LAYER_ENABLED = os.getenv("TEXT_LAYER_ENABLED", "true").strip().lower() in ("1", "true", "yes")
TEXT_LAYER_MIN_CHARS = int(os.getenv("TEXT_LAYER_MIN_CHARS", "20")) # below this per page -> no text layer
TEXT_LAYER_MAX_CHARS = int(os.getenv("TEXT_LAYER_MAX_CHARS", "12000")) # cap per sheet in extractor prompt
VERIFY_TEXT_MAX_CHARS = int(os.getenv("VERIFY_TEXT_MAX_CHARS", "8000"))# cap of excerpt in verify scope
VERIFY_HI_DPI_CROPS = os.getenv("VERIFY_HI_DPI_CROPS", "true").strip().lower() in ("1", "true", "yes")
VERIFY_CROP_DPI = int(os.getenv("VERIFY_CROP_DPI", "300"))
VERIFY_CROP_MARGIN_PTS = int(os.getenv("VERIFY_CROP_MARGIN_PTS", "36"))# padding around evidence bbox (PDF points)
# Agent-mode human-review gate. When on (default), Agent runs stop after the
# Brain merge and wait for human decisions before RFIs/final report/email go
# out. AGENT_REVIEW_AUDIT_SAMPLE caps how many clean clusters get added to the
+57 -8
View File
@@ -64,7 +64,8 @@ def discipline_from_sheet_number(sheet_number: Optional[str]) -> Optional[str]:
return None
def _is_grounded(value: str, source_text: str, graphical_basis: str = "") -> bool:
def _is_grounded(value: str, source_text: str, graphical_basis: str = "",
page_text: Optional[str] = None) -> bool:
"""
Keep an object only if its primary value is supported by its source_text,
OR it is a graphical object (has graphical_basis with no text to quote).
@@ -72,6 +73,9 @@ def _is_grounded(value: str, source_text: str, graphical_basis: str = "") -> boo
- If graphical_basis is set and source_text is absent, the object is valid.
- If the value contains digits, every distinct digit-run must appear in
source_text (catches invented dimensions/counts/elevations).
- Rescue tier: when page_text (the deterministic text layer) is given,
digit-runs absent from source_text but present in the page text are
still grounded - vision quoted imperfectly but the value is real.
- If the value has no digits, require some alphabetic-token overlap.
"""
# Graphical objects (no readable text on sheet) are always allowed through.
@@ -85,7 +89,11 @@ def _is_grounded(value: str, source_text: str, graphical_basis: str = "") -> boo
val_digits = set(_DIGITS_RE.findall(value))
if val_digits:
src_digits = set(_DIGITS_RE.findall(source_text))
return val_digits.issubset(src_digits)
if val_digits.issubset(src_digits):
return True
if page_text:
return val_digits.issubset(set(_DIGITS_RE.findall(page_text)))
return False
# No digits: text-based grounding.
val_norm = re.sub(r"[^a-z0-9]+", " ", value.lower()).strip()
@@ -109,7 +117,24 @@ def _primary_value(obj: Dict) -> str:
or obj.get("name") or obj.get("tag") or "")
def _normalize_sheet(parsed: Dict, page_number: int) -> Dict:
def _grounding_stamp(value: str, source_text: str,
page_text: Optional[str]) -> Optional[str]:
"""\"text_layer\" when the object survived only via the text-layer rescue
tier (digits absent from source_text but present in the page text)."""
if not page_text:
return None
val_digits = set(_DIGITS_RE.findall(str(value)))
if not val_digits:
return None
if val_digits.issubset(set(_DIGITS_RE.findall(source_text))):
return None
if val_digits.issubset(set(_DIGITS_RE.findall(page_text))):
return "text_layer"
return None
def _normalize_sheet(parsed: Dict, page_number: int,
page_text: Optional[str] = None) -> Dict:
"""
Validate + clean one parsed sheet result, attaching page_number and ids.
@@ -138,6 +163,7 @@ def _normalize_sheet(parsed: Dict, page_number: int) -> Dict:
raw_objects = parsed.get("objects") or parsed.get("assertions") or []
clean: List[Dict] = []
dropped = 0
rescued = 0
for idx, obj in enumerate(raw_objects):
if not isinstance(obj, dict):
@@ -149,9 +175,13 @@ def _normalize_sheet(parsed: Dict, page_number: int) -> Dict:
# Derive a primary value for the grounding check
primary_val = _primary_value(obj)
if not _is_grounded(primary_val, source_text, graphical_basis):
if not _is_grounded(primary_val, source_text, graphical_basis,
page_text=page_text):
dropped += 1
continue
grounding = _grounding_stamp(primary_val, source_text, page_text)
if grounding:
rescued += 1
# --- location_key: new schema is richer; map to legacy shape + extras ---
lk = obj.get("location_key")
@@ -203,10 +233,13 @@ def _normalize_sheet(parsed: Dict, page_number: int) -> Dict:
"object_attributes": attrs,
"graphical_basis": graphical_basis or None,
"review_uses": obj.get("review_uses") or [],
**({"grounding": grounding} if grounding else {}),
})
if dropped:
print(f"[Extract] Page {page_number} ({sheet_number}): dropped {dropped} ungrounded object(s)")
if dropped or rescued:
print(f"[Extract] Page {page_number} ({sheet_number}): "
f"dropped {dropped} ungrounded object(s)"
+ (f", rescued {rescued} via text layer" if rescued else ""))
unresolved = parsed.get("unresolved_items") or []
@@ -223,8 +256,23 @@ def _normalize_sheet(parsed: Dict, page_number: int) -> Dict:
}
def _text_layer_block(page: Dict) -> str:
"""
The TEXT LAYER block appended to the extractor instruction at call sites
(NOT a template placeholder - render() silently leaves missing keys as
literals). Empty string when the page has no usable text layer.
"""
text = (page.get("text_layer") or "").strip()
if not text:
return ""
return ("\n\nTEXT LAYER (authoritative for alphanumeric content — trust it "
"over the image for numbers, tags, and note text):\n"
+ text[:config.TEXT_LAYER_MAX_CHARS])
def _extract_one(page: Dict, sheet_hint: str = "") -> Dict:
user_text = EXTRACTOR_USER_INSTRUCTION.replace("{sheet_hint}", sheet_hint)
user_text = (EXTRACTOR_USER_INSTRUCTION.replace("{sheet_hint}", sheet_hint)
+ _text_layer_block(page))
parsed = call_json(
system_prompt=EXTRACTOR_SYSTEM_PROMPT,
user_text=user_text,
@@ -241,7 +289,8 @@ def _extract_one(page: Dict, sheet_hint: str = "") -> Dict:
"scale": None,
"assertions": [],
}
return _normalize_sheet(parsed, page["page_number"])
return _normalize_sheet(parsed, page["page_number"],
page_text=page.get("text_layer"))
def extract_assertions(pages: List[Dict], on_progress=None) -> List[Dict]:
+4
View File
@@ -27,6 +27,7 @@ from typing import Dict, Optional, Callable
from backend.pipeline.pdf_processor import convert_pdf_to_images
from backend.pipeline.extractor import extract_assertions
from backend.text_layer import attach_text_layers, coverage_gaps
from backend.pipeline.sheet_index import classify_sheets, derive_project_meta_from_cover
from backend.pipeline.jurisdiction import run_jurisdiction
from backend.pipeline.normalizer import normalize_assertions, build_project_intelligence
@@ -103,9 +104,12 @@ def _run_stages(
) -> Dict:
stage("PDF -> images")
pages = convert_pdf_to_images(pdf_path)
text_dir = os.path.join(out_dir, "text") if out_dir else None
attach_text_layers(pdf_path, pages, text_dir=text_dir)
stage("Extract assertions")
sheets = extract_assertions(pages)
coverage_gaps(pages, sheets) # classic: log-only recall signal
stage("Classify sheet index")
sheet_index = classify_sheets(sheets)
+4 -1
View File
@@ -230,6 +230,7 @@ Rules you must never break:
- Every object must include source_text copied verbatim from the sheet whenever text is available.
- If the object is graphical and has no text, describe it visually and mark confidence low or medium.
- Preserve tags, marks, room numbers, sheet numbers, detail references, and abbreviations exactly as shown.
TEXT LAYER GROUNDING: when a TEXT LAYER block is present in the user message, it is the sheet's deterministic PDF text layer and is authoritative for alphanumeric content (counts, dimensions, member tags, note text). Trust it over your reading of the image for numbers, tags, and note text; quote source_text from it verbatim. Use the image for geometry, symbols, linework, and anything absent from the text layer.
- Use null when information is not determinable.
- Keep objects atomic.
- Use plain ASCII only.
@@ -467,7 +468,9 @@ Respond only with valid JSON."""
VERIFY_USER_INSTRUCTION = """Verify this finding's evidence against the attached sheet images.
Respond ONLY with a valid JSON object - no markdown fences, no explanation:
{ "verdicts": [ { "sheet": "string", "source_text": "the evidence text judged", "verdict": "confirmed | corrected | not_found", "actual_text": "verbatim sheet text when corrected, else null", "notes": "string or null" } ] }
Finding: {finding}"""
Finding: {finding}
TEXT LAYER (deterministic page text extracted from the PDF - an oracle for alphanumeric content such as counts, dimensions, and member tags; when it disagrees with the extracted evidence, trust it and cite it as actual_text):
{text_layer}"""
# ---------------------------------------------------------------------------
+230
View File
@@ -0,0 +1,230 @@
"""
text_layer.py - deterministic PDF text-layer extraction (PyMuPDF, no LLM).
Most CAD-produced drawing sets carry a real vector text layer. We extract it
once per job and feed it to the extractor (grounding), the grounding guard
(rescue tier), and the wave-5b verifier (text oracle + high-DPI evidence
crops). Pages below TEXT_LAYER_MIN_CHARS of text are treated as having no
text layer (scanned/raster sheets stay vision-only).
If PyMuPDF is unavailable the module degrades gracefully: every public
function returns empty/None, equivalent to TEXT_LAYER_ENABLED=false.
"""
import re
from typing import Dict, List, Optional, Tuple
from backend import config
try: # PyMuPDF >= 1.24 prefers the pymupdf name; fitz works everywhere.
import pymupdf as fitz
except ImportError: # pragma: no cover - older PyMuPDF
try:
import fitz
except ImportError: # pragma: no cover - PyMuPDF not installed
fitz = None
_warned_unavailable = False
# Word token normalization for evidence matching: lowercase alphanumeric only.
_TOKEN_RE = re.compile(r"[^a-z0-9]+")
# Fuzzy match floor: fraction of needle tokens that must align with the page's
# word sequence for a bbox to count as a confident evidence location.
_FUZZY_MIN_RATIO = 0.6
def _fitz_or_none():
"""Return the fitz module, logging once if PyMuPDF is missing."""
global _warned_unavailable
if fitz is None and not _warned_unavailable:
print("[TextLayer] PyMuPDF not available - text-layer grounding disabled")
_warned_unavailable = True
return fitz
def extract_text_layers(pdf_path: str) -> Dict[int, Dict]:
"""
Extract the text layer of every page. Returns {1-based page_number:
{"text": str, "words": [{"text", "bbox": (x0,y0,x1,y1)}, ...],
"has_text_layer": bool}}. Returns {} when disabled or unavailable.
"""
if not config.TEXT_LAYER_ENABLED:
return {}
f = _fitz_or_none()
if f is None:
return {}
try:
doc = f.open(pdf_path)
except Exception as exc:
print(f"[TextLayer] could not open {pdf_path}: {exc}")
return {}
layers: Dict[int, Dict] = {}
try:
for index in range(doc.page_count):
page = doc[index]
text = page.get_text("text") or ""
words = [
{"text": w[4], "bbox": (w[0], w[1], w[2], w[3])}
for w in (page.get_text("words") or [])
]
has_text_layer = len(text.strip()) >= config.TEXT_LAYER_MIN_CHARS
if not has_text_layer:
print(f"[TextLayer] Page {index + 1}: {len(text.strip())} chars "
f"(< TEXT_LAYER_MIN_CHARS={config.TEXT_LAYER_MIN_CHARS}) - "
f"vision-only")
layers[index + 1] = {
"text": text,
"words": words,
"has_text_layer": has_text_layer,
}
finally:
doc.close()
return layers
def attach_text_layers(
pdf_path: str,
pages: List[Dict],
text_dir: Optional[str] = None,
) -> Dict[int, List[Dict]]:
"""
Attach page["text_layer"] (text or None) to each converted page dict and
return the runner-local {page_number: words} map (kept off page dicts -
those get serialized). When text_dir is set, dump one .txt per page there
(plain file writes; ProjectMemory is a closed registry).
"""
layers = extract_text_layers(pdf_path)
page_words: Dict[int, List[Dict]] = {}
for page in pages:
layer = layers.get(page["page_number"]) or {}
page["text_layer"] = layer.get("text") if layer.get("has_text_layer") else None
page_words[page["page_number"]] = layer.get("words") or []
if text_dir and layers:
import os
os.makedirs(text_dir, exist_ok=True)
for page_number, layer in layers.items():
if not layer.get("has_text_layer"):
continue
with open(os.path.join(text_dir, f"page-{page_number:03d}.txt"),
"w", encoding="utf-8") as fh:
fh.write(layer.get("text") or "")
return page_words
def _tokens(text: str) -> List[str]:
return [t for t in _TOKEN_RE.split(text.lower()) if t]
def _union_bbox(boxes: List[Tuple[float, float, float, float]]):
return (
min(b[0] for b in boxes),
min(b[1] for b in boxes),
max(b[2] for b in boxes),
max(b[3] for b in boxes),
)
def find_evidence_bbox(
words: List[Dict],
needle: str,
) -> Optional[Tuple[float, float, float, float]]:
"""
Best-effort fuzzy substring match of an evidence source_text against the
page's word sequence. Returns the union bbox of the matched words, or
None when nothing aligns confidently.
Exact contiguous token runs win; otherwise the best-scoring window with
>= _FUZZY_MIN_RATIO token alignment is accepted (vision quotes imperfectly
but the value is real page text).
"""
if not words or not needle:
return None
needle_tokens = _tokens(str(needle))
if not needle_tokens:
return None
page_tokens = [_tokens(w.get("text") or "") for w in words]
# Flatten multi-token words, remembering which word each token came from.
flat: List[Tuple[str, int]] = []
for word_index, parts in enumerate(page_tokens):
for part in parts:
flat.append((part, word_index))
if not flat:
return None
n = len(needle_tokens)
best_span = None
best_score = 0.0
for start in range(0, len(flat)):
window = flat[start:start + n]
if not window:
break
score = sum(1 for i, tok in enumerate(needle_tokens)
if i < len(window) and window[i][0] == tok) / n
if score > best_score:
best_score = score
best_span = window
if best_score == 1.0:
break
if best_span is None or best_score < _FUZZY_MIN_RATIO:
return None
word_indexes = {word_index for _, word_index in best_span}
return _union_bbox([words[i]["bbox"] for i in sorted(word_indexes)])
def render_crop(
pdf_path: str,
page_number: int,
bbox: Tuple[float, float, float, float],
dpi: Optional[int] = None,
margin_pts: Optional[float] = None,
) -> Optional[bytes]:
"""
Render a clip of one page around bbox (+ margin, clamped to the page) at
the given DPI and return JPEG bytes, or None on any failure.
"""
f = _fitz_or_none()
if f is None:
return None
dpi = dpi or config.VERIFY_CROP_DPI
margin_pts = config.VERIFY_CROP_MARGIN_PTS if margin_pts is None else margin_pts
try:
doc = f.open(pdf_path)
try:
page = doc[page_number - 1]
rect = f.Rect(
bbox[0] - margin_pts,
bbox[1] - margin_pts,
bbox[2] + margin_pts,
bbox[3] + margin_pts,
) & page.rect
if rect.is_empty:
return None
pix = page.get_pixmap(clip=rect, dpi=dpi)
return pix.tobytes("jpeg")
finally:
doc.close()
except Exception as exc:
print(f"[TextLayer] render_crop failed on page {page_number}: {exc}")
return None
def coverage_gaps(pages: List[Dict], sheets: List[Dict]) -> List[int]:
"""
Page numbers that have a text layer but whose extraction failed or
returned 0 objects - the silent extraction-loss signal. Logs one
[TextLayer] line per gap.
"""
by_page = {s.get("page_number"): s for s in sheets or []}
gaps: List[int] = []
for page in pages:
text = page.get("text_layer")
if not text:
continue
sheet = by_page.get(page["page_number"])
extracted = len(sheet.get("assertions") or []) if sheet else 0
if extracted == 0:
gaps.append(page["page_number"])
print(f"[TextLayer] Page {page['page_number']}: text layer present "
f"({len(text)} chars) but no objects extracted — possible "
f"extraction gap")
return gaps
+1
View File
@@ -2,6 +2,7 @@ fastapi==0.115.0
uvicorn[standard]==0.30.6
python-multipart==0.0.12
pdf2image==1.17.0
PyMuPDF>=1.23.0 # deterministic text-layer extraction (extractor grounding, verifier crops)
Pillow==10.4.0
openai==1.51.0
httpx==0.27.2 # openai 1.51 passes proxies= to httpx; >=0.28 dropped it
+128
View File
@@ -0,0 +1,128 @@
"""Runner-level text-layer flow: excerpt into verify scopes, hi-DPI crop
replacement with full-page fallback, and coverage-gap findings."""
import pytest
fitz = pytest.importorskip("pymupdf")
import backend.agents.runner as runner_mod
from backend.agents.base import AgentResult
from backend.agents.runner import run_agent_pipeline
PAGE_TEXT = "(5) 2X6 STUD PACK AT BEARING"
def _make_pdf(path):
doc = fitz.open()
page = doc.new_page(width=612, height=792)
page.insert_text((72, 72), PAGE_TEXT, fontsize=11)
doc.save(str(path))
doc.close()
return str(path)
def _finding(sheets, evidence_text):
return {
"issue_id": "C1", "severity": "critical", "confidence": "high",
"source_stage": "constructability", "sheets": sheets,
"description": "stud pack conflict",
"evidence": [{"sheet": sheets[0], "source_text": evidence_text}],
}
def _stub_agent(artifacts):
return lambda usage: type("S", (), {
"name": "stub",
"run": lambda self, scope: AgentResult(
scope_id=scope.scope_id, artifacts=list(artifacts)),
})()
def _patch_pipeline(monkeypatch, finding, verify_sink):
monkeypatch.setattr(
runner_mod, "convert_pdf_to_images",
lambda path: [{"page_number": 1, "base64": "QUJD"}])
monkeypatch.setattr(runner_mod, "SheetExtractorAgent", _stub_agent([
{"sheet_number": "S401", "page_number": 1, "level": "roof",
"discipline": "S", "assertions": [
{"text": "(5) 2X6 STUD PACK", "object_type": "framing"},
{"text": "HSS16X4 beam", "object_type": "framing"},
]},
]))
monkeypatch.setattr(runner_mod, "SheetIndexAgent", _stub_agent([{}]))
monkeypatch.setattr(runner_mod, "JurisdictionAgent", _stub_agent([{}]))
monkeypatch.setattr(runner_mod, "LinkerAgent", _stub_agent([
{"key": "c1", "location": "roof beam pocket", "assertions": []},
]))
monkeypatch.setattr(runner_mod, "ConflictCriticAgent", _stub_agent([]))
monkeypatch.setattr(runner_mod, "CodeAgent", _stub_agent([]))
monkeypatch.setattr(runner_mod, "ConstructabilityAgent",
_stub_agent([finding]))
monkeypatch.setattr(runner_mod, "CompletenessAgent", _stub_agent([]))
monkeypatch.setattr(
runner_mod, "BrainAgent",
lambda usage: type("B", (), {
"run": lambda self, findings, sheet_index, jurisdiction:
(list(findings), [])})())
class _RecordingVerifier:
name = "verify"
def __init__(self, usage):
pass
def run(self, scope):
verify_sink.append(scope.payload)
return AgentResult(scope_id=scope.scope_id, artifacts=[{
"finding_index": scope.payload["finding_index"],
"status": "confirmed",
"verdicts": [],
}])
monkeypatch.setattr(runner_mod, "EvidenceVerifierAgent",
lambda usage: _RecordingVerifier(usage))
def test_verify_scope_carries_text_excerpt_and_crop(monkeypatch, tmp_path):
"""Evidence text matches the page text layer -> excerpt present and the
full-page image is replaced by a hi-DPI crop."""
sink = []
_patch_pipeline(monkeypatch,
_finding(["S401"], "(5) 2X6 STUD PACK AT BEARING"), sink)
pdf = _make_pdf(tmp_path / "set.pdf")
run_agent_pipeline(pdf, out_dir=str(tmp_path), require_review=False)
assert len(sink) == 1
payload = sink[0]
assert "2X6 STUD PACK" in payload["text_layer_excerpt"]
assert payload["images_b64"], "crop must never drop all images"
assert payload["images_b64"][0] != "QUJD", "expected crop, not full page"
def test_verify_scope_falls_back_to_full_page(monkeypatch, tmp_path):
"""Evidence text not in the text layer -> keep the full-page image."""
sink = []
_patch_pipeline(monkeypatch,
_finding(["S401"], "PENTHOUSE EXHAUST FAN EF-9"), sink)
pdf = _make_pdf(tmp_path / "set.pdf")
run_agent_pipeline(pdf, out_dir=str(tmp_path), require_review=False)
assert len(sink) == 1
assert sink[0]["images_b64"] == ["QUJD"]
def test_coverage_gap_becomes_gap_finding(monkeypatch, tmp_path):
"""Text layer present but zero objects extracted -> failed-scope gap
finding survives into the report."""
sink = []
_patch_pipeline(monkeypatch, _finding(["S401"], PAGE_TEXT), sink)
# Extractor returns a sheet with NO objects despite a real text layer.
monkeypatch.setattr(runner_mod, "SheetExtractorAgent", _stub_agent([
{"sheet_number": "S401", "page_number": 1, "level": "roof",
"discipline": "S", "assertions": []},
]))
pdf = _make_pdf(tmp_path / "set.pdf")
report = run_agent_pipeline(pdf, out_dir=str(tmp_path),
require_review=False)
gaps = [f for f in (report.get("validated_issues") or [])
if f.get("category") == "analysis_gap"]
assert any("extraction gap" in (g.get("description") or "")
for g in gaps)
+86
View File
@@ -0,0 +1,86 @@
"""Grounding-guard rescue tier, text-layer prompt block, and render hygiene."""
from backend import config
from backend.pipeline._stage import render
from backend.pipeline.extractor import (
_is_grounded,
_normalize_sheet,
_text_layer_block,
)
from backend.prompts import EXTRACTOR_USER_INSTRUCTION, VERIFY_USER_INSTRUCTION
PAGE_TEXT = "NOTES: (5) 2X6 STUD PACK AT BEARING. HSS16X4 BEAM. 7'-0\" AFF."
def _parsed(value, source_text):
return {
"sheet": {"sheet_number": "S401"},
"objects": [{
"object_id": "o1",
"object_type": "framing",
"name": "stud pack",
"attributes": {"count": value},
"source_text": source_text,
}],
}
def test_rescue_tier_keeps_and_stamps():
"""Digits absent from source_text but present in the page text layer:
kept, stamped grounding=text_layer (vision quoted imperfectly)."""
sheet = _normalize_sheet(_parsed("(2)", "(2) 2x6 STUD PACK"), 1,
page_text=PAGE_TEXT)
# "(2)" is not grounded by its own source_text alone? it is - use a value
# whose digits differ from the quote to exercise the rescue path.
sheet = _normalize_sheet(_parsed("5", "(2) 2x6 STUD PACK"), 1,
page_text=PAGE_TEXT)
assert len(sheet["assertions"]) == 1
assert sheet["assertions"][0]["grounding"] == "text_layer"
def test_no_rescue_without_page_text():
sheet = _normalize_sheet(_parsed("5", "(2) 2x6 STUD PACK"), 1)
assert sheet["assertions"] == []
def test_still_dropped_when_digits_nowhere():
sheet = _normalize_sheet(_parsed("99", "(2) 2x6 STUD PACK"), 1,
page_text=PAGE_TEXT)
assert sheet["assertions"] == []
def test_is_grounded_backward_compatible():
assert _is_grounded("(5)", "(5) 2x6 STUD PACK") is True
# Digit-run guard is a set check: "(3)" has no support anywhere.
assert _is_grounded("(3)", "(5) 2x6 STUD PACK") is False
assert _is_grounded("(3)", "(5) 2x6 STUD PACK",
page_text="(3) 2x6 STUD PACK") is True
def test_text_layer_block_empty_without_layer():
assert _text_layer_block({"page_number": 1}) == ""
assert _text_layer_block({"page_number": 1, "text_layer": None}) == ""
def test_text_layer_block_appends_and_caps(monkeypatch):
block = _text_layer_block({"page_number": 1, "text_layer": PAGE_TEXT})
assert "TEXT LAYER" in block and "STUD PACK" in block
monkeypatch.setattr(config, "TEXT_LAYER_MAX_CHARS", 50)
block = _text_layer_block({"page_number": 1, "text_layer": "x" * 500})
assert len(block.split(":\n", 1)[1]) == 50
def test_verify_instruction_fully_rendered():
"""render() silently leaves missing keys as literals - both placeholders
must be substituted at the (single) verify render site."""
out = render(VERIFY_USER_INSTRUCTION,
{"finding": "FINDING_JSON", "text_layer": "PAGE_TEXT"})
assert "{finding}" not in out and "{text_layer}" not in out
assert "FINDING_JSON" in out and "PAGE_TEXT" in out
def test_extractor_instruction_fully_substituted():
page = {"page_number": 1, "text_layer": PAGE_TEXT}
out = (EXTRACTOR_USER_INSTRUCTION.replace("{sheet_hint}", "")
+ _text_layer_block(page))
assert "{sheet_hint}" not in out
+111
View File
@@ -0,0 +1,111 @@
"""Text-layer extraction, evidence bbox matching, and crop rendering."""
import os
import pytest
fitz = pytest.importorskip("pymupdf")
from backend import config
from backend.text_layer import (
attach_text_layers,
coverage_gaps,
extract_text_layers,
find_evidence_bbox,
render_crop,
)
EVIDENCE = "(5) 2X6 STUD PACK @ 16 IN O.C."
def _make_pdf(path, pages):
"""pages: list of str ('' = effectively blank page)."""
doc = fitz.open()
for text in pages:
page = doc.new_page(width=612, height=792)
if text:
page.insert_text((72, 72), text, fontsize=11)
doc.save(str(path))
doc.close()
return str(path)
@pytest.fixture
def text_pdf(tmp_path):
return _make_pdf(tmp_path / "set.pdf", [EVIDENCE, ""])
def test_extract_text_layers(text_pdf):
layers = extract_text_layers(text_pdf)
assert set(layers) == {1, 2}
assert layers[1]["has_text_layer"] is True
assert "2X6 STUD PACK" in layers[1]["text"]
assert layers[1]["words"], "expected word-level bboxes"
assert all("bbox" in w and len(w["bbox"]) == 4 for w in layers[1]["words"])
def test_blank_page_below_min_chars(text_pdf):
layers = extract_text_layers(text_pdf)
assert layers[2]["has_text_layer"] is False
def test_disabled_returns_empty(text_pdf, monkeypatch):
monkeypatch.setattr(config, "TEXT_LAYER_ENABLED", False)
assert extract_text_layers(text_pdf) == {}
def test_attach_text_layers(text_pdf, tmp_path):
pages = [{"page_number": 1}, {"page_number": 2}]
words = attach_text_layers(text_pdf, pages,
text_dir=str(tmp_path / "text"))
assert pages[0]["text_layer"] and "STUD PACK" in pages[0]["text_layer"]
assert pages[1]["text_layer"] is None
assert words[1] and not words[2]
assert os.path.isfile(tmp_path / "text" / "page-001.txt")
assert not os.path.exists(tmp_path / "text" / "page-002.txt")
def test_find_evidence_bbox_exact(text_pdf):
words = extract_text_layers(text_pdf)[1]["words"]
bbox = find_evidence_bbox(words, EVIDENCE)
assert bbox is not None
assert bbox[2] > bbox[0] and bbox[3] > bbox[1]
def test_find_evidence_bbox_fuzzy(text_pdf):
# Vision quotes imperfectly: wrong count token, rest exact.
words = extract_text_layers(text_pdf)[1]["words"]
bbox = find_evidence_bbox(words, "(2) 2X6 STUD PACK @ 16 IN O.C.")
assert bbox is not None
def test_find_evidence_bbox_miss(text_pdf):
words = extract_text_layers(text_pdf)[1]["words"]
assert find_evidence_bbox(words, "PENTHOUSE EXHAUST FAN EF-9") is None
assert find_evidence_bbox([], EVIDENCE) is None
assert find_evidence_bbox(words, "") is None
def test_render_crop(text_pdf):
words = extract_text_layers(text_pdf)[1]["words"]
bbox = find_evidence_bbox(words, EVIDENCE)
crop = render_crop(text_pdf, 1, bbox)
assert crop is not None
# Decodes as an image of plausible size (margin around the text line).
doc = fitz.open(stream=crop, filetype="jpeg")
pix = doc[0].get_pixmap()
assert pix.width > 100 and pix.height > 20
doc.close()
def test_render_crop_bad_page(text_pdf):
assert render_crop(text_pdf, 99, (0, 0, 10, 10)) is None
def test_coverage_gaps():
pages = [{"page_number": 1, "text_layer": "some real text"},
{"page_number": 2, "text_layer": "more text"},
{"page_number": 3, "text_layer": None}]
sheets = [{"page_number": 1, "assertions": [{"id": "a"}]},
{"page_number": 2, "assertions": []}]
assert coverage_gaps(pages, sheets) == [2]