Compare commits

..
25 Commits
Author SHA1 Message Date
woogi 570300324f feat: text-layer grounding (extractor authority, guard rescue tier, verifier oracle + hi-DPI crops)
Docker Release / build-and-push (push) Successful in 1m25s
Docker Release / release (push) Skipped
- backend/text_layer.py: PyMuPDF text-layer extraction, fuzzy evidence
  bbox matching, 300-DPI crop rendering, coverage-gap signal
- extractor (classic + agent): TEXT LAYER block appended at call sites;
  grounding guard gains text-layer rescue tier (grounding=text_layer stamp)
- verifier: {text_layer} oracle excerpt + evidence-located hi-DPI crops
  replacing full-page images (fallback preserved, I2 guard intact)
- coverage gaps: text-bearing pages with zero extraction -> failed-scope
  gap findings (agent) / log-only (classic)
- config knobs: TEXT_LAYER_ENABLED/MIN_CHARS/MAX_CHARS, VERIFY_TEXT_MAX_CHARS,
  VERIFY_HI_DPI_CROPS, VERIFY_CROP_DPI, VERIFY_CROP_MARGIN_PTS
- tests: 22 new (text_layer unit, grounding/render, runner-level flow)
Spec: docs/superpowers/specs/2026-08-12-text-layer-grounding-design.md
2026-08-12 14:27:00 -05:00
woogi 349b357e5c docs: text-layer grounding design spec 2026-08-12 13:00:38 -05:00
woogi b174b531cd fix: wave 5b review blockers - suppressed memory key, finalizer merge, zero-image guard
Docker Release / build-and-push (push) Successful in 59s
Docker Release / release (push) Skipped
2026-08-10 12:21:41 -05:00
woogi 3df359500c feat: wave 5b evidence verification - vision fact-check of cited sheet text 2026-08-10 11:30:28 -05:00
woogi 82952df307 fix: reasoning budget for conflict critic (wave-4 max_tokens truncation) 2026-08-10 11:29:03 -05:00
woogi c8af430143 fix: xref member-mark regex covers single-letter marks; recheck sheet diversity after cap 2026-08-10 11:23:06 -05:00
woogi f21eb5d912 feat: config knobs and prompts for evidence verification wave 2026-08-10 11:14:27 -05:00
woogi 4a3f33a245 feat: cross-level xref link scopes via detail_reference and member tag 2026-08-10 11:10:43 -05:00
woogi 15297038a2 fix: annotate clusters and substitute {disputes} in classic pipeline path 2026-08-10 11:06:15 -05:00
woogi d431a026ce feat: treat disputed extracted values as unverified in constructability and critic prompts 2026-08-10 10:56:53 -05:00
woogi 8631a26006 feat: surface disputed extracted values to critic and specialist prompts 2026-08-10 10:47:34 -05:00
woogi df4d15fd0c feat: deterministic disputed-value detection for cluster assertions 2026-08-10 10:40:10 -05:00
woogi 0ea0b0e897 Add non-technical pipeline overview: infographic + flow explainer doc
Docker Release / build-and-push (push) Successful in 59s
Docker Release / release (push) Skipped
2026-08-10 08:21:57 -05:00
woogi 228e8bd031 Sync .env.example with new extract token defaults
Docker Release / build-and-push (push) Successful in 57s
Docker Release / release (push) Skipped
2026-08-09 08:21:31 -05:00
woogi 3d7fce7bf9 Kill extract-wave truncation: 65k ceiling, hard thinking budget, reasoning-token telemetry
Docker Release / build-and-push (push) Successful in 57s
Docker Release / release (push) Skipped
Job 98194fa8d215 showed every extract call hitting the 32k cap with only
~20k chars visible despite reasoning effort=low - Gemini 2.5 Pro still
burned ~25k thinking tokens per sheet.

- EXTRACT_MAX_TOKENS default 32768 -> 65536 (model output ceiling)
- new EXTRACT_REASONING_MAX_TOKENS (default 2048): OpenRouter reasoning
  max_tokens / Gemini thinking_budget; takes precedence over effort
- log per-call reasoning token counts (usage.completion_tokens_details)
  and include thinking count in the finish_reason=length marker
2026-08-09 08:17:41 -05:00
woogi 76e0a52658 Fix sheet-extraction page loss: bare-list wrap, compact retry, reasoning cap
Docker Release / build-and-push (push) Successful in 1m1s
Docker Release / release (push) Skipped
- SheetExtractorAgent accepts top-level array responses as the objects
  array instead of discarding them (recovered the failure mode behind
  16/38 failed sheets on job 475a6f184dd1)
- Second-chance compact retry per page before declaring extraction failed
- call_json: reasoning_effort param (cloud-only extra_body), finish_reason
  capture + explicit max_tokens log line, finish_reason in raw dumps,
  cache key covers reasoning_effort
- EXTRACT_MAX_TOKENS default 16384 -> 32768 (Gemini thinking tokens count
  against the cap); new EXTRACT_REASONING_EFFORT=low default for extract
- tests: 5 new fallback-ladder tests
2026-08-07 07:35:53 -05:00
woogi 5c1fccfb35 Guard syncPipelineOptions when hybrid radio is removed.
Docker Release / build-and-push (push) Successful in 55s
Docker Release / release (push) Skipped
2026-08-05 16:07:34 -05:00
woogi 32544bc2af Add /models fetch timeout and health-derived default fallback for model dropdowns.
Docker Release / build-and-push (push) Successful in 57s
Docker Release / release (push) Skipped
2026-08-05 15:59:50 -05:00
woogi 4b3b62b3fa Add cache-busting meta tags and default build tag text.
Docker Release / build-and-push (push) Successful in 56s
Docker Release / release (push) Skipped
2026-08-05 15:50:35 -05:00
woogi 5305325d81 Fix model dropdown loading, cache models for 24h, and style build tag.
Docker Release / release (push) Skipped
Docker Release / build-and-push (push) Successful in 1m1s
2026-08-05 15:42:21 -05:00
woogi 6f10062b93 Agent-mode UI: remove hybrid compute option, add clickable sheet links in review queue.
Docker Release / build-and-push (push) Successful in 1m1s
Docker Release / release (push) Skipped
2026-08-05 15:28:09 -05:00
woogi 7488cf68c5 Verbose per-call LLM logging, raw request/response dumps, and end-of-log cost summary.
Docker Release / build-and-push (push) Successful in 1m13s
Docker Release / release (push) Skipped
- [LLM] line per call: stage, model, prompt size, output size, cost, parsed item counts
- LLM_RAW_DUMP: full prompt/response JSON per call under outputs/<job>/llm_raw/
- Cost block at tail of job.log (per-stage, per-model, cached vs live)
- Agent mode: reset llm cost counters per job; review finalization now teed into job.log + dumps
2026-08-05 15:17:17 -05:00
woogi f7e1b6bb7c Merge main: dual model dropdowns + richer job logs, adapted for agent-mode.
Docker Release / build-and-push (push) Successful in 1m0s
Docker Release / release (push) Skipped
- llm.py: set_model_overrides(vision, text) replaces the single job override;
  UI picks still beat per-call agent model args, but never name the hybrid
  local model (avoids main's hybrid footgun); local->cloud fallback uses the
  text pick.
- jobs.py: timestamped line-split tee (job_log.py), in-memory log + log_tail
  polls, full log on terminal states (done/error/needs_review/finalization_error),
  log-only disk recovery, error email links to the run log, and failed runs now
  append the full traceback to job.log. Keeps pipeline_mode, job.json, and the
  review gate.
- models.py: vision/text split via architecture modalities, pricing kept;
  /models returns {vision, text, defaults}; /check takes vision_model/text_model
  (replacing model); /health adds text_model. models_catalog.py dropped.
- UI: two priced dropdowns (OpenRouter compute only) + live run-log panel.
- Tests updated for dual overrides and the /models shape; new coverage for
  traceback capture and local-model immunity.
2026-08-02 09:55:09 -05:00
woogiandCursor bf508bfdf6 Expand session notes with API, job log, and model selection details.
Docker Release / build-and-push (push) Successful in 57s
Docker Release / release (push) Skipped
Keep NOTES.md current for new sessions after the job-log and dual-model UI work.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-31 20:21:35 -05:00
woogiandCursor a6b0c8fdfa Add per-job run logs and separate vision/text model selection.
Docker Release / build-and-push (push) Successful in 1m27s
Docker Release / release (push) Skipped
Capture pipeline stdout into job.log + API/UI so failed runs can be reviewed, and let users pick OpenRouter vision vs text models independently.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-31 14:56:48 -05:00
44 changed files with 4027 additions and 194 deletions
@@ -0,0 +1,936 @@
# Evidence Verification + Cross-Sheet Correlation Implementation Plan
> **For Hermes:** Use subagent-driven-development skill to implement this plan task-by-task.
**Goal:** Stop vision-extraction misreads (e.g. "(2) 2x6 STUD PACK" vs the actual "(5) 2x6") from becoming confident downstream findings, and correlate the same physical element across sheets (S101/S205/S401) so no stage reasons from one sheet's text in isolation.
**Architecture:** Three independently shippable phases on the agent pipeline (`backend/agents/runner.py`):
1. **Disputed-value detection** — deterministic post-link pass that flags contradictory extracted values inside a cluster and surfaces them to the critic/specialist prompts.
2. **Cross-sheet xref linking** — linker gains detail-reference/tag buckets that join assertions across levels (today `(level, family)` bucketing splits S101/S205/S401 apart).
3. **Evidence verification wave (5b)** — a bounded vision fact-check agent re-reads the cited sheet images for high-severity / disputed findings before the Brain merge, annotates or suppresses findings built on phantom text.
**Tech Stack:** Python 3.14, pytest (`tests/`), existing `call_json` LLM wrapper (supports `images_b64`, `reasoning_effort`, `reasoning_max_tokens`).
---
## Current context / root cause (from job 959e16407573)
Finding `validated_issues[3]` ("FRONT PERSPECTIVE detail on Sheet S401", HSS16x4 on
"(2) 2x6 STUD PACK", severity critical) is a **false positive built on a wave-1 vision
misread**. The sheet actually shows a (5) 2x6 stud pack (matching S205/S101). Chain of failure:
1. Wave 1 (`SheetExtractorAgent`) froze the misread into text. From then on it is "ground truth".
2. Wave 3 linker (`backend/agents/linker.py:33` `build_link_scopes`) buckets by
`(level, family)`. S101 (foundation), S205 (details), S401 (sections) get different
`level` values, so assertions about the same front-wall header never share a link scope
or cluster. No cross-sheet corroboration happened.
3. Wave 5 `ConstructabilityAgent` (`backend/agents/construct_agent.py:53`) calls
`call_json` **with no images** — in this job 120/120 constructability calls were `+0img`.
It reasoned arithmetically from the misread text ("2 x 1.5in = 3in < 4in -> unbuildable").
It even held the "(5) 2x6 STUD PACK" assertion in the same scope but labeled it
"Ambiguous column size specification" instead of arbitrating.
4. Nothing between wave 5 and the report ever looks at a sheet image again. Only the wave-4
conflict critic receives images, and only for its own cluster's pages.
Also confirmed in this log (separate known bug, fixed in Task 7 while we're here): wave-4
conflict critic truncates on Gemini thinking tokens because `conflict_critic.py:59` passes
`max_tokens=config.REASON_MAX_TOKENS` (4096) with no reasoning budget — 13/121 calls hit
`finish_reason=length`.
## Assumptions
- Assertions carry `id`, `attribute`, `value`, `source_text`, `location_key`
(`room`/`grid`/`detail_reference`/`tag`/`level`) — see `linker._payload` and
`_serialize.slim_assertion`.
- Extractor assertions already carry a `confidence` field (per test fixtures).
- `validate_issue` in `backend/pipeline/_stage.py` guarantees each finding an `issue_id`.
- Test convention: `unittest.mock.patch("backend.agents.<module>.call_json", ...)`
see `tests/agents/test_sheet_extractor_fallback.py`. Run tests with
`.venv/bin/python -m pytest tests/ -x -q`.
- `AgentResult.error` defaults to `""` (not None) in assertions.
---
## Phase 1 — Disputed-value detection + prompt hardening
### Task 1: `find_disputes` pure function (TDD)
**Objective:** Detect "same attribute, different values" inside one cluster's assertions.
**Files:**
- Create: `backend/agents/disputes.py`
- Test: `tests/agents/test_disputes.py`
**Step 1: Write failing test**
```python
# tests/agents/test_disputes.py
from backend.agents.disputes import annotate_clusters, find_disputes
def _a(id_, attribute, value):
return {"id": id_, "attribute": attribute, "value": value,
"source_text": value}
def test_find_disputes_flags_same_attribute_different_values():
assertions = [
_a("a1", "stud_pack_size", "(2) 2x6 STUD PACK"),
_a("a2", "stud_pack_size", "(5) 2x6 STUD PACK"),
_a("a3", "beam_size", "HSS16X4X5/8"),
]
disputes = find_disputes(assertions)
assert len(disputes) == 1
assert disputes[0]["attribute"] == "stud_pack_size"
assert disputes[0]["values"] == ["(2) 2x6 STUD PACK", "(5) 2x6 STUD PACK"]
assert disputes[0]["assertion_ids"] == ["a1", "a2"]
def test_find_disputes_ignores_agreeing_values_and_blanks():
assertions = [
_a("a1", "beam_size", "HSS16X4X5/8"),
_a("a2", "beam_size", " hss16x4x5/8 "), # same after normalize
_a("a3", "", "orphan"), # no attribute -> skipped
_a("a4", "beam_size", ""), # no value -> skipped
]
assert find_disputes(assertions) == []
def test_annotate_clusters_writes_disputed_attributes():
clusters = [
{"key": "c1", "assertions": [
_a("a1", "stud_pack_size", "(2) 2x6"),
_a("a2", "stud_pack_size", "(5) 2x6"),
]},
{"key": "c2", "assertions": [_a("a3", "x", "1"), _a("a4", "x", "1")]},
]
assert annotate_clusters(clusters) == 1
assert clusters[0]["disputed_attributes"][0]["attribute"] == "stud_pack_size"
assert "disputed_attributes" not in clusters[1]
```
**Step 2: Run test to verify failure**
Run: `.venv/bin/python -m pytest tests/agents/test_disputes.py -v`
Expected: FAIL — `ModuleNotFoundError: backend.agents.disputes`
**Step 3: Implement**
```python
# backend/agents/disputes.py
"""Deterministic detection of contradictory extracted values within a cluster.
Extraction is a vision pass: quantities and sizes can be misread ("(2) 2x6" vs
"(5) 2x6"). Cluster members are supposed to describe the same real-world
element, so two members asserting different values for the same attribute are
a probable misread. Flag these so downstream text-only stages treat the value
as unverified instead of reasoning from one reading.
"""
import re
from typing import Dict, List
def _norm(value) -> str:
return re.sub(r"\s+", " ", str(value or "").strip().lower())
def find_disputes(assertions: List[Dict]) -> List[Dict]:
"""Same attribute with >= 2 distinct normalized values = disputed."""
groups: Dict[str, Dict[str, set]] = {}
for assertion in assertions:
attribute = _norm(assertion.get("attribute"))
value = _norm(assertion.get("value"))
if not attribute or not value:
continue
groups.setdefault(attribute, {}).setdefault(value, set()).add(
assertion.get("id")
)
disputes = []
for attribute, values in sorted(groups.items()):
if len(values) < 2:
continue
disputes.append({
"attribute": attribute,
"values": sorted(values),
"assertion_ids": sorted(
aid for ids in values.values() for aid in ids if aid
),
})
return disputes
def annotate_clusters(clusters: List[Dict]) -> int:
"""Attach disputed_attributes to each cluster that has any. Returns count."""
annotated = 0
for cluster in clusters:
disputes = find_disputes(cluster.get("assertions") or [])
if disputes:
cluster["disputed_attributes"] = disputes
annotated += 1
return annotated
```
**Step 4: Run test to verify pass**
Run: `.venv/bin/python -m pytest tests/agents/test_disputes.py -v`
Expected: 3 passed
**Step 5: Commit**
```bash
git add backend/agents/disputes.py tests/agents/test_disputes.py
git commit -m "feat: deterministic disputed-value detection for cluster assertions"
```
---
### Task 2: Wire `annotate_clusters` into the runner + serialization
**Objective:** Disputes must be visible to the wave-4 critic and wave-5 constructability prompts.
**Files:**
- Modify: `backend/agents/runner.py` (after `memory.replace("clusters", clusters)`, ~line 116)
- Modify: `backend/pipeline/_serialize.py` (`slim_clusters`, line 46)
- Test: `tests/agents/test_disputes.py` (append)
**Step 1: Write failing test**
```python
def test_slim_clusters_preserves_disputed_attributes():
from backend.pipeline._serialize import slim_clusters
cluster = {"key": "c1", "assertions": [],
"disputed_attributes": [{"attribute": "a", "values": ["1", "2"],
"assertion_ids": ["x", "y"]}]}
slim = slim_clusters([cluster])[0]
assert slim["disputed_attributes"][0]["values"] == ["1", "2"]
```
**Step 2: Run test to verify failure**
Run: `.venv/bin/python -m pytest tests/agents/test_disputes.py::test_slim_clusters_preserves_disputed_attributes -v`
Expected: FAIL — `KeyError: 'disputed_attributes'`
**Step 3: Implement**
In `backend/pipeline/_serialize.py` `slim_clusters`, add the key:
```python
def slim_clusters(clusters: List[Dict]) -> List[Dict]:
return [
{
"key": c.get("key"),
"location": c.get("location"),
"disciplines": c.get("disciplines"),
"kind": c.get("kind"),
**({"disputed_attributes": c["disputed_attributes"]}
if c.get("disputed_attributes") else {}),
"assertions": [slim_assertion(a) for a in c.get("assertions", [])],
}
for c in clusters
]
```
In `backend/agents/runner.py`, right after `clusters = [...]` / `object_graph = build_object_graph(clusters)` (before `memory.replace("clusters", clusters)`):
```python
from backend.agents.disputes import annotate_clusters
...
object_graph = build_object_graph(clusters)
disputed_count = annotate_clusters(clusters)
if disputed_count:
orchestrator.log(
f"[Link] {disputed_count} clusters carry disputed extracted values"
)
```
(Check `Orchestrator` for the actual log method name — `orchestrator.stage(...)` exists;
if no `.log`, use the module's existing logging/print convention. Adjust to match.)
**Step 4: Run tests**
Run: `.venv/bin/python -m pytest tests/agents/ -v`
Expected: all pass (including existing `test_runner_review_gate.py`)
**Step 5: Commit**
```bash
git add backend/agents/runner.py backend/pipeline/_serialize.py tests/agents/test_disputes.py
git commit -m "feat: surface disputed extracted values to critic and specialist prompts"
```
---
### Task 3: Prompt hardening — extracted text is fallible
**Objective:** Tell text-only specialists how to handle disputed/unverified values so they stop asserting buildability conclusions from a single (possibly misread) number.
**Files:**
- Modify: `backend/prompts.py` `CONSTRUCTABILITY_SYSTEM_PROMPT` (line 533) and `CONSTRUCTABILITY_USER_INSTRUCTION` (line 551)
**Step 1: Edit prompts**
Append to `CONSTRUCTABILITY_SYSTEM_PROMPT` Rules list (after line 547, before "Use plain ASCII"):
```
- Assertions are machine-extracted from sheet images and may contain misread values,
especially quantities and member sizes (e.g. "(2) 2x6" vs "(5) 2x6").
- When the cluster lists disputed_attributes, or two evidence items disagree on a
numeric value, do NOT assert a buildability conclusion from one reading. Report the
ambiguity itself (category "detail_gap", confidence "low") and state that the value
needs verification against the sheet.
```
Append to `CONSTRUCTABILITY_USER_INSTRUCTION` after the `Cross-discipline conflicts already found: {conflicts}` line:
```
Disputed extracted values in this cluster (possible vision misreads - treat as unverified): {disputes}
```
**Step 2: Wire the `{disputes}` placeholder in `construct_agent.py`**
In `backend/agents/construct_agent.py` `run()`, extend the `substitutions` dict:
```python
substitutions = {
"assertions": dumps(cluster["assertions"]),
"clusters": dumps(slim_clusters([cluster])),
"conflicts": dumps(scope.payload.get("conflicts") or []),
"disputes": dumps(cluster.get("disputed_attributes") or []),
}
```
**Step 3: Run full test suite (prompt edits can break runner tests that snapshot prompts)**
Run: `.venv/bin/python -m pytest tests/ -q`
Expected: all pass
**Step 4: Commit**
```bash
git add backend/prompts.py backend/agents/construct_agent.py
git commit -m "feat: constructability prompt treats disputed extracted values as unverified"
```
---
## Phase 2 — Cross-sheet xref linking
### Task 4: detail-reference / tag xref buckets in the linker (TDD)
**Objective:** Assertions sharing a `detail_reference` or a member `tag` get linked across levels, so S101/S205/S401 details of the same physical element land in one scope.
**Files:**
- Modify: `backend/agents/linker.py` (`build_link_scopes`, line 33)
- Test: `tests/agents/test_linker_xref.py`
**Step 1: Write failing test**
```python
# tests/agents/test_linker_xref.py
from backend.agents.base import AgentScope
from backend.agents.linker import build_link_scopes
def _sheet(number, page, level, assertions):
return {"sheet_number": number, "page_number": page,
"discipline": "Structural", "level": level,
"assertions": assertions}
def _assertion(id_, ref=None, tag=None, level=None):
return {"id": id_, "attribute": "stud_pack_size", "value": "(5) 2x6",
"source_text": "(5) 2x6 STUD PACK",
"location_key": {"detail_reference": ref, "tag": tag,
"level": level}}
def test_xref_scope_joins_same_detail_reference_across_levels():
sheets = [
_sheet("S101", 10, "foundation", [_assertion("a1", ref="A/S205")]),
_sheet("S205", 20, "roof", [_assertion("a2", ref="A/S205")]),
_sheet("S401", 30, "roof", [_assertion("a3", ref="A/S205")]),
]
scopes = build_link_scopes(sheets)
xref = [s for s in scopes if s.scope_id.startswith("xref:")]
assert xref, "expected a cross-level detail-reference scope"
ids = {a["id"] for s in xref for a in s.payload["assertions"]}
assert ids == {"a1", "a2", "a3"}
def test_xref_scope_requires_two_distinct_sheets():
sheets = [
_sheet("S401", 30, "roof", [_assertion("a1", ref="A/S205"),
_assertion("a2", ref="A/S205")]),
]
scopes = build_link_scopes(sheets)
assert not [s for s in scopes if s.scope_id.startswith("xref:")]
def test_xref_scope_joins_shared_member_tag():
sheets = [
_sheet("S102", 5, "roof", [_assertion("a1", tag="HSS16X4X5/8")]),
_sheet("S401", 30, "unknown", [_assertion("a2", tag="HSS16X4X5/8")]),
]
scopes = build_link_scopes(sheets)
xref = [s for s in scopes if s.scope_id.startswith("xref:")]
assert xref
```
**Step 2: Run test to verify failure**
Run: `.venv/bin/python -m pytest tests/agents/test_linker_xref.py -v`
Expected: FAIL — no `xref:` scopes produced
**Step 3: Implement**
Rewrite `build_link_scopes` in `backend/agents/linker.py` (keep the existing
`(level, family)` bucketing, add the xref pass):
```python
def _xref_keys(assertion: Dict) -> List[str]:
"""Cross-level join keys: detail references and member tags."""
location = assertion.get("location_key") or {}
keys = []
ref = re.sub(r"\s+", "", str(location.get("detail_reference") or "")).upper()
if ref:
keys.append(f"detail:{ref}")
tag = re.sub(r"\s+", "", str(location.get("tag") or "")).upper()
if re.match(r"^[A-Z]{2,}\d", tag): # member marks: HSS16X4X5/8, W12X26, ...
keys.append(f"tag:{tag}")
return keys
def build_link_scopes(sheets: List[Dict]) -> List[AgentScope]:
"""Partition facts by level and object/tag family, then enforce a hard cap.
A second pass joins assertions that share a detail_reference or member tag
ACROSS levels, so plan/detail/section sheets describing the same physical
element are linked together even though their levels differ.
"""
buckets: Dict[Tuple[str, str], List[Dict]] = defaultdict(list)
xref: Dict[str, List[Dict]] = defaultdict(list)
for sheet in sheets:
for assertion in sheet.get("assertions", []):
enriched = {
**assertion,
"discipline": sheet.get("discipline") or "Unknown",
"sheet_number": sheet.get("sheet_number"),
"page_number": sheet.get("page_number"),
}
level = str((assertion.get("location_key") or {}).get("level")
or sheet.get("level") or "unknown").lower()
buckets[(level, _family(assertion))].append(enriched)
for key in _xref_keys(assertion):
xref[key].append(enriched)
scopes: List[AgentScope] = []
cap = max(2, config.AGENT_LINK_MAX_ASSERTIONS)
for (level, family), assertions in sorted(buckets.items()):
for offset in range(0, len(assertions), cap):
chunk = assertions[offset:offset + cap]
if len(chunk) < 2:
continue
scopes.append(AgentScope(
scope_id=f"{level}:{family}:{offset // cap + 1}",
payload={"assertions": chunk, "level": level, "family": family},
))
for key, assertions in sorted(xref.items()):
sheets_present = {a.get("sheet_number") for a in assertions}
if len(assertions) < 2 or len(sheets_present) < 2:
continue
scopes.append(AgentScope(
scope_id=f"xref:{key}",
payload={"assertions": assertions[:cap],
"level": "xref", "family": key},
))
return scopes
```
**Step 4: Run tests**
Run: `.venv/bin/python -m pytest tests/agents/test_linker_xref.py tests/agents/ -v`
Expected: all pass (watch existing runner tests for scope-count coupling)
**Step 5: Commit**
```bash
git add backend/agents/linker.py tests/agents/test_linker_xref.py
git commit -m "feat: cross-level xref link scopes via detail_reference and member tag"
```
---
## Phase 3 — Evidence verification wave (5b)
### Task 5: Config knobs + verify prompts
**Objective:** Add the tuning surface and prompts for the vision fact-check agent.
**Files:**
- Modify: `backend/config.py` (near line 44, with the other AGENT_* knobs)
- Modify: `backend/prompts.py` (append near the CONFLICT prompts, ~line 460)
- Modify: `backend/.env.example`
**Step 1: Add config knobs to `backend/config.py`**
```python
AGENT_VERIFY_MODEL = os.getenv("AGENT_VERIFY_MODEL", "") or MODEL
AGENT_VERIFY_CONCURRENCY = int(os.getenv("AGENT_VERIFY_CONCURRENCY", "4"))
AGENT_VERIFY_MAX_CHECKS = int(os.getenv("AGENT_VERIFY_MAX_CHECKS", "20"))
AGENT_VERIFY_SEVERITIES = {
s.strip().lower()
for s in os.getenv("AGENT_VERIFY_SEVERITIES", "critical,high").split(",")
if s.strip()
}
AGENT_VERIFY_REASONING_EFFORT = os.getenv("AGENT_VERIFY_REASONING_EFFORT", "low").strip()
VERIFY_MAX_TOKENS = int(os.getenv("VERIFY_MAX_TOKENS", "8192"))
```
Append to `backend/.env.example`:
```
# Wave 5b evidence verification (vision fact-check of cited sheet text)
AGENT_VERIFY_MAX_CHECKS=20
AGENT_VERIFY_SEVERITIES=critical,high
AGENT_VERIFY_REASONING_EFFORT=low
VERIFY_MAX_TOKENS=8192
```
**Step 2: Add prompts to `backend/prompts.py`**
```python
VERIFY_SYSTEM_PROMPT = """You are a meticulous construction document checker verifying machine-extracted evidence against the actual drawing sheet images.
For each evidence item you are given the sheet it was extracted from and the verbatim text the extractor claims appears there.
Judge each item against the images:
- confirmed: the text (or an obvious equivalent) appears on the cited sheet and means what the finding claims.
- corrected: the sheet shows a DIFFERENT value than the extracted text. Give the actual verbatim text.
- not_found: nothing like the extracted text appears on the cited sheet.
Be strict about numbers, quantities, and member sizes: "(2) 2x6" and "(5) 2x6" are different values. HSS16x4 and HSS16x16 are different values.
Use plain ASCII only.
Respond only with valid JSON."""
VERIFY_USER_INSTRUCTION = """Verify this finding's evidence against the attached sheet images.
Respond ONLY with a valid JSON object - no markdown fences, no explanation:
{ "verdicts": [ { "sheet": "string", "source_text": "the evidence text judged", "verdict": "confirmed | corrected | not_found", "actual_text": "verbatim sheet text when corrected, else null", "notes": "string or null" } ] }
Finding: {finding}"""
```
**Step 3: Sanity check**
Run: `.venv/bin/python -c "from backend import config, prompts; print(config.AGENT_VERIFY_MAX_CHECKS, config.AGENT_VERIFY_SEVERITIES); print(prompts.VERIFY_SYSTEM_PROMPT[:40])"`
Expected: `20 {'critical', 'high'}` and prompt text
**Step 4: Commit**
```bash
git add backend/config.py backend/prompts.py backend/.env.example
git commit -m "feat: config knobs and prompts for evidence verification wave"
```
---
### Task 6: `EvidenceVerifierAgent` + runner wave 5b (TDD)
**Objective:** Re-read cited sheet images for selected findings; annotate verified findings, suppress refuted ones before the Brain merge.
**Files:**
- Create: `backend/agents/verifier.py`
- Modify: `backend/agents/runner.py` (new wave between wave 5 and wave 6, ~line 167)
- Modify: `backend/agents/construct_agent.py` line 67 (stamp `cluster_key` for dispute-based selection)
- Test: `tests/agents/test_verifier.py`
**Step 1: Write failing test**
```python
# tests/agents/test_verifier.py
from unittest.mock import patch
from backend.agents.base import AgentScope, AgentUsage
from backend.agents.verifier import (
EvidenceVerifierAgent, apply_verdicts, select_findings,
)
def _finding(sev="critical", issue_id="i1", sheets=("S401",), cluster_key=None):
f = {"issue_id": issue_id, "severity": sev, "confidence": "high",
"source_stage": "constructability", "sheets": list(sheets),
"description": "HSS16x4 on (2) 2x6 STUD PACK is unbuildable",
"evidence": [{"sheet": "S401", "source_text": "(2) 2x6 STUD PACK",
"asserted_value": "3-inch width"}]}
if cluster_key:
f["cluster_key"] = cluster_key
return f
def test_select_findings_by_severity_and_dispute():
findings = [_finding("critical"), _finding("low", "i2"),
_finding("medium", "i3", cluster_key="c9")]
clusters = [{"key": "c9", "disputed_attributes": [{"attribute": "a"}]}]
selected = select_findings(findings, clusters, max_checks=20,
severities={"critical", "high"})
assert [f["issue_id"] for f in selected] == ["i1", "i3"]
def test_select_findings_respects_cap():
findings = [_finding("critical", f"i{n}") for n in range(30)]
selected = select_findings(findings, [], max_checks=5,
severities={"critical"})
assert len(selected) == 5
def test_run_attaches_verdicts_and_marks_refuted():
agent = EvidenceVerifierAgent(usage=AgentUsage())
scope = AgentScope(scope_id="verify:0", payload={
"finding_index": 0,
"finding": _finding(),
"images_b64": ["QUJD"],
})
verdicts = {"verdicts": [
{"sheet": "S401", "source_text": "(2) 2x6 STUD PACK",
"verdict": "corrected", "actual_text": "(5) 2x6 STUD PACK",
"notes": "callout reads (5)"},
]}
with patch("backend.agents.verifier.call_json", return_value=verdicts):
result = agent.run(scope)
assert not result.error
artifact = result.artifacts[0]
assert artifact["finding_index"] == 0
assert artifact["status"] == "refuted" # no evidence confirmed
assert artifact["verdicts"][0]["actual_text"] == "(5) 2x6 STUD PACK"
def test_apply_verdicts_annotates_and_suppresses():
findings = [_finding("critical", "i1"), _finding("high", "i2")]
from backend.agents.base import AgentResult
results = [AgentResult(scope_id="verify:0", artifacts=[
{"finding_index": 0, "status": "refuted", "verdicts": []},
{"finding_index": 1, "status": "confirmed", "verdicts": []},
])]
suppressed = apply_verdicts(findings, results)
assert suppressed == [findings[0]]
assert findings[0]["verification"]["status"] == "refuted"
assert findings[1]["verification"]["status"] == "confirmed"
```
**Step 2: Run test to verify failure**
Run: `.venv/bin/python -m pytest tests/agents/test_verifier.py -v`
Expected: FAIL — `ModuleNotFoundError: backend.agents.verifier`
**Step 3: Implement `backend/agents/verifier.py`**
```python
"""Wave 5b: vision fact-check of extracted evidence against cited sheet images.
Downstream specialists are text-only; a wave-1 vision misread ("(2) 2x6" vs
"(5) 2x6") otherwise becomes immutable ground truth. For high-severity or
dispute-linked findings, re-read the cited sheets and adjudicate each evidence
item: confirmed / corrected / not_found. Findings whose evidence is entirely
unconfirmed are suppressed before the Brain merge.
"""
from typing import Dict, List, Optional, Set
from backend import config
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
from backend.llm import call_json
from backend.pipeline._serialize import dumps
from backend.pipeline._stage import collect_list, render
from backend.prompts import VERIFY_SYSTEM_PROMPT, VERIFY_USER_INSTRUCTION
_SEVERITY_RANK = {"critical": 0, "high": 1, "medium": 2, "low": 3}
_VERDICTS = ("confirmed", "corrected", "not_found")
def select_findings(
findings: List[Dict],
clusters: List[Dict],
max_checks: int,
severities: Set[str],
) -> List[Dict]:
"""Severity-gated selection plus any finding tied to a disputed cluster."""
disputed_keys = {
cluster.get("key") for cluster in clusters
if cluster.get("disputed_attributes")
}
selected = [
finding for finding in findings
if str(finding.get("severity") or "").lower() in severities
or finding.get("cluster_key") in disputed_keys
]
selected.sort(key=lambda f: _SEVERITY_RANK.get(
str(f.get("severity") or "").lower(), 9))
return selected[:max_checks]
def _valid_verdict(item: Dict) -> Optional[Dict]:
if not isinstance(item, dict):
return None
verdict = str(item.get("verdict") or "").lower()
if verdict not in _VERDICTS:
return None
return {
"sheet": item.get("sheet") or "",
"source_text": item.get("source_text") or "",
"verdict": verdict,
"actual_text": item.get("actual_text"),
"notes": item.get("notes"),
}
def _status(verdicts: List[Dict]) -> str:
if not verdicts:
return "unverified"
confirmed = sum(1 for v in verdicts if v["verdict"] == "confirmed")
if confirmed == len(verdicts):
return "confirmed"
if confirmed == 0:
return "refuted"
return "mixed"
class EvidenceVerifierAgent:
name = "verify"
def __init__(self, usage: AgentUsage) -> None:
self.usage = usage
def run(self, scope: AgentScope) -> AgentResult:
try:
finding = scope.payload["finding"]
instruction = render(
VERIFY_USER_INSTRUCTION, {"finding": dumps(finding)}
)
parsed = call_json(
system_prompt=VERIFY_SYSTEM_PROMPT,
user_text=instruction,
images_b64=scope.payload.get("images_b64") or [],
max_tokens=config.VERIFY_MAX_TOKENS,
model=config.AGENT_VERIFY_MODEL,
reasoning_effort=config.AGENT_VERIFY_REASONING_EFFORT or None,
usage_tracker=self.usage,
usage_stage="agent.verify",
)
verdicts = collect_list(parsed, "verdicts", _valid_verdict)
return AgentResult(scope_id=scope.scope_id, artifacts=[{
"finding_index": scope.payload["finding_index"],
"status": _status(verdicts),
"verdicts": verdicts,
}])
except Exception as exc:
return failure(scope, exc)
def apply_verdicts(
findings: List[Dict], verify_results: List[AgentResult]
) -> List[Dict]:
"""Annotate findings with verification; return refuted ones to suppress."""
by_index: Dict[int, Dict] = {}
for result in verify_results:
for artifact in result.artifacts:
by_index[artifact["finding_index"]] = artifact
suppressed = []
for index, finding in enumerate(findings):
artifact = by_index.get(index)
if not artifact:
continue
finding["verification"] = {
"status": artifact["status"],
"verdicts": artifact["verdicts"],
}
if artifact["status"] == "refuted":
finding["confidence"] = "low"
suppressed.append(finding)
return suppressed
```
**Step 4: Stamp `cluster_key` on constructability findings**
In `backend/agents/construct_agent.py` line 66-67, change:
```python
for finding in findings:
finding.update(agent=self.name, scope_id=scope.scope_id)
```
to:
```python
for finding in findings:
finding.update(agent=self.name, scope_id=scope.scope_id,
cluster_key=cluster.get("key"))
```
**Step 5: Wire wave 5b into `backend/agents/runner.py`**
After `memory.extend("findings", specialist_findings)` (line 167) and before
`gap_findings` / wave 6:
```python
orchestrator.stage("Agent wave 5b: evidence verification")
sheet_to_page = {
sheet.get("sheet_number"): sheet.get("page_number") for sheet in sheets
}
verify_targets = select_findings(
specialist_findings, clusters,
max_checks=config.AGENT_VERIFY_MAX_CHECKS,
severities=config.AGENT_VERIFY_SEVERITIES,
)
target_indexes = {id(f): i for i, f in enumerate(specialist_findings)}
verify_scopes = [
AgentScope(
scope_id=f"verify:{target_indexes[id(finding)]}",
payload={
"finding_index": target_indexes[id(finding)],
"finding": finding,
"images_b64": [
page_to_b64[sheet_to_page[name]]
for name in (finding.get("sheets") or [])
[:config.AGENT_CONFLICT_MAX_IMAGES]
if sheet_to_page.get(name) in page_to_b64
],
},
)
for finding in verify_targets
]
verify_results = orchestrator.run_scopes(
EvidenceVerifierAgent(usage), verify_scopes,
config.AGENT_VERIFY_CONCURRENCY,
)
suppressed = apply_verdicts(specialist_findings, verify_results)
if suppressed:
suppressed_ids = {id(f) for f in suppressed}
specialist_findings = [
f for f in specialist_findings if id(f) not in suppressed_ids
]
memory.replace("suppressed", suppressed)
memory.extend("findings", specialist_findings) # see note below
```
NOTE for implementer: `memory.extend("findings", ...)` already ran with the
un-suppressed list. Adjust ordering so verification happens BEFORE
`memory.extend("findings", specialist_findings)` — i.e. move the extend to after
wave 5b — so the Brain never sees refuted findings. Keep `gap_findings` logic
unchanged. Also add imports at top of runner.py:
```python
from backend.agents.verifier import (
EvidenceVerifierAgent, apply_verdicts, select_findings,
)
```
And in the report dicts (both the `require_review` branch ~line 217-225 and the
wave-7 branch ~line 293-300), populate suppressed issues:
```python
"suppressed_issues": memory.snapshot().get("suppressed") or [],
```
Finally, `verify` results cost shows up as `agent.verify` in
`summary.cost_by_stage` automatically via `usage_stage="agent.verify"`.
**Step 6: Update existing runner tests**
`tests/agents/test_runner_review_gate.py` monkeypatches `BrainAgent` and
`convert_pdf_to_images` but lets waves 1-5 run against... check how LLM calls
are stubbed there (likely `call_json` returns None -> empty artifacts, which is
fine). The new wave must no-op cleanly when `select_findings` returns `[]`
(zero scopes -> `run_scopes` returns `[]` per orchestrator.py:60). Verify by
running the suite; if a runner test now fails because verification selects a
stubbed finding, monkeypatch `select_findings` to `lambda *a, **k: []` in that
test file's `_patch_brain` helper.
**Step 7: Run full suite**
Run: `.venv/bin/python -m pytest tests/ -q`
Expected: all pass
**Step 8: Commit**
```bash
git add backend/agents/verifier.py backend/agents/runner.py backend/agents/construct_agent.py tests/agents/test_verifier.py tests/agents/test_runner_review_gate.py
git commit -m "feat: wave 5b evidence verification - vision fact-check before Brain merge"
```
---
### Task 7 (small, related): reasoning budget for the conflict critic
**Objective:** Fix the wave-4 truncation found in this same job (13/121 calls hit `finish_reason=length` at the 4096 cap with ~3.7k thinking tokens).
**Files:**
- Modify: `backend/agents/conflict_critic.py:55-63`
- Modify: `backend/pipeline/conflict_checker.py:85` (same pattern, classic path)
**Step 1: Apply the extractor's reasoning-knob pattern**
```python
parsed = call_json(
system_prompt=CONFLICT_SYSTEM_PROMPT,
user_text=instruction,
images_b64=images,
max_tokens=config.REASON_MAX_TOKENS,
model=config.AGENT_CONFLICT_MODEL,
reasoning_effort=config.EXTRACT_REASONING_EFFORT or None,
reasoning_max_tokens=config.EXTRACT_REASONING_MAX_TOKENS or None,
usage_tracker=self.usage,
usage_stage="agent.conflict",
)
```
(Match the exact kwarg names `SheetExtractorAgent` uses — check
`backend/agents/extractors.py` for whether it passes `None` when the budget is
0, and mirror that guard.)
**Step 2: Run tests**
Run: `.venv/bin/python -m pytest tests/ -q`
Expected: all pass
**Step 3: Commit**
```bash
git add backend/agents/conflict_critic.py backend/pipeline/conflict_checker.py
git commit -m "fix: reasoning budget for conflict critic (wave-4 max_tokens truncation)"
```
---
## Tests / validation
1. `.venv/bin/python -m pytest tests/ -q` — full suite green.
2. **Targeted repro of the original failure:** pull the S401 page image from job
959e16407573's output dir on sits-docker (or re-render the PDF page), then run
one `EvidenceVerifierAgent` scope locally against the finding JSON from
`validated_issues[3]`. Expected: verdict `corrected`,
`actual_text: "(5) 2x6 STUD PACK"`, status `refuted`.
3. **End-to-end:** rerun the same Cypress TX PDF through the pipeline (local with
`LLM_CACHE`/`LLM_RAW_DUMP` per the conflict-checker skill). Expected:
- log shows `Agent wave 5b: evidence verification` with a bounded number of calls;
- the S401 stud-pack finding is either absent from `validated_issues` and present
in `suppressed_issues` with verification verdicts, or downgraded to low confidence;
- `agent.verify` appears in `summary.cost_by_stage`;
- zero `finish_reason=length` lines in wave 4 (Task 7).
4. Cost check: wave 5b adds at most `AGENT_VERIFY_MAX_CHECKS` (20) vision calls —
for this job's profile that is well under $1.
## Risks, tradeoffs, open questions
- **Dispute false positives:** cluster members with legitimately different values
(e.g. two doors in one door cluster) will produce `disputed_attributes`. Mitigation:
prompts treat disputes as "unverified", not "wrong"; only severity-gated findings
burn verification calls. Tune later by restricting `find_disputes` to numeric-ish
values if noise is high.
- **Verifier can also misread.** It is one model checking another with the same eyes.
Mitigation: verdict requires `actual_text` verbatim evidence for `corrected`, and
only fully-unconfirmed findings are suppressed (mixed keeps the finding with a note).
- **xref cost:** extra link scopes. Bounded by the >= 2 distinct sheets gate and the
existing assertion cap; expect a handful of extra scopes per set.
- **Suppression in review mode:** refuted findings land in `suppressed_issues` — the
review UI/finalizer must tolerate that list being non-empty (it is currently always
`[]` in agent mode). Open question: surface suppressed items in the human review
queue as informational, or keep them report-only?
- **Open question:** should wave-4 conflict findings (which already saw images) also be
verification-eligible? Plan says no (they had the pixels); revisit if critics show
the same misread pattern.
+163
View File
@@ -0,0 +1,163 @@
# Session Notes — Conflict Checker
Orientation for a new coding session. Setup and Docker details live in [README.md](README.md). This file tracks what the code actually does and what tends to waste time.
**Last updated:** 2026-08-02 · `agent-mode` branch (merged `main` tip `bf508bf`)
## What this is
Cross-discipline **design contradiction / senior-architect QAQC** for construction drawing PDFs (Arch, Struct, Mech, Elec, Plumb, FP, etc.). Flags disagreements *between disciplines* before a set goes to bid/permit.
**Not** IronBids scope-ownership conflict checker.
Source of truth: Scout IT Gitea — `gitea.scoutitsystems.com/woogi/Conflict_Checker`.
## Stack
| Layer | Detail |
|-------|--------|
| API | Python 3.12, FastAPI, Uvicorn ([backend/main.py](backend/main.py)) |
| UI | Single static file [frontend/index.html](frontend/index.html), served by FastAPI |
| Pipelines | **Classic:** [backend/pipeline/runner.py](backend/pipeline/runner.py) (shared by web + CLI). **Agent:** [backend/agents/runner.py](backend/agents/runner.py) — experimental scoped specialist agents, selected per job (`pipeline_mode`) |
| Review gate | Agent jobs stop at `needs_review` for human decisions before the report emails ([backend/review/](backend/review/)) |
| LLM | OpenRouter via `openai` SDK; default `google/gemini-2.5-pro`. Vision always OpenRouter; text stages can use local vLLM (classic/hybrid only) |
| PDF | `pdf2image` + system `poppler-utils` → JPEG page images |
| Jobs | In-memory threads ([backend/jobs.py](backend/jobs.py)) — no Redis/DB |
| Tests | `pytest tests/` (~82 tests; see Quick start) |
| Deploy | Docker Compose; app on port **8099**; public URL `https://conchecker.scoutitsystems.com` |
## Live pipeline (authoritative)
README still describes an older 5-stage extract-then-compare loop. **Trust `runner.py`.** Actual flow:
```
PDF → images → extract → sheet index → jurisdiction
→ normalize → project intelligence (GOIDs)
→ cluster → conflict reason
→ QAQC / code / constructability
→ dedup-validate → risk → RFIs → report
```
| Runner stage | Module | Notes |
|--------------|--------|--------|
| PDF → images | `pdf_processor` | Rasterize |
| Extract assertions | `extractor` | Vision, per sheet |
| Classify sheet index | `sheet_index` | LLM |
| Jurisdiction profile | `jurisdiction` | After cover meta |
| Normalize | `normalizer` | LLM batches |
| Project intelligence | `normalizer.build_project_intelligence` | GOIDs + relationships |
| Cluster | `llm_clusterer` or `clusterer` | Default `CLUSTERER=llm` |
| Conflicts | `conflict_checker` | Per-cluster vision reason |
| QAQC / code / constructability | `qaqc_review`, `code_review`, `constructability` | Full-set / batched |
| Validate & dedup | `validator` | Merges conflict + QAQC + code + construct issues |
| Risk / RFIs | `risk`, `rfi` | Text-only |
| Report | `report` | `conflicts.json` + `report.md` (+ stage JSON dumps when `out_dir` set) |
Design notes for Stage 2/3 engines also live under `Changes/*.docx`.
**Agent pipeline** (`pipeline_mode=agent`): scoped specialist agents in [backend/agents/](backend/agents/) (extractors, linker, brain, critics, RFI writer) run through `Orchestrator` + `ProjectMemory`; artifacts under `outputs/<job_id>/agent/`. Agent mode is OpenRouter-only (no hybrid) and, when `AGENT_REQUIRE_REVIEW=true`, stops at `needs_review` until a human saves decisions and finalizes via the review endpoints. Design docs: `docs/superpowers/`.
## HTTP API (current)
| Method | Path | Purpose |
|--------|------|---------|
| GET | `/health` | Liveness + `model` / `text_model` + `version` / `build` + key/email flags |
| GET | `/models` | OpenRouter catalog split into `vision[]` / `text[]` + `defaults`, with per-1M-token pricing (cached ~1h; **502** when OpenRouter is unreachable) |
| POST | `/check` | Upload PDF; returns `{job_id}` immediately |
| GET | `/jobs/{id}` | Status poll. Running: `stage` + `log_tail`. Done/error/needs_review/finalization_error: `report` and/or `error` + full `log` |
| GET | `/jobs/{id}/log` | Full run log as `text/plain` (404 when no log file) |
| GET | `/jobs/{id}/review` | Review queue + progress + saved decisions (agent jobs) |
| POST | `/jobs/{id}/review-decisions` | Save reviewer decisions (409 outside needs_review/reviewing) |
| POST | `/jobs/{id}/finalize-review` | Background finalize + send report (409 unless review gate passed) |
| GET | `/jobs/{id}/sheet-image/{page}` | JPEG of source PDF page for the sheet viewer |
| GET | `/` | Serves `frontend/index.html` |
### `/check` form fields
- Required: `file` (PDF)
- Optional: `notification_email`, `project_name`, `address`, `occupancy`, `work_type`
- Pipeline: `pipeline_mode` (`classic` default, or `agent`)
- Compute: `text_local` (`true` = hybrid local text; forced off for agent mode)
- Models: `vision_model`, `text_model` (OpenRouter ids; blank = config defaults)
## Job logs
Pipeline `print()` is teed for the job thread ([backend/job_log.py](backend/job_log.py)):
- Live: `GET /jobs/{id}``log_tail` (last 80 lines)
- Done/error/needs_review/finalization_error: same payload includes full `log`
- Disk: `backend/outputs/<job_id>/job.log` (survives restart; status registry does not)
- API: `GET /jobs/{id}/log``text/plain`
- UI: “Run log” panel updates while running; stays visible after finish/fail/review
- Failures append the **full traceback** to the log; each run starts with a header line (job id, mode, models, start time)
- `outputs/<job_id>/job.json` (written at start) carries email/mode/models so the disk fallback can rebuild a job after restart
- **Verbose LLM observability** (`LLM_VERBOSE=true`, default on): one `[LLM] #NNNN stage | model (backend) | in/out sizes | cost | parsed-item counts` line per call in `job.log` — list-valued keys show counts (`conflicts[3]`) so empty (miss) or invented (hallucination) results stand out
- **Raw dumps** (`LLM_RAW_DUMP=true`, default on): full prompt + raw response per call in `outputs/<job_id>/llm_raw/NNNN_stage_model.json` (base64 images excluded, `n_images` recorded). This is the source for tracing *why* the model missed/invented an item
- **End-of-log cost**: every run ends with an `=== Estimated LLM cost (this run) ===` block (total, per-stage, models); failed runs append a cost-so-far line. Local/hybrid calls have no $ accounting — total covers OpenRouter only
## Vision vs text models
Two models, not one:
| Role | Config | Stages | Backend |
|------|--------|--------|---------|
| Vision | `MODEL` | Extract, conflict reason (images) | Always OpenRouter |
| Text | `TEXT_MODEL` (falls back to `MODEL`) | Sheet index, jurisdiction, normalize, cluster(LLM), QAQC, code, construct, validate, risk, RFI | OpenRouter, or local when hybrid |
UI: two dropdowns (with per-1M pricing) filled from `GET /models` ([backend/models.py](backend/models.py)), shown only for OpenRouter compute. Per-run picks go through `set_model_overrides(vision, text)` in [backend/llm.py](backend/llm.py): classic runs pass them as `run_pipeline` kwargs (runner clears in `finally`); agent runs set them module-level around `run_agent_pipeline`. UI picks beat per-call agent `AGENT_*_MODEL` args but **never name the hybrid local model** — local stays on `LOCAL_TEXT_MODEL`; the text pick only covers the cloud fallback.
## Where to change what
| Concern | File |
|---------|------|
| Prompts, vocab, conflict taxonomy | [backend/prompts.py](backend/prompts.py) |
| Env knobs | [backend/config.py](backend/config.py), [backend/.env.example](backend/.env.example) |
| HTTP API | [backend/main.py](backend/main.py) |
| UI (upload, models, live log, results) | [frontend/index.html](frontend/index.html) |
| CLI tuning loop | [cli/run_check.py](cli/run_check.py) |
| LLM client, cache, cost, model overrides | [backend/llm.py](backend/llm.py) |
| OpenRouter catalog + pricing + vision/text split | [backend/models.py](backend/models.py) |
| Agent-mode pipeline | [backend/agents/](backend/agents/) (`runner.py` entry; agents call `llm.call_json` with per-call model args) |
| Human review gate (queue, decisions, finalize) | [backend/review/](backend/review/) |
| Job registry + stdout tee log | [backend/jobs.py](backend/jobs.py), [backend/job_log.py](backend/job_log.py) |
| Stage helpers (prompt render, issue validate) | [backend/pipeline/_stage.py](backend/pipeline/_stage.py) |
| Code text corpus (Stage 7) | [backend/code_corpus/](backend/code_corpus/) |
| Hybrid local LLM helper | [scripts/setup_vllm.sh](scripts/setup_vllm.sh) |
Older prompt snapshot: `backend/prompts.py.v1`.
## Gotchas
1. **Prompt placeholders** — Use `str.replace` via `_stage.render`, never `str.format`. Prompts contain literal `{` JSON braces.
2. **Clustering** — Default is LLM (`CLUSTERER=llm`); empty LLM result falls back to deterministic. Deterministic clusters need ≥2 disciplines (or schedule-vs-plan); single-discipline “missing” gaps are a known limit.
3. **Jobs are in-memory** — Process restart clears job status; `outputs/<job_id>/` (report + `job.log` + `job.json` + `source.pdf`) still reload via disk fallback, including `needs_review` recovery.
4. **Dependency pin**`httpx==0.27.2` with `openai==1.51.0`. httpx ≥0.28 breaks openais `proxies=` kwarg.
5. **Code corpus licensing** — Only `ada_2010.txt` is shipped. Do not paste IBC/IFC/IECC without a license (see `backend/code_corpus/README.md`).
6. **Dual assertion schema** — Newer `{sheet, objects[]}` is mapped to legacy `{assertions[]}` with `attribute`/`value` for older stages.
7. **Grounding guard** — Extractor drops objects whose numeric claims are not in `source_text` (graphical-only objects allowed).
8. **Module globals** — LLM cost counters, stdout tee, and model overrides are process-global; overlapping jobs can interleave (single-user tool assumption).
9. **Samples**`samples/*.pdf` are gitignored; drop PDFs locally for CLI runs.
10. **README drift** — Treat README for setup/CI; treat this file + `runner.py` for pipeline truth. `prompts.py` header may still say some prompts are unwired — they are wired through the runner.
11. **Git identity** — This box has no `user.name` / `user.email`; commits need `GIT_AUTHOR_*` / `GIT_COMMITTER_*` env vars (do not `git config`). Remote push to Gitea works.
12. **No local Python deps on host** — App is meant to run in Docker; bare `python3` imports may miss `dotenv` / `httpx`. Prefer `docker compose`. For tests on this box: venv + requirements, but unpin Pillow (`Pillow>=11`) — 10.4.0 doesn't build on Python 3.14 (Docker uses 3.12, where the pin is fine).
13. **Agent mode constraints** — OpenRouter-only (hybrid disabled in UI and forced off server-side); review gate statuses are `needs_review → reviewing → finalizing → done` (`finalization_error` on finalize failure); only terminal states include the full `log` in polls.
## Quick start pointers
- Full setup: [README.md](README.md) (`docker compose up -d --build` → http://localhost:8099).
- Local CLI: `python cli/run_check.py samples/your_set.pdf --out out/your_set`.
- Prompt iteration: set `LLM_CACHE=true` in `backend/.env` so unchanged stages replay for free; clear with `rm -rf backend/.llm_cache`.
- Artifacts per job: `assertions.json`, `clusters.json`, per-stage JSON, `conflicts.json`, `report.md`, `job.log`, `job.json`, `source.pdf` under `backend/outputs/<job_id>/`.
- Tests: `python -m pytest tests/` (needs the deps from `requirements.txt` + `pytest`; on this box use a venv, see gotcha #12).
## Recent work (2026-08-02, agent-mode)
Merged `main` tip (`a6b0c8f` + `bf508bf`) into `agent-mode`, reconciling with this branch's own earlier implementations:
- **Two model dropdowns** — main's vision/text split ported onto this branch's priced catalog (`models.py`); pickers stay OpenRouter-compute-only, and UI picks never override `LOCAL_TEXT_MODEL` (main's hybrid footgun avoided).
- **Better run logs** — main's timestamped line-splitting tee, `log_tail` polls, terminal-state full log, and log-only disk recovery merged with this branch's header line, `job.json` metadata, and review-gate states. Failed runs now also append the traceback to `job.log`.
- `backend/models_catalog.py` (main's unpriced catalog) intentionally dropped in favor of `models.py`.
## Conflict categories (taxonomy)
Defined in `backend/prompts.py`: `dimensional_disagreement`, `elevation_disagreement`, `location_mismatch`, `missing_element`, `schedule_vs_plan_mismatch`, `detail_vs_plan_mismatch`, `tag_or_reference_inconsistency`, `spatial_clash`, `note_or_spec_contradiction`.
+35 -1
View File
@@ -39,7 +39,13 @@ PDF_DPI=100
MAX_PAGES=60
MAX_DIMENSION=2400
LLM_TIMEOUT=180
EXTRACT_MAX_TOKENS=8192
EXTRACT_MAX_TOKENS=65536
# Reasoning effort for per-sheet extraction (low keeps Gemini thinking tokens
# from eating the output budget). Blank = don't send the parameter.
EXTRACT_REASONING_EFFORT=low
# Hard thinking-token budget for extraction (OpenRouter reasoning max_tokens /
# Gemini thinking_budget). Stronger than effort; 0 = fall back to effort only.
EXTRACT_REASONING_MAX_TOKENS=2048
REASON_MAX_TOKENS=4096
EXTRACT_CONCURRENCY=4
REASON_CONCURRENCY=4
@@ -48,6 +54,13 @@ REASON_CONCURRENCY=4
APP_BASE_URL=https://conchecker.scoutitsystems.com
# APP_BUILD is set by CI at image build time (sha-<short_sha>) - do not set manually.
# LLM observability (job-log verbosity + raw request/response dumps)
# LLM_VERBOSE: one line per LLM call in job.log (model, sizes, item counts, cost)
# LLM_RAW_DUMP: full prompt+response per call in outputs/<job_id>/llm_raw/
# (base64 images excluded). Both default on; set false to quiet down.
LLM_VERBOSE=true
LLM_RAW_DUMP=true
# Email notifications (optional). Leave SMTP_HOST blank to disable.
# Examples:
# Gmail: SMTP_HOST=smtp.gmail.com SMTP_PORT=587 (use an App Password)
@@ -59,3 +72,24 @@ SMTP_PASSWORD=
SMTP_FROM=
SMTP_USE_TLS=true
SMTP_USE_SSL=false
# Wave 5b evidence verification (vision fact-check of cited sheet text)
AGENT_VERIFY_MAX_CHECKS=20
AGENT_VERIFY_SEVERITIES=critical,high
AGENT_VERIFY_REASONING_EFFORT=low
VERIFY_MAX_TOKENS=8192
# Text-layer grounding (deterministic PDF text layer via PyMuPDF)
# TEXT_LAYER_ENABLED: master switch for text-layer extraction/grounding
# TEXT_LAYER_MIN_CHARS: below this per page the sheet stays vision-only
# TEXT_LAYER_MAX_CHARS: cap of text layer injected into the extractor prompt
# VERIFY_TEXT_MAX_CHARS: cap of the text-layer excerpt in verify scopes
# VERIFY_HI_DPI_CROPS: evidence-located high-DPI crops in the verifier
# VERIFY_CROP_DPI / VERIFY_CROP_MARGIN_PTS: crop render DPI / padding (PDF points)
TEXT_LAYER_ENABLED=true
TEXT_LAYER_MIN_CHARS=20
TEXT_LAYER_MAX_CHARS=12000
VERIFY_TEXT_MAX_CHARS=8000
VERIFY_HI_DPI_CROPS=true
VERIFY_CROP_DPI=300
VERIFY_CROP_MARGIN_PTS=36
+2
View File
@@ -60,6 +60,8 @@ class ConflictCriticAgent:
model=config.AGENT_CONFLICT_MODEL,
usage_tracker=self.usage,
usage_stage="agent.conflict",
reasoning_effort=config.EXTRACT_REASONING_EFFORT or None,
reasoning_max_tokens=config.EXTRACT_REASONING_MAX_TOKENS or None,
)
candidates = parsed if isinstance(parsed, list) else (
parsed.get("conflicts") if isinstance(parsed, dict) else []
+3 -1
View File
@@ -47,6 +47,7 @@ class ConstructabilityAgent:
"assertions": dumps(cluster["assertions"]),
"clusters": dumps(slim_clusters([cluster])),
"conflicts": dumps(scope.payload.get("conflicts") or []),
"disputes": dumps(cluster.get("disputed_attributes") or []),
}
for key, value in substitutions.items():
instruction = instruction.replace("{" + key + "}", value)
@@ -64,7 +65,8 @@ class ConstructabilityAgent:
lambda item: validate_issue(item, "constructability"),
)
for finding in findings:
finding.update(agent=self.name, scope_id=scope.scope_id)
finding.update(agent=self.name, scope_id=scope.scope_id,
cluster_key=cluster.get("key"))
return AgentResult(scope_id=scope.scope_id, artifacts=findings)
except Exception as exc:
return failure(scope, exc)
+55
View File
@@ -0,0 +1,55 @@
"""Deterministic detection of contradictory extracted values within a cluster.
Extraction is a vision pass: quantities and sizes can be misread ("(2) 2x6" vs
"(5) 2x6"). Cluster members are supposed to describe the same real-world
element, so two members asserting different values for the same attribute are
a probable misread. Flag these so downstream text-only stages treat the value
as unverified instead of reasoning from one reading.
"""
import re
from typing import Dict, List
def _norm(value) -> str:
return re.sub(r"\s+", " ", ("" if value is None else str(value)).strip().lower())
def find_disputes(assertions: List[Dict]) -> List[Dict]:
"""Same attribute with >= 2 distinct normalized values = disputed."""
groups: Dict[str, Dict[str, Dict]] = {}
for assertion in assertions:
attribute = _norm(assertion.get("attribute"))
value = _norm(assertion.get("value"))
if not attribute or not value:
continue
# Group on the normalized value, but keep the original (whitespace-
# collapsed) text so disputes read like the sheet, not a lowercase munge.
original = re.sub(r"\s+", " ", str(assertion.get("value")).strip())
bucket = groups.setdefault(attribute, {}).setdefault(
value, {"original": original, "ids": set()}
)
bucket["ids"].add(assertion.get("id"))
disputes = []
for attribute, values in sorted(groups.items()):
if len(values) < 2:
continue
disputes.append({
"attribute": attribute,
"values": sorted(v["original"] for v in values.values()),
"assertion_ids": sorted(
aid for v in values.values() for aid in v["ids"] if aid
),
})
return disputes
def annotate_clusters(clusters: List[Dict]) -> int:
"""Attach disputed_attributes to each cluster that has any. Returns count."""
annotated = 0
for cluster in clusters:
disputes = find_disputes(cluster.get("assertions") or [])
if disputes:
cluster["disputed_attributes"] = disputes
annotated += 1
return annotated
+51 -12
View File
@@ -6,7 +6,7 @@ from typing import Dict
from backend import config
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
from backend.llm import call_json
from backend.pipeline.extractor import _normalize_sheet
from backend.pipeline.extractor import _normalize_sheet, _text_layer_block
from backend.pipeline.sheet_index import _index_input
from backend.prompts import (
EXTRACTOR_SYSTEM_PROMPT,
@@ -18,30 +18,69 @@ from backend.prompts import (
)
# Appended to the extractor instruction on the second-chance retry. Dense plan
# sheets blow the output budget on the full schema; compact mode trades
# per-object verbosity for actually finishing the page.
_COMPACT_RETRY_SUFFIX = """
IMPORTANT - COMPACT RETRY: the first pass did not complete. Keep the SAME JSON
schema, but extract at most 40 objects, prioritizing coordination-relevant
items (equipment, fixtures, devices, keynotes, dimensions, markers/callouts,
schedule rows). Keep descriptions/attributes short; skip review_uses entries
you are unsure about. Finish the JSON - a smaller complete answer beats a
larger truncated one."""
def _wrap_bare_list(parsed, page_number: int):
"""Models sometimes skip the {sheet, objects} wrapper and return a bare
objects array (especially after truncation repair). Accept it - the sheet
header falls back to title-block deduction downstream."""
if isinstance(parsed, list):
print(f"[Extract] Page {page_number}: wrapping bare objects array "
f"({len(parsed)} items, no sheet header)")
return {"sheet": {}, "objects": parsed}
return parsed
class SheetExtractorAgent:
name = "sheet_extractor"
def __init__(self, usage: AgentUsage) -> None:
self.usage = usage
def _call(self, instruction: str, page: Dict):
return call_json(
system_prompt=EXTRACTOR_SYSTEM_PROMPT,
user_text=instruction,
images_b64=[page["base64"]],
max_tokens=config.EXTRACT_MAX_TOKENS,
model=config.AGENT_EXTRACT_MODEL,
usage_tracker=self.usage,
usage_stage="agent.extract",
reasoning_effort=config.EXTRACT_REASONING_EFFORT or None,
reasoning_max_tokens=config.EXTRACT_REASONING_MAX_TOKENS or None,
)
def run(self, scope: AgentScope) -> AgentResult:
try:
page = scope.payload["page"]
instruction = EXTRACTOR_USER_INSTRUCTION.replace(
"{sheet_hint}", str(scope.payload.get("sheet_hint") or "")
)
parsed = call_json(
system_prompt=EXTRACTOR_SYSTEM_PROMPT,
user_text=instruction,
images_b64=[page["base64"]],
max_tokens=config.EXTRACT_MAX_TOKENS,
model=config.AGENT_EXTRACT_MODEL,
usage_tracker=self.usage,
usage_stage="agent.extract",
)
) + _text_layer_block(page)
parsed = _wrap_bare_list(self._call(instruction, page),
page["page_number"])
if not isinstance(parsed, dict):
# Second chance: same page, compact instructions. Runs only
# when the full-schema pass returned nothing usable.
print(f"[Extract] Page {page['page_number']}: full extraction "
f"failed, retrying compact")
parsed = _wrap_bare_list(
self._call(instruction + _COMPACT_RETRY_SUFFIX, page),
page["page_number"],
)
if not isinstance(parsed, dict):
raise ValueError("no structured extraction returned")
sheet = _normalize_sheet(parsed, page["page_number"])
sheet = _normalize_sheet(parsed, page["page_number"],
page_text=page.get("text_layer"))
return AgentResult(scope_id=scope.scope_id, artifacts=[sheet])
except Exception as exc:
return failure(scope, exc)
+28
View File
@@ -30,9 +30,23 @@ def _family(assertion: Dict) -> str:
)
def _xref_keys(assertion: Dict) -> List[str]:
"""Cross-level join keys: detail references and member tags."""
location = assertion.get("location_key") or {}
keys = []
ref = re.sub(r"\s+", "", str(location.get("detail_reference") or "")).upper()
if ref:
keys.append(f"detail:{ref}")
tag = re.sub(r"\s+", "", str(location.get("tag") or "")).upper()
if re.match(r"^[A-Z]+\d", tag): # member marks: W12X26, HSS16X4X5/8, ...
keys.append(f"tag:{tag}")
return keys
def build_link_scopes(sheets: List[Dict]) -> List[AgentScope]:
"""Partition facts by level and object/tag family, then enforce a hard cap."""
buckets: Dict[Tuple[str, str], List[Dict]] = defaultdict(list)
xref: Dict[str, List[Dict]] = defaultdict(list)
for sheet in sheets:
for assertion in sheet.get("assertions", []):
enriched = {
@@ -44,6 +58,8 @@ def build_link_scopes(sheets: List[Dict]) -> List[AgentScope]:
level = str((assertion.get("location_key") or {}).get("level")
or sheet.get("level") or "unknown").lower()
buckets[(level, _family(assertion))].append(enriched)
for key in _xref_keys(assertion):
xref[key].append(enriched)
scopes: List[AgentScope] = []
cap = max(2, config.AGENT_LINK_MAX_ASSERTIONS)
@@ -56,6 +72,18 @@ def build_link_scopes(sheets: List[Dict]) -> List[AgentScope]:
scope_id=f"{level}:{family}:{offset // cap + 1}",
payload={"assertions": chunk, "level": level, "family": family},
))
for key, assertions in sorted(xref.items()):
sheets_present = {a.get("sheet_number") for a in assertions}
if len(assertions) < 2 or len(sheets_present) < 2:
continue
chunk = assertions[:cap]
if len({a.get("sheet_number") for a in chunk}) < 2:
continue # cap landed on a single sheet — xref adds nothing
scopes.append(AgentScope(
scope_id=f"xref:{key}",
payload={"assertions": chunk,
"level": "xref", "family": key},
))
return scopes
+2 -1
View File
@@ -7,7 +7,8 @@ import threading
from typing import Any, Dict, Iterable, Optional
_COLLECTION_KEYS = {"sheets", "clusters", "findings", "decisions", "rfis"}
_COLLECTION_KEYS = {"sheets", "clusters", "findings", "decisions", "rfis",
"suppressed"}
_MAPPING_KEYS = {"sheet_index", "jurisdiction", "object_graph"}
_MEMORY_KEYS = _COLLECTION_KEYS | _MAPPING_KEYS
+118 -1
View File
@@ -1,5 +1,6 @@
"""Public entry point for the scoped Agent-mode pipeline."""
import base64
import json
import os
from typing import Callable, Dict, Optional
@@ -11,6 +12,7 @@ from backend.agents.code_agent import CodeAgent, build_code_scopes
from backend.agents.completeness import CompletenessAgent, build_sheet_summaries
from backend.agents.conflict_critic import ConflictCriticAgent
from backend.agents.construct_agent import ConstructabilityAgent, build_construct_scopes
from backend.agents.disputes import annotate_clusters
from backend.agents.extractors import (
JurisdictionAgent,
SheetExtractorAgent,
@@ -20,11 +22,18 @@ from backend.agents.linker import LinkerAgent, build_link_scopes, build_object_g
from backend.agents.memory import ProjectMemory
from backend.agents.orchestrator import Orchestrator
from backend.agents.rfi_writer import RFIWriterAgent
from backend.agents.verifier import (
EvidenceVerifierAgent, apply_verdicts, select_findings,
)
from backend.llm import reset_cost
from backend.pipeline.pdf_processor import convert_pdf_to_images
from backend.pipeline.report import build_report, to_markdown
from backend.pipeline.sheet_index import derive_project_meta_from_cover
from backend.review.gate import build_review_queue
from backend.review.store import ReviewStore
from backend.text_layer import (
attach_text_layers, coverage_gaps, find_evidence_bbox, render_crop,
)
def run_agent_pipeline(
@@ -39,6 +48,10 @@ def run_agent_pipeline(
if not os.path.isfile(pdf_path):
raise FileNotFoundError(pdf_path)
# Keep the llm module's counters job-local (matches Classic): the job
# log's failure-path cost estimate in jobs.py reads llm.get_cost().
reset_cost()
agent_dir = os.path.join(out_dir, "agent") if out_dir else None
memory = ProjectMemory(artifact_dir=agent_dir)
orchestrator = Orchestrator(memory=memory, on_stage=on_stage)
@@ -48,6 +61,9 @@ def run_agent_pipeline(
orchestrator.stage("Agent ingest: PDF -> images")
pages = convert_pdf_to_images(pdf_path)
page_to_b64 = {page["page_number"]: page["base64"] for page in pages}
text_dir = os.path.join(agent_dir, "text") if agent_dir else None
page_words = attach_text_layers(pdf_path, pages, text_dir=text_dir)
page_to_text = {page["page_number"]: page.get("text_layer") for page in pages}
orchestrator.stage("Agent wave 1: extract sheets")
extract_scopes = [
@@ -68,6 +84,13 @@ def run_agent_pipeline(
sheets.sort(key=lambda sheet: sheet.get("page_number") or 0)
memory.replace("sheets", sheets)
memory.dump("01-extract.json")
# Coverage signal: text layer present but extraction failed/empty reuses
# the failed-scopes gap-finding path (finding built below wave 6).
for gap_page in coverage_gaps(pages, sheets):
orchestrator.stats.failed_scopes.append(
f"sheet_extractor:sheet:{gap_page}: extraction gap "
f"(text layer present, no objects extracted)"
)
cover_meta = derive_project_meta_from_cover(
sheets, source_name or os.path.basename(pdf_path)
@@ -108,6 +131,9 @@ def run_agent_pipeline(
for artifact in result.artifacts
][:config.CLUSTER_MAX]
object_graph = build_object_graph(clusters)
disputed_count = annotate_clusters(clusters)
if disputed_count:
orchestrator.stage(f"[Link] {disputed_count} clusters carry disputed extracted values")
memory.replace("clusters", clusters)
memory.replace("object_graph", object_graph)
memory.dump("03-link.json")
@@ -159,6 +185,56 @@ def run_agent_pipeline(
for result in code_results + construct_results + completeness_results
for artifact in result.artifacts
]
orchestrator.stage("Agent wave 5b: evidence verification")
sheet_to_page = {str(s.get("sheet_number")): s.get("page_number")
for s in sheets}
verify_targets = select_findings(
specialist_findings, clusters,
max_checks=config.AGENT_VERIFY_MAX_CHECKS,
severities=config.AGENT_VERIFY_SEVERITIES,
)
target_indexes = {id(f): i for i, f in enumerate(specialist_findings)}
verify_scopes = []
for finding in verify_targets:
cited_pages = [
sheet_to_page[str(name)]
for name in (finding.get("sheets") or [])
if sheet_to_page.get(str(name)) in page_to_b64
]
images = [
page_to_b64[p]
for p in cited_pages[:config.AGENT_CONFLICT_MAX_IMAGES]
]
if not images:
continue # never judge evidence against images we could not load
# Text oracle: concatenated text layer of the cited sheets, capped.
excerpt = "\n\n".join(
f"--- Page {p} ---\n{page_to_text[p]}"
for p in cited_pages
if page_to_text.get(p)
)[:config.VERIFY_TEXT_MAX_CHARS]
if config.VERIFY_HI_DPI_CROPS:
images = _evidence_crops(finding, cited_pages, sheet_to_page,
page_words, page_to_b64, pdf_path,
fallback=images)
verify_scopes.append(AgentScope(
scope_id=f"verify:{target_indexes[id(finding)]}",
payload={
"finding_index": target_indexes[id(finding)],
"finding": finding,
"images_b64": images,
"text_layer_excerpt": excerpt,
},
))
verify_results = orchestrator.run_scopes(
EvidenceVerifierAgent(usage), verify_scopes, config.AGENT_VERIFY_CONCURRENCY)
suppressed = apply_verdicts(specialist_findings, verify_results)
if suppressed:
suppressed_ids = {id(f) for f in suppressed}
specialist_findings = [f for f in specialist_findings if id(f) not in suppressed_ids]
memory.replace("suppressed", suppressed)
memory.extend("findings", specialist_findings)
gap_findings = [
{
@@ -216,7 +292,7 @@ def run_agent_pipeline(
"project_intelligence": object_graph,
"validated_issues": prioritized,
"rfis": [],
"suppressed_issues": [],
"suppressed_issues": memory.snapshot().get("suppressed") or [],
})
progress = store.progress(queue)
# Same usage/stats summary block as the wave-7 path (rfis: 0 — they
@@ -292,6 +368,7 @@ def run_agent_pipeline(
"project_intelligence": object_graph,
"validated_issues": prioritized,
"rfis": rfis,
"suppressed_issues": memory.snapshot().get("suppressed") or [],
})
cost = usage.snapshot()
orchestrator.stats.calls = cost["calls"]
@@ -346,6 +423,46 @@ def _dump(out_dir: str, name: str, value) -> None:
json.dump(value, f, indent=2)
def _evidence_crops(
finding: Dict,
cited_pages: list,
sheet_to_page: Dict,
page_words: Dict,
page_to_b64: Dict,
pdf_path: str,
fallback: list,
) -> list:
"""High-DPI crops around each evidence item's source_text, located via the
page text layer. Crops REPLACE full-page images when at least one evidence
location resolves confidently; otherwise the full-page fallback is kept.
Never returns an empty list when fallback is non-empty (I2 guard)."""
crops: list = []
for item in finding.get("evidence") or []:
if len(crops) >= config.AGENT_CONFLICT_MAX_IMAGES:
break
if not isinstance(item, dict):
continue
source_text = item.get("source_text") or ""
if not source_text:
continue
# Prefer the page named on the evidence item, then any cited page.
candidates = []
named_page = sheet_to_page.get(str(item.get("sheet") or ""))
if named_page in cited_pages:
candidates.append(named_page)
candidates.extend(p for p in cited_pages if p not in candidates)
for page in candidates:
bbox = find_evidence_bbox(page_words.get(page) or [], source_text)
if bbox is None:
continue
crop = render_crop(pdf_path, page, bbox)
if not crop:
continue
crops.append(base64.b64encode(crop).decode("utf-8"))
break
return crops or fallback
def _counts(items, key: str) -> Dict[str, int]:
counts: Dict[str, int] = {}
for item in items:
+99
View File
@@ -0,0 +1,99 @@
"""Wave 5b: vision fact-check of extracted evidence against cited sheet images."""
from backend import config
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
from backend.llm import call_json
from backend.pipeline._serialize import dumps
from backend.pipeline._stage import collect_list, render
from backend.prompts import VERIFY_SYSTEM_PROMPT, VERIFY_USER_INSTRUCTION
_SEVERITY_RANK = {"critical": 0, "high": 1, "medium": 2, "low": 3}
_VERDICTS = ("confirmed", "corrected", "not_found")
def select_findings(findings, clusters, max_checks, severities):
"""Severity-gated selection plus any finding tied to a disputed cluster."""
disputed_keys = {c.get("key") for c in clusters if c.get("disputed_attributes")}
selected = [f for f in findings
if str(f.get("severity") or "").lower() in severities
or f.get("cluster_key") in disputed_keys]
selected.sort(key=lambda f: _SEVERITY_RANK.get(
str(f.get("severity") or "").lower(), 9))
return selected[:max_checks]
def _valid_verdict(item):
if not isinstance(item, dict):
return None
verdict = str(item.get("verdict") or "").lower()
if verdict not in _VERDICTS:
return None
return {"sheet": item.get("sheet") or "",
"source_text": item.get("source_text") or "",
"verdict": verdict,
"actual_text": item.get("actual_text"),
"notes": item.get("notes")}
def _status(verdicts):
if not verdicts:
return "unverified"
confirmed = sum(1 for v in verdicts if v["verdict"] == "confirmed")
if confirmed == len(verdicts):
return "confirmed"
if confirmed == 0:
return "refuted"
return "mixed"
class EvidenceVerifierAgent:
name = "verify"
def __init__(self, usage: AgentUsage) -> None:
self.usage = usage
def run(self, scope: AgentScope) -> AgentResult:
try:
finding = scope.payload["finding"]
instruction = render(VERIFY_USER_INSTRUCTION, {
"finding": dumps(finding),
"text_layer": scope.payload.get("text_layer_excerpt")
or "(no text layer available for the cited sheets)",
})
parsed = call_json(
system_prompt=VERIFY_SYSTEM_PROMPT,
user_text=instruction,
images_b64=scope.payload.get("images_b64") or [],
max_tokens=config.VERIFY_MAX_TOKENS,
model=config.AGENT_VERIFY_MODEL,
reasoning_effort=config.AGENT_VERIFY_REASONING_EFFORT or None,
usage_tracker=self.usage,
usage_stage="agent.verify",
)
verdicts = collect_list(parsed, "verdicts", _valid_verdict)
return AgentResult(scope_id=scope.scope_id, artifacts=[{
"finding_index": scope.payload["finding_index"],
"status": _status(verdicts),
"verdicts": verdicts,
}])
except Exception as exc:
return failure(scope, exc)
def apply_verdicts(findings, verify_results):
"""Annotate findings with verification; return refuted ones to suppress."""
by_index = {}
for result in verify_results:
for artifact in result.artifacts:
by_index[artifact["finding_index"]] = artifact
suppressed = []
for index, finding in enumerate(findings):
artifact = by_index.get(index)
if not artifact:
continue
finding["verification"] = {"status": artifact["status"],
"verdicts": artifact["verdicts"]}
if artifact["status"] == "refuted":
finding["confidence"] = "low"
suppressed.append(finding)
return suppressed
+47 -1
View File
@@ -47,6 +47,30 @@ AGENT_CONFLICT_CONCURRENCY = int(os.getenv("AGENT_CONFLICT_CONCURRENCY", "4"))
AGENT_SPECIALIST_CONCURRENCY = int(os.getenv("AGENT_SPECIALIST_CONCURRENCY", "4"))
AGENT_RFI_CONCURRENCY = int(os.getenv("AGENT_RFI_CONCURRENCY", "4"))
# Wave 5b evidence verification (vision fact-check of cited sheet text)
AGENT_VERIFY_MODEL = os.getenv("AGENT_VERIFY_MODEL", "") or MODEL
AGENT_VERIFY_CONCURRENCY = int(os.getenv("AGENT_VERIFY_CONCURRENCY", "4"))
AGENT_VERIFY_MAX_CHECKS = int(os.getenv("AGENT_VERIFY_MAX_CHECKS", "20"))
AGENT_VERIFY_SEVERITIES = {
s.strip().lower()
for s in os.getenv("AGENT_VERIFY_SEVERITIES", "critical,high").split(",")
if s.strip()
}
AGENT_VERIFY_REASONING_EFFORT = os.getenv("AGENT_VERIFY_REASONING_EFFORT", "low").strip()
VERIFY_MAX_TOKENS = int(os.getenv("VERIFY_MAX_TOKENS", "8192"))
# -- Text-layer grounding (deterministic PDF text layer via PyMuPDF) ----
# The vector text layer is extracted once per job and grounds the extractor,
# rescues misquoted-but-real values in the grounding guard, and serves the
# wave-5b verifier as a text oracle plus high-DPI evidence crops.
TEXT_LAYER_ENABLED = os.getenv("TEXT_LAYER_ENABLED", "true").strip().lower() in ("1", "true", "yes")
TEXT_LAYER_MIN_CHARS = int(os.getenv("TEXT_LAYER_MIN_CHARS", "20")) # below this per page -> no text layer
TEXT_LAYER_MAX_CHARS = int(os.getenv("TEXT_LAYER_MAX_CHARS", "12000")) # cap per sheet in extractor prompt
VERIFY_TEXT_MAX_CHARS = int(os.getenv("VERIFY_TEXT_MAX_CHARS", "8000"))# cap of excerpt in verify scope
VERIFY_HI_DPI_CROPS = os.getenv("VERIFY_HI_DPI_CROPS", "true").strip().lower() in ("1", "true", "yes")
VERIFY_CROP_DPI = int(os.getenv("VERIFY_CROP_DPI", "300"))
VERIFY_CROP_MARGIN_PTS = int(os.getenv("VERIFY_CROP_MARGIN_PTS", "36"))# padding around evidence bbox (PDF points)
# Agent-mode human-review gate. When on (default), Agent runs stop after the
# Brain merge and wait for human decisions before RFIs/final report/email go
# out. AGENT_REVIEW_AUDIT_SAMPLE caps how many clean clusters get added to the
@@ -73,7 +97,21 @@ PDF_DPI = int(os.getenv("PDF_DPI", "100"))
MAX_PAGES = int(os.getenv("MAX_PAGES", "60"))
MAX_DIMENSION = int(os.getenv("MAX_DIMENSION", "2400")) # px cap on the long edge
LLM_TIMEOUT = int(os.getenv("LLM_TIMEOUT", "180")) # seconds per call
EXTRACT_MAX_TOKENS = int(os.getenv("EXTRACT_MAX_TOKENS", "16384"))
# Gemini 2.5 Pro counts thinking tokens against max_tokens, so the visible
# JSON budget is well under this number on dense sheets. 65536 is the model's
# output ceiling - give thinking all the room it wants so visible JSON never
# truncates; the thinking budget itself is capped separately below.
EXTRACT_MAX_TOKENS = int(os.getenv("EXTRACT_MAX_TOKENS", "65536"))
# Reasoning effort for the per-sheet extractor (OpenRouter reasoning knob).
# Extraction is perceptive, not deliberative - "low" keeps thinking tokens
# from eating the output budget. Empty string disables the parameter.
EXTRACT_REASONING_EFFORT = os.getenv("EXTRACT_REASONING_EFFORT", "low").strip()
# Hard thinking-token budget for the extractor (OpenRouter reasoning
# max_tokens -> Gemini thinking_budget). "low" effort alone still let Gemini
# burn ~25k thinking tokens per sheet (job 98194fa8d215); a hard cap forces
# the budget into visible output. 0 disables -> falls back to the effort knob.
# Mutually exclusive with effort when set (OpenRouter rejects both together).
EXTRACT_REASONING_MAX_TOKENS = int(os.getenv("EXTRACT_REASONING_MAX_TOKENS", "2048"))
REASON_MAX_TOKENS = int(os.getenv("REASON_MAX_TOKENS", "4096"))
# -- QAQC stage knobs (Stages 0-1, 3, 6-11) -------------------------
@@ -102,6 +140,14 @@ CLUSTER_MAX = int(os.getenv("CLUSTER_MAX", "120"))
LLM_CACHE = os.getenv("LLM_CACHE", "false").strip().lower() in ("1", "true", "yes")
LLM_CACHE_DIR = os.getenv("LLM_CACHE_DIR", os.path.join(_BASE_DIR, ".llm_cache"))
# Verbose LLM observability. Per call, one line lands in the job log (model,
# backend, prompt size, response size, parsed-item counts, per-call cost) and
# the full request/response is dumped to <out_dir>/llm_raw/ (base64 image
# payloads excluded; image count recorded instead) so missed or hallucinated
# items can be traced back to exactly what the model saw and returned.
LLM_VERBOSE = os.getenv("LLM_VERBOSE", "true").strip().lower() in ("1", "true", "yes")
LLM_RAW_DUMP = os.getenv("LLM_RAW_DUMP", "true").strip().lower() in ("1", "true", "yes")
# Parallelism (ThreadPoolExecutor workers)
EXTRACT_CONCURRENCY = int(os.getenv("EXTRACT_CONCURRENCY", "4"))
REASON_CONCURRENCY = int(os.getenv("REASON_CONCURRENCY", "4"))
+93
View File
@@ -0,0 +1,93 @@
"""
job_log.py - Capture pipeline stdout/stderr into a per-job log.
The pipeline already prints stage progress via print(). For a reviewable
post-run log we tee those lines into memory + outputs/<job_id>/job.log
without rewriting every call site.
"""
import sys
import time
from contextlib import contextmanager
from typing import Callable, Iterator, List, Optional, TextIO
class _LineSplitter:
"""Accumulate write() chunks and emit complete lines."""
def __init__(self, on_line: Callable[[str], None]):
self._buf = ""
self._on_line = on_line
def write(self, s: str) -> None:
if not s:
return
self._buf += s
while "\n" in self._buf:
line, self._buf = self._buf.split("\n", 1)
# Strip trailing CR from Windows-ish streams; keep content intact.
self._on_line(line.rstrip("\r"))
def flush_remainder(self) -> None:
if self._buf:
self._on_line(self._buf.rstrip("\r"))
self._buf = ""
class _Tee:
"""Mirror writes to the original stream and a line callback."""
def __init__(self, stream: TextIO, on_line: Callable[[str], None]):
self._stream = stream
self._lines = _LineSplitter(on_line)
def write(self, s: str) -> int:
n = self._stream.write(s)
self._stream.flush()
self._lines.write(s)
return n
def flush(self) -> None:
self._stream.flush()
def flush_remainder(self) -> None:
self._lines.flush_remainder()
def __getattr__(self, name: str):
return getattr(self._stream, name)
def stamp_line(line: str, t: Optional[float] = None) -> str:
"""Prefix a log line with HH:MM:SS."""
ts = time.strftime("%H:%M:%S", time.localtime(t if t is not None else time.time()))
return f"[{ts}] {line}"
@contextmanager
def capture_stdio(on_line: Callable[[str], None]) -> Iterator[None]:
"""
Tee sys.stdout and sys.stderr into on_line(raw_line) for the duration.
Safe for the single-job-at-a-time usage of this app; overlapping jobs
would interleave (same limitation as the LLM cost counters).
"""
old_out, old_err = sys.stdout, sys.stderr
tee_out = _Tee(old_out, on_line)
tee_err = _Tee(old_err, on_line)
sys.stdout = tee_out # type: ignore[assignment]
sys.stderr = tee_err # type: ignore[assignment]
try:
yield
finally:
tee_out.flush_remainder()
tee_err.flush_remainder()
sys.stdout = old_out
sys.stderr = old_err
def read_log_file(path: str) -> List[str]:
try:
with open(path, encoding="utf-8") as f:
return [ln.rstrip("\n") for ln in f]
except OSError:
return []
+204 -78
View File
@@ -9,61 +9,34 @@ for the completion email.
State is in-memory (fine for a single-user tool); the report is also persisted
to outputs/<job_id>/ so results survive a restart even though live status does
not. No external queue/DB.
A teed stdout/stderr log is kept in memory and written to outputs/<job_id>/job.log
so failed or suspicious runs can be reviewed after the fact.
"""
import json
import contextlib
import os
import sys
import time
import traceback
import uuid
import shutil
import threading
from typing import Dict, Optional
from contextlib import contextmanager
from typing import Dict, Iterator, List, Optional
from backend import config
from backend import llm
from backend.job_log import capture_stdio, read_log_file, stamp_line
from backend.agents.runner import run_agent_pipeline
from backend.pipeline.runner import run_pipeline
from backend.email_sender import send_conflict_report, send_review_required
_jobs: Dict[str, Dict] = {}
_lock = threading.Lock()
_LOG_TAIL = 80
PIPELINE_MODES = {"classic", "agent"}
class _Tee:
"""Write to both the real stream and the job log file."""
def __init__(self, stream, log_file) -> None:
self._stream = stream
self._log = log_file
def write(self, data):
self._stream.write(data)
self._log.write(data)
def flush(self):
self._stream.flush()
self._log.flush()
@contextlib.contextmanager
def _tee_log(log_path: str, header: str):
"""Mirror stdout/stderr into a per-job log file for the duration of a run.
sys.stdout is process-global, so two concurrent jobs would interleave in
each other's logs - acceptable for this single-user tool (same tradeoff as
the LLM cost globals in llm.py).
"""
with open(log_path, "a", encoding="utf-8") as log_file:
log_file.write(header + "\n")
real_out, real_err = sys.stdout, sys.stderr
sys.stdout, sys.stderr = _Tee(real_out, log_file), _Tee(real_err, log_file)
try:
yield
finally:
sys.stdout, sys.stderr = real_out, real_err
# States where the job will produce no more log output; polls get the full log.
_TERMINAL_STATES = {"done", "error", "needs_review", "finalization_error"}
def _set(job_id: str, **fields) -> None:
@@ -71,16 +44,39 @@ def _set(job_id: str, **fields) -> None:
_jobs[job_id].update(fields)
def create_job(pdf_path: str, source_filename: str, email: Optional[str] = None,
project_input: Optional[Dict] = None, text_local: bool = False,
pipeline_mode: str = "classic", model: Optional[str] = None) -> str:
def _append_log(job_id: str, raw_line: str, log_path: str) -> None:
"""Stamp, store, and append one captured stdout/stderr line."""
entry = stamp_line(raw_line)
with _lock:
job = _jobs.get(job_id)
if job is not None:
job.setdefault("log", []).append(entry)
try:
os.makedirs(os.path.dirname(log_path), exist_ok=True)
with open(log_path, "a", encoding="utf-8") as f:
f.write(entry + "\n")
except OSError:
# Don't fail the job over log I/O; avoid print() here — it would
# re-enter the stdio tee while a job is capturing.
pass
def create_job(
pdf_path: str,
source_filename: str,
email: Optional[str] = None,
project_input: Optional[Dict] = None,
text_local: bool = False,
pipeline_mode: str = "classic",
vision_model: Optional[str] = None,
text_model: Optional[str] = None,
) -> str:
"""Register a job and kick off its background thread. Returns the job_id."""
pipeline_mode = pipeline_mode.strip().lower()
if pipeline_mode not in PIPELINE_MODES:
raise ValueError(f"Unsupported pipeline mode: {pipeline_mode!r}")
# Agent mode v1 is OpenRouter-only.
text_local = bool(text_local and pipeline_mode == "classic")
model = (model or "").strip() or None
job_id = uuid.uuid4().hex[:12]
with _lock:
_jobs[job_id] = {
@@ -91,35 +87,67 @@ def create_job(pdf_path: str, source_filename: str, email: Optional[str] = None,
"project_input": project_input or {},
"text_local": text_local,
"pipeline_mode": pipeline_mode,
"model": model,
"vision_model": (vision_model or "").strip() or None,
"text_model": (text_model or "").strip() or None,
"stage": None,
"created_at": time.time(),
"finished_at": None,
"report": None,
"error": None,
"log": [],
}
threading.Thread(target=_run, args=(
job_id, pdf_path, project_input, text_local, pipeline_mode, model,
),
daemon=True).start()
threading.Thread(
target=_run,
args=(job_id, pdf_path, project_input, text_local, pipeline_mode,
vision_model, text_model),
daemon=True,
).start()
return job_id
def _run(job_id: str, pdf_path: str, project_input: Optional[Dict] = None,
text_local: bool = False, pipeline_mode: str = "classic",
model: Optional[str] = None) -> None:
def _run(
job_id: str,
pdf_path: str,
project_input: Optional[Dict] = None,
text_local: bool = False,
pipeline_mode: str = "classic",
vision_model: Optional[str] = None,
text_model: Optional[str] = None,
) -> None:
out_dir = os.path.join(config.OUTPUT_DIR, job_id)
log_path = os.path.join(out_dir, "job.log")
try:
_set(job_id, status="running")
# Keep a copy of the source PDF so its sheets can be viewed later.
os.makedirs(out_dir, exist_ok=True)
# Truncate any leftover log if job_id somehow collided (shouldn't).
with open(log_path, "w", encoding="utf-8"):
pass
header = (f"=== Job {job_id} | {pipeline_mode} | {_jobs[job_id].get('source')} | "
f"model={model or 'default'} | "
f"vision={vision_model or 'default'} text={text_model or 'default'} | "
f"started {time.strftime('%Y-%m-%d %H:%M:%S %Z', time.gmtime())} UTC ===")
with _tee_log(os.path.join(out_dir, "job.log"), header):
_append_log(job_id, header, log_path)
def on_line(raw: str) -> None:
_append_log(job_id, raw, log_path)
with capture_stdio(on_line):
_run_pipeline(job_id, pdf_path, out_dir, project_input, text_local,
pipeline_mode, model)
pipeline_mode, vision_model, text_model)
except Exception as e:
# Land the failure AND its traceback in the job log so failed runs can
# be diagnosed from the log alone (the stdio tee is already torn down).
try:
_append_log(job_id, f"[Jobs] Job {job_id} failed: {e}", log_path)
for ln in traceback.format_exc().rstrip().splitlines():
_append_log(job_id, ln, log_path)
# Cost-so-far for the failed run (counters are reset per job).
cost = llm.get_cost()
_append_log(job_id,
f"[Jobs] Estimated LLM cost before failure: "
f"${cost['usd']:.4f} over {cost['calls']} live calls "
f"({cost.get('cached', 0)} cached)", log_path)
except Exception:
pass
print(f"[Jobs] Job {job_id} failed: {e}")
_set(job_id, status="error", error=str(e), finished_at=time.time())
_notify_error(job_id)
@@ -132,7 +160,8 @@ def _run(job_id: str, pdf_path: str, project_input: Optional[Dict] = None,
def _run_pipeline(job_id: str, pdf_path: str, out_dir: str,
project_input: Optional[Dict], text_local: bool,
pipeline_mode: str, model: Optional[str]) -> None:
pipeline_mode: str, vision_model: Optional[str],
text_model: Optional[str]) -> None:
"""The body of a job run; executes inside the job's tee'd log capture."""
# Persist minimal job metadata so the disk fallback in get_job can
# recover the recipient email / pipeline mode after a server restart
@@ -143,8 +172,14 @@ def _run_pipeline(job_id: str, pdf_path: str, out_dir: str,
"email": _jobs[job_id].get("email"),
"pipeline_mode": pipeline_mode,
"source": _jobs[job_id].get("source"),
"vision_model": vision_model,
"text_model": text_model,
}, f, indent=2)
# Keep a copy of the source PDF so its sheets can be viewed later.
shutil.copy2(pdf_path, os.path.join(out_dir, "source.pdf"))
# Raw per-call LLM request/response dumps land in <out_dir>/llm_raw/.
if config.LLM_RAW_DUMP:
llm.set_raw_dump_dir(os.path.join(out_dir, "llm_raw"))
runner = run_agent_pipeline if pipeline_mode == "agent" else run_pipeline
runner_kwargs = {
"out_dir": out_dir,
@@ -153,18 +188,25 @@ def _run_pipeline(job_id: str, pdf_path: str, out_dir: str,
"source_name": _jobs[job_id].get("source"),
}
if pipeline_mode == "classic":
# run_pipeline takes the picks as params and clears them in finally.
runner_kwargs["text_local"] = text_local
runner_kwargs["vision_model"] = vision_model
runner_kwargs["text_model"] = text_model
else:
runner_kwargs["require_review"] = config.AGENT_REQUIRE_REVIEW
if model:
print(f"[Jobs] Model override for this run: {model}")
llm.set_model_override(model)
# The agent runner has no override params; set them module-level.
if vision_model or text_model:
print(f"[Jobs] Model overrides for this run: "
f"vision={vision_model or '(default)'} text={text_model or '(default)'}")
llm.set_model_overrides(vision_model, text_model)
try:
report = runner(pdf_path, **runner_kwargs)
finally:
if model:
llm.set_model_override(None)
llm.set_raw_dump_dir(None)
if pipeline_mode == "agent":
llm.set_model_overrides(None, None)
report.setdefault("summary", {})["pipeline_mode"] = pipeline_mode
_log_cost_summary(report.get("summary", {}))
if report["summary"].get("agent_status") == "needs_review":
# Human-review gate: hold the job, don't email the unreviewed report.
_set(job_id, status="needs_review", report=report,
@@ -178,6 +220,55 @@ def _run_pipeline(job_id: str, pdf_path: str, out_dir: str,
_notify(job_id, report, out_dir)
def _log_cost_summary(summary: Dict, label: str = "this run") -> None:
"""
End-of-log estimated LLM cost block, printed inside the job's stdio tee so
it lands at the tail of job.log. Both runners populate the same summary
fields (cost_usd / llm_calls / cached_calls / cost_by_stage / models_used).
Hybrid note: local text calls carry no usage accounting, so the dollar
figure covers OpenRouter calls only (local call counts still appear).
"""
if summary.get("cost_usd") is None and not summary.get("llm_calls"):
return
print(f"\n=== Estimated LLM cost ({label}) ===")
print(f" Total: ${summary.get('cost_usd', 0.0):.4f} across "
f"{summary.get('llm_calls', 0)} live calls "
f"({summary.get('cached_calls', 0)} cached at $0)")
for name, bucket in (summary.get("cost_by_stage") or {}).items():
print(f" {name}: ${bucket.get('usd', 0.0):.4f} "
f"({bucket.get('calls', 0)} live, {bucket.get('cached', 0)} cached)")
mu = summary.get("models_used") or {}
if mu.get("vision"):
print(f" Vision models: {', '.join(mu['vision'])}")
if mu.get("text_local"):
print(f" Text models (local): {', '.join(mu['text_local'])}")
if mu.get("text_cloud"):
print(f" Text models (cloud): {', '.join(mu['text_cloud'])}")
if mu.get("fallback_count"):
print(f" Local->cloud fallbacks: {mu['fallback_count']}")
if mu.get("text_local"):
print(" Note: local calls have no cost accounting; "
"the dollar total covers OpenRouter usage only.")
@contextmanager
def capture_job_output(job_id: str, out_dir: str) -> Iterator[None]:
"""
Re-open the stdio tee + raw LLM dump dir for post-run work that still
belongs to this job (review finalization): lines append to job.log and
the in-memory log, raw dumps resume under <out_dir>/llm_raw/. The tee is
process-global — same overlapping-job caveat as the main run.
"""
log_path = os.path.join(out_dir, "job.log")
if config.LLM_RAW_DUMP:
llm.set_raw_dump_dir(os.path.join(out_dir, "llm_raw"))
try:
with capture_stdio(lambda raw: _append_log(job_id, raw, log_path)):
yield
finally:
llm.set_raw_dump_dir(None)
def _notify(job_id: str, report: Dict, out_dir: str) -> None:
email = _jobs[job_id].get("email")
if not email:
@@ -197,11 +288,6 @@ def _notify_error(job_id: str) -> None:
email = job.get("email")
if not email:
return
# Reuse the report mailer with a minimal error-shaped payload.
err_report = {
"source": job.get("source", ""),
"summary": {"conflicts_found": 0, "by_severity": {}, "disciplines": []},
}
try:
from backend.email_sender import _smtp_ready, _send
from email.message import EmailMessage
@@ -216,6 +302,7 @@ def _notify_error(job_id: str) -> None:
"Your conflict check did not complete.\n\n"
f"Drawing set: {job.get('source','')}\n"
f"Error: {job.get('error','unknown')}\n\n"
f"Review the run log at: {config.APP_BASE_URL.rstrip('/')}/?job={job_id}\n\n"
"Generated by Conflict Checker"
)
_send(msg)
@@ -223,28 +310,63 @@ def _notify_error(job_id: str) -> None:
print(f"[Email] Failed to send error notice: {e}")
def _log_from_disk(job_id: str) -> List[str]:
return read_log_file(os.path.join(config.OUTPUT_DIR, job_id, "job.log"))
def get_job_log(job_id: str) -> Optional[List[str]]:
"""Full job log lines, from memory or disk. None if job unknown."""
with _lock:
job = _jobs.get(job_id)
if job is not None:
return list(job.get("log") or [])
log = _log_from_disk(job_id)
# Job exists on disk if we have a log or a report artifact.
report_path = os.path.join(config.OUTPUT_DIR, job_id, "conflicts.json")
if log or os.path.isfile(report_path):
return log
return None
def get_job(job_id: str) -> Optional[Dict]:
"""Public job view. Includes the full report only when done.
Falls back to the on-disk conflicts.json when the job isn't in the
in-memory registry (e.g. after a server restart).
Falls back to the on-disk artifacts (conflicts.json / job.log) when the
job isn't in the in-memory registry (e.g. after a server restart).
"""
with _lock:
job = _jobs.get(job_id)
if job:
return dict(job)
out = dict(job)
log = list(job.get("log") or [])
out["log_tail"] = log[-_LOG_TAIL:]
# Full log on terminal states so the UI can show it without a
# second fetch; keep polls light while running.
if out.get("status") in _TERMINAL_STATES:
out["log"] = log
else:
out.pop("log", None)
return out
# Try loading from disk
report_path = os.path.join(config.OUTPUT_DIR, job_id, "conflicts.json")
if not os.path.isfile(report_path):
log = _log_from_disk(job_id)
if not os.path.isfile(report_path) and not log:
return None
try:
with open(report_path, encoding="utf-8") as f:
report = json.load(f)
summary = report.get("summary", {})
# Recover the job's real state: a job that stopped at the review gate
# must come back as needs_review (not done) or it can never finalize.
status = "needs_review" if summary.get("agent_status") == "needs_review" else "done"
report = None
if os.path.isfile(report_path):
with open(report_path, encoding="utf-8") as f:
report = json.load(f)
summary = (report or {}).get("summary", {})
if report is None:
# Crashed before writing a report; the log is the only artifact.
status = "error"
else:
# Recover the job's real state: a job that stopped at the review gate
# must come back as needs_review (not done) or it can never finalize.
status = "needs_review" if summary.get("agent_status") == "needs_review" else "done"
# job.json (written at job start) carries the recipient email and
# pipeline mode so the final notification still fires after a restart.
# Missing/corrupt job.json degrades to the previous derivations.
@@ -261,16 +383,20 @@ def get_job(job_id: str) -> Optional[Dict]:
job = {
"job_id": job_id,
"status": status,
"source": meta.get("source") or report.get("source", os.path.basename(report_path)),
"source": meta.get("source") or (report or {}).get("source", os.path.basename(report_path)),
"email": meta.get("email"),
"project_input": report.get("project_input", {}),
"project_input": (report or {}).get("project_input", {}),
"text_local": summary.get("text_backend") == "local",
"pipeline_mode": meta.get("pipeline_mode") or summary.get("pipeline_mode", "classic"),
"vision_model": meta.get("vision_model"),
"text_model": meta.get("text_model"),
"stage": None,
"created_at": os.path.getmtime(source_pdf) if os.path.isfile(source_pdf) else None,
"finished_at": os.path.getmtime(report_path),
"finished_at": os.path.getmtime(report_path) if os.path.isfile(report_path) else None,
"report": report,
"error": None,
"error": None if report is not None else "Report missing; see job log",
"log": log,
"log_tail": log[-_LOG_TAIL:],
}
# Hydrate the in-memory registry so _set(...) transitions (reviewing,
# finalizing, done) work for restart-recovered jobs.
+187 -18
View File
@@ -8,6 +8,7 @@ with an image) and conflict-reasoning (Stage 3, with images) calls.
"""
import os
import re
import json
import hashlib
import threading
@@ -23,16 +24,115 @@ _clients: Dict[str, OpenAI] = {}
# call_json when routing a no-image (text) call. Module-global mirrors the
# set_stage/cost pattern (single-user tool).
_text_local = False
# Per-job model override (user picked a model in the UI). Same module-global
# Per-job model overrides (user picked models in the UI). Same module-global
# pattern: set by the job runner before the pipeline starts, cleared after.
_model_override: Optional[str] = None
# Vision applies to image calls, text to no-image calls on OpenRouter (and to
# the local->cloud fallback). The LOCAL endpoint's model name is never taken
# from these overrides - hybrid local keeps LOCAL_TEXT_MODEL.
_vision_model_override: Optional[str] = None
_text_model_override: Optional[str] = None
# Per-job raw request/response dumps (missed/hallucinated-item debugging).
# Set by the job runner to <out_dir>/llm_raw at job start, cleared after.
# Same module-global pattern as the model overrides (single-user tool).
_raw_dump_dir: Optional[str] = None
_seq_lock = threading.Lock()
_call_seq = 0
def set_model_override(model: Optional[str]) -> None:
"""Override the model for all OpenRouter calls (vision + text), or None to clear."""
global _model_override
_model_override = (model or "").strip() or None
def set_raw_dump_dir(path: Optional[str]) -> None:
"""Point raw LLM request/response dumps at a directory. None disables."""
global _raw_dump_dir, _call_seq
with _seq_lock:
_raw_dump_dir = path
_call_seq = 0
def _next_seq() -> int:
global _call_seq
with _seq_lock:
_call_seq += 1
return _call_seq
def _summarize_parsed(parsed: Any) -> str:
"""
Compact digest of a parsed response for the job log. List values become
item counts (e.g. conflicts[3]) so a stage that returned nothing (miss)
or invented items (hallucination) is visible without opening the raw dump.
"""
if isinstance(parsed, list):
return f"list[{len(parsed)}]"
if not isinstance(parsed, dict):
return type(parsed).__name__
parts = []
for k, v in parsed.items():
if isinstance(v, list):
parts.append(f"{k}[{len(v)}]")
elif isinstance(v, dict):
parts.append(f"{k}{{{len(v)}}}")
else:
s = str(v)
parts.append(f"{k}={s[:40]!r}{'...' if len(s) > 40 else ''}")
out = ", ".join(parts)
return out[:300] + ("..." if len(out) > 300 else "")
def _dump_raw(seq: int, be: Dict[str, Any], usage_stage: str,
system_prompt: str, user_text: str,
images_b64: Optional[List[str]], max_tokens: int,
raw: str, parsed: Any,
finish_reason: Optional[str] = None) -> None:
"""Write the full request/response for one call to the job's llm_raw dir."""
if not _raw_dump_dir:
return
try:
os.makedirs(_raw_dump_dir, exist_ok=True)
safe_stage = re.sub(r"[^A-Za-z0-9_.-]+", "_", usage_stage)[:40]
safe_model = re.sub(r"[^A-Za-z0-9_.-]+", "_", be["model"])
payload = {
"seq": seq,
"stage": usage_stage,
"model": be["model"],
"backend": "local" if be.get("local") else "cloud",
"max_tokens": max_tokens,
# base64 image payloads deliberately excluded (multi-MB each);
# the count + source.pdf in the job dir identify what was sent.
"n_images": len(images_b64 or []),
"system_prompt": system_prompt,
"user_text": user_text,
"finish_reason": finish_reason,
"raw_response": raw,
"parsed": parsed,
}
path = os.path.join(_raw_dump_dir, f"{seq:04d}_{safe_stage}_{safe_model}.json")
tmp = f"{path}.{threading.get_ident()}.tmp"
with open(tmp, "w", encoding="utf-8") as f:
json.dump(payload, f, indent=2)
os.replace(tmp, path) # atomic so thread-pooled stages can't tear it
except OSError as e:
print(f"[LLM] raw dump failed: {e}")
def _log_call(seq: int, be: Dict[str, Any], usage_stage: str,
user_text: str, images_b64: Optional[List[str]],
raw: str, parsed: Any, usd: Optional[float],
cached: bool = False,
reasoning_tokens: Optional[int] = None) -> None:
"""One verbose per-call line for the job log (tee'd by job_log.py)."""
if not config.LLM_VERBOSE:
return
backend = "local" if be.get("local") else "cloud"
if cached:
cost_str = "cache hit"
elif usd is not None:
cost_str = f"${usd:.4f}"
else:
cost_str = "cost n/a"
think_str = f"think {reasoning_tokens}tk | " if reasoning_tokens is not None else ""
print(f"[LLM] #{seq:04d} {usage_stage} | {be['model']} ({backend}) | "
f"in {len(user_text)}ch+{len(images_b64 or [])}img | "
f"out {len(raw)}ch | {think_str}{cost_str} | {_summarize_parsed(parsed)}")
def set_text_backend(local: bool) -> None:
@@ -40,6 +140,13 @@ def set_text_backend(local: bool) -> None:
global _text_local
_text_local = bool(local)
def set_model_overrides(vision: Optional[str] = None, text: Optional[str] = None) -> None:
"""Per-run OpenRouter vision/text model picks. None/blank clears to defaults."""
global _vision_model_override, _text_model_override
_vision_model_override = (vision or "").strip() or None
_text_model_override = (text or "").strip() or None
# --- per-job cost accounting -------------------------------------------------
# OpenRouter returns the real USD cost of each call when we request usage
# accounting. We accumulate it in a module-level counter; the runner resets it
@@ -125,9 +232,11 @@ def _add_cached() -> None:
def _cache_key(model: str, system_prompt: str, user_text: str,
images_b64: Optional[List[str]], max_tokens: int,
json_mode: bool) -> str:
json_mode: bool, reasoning_effort: Optional[str] = None,
reasoning_max_tokens: Optional[int] = None) -> str:
h = hashlib.sha256()
parts = [model, str(max_tokens), str(json_mode), system_prompt, user_text]
parts = [model, str(max_tokens), str(json_mode), str(reasoning_effort),
str(reasoning_max_tokens), system_prompt, user_text]
for b in (images_b64 or []):
parts.append(b)
for p in parts:
@@ -177,17 +286,22 @@ def _resolve_backend(has_images: bool, model_override: Optional[str]) -> Dict[st
return {
"base_url": config.LOCAL_BASE_URL,
"api_key": config.LOCAL_API_KEY,
# Local model name comes from per-call args or LOCAL_TEXT_MODEL —
# never the UI's OpenRouter picks, which a local server won't serve.
"model": model_override or config.LOCAL_TEXT_MODEL or config.TEXT_MODEL,
"usage": False, # local has no OpenRouter usage accounting
"local": True,
}
# Vision, or text-on-OpenRouter (default / fallback). A per-job override
# (user's UI model pick) wins over per-call and env defaults.
default_model = config.MODEL if has_images else config.TEXT_MODEL
if has_images:
model = _vision_model_override or model_override or config.MODEL
else:
model = _text_model_override or model_override or config.TEXT_MODEL
return {
"base_url": config.AI_BASE_URL,
"api_key": config.AI_API_KEY,
"model": _model_override or model_override or default_model,
"model": model,
"usage": True,
"local": False,
}
@@ -261,6 +375,21 @@ def _response_cost(response) -> Optional[float]:
return float(cost) if isinstance(cost, (int, float)) else None
def _reasoning_tokens(response) -> Optional[int]:
"""Hidden thinking tokens for this call (OpenRouter usage details).
Direct evidence of whether the reasoning knob is working - without it,
thinking burn can only be inferred from char counts vs the token cap."""
try:
dump = response.model_dump()
except Exception:
return None
usage = dump.get("usage") or {}
details = usage.get("completion_tokens_details") or {}
n = details.get("reasoning_tokens")
return int(n) if isinstance(n, (int, float)) else None
def _parse(raw: str) -> Optional[Dict[str, Any]]:
"""Parse model JSON, falling back to truncation repair."""
try:
@@ -280,6 +409,8 @@ def call_json(
model: Optional[str] = None,
usage_tracker: Optional[Any] = None,
usage_stage: str = "?",
reasoning_effort: Optional[str] = None,
reasoning_max_tokens: Optional[int] = None,
) -> Optional[Dict[str, Any]]:
"""
Send one chat completion expecting a JSON object back.
@@ -287,15 +418,23 @@ def call_json(
Uses the provider's JSON mode (response_format) so the model returns a bare
JSON object instead of prose/empty text, and recovers from max_tokens
truncation. images_b64: optional base64 JPEGs attached as high-detail image
parts. Returns the parsed dict, or None on a hard failure (caller degrades).
parts. reasoning_effort: optional OpenRouter reasoning knob ("low"/"medium"/
"high") - keeps thinking models from burning the output budget on hidden
reasoning. reasoning_max_tokens: optional hard thinking-token budget
(OpenRouter reasoning max_tokens -> Gemini thinking_budget); stronger than
effort, and takes precedence when both are given. Returns the parsed dict,
or None on a hard failure (caller degrades).
"""
has_images = bool(images_b64)
be = _resolve_backend(has_images, model)
seq = _next_seq()
cache_key = None
if config.LLM_CACHE:
cache_key = _cache_key(be["model"], system_prompt, user_text,
images_b64, max_tokens, json_mode=True)
images_b64, max_tokens, json_mode=True,
reasoning_effort=reasoning_effort,
reasoning_max_tokens=reasoning_max_tokens)
hit = _cache_get(cache_key)
if hit is not None:
_add_cached()
@@ -303,6 +442,8 @@ def call_json(
usage_tracker.record(
usage_stage, be["model"], cached=True, has_images=has_images
)
_log_call(seq, be, usage_stage, user_text, images_b64,
"", hit, None, cached=True)
return hit
content: List[Dict[str, Any]] = []
@@ -323,10 +464,21 @@ def call_json(
for attempt in range(2):
try:
client = get_client(be["base_url"], be["api_key"])
kwargs = dict(model=be["model"], messages=messages,
max_tokens=max_tokens, timeout=config.LLM_TIMEOUT)
kwargs: Dict[str, Any] = dict(model=be["model"], messages=messages,
max_tokens=max_tokens, timeout=config.LLM_TIMEOUT)
extra_body: Dict[str, Any] = {}
if be["usage"]:
kwargs["extra_body"] = {"usage": {"include": True}}
extra_body["usage"] = {"include": True}
# OpenRouter reasoning knob; only sent to cloud backends (local
# servers reject unknown fields). A hard thinking budget wins over
# the vaguer effort tier - OpenRouter treats them as exclusive.
if not be.get("local"):
if reasoning_max_tokens:
extra_body["reasoning"] = {"max_tokens": int(reasoning_max_tokens)}
elif reasoning_effort:
extra_body["reasoning"] = {"effort": reasoning_effort}
if extra_body:
kwargs["extra_body"] = extra_body
if use_json_mode:
kwargs["response_format"] = {"type": "json_object"}
response = client.chat.completions.create(**kwargs)
@@ -337,12 +489,28 @@ def call_json(
usage_tracker.record(
usage_stage, be["model"], usd=usd or 0.0, has_images=has_images
)
raw = _strip_fences(response.choices[0].message.content or "")
choice = response.choices[0]
finish_reason = getattr(choice, "finish_reason", None)
raw = _strip_fences(choice.message.content or "")
reasoning_tokens = _reasoning_tokens(response)
if finish_reason == "length":
# Hit max_tokens (thinking tokens included on reasoning models).
# Logged explicitly so silent truncation isn't mistaken for a
# parse problem; _parse below still salvages what it can.
print(f"[LLM] output hit max_tokens (finish_reason=length, "
f"{len(raw)}ch returned, "
f"thinking={reasoning_tokens}tk)")
parsed = _parse(raw)
if parsed is not None:
_record_model(be, has_images, fell_back)
if cache_key:
_cache_set(cache_key, parsed)
_log_call(seq, be, usage_stage, user_text, images_b64,
raw, parsed, usd, reasoning_tokens=reasoning_tokens)
if config.LLM_RAW_DUMP:
_dump_raw(seq, be, usage_stage, system_prompt, user_text,
images_b64, max_tokens, raw, parsed,
finish_reason=finish_reason)
return parsed
if attempt == 0:
print("[LLM] JSON parse error (retrying)")
@@ -362,7 +530,8 @@ def call_json(
_models["text_local"].add(be["model"])
_models["fallback_count"] += 1
be = {"base_url": config.AI_BASE_URL, "api_key": config.AI_API_KEY,
"model": config.TEXT_MODEL, "usage": True, "local": False}
"model": _text_model_override or config.TEXT_MODEL,
"usage": True, "local": False}
cache_key = None # don't cache fallback under the local-model key
fell_back = True
continue
+30 -7
View File
@@ -20,7 +20,7 @@ from fastapi.responses import HTMLResponse, JSONResponse, Response
from fastapi.staticfiles import StaticFiles
import backend.jobs
from backend import config
from backend import config, llm
from backend.jobs import PIPELINE_MODES, create_job, get_job, _set
from backend.pipeline.pdf_processor import render_page_jpeg
from backend.review.feedback import decision_to_label, write_label
@@ -35,6 +35,7 @@ _FRONTEND_DIR = os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", "
@app.get("/health")
def health():
return {"status": "ok", "model": config.MODEL,
"text_model": config.TEXT_MODEL,
"version": config.APP_VERSION,
"build": config.APP_BUILD,
"key_configured": bool(config.AI_API_KEY),
@@ -43,13 +44,15 @@ def health():
@app.get("/models")
def list_models():
"""Available OpenRouter models with per-1M-token pricing for the UI picker."""
from backend.models import fetch_models
"""Vision/text OpenRouter model lists with pricing for the UI dropdowns."""
from backend.models import fetch_models, split_vision_text
models = fetch_models()
if models is None:
raise HTTPException(status_code=502,
detail="Could not fetch the model list from OpenRouter")
return {"models": models, "default": config.MODEL, "default_text": config.TEXT_MODEL}
vision, text = split_vision_text(models)
return {"vision": vision, "text": text,
"defaults": {"vision": config.MODEL, "text": config.TEXT_MODEL}}
@app.get("/jobs/{job_id}/log")
@@ -72,7 +75,8 @@ async def check(
work_type: Optional[str] = Form(None),
text_local: bool = Form(False),
pipeline_mode: str = Form("classic"),
model: Optional[str] = Form(None),
vision_model: Optional[str] = Form(None),
text_model: Optional[str] = Form(None),
):
"""
Accept a PDF, start a background conflict check, and return a job_id
@@ -81,6 +85,9 @@ async def check(
Optional intake fields (project_name/address/occupancy/work_type) feed the
Stage 0 jurisdiction profile; anything left blank is derived from the cover
sheet.
vision_model / text_model override the configured defaults for this run
(vision always OpenRouter; text follows the OpenRouter vs hybrid choice).
"""
if not file.filename.lower().endswith(".pdf"):
raise HTTPException(status_code=400, detail="Please upload a PDF.")
@@ -106,9 +113,12 @@ async def check(
}.items()
if v and v.strip()
}
v_model = (vision_model or "").strip() or None
t_model = (text_model or "").strip() or None
job_id = create_job(tmp_path, source_filename=file.filename, email=email,
project_input=project_input, text_local=text_local,
pipeline_mode=pipeline_mode, model=model)
pipeline_mode=pipeline_mode, vision_model=v_model,
text_model=t_model)
return JSONResponse({
"job_id": job_id,
"status": "queued",
@@ -176,7 +186,20 @@ def save_review_decisions(job_id: str, payload: dict):
def _finalize_job(job_id: str, out_dir: str) -> None:
"""Background finalization: the ONE place the final report email may fire."""
try:
report = finalize_review(job_id, out_dir)
# Re-open the job's log tee + raw dump dir so the finalization LLM
# calls (clarification reruns, RFI drafting) land in job.log / llm_raw.
with backend.jobs.capture_job_output(job_id, out_dir):
print("\n=== Review finalization ===")
llm.reset_cost() # finalization-only cost attribution
report = finalize_review(job_id, out_dir)
cost = llm.get_cost()
backend.jobs._log_cost_summary({
"cost_usd": round(cost["usd"], 4),
"llm_calls": cost["calls"],
"cached_calls": cost.get("cached", 0),
"cost_by_stage": cost.get("by_stage", {}),
"models_used": cost.get("models", {}),
}, label="finalization")
except Exception as e:
try:
_set(job_id, status="finalization_error", error=str(e),
+27 -2
View File
@@ -3,11 +3,13 @@ models.py - Fetch the available OpenRouter model list with pricing (cached).
The /models endpoint is public (no API key needed). Results are normalized to
per-1M-token USD costs for display and cached in memory for an hour; callers
degrade gracefully when OpenRouter is unreachable.
degrade gracefully when OpenRouter is unreachable. Each entry also carries a
vision flag (accepts image input) so the UI can offer separate vision/text
model dropdowns.
"""
import time
from typing import List, Optional
from typing import List, Optional, Tuple
import httpx
@@ -25,6 +27,18 @@ def _per_mtok(rate) -> float:
return 0.0
def _is_vision(item: dict) -> bool:
"""True when the model accepts image input and produces text output."""
arch = item.get("architecture") or {}
inputs = arch.get("input_modalities") or []
outputs = arch.get("output_modalities") or []
# Legacy string form: "text+image->text"
modality = (arch.get("modality") or "").lower()
has_image_in = ("image" in inputs) or ("image" in modality.split("->")[0])
has_text_out = ("text" in outputs) or ("->text" in modality) or (not outputs and not modality)
return has_image_in and has_text_out
def _fetch_openrouter_models() -> Optional[List[dict]]:
"""Raw GET of the OpenRouter model list; None on any failure."""
try:
@@ -55,6 +69,7 @@ def fetch_models(force: bool = False) -> Optional[List[dict]]:
"prompt_usd_per_mtok": _per_mtok((item.get("pricing") or {}).get("prompt")),
"completion_usd_per_mtok": _per_mtok((item.get("pricing") or {}).get("completion")),
"context_length": item.get("context_length"),
"vision": _is_vision(item),
}
for item in data
if item.get("id")
@@ -63,3 +78,13 @@ def fetch_models(force: bool = False) -> Optional[List[dict]]:
_cache["models"] = models
_cache["at"] = time.time()
return models
def split_vision_text(models: List[dict]) -> Tuple[List[dict], List[dict]]:
"""Partition the normalized catalog into (vision, text) lists for the UI.
Every catalog model takes text in/out, so vision models appear in both
lists (same dicts, pricing included).
"""
vision = [m for m in models if m.get("vision")]
return vision, list(models)
+2
View File
@@ -54,6 +54,8 @@ def slim_clusters(clusters: List[Dict]) -> List[Dict]:
"location": c.get("location"),
"disciplines": c.get("disciplines"),
"kind": c.get("kind"),
**({"disputed_attributes": c["disputed_attributes"]}
if c.get("disputed_attributes") else {}),
"assertions": [slim_assertion(a) for a in c.get("assertions", [])],
}
for c in clusters
+12
View File
@@ -32,6 +32,16 @@ def _evidence_block(cluster: Dict) -> str:
f"{a.get('attribute','')} = {a.get('value','')} | "
f"\"{a.get('source_text','')}\""
)
disputes = cluster.get("disputed_attributes") or []
if disputes:
lines.append("")
for d in disputes:
lines.append(
"DISPUTED VALUE (possible extraction misread): "
f"attribute={d.get('attribute','')} "
f"values={' | '.join(d.get('values') or [])} "
f"(assertions {', '.join(d.get('assertion_ids') or [])})"
)
return "\n".join(lines)
@@ -83,6 +93,8 @@ def _check_one(cluster: Dict, page_to_b64: Dict[int, str]) -> List[Dict]:
user_text=user_text,
images_b64=_images_for(cluster, page_to_b64),
max_tokens=config.REASON_MAX_TOKENS,
reasoning_effort=config.EXTRACT_REASONING_EFFORT or None,
reasoning_max_tokens=config.EXTRACT_REASONING_MAX_TOKENS or None,
)
if isinstance(parsed, list):
candidates = parsed
+4
View File
@@ -27,6 +27,10 @@ def constructability_review(sheets: List[Dict], clusters: List[Dict],
"assertions": dumps(slim_sheets(sheets)),
"clusters": dumps(slim_clusters(clusters)),
"conflicts": dumps(conflicts),
"disputes": dumps([
d for cluster in clusters
for d in (cluster.get("disputed_attributes") or [])
]),
},
max_tokens=config.CONSTRUCT_MAX_TOKENS,
)
+57 -8
View File
@@ -64,7 +64,8 @@ def discipline_from_sheet_number(sheet_number: Optional[str]) -> Optional[str]:
return None
def _is_grounded(value: str, source_text: str, graphical_basis: str = "") -> bool:
def _is_grounded(value: str, source_text: str, graphical_basis: str = "",
page_text: Optional[str] = None) -> bool:
"""
Keep an object only if its primary value is supported by its source_text,
OR it is a graphical object (has graphical_basis with no text to quote).
@@ -72,6 +73,9 @@ def _is_grounded(value: str, source_text: str, graphical_basis: str = "") -> boo
- If graphical_basis is set and source_text is absent, the object is valid.
- If the value contains digits, every distinct digit-run must appear in
source_text (catches invented dimensions/counts/elevations).
- Rescue tier: when page_text (the deterministic text layer) is given,
digit-runs absent from source_text but present in the page text are
still grounded - vision quoted imperfectly but the value is real.
- If the value has no digits, require some alphabetic-token overlap.
"""
# Graphical objects (no readable text on sheet) are always allowed through.
@@ -85,7 +89,11 @@ def _is_grounded(value: str, source_text: str, graphical_basis: str = "") -> boo
val_digits = set(_DIGITS_RE.findall(value))
if val_digits:
src_digits = set(_DIGITS_RE.findall(source_text))
return val_digits.issubset(src_digits)
if val_digits.issubset(src_digits):
return True
if page_text:
return val_digits.issubset(set(_DIGITS_RE.findall(page_text)))
return False
# No digits: text-based grounding.
val_norm = re.sub(r"[^a-z0-9]+", " ", value.lower()).strip()
@@ -109,7 +117,24 @@ def _primary_value(obj: Dict) -> str:
or obj.get("name") or obj.get("tag") or "")
def _normalize_sheet(parsed: Dict, page_number: int) -> Dict:
def _grounding_stamp(value: str, source_text: str,
page_text: Optional[str]) -> Optional[str]:
"""\"text_layer\" when the object survived only via the text-layer rescue
tier (digits absent from source_text but present in the page text)."""
if not page_text:
return None
val_digits = set(_DIGITS_RE.findall(str(value)))
if not val_digits:
return None
if val_digits.issubset(set(_DIGITS_RE.findall(source_text))):
return None
if val_digits.issubset(set(_DIGITS_RE.findall(page_text))):
return "text_layer"
return None
def _normalize_sheet(parsed: Dict, page_number: int,
page_text: Optional[str] = None) -> Dict:
"""
Validate + clean one parsed sheet result, attaching page_number and ids.
@@ -138,6 +163,7 @@ def _normalize_sheet(parsed: Dict, page_number: int) -> Dict:
raw_objects = parsed.get("objects") or parsed.get("assertions") or []
clean: List[Dict] = []
dropped = 0
rescued = 0
for idx, obj in enumerate(raw_objects):
if not isinstance(obj, dict):
@@ -149,9 +175,13 @@ def _normalize_sheet(parsed: Dict, page_number: int) -> Dict:
# Derive a primary value for the grounding check
primary_val = _primary_value(obj)
if not _is_grounded(primary_val, source_text, graphical_basis):
if not _is_grounded(primary_val, source_text, graphical_basis,
page_text=page_text):
dropped += 1
continue
grounding = _grounding_stamp(primary_val, source_text, page_text)
if grounding:
rescued += 1
# --- location_key: new schema is richer; map to legacy shape + extras ---
lk = obj.get("location_key")
@@ -203,10 +233,13 @@ def _normalize_sheet(parsed: Dict, page_number: int) -> Dict:
"object_attributes": attrs,
"graphical_basis": graphical_basis or None,
"review_uses": obj.get("review_uses") or [],
**({"grounding": grounding} if grounding else {}),
})
if dropped:
print(f"[Extract] Page {page_number} ({sheet_number}): dropped {dropped} ungrounded object(s)")
if dropped or rescued:
print(f"[Extract] Page {page_number} ({sheet_number}): "
f"dropped {dropped} ungrounded object(s)"
+ (f", rescued {rescued} via text layer" if rescued else ""))
unresolved = parsed.get("unresolved_items") or []
@@ -223,8 +256,23 @@ def _normalize_sheet(parsed: Dict, page_number: int) -> Dict:
}
def _text_layer_block(page: Dict) -> str:
"""
The TEXT LAYER block appended to the extractor instruction at call sites
(NOT a template placeholder - render() silently leaves missing keys as
literals). Empty string when the page has no usable text layer.
"""
text = (page.get("text_layer") or "").strip()
if not text:
return ""
return ("\n\nTEXT LAYER (authoritative for alphanumeric content — trust it "
"over the image for numbers, tags, and note text):\n"
+ text[:config.TEXT_LAYER_MAX_CHARS])
def _extract_one(page: Dict, sheet_hint: str = "") -> Dict:
user_text = EXTRACTOR_USER_INSTRUCTION.replace("{sheet_hint}", sheet_hint)
user_text = (EXTRACTOR_USER_INSTRUCTION.replace("{sheet_hint}", sheet_hint)
+ _text_layer_block(page))
parsed = call_json(
system_prompt=EXTRACTOR_SYSTEM_PROMPT,
user_text=user_text,
@@ -241,7 +289,8 @@ def _extract_one(page: Dict, sheet_hint: str = "") -> Dict:
"scale": None,
"assertions": [],
}
return _normalize_sheet(parsed, page["page_number"])
return _normalize_sheet(parsed, page["page_number"],
page_text=page.get("text_layer"))
def extract_assertions(pages: List[Dict], on_progress=None) -> List[Dict]:
+39 -1
View File
@@ -27,12 +27,14 @@ from typing import Dict, Optional, Callable
from backend.pipeline.pdf_processor import convert_pdf_to_images
from backend.pipeline.extractor import extract_assertions
from backend.text_layer import attach_text_layers, coverage_gaps
from backend.pipeline.sheet_index import classify_sheets, derive_project_meta_from_cover
from backend.pipeline.jurisdiction import run_jurisdiction
from backend.pipeline.normalizer import normalize_assertions, build_project_intelligence
from backend.pipeline.clusterer import cluster_by_location
from backend.pipeline.llm_clusterer import cluster_by_location_llm
from backend import config
from backend.agents.disputes import annotate_clusters
from backend.pipeline.conflict_checker import check_conflicts
from backend.pipeline.qaqc_review import senior_review
from backend.pipeline.code_review import code_review
@@ -42,7 +44,9 @@ from backend.pipeline.risk import score_and_prioritize
from backend.pipeline.rfi import generate_rfis
from backend.pipeline.report import build_report, to_markdown
from backend.pipeline._stage import validate_issue
from backend.llm import reset_cost, get_cost, set_stage, set_text_backend
from backend.llm import (
reset_cost, get_cost, set_stage, set_text_backend, set_model_overrides,
)
def run_pipeline(
@@ -52,6 +56,8 @@ def run_pipeline(
project_input: Optional[Dict] = None,
source_name: Optional[str] = None,
text_local: bool = False,
vision_model: Optional[str] = None,
text_model: Optional[str] = None,
) -> Dict:
"""
Run the full QAQC pipeline on one PDF and return the report dict.
@@ -59,6 +65,9 @@ def run_pipeline(
project_input: optional intake fields (project_name, address, occupancy,
work_type). Cover-sheet-derived values fill any gaps; intake fields win.
vision_model / text_model: optional per-run OpenRouter (or local text)
model overrides from the UI. Blank/None keeps config defaults.
If out_dir is given, writes conflicts.json, report.md, and the intermediate
artifacts (assertions.json, clusters.json, and one json per QAQC stage).
"""
@@ -70,12 +79,37 @@ def run_pipeline(
reset_cost()
set_text_backend(text_local)
set_model_overrides(vision_model, text_model)
if vision_model or text_model:
print(f"[Runner] model overrides: vision={vision_model or '(default)'} "
f"text={text_model or '(default)'}")
try:
return _run_stages(
pdf_path, out_dir, stage, project_input, source_name, text_local,
)
finally:
# Don't leak per-run picks into a later overlapping/CLI call.
set_model_overrides(None, None)
set_text_backend(False)
def _run_stages(
pdf_path: str,
out_dir: Optional[str],
stage: Callable[[str], None],
project_input: Optional[Dict],
source_name: Optional[str],
text_local: bool,
) -> Dict:
stage("PDF -> images")
pages = convert_pdf_to_images(pdf_path)
text_dir = os.path.join(out_dir, "text") if out_dir else None
attach_text_layers(pdf_path, pages, text_dir=text_dir)
stage("Extract assertions")
sheets = extract_assertions(pages)
coverage_gaps(pages, sheets) # classic: log-only recall signal
stage("Classify sheet index")
sheet_index = classify_sheets(sheets)
@@ -100,6 +134,10 @@ def run_pipeline(
else:
clusters = cluster_by_location(sheets)
disputed_count = annotate_clusters(clusters)
if disputed_count:
print(f"[Cluster] {disputed_count} cluster(s) carry disputed extracted values")
stage("Reason over clusters (conflicts)")
conflicts = check_conflicts(clusters, pages)
+34 -1
View File
@@ -230,6 +230,7 @@ Rules you must never break:
- Every object must include source_text copied verbatim from the sheet whenever text is available.
- If the object is graphical and has no text, describe it visually and mark confidence low or medium.
- Preserve tags, marks, room numbers, sheet numbers, detail references, and abbreviations exactly as shown.
TEXT LAYER GROUNDING: when a TEXT LAYER block is present in the user message, it is the sheet's deterministic PDF text layer and is authoritative for alphanumeric content (counts, dimensions, member tags, note text). Trust it over your reading of the image for numbers, tags, and note text; quote source_text from it verbatim. Use the image for geometry, symbols, linework, and anything absent from the text layer.
- Use null when information is not determinable.
- Keep objects atomic.
- Use plain ASCII only.
@@ -410,6 +411,9 @@ What is NOT a conflict:
- Anything not supported with drawing evidence.
Be conservative:
- Only flag genuine disagreements.
- When a value is marked DISPUTED (possible extraction misread), verify it against
the sheet images before relying on either reading; if the images do not resolve
it, do not assert a conflict from one reading alone.
- A clean cluster with no contradiction must return an empty conflicts array.
- missing_element requires evidence that another discipline would reasonably be expected to show the missing item.
For each conflict:
@@ -447,6 +451,28 @@ Clustered assertions (evidence):
{evidence}"""
# ---------------------------------------------------------------------------
# Wave 5b - evidence verification (vision fact-check of cited sheet text)
# ---------------------------------------------------------------------------
VERIFY_SYSTEM_PROMPT = """You are a meticulous construction document checker verifying machine-extracted evidence against the actual drawing sheet images.
For each evidence item you are given the sheet it was extracted from and the verbatim text the extractor claims appears there.
Judge each item against the images:
- confirmed: the text (or an obvious equivalent) appears on the cited sheet and means what the finding claims.
- corrected: the sheet shows a DIFFERENT value than the extracted text. Give the actual verbatim text.
- not_found: nothing like the extracted text appears on the cited sheet.
Be strict about numbers, quantities, and member sizes: "(2) 2x6" and "(5) 2x6" are different values. HSS16x4 and HSS16x16 are different values.
Use plain ASCII only.
Respond only with valid JSON."""
VERIFY_USER_INSTRUCTION = """Verify this finding's evidence against the attached sheet images.
Respond ONLY with a valid JSON object - no markdown fences, no explanation:
{ "verdicts": [ { "sheet": "string", "source_text": "the evidence text judged", "verdict": "confirmed | corrected | not_found", "actual_text": "verbatim sheet text when corrected, else null", "notes": "string or null" } ] }
Finding: {finding}
TEXT LAYER (deterministic page text extracted from the PDF - an oracle for alphanumeric content such as counts, dimensions, and member tags; when it disagrees with the extracted evidence, trust it and cite it as actual_text):
{text_layer}"""
# ---------------------------------------------------------------------------
# Stage 6 - senior architect full-set QAQC review (NOT WIRED YET)
# ---------------------------------------------------------------------------
@@ -545,6 +571,12 @@ Flag:
Rules:
- Only flag issues supported by drawing evidence.
- Be specific about the location and why it is a constructability risk.
- Assertions are machine-extracted from sheet images and may contain misread values,
especially quantities and member sizes (e.g. "(2) 2x6" vs "(5) 2x6").
- When the cluster lists disputed_attributes, or two evidence items disagree on a
numeric value, do NOT assert a buildability conclusion from one reading. Report the
ambiguity itself (category "detail_gap", confidence "low") and state that the value
needs verification against the sheet.
- Use plain ASCII only.
Respond only with valid JSON."""
@@ -554,7 +586,8 @@ Respond ONLY with a valid JSON object - no markdown fences, no explanation:
If no constructability issues are found, return: { "issues": [] }
Extracted assertions: {assertions}
Clusters: {clusters}
Cross-discipline conflicts already found: {conflicts}"""
Cross-discipline conflicts already found: {conflicts}
Disputed extracted values in this cluster (possible vision misreads - treat as unverified): {disputes}"""
# ---------------------------------------------------------------------------
+1 -1
View File
@@ -216,7 +216,7 @@ def finalize_review(job_id: str, out_dir: str) -> dict:
rfis = _draft_rfis(kept)
report["validated_issues"] = kept
report["suppressed_issues"] = suppressed
report["suppressed_issues"] = (report.get("suppressed_issues") or []) + suppressed
report["rfis"] = rfis
summary = report.setdefault("summary", {})
summary["agent_status"] = "complete"
+230
View File
@@ -0,0 +1,230 @@
"""
text_layer.py - deterministic PDF text-layer extraction (PyMuPDF, no LLM).
Most CAD-produced drawing sets carry a real vector text layer. We extract it
once per job and feed it to the extractor (grounding), the grounding guard
(rescue tier), and the wave-5b verifier (text oracle + high-DPI evidence
crops). Pages below TEXT_LAYER_MIN_CHARS of text are treated as having no
text layer (scanned/raster sheets stay vision-only).
If PyMuPDF is unavailable the module degrades gracefully: every public
function returns empty/None, equivalent to TEXT_LAYER_ENABLED=false.
"""
import re
from typing import Dict, List, Optional, Tuple
from backend import config
try: # PyMuPDF >= 1.24 prefers the pymupdf name; fitz works everywhere.
import pymupdf as fitz
except ImportError: # pragma: no cover - older PyMuPDF
try:
import fitz
except ImportError: # pragma: no cover - PyMuPDF not installed
fitz = None
_warned_unavailable = False
# Word token normalization for evidence matching: lowercase alphanumeric only.
_TOKEN_RE = re.compile(r"[^a-z0-9]+")
# Fuzzy match floor: fraction of needle tokens that must align with the page's
# word sequence for a bbox to count as a confident evidence location.
_FUZZY_MIN_RATIO = 0.6
def _fitz_or_none():
"""Return the fitz module, logging once if PyMuPDF is missing."""
global _warned_unavailable
if fitz is None and not _warned_unavailable:
print("[TextLayer] PyMuPDF not available - text-layer grounding disabled")
_warned_unavailable = True
return fitz
def extract_text_layers(pdf_path: str) -> Dict[int, Dict]:
"""
Extract the text layer of every page. Returns {1-based page_number:
{"text": str, "words": [{"text", "bbox": (x0,y0,x1,y1)}, ...],
"has_text_layer": bool}}. Returns {} when disabled or unavailable.
"""
if not config.TEXT_LAYER_ENABLED:
return {}
f = _fitz_or_none()
if f is None:
return {}
try:
doc = f.open(pdf_path)
except Exception as exc:
print(f"[TextLayer] could not open {pdf_path}: {exc}")
return {}
layers: Dict[int, Dict] = {}
try:
for index in range(doc.page_count):
page = doc[index]
text = page.get_text("text") or ""
words = [
{"text": w[4], "bbox": (w[0], w[1], w[2], w[3])}
for w in (page.get_text("words") or [])
]
has_text_layer = len(text.strip()) >= config.TEXT_LAYER_MIN_CHARS
if not has_text_layer:
print(f"[TextLayer] Page {index + 1}: {len(text.strip())} chars "
f"(< TEXT_LAYER_MIN_CHARS={config.TEXT_LAYER_MIN_CHARS}) - "
f"vision-only")
layers[index + 1] = {
"text": text,
"words": words,
"has_text_layer": has_text_layer,
}
finally:
doc.close()
return layers
def attach_text_layers(
pdf_path: str,
pages: List[Dict],
text_dir: Optional[str] = None,
) -> Dict[int, List[Dict]]:
"""
Attach page["text_layer"] (text or None) to each converted page dict and
return the runner-local {page_number: words} map (kept off page dicts -
those get serialized). When text_dir is set, dump one .txt per page there
(plain file writes; ProjectMemory is a closed registry).
"""
layers = extract_text_layers(pdf_path)
page_words: Dict[int, List[Dict]] = {}
for page in pages:
layer = layers.get(page["page_number"]) or {}
page["text_layer"] = layer.get("text") if layer.get("has_text_layer") else None
page_words[page["page_number"]] = layer.get("words") or []
if text_dir and layers:
import os
os.makedirs(text_dir, exist_ok=True)
for page_number, layer in layers.items():
if not layer.get("has_text_layer"):
continue
with open(os.path.join(text_dir, f"page-{page_number:03d}.txt"),
"w", encoding="utf-8") as fh:
fh.write(layer.get("text") or "")
return page_words
def _tokens(text: str) -> List[str]:
return [t for t in _TOKEN_RE.split(text.lower()) if t]
def _union_bbox(boxes: List[Tuple[float, float, float, float]]):
return (
min(b[0] for b in boxes),
min(b[1] for b in boxes),
max(b[2] for b in boxes),
max(b[3] for b in boxes),
)
def find_evidence_bbox(
words: List[Dict],
needle: str,
) -> Optional[Tuple[float, float, float, float]]:
"""
Best-effort fuzzy substring match of an evidence source_text against the
page's word sequence. Returns the union bbox of the matched words, or
None when nothing aligns confidently.
Exact contiguous token runs win; otherwise the best-scoring window with
>= _FUZZY_MIN_RATIO token alignment is accepted (vision quotes imperfectly
but the value is real page text).
"""
if not words or not needle:
return None
needle_tokens = _tokens(str(needle))
if not needle_tokens:
return None
page_tokens = [_tokens(w.get("text") or "") for w in words]
# Flatten multi-token words, remembering which word each token came from.
flat: List[Tuple[str, int]] = []
for word_index, parts in enumerate(page_tokens):
for part in parts:
flat.append((part, word_index))
if not flat:
return None
n = len(needle_tokens)
best_span = None
best_score = 0.0
for start in range(0, len(flat)):
window = flat[start:start + n]
if not window:
break
score = sum(1 for i, tok in enumerate(needle_tokens)
if i < len(window) and window[i][0] == tok) / n
if score > best_score:
best_score = score
best_span = window
if best_score == 1.0:
break
if best_span is None or best_score < _FUZZY_MIN_RATIO:
return None
word_indexes = {word_index for _, word_index in best_span}
return _union_bbox([words[i]["bbox"] for i in sorted(word_indexes)])
def render_crop(
pdf_path: str,
page_number: int,
bbox: Tuple[float, float, float, float],
dpi: Optional[int] = None,
margin_pts: Optional[float] = None,
) -> Optional[bytes]:
"""
Render a clip of one page around bbox (+ margin, clamped to the page) at
the given DPI and return JPEG bytes, or None on any failure.
"""
f = _fitz_or_none()
if f is None:
return None
dpi = dpi or config.VERIFY_CROP_DPI
margin_pts = config.VERIFY_CROP_MARGIN_PTS if margin_pts is None else margin_pts
try:
doc = f.open(pdf_path)
try:
page = doc[page_number - 1]
rect = f.Rect(
bbox[0] - margin_pts,
bbox[1] - margin_pts,
bbox[2] + margin_pts,
bbox[3] + margin_pts,
) & page.rect
if rect.is_empty:
return None
pix = page.get_pixmap(clip=rect, dpi=dpi)
return pix.tobytes("jpeg")
finally:
doc.close()
except Exception as exc:
print(f"[TextLayer] render_crop failed on page {page_number}: {exc}")
return None
def coverage_gaps(pages: List[Dict], sheets: List[Dict]) -> List[int]:
"""
Page numbers that have a text layer but whose extraction failed or
returned 0 objects - the silent extraction-loss signal. Logs one
[TextLayer] line per gap.
"""
by_page = {s.get("page_number"): s for s in sheets or []}
gaps: List[int] = []
for page in pages:
text = page.get("text_layer")
if not text:
continue
sheet = by_page.get(page["page_number"])
extracted = len(sheet.get("assertions") or []) if sheet else 0
if extracted == 0:
gaps.append(page["page_number"])
print(f"[TextLayer] Page {page['page_number']}: text layer present "
f"({len(text)} chars) but no objects extracted — possible "
f"extraction gap")
return gaps
+187
View File
@@ -0,0 +1,187 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<title>Conflict Checker — How Your Plans Get Reviewed</title>
<style>
* { box-sizing: border-box; margin: 0; padding: 0; }
body {
font-family: "Segoe UI", "Helvetica Neue", Arial, sans-serif;
background: #f4f7fb;
color: #1f2d3d;
width: 1280px;
padding: 40px 48px;
}
header { text-align: center; margin-bottom: 10px; }
h1 { font-size: 34px; color: #123c6e; letter-spacing: 0.5px; }
.subtitle { font-size: 17px; color: #5a6b7f; margin-top: 8px; }
.blueprint {
background: #ffffff;
border: 2px solid #d5e3f2;
border-radius: 18px;
padding: 32px 36px;
margin-top: 24px;
background-image:
linear-gradient(#eef4fb 1px, transparent 1px),
linear-gradient(90deg, #eef4fb 1px, transparent 1px);
background-size: 28px 28px;
}
.row { display: flex; justify-content: center; align-items: stretch; gap: 0; }
.row + .connector-down { margin: 0; }
.card {
background: #ffffff;
border-radius: 14px;
border: 2px solid #cfdcec;
box-shadow: 0 3px 8px rgba(18,60,110,0.08);
width: 250px;
padding: 16px 16px 14px;
position: relative;
flex-shrink: 0;
}
.card .num {
position: absolute; top: -16px; left: -14px;
width: 36px; height: 36px; border-radius: 50%;
background: #123c6e; color: #fff;
font-weight: 700; font-size: 18px;
display: flex; align-items: center; justify-content: center;
box-shadow: 0 2px 5px rgba(0,0,0,0.2);
}
.card .icon { font-size: 34px; text-align: center; margin: 4px 0 6px; }
.card h2 { font-size: 17px; color: #123c6e; text-align: center; margin-bottom: 6px; }
.card p { font-size: 13px; line-height: 1.4; color: #42536a; text-align: center; }
.card.scan { border-color: #7fb3e0; background: #f0f7ff; }
.card.read { border-color: #7fb3e0; background: #f0f7ff; }
.card.lib { border-color: #8fd0a8; background: #f1faf4; }
.card.link { border-color: #8fd0a8; background: #f1faf4; }
.card.det { border-color: #f2b879; background: #fff8ef; }
.card.spec { border-color: #f2b879; background: #fff8ef; }
.card.brain { border-color: #c39bd3; background: #f9f3fc; }
.card.human { border-color: #e58f8f; background: #fdf1f1; }
.arrow {
display: flex; align-items: center; justify-content: center;
color: #123c6e; font-size: 30px; font-weight: bold;
width: 44px; flex-shrink: 0;
}
.connector-down {
text-align: center; color: #123c6e; font-size: 30px;
font-weight: bold; line-height: 1; padding: 6px 0;
}
.finish {
margin: 22px auto 0;
width: 560px;
background: #123c6e; color: #ffffff;
border-radius: 14px; padding: 18px 24px; text-align: center;
box-shadow: 0 4px 10px rgba(18,60,110,0.3);
}
.finish .big { font-size: 20px; font-weight: 700; }
.finish .small { font-size: 14px; margin-top: 6px; color: #cfe0f4; }
footer {
margin-top: 26px; text-align: center;
font-size: 13px; color: #7a8aa0;
}
.note {
margin: 18px auto 0; width: 900px; font-size: 13.5px; color: #42536a;
background: #ffffff; border-left: 4px solid #7fb3e0; border-radius: 6px;
padding: 10px 16px; line-height: 1.5;
}
</style>
</head>
<body>
<header>
<h1>🔍 CONFLICT CHECKER</h1>
<div class="subtitle">How your construction plans get reviewed — a team of AI assistants, each with one job, passing notes down the line.</div>
</header>
<div class="blueprint">
<!-- Row 1 -->
<div class="row">
<div class="card scan">
<div class="num">1</div>
<div class="icon">📄</div>
<h2>The Scanner</h2>
<p>Turns every page of your PDF blueprints into a picture the AI can read.</p>
</div>
<div class="arrow"></div>
<div class="card read">
<div class="num">2</div>
<div class="icon">👓</div>
<h2>The Readers</h2>
<p>One assistant per page, all working at once. Each writes down every fact: dimensions, notes, materials, callouts.</p>
</div>
<div class="arrow"></div>
<div class="card lib">
<div class="num">3</div>
<div class="icon">📚</div>
<h2>Librarian &amp; Code Scout</h2>
<p>Builds the table of contents (electrical, plumbing, structural…) and figures out <b>where</b> the project is, so the right building codes apply.</p>
</div>
<div class="arrow"></div>
<div class="card link">
<div class="num">4</div>
<div class="icon">🔗</div>
<h2>The Connector</h2>
<p>Connects the dots across sheets — "this water heater on the plumbing sheet is the same one on the electrical sheet" — and sorts facts into topic piles.</p>
</div>
</div>
<div class="connector-down"></div>
<!-- Row 2 -->
<div class="row">
<div class="card det">
<div class="num">5</div>
<div class="icon">🕵️</div>
<h2>The Detectives</h2>
<p>One per topic pile. Compares sheets that should agree and hunts for contradictions: "wall shown here, but not on the structural plan."</p>
</div>
<div class="arrow"></div>
<div class="card spec">
<div class="num">6</div>
<div class="icon">👷</div>
<h2>The Specialists</h2>
<p>Three experts at once: a <b>code inspector</b>, a veteran <b>builder</b> ("can this actually be built?"), and a <b>checklist keeper</b> ("is anything missing?").</p>
</div>
<div class="arrow"></div>
<div class="card brain">
<div class="num">7</div>
<div class="icon">🧠</div>
<h2>The Brain</h2>
<p>The senior reviewer. Collects every finding, merges duplicates, discards weak ones, and ranks the rest by how much trouble they'd cause.</p>
</div>
<div class="arrow"></div>
<div class="card human">
<div class="num">8</div>
<div class="icon"></div>
<h2>Human Review</h2>
<p>The important and uncertain findings land on <b>your</b> desk. You confirm, reject, or mark unsure — nothing goes out unapproved.</p>
</div>
</div>
<div class="connector-down"></div>
<div class="finish">
<div class="big">📋 Final Report + ✉️ Draft RFIs</div>
<div class="small">A prioritized list of every problem found — plus ready-to-send "please clarify" letters (Requests For Information) for the design team.</div>
</div>
<div class="note">
<b>Good to know:</b> everyone shares one notebook, so each step builds on the last.
If one page can't be read, the team keeps going and that page is flagged as a gap
instead of stopping the whole review. Every finding links back to the sheet it came from.
</div>
</div>
<footer>Conflict Checker · conchecker.scoutitsystems.com · Review your plans before they cost you money in the field.</footer>
</body>
</html>
Binary file not shown.

After

Width:  |  Height:  |  Size: 184 KiB

+174
View File
@@ -0,0 +1,174 @@
# Conflict Checker — How It Works (Plain Language)
**What it does:** You upload a set of construction drawings (a PDF of blueprints).
A team of AI assistants reads every page, compares everything against everything
else, and hands you a list of problems — contradictions, code violations, missing
information, and things that would be hard to build — before they cost you money
in the field.
Think of it like hiring a room full of specialist consultants to review your
plans overnight. Each one has a specific job, they pass their notes down the
table, and a senior reviewer at the end sorts it all into one clean report.
---
## The Big Picture (one sentence per step)
```
YOUR PDF OF BLUEPRINTS
|
v
+----------------------------------------------------------+
| 0. SCANNER |
| Turns every PDF page into a picture the AI can read |
+----------------------------------------------------------+
|
v
+----------------------------------------------------------+
| 1. READERS (one assistant per page, all at once) |
| Reads each sheet and writes down every fact: |
| dimensions, notes, materials, room names, callouts |
+----------------------------------------------------------+
|
v
+----------------------------------------------------------+
| 2. LIBRARIAN + LOCAL-CODE SCOUT (work side by side) |
| Librarian: builds the table of contents — which |
| sheets exist (electrical, plumbing, structural...) |
| Scout: figures out WHERE the project is, so we know |
| which building codes apply |
+----------------------------------------------------------+
|
v
+----------------------------------------------------------+
| 3. CONNECTOR |
| Connects the dots across sheets — e.g. "the water |
| heater on the plumbing sheet is the same one on the |
| electrical sheet" — and groups related facts into |
| topic piles (clusters) |
+----------------------------------------------------------+
|
v
+----------------------------------------------------------+
| 4. CONFLICT DETECTIVES (one per topic pile) |
| Compares sheets that should agree and looks for |
| contradictions: "Wall shown here on A-201 but not |
| on S-101", "Pipe runs through the duct" |
+----------------------------------------------------------+
|
v
+----------------------------------------------------------+
| 5. THREE SPECIALISTS (work side by side) |
| * Code Inspector — does anything break the local |
| building code? |
| * Builder — can this actually be built as |
| drawn? (access, clearances, sequencing) |
| * Completeness Checker — is anything MISSING from |
| the set? (sheets, schedules, required details) |
+----------------------------------------------------------+
|
v
+----------------------------------------------------------+
| 6. THE BRAIN (senior reviewer) |
| Collects EVERY finding from everyone, merges the |
| duplicates, throws out the weak ones, and ranks the |
| rest by how much trouble they'd cause |
+----------------------------------------------------------+
|
v
+----------------------------------------------------------+
| 7. HUMAN REVIEW GATE |
| The important/uncertain findings are queued for a |
| real person to Confirm / Reject / mark Unsure |
+----------------------------------------------------------+
|
v
+----------------------------------------------------------+
| 8. LETTER WRITER |
| Drafts a formal RFI (Request For Information — the |
| official "please clarify this" letter) for each |
| confirmed issue, ready to send to the design team |
+----------------------------------------------------------+
|
v
FINAL REPORT + DRAFT RFIs
```
---
## Who's Who (the "agents")
| # | Name | Analogy | What it actually does |
|---|------|---------|-----------------------|
| 0 | PDF Scanner | Photocopier | Converts each PDF page into an image the AI can "see" |
| 1 | Sheet Extractor | Speed-reader | Reads one page, writes structured notes (every page gets its own reader, in parallel) |
| 2 | Sheet Indexer | Librarian | Builds the table of contents of the drawing set |
| 2 | Jurisdiction Scout | Local guide | Identifies the project's location so the right building codes are used |
| 3 | Linker | Connector | Groups related facts from different sheets into topic clusters |
| 4 | Conflict Critic | Detective | Examines each cluster for contradictions between disciplines |
| 5 | Code Agent | Code inspector | Flags building-code violations, using the jurisdiction from step 2 |
| 5 | Constructability Agent | Veteran builder | Flags things that are drawn fine but can't be built practically |
| 5 | Completeness Agent | Checklist keeper | Flags missing sheets, missing details, gaps in the set |
| 6 | Brain | Chief estimator | Deduplicates, judges, and prioritizes all findings |
| 7 | Review Gate | Your desk | Presents the findings a human should approve before anything goes out |
| 8 | RFI Writer | Secretary | Writes the formal clarification letters for confirmed issues |
Everything the assistants learn is kept in a shared notebook (the "project
memory"), so each step builds on the last. If one reader fails on one page, the
rest of the team keeps going — that page is noted as a gap instead of crashing
the whole review.
---
## Where Improvements Could Be Made
### 1. Coverage — "make sure every page actually got read"
- Today, if a Reader fails on a page (the AI's answer gets cut off or comes back
garbled), that page quietly disappears from everything downstream. Worse, the
Completeness Checker can then report the sheet as "missing from the set" when
really it was there but unread — a false alarm.
- **Improvement:** retry failed pages with a backup model, and clearly separate
"sheet doesn't exist" from "sheet couldn't be read" in the report.
### 2. Speed — "the team waits in line more than it needs to"
- The steps run strictly one after another, but some could start earlier. The
Jurisdiction Scout only needs the cover page — it could run while the other
Readers are still working. The Letter Writer could start on high-confidence
findings instead of waiting for all human review.
- **Improvement:** overlap independent steps; start drafting letters for
confirmed/high-confidence findings sooner.
### 3. Cost — "smarter reading, fewer wasted words"
- Every page is read by a large, expensive AI model, and that model's
"thinking time" counts against its answer budget — we've seen it spend its
whole budget thinking and return a cut-off answer.
- **Improvement:** use cheaper models for simple pages (schedules, title
sheets), save the expensive model for dense drawings; keep tuning the
thinking budget knobs; reuse cached answers when the same plan set is
re-run.
### 4. Smarter grouping — "better piles, better detective work"
- The Connector caps how many topic piles it keeps (a fixed limit), so on big
sets some connections may never be made. The Detectives only see one pile at
a time, so a contradiction spanning two piles can slip through.
- **Improvement:** revisit the pile limit, and let the Brain (or a second-pass
Detective) look for conflicts that span multiple piles.
### 5. Human time — "review less, but review what matters"
- Today the review queue is built from rules about severity and confidence.
- **Improvement:** learn from your past Confirm/Reject decisions to sort the
queue better — the system already records your feedback, so it can get
smarter over time about what actually needs your eyes.
### 6. Trust — "show the receipts"
- Findings carry evidence, but a non-technical reader can't easily see *where
on the drawing* the problem is.
- **Improvement:** attach a cropped image snippet of the exact spot on the
sheet to each finding, so anyone can verify it in seconds.
---
*Technical reference for the curious: the pipeline lives in
`backend/agents/runner.py` (the waves above are the "Agent wave N" stages), the
team's shared notebook is `backend/agents/memory.py`, and the review queue is
`backend/review/gate.py` + `backend/review/finalizer.py`.*
@@ -0,0 +1,166 @@
# Text-Layer Grounding — Design Spec
**Date:** 2026-08-12 · **Branch:** `agent-mode` · **Status:** approved by user (2026-08-12)
## Problem
The pipeline is vision-only for extraction, but most CAD-produced drawing sets
carry a real vector text layer. Two worst documented failure modes are text
problems being solved with pixels:
1. **Wave-1 text misreads propagate immutably** — e.g. job `959e16407573`:
vision read "(2) 2x6 STUD PACK" where the sheet says "(5)"; text-only
downstream specialists treated the misread as ground truth → confident
false-positive findings.
2. **Silent extraction loss** — failed/under-extracted pages are invisible
(job `475a6f184dd1`: 42% extraction loss), producing false
`missing_expected_sheets` warnings and missed conflicts.
Priority (user, 2026-08-12): reduce false positives **and** missed items;
more accurate conflicts.
## Approach
Extract the PDF text layer deterministically (PyMuPDF) once per job, and make
it a first-class citizen at three points: extractor grounding, the grounding
guard, and the wave-5b verifier (as text oracle + high-DPI evidence crops).
Inspired by `hamzaabduljabbar/construction-drawing-analyzer` (patterns only —
its license is source-available/no-resale; all code here is original).
## Components
### 1. New module `backend/text_layer.py` (deterministic, no LLM)
- `extract_text_layers(pdf_path) -> Dict[int, dict]` — per 1-based page:
`{"text": str, "words": [{"text", "bbox": (x0,y0,x1,y1)}, ...],
"has_text_layer": bool}`. Pages with < `TEXT_LAYER_MIN_CHARS` of text are
`has_text_layer=False` (scanned/raster sheets stay vision-only; logged).
- `find_evidence_bbox(words, needle) -> bbox | None` — best-effort fuzzy
substring match of an evidence `source_text` against word sequence; returns
union rect of matched words.
- `render_crop(pdf_path, page_number, bbox, dpi, margin_pts) -> bytes`
PyMuPDF `page.get_pixmap(clip=rect, dpi=dpi)` → JPEG bytes.
Both runners call `extract_text_layers` right after `convert_pdf_to_images`
and attach `page["text_layer"] = <text or None>` to each page dict. Word
positions stay in a separate `page_words: Dict[int, list]` runner-local map
(not attached to page dicts — they get serialized).
### 2. Extractor grounding (both pipelines)
- Static paragraph added to `_EXTRACTOR_SYSTEM_TEMPLATE` in
`backend/prompts.py` (no new placeholder): when a TEXT LAYER block is
present in the user message it is **authoritative for alphanumeric content**
(counts, dimensions, member tags, notes); the image is for geometry,
symbols, linework, and anything absent from the text layer.
- Text-layer content is **appended programmatically** at each extractor call
site (classic `extractor.py::_extract_one`, agent
`extractors.py::SheetExtractorAgent.run`) — NOT a new `{placeholder}` in the
shared template (two-render-path trap: `render()` silently leaves missing
keys as literals). Block capped at `TEXT_LAYER_MAX_CHARS`.
Format: `\n\nTEXT LAYER (authoritative for alphanumeric content — trust it
over the image for numbers, tags, and note text):\n<text>`
### 3. Grounding guard rescue tier (`pipeline/extractor.py::_normalize_sheet`)
Current guard drops an object when its primary value's digit-runs aren't in
its own `source_text`. New tier, only when a text layer exists for the page:
- digits ⊆ source_text → keep (unchanged)
- digits ⊆ page text layer but ⊄ source_text → keep, stamp
`grounding: "text_layer"` on the assertion (recall rescue — vision quoted
imperfectly but the value is real page text)
- otherwise → drop (unchanged)
`_is_grounded` gains an optional `page_text` param; existing callers/tests
unaffected. Dropped/ rescued counts logged per page.
### 4. Verifier: text oracle + high-DPI crops (wave 5b)
Wherever verify scopes are built (agent runner confirmed; classic runner to be
checked — integrate at both if present):
- Scope payload gains `text_layer_excerpt`: concatenated text of the finding's
cited sheets, capped at `VERIFY_TEXT_MAX_CHARS`. `VERIFY_USER_INSTRUCTION`
gains a `{text_layer}` placeholder with instructions to treat it as
deterministic page text (verdicts may cite it as `actual_text`). **Both
render sites** (agent verifier + any classic-path render) must substitute it
— grep the template name across `backend/agents/` and `backend/pipeline/`.
- When `VERIFY_HI_DPI_CROPS` and the page has words: for each evidence item,
`find_evidence_bbox` on the cited page's words; on hit, `render_crop` at
`VERIFY_CROP_DPI` with margin → crop images replace full-page images (up to
`AGENT_CONFLICT_MAX_IMAGES`). On any miss/failure → fall back to the current
full-page image. Zero-resolved-images ⇒ scope skipped (I2 guard preserved).
### 5. Coverage signal (recall)
After extraction in both runners: for each page with `has_text_layer=True`
whose extraction failed or returned 0 objects, log
`[TextLayer] Page N: text layer present (M chars) but no objects extracted —
possible extraction gap` and add the page to the existing gap-finding path
(agent: `orchestrator.stats.failed_scopes`-style finding; classic: log only).
## Config knobs (`backend/config.py`, env-overridable, documented in `.env.example`)
| Key | Default | Effect |
|-----|---------|--------|
| `TEXT_LAYER_ENABLED` | `true` | Master switch |
| `TEXT_LAYER_MIN_CHARS` | `20` | Below this per page → `has_text_layer=False` |
| `TEXT_LAYER_MAX_CHARS` | `12000` | Cap per sheet injected into extractor prompt |
| `VERIFY_TEXT_MAX_CHARS` | `8000` | Cap of text-layer excerpt in verify scope |
| `VERIFY_HI_DPI_CROPS` | `true` | Evidence-located crops in verifier |
| `VERIFY_CROP_DPI` | `300` | Crop render DPI |
| `VERIFY_CROP_MARGIN_PTS` | `36` | Padding around evidence bbox (PDF points) |
## Known traps (from project history — designed around)
- **Two render paths:** no new `{placeholder}` in extractor templates; the one
new placeholder (`{text_layer}` in VERIFY_USER_INSTRUCTION) substituted at
every render site; a render test asserts no `{...}` literals remain.
- **ProjectMemory closed registry:** no new memory keys. Text artifacts dump
via plain file writes under `outputs/<job>/text/` (agent: under `agent/`).
- **`slim_clusters`:** no new cluster fields — unchanged.
- **I2 zero-image path:** crops replace full-page images only on confident
bbox match; never reduce image count to zero.
- **Base64 hygiene:** page dicts already carry base64; `text_layer` strings
must not leak into `clusters.json` dumps — reuse `_without_base64` pattern
if assertions ever carry page refs (they don't today).
## Dependencies
`PyMuPDF>=1.23` added to `requirements.txt` (Docker image rebuild picks it up;
pdf2image/poppler unchanged).
## Testing
- `tests/test_text_layer.py` — build tiny PDFs with PyMuPDF in-test:
extraction, `has_text_layer` thresholds, `find_evidence_bbox` hit/miss,
`render_crop` dimensions.
- Extractor guard: rescue-tier unit tests (keep-with-flag, still-drop,
unchanged behavior without text layer).
- Prompt render test: extractor + verify instructions fully substituted at
every site (both pipelines).
- Runner-level (pattern from `tests/agents/test_wave5b_suppression.py`):
stubbed waves, assert text layer reaches extract scopes and verify scopes
(excerpt present, crop fallback on no-match), full `run_agent_pipeline`.
- Full `pytest tests/` green before push.
## Validation (post-deploy)
Re-run the Cypress set (source PDF persists at
`/app/backend/outputs/959e16407573/source.pdf` on sits-docker) per the
documented re-run workflow. Success criteria:
1. The "(2) vs (5)"-class findings are not generated, or are verifier-refuted
with text-layer evidence cited.
2. Coverage-gap log lines appear for any page with text but no objects.
3. No new `finish_reason=length` in waves 1/4; cost delta reported vs
baseline job.
## Out of scope (future PRs)
- Legend/symbol-library wave injected into extractor + critic prompts.
- Deterministic schedule-row recall pass (text-layer tables → assertions).
- pdf-markup export for the review UI.
- Takeoff/polygon geometry (belongs to AI_Takeoffs, not this product).
+163 -36
View File
@@ -3,6 +3,9 @@
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1" />
<meta http-equiv="Cache-Control" content="no-cache, no-store, must-revalidate" />
<meta http-equiv="Pragma" content="no-cache" />
<meta http-equiv="Expires" content="0" />
<title>Conflict Checker</title>
<style>
:root {
@@ -15,6 +18,7 @@
header { padding:24px 28px; border-bottom:1px solid var(--line); }
h1 { margin:0; font-size:20px; letter-spacing:.2px; }
.sub { color:var(--muted); font-size:13px; margin-top:4px; }
#buildTag { display:inline-block; font-size:11px; background:rgba(91,140,255,.15); color:var(--accent); padding:2px 8px; border-radius:12px; margin-left:8px; vertical-align:middle; text-transform:none; letter-spacing:.3px; }
main { max-width:920px; margin:0 auto; padding:28px; }
.drop { border:1.5px dashed var(--line); border-radius:12px; padding:36px; text-align:center;
background:var(--panel); transition:border-color .15s; cursor:pointer; }
@@ -60,7 +64,15 @@
border:1px solid var(--line); background:#0c0e13; color:var(--text); font-size:16px; outline:none; }
.email-card input[type=email]::placeholder { color:#6b7280; }
.email-card input[type=email]:focus { border-color:var(--accent); box-shadow:0 0 0 3px rgba(91,140,255,.18); }
.email-card select { width:100%; padding:10px 12px; border-radius:8px; margin-top:6px;
border:1px solid var(--line); background:#0c0e13; color:var(--text); font-size:14px; }
.email-card .field { margin-top:12px; }
.email-card .field > span { display:block; font-size:13px; color:var(--muted); margin-bottom:2px; }
.btn.full { width:100%; padding:14px; font-size:15px; margin-top:0; }
.logbox { background:#0c0e13; border:1px solid var(--line); border-radius:8px; padding:12px 14px;
margin-top:10px; max-height:320px; overflow:auto; font:12px/1.45 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace;
color:#c6cdd8; white-space:pre-wrap; word-break:break-word; }
.logbox .empty-log { color:var(--muted); }
.note { background:var(--panel); border:1px solid var(--line); border-radius:10px;
padding:16px 18px; margin:18px 0; }
.note b { color:var(--text); }
@@ -87,7 +99,7 @@
<body>
<header>
<h1>Conflict Checker</h1>
<div class="sub">Cross-discipline design contradiction review for construction drawing sets<span id="buildTag"></span></div>
<div class="sub">Cross-discipline design contradiction review for construction drawing sets<span id="buildTag">build ...</span></div>
</header>
<main>
<div class="drop" id="drop">
@@ -118,15 +130,23 @@
<label>&#9881;&#65039; Compute <span class="opt">(text stages; vision always runs on the API)</span></label>
<label style="display:block;font-weight:400;margin-top:6px">
<input type="radio" name="compute" value="openrouter" checked> OpenRouter &mdash; all stages (fastest, paid)</label>
<label style="display:block;font-weight:400;margin-top:6px">
<input type="radio" name="compute" value="local"> Hybrid &mdash; text stages on local LLM (cheaper, slower)</label>
<div id="modelPick" style="margin-top:10px">
<label for="model" style="font-weight:400">Model <span class="opt" id="modelNote">loading...</span></label>
<select id="model" style="width:100%;margin-top:6px;padding:10px;border-radius:8px;border:1px solid var(--line);background:#0c0e13;color:var(--text)"></select>
<div class="field">
<span>Vision model <span class="opt">(image stages)</span> <span class="opt" id="modelNote">loading...</span></span>
<select id="vision_model" disabled><option value="">Loading models&hellip;</option></select>
</div>
<div class="field">
<span>Text model <span class="opt">(non-image stages)</span></span>
<select id="text_model" disabled><option value="">Loading models&hellip;</option></select>
</div>
</div>
</div>
<button class="btn full" id="run" disabled>Run conflict check</button>
<div class="status" id="status"></div>
<div id="liveLog" style="display:none" class="note">
<b>Run log</b> <span class="opt" id="logHint">(updates live)</span>
<pre class="logbox" id="logBox"><span class="empty-log">Waiting for output&hellip;</span></pre>
</div>
<div id="results"></div>
</main>
<div id="viewer">
@@ -144,8 +164,106 @@
const drop=document.getElementById('drop'), fileInput=document.getElementById('file'),
runBtn=document.getElementById('run'), statusEl=document.getElementById('status'),
results=document.getElementById('results'), dropLabel=document.getElementById('dropLabel'),
emailEl=document.getElementById('email');
emailEl=document.getElementById('email'),
visionSel=document.getElementById('vision_model'),
textSel=document.getElementById('text_model'),
liveLog=document.getElementById('liveLog'),
logBox=document.getElementById('logBox'),
logHint=document.getElementById('logHint');
let chosen=null, polling=null, currentJobId=null, sheetPage={}, viewerZoom=1, reviewDirty=false;
let defaultVisionModel='', defaultTextModel='';
function modelLabel(m){
// Include per-1M-token pricing when the catalog provides it.
let s=m.name||m.id;
if(m.prompt_usd_per_mtok!=null)
s+=' — $'+m.prompt_usd_per_mtok+' / $'+m.completion_usd_per_mtok+' per 1M tok';
return s;
}
function fillSelect(sel, items, preferred){
sel.innerHTML='';
(items||[]).forEach(m=>{
const opt=document.createElement('option');
opt.value=m.id; opt.textContent=modelLabel(m);
if(m.id===preferred) opt.selected=true;
sel.appendChild(opt);
});
if(!sel.options.length){
const opt=document.createElement('option');
opt.value=preferred||''; opt.textContent=preferred||'(no models)';
sel.appendChild(opt);
}
sel.disabled=false;
}
let modelsLoaded=false;
const MODELS_CACHE_KEY='cc_models_v1';
const MODELS_CACHE_TTL=24*60*60*1000;
function loadModelsCache(){
try{
const raw=localStorage.getItem(MODELS_CACHE_KEY);
if(!raw) return null;
const parsed=JSON.parse(raw);
if(!parsed.ts || Date.now()-parsed.ts > MODELS_CACHE_TTL) return null;
return parsed.data||null;
}catch(e){ return null; }
}
function saveModelsCache(data){
try{ localStorage.setItem(MODELS_CACHE_KEY, JSON.stringify({ts:Date.now(), data})); }
catch(e){}
}
function fetchWithTimeout(url, ms){
return Promise.race([
fetch(url, {cache:'no-store'}),
new Promise((_,reject)=>setTimeout(()=>reject(new Error('timeout')), ms))
]);
}
async function loadModels(){
const note=document.getElementById('modelNote');
const cached=loadModelsCache();
if(cached){
const defs=cached.defaults||{};
fillSelect(visionSel, cached.vision, defs.vision);
fillSelect(textSel, cached.text, defs.text);
modelsLoaded=true;
note.textContent='('+(cached.text||[]).length+' text / '+(cached.vision||[]).length+' vision cached)';
return;
}
try{
note.textContent='fetching models...';
const res=await fetchWithTimeout('/models', 5000);
if(!res.ok) throw new Error('models HTTP '+res.status);
const data=await res.json();
saveModelsCache(data);
const defs=data.defaults||{};
fillSelect(visionSel, data.vision, defs.vision);
fillSelect(textSel, data.text, defs.text);
modelsLoaded=true;
note.textContent='('+(data.text||[]).length+' text / '+(data.vision||[]).length+' vision available)';
}catch(err){
const fallback=[{id:defaultVisionModel||'', name:defaultVisionModel||'(default)', prompt_usd_per_mtok:null, completion_usd_per_mtok:null}];
fillSelect(visionSel, fallback.filter(m=>m.id), defaultVisionModel);
const fallbackText=[{id:defaultTextModel||'', name:defaultTextModel||'(default)', prompt_usd_per_mtok:null, completion_usd_per_mtok:null}];
fillSelect(textSel, fallbackText.filter(m=>m.id), defaultTextModel);
visionSel.disabled=false; textSel.disabled=false;
note.textContent='using configured defaults ('+(err.message==='timeout'?'fetch timed out':'list unavailable')+')';
console.warn('Could not load models:', err);
}
}
function showLog(lines, live){
liveLog.style.display='block';
logHint.textContent=live?'(updates live)':'(saved with this job)';
const arr=lines||[];
if(!arr.length){
logBox.innerHTML='<span class="empty-log">No log lines yet&hellip;</span>';
return;
}
logBox.textContent=arr.join('\n');
logBox.scrollTop=logBox.scrollHeight;
}
function setFile(f){ chosen=f; dropLabel.textContent=f?('Selected: '+f.name):'Drop a PDF drawing set here, or click to choose';
runBtn.disabled=!f; }
@@ -160,6 +278,7 @@ runBtn.addEventListener('click',async e=>{
if(!chosen) return;
runBtn.disabled=true; results.innerHTML='';
statusEl.innerHTML='<span class="spinner"></span>Uploading...';
showLog([], true);
const fd=new FormData(); fd.append('file',chosen);
const email=(emailEl.value||'').trim(); if(email) fd.append('notification_email',email);
['project_name','address','occupancy','work_type'].forEach(id=>{
@@ -168,8 +287,9 @@ runBtn.addEventListener('click',async e=>{
const compute=(document.querySelector('input[name="compute"]:checked')||{}).value;
fd.append('text_local', compute==='local' ? 'true' : 'false');
if(compute==='openrouter'){
const modelSel=document.getElementById('model');
if(modelSel.value) fd.append('model', modelSel.value);
// Model picks only apply to OpenRouter compute; hybrid keeps its local model.
if(visionSel.value) fd.append('vision_model', visionSel.value);
if(textSel.value) fd.append('text_model', textSel.value);
}
const pipelineMode=(document.querySelector('input[name="pipeline_mode"]:checked')||{}).value||'classic';
fd.append('pipeline_mode',pipelineMode);
@@ -196,21 +316,30 @@ function poll(jobId){
const res=await fetch('/jobs/'+jobId);
if(!res.ok) throw new Error('job not found');
const job=await res.json();
const live=['running','queued','finalizing'].includes(job.status);
if(job.log_tail && job.log_tail.length) showLog(job.log_tail, live);
if(job.status==='running'||job.status==='queued'){
statusEl.innerHTML='<span class="spinner"></span>'+esc(job.stage||'Working...')+
' &middot; you can leave this page';
} else if(job.status==='done'){
clearInterval(polling); polling=null; runBtn.disabled=false; render(job.report);
clearInterval(polling); polling=null; runBtn.disabled=false;
if(job.log && job.log.length) showLog(job.log, false);
render(job.report);
} else if(job.status==='needs_review'||job.status==='reviewing'){
clearInterval(polling); polling=null; runBtn.disabled=false; renderReview(job);
clearInterval(polling); polling=null; runBtn.disabled=false;
if(job.log && job.log.length) showLog(job.log, false);
renderReview(job);
} else if(job.status==='finalizing'){
statusEl.innerHTML='<span class="spinner"></span>Finalizing reviewed report...';
} else if(job.status==='finalization_error'){
clearInterval(polling); polling=null; runBtn.disabled=false;
statusEl.textContent='Finalization failed: '+(job.error||'unknown error');
if(job.log && job.log.length) showLog(job.log, false);
} else if(job.status==='error'){
clearInterval(polling); polling=null; runBtn.disabled=false;
statusEl.textContent='Run failed: '+(job.error||'unknown error');
if(job.log && job.log.length) showLog(job.log, false);
else if(job.log_tail && job.log_tail.length) showLog(job.log_tail, false);
}
}catch(err){ clearInterval(polling); polling=null; runBtn.disabled=false;
statusEl.textContent='Error: '+err.message; }
@@ -224,35 +353,19 @@ function escAttr(s){ return esc(s).replace(/"/g,'&quot;'); }
function syncPipelineOptions(){
const agent=(document.querySelector('input[name="pipeline_mode"]:checked')||{}).value==='agent';
const local=document.querySelector('input[name="compute"][value="local"]');
local.disabled=agent;
if(agent&&local.checked) document.querySelector('input[name="compute"][value="openrouter"]').checked=true;
if(local){
local.disabled=agent;
if(agent&&local.checked) document.querySelector('input[name="compute"][value="openrouter"]').checked=true;
}
}
document.querySelectorAll('input[name="pipeline_mode"]').forEach(el=>el.addEventListener('change',syncPipelineOptions));
syncPipelineOptions();
// --- model picker (OpenRouter compute only) ---
let modelList=null;
async function loadModels(){
const note=document.getElementById('modelNote'), sel=document.getElementById('model');
try{
const res=await fetch('/models');
if(!res.ok) throw new Error('list unavailable');
const data=await res.json();
modelList=data.models||[];
sel.innerHTML=modelList.map(m=>
'<option value="'+escAttr(m.id)+'"'+(m.id===data.default?' selected':'')+'>'+
esc(m.name||m.id)+' &mdash; $'+esc(m.prompt_usd_per_mtok)+' / $'+esc(m.completion_usd_per_mtok)+
' per 1M tok</option>').join('');
note.textContent='('+modelList.length+' available)';
}catch(e){
sel.innerHTML='';
note.textContent='using configured default (list unavailable)';
}
}
// --- model pickers (OpenRouter compute only) ---
function syncCompute(){
const openrouter=(document.querySelector('input[name="compute"]:checked')||{}).value==='openrouter';
document.getElementById('modelPick').style.display=openrouter?'block':'none';
if(openrouter&&!modelList) loadModels();
if(openrouter&&!modelsLoaded) loadModels();
}
document.querySelectorAll('input[name="compute"]').forEach(el=>el.addEventListener('change',syncCompute));
syncCompute();
@@ -394,6 +507,10 @@ function render(rep){
html+='</details>';
}
html+='<div class="meta" style="margin-top:14px">Full run log: <a href="/jobs/'+
esc(currentJobId)+'/log?plain=1" target="_blank" rel="noopener">/jobs/'+
esc(currentJobId)+'/log</a> (also saved as job.log on the server)</div>';
results.innerHTML=html;
}
function stat(v,l){ return '<div class="stat"><b>'+esc(v)+'</b><span>'+esc(l)+'</span></div>'; }
@@ -467,6 +584,9 @@ function reviewItemHtml(item,uid,prev){
esc(e.source_text)+'"</div>').join('')+'</div>';
}
}
if((p.sheets||[]).length){
html+='<div class="meta">Sheets: '+sheetList(p.sheets)+'</div>';
}
if((item.reasons||[]).length){
html+='<div class="meta">Review triggers: '+esc(item.reasons.join(', '))+'</div>';
}
@@ -566,10 +686,17 @@ async function finalizeReview(){
// If opened from an email link (/?job=<id>), load that job's results directly.
(function init(){
// Show the deployed build in the header so it's obvious which version is up.
fetch('/health').then(r=>r.ok?r.json():null).then(h=>{
if(h&&h.build) document.getElementById('buildTag').textContent=' · build '+h.build;
}).catch(()=>{});
// Fetch /health first: build tag for the header and default models as a
// fallback if the larger /models catalog fails or times out.
fetch('/health', {cache:'no-store'}).then(r=>r.ok?r.json():null).then(h=>{
if(h){
if(h.build) document.getElementById('buildTag').textContent=' · build '+h.build;
if(h.model) defaultVisionModel=h.model;
if(h.text_model) defaultTextModel=h.text_model;
}
}).catch(()=>{}).finally(()=>{
loadModels();
});
const jobId=new URLSearchParams(location.search).get('job');
if(jobId){ statusEl.innerHTML='<span class="spinner"></span>Loading job '+esc(jobId)+'...'; poll(jobId); }
})();
+1
View File
@@ -2,6 +2,7 @@ fastapi==0.115.0
uvicorn[standard]==0.30.6
python-multipart==0.0.12
pdf2image==1.17.0
PyMuPDF>=1.23.0 # deterministic text-layer extraction (extractor grounding, verifier crops)
Pillow==10.4.0
openai==1.51.0
httpx==0.27.2 # openai 1.51 passes proxies= to httpx; >=0.28 dropped it
@@ -0,0 +1,51 @@
"""Classic pipeline path must also satisfy the {disputes} placeholder added to
CONSTRUCTABILITY_USER_INSTRUCTION (agent path substitutes it in construct_agent.py;
the classic stage builds its own subs dict)."""
from unittest.mock import patch
from backend.pipeline._stage import render
from backend.pipeline.constructability import constructability_review
from backend.prompts import CONSTRUCTABILITY_USER_INSTRUCTION
def _cluster_with_dispute():
return {
"key": "c1",
"assertions": [],
"disputed_attributes": [{
"attribute": "stud_pack_size",
"values": ["(2) 2x6 STUD PACK", "(5) 2x6 STUD PACK"],
"assertion_ids": ["a1", "a2"],
}],
}
def test_classic_constructability_supplies_disputes_sub():
captured = {}
def fake_call_stage(system_prompt, user_instruction, subs=None, **kwargs):
captured["subs"] = subs or {}
return {"issues": []}
with patch("backend.pipeline.constructability.call_stage", fake_call_stage):
constructability_review([], [_cluster_with_dispute()], [])
assert "disputes" in captured["subs"], "classic path must substitute {disputes}"
rendered = render(CONSTRUCTABILITY_USER_INSTRUCTION, captured["subs"])
assert "{disputes}" not in rendered
assert "(5) 2x6 STUD PACK" in rendered
def test_classic_constructability_disputes_defaults_empty():
captured = {}
def fake_call_stage(system_prompt, user_instruction, subs=None, **kwargs):
captured["subs"] = subs or {}
return {"issues": []}
with patch("backend.pipeline.constructability.call_stage", fake_call_stage):
constructability_review([], [{"key": "c2", "assertions": []}], [])
rendered = render(CONSTRUCTABILITY_USER_INSTRUCTION, captured["subs"])
assert "{disputes}" not in rendered
+69
View File
@@ -0,0 +1,69 @@
from backend.agents.disputes import annotate_clusters, find_disputes
def _a(id_, attribute, value):
return {"id": id_, "attribute": attribute, "value": value,
"source_text": value}
def test_find_disputes_flags_same_attribute_different_values():
assertions = [
_a("a1", "stud_pack_size", "(2) 2x6 STUD PACK"),
_a("a2", "stud_pack_size", "(5) 2x6 STUD PACK"),
_a("a3", "beam_size", "HSS16X4X5/8"),
]
disputes = find_disputes(assertions)
assert len(disputes) == 1
assert disputes[0]["attribute"] == "stud_pack_size"
assert disputes[0]["values"] == ["(2) 2x6 STUD PACK", "(5) 2x6 STUD PACK"]
assert disputes[0]["assertion_ids"] == ["a1", "a2"]
def test_find_disputes_ignores_agreeing_values_and_blanks():
assertions = [
_a("a1", "beam_size", "HSS16X4X5/8"),
_a("a2", "beam_size", " hss16x4x5/8 "), # same after normalize
_a("a3", "", "orphan"), # no attribute -> skipped
_a("a4", "beam_size", ""), # no value -> skipped
]
assert find_disputes(assertions) == []
def test_annotate_clusters_writes_disputed_attributes():
clusters = [
{"key": "c1", "assertions": [
_a("a1", "stud_pack_size", "(2) 2x6"),
_a("a2", "stud_pack_size", "(5) 2x6"),
]},
{"key": "c2", "assertions": [_a("a3", "x", "1"), _a("a4", "x", "1")]},
]
assert annotate_clusters(clusters) == 1
assert clusters[0]["disputed_attributes"][0]["attribute"] == "stud_pack_size"
assert "disputed_attributes" not in clusters[1]
def test_slim_clusters_preserves_disputed_attributes():
from backend.pipeline._serialize import slim_clusters
cluster = {"key": "c1", "assertions": [],
"disputed_attributes": [{"attribute": "a", "values": ["1", "2"],
"assertion_ids": ["x", "y"]}]}
slim = slim_clusters([cluster])[0]
assert slim["disputed_attributes"][0]["values"] == ["1", "2"]
def test_find_disputes_handles_none_and_zero_values():
# None value/attribute -> skipped; numeric 0 is a real value, not blank
assertions = [
{"id": "a1", "attribute": "count", "value": 0},
{"id": "a2", "attribute": "count", "value": 1},
{"id": "a3", "attribute": None, "value": "x"},
{"id": "a4", "attribute": "count", "value": None},
]
disputes = find_disputes(assertions)
assert len(disputes) == 1
assert disputes[0]["values"] == ["0", "1"]
def test_find_disputes_empty_input():
assert find_disputes([]) == []
assert annotate_clusters([]) == 0
+74
View File
@@ -0,0 +1,74 @@
from backend.agents.base import AgentScope
from backend.agents.linker import build_link_scopes
def _sheet(number, page, level, assertions):
return {"sheet_number": number, "page_number": page,
"discipline": "Structural", "level": level,
"assertions": assertions}
def _assertion(id_, ref=None, tag=None, level=None):
return {"id": id_, "attribute": "stud_pack_size", "value": "(5) 2x6",
"source_text": "(5) 2x6 STUD PACK",
"location_key": {"detail_reference": ref, "tag": tag,
"level": level}}
def test_xref_scope_joins_same_detail_reference_across_levels():
sheets = [
_sheet("S101", 10, "foundation", [_assertion("a1", ref="A/S205")]),
_sheet("S205", 20, "roof", [_assertion("a2", ref="A/S205")]),
_sheet("S401", 30, "roof", [_assertion("a3", ref="A/S205")]),
]
scopes = build_link_scopes(sheets)
xref = [s for s in scopes if s.scope_id.startswith("xref:")]
assert xref, "expected a cross-level detail-reference scope"
ids = {a["id"] for s in xref for a in s.payload["assertions"]}
assert ids == {"a1", "a2", "a3"}
def test_xref_scope_requires_two_distinct_sheets():
sheets = [
_sheet("S401", 30, "roof", [_assertion("a1", ref="A/S205"),
_assertion("a2", ref="A/S205")]),
]
scopes = build_link_scopes(sheets)
assert not [s for s in scopes if s.scope_id.startswith("xref:")]
def test_xref_scope_joins_shared_member_tag():
sheets = [
_sheet("S102", 5, "roof", [_assertion("a1", tag="HSS16X4X5/8")]),
_sheet("S401", 30, "unknown", [_assertion("a2", tag="HSS16X4X5/8")]),
]
scopes = build_link_scopes(sheets)
xref = [s for s in scopes if s.scope_id.startswith("xref:")]
assert xref
def test_xref_scope_joins_single_letter_member_mark():
# W-shapes (W12X26) are the most common steel marks and have one leading letter
sheets = [
_sheet("S102", 5, "roof", [_assertion("a1", tag="W12X26")]),
_sheet("S401", 30, "unknown", [_assertion("a2", tag="W12X26")]),
]
scopes = build_link_scopes(sheets)
xref = [s for s in scopes if s.scope_id.startswith("xref:")]
assert xref, "single-letter member marks (W12X26) must join xref scopes"
def test_xref_scope_rechecks_sheet_diversity_after_cap(monkeypatch):
from backend import config
monkeypatch.setattr(config, "AGENT_LINK_MAX_ASSERTIONS", 2)
sheets = [
_sheet("S401", 30, "roof", [_assertion("a1", ref="A/S205"),
_assertion("a2", ref="A/S205")]),
_sheet("S205", 20, "roof", [_assertion("a3", ref="A/S205")]),
]
scopes = build_link_scopes(sheets)
xref = [s for s in scopes if s.scope_id.startswith("xref:")]
for scope in xref:
sheets_in_scope = {a["sheet_number"] for a in scope.payload["assertions"]}
assert len(sheets_in_scope) >= 2, \
"capped xref scope must still span two sheets"
@@ -0,0 +1,76 @@
"""SheetExtractorAgent fallback ladder tests (bare-list wrap + compact retry)."""
from unittest.mock import patch
from backend.agents.base import AgentScope, AgentUsage
from backend.agents.extractors import SheetExtractorAgent, _wrap_bare_list
def _scope():
return AgentScope(
scope_id="sheet:4",
payload={"page": {"page_number": 4, "base64": "QUJD"}},
)
def _objects(n=2):
return [
{
"object_id": f"obj-{i}",
"object_type": "equipment",
"category": "mechanical",
"name": f"RTU-{i}",
"source_text": f"RTU-{i}",
"confidence": "high",
}
for i in range(n)
]
def test_wrap_bare_list_builds_sheet_envelope():
wrapped = _wrap_bare_list(_objects(3), page_number=4)
assert wrapped["sheet"] == {}
assert len(wrapped["objects"]) == 3
def test_wrap_bare_list_passes_dicts_and_none_through():
assert _wrap_bare_list({"sheet": {}, "objects": []}, 1) == {"sheet": {}, "objects": []}
assert _wrap_bare_list(None, 1) is None
def test_run_accepts_bare_list_response():
agent = SheetExtractorAgent(usage=AgentUsage())
with patch("backend.agents.extractors.call_json",
return_value=_objects(5)) as mock_call:
result = agent.run(_scope())
assert not result.error
assert len(result.artifacts) == 1
sheet = result.artifacts[0]
assert sheet["page_number"] == 4
assert len(sheet["assertions"]) == 5
# No compact retry needed when the first call yields data.
assert mock_call.call_count == 1
# Reasoning knobs are forwarded (None when config is blank in tests).
assert "reasoning_effort" in mock_call.call_args.kwargs
assert "reasoning_max_tokens" in mock_call.call_args.kwargs
def test_run_compact_retry_after_hard_failure():
agent = SheetExtractorAgent(usage=AgentUsage())
with patch("backend.agents.extractors.call_json",
side_effect=[None, {"sheet": {"sheet_number": "A102"},
"objects": _objects(2)}]) as mock_call:
result = agent.run(_scope())
assert not result.error
assert result.artifacts[0]["sheet_number"] == "A102"
assert mock_call.call_count == 2
# Second call carried the compact suffix.
assert "COMPACT RETRY" in mock_call.call_args_list[1].kwargs["user_text"]
def test_run_fails_only_after_both_attempts_miss():
agent = SheetExtractorAgent(usage=AgentUsage())
with patch("backend.agents.extractors.call_json", return_value=None) as mock_call:
result = agent.run(_scope())
assert result.error == "no structured extraction returned"
assert mock_call.call_count == 2
+128
View File
@@ -0,0 +1,128 @@
"""Runner-level text-layer flow: excerpt into verify scopes, hi-DPI crop
replacement with full-page fallback, and coverage-gap findings."""
import pytest
fitz = pytest.importorskip("pymupdf")
import backend.agents.runner as runner_mod
from backend.agents.base import AgentResult
from backend.agents.runner import run_agent_pipeline
PAGE_TEXT = "(5) 2X6 STUD PACK AT BEARING"
def _make_pdf(path):
doc = fitz.open()
page = doc.new_page(width=612, height=792)
page.insert_text((72, 72), PAGE_TEXT, fontsize=11)
doc.save(str(path))
doc.close()
return str(path)
def _finding(sheets, evidence_text):
return {
"issue_id": "C1", "severity": "critical", "confidence": "high",
"source_stage": "constructability", "sheets": sheets,
"description": "stud pack conflict",
"evidence": [{"sheet": sheets[0], "source_text": evidence_text}],
}
def _stub_agent(artifacts):
return lambda usage: type("S", (), {
"name": "stub",
"run": lambda self, scope: AgentResult(
scope_id=scope.scope_id, artifacts=list(artifacts)),
})()
def _patch_pipeline(monkeypatch, finding, verify_sink):
monkeypatch.setattr(
runner_mod, "convert_pdf_to_images",
lambda path: [{"page_number": 1, "base64": "QUJD"}])
monkeypatch.setattr(runner_mod, "SheetExtractorAgent", _stub_agent([
{"sheet_number": "S401", "page_number": 1, "level": "roof",
"discipline": "S", "assertions": [
{"text": "(5) 2X6 STUD PACK", "object_type": "framing"},
{"text": "HSS16X4 beam", "object_type": "framing"},
]},
]))
monkeypatch.setattr(runner_mod, "SheetIndexAgent", _stub_agent([{}]))
monkeypatch.setattr(runner_mod, "JurisdictionAgent", _stub_agent([{}]))
monkeypatch.setattr(runner_mod, "LinkerAgent", _stub_agent([
{"key": "c1", "location": "roof beam pocket", "assertions": []},
]))
monkeypatch.setattr(runner_mod, "ConflictCriticAgent", _stub_agent([]))
monkeypatch.setattr(runner_mod, "CodeAgent", _stub_agent([]))
monkeypatch.setattr(runner_mod, "ConstructabilityAgent",
_stub_agent([finding]))
monkeypatch.setattr(runner_mod, "CompletenessAgent", _stub_agent([]))
monkeypatch.setattr(
runner_mod, "BrainAgent",
lambda usage: type("B", (), {
"run": lambda self, findings, sheet_index, jurisdiction:
(list(findings), [])})())
class _RecordingVerifier:
name = "verify"
def __init__(self, usage):
pass
def run(self, scope):
verify_sink.append(scope.payload)
return AgentResult(scope_id=scope.scope_id, artifacts=[{
"finding_index": scope.payload["finding_index"],
"status": "confirmed",
"verdicts": [],
}])
monkeypatch.setattr(runner_mod, "EvidenceVerifierAgent",
lambda usage: _RecordingVerifier(usage))
def test_verify_scope_carries_text_excerpt_and_crop(monkeypatch, tmp_path):
"""Evidence text matches the page text layer -> excerpt present and the
full-page image is replaced by a hi-DPI crop."""
sink = []
_patch_pipeline(monkeypatch,
_finding(["S401"], "(5) 2X6 STUD PACK AT BEARING"), sink)
pdf = _make_pdf(tmp_path / "set.pdf")
run_agent_pipeline(pdf, out_dir=str(tmp_path), require_review=False)
assert len(sink) == 1
payload = sink[0]
assert "2X6 STUD PACK" in payload["text_layer_excerpt"]
assert payload["images_b64"], "crop must never drop all images"
assert payload["images_b64"][0] != "QUJD", "expected crop, not full page"
def test_verify_scope_falls_back_to_full_page(monkeypatch, tmp_path):
"""Evidence text not in the text layer -> keep the full-page image."""
sink = []
_patch_pipeline(monkeypatch,
_finding(["S401"], "PENTHOUSE EXHAUST FAN EF-9"), sink)
pdf = _make_pdf(tmp_path / "set.pdf")
run_agent_pipeline(pdf, out_dir=str(tmp_path), require_review=False)
assert len(sink) == 1
assert sink[0]["images_b64"] == ["QUJD"]
def test_coverage_gap_becomes_gap_finding(monkeypatch, tmp_path):
"""Text layer present but zero objects extracted -> failed-scope gap
finding survives into the report."""
sink = []
_patch_pipeline(monkeypatch, _finding(["S401"], PAGE_TEXT), sink)
# Extractor returns a sheet with NO objects despite a real text layer.
monkeypatch.setattr(runner_mod, "SheetExtractorAgent", _stub_agent([
{"sheet_number": "S401", "page_number": 1, "level": "roof",
"discipline": "S", "assertions": []},
]))
pdf = _make_pdf(tmp_path / "set.pdf")
report = run_agent_pipeline(pdf, out_dir=str(tmp_path),
require_review=False)
gaps = [f for f in (report.get("validated_issues") or [])
if f.get("category") == "analysis_gap"]
assert any("extraction gap" in (g.get("description") or "")
for g in gaps)
+68
View File
@@ -0,0 +1,68 @@
from unittest.mock import patch
from backend.agents.base import AgentScope, AgentUsage
from backend.agents.verifier import (
EvidenceVerifierAgent, apply_verdicts, select_findings,
)
def _finding(sev="critical", issue_id="i1", sheets=("S401",), cluster_key=None):
f = {"issue_id": issue_id, "severity": sev, "confidence": "high",
"source_stage": "constructability", "sheets": list(sheets),
"description": "HSS16x4 on (2) 2x6 STUD PACK is unbuildable",
"evidence": [{"sheet": "S401", "source_text": "(2) 2x6 STUD PACK",
"asserted_value": "3-inch width"}]}
if cluster_key:
f["cluster_key"] = cluster_key
return f
def test_select_findings_by_severity_and_dispute():
findings = [_finding("critical"), _finding("low", "i2"),
_finding("medium", "i3", cluster_key="c9")]
clusters = [{"key": "c9", "disputed_attributes": [{"attribute": "a"}]}]
selected = select_findings(findings, clusters, max_checks=20,
severities={"critical", "high"})
assert [f["issue_id"] for f in selected] == ["i1", "i3"]
def test_select_findings_respects_cap():
findings = [_finding("critical", f"i{n}") for n in range(30)]
selected = select_findings(findings, [], max_checks=5,
severities={"critical"})
assert len(selected) == 5
def test_run_attaches_verdicts_and_marks_refuted():
agent = EvidenceVerifierAgent(usage=AgentUsage())
scope = AgentScope(scope_id="verify:0", payload={
"finding_index": 0,
"finding": _finding(),
"images_b64": ["QUJD"],
})
verdicts = {"verdicts": [
{"sheet": "S401", "source_text": "(2) 2x6 STUD PACK",
"verdict": "corrected", "actual_text": "(5) 2x6 STUD PACK",
"notes": "callout reads (5)"},
]}
with patch("backend.agents.verifier.call_json", return_value=verdicts):
result = agent.run(scope)
assert not result.error
artifact = result.artifacts[0]
assert artifact["finding_index"] == 0
assert artifact["status"] == "refuted" # no evidence confirmed
assert artifact["verdicts"][0]["actual_text"] == "(5) 2x6 STUD PACK"
def test_apply_verdicts_annotates_and_suppresses():
from backend.agents.base import AgentResult
findings = [_finding("critical", "i1"), _finding("high", "i2")]
results = [AgentResult(scope_id="verify:0", artifacts=[
{"finding_index": 0, "status": "refuted", "verdicts": []},
{"finding_index": 1, "status": "confirmed", "verdicts": []},
])]
suppressed = apply_verdicts(findings, results)
assert suppressed == [findings[0]]
assert findings[0]["verification"]["status"] == "refuted"
assert findings[1]["verification"]["status"] == "confirmed"
+86
View File
@@ -0,0 +1,86 @@
"""Runner-level wave-5b tests: suppression path and zero-image guard."""
import backend.agents.runner as runner_mod
from backend.agents.base import AgentResult
from backend.agents.runner import run_agent_pipeline
def _finding(sheets):
return {
"issue_id": "C1", "severity": "critical", "confidence": "high",
"source_stage": "constructability", "sheets": sheets,
"description": "HSS16x4 on (2) 2x6 STUD PACK is unbuildable",
"evidence": [{"sheet": sheets[0], "source_text": "(2) 2x6 STUD PACK"}],
}
def _stub_agent(artifacts):
return lambda usage: type("S", (), {
"name": "stub",
"run": lambda self, scope: AgentResult(
scope_id=scope.scope_id, artifacts=list(artifacts)),
})()
def _patch_pipeline(monkeypatch, finding):
monkeypatch.setattr(
runner_mod, "convert_pdf_to_images",
lambda path: [{"page_number": 1, "base64": "QUJD"}])
monkeypatch.setattr(runner_mod, "SheetExtractorAgent", _stub_agent([
{"sheet_number": "S401", "page_number": 1, "level": "roof",
"discipline": "S", "assertions": [
{"text": "(2) 2x6 STUD PACK", "object_type": "framing"},
{"text": "HSS16X4 beam", "object_type": "framing"},
]},
]))
monkeypatch.setattr(runner_mod, "SheetIndexAgent", _stub_agent([{}]))
monkeypatch.setattr(runner_mod, "JurisdictionAgent", _stub_agent([{}]))
monkeypatch.setattr(runner_mod, "LinkerAgent", _stub_agent([
{"key": "c1", "location": "roof beam pocket", "assertions": []},
]))
monkeypatch.setattr(runner_mod, "ConflictCriticAgent", _stub_agent([]))
monkeypatch.setattr(runner_mod, "CodeAgent", _stub_agent([]))
monkeypatch.setattr(runner_mod, "ConstructabilityAgent", _stub_agent([finding]))
monkeypatch.setattr(runner_mod, "CompletenessAgent", _stub_agent([]))
monkeypatch.setattr(
runner_mod, "BrainAgent",
lambda usage: type("B", (), {
"run": lambda self, findings, sheet_index, jurisdiction:
(list(findings), [])})())
def test_refuted_finding_is_suppressed_not_crash(monkeypatch, tmp_path):
"""Regression: memory.replace("suppressed", ...) must not KeyError."""
_patch_pipeline(monkeypatch, _finding(["S401"]))
monkeypatch.setattr(
"backend.agents.verifier.call_json",
lambda **kwargs: {"verdicts": [
{"sheet": "S401", "source_text": "(2) 2x6 STUD PACK",
"verdict": "corrected", "actual_text": "(5) 2x6 STUD PACK",
"notes": "callout reads (5)"},
]})
pdf = tmp_path / "dummy.pdf"
pdf.write_bytes(b"%PDF-1.4\n")
report = run_agent_pipeline(str(pdf), out_dir=str(tmp_path),
require_review=False)
assert [f["issue_id"] for f in report["suppressed_issues"]] == ["C1"]
assert report["suppressed_issues"][0]["verification"]["status"] == "refuted"
def test_zero_image_finding_is_not_suppressed(monkeypatch, tmp_path):
"""A finding whose sheets resolve to no page images must not be judged
(and must never be refuted) without pixels."""
_patch_pipeline(monkeypatch, _finding(["S999"])) # no such sheet
monkeypatch.setattr(
"backend.agents.verifier.call_json",
lambda **kwargs: {"verdicts": [
{"sheet": "S999", "source_text": "(2) 2x6 STUD PACK",
"verdict": "not_found", "actual_text": None, "notes": None},
]})
pdf = tmp_path / "dummy.pdf"
pdf.write_bytes(b"%PDF-1.4\n")
report = run_agent_pipeline(str(pdf), out_dir=str(tmp_path),
require_review=False)
assert report["suppressed_issues"] == []
validated = report.get("validated_issues") or []
assert any(f.get("issue_id") == "C1" for f in validated)
+45 -7
View File
@@ -60,18 +60,56 @@ def test_job_log_endpoint_serves_log_and_404s(job_env, monkeypatch):
assert client.get("/jobs/nope/log").status_code == 404
def test_model_override_set_and_cleared_around_run(job_env, monkeypatch):
from backend import llm
def test_model_overrides_passed_to_classic_runner(job_env, monkeypatch):
"""Classic mode: per-run picks travel as run_pipeline kwargs (the runner
sets and clears llm.set_model_overrides itself)."""
seen = {}
def fake_runner(pdf_path, **kwargs):
seen["override"] = llm._model_override
seen.update(kwargs)
return {"source": "set.pdf", "summary": {}}
monkeypatch.setattr("backend.jobs.run_pipeline", fake_runner)
jobs.create_job(str(job_env / "set.pdf"), "set.pdf",
pipeline_mode="classic", model="openai/gpt-4o")
pipeline_mode="classic",
vision_model="openai/gpt-4o", text_model="openai/gpt-4o-mini")
assert seen["override"] == "openai/gpt-4o"
assert llm._model_override is None # cleared after the run
assert seen["vision_model"] == "openai/gpt-4o"
assert seen["text_model"] == "openai/gpt-4o-mini"
def test_model_overrides_set_and_cleared_around_agent_run(job_env, monkeypatch):
"""Agent mode: the agent runner has no override params, so jobs.py sets
them module-level for the duration of the run."""
from backend import llm
seen = {}
def fake_agent_runner(pdf_path, **kwargs):
seen["vision"] = llm._vision_model_override
seen["text"] = llm._text_model_override
return {"source": "set.pdf", "summary": {}}
monkeypatch.setattr("backend.jobs.run_agent_pipeline", fake_agent_runner)
jobs.create_job(str(job_env / "set.pdf"), "set.pdf",
pipeline_mode="agent",
vision_model="openai/gpt-4o", text_model="openai/gpt-4o-mini")
assert seen["vision"] == "openai/gpt-4o"
assert seen["text"] == "openai/gpt-4o-mini"
assert llm._vision_model_override is None # cleared after the run
assert llm._text_model_override is None
def test_failed_run_logs_traceback(job_env, monkeypatch):
"""A crashed job must leave the traceback in job.log, not just str(e)."""
def boom(pdf_path, **kwargs):
raise RuntimeError("kaboom-stage-failure")
monkeypatch.setattr("backend.jobs.run_pipeline", boom)
job_id = jobs.create_job(str(job_env / "set.pdf"), "set.pdf", pipeline_mode="classic")
assert jobs._jobs[job_id]["status"] == "error"
content = (job_env / job_id / "job.log").read_text()
assert "Traceback (most recent call last)" in content
assert "RuntimeError: kaboom-stage-failure" in content
+27 -3
View File
@@ -11,12 +11,23 @@ _PAYLOAD = {
"name": "GPT-4o",
"pricing": {"prompt": "0.0000025", "completion": "0.00001"},
"context_length": 128000,
"architecture": {"input_modalities": ["text", "image"],
"output_modalities": ["text"]},
},
{
"id": "google/gemini-2.5-pro",
"name": "Gemini 2.5 Pro",
"pricing": {"prompt": "0.00000125", "completion": "0.00001"},
"context_length": 1000000,
"architecture": {"modality": "text+image->text"},
},
{
"id": "meta-llama/llama-3.1-70b-instruct",
"name": "Llama 3.1 70B Instruct",
"pricing": {"prompt": "0.0000005", "completion": "0.0000008"},
"context_length": 131072,
"architecture": {"input_modalities": ["text"],
"output_modalities": ["text"]},
},
]
}
@@ -34,14 +45,27 @@ def test_models_endpoint_normalizes_pricing(monkeypatch):
response = client.get("/models")
assert response.status_code == 200
body = response.json()
assert body["default"] == config.MODEL
assert body["default_text"] == config.TEXT_MODEL
by_id = {m["id"]: m for m in body["models"]}
assert body["defaults"] == {"vision": config.MODEL, "text": config.TEXT_MODEL}
by_id = {m["id"]: m for m in body["text"]}
assert by_id["openai/gpt-4o"]["prompt_usd_per_mtok"] == 2.5
assert by_id["openai/gpt-4o"]["completion_usd_per_mtok"] == 10.0
assert by_id["openai/gpt-4o"]["context_length"] == 128000
def test_models_endpoint_splits_vision_and_text(monkeypatch):
_reset_cache()
monkeypatch.setattr(models, "_fetch_openrouter_models", lambda: _PAYLOAD["data"])
client = TestClient(app)
body = client.get("/models").json()
vision_ids = {m["id"] for m in body["vision"]}
text_ids = {m["id"] for m in body["text"]}
# Both modality shapes (structured and legacy string) are recognized.
assert vision_ids == {"openai/gpt-4o", "google/gemini-2.5-pro"}
# Text list is the full catalog; vision models appear in both.
assert text_ids == {"openai/gpt-4o", "google/gemini-2.5-pro",
"meta-llama/llama-3.1-70b-instruct"}
def test_models_endpoint_caches(monkeypatch):
_reset_cache()
calls = []
+23
View File
@@ -85,6 +85,29 @@ def test_finalize_confirm_keeps_confirmed(monkeypatch, tmp_path):
assert report["summary"]["agent_status"] == "complete"
def test_finalize_preserves_verifier_suppressed(monkeypatch, tmp_path):
"""Wave-5b (verifier) suppressions must survive review finalization and
merge with review-rejected suppressions."""
monkeypatch.setattr("backend.review.finalizer._draft_rfis", lambda kept: [])
_write_job(
str(tmp_path),
prioritized=[{"issue_id": "AGENT-0001", "severity": "high"}],
queue=[_blocking_item("AGENT-0001")],
decisions=[{"review_item_id": "finding:AGENT-0001",
"decision": "reject", "reason_code": "not_a_contradiction"}],
)
path = os.path.join(str(tmp_path), "conflicts.json")
with open(path, encoding="utf-8") as f:
report = json.load(f)
report["suppressed_issues"] = [
{"issue_id": "C1", "verification": {"status": "refuted"}}]
with open(path, "w", encoding="utf-8") as f:
json.dump(report, f)
final = finalize_review("job1", str(tmp_path))
ids = [f["issue_id"] for f in final["suppressed_issues"]]
assert ids == ["C1", "AGENT-0001"]
def test_finalize_no_decision_keeps_unreviewed(monkeypatch, tmp_path):
"""Non-blocking (audit) items don't need a decision; issue stays unreviewed."""
monkeypatch.setattr("backend.review.finalizer._draft_rfis", lambda kept: [])
+86
View File
@@ -0,0 +1,86 @@
"""Grounding-guard rescue tier, text-layer prompt block, and render hygiene."""
from backend import config
from backend.pipeline._stage import render
from backend.pipeline.extractor import (
_is_grounded,
_normalize_sheet,
_text_layer_block,
)
from backend.prompts import EXTRACTOR_USER_INSTRUCTION, VERIFY_USER_INSTRUCTION
PAGE_TEXT = "NOTES: (5) 2X6 STUD PACK AT BEARING. HSS16X4 BEAM. 7'-0\" AFF."
def _parsed(value, source_text):
return {
"sheet": {"sheet_number": "S401"},
"objects": [{
"object_id": "o1",
"object_type": "framing",
"name": "stud pack",
"attributes": {"count": value},
"source_text": source_text,
}],
}
def test_rescue_tier_keeps_and_stamps():
"""Digits absent from source_text but present in the page text layer:
kept, stamped grounding=text_layer (vision quoted imperfectly)."""
sheet = _normalize_sheet(_parsed("(2)", "(2) 2x6 STUD PACK"), 1,
page_text=PAGE_TEXT)
# "(2)" is not grounded by its own source_text alone? it is - use a value
# whose digits differ from the quote to exercise the rescue path.
sheet = _normalize_sheet(_parsed("5", "(2) 2x6 STUD PACK"), 1,
page_text=PAGE_TEXT)
assert len(sheet["assertions"]) == 1
assert sheet["assertions"][0]["grounding"] == "text_layer"
def test_no_rescue_without_page_text():
sheet = _normalize_sheet(_parsed("5", "(2) 2x6 STUD PACK"), 1)
assert sheet["assertions"] == []
def test_still_dropped_when_digits_nowhere():
sheet = _normalize_sheet(_parsed("99", "(2) 2x6 STUD PACK"), 1,
page_text=PAGE_TEXT)
assert sheet["assertions"] == []
def test_is_grounded_backward_compatible():
assert _is_grounded("(5)", "(5) 2x6 STUD PACK") is True
# Digit-run guard is a set check: "(3)" has no support anywhere.
assert _is_grounded("(3)", "(5) 2x6 STUD PACK") is False
assert _is_grounded("(3)", "(5) 2x6 STUD PACK",
page_text="(3) 2x6 STUD PACK") is True
def test_text_layer_block_empty_without_layer():
assert _text_layer_block({"page_number": 1}) == ""
assert _text_layer_block({"page_number": 1, "text_layer": None}) == ""
def test_text_layer_block_appends_and_caps(monkeypatch):
block = _text_layer_block({"page_number": 1, "text_layer": PAGE_TEXT})
assert "TEXT LAYER" in block and "STUD PACK" in block
monkeypatch.setattr(config, "TEXT_LAYER_MAX_CHARS", 50)
block = _text_layer_block({"page_number": 1, "text_layer": "x" * 500})
assert len(block.split(":\n", 1)[1]) == 50
def test_verify_instruction_fully_rendered():
"""render() silently leaves missing keys as literals - both placeholders
must be substituted at the (single) verify render site."""
out = render(VERIFY_USER_INSTRUCTION,
{"finding": "FINDING_JSON", "text_layer": "PAGE_TEXT"})
assert "{finding}" not in out and "{text_layer}" not in out
assert "FINDING_JSON" in out and "PAGE_TEXT" in out
def test_extractor_instruction_fully_substituted():
page = {"page_number": 1, "text_layer": PAGE_TEXT}
out = (EXTRACTOR_USER_INSTRUCTION.replace("{sheet_hint}", "")
+ _text_layer_block(page))
assert "{sheet_hint}" not in out
+33 -15
View File
@@ -1,26 +1,44 @@
from backend import config
from backend.llm import _resolve_backend, set_model_override
from backend.llm import _resolve_backend, set_model_overrides, set_text_backend
def test_override_wins_for_vision_and_text():
set_model_override("openai/gpt-4o")
try:
assert _resolve_backend(has_images=True, model_override=None)["model"] == "openai/gpt-4o"
assert _resolve_backend(has_images=False, model_override=None)["model"] == "openai/gpt-4o"
finally:
set_model_override(None)
def teardown_function():
set_model_overrides(None, None)
set_text_backend(False)
def test_vision_override_wins_for_vision_only():
set_model_overrides(vision="openai/gpt-4o", text=None)
assert _resolve_backend(has_images=True, model_override=None)["model"] == "openai/gpt-4o"
assert _resolve_backend(has_images=False, model_override=None)["model"] == config.TEXT_MODEL
def test_text_override_wins_for_text_only():
set_model_overrides(vision=None, text="anthropic/claude-sonnet-4")
assert _resolve_backend(has_images=False, model_override=None)["model"] == "anthropic/claude-sonnet-4"
assert _resolve_backend(has_images=True, model_override=None)["model"] == config.MODEL
def test_override_beats_per_call_model_arg():
set_model_override("openai/gpt-4o")
try:
# Agents pass their AGENT_*_MODEL per call; the user's job pick wins.
assert _resolve_backend(has_images=False, model_override="other/model")["model"] == "openai/gpt-4o"
finally:
set_model_override(None)
set_model_overrides(vision="openai/gpt-4o", text="openai/gpt-4o-mini")
# Agents pass their AGENT_*_MODEL per call; the user's job pick wins.
assert _resolve_backend(has_images=True, model_override="other/model")["model"] == "openai/gpt-4o"
assert _resolve_backend(has_images=False, model_override="other/model")["model"] == "openai/gpt-4o-mini"
def test_no_override_keeps_defaults():
set_model_override(None)
set_model_overrides(None, None)
assert _resolve_backend(has_images=True, model_override=None)["model"] == config.MODEL
assert _resolve_backend(has_images=False, model_override=None)["model"] == config.TEXT_MODEL
def test_ui_picks_never_name_the_local_model(monkeypatch):
"""Hybrid runs keep LOCAL_TEXT_MODEL; OpenRouter picks must not leak into
the local endpoint (a vLLM server won't serve OpenRouter model ids)."""
monkeypatch.setattr(config, "LOCAL_BASE_URL", "http://localhost:8000/v1")
monkeypatch.setattr(config, "LOCAL_TEXT_MODEL", "qwen/local-instruct")
set_text_backend(True)
set_model_overrides(vision="openai/gpt-4o", text="anthropic/claude-sonnet-4")
be = _resolve_backend(has_images=False, model_override=None)
assert be["local"] is True
assert be["model"] == "qwen/local-instruct"
+111
View File
@@ -0,0 +1,111 @@
"""Text-layer extraction, evidence bbox matching, and crop rendering."""
import os
import pytest
fitz = pytest.importorskip("pymupdf")
from backend import config
from backend.text_layer import (
attach_text_layers,
coverage_gaps,
extract_text_layers,
find_evidence_bbox,
render_crop,
)
EVIDENCE = "(5) 2X6 STUD PACK @ 16 IN O.C."
def _make_pdf(path, pages):
"""pages: list of str ('' = effectively blank page)."""
doc = fitz.open()
for text in pages:
page = doc.new_page(width=612, height=792)
if text:
page.insert_text((72, 72), text, fontsize=11)
doc.save(str(path))
doc.close()
return str(path)
@pytest.fixture
def text_pdf(tmp_path):
return _make_pdf(tmp_path / "set.pdf", [EVIDENCE, ""])
def test_extract_text_layers(text_pdf):
layers = extract_text_layers(text_pdf)
assert set(layers) == {1, 2}
assert layers[1]["has_text_layer"] is True
assert "2X6 STUD PACK" in layers[1]["text"]
assert layers[1]["words"], "expected word-level bboxes"
assert all("bbox" in w and len(w["bbox"]) == 4 for w in layers[1]["words"])
def test_blank_page_below_min_chars(text_pdf):
layers = extract_text_layers(text_pdf)
assert layers[2]["has_text_layer"] is False
def test_disabled_returns_empty(text_pdf, monkeypatch):
monkeypatch.setattr(config, "TEXT_LAYER_ENABLED", False)
assert extract_text_layers(text_pdf) == {}
def test_attach_text_layers(text_pdf, tmp_path):
pages = [{"page_number": 1}, {"page_number": 2}]
words = attach_text_layers(text_pdf, pages,
text_dir=str(tmp_path / "text"))
assert pages[0]["text_layer"] and "STUD PACK" in pages[0]["text_layer"]
assert pages[1]["text_layer"] is None
assert words[1] and not words[2]
assert os.path.isfile(tmp_path / "text" / "page-001.txt")
assert not os.path.exists(tmp_path / "text" / "page-002.txt")
def test_find_evidence_bbox_exact(text_pdf):
words = extract_text_layers(text_pdf)[1]["words"]
bbox = find_evidence_bbox(words, EVIDENCE)
assert bbox is not None
assert bbox[2] > bbox[0] and bbox[3] > bbox[1]
def test_find_evidence_bbox_fuzzy(text_pdf):
# Vision quotes imperfectly: wrong count token, rest exact.
words = extract_text_layers(text_pdf)[1]["words"]
bbox = find_evidence_bbox(words, "(2) 2X6 STUD PACK @ 16 IN O.C.")
assert bbox is not None
def test_find_evidence_bbox_miss(text_pdf):
words = extract_text_layers(text_pdf)[1]["words"]
assert find_evidence_bbox(words, "PENTHOUSE EXHAUST FAN EF-9") is None
assert find_evidence_bbox([], EVIDENCE) is None
assert find_evidence_bbox(words, "") is None
def test_render_crop(text_pdf):
words = extract_text_layers(text_pdf)[1]["words"]
bbox = find_evidence_bbox(words, EVIDENCE)
crop = render_crop(text_pdf, 1, bbox)
assert crop is not None
# Decodes as an image of plausible size (margin around the text line).
doc = fitz.open(stream=crop, filetype="jpeg")
pix = doc[0].get_pixmap()
assert pix.width > 100 and pix.height > 20
doc.close()
def test_render_crop_bad_page(text_pdf):
assert render_crop(text_pdf, 99, (0, 0, 10, 10)) is None
def test_coverage_gaps():
pages = [{"page_number": 1, "text_layer": "some real text"},
{"page_number": 2, "text_layer": "more text"},
{"page_number": 3, "text_layer": None}]
sheets = [{"page_number": 1, "assertions": [{"id": "a"}]},
{"page_number": 2, "assertions": []}]
assert coverage_gaps(pages, sheets) == [2]