Compare commits
34
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
23e19f53b2 | ||
|
|
46db871152 | ||
|
|
d37ac8c1c7 | ||
|
|
bae608a505 | ||
|
|
fe09e4a66b | ||
|
|
48fefa4007 | ||
|
|
23f6d7fe89 | ||
|
|
06e108142e | ||
|
|
0d109fb5cd | ||
|
|
570300324f | ||
|
|
349b357e5c | ||
|
|
b174b531cd | ||
|
|
3df359500c | ||
|
|
82952df307 | ||
|
|
c8af430143 | ||
|
|
f21eb5d912 | ||
|
|
4a3f33a245 | ||
|
|
15297038a2 | ||
|
|
d431a026ce | ||
|
|
8631a26006 | ||
|
|
df4d15fd0c | ||
|
|
0ea0b0e897 | ||
|
|
228e8bd031 | ||
|
|
3d7fce7bf9 | ||
|
|
76e0a52658 | ||
|
|
5c1fccfb35 | ||
|
|
32544bc2af | ||
|
|
4b3b62b3fa | ||
|
|
5305325d81 | ||
|
|
6f10062b93 | ||
|
|
7488cf68c5 | ||
|
|
f7e1b6bb7c | ||
|
|
bf508bfdf6 | ||
|
|
a6b0c8fdfa |
No files matched your search
@@ -0,0 +1,936 @@
|
|||||||
|
# Evidence Verification + Cross-Sheet Correlation Implementation Plan
|
||||||
|
|
||||||
|
> **For Hermes:** Use subagent-driven-development skill to implement this plan task-by-task.
|
||||||
|
|
||||||
|
**Goal:** Stop vision-extraction misreads (e.g. "(2) 2x6 STUD PACK" vs the actual "(5) 2x6") from becoming confident downstream findings, and correlate the same physical element across sheets (S101/S205/S401) so no stage reasons from one sheet's text in isolation.
|
||||||
|
|
||||||
|
**Architecture:** Three independently shippable phases on the agent pipeline (`backend/agents/runner.py`):
|
||||||
|
1. **Disputed-value detection** — deterministic post-link pass that flags contradictory extracted values inside a cluster and surfaces them to the critic/specialist prompts.
|
||||||
|
2. **Cross-sheet xref linking** — linker gains detail-reference/tag buckets that join assertions across levels (today `(level, family)` bucketing splits S101/S205/S401 apart).
|
||||||
|
3. **Evidence verification wave (5b)** — a bounded vision fact-check agent re-reads the cited sheet images for high-severity / disputed findings before the Brain merge, annotates or suppresses findings built on phantom text.
|
||||||
|
|
||||||
|
**Tech Stack:** Python 3.14, pytest (`tests/`), existing `call_json` LLM wrapper (supports `images_b64`, `reasoning_effort`, `reasoning_max_tokens`).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Current context / root cause (from job 959e16407573)
|
||||||
|
|
||||||
|
Finding `validated_issues[3]` ("FRONT PERSPECTIVE detail on Sheet S401", HSS16x4 on
|
||||||
|
"(2) 2x6 STUD PACK", severity critical) is a **false positive built on a wave-1 vision
|
||||||
|
misread**. The sheet actually shows a (5) 2x6 stud pack (matching S205/S101). Chain of failure:
|
||||||
|
|
||||||
|
1. Wave 1 (`SheetExtractorAgent`) froze the misread into text. From then on it is "ground truth".
|
||||||
|
2. Wave 3 linker (`backend/agents/linker.py:33` `build_link_scopes`) buckets by
|
||||||
|
`(level, family)`. S101 (foundation), S205 (details), S401 (sections) get different
|
||||||
|
`level` values, so assertions about the same front-wall header never share a link scope
|
||||||
|
or cluster. No cross-sheet corroboration happened.
|
||||||
|
3. Wave 5 `ConstructabilityAgent` (`backend/agents/construct_agent.py:53`) calls
|
||||||
|
`call_json` **with no images** — in this job 120/120 constructability calls were `+0img`.
|
||||||
|
It reasoned arithmetically from the misread text ("2 x 1.5in = 3in < 4in -> unbuildable").
|
||||||
|
It even held the "(5) 2x6 STUD PACK" assertion in the same scope but labeled it
|
||||||
|
"Ambiguous column size specification" instead of arbitrating.
|
||||||
|
4. Nothing between wave 5 and the report ever looks at a sheet image again. Only the wave-4
|
||||||
|
conflict critic receives images, and only for its own cluster's pages.
|
||||||
|
|
||||||
|
Also confirmed in this log (separate known bug, fixed in Task 7 while we're here): wave-4
|
||||||
|
conflict critic truncates on Gemini thinking tokens because `conflict_critic.py:59` passes
|
||||||
|
`max_tokens=config.REASON_MAX_TOKENS` (4096) with no reasoning budget — 13/121 calls hit
|
||||||
|
`finish_reason=length`.
|
||||||
|
|
||||||
|
## Assumptions
|
||||||
|
|
||||||
|
- Assertions carry `id`, `attribute`, `value`, `source_text`, `location_key`
|
||||||
|
(`room`/`grid`/`detail_reference`/`tag`/`level`) — see `linker._payload` and
|
||||||
|
`_serialize.slim_assertion`.
|
||||||
|
- Extractor assertions already carry a `confidence` field (per test fixtures).
|
||||||
|
- `validate_issue` in `backend/pipeline/_stage.py` guarantees each finding an `issue_id`.
|
||||||
|
- Test convention: `unittest.mock.patch("backend.agents.<module>.call_json", ...)` —
|
||||||
|
see `tests/agents/test_sheet_extractor_fallback.py`. Run tests with
|
||||||
|
`.venv/bin/python -m pytest tests/ -x -q`.
|
||||||
|
- `AgentResult.error` defaults to `""` (not None) in assertions.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 1 — Disputed-value detection + prompt hardening
|
||||||
|
|
||||||
|
### Task 1: `find_disputes` pure function (TDD)
|
||||||
|
|
||||||
|
**Objective:** Detect "same attribute, different values" inside one cluster's assertions.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Create: `backend/agents/disputes.py`
|
||||||
|
- Test: `tests/agents/test_disputes.py`
|
||||||
|
|
||||||
|
**Step 1: Write failing test**
|
||||||
|
|
||||||
|
```python
|
||||||
|
# tests/agents/test_disputes.py
|
||||||
|
from backend.agents.disputes import annotate_clusters, find_disputes
|
||||||
|
|
||||||
|
|
||||||
|
def _a(id_, attribute, value):
|
||||||
|
return {"id": id_, "attribute": attribute, "value": value,
|
||||||
|
"source_text": value}
|
||||||
|
|
||||||
|
|
||||||
|
def test_find_disputes_flags_same_attribute_different_values():
|
||||||
|
assertions = [
|
||||||
|
_a("a1", "stud_pack_size", "(2) 2x6 STUD PACK"),
|
||||||
|
_a("a2", "stud_pack_size", "(5) 2x6 STUD PACK"),
|
||||||
|
_a("a3", "beam_size", "HSS16X4X5/8"),
|
||||||
|
]
|
||||||
|
disputes = find_disputes(assertions)
|
||||||
|
assert len(disputes) == 1
|
||||||
|
assert disputes[0]["attribute"] == "stud_pack_size"
|
||||||
|
assert disputes[0]["values"] == ["(2) 2x6 STUD PACK", "(5) 2x6 STUD PACK"]
|
||||||
|
assert disputes[0]["assertion_ids"] == ["a1", "a2"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_find_disputes_ignores_agreeing_values_and_blanks():
|
||||||
|
assertions = [
|
||||||
|
_a("a1", "beam_size", "HSS16X4X5/8"),
|
||||||
|
_a("a2", "beam_size", " hss16x4x5/8 "), # same after normalize
|
||||||
|
_a("a3", "", "orphan"), # no attribute -> skipped
|
||||||
|
_a("a4", "beam_size", ""), # no value -> skipped
|
||||||
|
]
|
||||||
|
assert find_disputes(assertions) == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_annotate_clusters_writes_disputed_attributes():
|
||||||
|
clusters = [
|
||||||
|
{"key": "c1", "assertions": [
|
||||||
|
_a("a1", "stud_pack_size", "(2) 2x6"),
|
||||||
|
_a("a2", "stud_pack_size", "(5) 2x6"),
|
||||||
|
]},
|
||||||
|
{"key": "c2", "assertions": [_a("a3", "x", "1"), _a("a4", "x", "1")]},
|
||||||
|
]
|
||||||
|
assert annotate_clusters(clusters) == 1
|
||||||
|
assert clusters[0]["disputed_attributes"][0]["attribute"] == "stud_pack_size"
|
||||||
|
assert "disputed_attributes" not in clusters[1]
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 2: Run test to verify failure**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/agents/test_disputes.py -v`
|
||||||
|
Expected: FAIL — `ModuleNotFoundError: backend.agents.disputes`
|
||||||
|
|
||||||
|
**Step 3: Implement**
|
||||||
|
|
||||||
|
```python
|
||||||
|
# backend/agents/disputes.py
|
||||||
|
"""Deterministic detection of contradictory extracted values within a cluster.
|
||||||
|
|
||||||
|
Extraction is a vision pass: quantities and sizes can be misread ("(2) 2x6" vs
|
||||||
|
"(5) 2x6"). Cluster members are supposed to describe the same real-world
|
||||||
|
element, so two members asserting different values for the same attribute are
|
||||||
|
a probable misread. Flag these so downstream text-only stages treat the value
|
||||||
|
as unverified instead of reasoning from one reading.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import re
|
||||||
|
from typing import Dict, List
|
||||||
|
|
||||||
|
|
||||||
|
def _norm(value) -> str:
|
||||||
|
return re.sub(r"\s+", " ", str(value or "").strip().lower())
|
||||||
|
|
||||||
|
|
||||||
|
def find_disputes(assertions: List[Dict]) -> List[Dict]:
|
||||||
|
"""Same attribute with >= 2 distinct normalized values = disputed."""
|
||||||
|
groups: Dict[str, Dict[str, set]] = {}
|
||||||
|
for assertion in assertions:
|
||||||
|
attribute = _norm(assertion.get("attribute"))
|
||||||
|
value = _norm(assertion.get("value"))
|
||||||
|
if not attribute or not value:
|
||||||
|
continue
|
||||||
|
groups.setdefault(attribute, {}).setdefault(value, set()).add(
|
||||||
|
assertion.get("id")
|
||||||
|
)
|
||||||
|
disputes = []
|
||||||
|
for attribute, values in sorted(groups.items()):
|
||||||
|
if len(values) < 2:
|
||||||
|
continue
|
||||||
|
disputes.append({
|
||||||
|
"attribute": attribute,
|
||||||
|
"values": sorted(values),
|
||||||
|
"assertion_ids": sorted(
|
||||||
|
aid for ids in values.values() for aid in ids if aid
|
||||||
|
),
|
||||||
|
})
|
||||||
|
return disputes
|
||||||
|
|
||||||
|
|
||||||
|
def annotate_clusters(clusters: List[Dict]) -> int:
|
||||||
|
"""Attach disputed_attributes to each cluster that has any. Returns count."""
|
||||||
|
annotated = 0
|
||||||
|
for cluster in clusters:
|
||||||
|
disputes = find_disputes(cluster.get("assertions") or [])
|
||||||
|
if disputes:
|
||||||
|
cluster["disputed_attributes"] = disputes
|
||||||
|
annotated += 1
|
||||||
|
return annotated
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 4: Run test to verify pass**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/agents/test_disputes.py -v`
|
||||||
|
Expected: 3 passed
|
||||||
|
|
||||||
|
**Step 5: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add backend/agents/disputes.py tests/agents/test_disputes.py
|
||||||
|
git commit -m "feat: deterministic disputed-value detection for cluster assertions"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Task 2: Wire `annotate_clusters` into the runner + serialization
|
||||||
|
|
||||||
|
**Objective:** Disputes must be visible to the wave-4 critic and wave-5 constructability prompts.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `backend/agents/runner.py` (after `memory.replace("clusters", clusters)`, ~line 116)
|
||||||
|
- Modify: `backend/pipeline/_serialize.py` (`slim_clusters`, line 46)
|
||||||
|
- Test: `tests/agents/test_disputes.py` (append)
|
||||||
|
|
||||||
|
**Step 1: Write failing test**
|
||||||
|
|
||||||
|
```python
|
||||||
|
def test_slim_clusters_preserves_disputed_attributes():
|
||||||
|
from backend.pipeline._serialize import slim_clusters
|
||||||
|
cluster = {"key": "c1", "assertions": [],
|
||||||
|
"disputed_attributes": [{"attribute": "a", "values": ["1", "2"],
|
||||||
|
"assertion_ids": ["x", "y"]}]}
|
||||||
|
slim = slim_clusters([cluster])[0]
|
||||||
|
assert slim["disputed_attributes"][0]["values"] == ["1", "2"]
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 2: Run test to verify failure**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/agents/test_disputes.py::test_slim_clusters_preserves_disputed_attributes -v`
|
||||||
|
Expected: FAIL — `KeyError: 'disputed_attributes'`
|
||||||
|
|
||||||
|
**Step 3: Implement**
|
||||||
|
|
||||||
|
In `backend/pipeline/_serialize.py` `slim_clusters`, add the key:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def slim_clusters(clusters: List[Dict]) -> List[Dict]:
|
||||||
|
return [
|
||||||
|
{
|
||||||
|
"key": c.get("key"),
|
||||||
|
"location": c.get("location"),
|
||||||
|
"disciplines": c.get("disciplines"),
|
||||||
|
"kind": c.get("kind"),
|
||||||
|
**({"disputed_attributes": c["disputed_attributes"]}
|
||||||
|
if c.get("disputed_attributes") else {}),
|
||||||
|
"assertions": [slim_assertion(a) for a in c.get("assertions", [])],
|
||||||
|
}
|
||||||
|
for c in clusters
|
||||||
|
]
|
||||||
|
```
|
||||||
|
|
||||||
|
In `backend/agents/runner.py`, right after `clusters = [...]` / `object_graph = build_object_graph(clusters)` (before `memory.replace("clusters", clusters)`):
|
||||||
|
|
||||||
|
```python
|
||||||
|
from backend.agents.disputes import annotate_clusters
|
||||||
|
...
|
||||||
|
object_graph = build_object_graph(clusters)
|
||||||
|
disputed_count = annotate_clusters(clusters)
|
||||||
|
if disputed_count:
|
||||||
|
orchestrator.log(
|
||||||
|
f"[Link] {disputed_count} clusters carry disputed extracted values"
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
(Check `Orchestrator` for the actual log method name — `orchestrator.stage(...)` exists;
|
||||||
|
if no `.log`, use the module's existing logging/print convention. Adjust to match.)
|
||||||
|
|
||||||
|
**Step 4: Run tests**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/agents/ -v`
|
||||||
|
Expected: all pass (including existing `test_runner_review_gate.py`)
|
||||||
|
|
||||||
|
**Step 5: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add backend/agents/runner.py backend/pipeline/_serialize.py tests/agents/test_disputes.py
|
||||||
|
git commit -m "feat: surface disputed extracted values to critic and specialist prompts"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Task 3: Prompt hardening — extracted text is fallible
|
||||||
|
|
||||||
|
**Objective:** Tell text-only specialists how to handle disputed/unverified values so they stop asserting buildability conclusions from a single (possibly misread) number.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `backend/prompts.py` `CONSTRUCTABILITY_SYSTEM_PROMPT` (line 533) and `CONSTRUCTABILITY_USER_INSTRUCTION` (line 551)
|
||||||
|
|
||||||
|
**Step 1: Edit prompts**
|
||||||
|
|
||||||
|
Append to `CONSTRUCTABILITY_SYSTEM_PROMPT` Rules list (after line 547, before "Use plain ASCII"):
|
||||||
|
|
||||||
|
```
|
||||||
|
- Assertions are machine-extracted from sheet images and may contain misread values,
|
||||||
|
especially quantities and member sizes (e.g. "(2) 2x6" vs "(5) 2x6").
|
||||||
|
- When the cluster lists disputed_attributes, or two evidence items disagree on a
|
||||||
|
numeric value, do NOT assert a buildability conclusion from one reading. Report the
|
||||||
|
ambiguity itself (category "detail_gap", confidence "low") and state that the value
|
||||||
|
needs verification against the sheet.
|
||||||
|
```
|
||||||
|
|
||||||
|
Append to `CONSTRUCTABILITY_USER_INSTRUCTION` after the `Cross-discipline conflicts already found: {conflicts}` line:
|
||||||
|
|
||||||
|
```
|
||||||
|
Disputed extracted values in this cluster (possible vision misreads - treat as unverified): {disputes}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 2: Wire the `{disputes}` placeholder in `construct_agent.py`**
|
||||||
|
|
||||||
|
In `backend/agents/construct_agent.py` `run()`, extend the `substitutions` dict:
|
||||||
|
|
||||||
|
```python
|
||||||
|
substitutions = {
|
||||||
|
"assertions": dumps(cluster["assertions"]),
|
||||||
|
"clusters": dumps(slim_clusters([cluster])),
|
||||||
|
"conflicts": dumps(scope.payload.get("conflicts") or []),
|
||||||
|
"disputes": dumps(cluster.get("disputed_attributes") or []),
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 3: Run full test suite (prompt edits can break runner tests that snapshot prompts)**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/ -q`
|
||||||
|
Expected: all pass
|
||||||
|
|
||||||
|
**Step 4: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add backend/prompts.py backend/agents/construct_agent.py
|
||||||
|
git commit -m "feat: constructability prompt treats disputed extracted values as unverified"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 2 — Cross-sheet xref linking
|
||||||
|
|
||||||
|
### Task 4: detail-reference / tag xref buckets in the linker (TDD)
|
||||||
|
|
||||||
|
**Objective:** Assertions sharing a `detail_reference` or a member `tag` get linked across levels, so S101/S205/S401 details of the same physical element land in one scope.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `backend/agents/linker.py` (`build_link_scopes`, line 33)
|
||||||
|
- Test: `tests/agents/test_linker_xref.py`
|
||||||
|
|
||||||
|
**Step 1: Write failing test**
|
||||||
|
|
||||||
|
```python
|
||||||
|
# tests/agents/test_linker_xref.py
|
||||||
|
from backend.agents.base import AgentScope
|
||||||
|
from backend.agents.linker import build_link_scopes
|
||||||
|
|
||||||
|
|
||||||
|
def _sheet(number, page, level, assertions):
|
||||||
|
return {"sheet_number": number, "page_number": page,
|
||||||
|
"discipline": "Structural", "level": level,
|
||||||
|
"assertions": assertions}
|
||||||
|
|
||||||
|
|
||||||
|
def _assertion(id_, ref=None, tag=None, level=None):
|
||||||
|
return {"id": id_, "attribute": "stud_pack_size", "value": "(5) 2x6",
|
||||||
|
"source_text": "(5) 2x6 STUD PACK",
|
||||||
|
"location_key": {"detail_reference": ref, "tag": tag,
|
||||||
|
"level": level}}
|
||||||
|
|
||||||
|
|
||||||
|
def test_xref_scope_joins_same_detail_reference_across_levels():
|
||||||
|
sheets = [
|
||||||
|
_sheet("S101", 10, "foundation", [_assertion("a1", ref="A/S205")]),
|
||||||
|
_sheet("S205", 20, "roof", [_assertion("a2", ref="A/S205")]),
|
||||||
|
_sheet("S401", 30, "roof", [_assertion("a3", ref="A/S205")]),
|
||||||
|
]
|
||||||
|
scopes = build_link_scopes(sheets)
|
||||||
|
xref = [s for s in scopes if s.scope_id.startswith("xref:")]
|
||||||
|
assert xref, "expected a cross-level detail-reference scope"
|
||||||
|
ids = {a["id"] for s in xref for a in s.payload["assertions"]}
|
||||||
|
assert ids == {"a1", "a2", "a3"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_xref_scope_requires_two_distinct_sheets():
|
||||||
|
sheets = [
|
||||||
|
_sheet("S401", 30, "roof", [_assertion("a1", ref="A/S205"),
|
||||||
|
_assertion("a2", ref="A/S205")]),
|
||||||
|
]
|
||||||
|
scopes = build_link_scopes(sheets)
|
||||||
|
assert not [s for s in scopes if s.scope_id.startswith("xref:")]
|
||||||
|
|
||||||
|
|
||||||
|
def test_xref_scope_joins_shared_member_tag():
|
||||||
|
sheets = [
|
||||||
|
_sheet("S102", 5, "roof", [_assertion("a1", tag="HSS16X4X5/8")]),
|
||||||
|
_sheet("S401", 30, "unknown", [_assertion("a2", tag="HSS16X4X5/8")]),
|
||||||
|
]
|
||||||
|
scopes = build_link_scopes(sheets)
|
||||||
|
xref = [s for s in scopes if s.scope_id.startswith("xref:")]
|
||||||
|
assert xref
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 2: Run test to verify failure**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/agents/test_linker_xref.py -v`
|
||||||
|
Expected: FAIL — no `xref:` scopes produced
|
||||||
|
|
||||||
|
**Step 3: Implement**
|
||||||
|
|
||||||
|
Rewrite `build_link_scopes` in `backend/agents/linker.py` (keep the existing
|
||||||
|
`(level, family)` bucketing, add the xref pass):
|
||||||
|
|
||||||
|
```python
|
||||||
|
def _xref_keys(assertion: Dict) -> List[str]:
|
||||||
|
"""Cross-level join keys: detail references and member tags."""
|
||||||
|
location = assertion.get("location_key") or {}
|
||||||
|
keys = []
|
||||||
|
ref = re.sub(r"\s+", "", str(location.get("detail_reference") or "")).upper()
|
||||||
|
if ref:
|
||||||
|
keys.append(f"detail:{ref}")
|
||||||
|
tag = re.sub(r"\s+", "", str(location.get("tag") or "")).upper()
|
||||||
|
if re.match(r"^[A-Z]{2,}\d", tag): # member marks: HSS16X4X5/8, W12X26, ...
|
||||||
|
keys.append(f"tag:{tag}")
|
||||||
|
return keys
|
||||||
|
|
||||||
|
|
||||||
|
def build_link_scopes(sheets: List[Dict]) -> List[AgentScope]:
|
||||||
|
"""Partition facts by level and object/tag family, then enforce a hard cap.
|
||||||
|
|
||||||
|
A second pass joins assertions that share a detail_reference or member tag
|
||||||
|
ACROSS levels, so plan/detail/section sheets describing the same physical
|
||||||
|
element are linked together even though their levels differ.
|
||||||
|
"""
|
||||||
|
buckets: Dict[Tuple[str, str], List[Dict]] = defaultdict(list)
|
||||||
|
xref: Dict[str, List[Dict]] = defaultdict(list)
|
||||||
|
for sheet in sheets:
|
||||||
|
for assertion in sheet.get("assertions", []):
|
||||||
|
enriched = {
|
||||||
|
**assertion,
|
||||||
|
"discipline": sheet.get("discipline") or "Unknown",
|
||||||
|
"sheet_number": sheet.get("sheet_number"),
|
||||||
|
"page_number": sheet.get("page_number"),
|
||||||
|
}
|
||||||
|
level = str((assertion.get("location_key") or {}).get("level")
|
||||||
|
or sheet.get("level") or "unknown").lower()
|
||||||
|
buckets[(level, _family(assertion))].append(enriched)
|
||||||
|
for key in _xref_keys(assertion):
|
||||||
|
xref[key].append(enriched)
|
||||||
|
|
||||||
|
scopes: List[AgentScope] = []
|
||||||
|
cap = max(2, config.AGENT_LINK_MAX_ASSERTIONS)
|
||||||
|
for (level, family), assertions in sorted(buckets.items()):
|
||||||
|
for offset in range(0, len(assertions), cap):
|
||||||
|
chunk = assertions[offset:offset + cap]
|
||||||
|
if len(chunk) < 2:
|
||||||
|
continue
|
||||||
|
scopes.append(AgentScope(
|
||||||
|
scope_id=f"{level}:{family}:{offset // cap + 1}",
|
||||||
|
payload={"assertions": chunk, "level": level, "family": family},
|
||||||
|
))
|
||||||
|
for key, assertions in sorted(xref.items()):
|
||||||
|
sheets_present = {a.get("sheet_number") for a in assertions}
|
||||||
|
if len(assertions) < 2 or len(sheets_present) < 2:
|
||||||
|
continue
|
||||||
|
scopes.append(AgentScope(
|
||||||
|
scope_id=f"xref:{key}",
|
||||||
|
payload={"assertions": assertions[:cap],
|
||||||
|
"level": "xref", "family": key},
|
||||||
|
))
|
||||||
|
return scopes
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 4: Run tests**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/agents/test_linker_xref.py tests/agents/ -v`
|
||||||
|
Expected: all pass (watch existing runner tests for scope-count coupling)
|
||||||
|
|
||||||
|
**Step 5: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add backend/agents/linker.py tests/agents/test_linker_xref.py
|
||||||
|
git commit -m "feat: cross-level xref link scopes via detail_reference and member tag"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 3 — Evidence verification wave (5b)
|
||||||
|
|
||||||
|
### Task 5: Config knobs + verify prompts
|
||||||
|
|
||||||
|
**Objective:** Add the tuning surface and prompts for the vision fact-check agent.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `backend/config.py` (near line 44, with the other AGENT_* knobs)
|
||||||
|
- Modify: `backend/prompts.py` (append near the CONFLICT prompts, ~line 460)
|
||||||
|
- Modify: `backend/.env.example`
|
||||||
|
|
||||||
|
**Step 1: Add config knobs to `backend/config.py`**
|
||||||
|
|
||||||
|
```python
|
||||||
|
AGENT_VERIFY_MODEL = os.getenv("AGENT_VERIFY_MODEL", "") or MODEL
|
||||||
|
AGENT_VERIFY_CONCURRENCY = int(os.getenv("AGENT_VERIFY_CONCURRENCY", "4"))
|
||||||
|
AGENT_VERIFY_MAX_CHECKS = int(os.getenv("AGENT_VERIFY_MAX_CHECKS", "20"))
|
||||||
|
AGENT_VERIFY_SEVERITIES = {
|
||||||
|
s.strip().lower()
|
||||||
|
for s in os.getenv("AGENT_VERIFY_SEVERITIES", "critical,high").split(",")
|
||||||
|
if s.strip()
|
||||||
|
}
|
||||||
|
AGENT_VERIFY_REASONING_EFFORT = os.getenv("AGENT_VERIFY_REASONING_EFFORT", "low").strip()
|
||||||
|
VERIFY_MAX_TOKENS = int(os.getenv("VERIFY_MAX_TOKENS", "8192"))
|
||||||
|
```
|
||||||
|
|
||||||
|
Append to `backend/.env.example`:
|
||||||
|
|
||||||
|
```
|
||||||
|
# Wave 5b evidence verification (vision fact-check of cited sheet text)
|
||||||
|
AGENT_VERIFY_MAX_CHECKS=20
|
||||||
|
AGENT_VERIFY_SEVERITIES=critical,high
|
||||||
|
AGENT_VERIFY_REASONING_EFFORT=low
|
||||||
|
VERIFY_MAX_TOKENS=8192
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 2: Add prompts to `backend/prompts.py`**
|
||||||
|
|
||||||
|
```python
|
||||||
|
VERIFY_SYSTEM_PROMPT = """You are a meticulous construction document checker verifying machine-extracted evidence against the actual drawing sheet images.
|
||||||
|
For each evidence item you are given the sheet it was extracted from and the verbatim text the extractor claims appears there.
|
||||||
|
Judge each item against the images:
|
||||||
|
- confirmed: the text (or an obvious equivalent) appears on the cited sheet and means what the finding claims.
|
||||||
|
- corrected: the sheet shows a DIFFERENT value than the extracted text. Give the actual verbatim text.
|
||||||
|
- not_found: nothing like the extracted text appears on the cited sheet.
|
||||||
|
Be strict about numbers, quantities, and member sizes: "(2) 2x6" and "(5) 2x6" are different values. HSS16x4 and HSS16x16 are different values.
|
||||||
|
Use plain ASCII only.
|
||||||
|
Respond only with valid JSON."""
|
||||||
|
|
||||||
|
VERIFY_USER_INSTRUCTION = """Verify this finding's evidence against the attached sheet images.
|
||||||
|
Respond ONLY with a valid JSON object - no markdown fences, no explanation:
|
||||||
|
{ "verdicts": [ { "sheet": "string", "source_text": "the evidence text judged", "verdict": "confirmed | corrected | not_found", "actual_text": "verbatim sheet text when corrected, else null", "notes": "string or null" } ] }
|
||||||
|
Finding: {finding}"""
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 3: Sanity check**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -c "from backend import config, prompts; print(config.AGENT_VERIFY_MAX_CHECKS, config.AGENT_VERIFY_SEVERITIES); print(prompts.VERIFY_SYSTEM_PROMPT[:40])"`
|
||||||
|
Expected: `20 {'critical', 'high'}` and prompt text
|
||||||
|
|
||||||
|
**Step 4: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add backend/config.py backend/prompts.py backend/.env.example
|
||||||
|
git commit -m "feat: config knobs and prompts for evidence verification wave"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Task 6: `EvidenceVerifierAgent` + runner wave 5b (TDD)
|
||||||
|
|
||||||
|
**Objective:** Re-read cited sheet images for selected findings; annotate verified findings, suppress refuted ones before the Brain merge.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Create: `backend/agents/verifier.py`
|
||||||
|
- Modify: `backend/agents/runner.py` (new wave between wave 5 and wave 6, ~line 167)
|
||||||
|
- Modify: `backend/agents/construct_agent.py` line 67 (stamp `cluster_key` for dispute-based selection)
|
||||||
|
- Test: `tests/agents/test_verifier.py`
|
||||||
|
|
||||||
|
**Step 1: Write failing test**
|
||||||
|
|
||||||
|
```python
|
||||||
|
# tests/agents/test_verifier.py
|
||||||
|
from unittest.mock import patch
|
||||||
|
|
||||||
|
from backend.agents.base import AgentScope, AgentUsage
|
||||||
|
from backend.agents.verifier import (
|
||||||
|
EvidenceVerifierAgent, apply_verdicts, select_findings,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _finding(sev="critical", issue_id="i1", sheets=("S401",), cluster_key=None):
|
||||||
|
f = {"issue_id": issue_id, "severity": sev, "confidence": "high",
|
||||||
|
"source_stage": "constructability", "sheets": list(sheets),
|
||||||
|
"description": "HSS16x4 on (2) 2x6 STUD PACK is unbuildable",
|
||||||
|
"evidence": [{"sheet": "S401", "source_text": "(2) 2x6 STUD PACK",
|
||||||
|
"asserted_value": "3-inch width"}]}
|
||||||
|
if cluster_key:
|
||||||
|
f["cluster_key"] = cluster_key
|
||||||
|
return f
|
||||||
|
|
||||||
|
|
||||||
|
def test_select_findings_by_severity_and_dispute():
|
||||||
|
findings = [_finding("critical"), _finding("low", "i2"),
|
||||||
|
_finding("medium", "i3", cluster_key="c9")]
|
||||||
|
clusters = [{"key": "c9", "disputed_attributes": [{"attribute": "a"}]}]
|
||||||
|
selected = select_findings(findings, clusters, max_checks=20,
|
||||||
|
severities={"critical", "high"})
|
||||||
|
assert [f["issue_id"] for f in selected] == ["i1", "i3"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_select_findings_respects_cap():
|
||||||
|
findings = [_finding("critical", f"i{n}") for n in range(30)]
|
||||||
|
selected = select_findings(findings, [], max_checks=5,
|
||||||
|
severities={"critical"})
|
||||||
|
assert len(selected) == 5
|
||||||
|
|
||||||
|
|
||||||
|
def test_run_attaches_verdicts_and_marks_refuted():
|
||||||
|
agent = EvidenceVerifierAgent(usage=AgentUsage())
|
||||||
|
scope = AgentScope(scope_id="verify:0", payload={
|
||||||
|
"finding_index": 0,
|
||||||
|
"finding": _finding(),
|
||||||
|
"images_b64": ["QUJD"],
|
||||||
|
})
|
||||||
|
verdicts = {"verdicts": [
|
||||||
|
{"sheet": "S401", "source_text": "(2) 2x6 STUD PACK",
|
||||||
|
"verdict": "corrected", "actual_text": "(5) 2x6 STUD PACK",
|
||||||
|
"notes": "callout reads (5)"},
|
||||||
|
]}
|
||||||
|
with patch("backend.agents.verifier.call_json", return_value=verdicts):
|
||||||
|
result = agent.run(scope)
|
||||||
|
assert not result.error
|
||||||
|
artifact = result.artifacts[0]
|
||||||
|
assert artifact["finding_index"] == 0
|
||||||
|
assert artifact["status"] == "refuted" # no evidence confirmed
|
||||||
|
assert artifact["verdicts"][0]["actual_text"] == "(5) 2x6 STUD PACK"
|
||||||
|
|
||||||
|
|
||||||
|
def test_apply_verdicts_annotates_and_suppresses():
|
||||||
|
findings = [_finding("critical", "i1"), _finding("high", "i2")]
|
||||||
|
from backend.agents.base import AgentResult
|
||||||
|
results = [AgentResult(scope_id="verify:0", artifacts=[
|
||||||
|
{"finding_index": 0, "status": "refuted", "verdicts": []},
|
||||||
|
{"finding_index": 1, "status": "confirmed", "verdicts": []},
|
||||||
|
])]
|
||||||
|
suppressed = apply_verdicts(findings, results)
|
||||||
|
assert suppressed == [findings[0]]
|
||||||
|
assert findings[0]["verification"]["status"] == "refuted"
|
||||||
|
assert findings[1]["verification"]["status"] == "confirmed"
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 2: Run test to verify failure**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/agents/test_verifier.py -v`
|
||||||
|
Expected: FAIL — `ModuleNotFoundError: backend.agents.verifier`
|
||||||
|
|
||||||
|
**Step 3: Implement `backend/agents/verifier.py`**
|
||||||
|
|
||||||
|
```python
|
||||||
|
"""Wave 5b: vision fact-check of extracted evidence against cited sheet images.
|
||||||
|
|
||||||
|
Downstream specialists are text-only; a wave-1 vision misread ("(2) 2x6" vs
|
||||||
|
"(5) 2x6") otherwise becomes immutable ground truth. For high-severity or
|
||||||
|
dispute-linked findings, re-read the cited sheets and adjudicate each evidence
|
||||||
|
item: confirmed / corrected / not_found. Findings whose evidence is entirely
|
||||||
|
unconfirmed are suppressed before the Brain merge.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from typing import Dict, List, Optional, Set
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
|
||||||
|
from backend.llm import call_json
|
||||||
|
from backend.pipeline._serialize import dumps
|
||||||
|
from backend.pipeline._stage import collect_list, render
|
||||||
|
from backend.prompts import VERIFY_SYSTEM_PROMPT, VERIFY_USER_INSTRUCTION
|
||||||
|
|
||||||
|
_SEVERITY_RANK = {"critical": 0, "high": 1, "medium": 2, "low": 3}
|
||||||
|
_VERDICTS = ("confirmed", "corrected", "not_found")
|
||||||
|
|
||||||
|
|
||||||
|
def select_findings(
|
||||||
|
findings: List[Dict],
|
||||||
|
clusters: List[Dict],
|
||||||
|
max_checks: int,
|
||||||
|
severities: Set[str],
|
||||||
|
) -> List[Dict]:
|
||||||
|
"""Severity-gated selection plus any finding tied to a disputed cluster."""
|
||||||
|
disputed_keys = {
|
||||||
|
cluster.get("key") for cluster in clusters
|
||||||
|
if cluster.get("disputed_attributes")
|
||||||
|
}
|
||||||
|
selected = [
|
||||||
|
finding for finding in findings
|
||||||
|
if str(finding.get("severity") or "").lower() in severities
|
||||||
|
or finding.get("cluster_key") in disputed_keys
|
||||||
|
]
|
||||||
|
selected.sort(key=lambda f: _SEVERITY_RANK.get(
|
||||||
|
str(f.get("severity") or "").lower(), 9))
|
||||||
|
return selected[:max_checks]
|
||||||
|
|
||||||
|
|
||||||
|
def _valid_verdict(item: Dict) -> Optional[Dict]:
|
||||||
|
if not isinstance(item, dict):
|
||||||
|
return None
|
||||||
|
verdict = str(item.get("verdict") or "").lower()
|
||||||
|
if verdict not in _VERDICTS:
|
||||||
|
return None
|
||||||
|
return {
|
||||||
|
"sheet": item.get("sheet") or "",
|
||||||
|
"source_text": item.get("source_text") or "",
|
||||||
|
"verdict": verdict,
|
||||||
|
"actual_text": item.get("actual_text"),
|
||||||
|
"notes": item.get("notes"),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _status(verdicts: List[Dict]) -> str:
|
||||||
|
if not verdicts:
|
||||||
|
return "unverified"
|
||||||
|
confirmed = sum(1 for v in verdicts if v["verdict"] == "confirmed")
|
||||||
|
if confirmed == len(verdicts):
|
||||||
|
return "confirmed"
|
||||||
|
if confirmed == 0:
|
||||||
|
return "refuted"
|
||||||
|
return "mixed"
|
||||||
|
|
||||||
|
|
||||||
|
class EvidenceVerifierAgent:
|
||||||
|
name = "verify"
|
||||||
|
|
||||||
|
def __init__(self, usage: AgentUsage) -> None:
|
||||||
|
self.usage = usage
|
||||||
|
|
||||||
|
def run(self, scope: AgentScope) -> AgentResult:
|
||||||
|
try:
|
||||||
|
finding = scope.payload["finding"]
|
||||||
|
instruction = render(
|
||||||
|
VERIFY_USER_INSTRUCTION, {"finding": dumps(finding)}
|
||||||
|
)
|
||||||
|
parsed = call_json(
|
||||||
|
system_prompt=VERIFY_SYSTEM_PROMPT,
|
||||||
|
user_text=instruction,
|
||||||
|
images_b64=scope.payload.get("images_b64") or [],
|
||||||
|
max_tokens=config.VERIFY_MAX_TOKENS,
|
||||||
|
model=config.AGENT_VERIFY_MODEL,
|
||||||
|
reasoning_effort=config.AGENT_VERIFY_REASONING_EFFORT or None,
|
||||||
|
usage_tracker=self.usage,
|
||||||
|
usage_stage="agent.verify",
|
||||||
|
)
|
||||||
|
verdicts = collect_list(parsed, "verdicts", _valid_verdict)
|
||||||
|
return AgentResult(scope_id=scope.scope_id, artifacts=[{
|
||||||
|
"finding_index": scope.payload["finding_index"],
|
||||||
|
"status": _status(verdicts),
|
||||||
|
"verdicts": verdicts,
|
||||||
|
}])
|
||||||
|
except Exception as exc:
|
||||||
|
return failure(scope, exc)
|
||||||
|
|
||||||
|
|
||||||
|
def apply_verdicts(
|
||||||
|
findings: List[Dict], verify_results: List[AgentResult]
|
||||||
|
) -> List[Dict]:
|
||||||
|
"""Annotate findings with verification; return refuted ones to suppress."""
|
||||||
|
by_index: Dict[int, Dict] = {}
|
||||||
|
for result in verify_results:
|
||||||
|
for artifact in result.artifacts:
|
||||||
|
by_index[artifact["finding_index"]] = artifact
|
||||||
|
suppressed = []
|
||||||
|
for index, finding in enumerate(findings):
|
||||||
|
artifact = by_index.get(index)
|
||||||
|
if not artifact:
|
||||||
|
continue
|
||||||
|
finding["verification"] = {
|
||||||
|
"status": artifact["status"],
|
||||||
|
"verdicts": artifact["verdicts"],
|
||||||
|
}
|
||||||
|
if artifact["status"] == "refuted":
|
||||||
|
finding["confidence"] = "low"
|
||||||
|
suppressed.append(finding)
|
||||||
|
return suppressed
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 4: Stamp `cluster_key` on constructability findings**
|
||||||
|
|
||||||
|
In `backend/agents/construct_agent.py` line 66-67, change:
|
||||||
|
|
||||||
|
```python
|
||||||
|
for finding in findings:
|
||||||
|
finding.update(agent=self.name, scope_id=scope.scope_id)
|
||||||
|
```
|
||||||
|
|
||||||
|
to:
|
||||||
|
|
||||||
|
```python
|
||||||
|
for finding in findings:
|
||||||
|
finding.update(agent=self.name, scope_id=scope.scope_id,
|
||||||
|
cluster_key=cluster.get("key"))
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 5: Wire wave 5b into `backend/agents/runner.py`**
|
||||||
|
|
||||||
|
After `memory.extend("findings", specialist_findings)` (line 167) and before
|
||||||
|
`gap_findings` / wave 6:
|
||||||
|
|
||||||
|
```python
|
||||||
|
orchestrator.stage("Agent wave 5b: evidence verification")
|
||||||
|
sheet_to_page = {
|
||||||
|
sheet.get("sheet_number"): sheet.get("page_number") for sheet in sheets
|
||||||
|
}
|
||||||
|
verify_targets = select_findings(
|
||||||
|
specialist_findings, clusters,
|
||||||
|
max_checks=config.AGENT_VERIFY_MAX_CHECKS,
|
||||||
|
severities=config.AGENT_VERIFY_SEVERITIES,
|
||||||
|
)
|
||||||
|
target_indexes = {id(f): i for i, f in enumerate(specialist_findings)}
|
||||||
|
verify_scopes = [
|
||||||
|
AgentScope(
|
||||||
|
scope_id=f"verify:{target_indexes[id(finding)]}",
|
||||||
|
payload={
|
||||||
|
"finding_index": target_indexes[id(finding)],
|
||||||
|
"finding": finding,
|
||||||
|
"images_b64": [
|
||||||
|
page_to_b64[sheet_to_page[name]]
|
||||||
|
for name in (finding.get("sheets") or [])
|
||||||
|
[:config.AGENT_CONFLICT_MAX_IMAGES]
|
||||||
|
if sheet_to_page.get(name) in page_to_b64
|
||||||
|
],
|
||||||
|
},
|
||||||
|
)
|
||||||
|
for finding in verify_targets
|
||||||
|
]
|
||||||
|
verify_results = orchestrator.run_scopes(
|
||||||
|
EvidenceVerifierAgent(usage), verify_scopes,
|
||||||
|
config.AGENT_VERIFY_CONCURRENCY,
|
||||||
|
)
|
||||||
|
suppressed = apply_verdicts(specialist_findings, verify_results)
|
||||||
|
if suppressed:
|
||||||
|
suppressed_ids = {id(f) for f in suppressed}
|
||||||
|
specialist_findings = [
|
||||||
|
f for f in specialist_findings if id(f) not in suppressed_ids
|
||||||
|
]
|
||||||
|
memory.replace("suppressed", suppressed)
|
||||||
|
memory.extend("findings", specialist_findings) # see note below
|
||||||
|
```
|
||||||
|
|
||||||
|
NOTE for implementer: `memory.extend("findings", ...)` already ran with the
|
||||||
|
un-suppressed list. Adjust ordering so verification happens BEFORE
|
||||||
|
`memory.extend("findings", specialist_findings)` — i.e. move the extend to after
|
||||||
|
wave 5b — so the Brain never sees refuted findings. Keep `gap_findings` logic
|
||||||
|
unchanged. Also add imports at top of runner.py:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from backend.agents.verifier import (
|
||||||
|
EvidenceVerifierAgent, apply_verdicts, select_findings,
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
And in the report dicts (both the `require_review` branch ~line 217-225 and the
|
||||||
|
wave-7 branch ~line 293-300), populate suppressed issues:
|
||||||
|
|
||||||
|
```python
|
||||||
|
"suppressed_issues": memory.snapshot().get("suppressed") or [],
|
||||||
|
```
|
||||||
|
|
||||||
|
Finally, `verify` results cost shows up as `agent.verify` in
|
||||||
|
`summary.cost_by_stage` automatically via `usage_stage="agent.verify"`.
|
||||||
|
|
||||||
|
**Step 6: Update existing runner tests**
|
||||||
|
|
||||||
|
`tests/agents/test_runner_review_gate.py` monkeypatches `BrainAgent` and
|
||||||
|
`convert_pdf_to_images` but lets waves 1-5 run against... check how LLM calls
|
||||||
|
are stubbed there (likely `call_json` returns None -> empty artifacts, which is
|
||||||
|
fine). The new wave must no-op cleanly when `select_findings` returns `[]`
|
||||||
|
(zero scopes -> `run_scopes` returns `[]` per orchestrator.py:60). Verify by
|
||||||
|
running the suite; if a runner test now fails because verification selects a
|
||||||
|
stubbed finding, monkeypatch `select_findings` to `lambda *a, **k: []` in that
|
||||||
|
test file's `_patch_brain` helper.
|
||||||
|
|
||||||
|
**Step 7: Run full suite**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/ -q`
|
||||||
|
Expected: all pass
|
||||||
|
|
||||||
|
**Step 8: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add backend/agents/verifier.py backend/agents/runner.py backend/agents/construct_agent.py tests/agents/test_verifier.py tests/agents/test_runner_review_gate.py
|
||||||
|
git commit -m "feat: wave 5b evidence verification - vision fact-check before Brain merge"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Task 7 (small, related): reasoning budget for the conflict critic
|
||||||
|
|
||||||
|
**Objective:** Fix the wave-4 truncation found in this same job (13/121 calls hit `finish_reason=length` at the 4096 cap with ~3.7k thinking tokens).
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `backend/agents/conflict_critic.py:55-63`
|
||||||
|
- Modify: `backend/pipeline/conflict_checker.py:85` (same pattern, classic path)
|
||||||
|
|
||||||
|
**Step 1: Apply the extractor's reasoning-knob pattern**
|
||||||
|
|
||||||
|
```python
|
||||||
|
parsed = call_json(
|
||||||
|
system_prompt=CONFLICT_SYSTEM_PROMPT,
|
||||||
|
user_text=instruction,
|
||||||
|
images_b64=images,
|
||||||
|
max_tokens=config.REASON_MAX_TOKENS,
|
||||||
|
model=config.AGENT_CONFLICT_MODEL,
|
||||||
|
reasoning_effort=config.EXTRACT_REASONING_EFFORT or None,
|
||||||
|
reasoning_max_tokens=config.EXTRACT_REASONING_MAX_TOKENS or None,
|
||||||
|
usage_tracker=self.usage,
|
||||||
|
usage_stage="agent.conflict",
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
(Match the exact kwarg names `SheetExtractorAgent` uses — check
|
||||||
|
`backend/agents/extractors.py` for whether it passes `None` when the budget is
|
||||||
|
0, and mirror that guard.)
|
||||||
|
|
||||||
|
**Step 2: Run tests**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/ -q`
|
||||||
|
Expected: all pass
|
||||||
|
|
||||||
|
**Step 3: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add backend/agents/conflict_critic.py backend/pipeline/conflict_checker.py
|
||||||
|
git commit -m "fix: reasoning budget for conflict critic (wave-4 max_tokens truncation)"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Tests / validation
|
||||||
|
|
||||||
|
1. `.venv/bin/python -m pytest tests/ -q` — full suite green.
|
||||||
|
2. **Targeted repro of the original failure:** pull the S401 page image from job
|
||||||
|
959e16407573's output dir on sits-docker (or re-render the PDF page), then run
|
||||||
|
one `EvidenceVerifierAgent` scope locally against the finding JSON from
|
||||||
|
`validated_issues[3]`. Expected: verdict `corrected`,
|
||||||
|
`actual_text: "(5) 2x6 STUD PACK"`, status `refuted`.
|
||||||
|
3. **End-to-end:** rerun the same Cypress TX PDF through the pipeline (local with
|
||||||
|
`LLM_CACHE`/`LLM_RAW_DUMP` per the conflict-checker skill). Expected:
|
||||||
|
- log shows `Agent wave 5b: evidence verification` with a bounded number of calls;
|
||||||
|
- the S401 stud-pack finding is either absent from `validated_issues` and present
|
||||||
|
in `suppressed_issues` with verification verdicts, or downgraded to low confidence;
|
||||||
|
- `agent.verify` appears in `summary.cost_by_stage`;
|
||||||
|
- zero `finish_reason=length` lines in wave 4 (Task 7).
|
||||||
|
4. Cost check: wave 5b adds at most `AGENT_VERIFY_MAX_CHECKS` (20) vision calls —
|
||||||
|
for this job's profile that is well under $1.
|
||||||
|
|
||||||
|
## Risks, tradeoffs, open questions
|
||||||
|
|
||||||
|
- **Dispute false positives:** cluster members with legitimately different values
|
||||||
|
(e.g. two doors in one door cluster) will produce `disputed_attributes`. Mitigation:
|
||||||
|
prompts treat disputes as "unverified", not "wrong"; only severity-gated findings
|
||||||
|
burn verification calls. Tune later by restricting `find_disputes` to numeric-ish
|
||||||
|
values if noise is high.
|
||||||
|
- **Verifier can also misread.** It is one model checking another with the same eyes.
|
||||||
|
Mitigation: verdict requires `actual_text` verbatim evidence for `corrected`, and
|
||||||
|
only fully-unconfirmed findings are suppressed (mixed keeps the finding with a note).
|
||||||
|
- **xref cost:** extra link scopes. Bounded by the >= 2 distinct sheets gate and the
|
||||||
|
existing assertion cap; expect a handful of extra scopes per set.
|
||||||
|
- **Suppression in review mode:** refuted findings land in `suppressed_issues` — the
|
||||||
|
review UI/finalizer must tolerate that list being non-empty (it is currently always
|
||||||
|
`[]` in agent mode). Open question: surface suppressed items in the human review
|
||||||
|
queue as informational, or keep them report-only?
|
||||||
|
- **Open question:** should wave-4 conflict findings (which already saw images) also be
|
||||||
|
verification-eligible? Plan says no (they had the pixels); revisit if critics show
|
||||||
|
the same misread pattern.
|
||||||
@@ -0,0 +1,715 @@
|
|||||||
|
# Extraction Coverage Guarantee — Implementation Plan
|
||||||
|
|
||||||
|
> **For Hermes:** Use subagent-driven-development skill to implement this plan task-by-task.
|
||||||
|
|
||||||
|
**Goal:** Eliminate dark sheets (missed pages) and vision-misread content by making wave-1 extraction coverage-guaranteed: deterministic coverage measurement, a text-first retry ladder, deterministic fallback extraction, and text-layer sheet identity recovery.
|
||||||
|
|
||||||
|
**Architecture:** For text-bearing sheets the authoritative alphanumeric content already exists in the PyMuPDF text layer (backend/text_layer.py). Today the LLM transcribes from pixels and we merely *detect* failure post-hoc (coverage_gaps logs; nothing retries). This plan flips wave 1 to: run the vision pass (unchanged, always, on every page) → measure text coverage deterministically per page → if below floor, ADD a text-only structuring pass (LLM segments the text layer, no image, no misreads possible) and MERGE its objects into the vision results — vision keeps everything it found, text structuring fills what it missed → if still below floor, emit deterministic stub objects straight from the text layer so NO text-bearing page ever contributes zero objects. Sheet identity is recovered from the text layer when the LLM drops the header. Vision stays the only source for graphical content (symbols, geometry, line work) and the only path for scanned pages.
|
||||||
|
|
||||||
|
**Tech Stack:** Python 3.14, PyMuPDF (already a dep), existing call_json LLM plumbing, pytest.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Root-Cause Diagnosis (why this keeps happening)
|
||||||
|
|
||||||
|
Confirmed against Cypress job 3e01d5baba32 (38-page Verizon set) and code:
|
||||||
|
|
||||||
|
**Missed sheets (pages 8 = S202 wood notes, 10 = S204 lap-splice tables, 18 = A102 REFLECTED CEILING PLAN — zero assertions each):**
|
||||||
|
1. The extractor prompt (backend/prompts.py:277) is biased toward physical "construction objects" (rooms, doors, fixtures). Notes/table-dense sheets have few, so the model returns a bare array with ONE generic summary object (log: `wrapping bare objects array (1 items, no sheet header)`).
|
||||||
|
2. The grounding guard (backend/pipeline/extractor.py:178 `_is_grounded`) drops that summary object as ungrounded → 0 objects.
|
||||||
|
3. `_wrap_bare_list` (backend/agents/extractors.py:33) converts the 1-item bare array into a valid dict, so the compact retry (extractors.py:71) NEVER fires — it only triggers when parsing fully fails. A 1-object page counts as "success".
|
||||||
|
4. `coverage_gaps()` (backend/text_layer.py:211) only LOGS the gap and adds a failed_scope note. No retry, no fallback. The page is silently dark for every downstream wave.
|
||||||
|
5. Sheet identity comes ONLY from the LLM reading the title block in the image. 7/38 Cypress pages ended with `sheet_number=None` (4 of them WITH assertions: pages 22, 30, 31, 37), so they can't join sheet-keyed scopes and corrupt `missing_expected_sheets` downstream.
|
||||||
|
|
||||||
|
**Completely incorrect information:**
|
||||||
|
1. Vision misreads of dense alphanumeric content (the "(2) vs (5) 2x6 STUD PACK" family). The wave-1.5 rescue tier catches invented numbers but is a SET subset test — it cannot catch SWAPPED numbers (documented in docs/superpowers/specs/2026-08-12-text-layer-grounding-design.md).
|
||||||
|
2. Gemini thinking tokens count against max_tokens → `recovered truncated JSON` silently drops tail objects (bottom/right of sheet vanishes). Nothing flags the page as degraded.
|
||||||
|
3. The model paraphrases `source_text`; the guard only checks digit-run/token overlap, so plausible-but-wrong values pass.
|
||||||
|
4. `JSON parse error (giving up)` → classic path returns an empty "extraction failed" sheet (extractor.py:282-291); the page vanishes from analysis while `sheets_analyzed` still counts it.
|
||||||
|
|
||||||
|
**Cornerstone principles for the fix:**
|
||||||
|
1. If a page has a text layer, the truth is already deterministic and free. The LLM's job on such pages is STRUCTURING, not TRANSCRIPTION. Every extracted alphanumeric claim must trace to the text layer; anything that can't is vision-only and gets stamped as such.
|
||||||
|
2. The vision pass is never skipped and never replaced. These are construction documents: symbols, device/fixture locations, geometry, and line work exist only in the image. The text-only rung and the fallback rung are strictly ADDITIVE — they merge into the vision results (deduped by normalized source_text), so a rescue can only add coverage, never subtract graphical content.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 1: Coverage metric module (backend/text_coverage.py)
|
||||||
|
|
||||||
|
**Objective:** Deterministic per-page coverage measurement: what fraction of the text layer is actually represented in extracted objects.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Create: `backend/text_coverage.py`
|
||||||
|
- Test: `tests/test_text_coverage.py`
|
||||||
|
|
||||||
|
**Step 1: Write failing test**
|
||||||
|
|
||||||
|
```python
|
||||||
|
# tests/test_text_coverage.py
|
||||||
|
from backend.text_coverage import text_coverage, segment_text_layer, fallback_objects
|
||||||
|
|
||||||
|
def test_coverage_full():
|
||||||
|
text = "NOTE 1\nALL LUMBER NO. 2 SOUTHERN PINE\nNOTE 2\nUSE 5/8\" PLYWOOD"
|
||||||
|
objects = [{"source_text": "ALL LUMBER NO. 2 SOUTHERN PINE"},
|
||||||
|
{"source_text": "USE 5/8\" PLYWOOD"}]
|
||||||
|
cov = text_coverage(text, objects)
|
||||||
|
assert cov["covered_lines"] == 2
|
||||||
|
assert cov["total_lines"] == 2
|
||||||
|
assert cov["ratio"] == 1.0
|
||||||
|
|
||||||
|
def test_coverage_zero_on_empty_objects():
|
||||||
|
cov = text_coverage("LINE A\nLINE B\nLINE C", [])
|
||||||
|
assert cov["ratio"] == 0.0 and cov["total_lines"] == 3
|
||||||
|
|
||||||
|
def test_coverage_ignores_short_and_numeric_noise_lines():
|
||||||
|
text = "15\"\n19\"\nA\nB\nREAL NOTE ABOUT FRAMING HERE"
|
||||||
|
cov = text_coverage(text, [{"source_text": "REAL NOTE ABOUT FRAMING HERE"}])
|
||||||
|
# short/noise lines (< MIN_LINE_CHARS or pure dimension ticks) excluded
|
||||||
|
assert cov["total_lines"] == 1 and cov["ratio"] == 1.0
|
||||||
|
|
||||||
|
def test_segment_notes_and_rows():
|
||||||
|
text = "WOOD CONSTRUCTION\n1. \nALL SAWN LUMBER TO BE SOUTHERN PINE.\n2. \nROOF SHEATHING 5/8\" PLYWOOD."
|
||||||
|
segs = segment_text_layer(text)
|
||||||
|
assert any("ALL SAWN LUMBER" in s for s in segs)
|
||||||
|
assert any("ROOF SHEATHING" in s for s in segs)
|
||||||
|
|
||||||
|
def test_fallback_objects_verbatim_and_stamped():
|
||||||
|
objs = fallback_objects("1. \nALL SAWN LUMBER TO BE SOUTHERN PINE.", page_number=8)
|
||||||
|
assert len(objs) == 1
|
||||||
|
assert objs[0]["source_text"] == "ALL SAWN LUMBER TO BE SOUTHERN PINE."
|
||||||
|
assert objs[0]["grounding"] == "text_layer_fallback"
|
||||||
|
assert objs[0]["confidence"] == "low"
|
||||||
|
|
||||||
|
def test_merge_objects_keeps_vision_and_unions_text():
|
||||||
|
vision = [
|
||||||
|
{"source_text": "2X6 WD STUD @ 16\" O.C.", "object_type": "wall"},
|
||||||
|
{"source_text": None, "graphical_basis": "light fixture symbol, grid C-4",
|
||||||
|
"object_type": "lighting_fixture"}, # graphical: exists only in image
|
||||||
|
]
|
||||||
|
text = [
|
||||||
|
{"source_text": "2X6 WD STUD @ 16\" O.C.", "object_type": "wall"}, # dup
|
||||||
|
{"source_text": "ALL LUMBER NO. 2 SOUTHERN PINE", "object_type": "general_note"},
|
||||||
|
]
|
||||||
|
merged = merge_objects(vision, text)
|
||||||
|
assert len(merged) == 3 # dup dropped, note added
|
||||||
|
assert any(o.get("graphical_basis") for o in merged) # graphical kept
|
||||||
|
assert merged[0]["object_type"] == "wall" # vision order preserved
|
||||||
|
|
||||||
|
def test_merge_objects_dedupes_by_normalized_text():
|
||||||
|
a = [{"source_text": "RTU-1: 5 TON, 1600 CFM"}]
|
||||||
|
b = [{"source_text": "rtu 1 5 ton 1600 cfm"}] # same content, different case/punct
|
||||||
|
assert len(merge_objects(a, b)) == 1
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 2: Run test to verify failure**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/test_text_coverage.py -v`
|
||||||
|
Expected: FAIL — ModuleNotFoundError: backend.text_coverage
|
||||||
|
|
||||||
|
**Step 3: Implement `backend/text_coverage.py`**
|
||||||
|
|
||||||
|
```python
|
||||||
|
"""text_coverage.py - deterministic extraction-coverage measurement.
|
||||||
|
|
||||||
|
The coverage guarantee: for any page with a usable text layer, measure how
|
||||||
|
much of that layer ended up represented in extracted objects. Pages below
|
||||||
|
the floor route into the extraction retry ladder (agents/extractors.py and
|
||||||
|
pipeline/extractor.py). fallback_objects() is the last rung: stub objects
|
||||||
|
segmented straight from the text layer so no text-bearing page goes dark.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import re
|
||||||
|
from typing import Dict, List
|
||||||
|
|
||||||
|
# Lines below this many meaningful chars are noise (dimension ticks, grid
|
||||||
|
# bubbles, single letters) and excluded from the coverage denominator.
|
||||||
|
MIN_LINE_CHARS = 12
|
||||||
|
# Pure dimension/elevation ticks like 15" or 8' - 0" carry no prose content.
|
||||||
|
_TICK_RE = re.compile(r"^[\d\s'\"/.,-]+$")
|
||||||
|
_WORD_RE = re.compile(r"[a-z0-9]+")
|
||||||
|
|
||||||
|
|
||||||
|
def _meaningful_lines(text: str) -> List[str]:
|
||||||
|
lines = []
|
||||||
|
for raw in (text or "").splitlines():
|
||||||
|
line = " ".join(raw.split())
|
||||||
|
if len(line) < MIN_LINE_CHARS or _TICK_RE.match(line):
|
||||||
|
continue
|
||||||
|
lines.append(line)
|
||||||
|
return lines
|
||||||
|
|
||||||
|
|
||||||
|
def _norm(text: str) -> str:
|
||||||
|
return " ".join(_WORD_RE.findall((text or "").lower()))
|
||||||
|
|
||||||
|
|
||||||
|
def text_coverage(page_text: str, objects: List[Dict]) -> Dict:
|
||||||
|
"""Fraction of meaningful text-layer lines whose normalized form appears
|
||||||
|
in the concatenated normalized source_text of extracted objects."""
|
||||||
|
lines = _meaningful_lines(page_text)
|
||||||
|
if not lines:
|
||||||
|
return {"total_lines": 0, "covered_lines": 0, "ratio": 1.0}
|
||||||
|
haystack = " ".join(
|
||||||
|
_norm(str(o.get("source_text") or o.get("object_description")
|
||||||
|
or o.get("value") or ""))
|
||||||
|
for o in objects if isinstance(o, dict)
|
||||||
|
)
|
||||||
|
covered = sum(1 for ln in lines if _norm(ln) and _norm(ln) in haystack)
|
||||||
|
return {
|
||||||
|
"total_lines": len(lines),
|
||||||
|
"covered_lines": covered,
|
||||||
|
"ratio": covered / len(lines) if lines else 1.0,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def segment_text_layer(text: str) -> List[str]:
|
||||||
|
"""Segment a page text layer into note-sized blocks: numbered notes and
|
||||||
|
contiguous prose runs. PyMuPDF emits each note number on its own line
|
||||||
|
('1. ', '2. ') followed by wrapped text lines; rejoin number->body and
|
||||||
|
merge continuation lines until the next number or blank-line break."""
|
||||||
|
segments: List[str] = []
|
||||||
|
buf: List[str] = []
|
||||||
|
number_re = re.compile(r"^(\d{1,2}[.)]?|[A-Z]\d{0,2}[.)]?)\s*$")
|
||||||
|
|
||||||
|
def flush():
|
||||||
|
joined = " ".join(buf).strip()
|
||||||
|
if len(joined) >= MIN_LINE_CHARS:
|
||||||
|
segments.append(joined)
|
||||||
|
buf.clear()
|
||||||
|
|
||||||
|
for raw in (text or "").splitlines():
|
||||||
|
line = raw.strip()
|
||||||
|
if not line:
|
||||||
|
flush()
|
||||||
|
continue
|
||||||
|
if number_re.match(line):
|
||||||
|
flush()
|
||||||
|
buf.append(line.rstrip(".)"))
|
||||||
|
continue
|
||||||
|
buf.append(line)
|
||||||
|
# wrapped-note heuristic: a line starting a new sentence after a
|
||||||
|
# period ends the segment
|
||||||
|
if line.endswith(".") and len(" ".join(buf)) > 120:
|
||||||
|
flush()
|
||||||
|
flush()
|
||||||
|
return segments
|
||||||
|
|
||||||
|
|
||||||
|
def fallback_objects(page_text: str, page_number: int,
|
||||||
|
max_objects: int = 200) -> List[Dict]:
|
||||||
|
"""Last-rung deterministic extraction: one stub object per text segment,
|
||||||
|
source_text verbatim from the text layer. confidence=low and
|
||||||
|
grounding=text_layer_fallback make their provenance explicit downstream."""
|
||||||
|
objs = []
|
||||||
|
for idx, seg in enumerate(segment_text_layer(page_text)[:max_objects]):
|
||||||
|
objs.append({
|
||||||
|
"object_id": f"p{page_number}-tl{idx}",
|
||||||
|
"object_type": "general_note",
|
||||||
|
"category": "general",
|
||||||
|
"tag": None,
|
||||||
|
"name": seg[:80],
|
||||||
|
"description": seg,
|
||||||
|
"attributes": {},
|
||||||
|
"location_key": {},
|
||||||
|
"source_text": seg,
|
||||||
|
"graphical_basis": None,
|
||||||
|
"review_uses": ["code_review", "constructability_review"],
|
||||||
|
"confidence": "low",
|
||||||
|
"grounding": "text_layer_fallback",
|
||||||
|
})
|
||||||
|
return objs
|
||||||
|
|
||||||
|
|
||||||
|
def merge_objects(vision_objs: List[Dict], text_objs: List[Dict]) -> List[Dict]:
|
||||||
|
"""Union of vision and text-structured objects. Vision results come first
|
||||||
|
and are never dropped (graphical_basis objects exist only in the image).
|
||||||
|
Text objects are appended unless their normalized source_text is already
|
||||||
|
represented. A merge can only add coverage, never subtract it."""
|
||||||
|
merged = list(vision_objs or [])
|
||||||
|
seen = {_norm(str(o.get("source_text") or ""))
|
||||||
|
for o in merged if isinstance(o, dict)}
|
||||||
|
seen.discard("")
|
||||||
|
for obj in text_objs or []:
|
||||||
|
if not isinstance(obj, dict):
|
||||||
|
continue
|
||||||
|
key = _norm(str(obj.get("source_text") or ""))
|
||||||
|
if key and key in seen:
|
||||||
|
continue
|
||||||
|
seen.add(key)
|
||||||
|
merged.append(obj)
|
||||||
|
return merged
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 4: Run test to verify pass**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/test_text_coverage.py -v`
|
||||||
|
Expected: 7 passed
|
||||||
|
|
||||||
|
**Step 5: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add backend/text_coverage.py tests/test_text_coverage.py
|
||||||
|
git commit -m "feat: deterministic text-layer coverage metric + fallback extraction"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 2: Sheet identity recovery from the text layer
|
||||||
|
|
||||||
|
**Objective:** When the LLM drops/misreads the sheet header, recover `sheet_number` (and discipline via existing `discipline_from_sheet_number`) deterministically from the text layer instead of leaving None.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `backend/text_coverage.py` (add `recover_sheet_number`)
|
||||||
|
- Test: `tests/test_text_coverage.py` (add tests)
|
||||||
|
|
||||||
|
**Step 1: Write failing test**
|
||||||
|
|
||||||
|
```python
|
||||||
|
def test_recover_sheet_number_from_title_block():
|
||||||
|
text = ("WALL SECTIONS\n...\nSheet Information\nS301\n"
|
||||||
|
"Issue Date 05.29.26\nProject Number 25177")
|
||||||
|
assert recover_sheet_number(text) == "S301"
|
||||||
|
|
||||||
|
def test_recover_sheet_number_none_when_absent():
|
||||||
|
assert recover_sheet_number("just some notes about lumber") is None
|
||||||
|
|
||||||
|
def test_recover_prefers_discipline_pattern_over_dates():
|
||||||
|
# 05.29.26 and 25177 must never match
|
||||||
|
text = "Issue Date 05.29.26\nProject Number 25177\nA102 REFLECTED CEILING PLAN"
|
||||||
|
assert recover_sheet_number(text) == "A102"
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 2: Run to verify failure**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/test_text_coverage.py::test_recover_sheet_number_from_title_block -v`
|
||||||
|
Expected: FAIL — ImportError
|
||||||
|
|
||||||
|
**Step 3: Implement in `backend/text_coverage.py`**
|
||||||
|
|
||||||
|
```python
|
||||||
|
# Sheet ids: 1-2 uppercase letters + 2-3 digits + optional decimal suffix
|
||||||
|
# (S301, A102, M200, E500, LS101, P100, G000). Deliberately excludes pure
|
||||||
|
# numbers (dates, project numbers) and long alphanumerics (member marks).
|
||||||
|
_SHEET_ID_RE = re.compile(r"\b([A-Z]{1,2}\d{2,3}(?:\.\d+)?)\b")
|
||||||
|
_TITLE_HINT_RE = re.compile(
|
||||||
|
r"(?i)sheet\s*(?:information|no|number)?|"
|
||||||
|
r"(floor plan|ceiling plan|elevations?|sections?|details?|schedule|"
|
||||||
|
r"notes|legend|plan)")
|
||||||
|
|
||||||
|
def recover_sheet_number(page_text: str) -> Optional[str]:
|
||||||
|
"""Deterministic sheet id from the text layer. Strategy: collect every
|
||||||
|
sheet-id-shaped token, prefer ones appearing near title words or in the
|
||||||
|
last ~15%% of the page (title block lives at the drawing edge)."""
|
||||||
|
text = page_text or ""
|
||||||
|
cands = _SHEET_ID_RE.findall(text)
|
||||||
|
if not cands:
|
||||||
|
return None
|
||||||
|
tail = text[int(len(text) * 0.85):]
|
||||||
|
for cand in reversed(_SHEET_ID_RE.findall(tail)):
|
||||||
|
return cand
|
||||||
|
return cands[0]
|
||||||
|
```
|
||||||
|
|
||||||
|
(Add `from typing import Optional` to the imports.)
|
||||||
|
|
||||||
|
**Step 4: Run to verify pass**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/test_text_coverage.py -v`
|
||||||
|
Expected: all pass (10 tests)
|
||||||
|
|
||||||
|
**Step 5: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add backend/text_coverage.py tests/test_text_coverage.py
|
||||||
|
git commit -m "feat: deterministic sheet-number recovery from text layer"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 3: Text-only structuring prompt (no image)
|
||||||
|
|
||||||
|
**Objective:** Second rung of the ladder: give the LLM the raw text layer and ask it to segment EVERY note/row/callout into objects with verbatim source_text. No image = no vision misreads for alphanumerics; far cheaper than the vision pass.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `backend/prompts.py` (append after EXTRACTOR_USER_INSTRUCTION, ~line 282)
|
||||||
|
- Test: `tests/agents/test_extraction_ladder.py` (prompt-content assertions only; rendering tested in Task 4)
|
||||||
|
|
||||||
|
**Step 1: Write failing test**
|
||||||
|
|
||||||
|
```python
|
||||||
|
# tests/agents/test_extraction_ladder.py
|
||||||
|
from backend.prompts import TEXT_STRUCTURING_SYSTEM_PROMPT, TEXT_STRUCTURING_USER_INSTRUCTION
|
||||||
|
|
||||||
|
def test_text_structuring_prompt_demands_verbatim_and_completeness():
|
||||||
|
assert "verbatim" in TEXT_STRUCTURING_USER_INSTRUCTION.lower()
|
||||||
|
assert "every" in TEXT_STRUCTURING_USER_INSTRUCTION.lower()
|
||||||
|
assert "{text_layer}" in TEXT_STRUCTURING_USER_INSTRUCTION
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 2: Run to verify failure**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/agents/test_extraction_ladder.py -v`
|
||||||
|
Expected: FAIL — ImportError
|
||||||
|
|
||||||
|
**Step 3: Append to `backend/prompts.py`**
|
||||||
|
|
||||||
|
```python
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Stage 2b - text-only structuring (extraction retry ladder, rung 2)
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
TEXT_STRUCTURING_SYSTEM_PROMPT = """You are a construction document structuring engine.
|
||||||
|
You receive the deterministic text layer extracted from one drawing sheet. It is complete and authoritative.
|
||||||
|
Your ONLY job is to segment it into structured objects. You are NOT reading an image. You must NOT invent, complete, or correct any text.
|
||||||
|
Rules:
|
||||||
|
- Every numbered note, schedule row, callout, tag, legend entry, and title-block field becomes its own object.
|
||||||
|
- source_text must be copied VERBATIM from the input, character-for-character. Never paraphrase.
|
||||||
|
- Cover the ENTIRE input. Omitting a note is a failure. When unsure of an object's type, use general_note with confidence low.
|
||||||
|
- Numbers, model numbers, dimensions, and tags must appear in source_text exactly as in the input.
|
||||||
|
Respond only with valid JSON."""
|
||||||
|
|
||||||
|
TEXT_STRUCTURING_USER_INSTRUCTION = """Segment this sheet's text layer into structured construction objects.
|
||||||
|
Respond ONLY with a valid JSON object - no markdown fences:
|
||||||
|
{ "sheet": { "sheet_number": "string or null", "sheet_title": "string or null", "discipline": "string or null", "drawing_type": "string or null", "level": "string or null", "scale": "string or null" }, "objects": [ { "object_id": "string", "object_type": "room | door | window | wall | finish | ceiling | dimension | grid | callout | keynote | general_note | equipment | plumbing_fixture | mechanical_equipment | electrical_device | lighting_fixture | structural_element | schedule_reference | symbol | abbreviation", "category": "architectural | structural | mechanical | electrical | plumbing | code | general", "tag": "string or null", "name": "string or null", "description": "string or null", "attributes": { "attribute_name": "attribute_value" }, "location_key": { "room_number": "string or null", "grid": "string or null", "detail_reference": "string or null" }, "source_text": "VERBATIM text copied from the input", "graphical_basis": null, "review_uses": [ "schedule_comparison", "cross_discipline_coordination", "code_review", "constructability_review" ], "confidence": "high | medium | low" } ], "unresolved_items": [] }
|
||||||
|
Optional sheet hint: {sheet_hint}
|
||||||
|
|
||||||
|
TEXT LAYER (segment ALL of it):
|
||||||
|
{text_layer}"""
|
||||||
|
```
|
||||||
|
|
||||||
|
NOTE the two render sites you will add in Tasks 4-5 substitute `{sheet_hint}` and `{text_layer}` with str.replace directly (NOT via render()/call_stage) — this matches the wave-1.5 pattern and avoids the classic-path literal-placeholder leak documented in the project pitfalls.
|
||||||
|
|
||||||
|
**Step 4: Run to verify pass**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/agents/test_extraction_ladder.py -v`
|
||||||
|
Expected: 1 passed
|
||||||
|
|
||||||
|
**Step 5: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add backend/prompts.py tests/agents/test_extraction_ladder.py
|
||||||
|
git commit -m "feat: text-only structuring prompt for extraction retry ladder"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 4: Retry ladder in the AGENT path (SheetExtractorAgent)
|
||||||
|
|
||||||
|
**Objective:** Replace the binary parse-fail retry with a coverage-driven ladder: vision pass → coverage check → text-only structuring pass → deterministic fallback. Also recover sheet identity and mark truncation-degraded pages.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `backend/agents/extractors.py:63-86` (SheetExtractorAgent.run)
|
||||||
|
- Modify: `backend/config.py` (new knobs, below)
|
||||||
|
- Test: `tests/agents/test_extraction_ladder.py`
|
||||||
|
|
||||||
|
**New config knobs (backend/config.py, follow existing env pattern):**
|
||||||
|
|
||||||
|
```python
|
||||||
|
EXTRACT_COVERAGE_FLOOR = float(os.getenv("EXTRACT_COVERAGE_FLOOR", "0.6"))
|
||||||
|
EXTRACT_TEXT_RETRY_ENABLED = os.getenv("EXTRACT_TEXT_RETRY_ENABLED", "true").lower() == "true"
|
||||||
|
EXTRACT_FALLBACK_ENABLED = os.getenv("EXTRACT_FALLBACK_ENABLED", "true").lower() == "true"
|
||||||
|
EXTRACT_FALLBACK_MAX_OBJECTS = int(os.getenv("EXTRACT_FALLBACK_MAX_OBJECTS", "200"))
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 1: Write failing test**
|
||||||
|
|
||||||
|
```python
|
||||||
|
from backend.agents.extractors import SheetExtractorAgent
|
||||||
|
from backend.agents.base import AgentScope, AgentUsage
|
||||||
|
|
||||||
|
def _page(n=8, text="1. \nALL SAWN LUMBER IN CONTACT WITH SOIL TO BE SOUTHERN PINE, PRESSURE TREATED.\n2. \nROOF SHEATHING: 5/8\" PLYWOOD, C-D GRADE, STRUCTURAL I."):
|
||||||
|
return {"page_number": n, "base64": "AAAA", "text_layer": text}
|
||||||
|
|
||||||
|
def test_ladder_falls_back_when_vision_returns_nothing(agent_monkeypatch):
|
||||||
|
# vision pass returns 1 summary object that the guard drops;
|
||||||
|
# text-structuring disabled to exercise the deterministic rung
|
||||||
|
agent_monkeypatch.setattr("backend.agents.extractors.call_json",
|
||||||
|
lambda **kw: [{"name": "general notes", "value": "notes"}])
|
||||||
|
agent_monkeypatch.setattr("backend.config.EXTRACT_TEXT_RETRY_ENABLED", False)
|
||||||
|
agent = SheetExtractorAgent(AgentUsage())
|
||||||
|
scope = AgentScope(scope_id="sheet:8", payload={"page": _page(), "sheet_hint": ""})
|
||||||
|
result = agent.run(scope)
|
||||||
|
sheet = result.artifacts[0]
|
||||||
|
assert sheet["assertions"], "dark sheet must be impossible with fallback enabled"
|
||||||
|
assert all(a.get("grounding") == "text_layer_fallback" for a in sheet["assertions"])
|
||||||
|
assert sheet["coverage"]["ratio"] >= 0.6
|
||||||
|
|
||||||
|
def test_ladder_merge_preserves_graphical_objects(agent_monkeypatch):
|
||||||
|
# vision finds a graphical symbol + misreads nothing; text rung adds notes.
|
||||||
|
# The graphical object MUST survive the merge.
|
||||||
|
calls = {"n": 0}
|
||||||
|
def fake_call_json(**kw):
|
||||||
|
calls["n"] += 1
|
||||||
|
if kw.get("images_b64"): # vision pass
|
||||||
|
return {"sheet": {}, "objects": [
|
||||||
|
{"object_id": "g1", "object_type": "lighting_fixture",
|
||||||
|
"name": "pendant at grid C-4", "source_text": None,
|
||||||
|
"graphical_basis": "16in pendant symbol at grid C-4"}]}
|
||||||
|
return {"sheet": {}, "objects": [ # text-structuring pass
|
||||||
|
{"object_id": "t1", "object_type": "general_note",
|
||||||
|
"source_text": "ALL SAWN LUMBER IN CONTACT WITH SOIL TO BE SOUTHERN PINE, PRESSURE TREATED.",
|
||||||
|
"name": "lumber note"}]}
|
||||||
|
agent_monkeypatch.setattr("backend.agents.extractors.call_json", fake_call_json)
|
||||||
|
agent = SheetExtractorAgent(AgentUsage())
|
||||||
|
scope = AgentScope(scope_id="sheet:8", payload={"page": _page(), "sheet_hint": ""})
|
||||||
|
sheet = agent.run(scope).artifacts[0]
|
||||||
|
assert any(a.get("graphical_basis") for a in sheet["assertions"])
|
||||||
|
assert any("SAWN LUMBER" in (a.get("source_text") or "") for a in sheet["assertions"])
|
||||||
|
|
||||||
|
def test_ladder_recovers_sheet_number_from_text_layer(agent_monkeypatch):
|
||||||
|
agent_monkeypatch.setattr(
|
||||||
|
"backend.agents.extractors.call_json",
|
||||||
|
lambda **kw: {"sheet": {}, "objects": [
|
||||||
|
{"object_id": "o1", "name": "RCP note",
|
||||||
|
"source_text": "GYP. BD. CEILING 8'-11 3/8\" A.F.F. TYP. FOR ALL STOREFRONT",
|
||||||
|
"attributes": {"height": "8'-11 3/8\""}}]})
|
||||||
|
agent = SheetExtractorAgent(AgentUsage())
|
||||||
|
scope = AgentScope(scope_id="sheet:18",
|
||||||
|
payload={"page": _page(18, "REFLECTED CEILING PLAN\nA102\nGYP. BD. CEILING 8'-11 3/8\" A.F.F. TYP. FOR ALL STOREFRONT"),
|
||||||
|
"sheet_hint": ""})
|
||||||
|
sheet = agent.run(scope).artifacts[0]
|
||||||
|
assert sheet["sheet_number"] == "A102"
|
||||||
|
```
|
||||||
|
|
||||||
|
(Monkeypatch fixture: plain `unittest.mock.patch` context or pytest `monkeypatch`; follow tests/agents/test_text_layer_flow.py patterns for scope/result construction — check AgentScope/AgentResult signatures in backend/agents/base.py before writing.)
|
||||||
|
|
||||||
|
**Step 2: Run to verify failure**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/agents/test_extraction_ladder.py -v`
|
||||||
|
Expected: FAIL — assertions on coverage/sheet_number fail (ladder not implemented)
|
||||||
|
|
||||||
|
**Step 3: Implement the ladder in `backend/agents/extractors.py`**
|
||||||
|
|
||||||
|
Replace `SheetExtractorAgent.run` (lines 63-86) with:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def _text_structuring_call(self, page, sheet_hint):
|
||||||
|
from backend.prompts import (TEXT_STRUCTURING_SYSTEM_PROMPT,
|
||||||
|
TEXT_STRUCTURING_USER_INSTRUCTION)
|
||||||
|
instruction = (TEXT_STRUCTURING_USER_INSTRUCTION
|
||||||
|
.replace("{sheet_hint}", str(sheet_hint or ""))
|
||||||
|
.replace("{text_layer}",
|
||||||
|
(page.get("text_layer") or "")
|
||||||
|
[:config.TEXT_LAYER_MAX_CHARS]))
|
||||||
|
return call_json(
|
||||||
|
system_prompt=TEXT_STRUCTURING_SYSTEM_PROMPT,
|
||||||
|
user_text=instruction,
|
||||||
|
images_b64=None,
|
||||||
|
max_tokens=config.EXTRACT_MAX_TOKENS,
|
||||||
|
model=config.AGENT_EXTRACT_MODEL,
|
||||||
|
usage_tracker=self.usage,
|
||||||
|
usage_stage="agent.extract_text",
|
||||||
|
reasoning_effort=config.EXTRACT_REASONING_EFFORT or None,
|
||||||
|
reasoning_max_tokens=config.EXTRACT_REASONING_MAX_TOKENS or None,
|
||||||
|
)
|
||||||
|
|
||||||
|
def run(self, scope: AgentScope) -> AgentResult:
|
||||||
|
from backend.text_coverage import (fallback_objects, merge_objects,
|
||||||
|
recover_sheet_number, text_coverage)
|
||||||
|
try:
|
||||||
|
page = scope.payload["page"]
|
||||||
|
hint = scope.payload.get("sheet_hint") or ""
|
||||||
|
page_text = page.get("text_layer")
|
||||||
|
instruction = EXTRACTOR_USER_INSTRUCTION.replace(
|
||||||
|
"{sheet_hint}", str(hint)) + _text_layer_block(page)
|
||||||
|
|
||||||
|
# Rung 1: vision pass (unchanged behaviour, incl. compact retry)
|
||||||
|
parsed = _wrap_bare_list(self._call(instruction, page),
|
||||||
|
page["page_number"])
|
||||||
|
if not isinstance(parsed, dict):
|
||||||
|
print(f"[Extract] Page {page['page_number']}: full extraction "
|
||||||
|
f"failed, retrying compact")
|
||||||
|
parsed = _wrap_bare_list(
|
||||||
|
self._call(instruction + _COMPACT_RETRY_SUFFIX, page),
|
||||||
|
page["page_number"])
|
||||||
|
if not isinstance(parsed, dict):
|
||||||
|
parsed = {"sheet": {}, "objects": []}
|
||||||
|
|
||||||
|
sheet = _normalize_sheet(parsed, page["page_number"],
|
||||||
|
page_text=page_text)
|
||||||
|
cov = text_coverage(page_text or "", sheet["assertions"])
|
||||||
|
sheet["coverage"] = cov
|
||||||
|
|
||||||
|
# Rung 2: text-only structuring when coverage is below floor.
|
||||||
|
# MERGE, never replace: vision keeps every object it found
|
||||||
|
# (graphical_basis content exists only in the image); the text
|
||||||
|
# pass fills in the text content the vision pass missed.
|
||||||
|
if (page_text and config.EXTRACT_TEXT_RETRY_ENABLED
|
||||||
|
and cov["ratio"] < config.EXTRACT_COVERAGE_FLOOR):
|
||||||
|
print(f"[Extract] Page {page['page_number']}: coverage "
|
||||||
|
f"{cov['ratio']:.0%} < floor - text-only structuring pass")
|
||||||
|
parsed2 = _wrap_bare_list(
|
||||||
|
self._text_structuring_call(page, hint), page["page_number"])
|
||||||
|
if isinstance(parsed2, dict):
|
||||||
|
sheet2 = _normalize_sheet(parsed2, page["page_number"],
|
||||||
|
page_text=page_text)
|
||||||
|
before = len(sheet["assertions"])
|
||||||
|
sheet["assertions"] = merge_objects(sheet["assertions"],
|
||||||
|
sheet2["assertions"])
|
||||||
|
# Fill header gaps the vision pass left null
|
||||||
|
for key in ("sheet_number", "sheet_title", "discipline",
|
||||||
|
"level", "scale", "drawing_type"):
|
||||||
|
if not sheet.get(key) and sheet2.get(key):
|
||||||
|
sheet[key] = sheet2[key]
|
||||||
|
cov = text_coverage(page_text, sheet["assertions"])
|
||||||
|
sheet["coverage"] = cov
|
||||||
|
print(f"[Extract] Page {page['page_number']}: merged "
|
||||||
|
f"{len(sheet['assertions']) - before} text-structured "
|
||||||
|
f"object(s), coverage now {cov['ratio']:.0%}")
|
||||||
|
|
||||||
|
# Rung 3: deterministic fallback - dark sheets are impossible.
|
||||||
|
# Also merged (deduped) so stub notes never double up with
|
||||||
|
# objects the earlier rungs already captured.
|
||||||
|
if (page_text and config.EXTRACT_FALLBACK_ENABLED
|
||||||
|
and cov["ratio"] < config.EXTRACT_COVERAGE_FLOOR):
|
||||||
|
stubs = fallback_objects(page_text, page["page_number"],
|
||||||
|
config.EXTRACT_FALLBACK_MAX_OBJECTS)
|
||||||
|
stubs = _normalize_sheet({"sheet": {}, "objects": stubs},
|
||||||
|
page["page_number"],
|
||||||
|
page_text=page_text)["assertions"]
|
||||||
|
before = len(sheet["assertions"])
|
||||||
|
sheet["assertions"] = merge_objects(sheet["assertions"], stubs)
|
||||||
|
print(f"[Extract] Page {page['page_number']}: fallback merged "
|
||||||
|
f"{len(sheet['assertions']) - before} text-layer stub(s)")
|
||||||
|
sheet["coverage"] = text_coverage(page_text,
|
||||||
|
sheet["assertions"])
|
||||||
|
|
||||||
|
# Identity recovery: never leave a text-bearing page sheet-less
|
||||||
|
if not sheet.get("sheet_number") and page_text:
|
||||||
|
recovered = recover_sheet_number(page_text)
|
||||||
|
if recovered:
|
||||||
|
sheet["sheet_number"] = recovered
|
||||||
|
sheet["discipline"] = (
|
||||||
|
__import__("backend.pipeline.extractor",
|
||||||
|
fromlist=["discipline_from_sheet_number"])
|
||||||
|
.discipline_from_sheet_number(recovered)
|
||||||
|
or sheet.get("discipline") or "Unknown")
|
||||||
|
print(f"[Extract] Page {page['page_number']}: sheet number "
|
||||||
|
f"recovered from text layer -> {recovered}")
|
||||||
|
|
||||||
|
return AgentResult(scope_id=scope.scope_id, artifacts=[sheet])
|
||||||
|
except Exception as exc:
|
||||||
|
return failure(scope, exc)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 4: Run to verify pass**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/agents/test_extraction_ladder.py -v`
|
||||||
|
Expected: all pass
|
||||||
|
|
||||||
|
**Step 5: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add backend/agents/extractors.py backend/config.py tests/agents/test_extraction_ladder.py
|
||||||
|
git commit -m "feat: coverage-driven extraction retry ladder (agent path)"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 5: Same ladder in the CLASSIC path (pipeline/extractor.py)
|
||||||
|
|
||||||
|
**Objective:** The classic pipeline (`_extract_one`, backend/pipeline/extractor.py:273-293) must get the identical ladder — two render paths share everything, per the documented project trap.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `backend/pipeline/extractor.py:273-293`
|
||||||
|
- Test: `tests/test_text_layer_flow.py` or new `tests/test_extraction_ladder_classic.py`
|
||||||
|
|
||||||
|
**Step 1: Write failing test** — mirror Task 4's tests against `_extract_one` directly (monkeypatch `backend.pipeline.extractor.call_json`).
|
||||||
|
|
||||||
|
**Step 2: Run to verify failure**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/test_extraction_ladder_classic.py -v`
|
||||||
|
Expected: FAIL
|
||||||
|
|
||||||
|
**Step 3: Implement** — same ladder shape as Task 4 but inside `_extract_one`; the text-structuring call here uses default model (no `model=` kwarg, matching existing `_extract_one` call_json usage). Keep the existing "extraction failed" empty-sheet shape for pages with NO text layer (scanned pages stay vision-only and may legitimately return empty).
|
||||||
|
|
||||||
|
**Step 4: Run to verify pass**
|
||||||
|
|
||||||
|
Run: `.venv/bin/python -m pytest tests/test_extraction_ladder_classic.py -v`
|
||||||
|
Expected: all pass
|
||||||
|
|
||||||
|
**Step 5: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add backend/pipeline/extractor.py tests/test_extraction_ladder_classic.py
|
||||||
|
git commit -m "feat: coverage-driven extraction retry ladder (classic path)"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 6: Verbatim-source stamping upgrade
|
||||||
|
|
||||||
|
**Objective:** When a text layer exists, check each object's source_text against the page text with the existing fuzzy machinery; stamp `grounding="vision_unverified"` when it doesn't match so the wave-5b verifier prioritizes it. Cheap upgrade, reuses text_layer._tokens — no new call sites.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `backend/pipeline/extractor.py` (`_normalize_sheet`, ~line 182 where `grounding` is stamped)
|
||||||
|
- Test: extend `tests/test_extractor_text_grounding.py`
|
||||||
|
|
||||||
|
**Step 1: Failing test** — object whose source_text is NOT a fuzzy substring of the page text keeps the object (guard passes via digits) but gets stamped `vision_unverified`.
|
||||||
|
|
||||||
|
**Step 2:** Run, expect FAIL.
|
||||||
|
|
||||||
|
**Step 3: Implement** — in `_normalize_sheet`, when `page_text` is present and no `grounding` stamp yet: normalized source_text (via `backend.text_coverage._norm`) not substring of normalized page text → `grounding = "vision_unverified"` (counted in the existing log line as a third counter).
|
||||||
|
|
||||||
|
**Step 4:** Run, expect PASS.
|
||||||
|
|
||||||
|
**Step 5: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add backend/pipeline/extractor.py tests/test_extractor_text_grounding.py
|
||||||
|
git commit -m "feat: stamp vision-unverified source_text against text layer"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 7: Surface coverage in the report
|
||||||
|
|
||||||
|
**Objective:** `report.summary` gains per-job extraction-quality visibility so "is extraction healthy?" is answerable without log spelunking.
|
||||||
|
|
||||||
|
**Files:**
|
||||||
|
- Modify: `backend/agents/runner.py` (where summary is assembled) and/or `backend/pipeline/report.py`
|
||||||
|
- Test: extend the runner-level stub test (tests/agents/test_wave5b_suppression.py pattern)
|
||||||
|
|
||||||
|
**Step 1: Failing test** — runner-level: summary contains `extraction_coverage = {"pages_below_floor": [...], "mean_ratio": float, "fallback_pages": [...]}`.
|
||||||
|
|
||||||
|
**Step 2-4:** Implement by aggregating the `coverage` dicts Task 4/5 attach to each sheet; NO new ProjectMemory keys (closed registry trap) — compute at report assembly from the sheets list already in scope.
|
||||||
|
|
||||||
|
**Step 5: Commit**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git add backend/agents/runner.py backend/pipeline/report.py tests/
|
||||||
|
git commit -m "feat: extraction coverage summary in report"
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Task 8: Full suite + Cypress validation run
|
||||||
|
|
||||||
|
**Step 1:** `.venv/bin/python -m pytest tests/ -q` — expected: all pass (128 + new).
|
||||||
|
|
||||||
|
**Step 2:** Push branch, wait for Gitea Actions sha-<short> build, deploy per the skill's deploy runbook (compose pull + up -d --force-recreate).
|
||||||
|
|
||||||
|
**Step 3:** Resubmit the exact Cypress PDF (`docker cp`'d source.pdf preserved at /tmp/cypress-source.pdf on sits-docker):
|
||||||
|
`curl -F file=@source.pdf -F pipeline_mode=agent https://conchecker.scoutitsystems.com/check`
|
||||||
|
|
||||||
|
**Step 4: Acceptance criteria (compare against job 3e01d5baba32):**
|
||||||
|
- Zero text-bearing pages with 0 assertions (was: pages 8, 10, 18).
|
||||||
|
- `sheet_number` present on >= 37/38 pages (was: 31/38).
|
||||||
|
- A102 RCP content (ceiling heights, tape lights, sconces) present in assertions.
|
||||||
|
- Spot-check: no regression in validated-issue quality — suppressed_issues and validated_issues counts within noise of the prior run; cost delta reported (expect +1 cheap text-only call per low-coverage page, ~$0 on healthy pages).
|
||||||
|
- `report.summary.extraction_coverage.pages_below_floor` is empty or every entry is a genuinely scanned page.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Files Touched (summary)
|
||||||
|
|
||||||
|
- Create: `backend/text_coverage.py`
|
||||||
|
- Modify: `backend/prompts.py`, `backend/config.py`, `backend/agents/extractors.py`, `backend/pipeline/extractor.py`, `backend/agents/runner.py`, `backend/pipeline/report.py`
|
||||||
|
- Tests: `tests/test_text_coverage.py`, `tests/agents/test_extraction_ladder.py`, `tests/test_extraction_ladder_classic.py`, extensions to `tests/test_extractor_text_grounding.py` and the runner-level stub test
|
||||||
|
|
||||||
|
## Risks, Tradeoffs, Open Questions
|
||||||
|
|
||||||
|
- **Fallback flood risk:** 200 low-confidence stubs/page could flood downstream scopes. Mitigations: EXTRACT_FALLBACK_MAX_OBJECTS cap, confidence=low (specialists already weight confidence), fallback only fires below the coverage floor (3/38 pages on Cypress). If Brain merge gets noisy, lower the cap or restrict fallback to pages where rungs 1+2 BOTH return 0 objects.
|
||||||
|
- **Cost:** rung 2 adds one text-only call per low-coverage page (~12k input chars, no image) — negligible vs the 65k-token vision pass. Healthy pages skip it entirely.
|
||||||
|
- **Sheet-id regex false positives:** member marks like W12X26 are excluded by the 2-3-digit shape, but "S301" inside a detail reference ("2/S301") will match. Tail-of-page preference mitigates; wrong-but-present sheet_number is still strictly better than None for scope keying (sheet_index wave can correct it).
|
||||||
|
- **Open question:** should rung 2 route to the local model (aimax LM Studio) instead of the cloud extractor model to make retries free? Config knob `AGENT_EXTRACT_TEXT_MODEL` would allow it; not included in this plan (YAGNI until cost data from the validation run says otherwise).
|
||||||
|
- **Explicit non-goal:** graphical-only content (symbol geometry, line work) stays vision-based — text-layer-first cannot see it. Scanned PDFs (no text layer) keep today's behaviour plus the existing failed_scopes gap note.
|
||||||
@@ -0,0 +1,60 @@
|
|||||||
|
# Plan: Brain-directed clarification pass (bounded hub-and-spoke)
|
||||||
|
|
||||||
|
Date: 2026-08-20
|
||||||
|
Branch: agent-mode
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
Let the Brain actively chase weak/ambiguous findings instead of only judging
|
||||||
|
the finished pile once. Bounded, traceable, reuses the wave-5b verifier as the
|
||||||
|
"answer" channel. NOT a free agentic loop.
|
||||||
|
|
||||||
|
## Shape (agent pipeline)
|
||||||
|
Insert **wave 6.5: Brain-directed clarification** between the wave-6 Brain merge
|
||||||
|
and the review-gate / wave-7 branches, so BOTH paths benefit.
|
||||||
|
|
||||||
|
1. `BrainAgent.plan_clarifications(prioritized)` — one focused LLM call. Brain
|
||||||
|
names findings it is unsure about and emits TYPED requests:
|
||||||
|
`{issue_id, request_type, reason}`. v1 executes only `verify_evidence`;
|
||||||
|
the router accepts other types but logs them as "planned, not executed"
|
||||||
|
(extensible without a rewrite). Capped at `BRAIN_CLARIFY_MAX_REQUESTS`.
|
||||||
|
Brain is told which findings already carry `verification` (from 5b) so it
|
||||||
|
does not re-request them.
|
||||||
|
2. Route `verify_evidence` requests → build verify scopes for exactly those
|
||||||
|
findings (reuse the SAME scope builder as wave 5b: fresh page images +
|
||||||
|
hi-DPI evidence crops + text-layer oracle) → `EvidenceVerifierAgent` →
|
||||||
|
`apply_verdicts(prioritized, ...)`. Refuted findings are annotated,
|
||||||
|
demoted, removed from `prioritized`, and pushed into `memory["suppressed"]`
|
||||||
|
(existing key — no memory-registry crash). Clarify decisions recorded in
|
||||||
|
`memory["decisions"]`.
|
||||||
|
3. No second full Brain merge: Brain ASKED (step 1) and the verifier ANSWERED
|
||||||
|
(step 2); the answer prunes/annotates the list. This keeps issue_ids stable
|
||||||
|
for the review queue and adds at most 1 + N calls. One iteration only.
|
||||||
|
|
||||||
|
## Bounds / knobs (config.py, all env-overridable)
|
||||||
|
- `ENABLE_BRAIN_CLARIFY` (_flag, default true)
|
||||||
|
- `BRAIN_CLARIFY_MAX_REQUESTS` (default 8)
|
||||||
|
- reuse `AGENT_VERIFY_CONCURRENCY`, `VERIFY_MAX_TOKENS`,
|
||||||
|
`AGENT_VERIFY_REASONING_EFFORT`, `AGENT_CONFLICT_MAX_IMAGES`,
|
||||||
|
`VERIFY_HI_DPI_CROPS`.
|
||||||
|
|
||||||
|
## Reuse / refactor
|
||||||
|
- Extract the inline wave-5b verify-scope construction into
|
||||||
|
`_build_verify_scopes(findings, targets, sheet_to_page, page_to_b64,
|
||||||
|
page_to_text, page_words, pdf_path, prefix)` so wave 5b and wave 6.5 share
|
||||||
|
it. Preserve wave-5b behavior exactly (its tests guard this).
|
||||||
|
|
||||||
|
## Classic pipeline
|
||||||
|
Out of scope for v1 — the verifier/crops live only in the agent path. Classic
|
||||||
|
keeps its single dedup_validate. Documented as agent-only.
|
||||||
|
|
||||||
|
## Tests
|
||||||
|
- `plan_clarifications` parses/caps/skips-already-verified (stub call_json).
|
||||||
|
- Router executes verify_evidence, ignores unknown types.
|
||||||
|
- Runner smoke: a low-confidence finding Brain flags gets refuted → moves to
|
||||||
|
suppressed_issues; stub verifier.call_json (no live calls).
|
||||||
|
|
||||||
|
## Pitfalls to respect
|
||||||
|
- ProjectMemory keys are a closed registry — only use existing `suppressed` /
|
||||||
|
`decisions`. (Skill defect C1.)
|
||||||
|
- Stub `backend.agents.verifier.call_json` in runner tests or it hits the net.
|
||||||
|
- Extractor stub sheets need >= 2 assertions or no clusters form.
|
||||||
@@ -0,0 +1,163 @@
|
|||||||
|
# Session Notes — Conflict Checker
|
||||||
|
|
||||||
|
Orientation for a new coding session. Setup and Docker details live in [README.md](README.md). This file tracks what the code actually does and what tends to waste time.
|
||||||
|
|
||||||
|
**Last updated:** 2026-08-02 · `agent-mode` branch (merged `main` tip `bf508bf`)
|
||||||
|
|
||||||
|
## What this is
|
||||||
|
|
||||||
|
Cross-discipline **design contradiction / senior-architect QAQC** for construction drawing PDFs (Arch, Struct, Mech, Elec, Plumb, FP, etc.). Flags disagreements *between disciplines* before a set goes to bid/permit.
|
||||||
|
|
||||||
|
**Not** IronBid’s scope-ownership conflict checker.
|
||||||
|
|
||||||
|
Source of truth: Scout IT Gitea — `gitea.scoutitsystems.com/woogi/Conflict_Checker`.
|
||||||
|
|
||||||
|
## Stack
|
||||||
|
|
||||||
|
| Layer | Detail |
|
||||||
|
|-------|--------|
|
||||||
|
| API | Python 3.12, FastAPI, Uvicorn ([backend/main.py](backend/main.py)) |
|
||||||
|
| UI | Single static file [frontend/index.html](frontend/index.html), served by FastAPI |
|
||||||
|
| Pipelines | **Classic:** [backend/pipeline/runner.py](backend/pipeline/runner.py) (shared by web + CLI). **Agent:** [backend/agents/runner.py](backend/agents/runner.py) — experimental scoped specialist agents, selected per job (`pipeline_mode`) |
|
||||||
|
| Review gate | Agent jobs stop at `needs_review` for human decisions before the report emails ([backend/review/](backend/review/)) |
|
||||||
|
| LLM | OpenRouter via `openai` SDK; default `google/gemini-2.5-pro`. Vision always OpenRouter; text stages can use local vLLM (classic/hybrid only) |
|
||||||
|
| PDF | `pdf2image` + system `poppler-utils` → JPEG page images |
|
||||||
|
| Jobs | In-memory threads ([backend/jobs.py](backend/jobs.py)) — no Redis/DB |
|
||||||
|
| Tests | `pytest tests/` (~82 tests; see Quick start) |
|
||||||
|
| Deploy | Docker Compose; app on port **8099**; public URL `https://conchecker.scoutitsystems.com` |
|
||||||
|
|
||||||
|
## Live pipeline (authoritative)
|
||||||
|
|
||||||
|
README still describes an older 5-stage extract-then-compare loop. **Trust `runner.py`.** Actual flow:
|
||||||
|
|
||||||
|
```
|
||||||
|
PDF → images → extract → sheet index → jurisdiction
|
||||||
|
→ normalize → project intelligence (GOIDs)
|
||||||
|
→ cluster → conflict reason
|
||||||
|
→ QAQC / code / constructability
|
||||||
|
→ dedup-validate → risk → RFIs → report
|
||||||
|
```
|
||||||
|
|
||||||
|
| Runner stage | Module | Notes |
|
||||||
|
|--------------|--------|--------|
|
||||||
|
| PDF → images | `pdf_processor` | Rasterize |
|
||||||
|
| Extract assertions | `extractor` | Vision, per sheet |
|
||||||
|
| Classify sheet index | `sheet_index` | LLM |
|
||||||
|
| Jurisdiction profile | `jurisdiction` | After cover meta |
|
||||||
|
| Normalize | `normalizer` | LLM batches |
|
||||||
|
| Project intelligence | `normalizer.build_project_intelligence` | GOIDs + relationships |
|
||||||
|
| Cluster | `llm_clusterer` or `clusterer` | Default `CLUSTERER=llm` |
|
||||||
|
| Conflicts | `conflict_checker` | Per-cluster vision reason |
|
||||||
|
| QAQC / code / constructability | `qaqc_review`, `code_review`, `constructability` | Full-set / batched |
|
||||||
|
| Validate & dedup | `validator` | Merges conflict + QAQC + code + construct issues |
|
||||||
|
| Risk / RFIs | `risk`, `rfi` | Text-only |
|
||||||
|
| Report | `report` | `conflicts.json` + `report.md` (+ stage JSON dumps when `out_dir` set) |
|
||||||
|
|
||||||
|
Design notes for Stage 2/3 engines also live under `Changes/*.docx`.
|
||||||
|
|
||||||
|
**Agent pipeline** (`pipeline_mode=agent`): scoped specialist agents in [backend/agents/](backend/agents/) (extractors, linker, brain, critics, RFI writer) run through `Orchestrator` + `ProjectMemory`; artifacts under `outputs/<job_id>/agent/`. Agent mode is OpenRouter-only (no hybrid) and, when `AGENT_REQUIRE_REVIEW=true`, stops at `needs_review` until a human saves decisions and finalizes via the review endpoints. Design docs: `docs/superpowers/`.
|
||||||
|
|
||||||
|
## HTTP API (current)
|
||||||
|
|
||||||
|
| Method | Path | Purpose |
|
||||||
|
|--------|------|---------|
|
||||||
|
| GET | `/health` | Liveness + `model` / `text_model` + `version` / `build` + key/email flags |
|
||||||
|
| GET | `/models` | OpenRouter catalog split into `vision[]` / `text[]` + `defaults`, with per-1M-token pricing (cached ~1h; **502** when OpenRouter is unreachable) |
|
||||||
|
| POST | `/check` | Upload PDF; returns `{job_id}` immediately |
|
||||||
|
| GET | `/jobs/{id}` | Status poll. Running: `stage` + `log_tail`. Done/error/needs_review/finalization_error: `report` and/or `error` + full `log` |
|
||||||
|
| GET | `/jobs/{id}/log` | Full run log as `text/plain` (404 when no log file) |
|
||||||
|
| GET | `/jobs/{id}/review` | Review queue + progress + saved decisions (agent jobs) |
|
||||||
|
| POST | `/jobs/{id}/review-decisions` | Save reviewer decisions (409 outside needs_review/reviewing) |
|
||||||
|
| POST | `/jobs/{id}/finalize-review` | Background finalize + send report (409 unless review gate passed) |
|
||||||
|
| GET | `/jobs/{id}/sheet-image/{page}` | JPEG of source PDF page for the sheet viewer |
|
||||||
|
| GET | `/` | Serves `frontend/index.html` |
|
||||||
|
|
||||||
|
### `/check` form fields
|
||||||
|
|
||||||
|
- Required: `file` (PDF)
|
||||||
|
- Optional: `notification_email`, `project_name`, `address`, `occupancy`, `work_type`
|
||||||
|
- Pipeline: `pipeline_mode` (`classic` default, or `agent`)
|
||||||
|
- Compute: `text_local` (`true` = hybrid local text; forced off for agent mode)
|
||||||
|
- Models: `vision_model`, `text_model` (OpenRouter ids; blank = config defaults)
|
||||||
|
|
||||||
|
## Job logs
|
||||||
|
|
||||||
|
Pipeline `print()` is teed for the job thread ([backend/job_log.py](backend/job_log.py)):
|
||||||
|
|
||||||
|
- Live: `GET /jobs/{id}` → `log_tail` (last 80 lines)
|
||||||
|
- Done/error/needs_review/finalization_error: same payload includes full `log`
|
||||||
|
- Disk: `backend/outputs/<job_id>/job.log` (survives restart; status registry does not)
|
||||||
|
- API: `GET /jobs/{id}/log` → `text/plain`
|
||||||
|
- UI: “Run log” panel updates while running; stays visible after finish/fail/review
|
||||||
|
- Failures append the **full traceback** to the log; each run starts with a header line (job id, mode, models, start time)
|
||||||
|
- `outputs/<job_id>/job.json` (written at start) carries email/mode/models so the disk fallback can rebuild a job after restart
|
||||||
|
- **Verbose LLM observability** (`LLM_VERBOSE=true`, default on): one `[LLM] #NNNN stage | model (backend) | in/out sizes | cost | parsed-item counts` line per call in `job.log` — list-valued keys show counts (`conflicts[3]`) so empty (miss) or invented (hallucination) results stand out
|
||||||
|
- **Raw dumps** (`LLM_RAW_DUMP=true`, default on): full prompt + raw response per call in `outputs/<job_id>/llm_raw/NNNN_stage_model.json` (base64 images excluded, `n_images` recorded). This is the source for tracing *why* the model missed/invented an item
|
||||||
|
- **End-of-log cost**: every run ends with an `=== Estimated LLM cost (this run) ===` block (total, per-stage, models); failed runs append a cost-so-far line. Local/hybrid calls have no $ accounting — total covers OpenRouter only
|
||||||
|
|
||||||
|
## Vision vs text models
|
||||||
|
|
||||||
|
Two models, not one:
|
||||||
|
|
||||||
|
| Role | Config | Stages | Backend |
|
||||||
|
|------|--------|--------|---------|
|
||||||
|
| Vision | `MODEL` | Extract, conflict reason (images) | Always OpenRouter |
|
||||||
|
| Text | `TEXT_MODEL` (falls back to `MODEL`) | Sheet index, jurisdiction, normalize, cluster(LLM), QAQC, code, construct, validate, risk, RFI | OpenRouter, or local when hybrid |
|
||||||
|
|
||||||
|
UI: two dropdowns (with per-1M pricing) filled from `GET /models` ([backend/models.py](backend/models.py)), shown only for OpenRouter compute. Per-run picks go through `set_model_overrides(vision, text)` in [backend/llm.py](backend/llm.py): classic runs pass them as `run_pipeline` kwargs (runner clears in `finally`); agent runs set them module-level around `run_agent_pipeline`. UI picks beat per-call agent `AGENT_*_MODEL` args but **never name the hybrid local model** — local stays on `LOCAL_TEXT_MODEL`; the text pick only covers the cloud fallback.
|
||||||
|
|
||||||
|
## Where to change what
|
||||||
|
|
||||||
|
| Concern | File |
|
||||||
|
|---------|------|
|
||||||
|
| Prompts, vocab, conflict taxonomy | [backend/prompts.py](backend/prompts.py) |
|
||||||
|
| Env knobs | [backend/config.py](backend/config.py), [backend/.env.example](backend/.env.example) |
|
||||||
|
| HTTP API | [backend/main.py](backend/main.py) |
|
||||||
|
| UI (upload, models, live log, results) | [frontend/index.html](frontend/index.html) |
|
||||||
|
| CLI tuning loop | [cli/run_check.py](cli/run_check.py) |
|
||||||
|
| LLM client, cache, cost, model overrides | [backend/llm.py](backend/llm.py) |
|
||||||
|
| OpenRouter catalog + pricing + vision/text split | [backend/models.py](backend/models.py) |
|
||||||
|
| Agent-mode pipeline | [backend/agents/](backend/agents/) (`runner.py` entry; agents call `llm.call_json` with per-call model args) |
|
||||||
|
| Human review gate (queue, decisions, finalize) | [backend/review/](backend/review/) |
|
||||||
|
| Job registry + stdout tee log | [backend/jobs.py](backend/jobs.py), [backend/job_log.py](backend/job_log.py) |
|
||||||
|
| Stage helpers (prompt render, issue validate) | [backend/pipeline/_stage.py](backend/pipeline/_stage.py) |
|
||||||
|
| Code text corpus (Stage 7) | [backend/code_corpus/](backend/code_corpus/) |
|
||||||
|
| Hybrid local LLM helper | [scripts/setup_vllm.sh](scripts/setup_vllm.sh) |
|
||||||
|
|
||||||
|
Older prompt snapshot: `backend/prompts.py.v1`.
|
||||||
|
|
||||||
|
## Gotchas
|
||||||
|
|
||||||
|
1. **Prompt placeholders** — Use `str.replace` via `_stage.render`, never `str.format`. Prompts contain literal `{` JSON braces.
|
||||||
|
2. **Clustering** — Default is LLM (`CLUSTERER=llm`); empty LLM result falls back to deterministic. Deterministic clusters need ≥2 disciplines (or schedule-vs-plan); single-discipline “missing” gaps are a known limit.
|
||||||
|
3. **Jobs are in-memory** — Process restart clears job status; `outputs/<job_id>/` (report + `job.log` + `job.json` + `source.pdf`) still reload via disk fallback, including `needs_review` recovery.
|
||||||
|
4. **Dependency pin** — `httpx==0.27.2` with `openai==1.51.0`. httpx ≥0.28 breaks openai’s `proxies=` kwarg.
|
||||||
|
5. **Code corpus licensing** — Only `ada_2010.txt` is shipped. Do not paste IBC/IFC/IECC without a license (see `backend/code_corpus/README.md`).
|
||||||
|
6. **Dual assertion schema** — Newer `{sheet, objects[]}` is mapped to legacy `{assertions[]}` with `attribute`/`value` for older stages.
|
||||||
|
7. **Grounding guard** — Extractor drops objects whose numeric claims are not in `source_text` (graphical-only objects allowed).
|
||||||
|
8. **Module globals** — LLM cost counters, stdout tee, and model overrides are process-global; overlapping jobs can interleave (single-user tool assumption).
|
||||||
|
9. **Samples** — `samples/*.pdf` are gitignored; drop PDFs locally for CLI runs.
|
||||||
|
10. **README drift** — Treat README for setup/CI; treat this file + `runner.py` for pipeline truth. `prompts.py` header may still say some prompts are unwired — they are wired through the runner.
|
||||||
|
11. **Git identity** — This box has no `user.name` / `user.email`; commits need `GIT_AUTHOR_*` / `GIT_COMMITTER_*` env vars (do not `git config`). Remote push to Gitea works.
|
||||||
|
12. **No local Python deps on host** — App is meant to run in Docker; bare `python3` imports may miss `dotenv` / `httpx`. Prefer `docker compose`. For tests on this box: venv + requirements, but unpin Pillow (`Pillow>=11`) — 10.4.0 doesn't build on Python 3.14 (Docker uses 3.12, where the pin is fine).
|
||||||
|
13. **Agent mode constraints** — OpenRouter-only (hybrid disabled in UI and forced off server-side); review gate statuses are `needs_review → reviewing → finalizing → done` (`finalization_error` on finalize failure); only terminal states include the full `log` in polls.
|
||||||
|
|
||||||
|
## Quick start pointers
|
||||||
|
|
||||||
|
- Full setup: [README.md](README.md) (`docker compose up -d --build` → http://localhost:8099).
|
||||||
|
- Local CLI: `python cli/run_check.py samples/your_set.pdf --out out/your_set`.
|
||||||
|
- Prompt iteration: set `LLM_CACHE=true` in `backend/.env` so unchanged stages replay for free; clear with `rm -rf backend/.llm_cache`.
|
||||||
|
- Artifacts per job: `assertions.json`, `clusters.json`, per-stage JSON, `conflicts.json`, `report.md`, `job.log`, `job.json`, `source.pdf` under `backend/outputs/<job_id>/`.
|
||||||
|
- Tests: `python -m pytest tests/` (needs the deps from `requirements.txt` + `pytest`; on this box use a venv, see gotcha #12).
|
||||||
|
|
||||||
|
## Recent work (2026-08-02, agent-mode)
|
||||||
|
|
||||||
|
Merged `main` tip (`a6b0c8f` + `bf508bf`) into `agent-mode`, reconciling with this branch's own earlier implementations:
|
||||||
|
|
||||||
|
- **Two model dropdowns** — main's vision/text split ported onto this branch's priced catalog (`models.py`); pickers stay OpenRouter-compute-only, and UI picks never override `LOCAL_TEXT_MODEL` (main's hybrid footgun avoided).
|
||||||
|
- **Better run logs** — main's timestamped line-splitting tee, `log_tail` polls, terminal-state full log, and log-only disk recovery merged with this branch's header line, `job.json` metadata, and review-gate states. Failed runs now also append the traceback to `job.log`.
|
||||||
|
- `backend/models_catalog.py` (main's unpriced catalog) intentionally dropped in favor of `models.py`.
|
||||||
|
|
||||||
|
## Conflict categories (taxonomy)
|
||||||
|
|
||||||
|
Defined in `backend/prompts.py`: `dimensional_disagreement`, `elevation_disagreement`, `location_mismatch`, `missing_element`, `schedule_vs_plan_mismatch`, `detail_vs_plan_mismatch`, `tag_or_reference_inconsistency`, `spatial_clash`, `note_or_spec_contradiction`.
|
||||||
@@ -212,10 +212,54 @@ From the CLI, `--no-review` bypasses the gate for that run (it overrides
|
|||||||
python cli/run_check.py samples/your_set.pdf --mode agent --no-review --out out/agent-run
|
python cli/run_check.py samples/your_set.pdf --mode agent --no-review --out out/agent-run
|
||||||
```
|
```
|
||||||
|
|
||||||
|
### Asking the run why: review chat
|
||||||
|
|
||||||
|
Each item on the review screen has an **Ask about this finding** panel, and the
|
||||||
|
screen carries one **Ask about this run** panel for questions that are not about
|
||||||
|
a single finding. The chat answers from the job's own artifacts — the finding's
|
||||||
|
evidence, the cluster it came from, the raw per-sheet extraction, the
|
||||||
|
verification verdict, the Brain's merge decision, the sheet index, the cover-index
|
||||||
|
reconciliation, and matching `job.log` lines.
|
||||||
|
|
||||||
|
```
|
||||||
|
"why does it think the AC unit is mounted on the ground?" -> item scope
|
||||||
|
"why didn't it pick up on the Civil set?" -> run scope
|
||||||
|
```
|
||||||
|
|
||||||
|
The chat is **read-only**. It cannot change a finding, a severity, a decision,
|
||||||
|
or the report, and the prompt forbids it from proposing code or config changes —
|
||||||
|
the radio buttons remain the only thing that alters review state. When the
|
||||||
|
artifacts do not contain the answer, it says so and names what is missing rather
|
||||||
|
than guessing.
|
||||||
|
|
||||||
|
Every turn is logged twice:
|
||||||
|
|
||||||
|
- `outputs/<job_id>/review/chat_log.jsonl` — the auditable record: the issue as
|
||||||
|
it stood when asked about, the question, the answer, the determinations, and
|
||||||
|
the evidence quoted. Readable as a transcript at
|
||||||
|
`GET /jobs/{id}/review-chat/log`.
|
||||||
|
- `REVIEW_FEEDBACK_DIR/chat_turns.jsonl` — the cross-job roll-up, alongside
|
||||||
|
`decisions.jsonl`. When a reviewer corrects a misidentification in
|
||||||
|
conversation ("that is not a floor drain, it is a power floor box"), the
|
||||||
|
correction is captured as `suggested_category_correction` rather than dying in
|
||||||
|
free text. Nothing reads this store yet; writing it is what makes priming a
|
||||||
|
future run on past corrections possible.
|
||||||
|
|
||||||
|
| Key | Default | Effect |
|
||||||
|
|-----|---------|--------|
|
||||||
|
| `ENABLE_REVIEW_CHAT` | `true` | `false` = the chat endpoints refuse and the panels stay empty |
|
||||||
|
| `REVIEW_CHAT_MODEL` | `TEXT_MODEL` | Model for chat answers |
|
||||||
|
| `REVIEW_CHAT_MAX_TOKENS` | `4096` | Answer budget |
|
||||||
|
| `REVIEW_CHAT_HISTORY_TURNS` | `6` | Prior turns replayed into a thread's prompt |
|
||||||
|
| `REVIEW_CHAT_LOG_LINES` | `40` | Max `job.log` lines pulled into the context bundle |
|
||||||
|
| `REVIEW_FEEDBACK_DIR` | `backend/outputs/_feedback` | Cross-job decision + chat feedback store |
|
||||||
|
|
||||||
**Deployment note:** the review endpoints (`/jobs/{id}/review-decisions`,
|
**Deployment note:** the review endpoints (`/jobs/{id}/review-decisions`,
|
||||||
`/jobs/{id}/finalize-review`) are **state-changing and sensitive** — they accept
|
`/jobs/{id}/finalize-review`) are **state-changing and sensitive** — they accept
|
||||||
human decisions that alter the final report. Do **not** expose the UI/API
|
human decisions that alter the final report. `/jobs/{id}/review-chat` does not
|
||||||
publicly without reverse-proxy auth or a shared access token in front of it.
|
change review state, but it does spend model budget and returns drawing
|
||||||
|
evidence. Do **not** expose the UI/API publicly without reverse-proxy auth or a
|
||||||
|
shared access token in front of it.
|
||||||
|
|
||||||
Web UI (upload + view):
|
Web UI (upload + view):
|
||||||
|
|
||||||
|
|||||||
+75
-1
@@ -27,6 +27,23 @@ AGENT_CONFLICT_CONCURRENCY=4
|
|||||||
AGENT_SPECIALIST_CONCURRENCY=4
|
AGENT_SPECIALIST_CONCURRENCY=4
|
||||||
AGENT_RFI_CONCURRENCY=4
|
AGENT_RFI_CONCURRENCY=4
|
||||||
|
|
||||||
|
# -- Review focus toggles -------------------------------------------
|
||||||
|
# ENABLE_CODE_REVIEW: run the code/ADA/jurisdiction review path (both pipelines).
|
||||||
|
# Default OFF - the product focuses on drawing integrity and cross-discipline
|
||||||
|
# coordination, not code/accessibility compliance. Set to 1 to restore it.
|
||||||
|
ENABLE_CODE_REVIEW=false
|
||||||
|
# ENABLE_DRAWING_INTEGRITY: per-sheet Drawing Integrity QA wave (both pipelines).
|
||||||
|
# The drawing-focused pass - dangling references, on-sheet contradictions,
|
||||||
|
# dimension sanity, missing sheet essentials, tag hygiene. Default ON.
|
||||||
|
ENABLE_DRAWING_INTEGRITY=true
|
||||||
|
AGENT_INTEGRITY_MODEL=
|
||||||
|
AGENT_INTEGRITY_CONCURRENCY=4
|
||||||
|
AGENT_INTEGRITY_MAX_IMAGES=1
|
||||||
|
AGENT_INTEGRITY_MAX_ASSERTIONS=80
|
||||||
|
INTEGRITY_MAX_TOKENS=16384
|
||||||
|
# Skip sheets with fewer than this many extracted objects (too sparse to check)
|
||||||
|
INTEGRITY_MIN_ASSERTIONS=3
|
||||||
|
|
||||||
# Agent-mode human-review gate (pipeline stops after Brain until a human reviews)
|
# Agent-mode human-review gate (pipeline stops after Brain until a human reviews)
|
||||||
AGENT_REQUIRE_REVIEW=true
|
AGENT_REQUIRE_REVIEW=true
|
||||||
# Max clean clusters added to the review queue as non-blocking spot-checks
|
# Max clean clusters added to the review queue as non-blocking spot-checks
|
||||||
@@ -34,12 +51,33 @@ AGENT_REVIEW_AUDIT_SAMPLE=5
|
|||||||
# Allow future cross-job review-feedback aggregation to include source_text/images/comments
|
# Allow future cross-job review-feedback aggregation to include source_text/images/comments
|
||||||
REVIEW_AGGREGATE_INCLUDE_TEXT=false
|
REVIEW_AGGREGATE_INCLUDE_TEXT=false
|
||||||
|
|
||||||
|
# Review chat: read-only Q&A about findings and coverage on the review screen.
|
||||||
|
# It explains what the run did from the job's artifacts; it never changes a
|
||||||
|
# finding, a decision, or the report.
|
||||||
|
ENABLE_REVIEW_CHAT=true
|
||||||
|
# Model for chat answers (blank inherits TEXT_MODEL)
|
||||||
|
REVIEW_CHAT_MODEL=
|
||||||
|
REVIEW_CHAT_MAX_TOKENS=4096
|
||||||
|
# Prior turns replayed into a thread's prompt
|
||||||
|
REVIEW_CHAT_HISTORY_TURNS=6
|
||||||
|
# Max job.log lines searched into the chat's context bundle
|
||||||
|
REVIEW_CHAT_LOG_LINES=40
|
||||||
|
REVIEW_CHAT_MAX_QUESTION_CHARS=2000
|
||||||
|
# Cross-job store for review decisions + chat turns (blank = backend/outputs/_feedback)
|
||||||
|
REVIEW_FEEDBACK_DIR=
|
||||||
|
|
||||||
# Pipeline tuning
|
# Pipeline tuning
|
||||||
PDF_DPI=100
|
PDF_DPI=100
|
||||||
MAX_PAGES=60
|
MAX_PAGES=60
|
||||||
MAX_DIMENSION=2400
|
MAX_DIMENSION=2400
|
||||||
LLM_TIMEOUT=180
|
LLM_TIMEOUT=180
|
||||||
EXTRACT_MAX_TOKENS=8192
|
EXTRACT_MAX_TOKENS=65536
|
||||||
|
# Reasoning effort for per-sheet extraction (low keeps Gemini thinking tokens
|
||||||
|
# from eating the output budget). Blank = don't send the parameter.
|
||||||
|
EXTRACT_REASONING_EFFORT=low
|
||||||
|
# Hard thinking-token budget for extraction (OpenRouter reasoning max_tokens /
|
||||||
|
# Gemini thinking_budget). Stronger than effort; 0 = fall back to effort only.
|
||||||
|
EXTRACT_REASONING_MAX_TOKENS=2048
|
||||||
REASON_MAX_TOKENS=4096
|
REASON_MAX_TOKENS=4096
|
||||||
EXTRACT_CONCURRENCY=4
|
EXTRACT_CONCURRENCY=4
|
||||||
REASON_CONCURRENCY=4
|
REASON_CONCURRENCY=4
|
||||||
@@ -48,6 +86,13 @@ REASON_CONCURRENCY=4
|
|||||||
APP_BASE_URL=https://conchecker.scoutitsystems.com
|
APP_BASE_URL=https://conchecker.scoutitsystems.com
|
||||||
# APP_BUILD is set by CI at image build time (sha-<short_sha>) - do not set manually.
|
# APP_BUILD is set by CI at image build time (sha-<short_sha>) - do not set manually.
|
||||||
|
|
||||||
|
# LLM observability (job-log verbosity + raw request/response dumps)
|
||||||
|
# LLM_VERBOSE: one line per LLM call in job.log (model, sizes, item counts, cost)
|
||||||
|
# LLM_RAW_DUMP: full prompt+response per call in outputs/<job_id>/llm_raw/
|
||||||
|
# (base64 images excluded). Both default on; set false to quiet down.
|
||||||
|
LLM_VERBOSE=true
|
||||||
|
LLM_RAW_DUMP=true
|
||||||
|
|
||||||
# Email notifications (optional). Leave SMTP_HOST blank to disable.
|
# Email notifications (optional). Leave SMTP_HOST blank to disable.
|
||||||
# Examples:
|
# Examples:
|
||||||
# Gmail: SMTP_HOST=smtp.gmail.com SMTP_PORT=587 (use an App Password)
|
# Gmail: SMTP_HOST=smtp.gmail.com SMTP_PORT=587 (use an App Password)
|
||||||
@@ -59,3 +104,32 @@ SMTP_PASSWORD=
|
|||||||
SMTP_FROM=
|
SMTP_FROM=
|
||||||
SMTP_USE_TLS=true
|
SMTP_USE_TLS=true
|
||||||
SMTP_USE_SSL=false
|
SMTP_USE_SSL=false
|
||||||
|
|
||||||
|
# Wave 5b evidence verification (vision fact-check of cited sheet text)
|
||||||
|
AGENT_VERIFY_MAX_CHECKS=20
|
||||||
|
AGENT_VERIFY_SEVERITIES=critical,high
|
||||||
|
AGENT_VERIFY_REASONING_EFFORT=low
|
||||||
|
VERIFY_MAX_TOKENS=8192
|
||||||
|
|
||||||
|
# Wave 6.5 Brain-directed clarification (bounded hub-and-spoke). After the Brain
|
||||||
|
# merge, the Brain names findings it is unsure about; verify_evidence requests
|
||||||
|
# route back through the wave-5b verifier. One planning call + at most
|
||||||
|
# BRAIN_CLARIFY_MAX_REQUESTS verifications, single iteration. Default ON.
|
||||||
|
ENABLE_BRAIN_CLARIFY=true
|
||||||
|
BRAIN_CLARIFY_MAX_REQUESTS=8
|
||||||
|
BRAIN_CLARIFY_MAX_TOKENS=4096
|
||||||
|
|
||||||
|
# Text-layer grounding (deterministic PDF text layer via PyMuPDF)
|
||||||
|
# TEXT_LAYER_ENABLED: master switch for text-layer extraction/grounding
|
||||||
|
# TEXT_LAYER_MIN_CHARS: below this per page the sheet stays vision-only
|
||||||
|
# TEXT_LAYER_MAX_CHARS: cap of text layer injected into the extractor prompt
|
||||||
|
# VERIFY_TEXT_MAX_CHARS: cap of the text-layer excerpt in verify scopes
|
||||||
|
# VERIFY_HI_DPI_CROPS: evidence-located high-DPI crops in the verifier
|
||||||
|
# VERIFY_CROP_DPI / VERIFY_CROP_MARGIN_PTS: crop render DPI / padding (PDF points)
|
||||||
|
TEXT_LAYER_ENABLED=true
|
||||||
|
TEXT_LAYER_MIN_CHARS=20
|
||||||
|
TEXT_LAYER_MAX_CHARS=12000
|
||||||
|
VERIFY_TEXT_MAX_CHARS=8000
|
||||||
|
VERIFY_HI_DPI_CROPS=true
|
||||||
|
VERIFY_CROP_DPI=300
|
||||||
|
VERIFY_CROP_MARGIN_PTS=36
|
||||||
+72
-1
@@ -6,7 +6,12 @@ from typing import Dict, List, Tuple
|
|||||||
|
|
||||||
from backend import config
|
from backend import config
|
||||||
from backend.agents.base import AgentUsage
|
from backend.agents.base import AgentUsage
|
||||||
from backend.agents.prompts import BRAIN_SYSTEM_PROMPT, BRAIN_USER_PROMPT
|
from backend.agents.prompts import (
|
||||||
|
BRAIN_CLARIFY_SYSTEM_PROMPT,
|
||||||
|
BRAIN_CLARIFY_USER_PROMPT,
|
||||||
|
BRAIN_SYSTEM_PROMPT,
|
||||||
|
BRAIN_USER_PROMPT,
|
||||||
|
)
|
||||||
from backend.llm import call_json
|
from backend.llm import call_json
|
||||||
from backend.pipeline._stage import collect_list, validate_issue
|
from backend.pipeline._stage import collect_list, validate_issue
|
||||||
|
|
||||||
@@ -127,3 +132,69 @@ class BrainAgent:
|
|||||||
issues.sort(key=lambda item: -int(item.get("risk_score") or 0))
|
issues.sort(key=lambda item: -int(item.get("risk_score") or 0))
|
||||||
decisions = parsed.get("decisions") or []
|
decisions = parsed.get("decisions") or []
|
||||||
return issues, [item for item in decisions if isinstance(item, dict)]
|
return issues, [item for item in decisions if isinstance(item, dict)]
|
||||||
|
|
||||||
|
def plan_clarifications(self, prioritized: List[Dict]) -> List[Dict]:
|
||||||
|
"""Wave 6.5 planning call: name kept findings the Brain wants to
|
||||||
|
double-check before publishing, as typed clarification requests.
|
||||||
|
|
||||||
|
Returns a capped list of {issue_id, request_type, reason}. Only
|
||||||
|
findings that carry an issue_id and do NOT already have a verification
|
||||||
|
result are offered to the model; anything the model names outside that
|
||||||
|
set, or with an unknown request_type, is dropped by the caller/router.
|
||||||
|
Never raises — a failed/empty plan just yields no requests.
|
||||||
|
"""
|
||||||
|
max_requests = config.BRAIN_CLARIFY_MAX_REQUESTS
|
||||||
|
if not prioritized or max_requests <= 0:
|
||||||
|
return []
|
||||||
|
candidates = [
|
||||||
|
{
|
||||||
|
"issue_id": f.get("issue_id"),
|
||||||
|
"severity": f.get("severity"),
|
||||||
|
"confidence": f.get("confidence"),
|
||||||
|
"source_stage": f.get("source_stage"),
|
||||||
|
"description": (f.get("description") or "")[:400],
|
||||||
|
"evidence": f.get("evidence") or [],
|
||||||
|
"already_verified": bool(f.get("verification")),
|
||||||
|
}
|
||||||
|
for f in prioritized
|
||||||
|
if f.get("issue_id") and not f.get("verification")
|
||||||
|
]
|
||||||
|
if not candidates:
|
||||||
|
return []
|
||||||
|
instruction = (
|
||||||
|
BRAIN_CLARIFY_USER_PROMPT
|
||||||
|
.replace("{max_requests}", str(max_requests))
|
||||||
|
.replace("{findings}", json.dumps(candidates, ensure_ascii=True))
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
parsed = call_json(
|
||||||
|
system_prompt=BRAIN_CLARIFY_SYSTEM_PROMPT,
|
||||||
|
user_text=instruction,
|
||||||
|
max_tokens=config.BRAIN_CLARIFY_MAX_TOKENS,
|
||||||
|
model=config.AGENT_BRAIN_MODEL,
|
||||||
|
usage_tracker=self.usage,
|
||||||
|
usage_stage="agent.brain_clarify",
|
||||||
|
)
|
||||||
|
except Exception:
|
||||||
|
return []
|
||||||
|
raw = parsed.get("requests") if isinstance(parsed, dict) else parsed
|
||||||
|
if not isinstance(raw, list):
|
||||||
|
return []
|
||||||
|
valid_ids = {c["issue_id"] for c in candidates}
|
||||||
|
requests: List[Dict] = []
|
||||||
|
seen: set = set()
|
||||||
|
for item in raw:
|
||||||
|
if not isinstance(item, dict):
|
||||||
|
continue
|
||||||
|
issue_id = item.get("issue_id")
|
||||||
|
if issue_id not in valid_ids or issue_id in seen:
|
||||||
|
continue
|
||||||
|
requests.append({
|
||||||
|
"issue_id": issue_id,
|
||||||
|
"request_type": (item.get("request_type") or "verify_evidence").strip(),
|
||||||
|
"reason": (item.get("reason") or "").strip(),
|
||||||
|
})
|
||||||
|
seen.add(issue_id)
|
||||||
|
if len(requests) >= max_requests:
|
||||||
|
break
|
||||||
|
return requests
|
||||||
@@ -60,6 +60,8 @@ class ConflictCriticAgent:
|
|||||||
model=config.AGENT_CONFLICT_MODEL,
|
model=config.AGENT_CONFLICT_MODEL,
|
||||||
usage_tracker=self.usage,
|
usage_tracker=self.usage,
|
||||||
usage_stage="agent.conflict",
|
usage_stage="agent.conflict",
|
||||||
|
reasoning_effort=config.EXTRACT_REASONING_EFFORT or None,
|
||||||
|
reasoning_max_tokens=config.EXTRACT_REASONING_MAX_TOKENS or None,
|
||||||
)
|
)
|
||||||
candidates = parsed if isinstance(parsed, list) else (
|
candidates = parsed if isinstance(parsed, list) else (
|
||||||
parsed.get("conflicts") if isinstance(parsed, dict) else []
|
parsed.get("conflicts") if isinstance(parsed, dict) else []
|
||||||
|
|||||||
@@ -47,6 +47,7 @@ class ConstructabilityAgent:
|
|||||||
"assertions": dumps(cluster["assertions"]),
|
"assertions": dumps(cluster["assertions"]),
|
||||||
"clusters": dumps(slim_clusters([cluster])),
|
"clusters": dumps(slim_clusters([cluster])),
|
||||||
"conflicts": dumps(scope.payload.get("conflicts") or []),
|
"conflicts": dumps(scope.payload.get("conflicts") or []),
|
||||||
|
"disputes": dumps(cluster.get("disputed_attributes") or []),
|
||||||
}
|
}
|
||||||
for key, value in substitutions.items():
|
for key, value in substitutions.items():
|
||||||
instruction = instruction.replace("{" + key + "}", value)
|
instruction = instruction.replace("{" + key + "}", value)
|
||||||
@@ -64,7 +65,8 @@ class ConstructabilityAgent:
|
|||||||
lambda item: validate_issue(item, "constructability"),
|
lambda item: validate_issue(item, "constructability"),
|
||||||
)
|
)
|
||||||
for finding in findings:
|
for finding in findings:
|
||||||
finding.update(agent=self.name, scope_id=scope.scope_id)
|
finding.update(agent=self.name, scope_id=scope.scope_id,
|
||||||
|
cluster_key=cluster.get("key"))
|
||||||
return AgentResult(scope_id=scope.scope_id, artifacts=findings)
|
return AgentResult(scope_id=scope.scope_id, artifacts=findings)
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
return failure(scope, exc)
|
return failure(scope, exc)
|
||||||
@@ -0,0 +1,55 @@
|
|||||||
|
"""Deterministic detection of contradictory extracted values within a cluster.
|
||||||
|
|
||||||
|
Extraction is a vision pass: quantities and sizes can be misread ("(2) 2x6" vs
|
||||||
|
"(5) 2x6"). Cluster members are supposed to describe the same real-world
|
||||||
|
element, so two members asserting different values for the same attribute are
|
||||||
|
a probable misread. Flag these so downstream text-only stages treat the value
|
||||||
|
as unverified instead of reasoning from one reading.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import re
|
||||||
|
from typing import Dict, List
|
||||||
|
|
||||||
|
|
||||||
|
def _norm(value) -> str:
|
||||||
|
return re.sub(r"\s+", " ", ("" if value is None else str(value)).strip().lower())
|
||||||
|
|
||||||
|
|
||||||
|
def find_disputes(assertions: List[Dict]) -> List[Dict]:
|
||||||
|
"""Same attribute with >= 2 distinct normalized values = disputed."""
|
||||||
|
groups: Dict[str, Dict[str, Dict]] = {}
|
||||||
|
for assertion in assertions:
|
||||||
|
attribute = _norm(assertion.get("attribute"))
|
||||||
|
value = _norm(assertion.get("value"))
|
||||||
|
if not attribute or not value:
|
||||||
|
continue
|
||||||
|
# Group on the normalized value, but keep the original (whitespace-
|
||||||
|
# collapsed) text so disputes read like the sheet, not a lowercase munge.
|
||||||
|
original = re.sub(r"\s+", " ", str(assertion.get("value")).strip())
|
||||||
|
bucket = groups.setdefault(attribute, {}).setdefault(
|
||||||
|
value, {"original": original, "ids": set()}
|
||||||
|
)
|
||||||
|
bucket["ids"].add(assertion.get("id"))
|
||||||
|
disputes = []
|
||||||
|
for attribute, values in sorted(groups.items()):
|
||||||
|
if len(values) < 2:
|
||||||
|
continue
|
||||||
|
disputes.append({
|
||||||
|
"attribute": attribute,
|
||||||
|
"values": sorted(v["original"] for v in values.values()),
|
||||||
|
"assertion_ids": sorted(
|
||||||
|
aid for v in values.values() for aid in v["ids"] if aid
|
||||||
|
),
|
||||||
|
})
|
||||||
|
return disputes
|
||||||
|
|
||||||
|
|
||||||
|
def annotate_clusters(clusters: List[Dict]) -> int:
|
||||||
|
"""Attach disputed_attributes to each cluster that has any. Returns count."""
|
||||||
|
annotated = 0
|
||||||
|
for cluster in clusters:
|
||||||
|
disputes = find_disputes(cluster.get("assertions") or [])
|
||||||
|
if disputes:
|
||||||
|
cluster["disputed_attributes"] = disputes
|
||||||
|
annotated += 1
|
||||||
|
return annotated
|
||||||
+145
-10
@@ -6,7 +6,11 @@ from typing import Dict
|
|||||||
from backend import config
|
from backend import config
|
||||||
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
|
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
|
||||||
from backend.llm import call_json
|
from backend.llm import call_json
|
||||||
from backend.pipeline.extractor import _normalize_sheet
|
from backend.pipeline.extractor import (
|
||||||
|
_normalize_sheet,
|
||||||
|
_text_layer_block,
|
||||||
|
discipline_from_sheet_number,
|
||||||
|
)
|
||||||
from backend.pipeline.sheet_index import _index_input
|
from backend.pipeline.sheet_index import _index_input
|
||||||
from backend.prompts import (
|
from backend.prompts import (
|
||||||
EXTRACTOR_SYSTEM_PROMPT,
|
EXTRACTOR_SYSTEM_PROMPT,
|
||||||
@@ -18,19 +22,37 @@ from backend.prompts import (
|
|||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# Appended to the extractor instruction on the second-chance retry. Dense plan
|
||||||
|
# sheets blow the output budget on the full schema; compact mode trades
|
||||||
|
# per-object verbosity for actually finishing the page.
|
||||||
|
_COMPACT_RETRY_SUFFIX = """
|
||||||
|
IMPORTANT - COMPACT RETRY: the first pass did not complete. Keep the SAME JSON
|
||||||
|
schema, but extract at most 40 objects, prioritizing coordination-relevant
|
||||||
|
items (equipment, fixtures, devices, keynotes, dimensions, markers/callouts,
|
||||||
|
schedule rows). Keep descriptions/attributes short; skip review_uses entries
|
||||||
|
you are unsure about. Finish the JSON - a smaller complete answer beats a
|
||||||
|
larger truncated one."""
|
||||||
|
|
||||||
|
|
||||||
|
def _wrap_bare_list(parsed, page_number: int):
|
||||||
|
"""Models sometimes skip the {sheet, objects} wrapper and return a bare
|
||||||
|
objects array (especially after truncation repair). Accept it - the sheet
|
||||||
|
header falls back to title-block deduction downstream."""
|
||||||
|
if isinstance(parsed, list):
|
||||||
|
print(f"[Extract] Page {page_number}: wrapping bare objects array "
|
||||||
|
f"({len(parsed)} items, no sheet header)")
|
||||||
|
return {"sheet": {}, "objects": parsed}
|
||||||
|
return parsed
|
||||||
|
|
||||||
|
|
||||||
class SheetExtractorAgent:
|
class SheetExtractorAgent:
|
||||||
name = "sheet_extractor"
|
name = "sheet_extractor"
|
||||||
|
|
||||||
def __init__(self, usage: AgentUsage) -> None:
|
def __init__(self, usage: AgentUsage) -> None:
|
||||||
self.usage = usage
|
self.usage = usage
|
||||||
|
|
||||||
def run(self, scope: AgentScope) -> AgentResult:
|
def _call(self, instruction: str, page: Dict):
|
||||||
try:
|
return call_json(
|
||||||
page = scope.payload["page"]
|
|
||||||
instruction = EXTRACTOR_USER_INSTRUCTION.replace(
|
|
||||||
"{sheet_hint}", str(scope.payload.get("sheet_hint") or "")
|
|
||||||
)
|
|
||||||
parsed = call_json(
|
|
||||||
system_prompt=EXTRACTOR_SYSTEM_PROMPT,
|
system_prompt=EXTRACTOR_SYSTEM_PROMPT,
|
||||||
user_text=instruction,
|
user_text=instruction,
|
||||||
images_b64=[page["base64"]],
|
images_b64=[page["base64"]],
|
||||||
@@ -38,10 +60,123 @@ class SheetExtractorAgent:
|
|||||||
model=config.AGENT_EXTRACT_MODEL,
|
model=config.AGENT_EXTRACT_MODEL,
|
||||||
usage_tracker=self.usage,
|
usage_tracker=self.usage,
|
||||||
usage_stage="agent.extract",
|
usage_stage="agent.extract",
|
||||||
|
reasoning_effort=config.EXTRACT_REASONING_EFFORT or None,
|
||||||
|
reasoning_max_tokens=config.EXTRACT_REASONING_MAX_TOKENS or None,
|
||||||
|
)
|
||||||
|
|
||||||
|
def _text_structuring_call(self, instruction_page: Dict, sheet_hint: str):
|
||||||
|
"""Rung 2: text-only structuring pass over the page's text layer
|
||||||
|
(no image). Recovers text content the vision pass missed."""
|
||||||
|
from backend.prompts import (TEXT_STRUCTURING_SYSTEM_PROMPT,
|
||||||
|
TEXT_STRUCTURING_USER_INSTRUCTION)
|
||||||
|
instruction = (TEXT_STRUCTURING_USER_INSTRUCTION
|
||||||
|
.replace("{sheet_hint}", str(sheet_hint or ""))
|
||||||
|
.replace("{text_layer}",
|
||||||
|
(instruction_page.get("text_layer") or "")
|
||||||
|
[:config.TEXT_LAYER_MAX_CHARS]))
|
||||||
|
return call_json(
|
||||||
|
system_prompt=TEXT_STRUCTURING_SYSTEM_PROMPT,
|
||||||
|
user_text=instruction,
|
||||||
|
images_b64=None,
|
||||||
|
max_tokens=config.EXTRACT_MAX_TOKENS,
|
||||||
|
model=config.AGENT_EXTRACT_MODEL,
|
||||||
|
usage_tracker=self.usage,
|
||||||
|
usage_stage="agent.extract_text",
|
||||||
|
reasoning_effort=config.EXTRACT_REASONING_EFFORT or None,
|
||||||
|
reasoning_max_tokens=config.EXTRACT_REASONING_MAX_TOKENS or None,
|
||||||
|
)
|
||||||
|
|
||||||
|
def run(self, scope: AgentScope) -> AgentResult:
|
||||||
|
from backend.text_coverage import (fallback_objects, merge_objects,
|
||||||
|
recover_sheet_number, text_coverage)
|
||||||
|
try:
|
||||||
|
page = scope.payload["page"]
|
||||||
|
hint = scope.payload.get("sheet_hint") or ""
|
||||||
|
page_text = page.get("text_layer")
|
||||||
|
instruction = EXTRACTOR_USER_INSTRUCTION.replace(
|
||||||
|
"{sheet_hint}", str(hint)) + _text_layer_block(page)
|
||||||
|
|
||||||
|
# Rung 1: vision pass (unchanged behaviour, incl. compact retry)
|
||||||
|
parsed = _wrap_bare_list(self._call(instruction, page),
|
||||||
|
page["page_number"])
|
||||||
|
if not isinstance(parsed, dict):
|
||||||
|
# Second chance: same page, compact instructions. Runs only
|
||||||
|
# when the full-schema pass returned nothing usable.
|
||||||
|
print(f"[Extract] Page {page['page_number']}: full extraction "
|
||||||
|
f"failed, retrying compact")
|
||||||
|
parsed = _wrap_bare_list(
|
||||||
|
self._call(instruction + _COMPACT_RETRY_SUFFIX, page),
|
||||||
|
page["page_number"],
|
||||||
)
|
)
|
||||||
if not isinstance(parsed, dict):
|
if not isinstance(parsed, dict):
|
||||||
raise ValueError("no structured extraction returned")
|
# Don't give up on the page - the ladder below can still
|
||||||
sheet = _normalize_sheet(parsed, page["page_number"])
|
# rescue it from the text layer.
|
||||||
|
parsed = {"sheet": {}, "objects": []}
|
||||||
|
|
||||||
|
sheet = _normalize_sheet(parsed, page["page_number"],
|
||||||
|
page_text=page_text)
|
||||||
|
cov = text_coverage(page_text or "", sheet["assertions"])
|
||||||
|
sheet["coverage"] = cov
|
||||||
|
|
||||||
|
# Rung 2: text-only structuring when coverage is below floor.
|
||||||
|
# MERGE, never replace: vision keeps every object it found
|
||||||
|
# (graphical_basis content exists only in the image); the text
|
||||||
|
# pass fills in the text content the vision pass missed.
|
||||||
|
if (page_text and config.EXTRACT_TEXT_RETRY_ENABLED
|
||||||
|
and cov["ratio"] < config.EXTRACT_COVERAGE_FLOOR):
|
||||||
|
print(f"[Extract] Page {page['page_number']}: coverage "
|
||||||
|
f"{cov['ratio']:.0%} < floor - text-only structuring pass")
|
||||||
|
parsed2 = _wrap_bare_list(
|
||||||
|
self._text_structuring_call(page, hint), page["page_number"])
|
||||||
|
if isinstance(parsed2, dict):
|
||||||
|
sheet2 = _normalize_sheet(parsed2, page["page_number"],
|
||||||
|
page_text=page_text)
|
||||||
|
before = len(sheet["assertions"])
|
||||||
|
sheet["assertions"] = merge_objects(sheet["assertions"],
|
||||||
|
sheet2["assertions"])
|
||||||
|
# Fill header gaps the vision pass left null
|
||||||
|
for key in ("sheet_number", "sheet_title", "discipline",
|
||||||
|
"level", "scale", "drawing_type"):
|
||||||
|
if not sheet.get(key) and sheet2.get(key):
|
||||||
|
sheet[key] = sheet2[key]
|
||||||
|
cov = text_coverage(page_text, sheet["assertions"])
|
||||||
|
sheet["coverage"] = cov
|
||||||
|
print(f"[Extract] Page {page['page_number']}: merged "
|
||||||
|
f"{len(sheet['assertions']) - before} text-structured "
|
||||||
|
f"object(s), coverage now {cov['ratio']:.0%}")
|
||||||
|
|
||||||
|
# Rung 3: deterministic fallback - dark sheets are impossible.
|
||||||
|
# Also merged (deduped) so stub notes never double up with
|
||||||
|
# objects the earlier rungs already captured.
|
||||||
|
if (page_text and config.EXTRACT_FALLBACK_ENABLED
|
||||||
|
and cov["ratio"] < config.EXTRACT_COVERAGE_FLOOR):
|
||||||
|
stubs = fallback_objects(page_text, page["page_number"],
|
||||||
|
config.EXTRACT_FALLBACK_MAX_OBJECTS)
|
||||||
|
stubs = _normalize_sheet({"sheet": {}, "objects": stubs},
|
||||||
|
page["page_number"],
|
||||||
|
page_text=page_text)["assertions"]
|
||||||
|
# _normalize_sheet only stamps its own "text_layer" rescue
|
||||||
|
# grounding; restore the explicit fallback provenance.
|
||||||
|
for stub in stubs:
|
||||||
|
stub["grounding"] = "text_layer_fallback"
|
||||||
|
before = len(sheet["assertions"])
|
||||||
|
sheet["assertions"] = merge_objects(sheet["assertions"], stubs)
|
||||||
|
print(f"[Extract] Page {page['page_number']}: fallback merged "
|
||||||
|
f"{len(sheet['assertions']) - before} text-layer stub(s)")
|
||||||
|
sheet["coverage"] = text_coverage(page_text,
|
||||||
|
sheet["assertions"])
|
||||||
|
|
||||||
|
# Identity recovery: never leave a text-bearing page sheet-less
|
||||||
|
if not sheet.get("sheet_number") and page_text:
|
||||||
|
recovered = recover_sheet_number(page_text)
|
||||||
|
if recovered:
|
||||||
|
sheet["sheet_number"] = recovered
|
||||||
|
sheet["discipline"] = (
|
||||||
|
discipline_from_sheet_number(recovered)
|
||||||
|
or sheet.get("discipline") or "Unknown")
|
||||||
|
print(f"[Extract] Page {page['page_number']}: sheet number "
|
||||||
|
f"recovered from text layer -> {recovered}")
|
||||||
|
|
||||||
return AgentResult(scope_id=scope.scope_id, artifacts=[sheet])
|
return AgentResult(scope_id=scope.scope_id, artifacts=[sheet])
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
return failure(scope, exc)
|
return failure(scope, exc)
|
||||||
|
|||||||
@@ -0,0 +1,111 @@
|
|||||||
|
"""Per-sheet Drawing Integrity QA agent.
|
||||||
|
|
||||||
|
Reads ONE sheet's own extracted objects + sheet image + deterministic text
|
||||||
|
layer and flags defects internal to that single sheet: dangling detail/
|
||||||
|
callout/keynote references, schedule-vs-plan/legend disagreements on the same
|
||||||
|
sheet, dimension strings that do not sum, missing title-block/scale/north
|
||||||
|
essentials, and duplicate/inconsistent tags. This is the drawing-focused pass
|
||||||
|
that complements the cross-sheet conflict critic; it never does code/ADA or
|
||||||
|
cross-sheet coordination.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from typing import Dict, List
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
|
||||||
|
from backend.agents.prompts import (
|
||||||
|
DRAWING_INTEGRITY_SYSTEM_PROMPT,
|
||||||
|
DRAWING_INTEGRITY_USER_PROMPT,
|
||||||
|
)
|
||||||
|
from backend.llm import call_json
|
||||||
|
from backend.pipeline._serialize import dumps
|
||||||
|
from backend.pipeline._stage import collect_list, validate_issue
|
||||||
|
|
||||||
|
|
||||||
|
def _sheet_meta(sheet: Dict) -> Dict:
|
||||||
|
"""Compact title-block-ish descriptor of the sheet (no raw assertions)."""
|
||||||
|
return {
|
||||||
|
"sheet_number": sheet.get("sheet_number"),
|
||||||
|
"sheet_title": sheet.get("sheet_title"),
|
||||||
|
"discipline": sheet.get("discipline"),
|
||||||
|
"drawing_type": sheet.get("drawing_type"),
|
||||||
|
"level": sheet.get("level"),
|
||||||
|
"scale": sheet.get("scale"),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def build_integrity_scopes(
|
||||||
|
sheets: List[Dict], page_to_b64: Dict, page_to_text: Dict
|
||||||
|
) -> List[AgentScope]:
|
||||||
|
"""One scope per sheet that carries enough objects to judge internal
|
||||||
|
consistency. Sheets below INTEGRITY_MIN_ASSERTIONS are skipped as too
|
||||||
|
sparse for a meaningful single-sheet back-check."""
|
||||||
|
scopes: List[AgentScope] = []
|
||||||
|
for sheet in sheets:
|
||||||
|
assertions = sheet.get("assertions") or []
|
||||||
|
if len(assertions) < config.INTEGRITY_MIN_ASSERTIONS:
|
||||||
|
continue
|
||||||
|
page_number = sheet.get("page_number")
|
||||||
|
scopes.append(AgentScope(
|
||||||
|
scope_id=f"integrity:{page_number}",
|
||||||
|
payload={
|
||||||
|
"sheet": sheet,
|
||||||
|
"page_number": page_number,
|
||||||
|
"image_b64": page_to_b64.get(page_number),
|
||||||
|
"text_layer": page_to_text.get(page_number) or "",
|
||||||
|
},
|
||||||
|
))
|
||||||
|
return scopes
|
||||||
|
|
||||||
|
|
||||||
|
class DrawingIntegrityAgent:
|
||||||
|
name = "drawing_integrity"
|
||||||
|
|
||||||
|
def __init__(self, usage: AgentUsage) -> None:
|
||||||
|
self.usage = usage
|
||||||
|
|
||||||
|
def run(self, scope: AgentScope) -> AgentResult:
|
||||||
|
try:
|
||||||
|
sheet = dict(scope.payload["sheet"])
|
||||||
|
assertions = (
|
||||||
|
sheet.get("assertions") or []
|
||||||
|
)[:config.AGENT_INTEGRITY_MAX_ASSERTIONS]
|
||||||
|
text_layer = (scope.payload.get("text_layer") or "")[
|
||||||
|
:config.TEXT_LAYER_MAX_CHARS
|
||||||
|
]
|
||||||
|
image_b64 = scope.payload.get("image_b64")
|
||||||
|
images = [image_b64] if image_b64 else []
|
||||||
|
images = images[:config.AGENT_INTEGRITY_MAX_IMAGES]
|
||||||
|
|
||||||
|
instruction = DRAWING_INTEGRITY_USER_PROMPT
|
||||||
|
for key, value in {
|
||||||
|
"sheet_meta": dumps(_sheet_meta(sheet)),
|
||||||
|
"assertions": dumps(assertions),
|
||||||
|
"text_layer": text_layer,
|
||||||
|
}.items():
|
||||||
|
instruction = instruction.replace("{" + key + "}", value)
|
||||||
|
|
||||||
|
parsed = call_json(
|
||||||
|
system_prompt=DRAWING_INTEGRITY_SYSTEM_PROMPT,
|
||||||
|
user_text=instruction,
|
||||||
|
images_b64=images,
|
||||||
|
max_tokens=config.INTEGRITY_MAX_TOKENS,
|
||||||
|
model=config.AGENT_INTEGRITY_MODEL,
|
||||||
|
usage_tracker=self.usage,
|
||||||
|
usage_stage="agent.drawing_integrity",
|
||||||
|
reasoning_effort=config.EXTRACT_REASONING_EFFORT or None,
|
||||||
|
reasoning_max_tokens=config.EXTRACT_REASONING_MAX_TOKENS or None,
|
||||||
|
)
|
||||||
|
findings = collect_list(
|
||||||
|
parsed, "issues",
|
||||||
|
lambda item: validate_issue(item, "drawing_integrity"),
|
||||||
|
)
|
||||||
|
sheet_number = sheet.get("sheet_number")
|
||||||
|
for finding in findings:
|
||||||
|
finding.update(agent=self.name, scope_id=scope.scope_id)
|
||||||
|
# Anchor the finding to this sheet if the model left it blank.
|
||||||
|
if not finding.get("sheets") and sheet_number:
|
||||||
|
finding["sheets"] = [sheet_number]
|
||||||
|
return AgentResult(scope_id=scope.scope_id, artifacts=findings)
|
||||||
|
except Exception as exc:
|
||||||
|
return failure(scope, exc)
|
||||||
@@ -30,9 +30,23 @@ def _family(assertion: Dict) -> str:
|
|||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _xref_keys(assertion: Dict) -> List[str]:
|
||||||
|
"""Cross-level join keys: detail references and member tags."""
|
||||||
|
location = assertion.get("location_key") or {}
|
||||||
|
keys = []
|
||||||
|
ref = re.sub(r"\s+", "", str(location.get("detail_reference") or "")).upper()
|
||||||
|
if ref:
|
||||||
|
keys.append(f"detail:{ref}")
|
||||||
|
tag = re.sub(r"\s+", "", str(location.get("tag") or "")).upper()
|
||||||
|
if re.match(r"^[A-Z]+\d", tag): # member marks: W12X26, HSS16X4X5/8, ...
|
||||||
|
keys.append(f"tag:{tag}")
|
||||||
|
return keys
|
||||||
|
|
||||||
|
|
||||||
def build_link_scopes(sheets: List[Dict]) -> List[AgentScope]:
|
def build_link_scopes(sheets: List[Dict]) -> List[AgentScope]:
|
||||||
"""Partition facts by level and object/tag family, then enforce a hard cap."""
|
"""Partition facts by level and object/tag family, then enforce a hard cap."""
|
||||||
buckets: Dict[Tuple[str, str], List[Dict]] = defaultdict(list)
|
buckets: Dict[Tuple[str, str], List[Dict]] = defaultdict(list)
|
||||||
|
xref: Dict[str, List[Dict]] = defaultdict(list)
|
||||||
for sheet in sheets:
|
for sheet in sheets:
|
||||||
for assertion in sheet.get("assertions", []):
|
for assertion in sheet.get("assertions", []):
|
||||||
enriched = {
|
enriched = {
|
||||||
@@ -44,6 +58,8 @@ def build_link_scopes(sheets: List[Dict]) -> List[AgentScope]:
|
|||||||
level = str((assertion.get("location_key") or {}).get("level")
|
level = str((assertion.get("location_key") or {}).get("level")
|
||||||
or sheet.get("level") or "unknown").lower()
|
or sheet.get("level") or "unknown").lower()
|
||||||
buckets[(level, _family(assertion))].append(enriched)
|
buckets[(level, _family(assertion))].append(enriched)
|
||||||
|
for key in _xref_keys(assertion):
|
||||||
|
xref[key].append(enriched)
|
||||||
|
|
||||||
scopes: List[AgentScope] = []
|
scopes: List[AgentScope] = []
|
||||||
cap = max(2, config.AGENT_LINK_MAX_ASSERTIONS)
|
cap = max(2, config.AGENT_LINK_MAX_ASSERTIONS)
|
||||||
@@ -56,6 +72,18 @@ def build_link_scopes(sheets: List[Dict]) -> List[AgentScope]:
|
|||||||
scope_id=f"{level}:{family}:{offset // cap + 1}",
|
scope_id=f"{level}:{family}:{offset // cap + 1}",
|
||||||
payload={"assertions": chunk, "level": level, "family": family},
|
payload={"assertions": chunk, "level": level, "family": family},
|
||||||
))
|
))
|
||||||
|
for key, assertions in sorted(xref.items()):
|
||||||
|
sheets_present = {a.get("sheet_number") for a in assertions}
|
||||||
|
if len(assertions) < 2 or len(sheets_present) < 2:
|
||||||
|
continue
|
||||||
|
chunk = assertions[:cap]
|
||||||
|
if len({a.get("sheet_number") for a in chunk}) < 2:
|
||||||
|
continue # cap landed on a single sheet — xref adds nothing
|
||||||
|
scopes.append(AgentScope(
|
||||||
|
scope_id=f"xref:{key}",
|
||||||
|
payload={"assertions": chunk,
|
||||||
|
"level": "xref", "family": key},
|
||||||
|
))
|
||||||
return scopes
|
return scopes
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -7,7 +7,8 @@ import threading
|
|||||||
from typing import Any, Dict, Iterable, Optional
|
from typing import Any, Dict, Iterable, Optional
|
||||||
|
|
||||||
|
|
||||||
_COLLECTION_KEYS = {"sheets", "clusters", "findings", "decisions", "rfis"}
|
_COLLECTION_KEYS = {"sheets", "clusters", "findings", "decisions", "rfis",
|
||||||
|
"suppressed"}
|
||||||
_MAPPING_KEYS = {"sheet_index", "jurisdiction", "object_graph"}
|
_MAPPING_KEYS = {"sheet_index", "jurisdiction", "object_graph"}
|
||||||
_MEMORY_KEYS = _COLLECTION_KEYS | _MAPPING_KEYS
|
_MEMORY_KEYS = _COLLECTION_KEYS | _MAPPING_KEYS
|
||||||
|
|
||||||
|
|||||||
@@ -12,6 +12,42 @@ Sheet index: {sheet_index}
|
|||||||
Aggregate sheet summaries: {sheet_summaries}
|
Aggregate sheet summaries: {sheet_summaries}
|
||||||
Cluster summary: {cluster_summary}"""
|
Cluster summary: {cluster_summary}"""
|
||||||
|
|
||||||
|
DRAWING_INTEGRITY_SYSTEM_PROMPT = """You are a Senior Architect performing a single-sheet QAQC back-check of ONE construction drawing before the set is issued for bid, permit, or construction.
|
||||||
|
You are given the extracted construction objects for this one sheet, plus the sheet image and its deterministic PDF text layer.
|
||||||
|
Your job is to find problems INTERNAL TO THIS SHEET - defects a human checker would red-line on this drawing by itself, without needing any other sheet.
|
||||||
|
You are NOT performing code review. You are NOT checking ADA/accessibility. You are NOT doing cross-sheet coordination (a separate reviewer handles conflicts between sheets). You are NOT estimating cost. You are NOT redesigning anything.
|
||||||
|
What IS a drawing-integrity issue on this sheet:
|
||||||
|
- Dangling reference: a detail callout, section marker, elevation marker, keynote, or sheet reference that points to a target that does not exist on this sheet AND is not resolved by an explicit off-sheet reference (e.g. "SIM 5/A501" when this is A501 and it has no detail 5; a keynote number called out in the plan but absent from the keynote legend on the same sheet).
|
||||||
|
- On-sheet contradiction: the plan disagrees with a schedule or legend printed on the SAME sheet; two notes on the sheet contradict each other; a tag in the plan is not in the sheet's own schedule/legend (or vice versa); the title block discipline/level disagrees with the drawing content.
|
||||||
|
- Dimension sanity: a dimension string whose segments do not sum to the stated overall; an overall dimension that contradicts a repeated/typical dimension on the same sheet; obviously impossible or missing critical dimensions on a dimensioned plan.
|
||||||
|
- Missing sheet essentials: no scale, no north arrow on a plan that needs one, missing sheet number/title in the title block, a schedule with header columns but no rows, a legend referenced but not present.
|
||||||
|
- Label/tag hygiene: duplicate tags that should be unique on this sheet (two different doors both tagged 101A), a room shown with no room number/name where the sheet otherwise numbers rooms, inconsistent tag formatting that breaks a reference.
|
||||||
|
What is NOT a drawing-integrity issue:
|
||||||
|
- Anything requiring another sheet to judge (that is cross-sheet coordination, handled elsewhere).
|
||||||
|
- A code, ADA, or accessibility requirement.
|
||||||
|
- A design preference or cost concern.
|
||||||
|
- A value simply not repeated where repetition is optional.
|
||||||
|
- Anything you cannot support with text or a clear visual from THIS sheet.
|
||||||
|
Be conservative and evidence-bound:
|
||||||
|
- Only flag defects you can point to with verbatim source_text from this sheet or a clear description of what the image shows.
|
||||||
|
- Trust the TEXT LAYER for alphanumeric content (numbers, tags, note text, dimensions); use the image for geometry, symbols, linework, and whether a referenced target actually appears.
|
||||||
|
- When a value is marked DISPUTED (possible extraction misread), verify against the image before relying on it.
|
||||||
|
- If the sheet is internally clean, return an empty issues array.
|
||||||
|
Severity (use exactly one of critical, high, medium, low):
|
||||||
|
- high = a defect that would cause rework, a wrong build, or a stop at permit/bid if issued as-is (missing critical dimension, dangling reference to a nonexistent detail that drives construction).
|
||||||
|
- medium = a real drawing defect needing correction before issue.
|
||||||
|
- low = minor cleanup/clarification.
|
||||||
|
Use plain ASCII only. Respond only with valid JSON."""
|
||||||
|
|
||||||
|
DRAWING_INTEGRITY_USER_PROMPT = """Back-check this single sheet for internal drawing-integrity defects.
|
||||||
|
Respond ONLY with a valid JSON object - no markdown fences, no explanation:
|
||||||
|
{"issues":[{"issue_id":"string","source_stage":"drawing_integrity","category":"dangling_reference | on_sheet_contradiction | dimension_error | missing_sheet_essential | tag_or_label_error | other","severity":"critical | high | medium | low","confidence":"high | medium | low","location":"where on the sheet, e.g. 'Room 124 / detail callout 5' or 'door schedule'","disciplines":["string"],"sheets":["this sheet number"],"description":"senior architect explanation of the defect and why it matters","evidence":[{"discipline":"string","sheet":"string","source_text":"verbatim text from this sheet","asserted_value":"string"}],"recommended_resolution":"coordinate drawing | correct dimension | add missing detail | issue RFI | verify with architect | verify with engineer","code_reference":null}]}
|
||||||
|
If the sheet is internally clean, return {"issues":[]}.
|
||||||
|
Sheet: {sheet_meta}
|
||||||
|
Extracted objects on this sheet: {assertions}
|
||||||
|
TEXT LAYER (deterministic page text - authoritative for alphanumeric content):
|
||||||
|
{text_layer}"""
|
||||||
|
|
||||||
BRAIN_SYSTEM_PROMPT = """You are the central decision layer for a construction drawing
|
BRAIN_SYSTEM_PROMPT = """You are the central decision layer for a construction drawing
|
||||||
review. Merge duplicate specialist findings, reject vague or unsupported findings,
|
review. Merge duplicate specialist findings, reject vague or unsupported findings,
|
||||||
preserve verbatim evidence, and prioritize the kept issues. Do not create new issues.
|
preserve verbatim evidence, and prioritize the kept issues. Do not create new issues.
|
||||||
@@ -19,7 +55,23 @@ Conflicts need drawing evidence; completeness findings may instead cite an expli
|
|||||||
missing item from the sheet index. Return only valid JSON."""
|
missing item from the sheet index. Return only valid JSON."""
|
||||||
|
|
||||||
BRAIN_USER_PROMPT = """Judge and consolidate these scoped specialist findings.
|
BRAIN_USER_PROMPT = """Judge and consolidate these scoped specialist findings.
|
||||||
Return {"issues":[{"issue_id":"string","source_stage":"conflict | qaqc | code | constructability","category":"string","severity":"critical | high | medium | low","confidence":"high | medium | low","location":"string","disciplines":["string"],"sheets":["string"],"description":"string","evidence":[{"discipline":"string","sheet":"string","source_text":"string","asserted_value":"string"}],"recommended_resolution":"string","code_reference":"string or null","risk_score":1,"recommended_priority":"immediate | before_bid | before_permit | before_construction | track_only"}],"decisions":[{"finding_refs":["string"],"action":"kept | merged | dropped","reason":"string","kept_issue_id":"string or null"}]}.
|
Return {"issues":[{"issue_id":"string","source_stage":"conflict | drawing_integrity | qaqc | code | constructability","category":"string","severity":"critical | high | medium | low","confidence":"high | medium | low","location":"string","disciplines":["string"],"sheets":["string"],"description":"string","evidence":[{"discipline":"string","sheet":"string","source_text":"string","asserted_value":"string"}],"recommended_resolution":"string","code_reference":"string or null","risk_score":1,"recommended_priority":"immediate | before_bid | before_permit | before_construction | track_only"}],"decisions":[{"finding_refs":["string"],"action":"kept | merged | dropped","reason":"string","kept_issue_id":"string or null"}]}.
|
||||||
Sheet index: {sheet_index}
|
Sheet index: {sheet_index}
|
||||||
Jurisdiction summary: {jurisdiction}
|
Jurisdiction summary: {jurisdiction}
|
||||||
Specialist findings: {findings}"""
|
Specialist findings: {findings}"""
|
||||||
|
|
||||||
|
BRAIN_CLARIFY_SYSTEM_PROMPT = """You are the central decision layer for a construction drawing review, deciding which of your kept findings you are NOT yet confident enough to publish.
|
||||||
|
You have already merged and prioritized the findings. Now, for the borderline ones, you may request ONE targeted clarification each before the report is finalized.
|
||||||
|
Request a clarification only when a finding's evidence is thin, ambiguous, possibly a misread of the drawing, or internally inconsistent - the kind of finding a senior reviewer would double-check against the sheet before signing off. Do NOT request clarification for findings that are already clearly supported by verbatim evidence, and do NOT re-request a finding that already carries a verification result.
|
||||||
|
The only request type available right now is:
|
||||||
|
- verify_evidence: re-check this finding's quoted evidence against the actual sheet images and deterministic text layer (catches wave-1 vision misreads such as "(2)" vs "(5)" and dangling references that do not actually appear on the sheet).
|
||||||
|
Be selective. Requesting everything wastes the budget and slows the review; request only the findings where a second look would actually change your decision.
|
||||||
|
Use plain ASCII only. Respond only with valid JSON."""
|
||||||
|
|
||||||
|
BRAIN_CLARIFY_USER_PROMPT = """Decide which of these kept findings you want to double-check before publishing.
|
||||||
|
You may request at most {max_requests} clarifications. Choose the findings where a second look at the sheet would most likely change your keep/drop/severity decision.
|
||||||
|
Respond ONLY with a valid JSON object - no markdown fences, no explanation:
|
||||||
|
{"requests":[{"issue_id":"the issue_id of the finding to check","request_type":"verify_evidence","reason":"one sentence: why this finding is uncertain"}]}
|
||||||
|
If every finding is already well supported, return {"requests":[]}.
|
||||||
|
Findings (each shows issue_id, severity, confidence, evidence, and whether it already has a verification result):
|
||||||
|
{findings}"""
|
||||||
+278
-2
@@ -1,5 +1,6 @@
|
|||||||
"""Public entry point for the scoped Agent-mode pipeline."""
|
"""Public entry point for the scoped Agent-mode pipeline."""
|
||||||
|
|
||||||
|
import base64
|
||||||
import json
|
import json
|
||||||
import os
|
import os
|
||||||
from typing import Callable, Dict, Optional
|
from typing import Callable, Dict, Optional
|
||||||
@@ -11,20 +12,32 @@ from backend.agents.code_agent import CodeAgent, build_code_scopes
|
|||||||
from backend.agents.completeness import CompletenessAgent, build_sheet_summaries
|
from backend.agents.completeness import CompletenessAgent, build_sheet_summaries
|
||||||
from backend.agents.conflict_critic import ConflictCriticAgent
|
from backend.agents.conflict_critic import ConflictCriticAgent
|
||||||
from backend.agents.construct_agent import ConstructabilityAgent, build_construct_scopes
|
from backend.agents.construct_agent import ConstructabilityAgent, build_construct_scopes
|
||||||
|
from backend.agents.disputes import annotate_clusters
|
||||||
from backend.agents.extractors import (
|
from backend.agents.extractors import (
|
||||||
JurisdictionAgent,
|
JurisdictionAgent,
|
||||||
SheetExtractorAgent,
|
SheetExtractorAgent,
|
||||||
SheetIndexAgent,
|
SheetIndexAgent,
|
||||||
)
|
)
|
||||||
|
from backend.agents.integrity_agent import (
|
||||||
|
DrawingIntegrityAgent, build_integrity_scopes,
|
||||||
|
)
|
||||||
from backend.agents.linker import LinkerAgent, build_link_scopes, build_object_graph
|
from backend.agents.linker import LinkerAgent, build_link_scopes, build_object_graph
|
||||||
from backend.agents.memory import ProjectMemory
|
from backend.agents.memory import ProjectMemory
|
||||||
from backend.agents.orchestrator import Orchestrator
|
from backend.agents.orchestrator import Orchestrator
|
||||||
from backend.agents.rfi_writer import RFIWriterAgent
|
from backend.agents.rfi_writer import RFIWriterAgent
|
||||||
|
from backend.agents.verifier import (
|
||||||
|
EvidenceVerifierAgent, apply_verdicts, select_findings,
|
||||||
|
)
|
||||||
|
from backend.llm import reset_cost
|
||||||
from backend.pipeline.pdf_processor import convert_pdf_to_images
|
from backend.pipeline.pdf_processor import convert_pdf_to_images
|
||||||
from backend.pipeline.report import build_report, to_markdown
|
from backend.pipeline.report import build_report, to_markdown
|
||||||
from backend.pipeline.sheet_index import derive_project_meta_from_cover
|
from backend.pipeline.sheet_index import derive_project_meta_from_cover
|
||||||
from backend.review.gate import build_review_queue
|
from backend.review.gate import build_review_queue
|
||||||
from backend.review.store import ReviewStore
|
from backend.review.store import ReviewStore
|
||||||
|
from backend.sheet_reconcile import declared_sheet_list, reconcile_sheets
|
||||||
|
from backend.text_layer import (
|
||||||
|
attach_text_layers, coverage_gaps, find_evidence_bbox, render_crop,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def run_agent_pipeline(
|
def run_agent_pipeline(
|
||||||
@@ -39,6 +52,10 @@ def run_agent_pipeline(
|
|||||||
if not os.path.isfile(pdf_path):
|
if not os.path.isfile(pdf_path):
|
||||||
raise FileNotFoundError(pdf_path)
|
raise FileNotFoundError(pdf_path)
|
||||||
|
|
||||||
|
# Keep the llm module's counters job-local (matches Classic): the job
|
||||||
|
# log's failure-path cost estimate in jobs.py reads llm.get_cost().
|
||||||
|
reset_cost()
|
||||||
|
|
||||||
agent_dir = os.path.join(out_dir, "agent") if out_dir else None
|
agent_dir = os.path.join(out_dir, "agent") if out_dir else None
|
||||||
memory = ProjectMemory(artifact_dir=agent_dir)
|
memory = ProjectMemory(artifact_dir=agent_dir)
|
||||||
orchestrator = Orchestrator(memory=memory, on_stage=on_stage)
|
orchestrator = Orchestrator(memory=memory, on_stage=on_stage)
|
||||||
@@ -48,6 +65,9 @@ def run_agent_pipeline(
|
|||||||
orchestrator.stage("Agent ingest: PDF -> images")
|
orchestrator.stage("Agent ingest: PDF -> images")
|
||||||
pages = convert_pdf_to_images(pdf_path)
|
pages = convert_pdf_to_images(pdf_path)
|
||||||
page_to_b64 = {page["page_number"]: page["base64"] for page in pages}
|
page_to_b64 = {page["page_number"]: page["base64"] for page in pages}
|
||||||
|
text_dir = os.path.join(agent_dir, "text") if agent_dir else None
|
||||||
|
page_words = attach_text_layers(pdf_path, pages, text_dir=text_dir)
|
||||||
|
page_to_text = {page["page_number"]: page.get("text_layer") for page in pages}
|
||||||
|
|
||||||
orchestrator.stage("Agent wave 1: extract sheets")
|
orchestrator.stage("Agent wave 1: extract sheets")
|
||||||
extract_scopes = [
|
extract_scopes = [
|
||||||
@@ -69,6 +89,27 @@ def run_agent_pipeline(
|
|||||||
memory.replace("sheets", sheets)
|
memory.replace("sheets", sheets)
|
||||||
memory.dump("01-extract.json")
|
memory.dump("01-extract.json")
|
||||||
|
|
||||||
|
# Deterministic reconciliation: the cover sheet's own sheet index
|
||||||
|
# declares what the set should contain; compare against what wave 1
|
||||||
|
# identified (catches missed sheets AND phantom/misread sheet numbers).
|
||||||
|
sheet_recon = reconcile_sheets(sheets, declared_sheet_list(page_to_text))
|
||||||
|
if sheet_recon["declared_total"]:
|
||||||
|
print(f"[SheetIndex] cover declares {sheet_recon['declared_total']} "
|
||||||
|
f"sheets; {sheet_recon['found_total']} identified in set")
|
||||||
|
if sheet_recon["declared_not_in_set"]:
|
||||||
|
print(f"[SheetIndex] declared but not in set: "
|
||||||
|
f"{', '.join(sheet_recon['declared_not_in_set'][:20])}")
|
||||||
|
if sheet_recon["in_set_not_declared"]:
|
||||||
|
print(f"[SheetIndex] in set but not declared: "
|
||||||
|
f"{', '.join(sheet_recon['in_set_not_declared'][:20])}")
|
||||||
|
# Coverage signal: text layer present but extraction failed/empty reuses
|
||||||
|
# the failed-scopes gap-finding path (finding built below wave 6).
|
||||||
|
for gap_page in coverage_gaps(pages, sheets):
|
||||||
|
orchestrator.stats.failed_scopes.append(
|
||||||
|
f"sheet_extractor:sheet:{gap_page}: extraction gap "
|
||||||
|
f"(text layer present, no objects extracted)"
|
||||||
|
)
|
||||||
|
|
||||||
cover_meta = derive_project_meta_from_cover(
|
cover_meta = derive_project_meta_from_cover(
|
||||||
sheets, source_name or os.path.basename(pdf_path)
|
sheets, source_name or os.path.basename(pdf_path)
|
||||||
)
|
)
|
||||||
@@ -108,6 +149,9 @@ def run_agent_pipeline(
|
|||||||
for artifact in result.artifacts
|
for artifact in result.artifacts
|
||||||
][:config.CLUSTER_MAX]
|
][:config.CLUSTER_MAX]
|
||||||
object_graph = build_object_graph(clusters)
|
object_graph = build_object_graph(clusters)
|
||||||
|
disputed_count = annotate_clusters(clusters)
|
||||||
|
if disputed_count:
|
||||||
|
orchestrator.stage(f"[Link] {disputed_count} clusters carry disputed extracted values")
|
||||||
memory.replace("clusters", clusters)
|
memory.replace("clusters", clusters)
|
||||||
memory.replace("object_graph", object_graph)
|
memory.replace("object_graph", object_graph)
|
||||||
memory.dump("03-link.json")
|
memory.dump("03-link.json")
|
||||||
@@ -133,11 +177,23 @@ def run_agent_pipeline(
|
|||||||
memory.extend("findings", conflict_findings)
|
memory.extend("findings", conflict_findings)
|
||||||
|
|
||||||
orchestrator.stage("Agent wave 5: scoped specialists")
|
orchestrator.stage("Agent wave 5: scoped specialists")
|
||||||
|
if config.ENABLE_CODE_REVIEW:
|
||||||
code_results = orchestrator.run_scopes(
|
code_results = orchestrator.run_scopes(
|
||||||
CodeAgent(usage),
|
CodeAgent(usage),
|
||||||
build_code_scopes(sheets, jurisdiction, sheet_index),
|
build_code_scopes(sheets, jurisdiction, sheet_index),
|
||||||
config.AGENT_SPECIALIST_CONCURRENCY,
|
config.AGENT_SPECIALIST_CONCURRENCY,
|
||||||
)
|
)
|
||||||
|
else:
|
||||||
|
orchestrator.stage("[wave 5] code/ADA review disabled (ENABLE_CODE_REVIEW=0)")
|
||||||
|
code_results = []
|
||||||
|
if config.ENABLE_DRAWING_INTEGRITY:
|
||||||
|
integrity_results = orchestrator.run_scopes(
|
||||||
|
DrawingIntegrityAgent(usage),
|
||||||
|
build_integrity_scopes(sheets, page_to_b64, page_to_text),
|
||||||
|
config.AGENT_INTEGRITY_CONCURRENCY,
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
integrity_results = []
|
||||||
construct_results = orchestrator.run_scopes(
|
construct_results = orchestrator.run_scopes(
|
||||||
ConstructabilityAgent(usage),
|
ConstructabilityAgent(usage),
|
||||||
build_construct_scopes(clusters, conflict_findings),
|
build_construct_scopes(clusters, conflict_findings),
|
||||||
@@ -156,9 +212,34 @@ def run_agent_pipeline(
|
|||||||
)
|
)
|
||||||
specialist_findings = [
|
specialist_findings = [
|
||||||
artifact
|
artifact
|
||||||
for result in code_results + construct_results + completeness_results
|
for result in (code_results + integrity_results
|
||||||
|
+ construct_results + completeness_results)
|
||||||
for artifact in result.artifacts
|
for artifact in result.artifacts
|
||||||
]
|
]
|
||||||
|
|
||||||
|
orchestrator.stage("Agent wave 5b: evidence verification")
|
||||||
|
sheet_to_page = {str(s.get("sheet_number")): s.get("page_number")
|
||||||
|
for s in sheets}
|
||||||
|
verify_targets = select_findings(
|
||||||
|
specialist_findings, clusters,
|
||||||
|
max_checks=config.AGENT_VERIFY_MAX_CHECKS,
|
||||||
|
severities=config.AGENT_VERIFY_SEVERITIES,
|
||||||
|
)
|
||||||
|
target_indexes = {id(f): i for i, f in enumerate(specialist_findings)}
|
||||||
|
verify_scopes = _build_verify_scopes(
|
||||||
|
verify_targets,
|
||||||
|
index_for=lambda f: target_indexes[id(f)],
|
||||||
|
sheet_to_page=sheet_to_page, page_to_b64=page_to_b64,
|
||||||
|
page_to_text=page_to_text, page_words=page_words, pdf_path=pdf_path,
|
||||||
|
)
|
||||||
|
verify_results = orchestrator.run_scopes(
|
||||||
|
EvidenceVerifierAgent(usage), verify_scopes, config.AGENT_VERIFY_CONCURRENCY)
|
||||||
|
suppressed = apply_verdicts(specialist_findings, verify_results)
|
||||||
|
if suppressed:
|
||||||
|
suppressed_ids = {id(f) for f in suppressed}
|
||||||
|
specialist_findings = [f for f in specialist_findings if id(f) not in suppressed_ids]
|
||||||
|
memory.replace("suppressed", suppressed)
|
||||||
|
|
||||||
memory.extend("findings", specialist_findings)
|
memory.extend("findings", specialist_findings)
|
||||||
gap_findings = [
|
gap_findings = [
|
||||||
{
|
{
|
||||||
@@ -195,6 +276,18 @@ def run_agent_pipeline(
|
|||||||
1 for decision in decisions if decision.get("action") == "merged"
|
1 for decision in decisions if decision.get("action") == "merged"
|
||||||
)
|
)
|
||||||
|
|
||||||
|
# Wave 6.5 — Brain-directed clarification (bounded hub-and-spoke). The Brain
|
||||||
|
# names kept findings it is unsure about; verify_evidence requests route
|
||||||
|
# back through the wave-5b verifier (fresh images + text-layer oracle).
|
||||||
|
# Refuted findings are demoted, dropped from `prioritized`, and moved to
|
||||||
|
# memory["suppressed"]. One planning call, one bounded verify wave, no loop.
|
||||||
|
if config.ENABLE_BRAIN_CLARIFY and prioritized:
|
||||||
|
prioritized = _brain_clarification_pass(
|
||||||
|
orchestrator, usage, memory, prioritized,
|
||||||
|
sheet_to_page=sheet_to_page, page_to_b64=page_to_b64,
|
||||||
|
page_to_text=page_to_text, page_words=page_words, pdf_path=pdf_path,
|
||||||
|
)
|
||||||
|
|
||||||
if require_review:
|
if require_review:
|
||||||
orchestrator.stage("Agent review gate: build human-review queue")
|
orchestrator.stage("Agent review gate: build human-review queue")
|
||||||
memory_snapshot = memory.snapshot()
|
memory_snapshot = memory.snapshot()
|
||||||
@@ -213,10 +306,11 @@ def run_agent_pipeline(
|
|||||||
"project_input": merged_input,
|
"project_input": merged_input,
|
||||||
"jurisdiction": jurisdiction,
|
"jurisdiction": jurisdiction,
|
||||||
"sheet_index": sheet_index,
|
"sheet_index": sheet_index,
|
||||||
|
"sheet_reconciliation": sheet_recon,
|
||||||
"project_intelligence": object_graph,
|
"project_intelligence": object_graph,
|
||||||
"validated_issues": prioritized,
|
"validated_issues": prioritized,
|
||||||
"rfis": [],
|
"rfis": [],
|
||||||
"suppressed_issues": [],
|
"suppressed_issues": memory.snapshot().get("suppressed") or [],
|
||||||
})
|
})
|
||||||
progress = store.progress(queue)
|
progress = store.progress(queue)
|
||||||
# Same usage/stats summary block as the wave-7 path (rfis: 0 — they
|
# Same usage/stats summary block as the wave-7 path (rfis: 0 — they
|
||||||
@@ -239,6 +333,10 @@ def run_agent_pipeline(
|
|||||||
1 for item in specialist_findings
|
1 for item in specialist_findings
|
||||||
if item.get("source_stage") == "code"
|
if item.get("source_stage") == "code"
|
||||||
),
|
),
|
||||||
|
"drawing_integrity": sum(
|
||||||
|
1 for item in specialist_findings
|
||||||
|
if item.get("source_stage") == "drawing_integrity"
|
||||||
|
),
|
||||||
"constructability": sum(
|
"constructability": sum(
|
||||||
1 for item in specialist_findings
|
1 for item in specialist_findings
|
||||||
if item.get("source_stage") == "constructability"
|
if item.get("source_stage") == "constructability"
|
||||||
@@ -289,9 +387,11 @@ def run_agent_pipeline(
|
|||||||
"project_input": merged_input,
|
"project_input": merged_input,
|
||||||
"jurisdiction": jurisdiction,
|
"jurisdiction": jurisdiction,
|
||||||
"sheet_index": sheet_index,
|
"sheet_index": sheet_index,
|
||||||
|
"sheet_reconciliation": sheet_recon,
|
||||||
"project_intelligence": object_graph,
|
"project_intelligence": object_graph,
|
||||||
"validated_issues": prioritized,
|
"validated_issues": prioritized,
|
||||||
"rfis": rfis,
|
"rfis": rfis,
|
||||||
|
"suppressed_issues": memory.snapshot().get("suppressed") or [],
|
||||||
})
|
})
|
||||||
cost = usage.snapshot()
|
cost = usage.snapshot()
|
||||||
orchestrator.stats.calls = cost["calls"]
|
orchestrator.stats.calls = cost["calls"]
|
||||||
@@ -310,6 +410,10 @@ def run_agent_pipeline(
|
|||||||
1 for item in specialist_findings
|
1 for item in specialist_findings
|
||||||
if item.get("source_stage") == "code"
|
if item.get("source_stage") == "code"
|
||||||
),
|
),
|
||||||
|
"drawing_integrity": sum(
|
||||||
|
1 for item in specialist_findings
|
||||||
|
if item.get("source_stage") == "drawing_integrity"
|
||||||
|
),
|
||||||
"constructability": sum(
|
"constructability": sum(
|
||||||
1 for item in specialist_findings
|
1 for item in specialist_findings
|
||||||
if item.get("source_stage") == "constructability"
|
if item.get("source_stage") == "constructability"
|
||||||
@@ -346,6 +450,178 @@ def _dump(out_dir: str, name: str, value) -> None:
|
|||||||
json.dump(value, f, indent=2)
|
json.dump(value, f, indent=2)
|
||||||
|
|
||||||
|
|
||||||
|
def _brain_clarification_pass(
|
||||||
|
orchestrator,
|
||||||
|
usage,
|
||||||
|
memory,
|
||||||
|
prioritized,
|
||||||
|
sheet_to_page,
|
||||||
|
page_to_b64,
|
||||||
|
page_to_text,
|
||||||
|
page_words,
|
||||||
|
pdf_path,
|
||||||
|
):
|
||||||
|
"""Wave 6.5: let the Brain request targeted clarifications, execute the
|
||||||
|
verify_evidence ones through the wave-5b verifier, and prune refuted
|
||||||
|
findings out of `prioritized` into memory["suppressed"].
|
||||||
|
|
||||||
|
Bounded and non-looping: one Brain planning call, at most
|
||||||
|
BRAIN_CLARIFY_MAX_REQUESTS verifications, a single pass. Returns the
|
||||||
|
(possibly shortened) prioritized list. Any request type other than
|
||||||
|
verify_evidence is logged as planned-but-not-executed and left untouched.
|
||||||
|
"""
|
||||||
|
requests = BrainAgent(usage).plan_clarifications(prioritized)
|
||||||
|
if not requests:
|
||||||
|
return prioritized
|
||||||
|
by_id = {f.get("issue_id"): f for f in prioritized}
|
||||||
|
verify_findings = []
|
||||||
|
unsupported = 0
|
||||||
|
for req in requests:
|
||||||
|
if req.get("request_type") != "verify_evidence":
|
||||||
|
unsupported += 1
|
||||||
|
continue
|
||||||
|
finding = by_id.get(req.get("issue_id"))
|
||||||
|
if finding is not None and finding not in verify_findings:
|
||||||
|
verify_findings.append(finding)
|
||||||
|
orchestrator.stage(
|
||||||
|
f"Agent wave 6.5: Brain-directed clarification "
|
||||||
|
f"({len(verify_findings)} verify, {unsupported} other)"
|
||||||
|
)
|
||||||
|
if unsupported:
|
||||||
|
for req in requests:
|
||||||
|
if req.get("request_type") != "verify_evidence":
|
||||||
|
orchestrator.stats.failed_scopes.append(
|
||||||
|
f"brain_clarify:{req.get('issue_id')}: "
|
||||||
|
f"request_type '{req.get('request_type')}' planned, "
|
||||||
|
f"not executed (v1 supports verify_evidence only)"
|
||||||
|
)
|
||||||
|
if not verify_findings:
|
||||||
|
return prioritized
|
||||||
|
index_of = {id(f): i for i, f in enumerate(prioritized)}
|
||||||
|
verify_scopes = _build_verify_scopes(
|
||||||
|
verify_findings,
|
||||||
|
index_for=lambda f: index_of[id(f)],
|
||||||
|
sheet_to_page=sheet_to_page, page_to_b64=page_to_b64,
|
||||||
|
page_to_text=page_to_text, page_words=page_words, pdf_path=pdf_path,
|
||||||
|
)
|
||||||
|
if not verify_scopes:
|
||||||
|
return prioritized
|
||||||
|
verify_results = orchestrator.run_scopes(
|
||||||
|
EvidenceVerifierAgent(usage), verify_scopes,
|
||||||
|
config.AGENT_VERIFY_CONCURRENCY,
|
||||||
|
)
|
||||||
|
suppressed = apply_verdicts(prioritized, verify_results)
|
||||||
|
if suppressed:
|
||||||
|
suppressed_ids = {id(f) for f in suppressed}
|
||||||
|
prioritized = [f for f in prioritized if id(f) not in suppressed_ids]
|
||||||
|
existing = memory.snapshot().get("suppressed") or []
|
||||||
|
memory.replace("suppressed", existing + suppressed)
|
||||||
|
memory.extend("decisions", [
|
||||||
|
{
|
||||||
|
"finding_refs": [f.get("issue_id")],
|
||||||
|
"action": "dropped",
|
||||||
|
"reason": "Brain-directed clarification: evidence refuted on re-check",
|
||||||
|
"kept_issue_id": None,
|
||||||
|
}
|
||||||
|
for f in suppressed
|
||||||
|
])
|
||||||
|
return prioritized
|
||||||
|
|
||||||
|
|
||||||
|
def _build_verify_scopes(
|
||||||
|
targets,
|
||||||
|
index_for,
|
||||||
|
sheet_to_page,
|
||||||
|
page_to_b64,
|
||||||
|
page_to_text,
|
||||||
|
page_words,
|
||||||
|
pdf_path,
|
||||||
|
):
|
||||||
|
"""Build EvidenceVerifierAgent scopes for a set of findings.
|
||||||
|
|
||||||
|
Shared by wave 5b (severity-gated) and wave 6.5 (Brain-directed): each
|
||||||
|
finding's cited sheets are mapped to page images (hi-DPI evidence crops
|
||||||
|
when enabled, else full pages) plus a capped text-layer oracle. Findings
|
||||||
|
whose sheets resolve to NO loadable image are skipped (I2 guard) — never
|
||||||
|
judge evidence against images we could not load. index_for(finding) yields
|
||||||
|
the finding_index the verifier echoes back for apply_verdicts alignment.
|
||||||
|
"""
|
||||||
|
scopes = []
|
||||||
|
for finding in targets:
|
||||||
|
cited_pages = [
|
||||||
|
sheet_to_page[str(name)]
|
||||||
|
for name in (finding.get("sheets") or [])
|
||||||
|
if sheet_to_page.get(str(name)) in page_to_b64
|
||||||
|
]
|
||||||
|
images = [
|
||||||
|
page_to_b64[p]
|
||||||
|
for p in cited_pages[:config.AGENT_CONFLICT_MAX_IMAGES]
|
||||||
|
]
|
||||||
|
if not images:
|
||||||
|
continue # never judge evidence against images we could not load
|
||||||
|
# Text oracle: concatenated text layer of the cited sheets, capped.
|
||||||
|
excerpt = "\n\n".join(
|
||||||
|
f"--- Page {p} ---\n{page_to_text[p]}"
|
||||||
|
for p in cited_pages
|
||||||
|
if page_to_text.get(p)
|
||||||
|
)[:config.VERIFY_TEXT_MAX_CHARS]
|
||||||
|
if config.VERIFY_HI_DPI_CROPS:
|
||||||
|
images = _evidence_crops(finding, cited_pages, sheet_to_page,
|
||||||
|
page_words, page_to_b64, pdf_path,
|
||||||
|
fallback=images)
|
||||||
|
finding_index = index_for(finding)
|
||||||
|
scopes.append(AgentScope(
|
||||||
|
scope_id=f"verify:{finding_index}",
|
||||||
|
payload={
|
||||||
|
"finding_index": finding_index,
|
||||||
|
"finding": finding,
|
||||||
|
"images_b64": images,
|
||||||
|
"text_layer_excerpt": excerpt,
|
||||||
|
},
|
||||||
|
))
|
||||||
|
return scopes
|
||||||
|
|
||||||
|
|
||||||
|
def _evidence_crops(
|
||||||
|
finding: Dict,
|
||||||
|
cited_pages: list,
|
||||||
|
sheet_to_page: Dict,
|
||||||
|
page_words: Dict,
|
||||||
|
page_to_b64: Dict,
|
||||||
|
pdf_path: str,
|
||||||
|
fallback: list,
|
||||||
|
) -> list:
|
||||||
|
"""High-DPI crops around each evidence item's source_text, located via the
|
||||||
|
page text layer. Crops REPLACE full-page images when at least one evidence
|
||||||
|
location resolves confidently; otherwise the full-page fallback is kept.
|
||||||
|
Never returns an empty list when fallback is non-empty (I2 guard)."""
|
||||||
|
crops: list = []
|
||||||
|
for item in finding.get("evidence") or []:
|
||||||
|
if len(crops) >= config.AGENT_CONFLICT_MAX_IMAGES:
|
||||||
|
break
|
||||||
|
if not isinstance(item, dict):
|
||||||
|
continue
|
||||||
|
source_text = item.get("source_text") or ""
|
||||||
|
if not source_text:
|
||||||
|
continue
|
||||||
|
# Prefer the page named on the evidence item, then any cited page.
|
||||||
|
candidates = []
|
||||||
|
named_page = sheet_to_page.get(str(item.get("sheet") or ""))
|
||||||
|
if named_page in cited_pages:
|
||||||
|
candidates.append(named_page)
|
||||||
|
candidates.extend(p for p in cited_pages if p not in candidates)
|
||||||
|
for page in candidates:
|
||||||
|
bbox = find_evidence_bbox(page_words.get(page) or [], source_text)
|
||||||
|
if bbox is None:
|
||||||
|
continue
|
||||||
|
crop = render_crop(pdf_path, page, bbox)
|
||||||
|
if not crop:
|
||||||
|
continue
|
||||||
|
crops.append(base64.b64encode(crop).decode("utf-8"))
|
||||||
|
break
|
||||||
|
return crops or fallback
|
||||||
|
|
||||||
|
|
||||||
def _counts(items, key: str) -> Dict[str, int]:
|
def _counts(items, key: str) -> Dict[str, int]:
|
||||||
counts: Dict[str, int] = {}
|
counts: Dict[str, int] = {}
|
||||||
for item in items:
|
for item in items:
|
||||||
|
|||||||
@@ -0,0 +1,108 @@
|
|||||||
|
"""Wave 5b: vision fact-check of extracted evidence against cited sheet images."""
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
|
||||||
|
from backend.llm import call_json
|
||||||
|
from backend.pipeline._serialize import dumps
|
||||||
|
from backend.pipeline._stage import collect_list, render
|
||||||
|
from backend.prompts import VERIFY_SYSTEM_PROMPT, VERIFY_USER_INSTRUCTION
|
||||||
|
|
||||||
|
_SEVERITY_RANK = {"critical": 0, "high": 1, "medium": 2, "low": 3}
|
||||||
|
_VERDICTS = ("confirmed", "corrected", "not_found")
|
||||||
|
|
||||||
|
|
||||||
|
def select_findings(findings, clusters, max_checks, severities):
|
||||||
|
"""Severity-gated selection plus any finding tied to a disputed cluster."""
|
||||||
|
disputed_keys = {c.get("key") for c in clusters if c.get("disputed_attributes")}
|
||||||
|
selected = [f for f in findings
|
||||||
|
if str(f.get("severity") or "").lower() in severities
|
||||||
|
or f.get("cluster_key") in disputed_keys]
|
||||||
|
selected.sort(key=lambda f: _SEVERITY_RANK.get(
|
||||||
|
str(f.get("severity") or "").lower(), 9))
|
||||||
|
return selected[:max_checks]
|
||||||
|
|
||||||
|
|
||||||
|
def _valid_verdict(item):
|
||||||
|
if not isinstance(item, dict):
|
||||||
|
return None
|
||||||
|
verdict = str(item.get("verdict") or "").lower()
|
||||||
|
if verdict not in _VERDICTS:
|
||||||
|
return None
|
||||||
|
return {"sheet": item.get("sheet") or "",
|
||||||
|
"source_text": item.get("source_text") or "",
|
||||||
|
"verdict": verdict,
|
||||||
|
"actual_text": item.get("actual_text"),
|
||||||
|
"notes": item.get("notes")}
|
||||||
|
|
||||||
|
|
||||||
|
def _status(verdicts):
|
||||||
|
"""Roll per-evidence verdicts up to a finding-level status.
|
||||||
|
|
||||||
|
NOTE on "corrected": it is deliberately NON-confirming. The canonical case
|
||||||
|
(job 959e16407573) is evidence quoting "(2) 2x6 STUD PACK" against a sheet
|
||||||
|
that reads "(5)" — the text exists but the VALUE the finding rests on was a
|
||||||
|
wave-1 misread, so the finding's basis is gone. Hence refuted = zero
|
||||||
|
CONFIRMED verdicts, not zero not_found ones. Do not "fix" this to treat
|
||||||
|
corrected as supporting; see tests/agents/test_verifier.py.
|
||||||
|
"""
|
||||||
|
if not verdicts:
|
||||||
|
return "unverified"
|
||||||
|
confirmed = sum(1 for v in verdicts if v["verdict"] == "confirmed")
|
||||||
|
if confirmed == len(verdicts):
|
||||||
|
return "confirmed"
|
||||||
|
if confirmed == 0:
|
||||||
|
return "refuted"
|
||||||
|
return "mixed"
|
||||||
|
|
||||||
|
|
||||||
|
class EvidenceVerifierAgent:
|
||||||
|
name = "verify"
|
||||||
|
|
||||||
|
def __init__(self, usage: AgentUsage) -> None:
|
||||||
|
self.usage = usage
|
||||||
|
|
||||||
|
def run(self, scope: AgentScope) -> AgentResult:
|
||||||
|
try:
|
||||||
|
finding = scope.payload["finding"]
|
||||||
|
instruction = render(VERIFY_USER_INSTRUCTION, {
|
||||||
|
"finding": dumps(finding),
|
||||||
|
"text_layer": scope.payload.get("text_layer_excerpt")
|
||||||
|
or "(no text layer available for the cited sheets)",
|
||||||
|
})
|
||||||
|
parsed = call_json(
|
||||||
|
system_prompt=VERIFY_SYSTEM_PROMPT,
|
||||||
|
user_text=instruction,
|
||||||
|
images_b64=scope.payload.get("images_b64") or [],
|
||||||
|
max_tokens=config.VERIFY_MAX_TOKENS,
|
||||||
|
model=config.AGENT_VERIFY_MODEL,
|
||||||
|
reasoning_effort=config.AGENT_VERIFY_REASONING_EFFORT or None,
|
||||||
|
usage_tracker=self.usage,
|
||||||
|
usage_stage="agent.verify",
|
||||||
|
)
|
||||||
|
verdicts = collect_list(parsed, "verdicts", _valid_verdict)
|
||||||
|
return AgentResult(scope_id=scope.scope_id, artifacts=[{
|
||||||
|
"finding_index": scope.payload["finding_index"],
|
||||||
|
"status": _status(verdicts),
|
||||||
|
"verdicts": verdicts,
|
||||||
|
}])
|
||||||
|
except Exception as exc:
|
||||||
|
return failure(scope, exc)
|
||||||
|
|
||||||
|
|
||||||
|
def apply_verdicts(findings, verify_results):
|
||||||
|
"""Annotate findings with verification; return refuted ones to suppress."""
|
||||||
|
by_index = {}
|
||||||
|
for result in verify_results:
|
||||||
|
for artifact in result.artifacts:
|
||||||
|
by_index[artifact["finding_index"]] = artifact
|
||||||
|
suppressed = []
|
||||||
|
for index, finding in enumerate(findings):
|
||||||
|
artifact = by_index.get(index)
|
||||||
|
if not artifact:
|
||||||
|
continue
|
||||||
|
finding["verification"] = {"status": artifact["status"],
|
||||||
|
"verdicts": artifact["verdicts"]}
|
||||||
|
if artifact["status"] == "refuted":
|
||||||
|
finding["confidence"] = "low"
|
||||||
|
suppressed.append(finding)
|
||||||
|
return suppressed
|
||||||
+125
-3
@@ -12,6 +12,15 @@ load_dotenv(os.path.join(os.path.dirname(os.path.abspath(__file__)), ".env"))
|
|||||||
|
|
||||||
_BASE_DIR = os.path.dirname(os.path.abspath(__file__))
|
_BASE_DIR = os.path.dirname(os.path.abspath(__file__))
|
||||||
|
|
||||||
|
_TRUTHY = ("1", "true", "yes", "on")
|
||||||
|
|
||||||
|
|
||||||
|
def _flag(name: str, default: str) -> bool:
|
||||||
|
"""Parse a boolean env knob. Accepts 1/true/yes/on (case-insensitive) so a
|
||||||
|
knob set to "1" behaves the same as one set to "true" — mixing bare
|
||||||
|
`== "true"` comparisons with this set silently disabled features."""
|
||||||
|
return os.getenv(name, default).strip().lower() in _TRUTHY
|
||||||
|
|
||||||
# -- AI Backend (OpenRouter) ----------------------------------------
|
# -- AI Backend (OpenRouter) ----------------------------------------
|
||||||
# One multimodal model does both extraction (Stage 1) and conflict
|
# One multimodal model does both extraction (Stage 1) and conflict
|
||||||
# reasoning (Stage 3). Override MODEL per-stage if you ever split them.
|
# reasoning (Stage 3). Override MODEL per-stage if you ever split them.
|
||||||
@@ -47,6 +56,68 @@ AGENT_CONFLICT_CONCURRENCY = int(os.getenv("AGENT_CONFLICT_CONCURRENCY", "4"))
|
|||||||
AGENT_SPECIALIST_CONCURRENCY = int(os.getenv("AGENT_SPECIALIST_CONCURRENCY", "4"))
|
AGENT_SPECIALIST_CONCURRENCY = int(os.getenv("AGENT_SPECIALIST_CONCURRENCY", "4"))
|
||||||
AGENT_RFI_CONCURRENCY = int(os.getenv("AGENT_RFI_CONCURRENCY", "4"))
|
AGENT_RFI_CONCURRENCY = int(os.getenv("AGENT_RFI_CONCURRENCY", "4"))
|
||||||
|
|
||||||
|
# -- Review focus toggles -------------------------------------------
|
||||||
|
# ENABLE_CODE_REVIEW gates the code/ADA/jurisdiction review path in BOTH
|
||||||
|
# pipelines. Default OFF: the product's focus is drawing-integrity and
|
||||||
|
# cross-discipline coordination, not code/accessibility compliance. When
|
||||||
|
# False the CodeAgent wave (agent) and the Code/ADA stage (classic) are
|
||||||
|
# skipped entirely, by_stage.code reports 0, and nothing in the ADA corpus
|
||||||
|
# or jurisdiction meta is deleted so the path can be re-enabled with one env
|
||||||
|
# flag. Set ENABLE_CODE_REVIEW=1 to restore code/ADA findings.
|
||||||
|
ENABLE_CODE_REVIEW = _flag("ENABLE_CODE_REVIEW", "false")
|
||||||
|
|
||||||
|
# Per-sheet Drawing Integrity QA wave (agent + classic). This is the
|
||||||
|
# drawing-focused pass: it reads ONE sheet's own objects + image + text layer
|
||||||
|
# and flags problems internal to that sheet -- dangling detail/callout/keynote
|
||||||
|
# references, schedule-vs-plan or legend disagreements on the same sheet,
|
||||||
|
# dimension strings that do not sum, missing title-block/scale/north-arrow,
|
||||||
|
# and notes that contradict each other. It complements (does not replace) the
|
||||||
|
# cross-sheet conflict critic. Default ON.
|
||||||
|
ENABLE_DRAWING_INTEGRITY = _flag("ENABLE_DRAWING_INTEGRITY", "true")
|
||||||
|
AGENT_INTEGRITY_MODEL = os.getenv("AGENT_INTEGRITY_MODEL", "") or MODEL
|
||||||
|
AGENT_INTEGRITY_CONCURRENCY = int(os.getenv("AGENT_INTEGRITY_CONCURRENCY", "4"))
|
||||||
|
AGENT_INTEGRITY_MAX_IMAGES = int(os.getenv("AGENT_INTEGRITY_MAX_IMAGES", "1"))
|
||||||
|
AGENT_INTEGRITY_MAX_ASSERTIONS = int(os.getenv("AGENT_INTEGRITY_MAX_ASSERTIONS", "80"))
|
||||||
|
INTEGRITY_MAX_TOKENS = int(os.getenv("INTEGRITY_MAX_TOKENS", "16384"))
|
||||||
|
# Skip sheets with fewer than this many extracted objects -- too sparse for a
|
||||||
|
# meaningful internal-consistency pass (avoids burning a call on near-empty pages).
|
||||||
|
INTEGRITY_MIN_ASSERTIONS = int(os.getenv("INTEGRITY_MIN_ASSERTIONS", "3"))
|
||||||
|
|
||||||
|
# Wave 5b evidence verification (vision fact-check of cited sheet text)
|
||||||
|
AGENT_VERIFY_MODEL = os.getenv("AGENT_VERIFY_MODEL", "") or MODEL
|
||||||
|
AGENT_VERIFY_CONCURRENCY = int(os.getenv("AGENT_VERIFY_CONCURRENCY", "4"))
|
||||||
|
AGENT_VERIFY_MAX_CHECKS = int(os.getenv("AGENT_VERIFY_MAX_CHECKS", "20"))
|
||||||
|
AGENT_VERIFY_SEVERITIES = {
|
||||||
|
s.strip().lower()
|
||||||
|
for s in os.getenv("AGENT_VERIFY_SEVERITIES", "critical,high").split(",")
|
||||||
|
if s.strip()
|
||||||
|
}
|
||||||
|
AGENT_VERIFY_REASONING_EFFORT = os.getenv("AGENT_VERIFY_REASONING_EFFORT", "low").strip()
|
||||||
|
VERIFY_MAX_TOKENS = int(os.getenv("VERIFY_MAX_TOKENS", "8192"))
|
||||||
|
|
||||||
|
# Wave 6.5 Brain-directed clarification. After the Brain merge, the Brain may
|
||||||
|
# name findings it is unsure about and emit typed clarification requests; v1
|
||||||
|
# executes verify_evidence requests by routing them back through the wave-5b
|
||||||
|
# EvidenceVerifierAgent (fresh page images + hi-DPI evidence crops + text-layer
|
||||||
|
# oracle). Bounded: one planning call, at most BRAIN_CLARIFY_MAX_REQUESTS
|
||||||
|
# verifications, a single iteration. Reuses AGENT_VERIFY_* / VERIFY_* knobs for
|
||||||
|
# the verification calls. Default ON.
|
||||||
|
ENABLE_BRAIN_CLARIFY = _flag("ENABLE_BRAIN_CLARIFY", "true")
|
||||||
|
BRAIN_CLARIFY_MAX_REQUESTS = int(os.getenv("BRAIN_CLARIFY_MAX_REQUESTS", "8"))
|
||||||
|
BRAIN_CLARIFY_MAX_TOKENS = int(os.getenv("BRAIN_CLARIFY_MAX_TOKENS", "4096"))
|
||||||
|
|
||||||
|
# -- Text-layer grounding (deterministic PDF text layer via PyMuPDF) ----
|
||||||
|
# The vector text layer is extracted once per job and grounds the extractor,
|
||||||
|
# rescues misquoted-but-real values in the grounding guard, and serves the
|
||||||
|
# wave-5b verifier as a text oracle plus high-DPI evidence crops.
|
||||||
|
TEXT_LAYER_ENABLED = os.getenv("TEXT_LAYER_ENABLED", "true").strip().lower() in ("1", "true", "yes")
|
||||||
|
TEXT_LAYER_MIN_CHARS = int(os.getenv("TEXT_LAYER_MIN_CHARS", "20")) # below this per page -> no text layer
|
||||||
|
TEXT_LAYER_MAX_CHARS = int(os.getenv("TEXT_LAYER_MAX_CHARS", "12000")) # cap per sheet in extractor prompt
|
||||||
|
VERIFY_TEXT_MAX_CHARS = int(os.getenv("VERIFY_TEXT_MAX_CHARS", "8000"))# cap of excerpt in verify scope
|
||||||
|
VERIFY_HI_DPI_CROPS = os.getenv("VERIFY_HI_DPI_CROPS", "true").strip().lower() in ("1", "true", "yes")
|
||||||
|
VERIFY_CROP_DPI = int(os.getenv("VERIFY_CROP_DPI", "300"))
|
||||||
|
VERIFY_CROP_MARGIN_PTS = int(os.getenv("VERIFY_CROP_MARGIN_PTS", "36"))# padding around evidence bbox (PDF points)
|
||||||
|
|
||||||
# Agent-mode human-review gate. When on (default), Agent runs stop after the
|
# Agent-mode human-review gate. When on (default), Agent runs stop after the
|
||||||
# Brain merge and wait for human decisions before RFIs/final report/email go
|
# Brain merge and wait for human decisions before RFIs/final report/email go
|
||||||
# out. AGENT_REVIEW_AUDIT_SAMPLE caps how many clean clusters get added to the
|
# out. AGENT_REVIEW_AUDIT_SAMPLE caps how many clean clusters get added to the
|
||||||
@@ -58,6 +129,26 @@ AGENT_REVIEW_AUDIT_SAMPLE = int(os.getenv("AGENT_REVIEW_AUDIT_SAMPLE", "5"))
|
|||||||
# NOTE: currently unwired - reserved for future cross-job aggregation tooling.
|
# NOTE: currently unwired - reserved for future cross-job aggregation tooling.
|
||||||
REVIEW_AGGREGATE_INCLUDE_TEXT = os.getenv("REVIEW_AGGREGATE_INCLUDE_TEXT", "false").strip().lower() in ("1", "true", "yes")
|
REVIEW_AGGREGATE_INCLUDE_TEXT = os.getenv("REVIEW_AGGREGATE_INCLUDE_TEXT", "false").strip().lower() in ("1", "true", "yes")
|
||||||
|
|
||||||
|
# -- Review chat (ask-the-run Q&A on the review screen) --------------
|
||||||
|
# A read-only explainer: it answers "why did the run decide X?" from the job's
|
||||||
|
# own artifacts and never mutates findings, decisions, or code. Every turn is
|
||||||
|
# appended to <out_dir>/review/chat_log.jsonl AND to the cross-job feedback
|
||||||
|
# store (REVIEW_FEEDBACK_DIR) so answers are available to future prompt priors.
|
||||||
|
# HISTORY_TURNS caps how much of a thread is replayed into the prompt;
|
||||||
|
# LOG_LINES caps how many job.log lines are searched into the context bundle.
|
||||||
|
ENABLE_REVIEW_CHAT = _flag("ENABLE_REVIEW_CHAT", "true")
|
||||||
|
REVIEW_CHAT_MODEL = os.getenv("REVIEW_CHAT_MODEL", "") or TEXT_MODEL
|
||||||
|
REVIEW_CHAT_MAX_TOKENS = int(os.getenv("REVIEW_CHAT_MAX_TOKENS", "4096"))
|
||||||
|
REVIEW_CHAT_HISTORY_TURNS = int(os.getenv("REVIEW_CHAT_HISTORY_TURNS", "6"))
|
||||||
|
REVIEW_CHAT_LOG_LINES = int(os.getenv("REVIEW_CHAT_LOG_LINES", "40"))
|
||||||
|
REVIEW_CHAT_MAX_QUESTION_CHARS = int(os.getenv("REVIEW_CHAT_MAX_QUESTION_CHARS", "2000"))
|
||||||
|
# Cross-job feedback store: where review decisions and chat turns accumulate so
|
||||||
|
# a future run can be primed with "what humans corrected last time". Job-local
|
||||||
|
# artifacts stay the source of truth; this is the append-only roll-up.
|
||||||
|
# (OUTPUT_DIR is defined further down; keep this in sync with it.)
|
||||||
|
REVIEW_FEEDBACK_DIR = os.getenv("REVIEW_FEEDBACK_DIR", "") or os.path.join(
|
||||||
|
_BASE_DIR, "outputs", "_feedback")
|
||||||
|
|
||||||
# -- Hybrid (local text LLM) ----------------------------------------
|
# -- Hybrid (local text LLM) ----------------------------------------
|
||||||
# Optional OpenAI-compatible local endpoint (e.g. a vLLM box) for the text-only
|
# Optional OpenAI-compatible local endpoint (e.g. a vLLM box) for the text-only
|
||||||
# QAQC stages. Vision stages ALWAYS use OpenRouter. The user picks hybrid per
|
# QAQC stages. Vision stages ALWAYS use OpenRouter. The user picks hybrid per
|
||||||
@@ -73,7 +164,30 @@ PDF_DPI = int(os.getenv("PDF_DPI", "100"))
|
|||||||
MAX_PAGES = int(os.getenv("MAX_PAGES", "60"))
|
MAX_PAGES = int(os.getenv("MAX_PAGES", "60"))
|
||||||
MAX_DIMENSION = int(os.getenv("MAX_DIMENSION", "2400")) # px cap on the long edge
|
MAX_DIMENSION = int(os.getenv("MAX_DIMENSION", "2400")) # px cap on the long edge
|
||||||
LLM_TIMEOUT = int(os.getenv("LLM_TIMEOUT", "180")) # seconds per call
|
LLM_TIMEOUT = int(os.getenv("LLM_TIMEOUT", "180")) # seconds per call
|
||||||
EXTRACT_MAX_TOKENS = int(os.getenv("EXTRACT_MAX_TOKENS", "16384"))
|
# Gemini 2.5 Pro counts thinking tokens against max_tokens, so the visible
|
||||||
|
# JSON budget is well under this number on dense sheets. 65536 is the model's
|
||||||
|
# output ceiling - give thinking all the room it wants so visible JSON never
|
||||||
|
# truncates; the thinking budget itself is capped separately below.
|
||||||
|
EXTRACT_MAX_TOKENS = int(os.getenv("EXTRACT_MAX_TOKENS", "65536"))
|
||||||
|
# Reasoning effort for the per-sheet extractor (OpenRouter reasoning knob).
|
||||||
|
# Extraction is perceptive, not deliberative - "low" keeps thinking tokens
|
||||||
|
# from eating the output budget. Empty string disables the parameter.
|
||||||
|
EXTRACT_REASONING_EFFORT = os.getenv("EXTRACT_REASONING_EFFORT", "low").strip()
|
||||||
|
# Hard thinking-token budget for the extractor (OpenRouter reasoning
|
||||||
|
# max_tokens -> Gemini thinking_budget). "low" effort alone still let Gemini
|
||||||
|
# burn ~25k thinking tokens per sheet (job 98194fa8d215); a hard cap forces
|
||||||
|
# the budget into visible output. 0 disables -> falls back to the effort knob.
|
||||||
|
# Mutually exclusive with effort when set (OpenRouter rejects both together).
|
||||||
|
EXTRACT_REASONING_MAX_TOKENS = int(os.getenv("EXTRACT_REASONING_MAX_TOKENS", "2048"))
|
||||||
|
# Coverage-driven extraction retry ladder. After the vision pass, the fraction
|
||||||
|
# of meaningful text-layer lines represented in extracted objects is measured;
|
||||||
|
# below EXTRACT_COVERAGE_FLOOR the page climbs the ladder: text-only
|
||||||
|
# structuring pass (rung 2), then deterministic text-layer fallback stubs
|
||||||
|
# (rung 3) so no text-bearing page goes dark.
|
||||||
|
EXTRACT_COVERAGE_FLOOR = float(os.getenv("EXTRACT_COVERAGE_FLOOR", "0.6"))
|
||||||
|
EXTRACT_TEXT_RETRY_ENABLED = _flag("EXTRACT_TEXT_RETRY_ENABLED", "true")
|
||||||
|
EXTRACT_FALLBACK_ENABLED = _flag("EXTRACT_FALLBACK_ENABLED", "true")
|
||||||
|
EXTRACT_FALLBACK_MAX_OBJECTS = int(os.getenv("EXTRACT_FALLBACK_MAX_OBJECTS", "200"))
|
||||||
REASON_MAX_TOKENS = int(os.getenv("REASON_MAX_TOKENS", "4096"))
|
REASON_MAX_TOKENS = int(os.getenv("REASON_MAX_TOKENS", "4096"))
|
||||||
|
|
||||||
# -- QAQC stage knobs (Stages 0-1, 3, 6-11) -------------------------
|
# -- QAQC stage knobs (Stages 0-1, 3, 6-11) -------------------------
|
||||||
@@ -102,6 +216,14 @@ CLUSTER_MAX = int(os.getenv("CLUSTER_MAX", "120"))
|
|||||||
LLM_CACHE = os.getenv("LLM_CACHE", "false").strip().lower() in ("1", "true", "yes")
|
LLM_CACHE = os.getenv("LLM_CACHE", "false").strip().lower() in ("1", "true", "yes")
|
||||||
LLM_CACHE_DIR = os.getenv("LLM_CACHE_DIR", os.path.join(_BASE_DIR, ".llm_cache"))
|
LLM_CACHE_DIR = os.getenv("LLM_CACHE_DIR", os.path.join(_BASE_DIR, ".llm_cache"))
|
||||||
|
|
||||||
|
# Verbose LLM observability. Per call, one line lands in the job log (model,
|
||||||
|
# backend, prompt size, response size, parsed-item counts, per-call cost) and
|
||||||
|
# the full request/response is dumped to <out_dir>/llm_raw/ (base64 image
|
||||||
|
# payloads excluded; image count recorded instead) so missed or hallucinated
|
||||||
|
# items can be traced back to exactly what the model saw and returned.
|
||||||
|
LLM_VERBOSE = os.getenv("LLM_VERBOSE", "true").strip().lower() in ("1", "true", "yes")
|
||||||
|
LLM_RAW_DUMP = os.getenv("LLM_RAW_DUMP", "true").strip().lower() in ("1", "true", "yes")
|
||||||
|
|
||||||
# Parallelism (ThreadPoolExecutor workers)
|
# Parallelism (ThreadPoolExecutor workers)
|
||||||
EXTRACT_CONCURRENCY = int(os.getenv("EXTRACT_CONCURRENCY", "4"))
|
EXTRACT_CONCURRENCY = int(os.getenv("EXTRACT_CONCURRENCY", "4"))
|
||||||
REASON_CONCURRENCY = int(os.getenv("REASON_CONCURRENCY", "4"))
|
REASON_CONCURRENCY = int(os.getenv("REASON_CONCURRENCY", "4"))
|
||||||
@@ -136,5 +258,5 @@ SMTP_PORT = int(os.getenv("SMTP_PORT", "587"))
|
|||||||
SMTP_USER = os.getenv("SMTP_USER", "")
|
SMTP_USER = os.getenv("SMTP_USER", "")
|
||||||
SMTP_PASSWORD = os.getenv("SMTP_PASSWORD", "")
|
SMTP_PASSWORD = os.getenv("SMTP_PASSWORD", "")
|
||||||
SMTP_FROM = os.getenv("SMTP_FROM", "")
|
SMTP_FROM = os.getenv("SMTP_FROM", "")
|
||||||
SMTP_USE_TLS = os.getenv("SMTP_USE_TLS", "true").lower() == "true"
|
SMTP_USE_TLS = _flag("SMTP_USE_TLS", "true")
|
||||||
SMTP_USE_SSL = os.getenv("SMTP_USE_SSL", "false").lower() == "true"
|
SMTP_USE_SSL = _flag("SMTP_USE_SSL", "false")
|
||||||
@@ -0,0 +1,93 @@
|
|||||||
|
"""
|
||||||
|
job_log.py - Capture pipeline stdout/stderr into a per-job log.
|
||||||
|
|
||||||
|
The pipeline already prints stage progress via print(). For a reviewable
|
||||||
|
post-run log we tee those lines into memory + outputs/<job_id>/job.log
|
||||||
|
without rewriting every call site.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import sys
|
||||||
|
import time
|
||||||
|
from contextlib import contextmanager
|
||||||
|
from typing import Callable, Iterator, List, Optional, TextIO
|
||||||
|
|
||||||
|
|
||||||
|
class _LineSplitter:
|
||||||
|
"""Accumulate write() chunks and emit complete lines."""
|
||||||
|
|
||||||
|
def __init__(self, on_line: Callable[[str], None]):
|
||||||
|
self._buf = ""
|
||||||
|
self._on_line = on_line
|
||||||
|
|
||||||
|
def write(self, s: str) -> None:
|
||||||
|
if not s:
|
||||||
|
return
|
||||||
|
self._buf += s
|
||||||
|
while "\n" in self._buf:
|
||||||
|
line, self._buf = self._buf.split("\n", 1)
|
||||||
|
# Strip trailing CR from Windows-ish streams; keep content intact.
|
||||||
|
self._on_line(line.rstrip("\r"))
|
||||||
|
|
||||||
|
def flush_remainder(self) -> None:
|
||||||
|
if self._buf:
|
||||||
|
self._on_line(self._buf.rstrip("\r"))
|
||||||
|
self._buf = ""
|
||||||
|
|
||||||
|
|
||||||
|
class _Tee:
|
||||||
|
"""Mirror writes to the original stream and a line callback."""
|
||||||
|
|
||||||
|
def __init__(self, stream: TextIO, on_line: Callable[[str], None]):
|
||||||
|
self._stream = stream
|
||||||
|
self._lines = _LineSplitter(on_line)
|
||||||
|
|
||||||
|
def write(self, s: str) -> int:
|
||||||
|
n = self._stream.write(s)
|
||||||
|
self._stream.flush()
|
||||||
|
self._lines.write(s)
|
||||||
|
return n
|
||||||
|
|
||||||
|
def flush(self) -> None:
|
||||||
|
self._stream.flush()
|
||||||
|
|
||||||
|
def flush_remainder(self) -> None:
|
||||||
|
self._lines.flush_remainder()
|
||||||
|
|
||||||
|
def __getattr__(self, name: str):
|
||||||
|
return getattr(self._stream, name)
|
||||||
|
|
||||||
|
|
||||||
|
def stamp_line(line: str, t: Optional[float] = None) -> str:
|
||||||
|
"""Prefix a log line with HH:MM:SS."""
|
||||||
|
ts = time.strftime("%H:%M:%S", time.localtime(t if t is not None else time.time()))
|
||||||
|
return f"[{ts}] {line}"
|
||||||
|
|
||||||
|
|
||||||
|
@contextmanager
|
||||||
|
def capture_stdio(on_line: Callable[[str], None]) -> Iterator[None]:
|
||||||
|
"""
|
||||||
|
Tee sys.stdout and sys.stderr into on_line(raw_line) for the duration.
|
||||||
|
|
||||||
|
Safe for the single-job-at-a-time usage of this app; overlapping jobs
|
||||||
|
would interleave (same limitation as the LLM cost counters).
|
||||||
|
"""
|
||||||
|
old_out, old_err = sys.stdout, sys.stderr
|
||||||
|
tee_out = _Tee(old_out, on_line)
|
||||||
|
tee_err = _Tee(old_err, on_line)
|
||||||
|
sys.stdout = tee_out # type: ignore[assignment]
|
||||||
|
sys.stderr = tee_err # type: ignore[assignment]
|
||||||
|
try:
|
||||||
|
yield
|
||||||
|
finally:
|
||||||
|
tee_out.flush_remainder()
|
||||||
|
tee_err.flush_remainder()
|
||||||
|
sys.stdout = old_out
|
||||||
|
sys.stderr = old_err
|
||||||
|
|
||||||
|
|
||||||
|
def read_log_file(path: str) -> List[str]:
|
||||||
|
try:
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
return [ln.rstrip("\n") for ln in f]
|
||||||
|
except OSError:
|
||||||
|
return []
|
||||||
+199
-73
@@ -9,61 +9,34 @@ for the completion email.
|
|||||||
State is in-memory (fine for a single-user tool); the report is also persisted
|
State is in-memory (fine for a single-user tool); the report is also persisted
|
||||||
to outputs/<job_id>/ so results survive a restart even though live status does
|
to outputs/<job_id>/ so results survive a restart even though live status does
|
||||||
not. No external queue/DB.
|
not. No external queue/DB.
|
||||||
|
|
||||||
|
A teed stdout/stderr log is kept in memory and written to outputs/<job_id>/job.log
|
||||||
|
so failed or suspicious runs can be reviewed after the fact.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
import json
|
import json
|
||||||
import contextlib
|
|
||||||
import os
|
import os
|
||||||
import sys
|
|
||||||
import time
|
import time
|
||||||
|
import traceback
|
||||||
import uuid
|
import uuid
|
||||||
import shutil
|
import shutil
|
||||||
import threading
|
import threading
|
||||||
from typing import Dict, Optional
|
from contextlib import contextmanager
|
||||||
|
from typing import Dict, Iterator, List, Optional
|
||||||
|
|
||||||
from backend import config
|
from backend import config
|
||||||
from backend import llm
|
from backend import llm
|
||||||
|
from backend.job_log import capture_stdio, read_log_file, stamp_line
|
||||||
from backend.agents.runner import run_agent_pipeline
|
from backend.agents.runner import run_agent_pipeline
|
||||||
from backend.pipeline.runner import run_pipeline
|
from backend.pipeline.runner import run_pipeline
|
||||||
from backend.email_sender import send_conflict_report, send_review_required
|
from backend.email_sender import send_conflict_report, send_review_required
|
||||||
|
|
||||||
_jobs: Dict[str, Dict] = {}
|
_jobs: Dict[str, Dict] = {}
|
||||||
_lock = threading.Lock()
|
_lock = threading.Lock()
|
||||||
|
_LOG_TAIL = 80
|
||||||
PIPELINE_MODES = {"classic", "agent"}
|
PIPELINE_MODES = {"classic", "agent"}
|
||||||
|
# States where the job will produce no more log output; polls get the full log.
|
||||||
|
_TERMINAL_STATES = {"done", "error", "needs_review", "finalization_error"}
|
||||||
class _Tee:
|
|
||||||
"""Write to both the real stream and the job log file."""
|
|
||||||
|
|
||||||
def __init__(self, stream, log_file) -> None:
|
|
||||||
self._stream = stream
|
|
||||||
self._log = log_file
|
|
||||||
|
|
||||||
def write(self, data):
|
|
||||||
self._stream.write(data)
|
|
||||||
self._log.write(data)
|
|
||||||
|
|
||||||
def flush(self):
|
|
||||||
self._stream.flush()
|
|
||||||
self._log.flush()
|
|
||||||
|
|
||||||
|
|
||||||
@contextlib.contextmanager
|
|
||||||
def _tee_log(log_path: str, header: str):
|
|
||||||
"""Mirror stdout/stderr into a per-job log file for the duration of a run.
|
|
||||||
|
|
||||||
sys.stdout is process-global, so two concurrent jobs would interleave in
|
|
||||||
each other's logs - acceptable for this single-user tool (same tradeoff as
|
|
||||||
the LLM cost globals in llm.py).
|
|
||||||
"""
|
|
||||||
with open(log_path, "a", encoding="utf-8") as log_file:
|
|
||||||
log_file.write(header + "\n")
|
|
||||||
real_out, real_err = sys.stdout, sys.stderr
|
|
||||||
sys.stdout, sys.stderr = _Tee(real_out, log_file), _Tee(real_err, log_file)
|
|
||||||
try:
|
|
||||||
yield
|
|
||||||
finally:
|
|
||||||
sys.stdout, sys.stderr = real_out, real_err
|
|
||||||
|
|
||||||
|
|
||||||
def _set(job_id: str, **fields) -> None:
|
def _set(job_id: str, **fields) -> None:
|
||||||
@@ -71,16 +44,39 @@ def _set(job_id: str, **fields) -> None:
|
|||||||
_jobs[job_id].update(fields)
|
_jobs[job_id].update(fields)
|
||||||
|
|
||||||
|
|
||||||
def create_job(pdf_path: str, source_filename: str, email: Optional[str] = None,
|
def _append_log(job_id: str, raw_line: str, log_path: str) -> None:
|
||||||
project_input: Optional[Dict] = None, text_local: bool = False,
|
"""Stamp, store, and append one captured stdout/stderr line."""
|
||||||
pipeline_mode: str = "classic", model: Optional[str] = None) -> str:
|
entry = stamp_line(raw_line)
|
||||||
|
with _lock:
|
||||||
|
job = _jobs.get(job_id)
|
||||||
|
if job is not None:
|
||||||
|
job.setdefault("log", []).append(entry)
|
||||||
|
try:
|
||||||
|
os.makedirs(os.path.dirname(log_path), exist_ok=True)
|
||||||
|
with open(log_path, "a", encoding="utf-8") as f:
|
||||||
|
f.write(entry + "\n")
|
||||||
|
except OSError:
|
||||||
|
# Don't fail the job over log I/O; avoid print() here — it would
|
||||||
|
# re-enter the stdio tee while a job is capturing.
|
||||||
|
pass
|
||||||
|
|
||||||
|
|
||||||
|
def create_job(
|
||||||
|
pdf_path: str,
|
||||||
|
source_filename: str,
|
||||||
|
email: Optional[str] = None,
|
||||||
|
project_input: Optional[Dict] = None,
|
||||||
|
text_local: bool = False,
|
||||||
|
pipeline_mode: str = "classic",
|
||||||
|
vision_model: Optional[str] = None,
|
||||||
|
text_model: Optional[str] = None,
|
||||||
|
) -> str:
|
||||||
"""Register a job and kick off its background thread. Returns the job_id."""
|
"""Register a job and kick off its background thread. Returns the job_id."""
|
||||||
pipeline_mode = pipeline_mode.strip().lower()
|
pipeline_mode = pipeline_mode.strip().lower()
|
||||||
if pipeline_mode not in PIPELINE_MODES:
|
if pipeline_mode not in PIPELINE_MODES:
|
||||||
raise ValueError(f"Unsupported pipeline mode: {pipeline_mode!r}")
|
raise ValueError(f"Unsupported pipeline mode: {pipeline_mode!r}")
|
||||||
# Agent mode v1 is OpenRouter-only.
|
# Agent mode v1 is OpenRouter-only.
|
||||||
text_local = bool(text_local and pipeline_mode == "classic")
|
text_local = bool(text_local and pipeline_mode == "classic")
|
||||||
model = (model or "").strip() or None
|
|
||||||
job_id = uuid.uuid4().hex[:12]
|
job_id = uuid.uuid4().hex[:12]
|
||||||
with _lock:
|
with _lock:
|
||||||
_jobs[job_id] = {
|
_jobs[job_id] = {
|
||||||
@@ -91,35 +87,67 @@ def create_job(pdf_path: str, source_filename: str, email: Optional[str] = None,
|
|||||||
"project_input": project_input or {},
|
"project_input": project_input or {},
|
||||||
"text_local": text_local,
|
"text_local": text_local,
|
||||||
"pipeline_mode": pipeline_mode,
|
"pipeline_mode": pipeline_mode,
|
||||||
"model": model,
|
"vision_model": (vision_model or "").strip() or None,
|
||||||
|
"text_model": (text_model or "").strip() or None,
|
||||||
"stage": None,
|
"stage": None,
|
||||||
"created_at": time.time(),
|
"created_at": time.time(),
|
||||||
"finished_at": None,
|
"finished_at": None,
|
||||||
"report": None,
|
"report": None,
|
||||||
"error": None,
|
"error": None,
|
||||||
|
"log": [],
|
||||||
}
|
}
|
||||||
threading.Thread(target=_run, args=(
|
threading.Thread(
|
||||||
job_id, pdf_path, project_input, text_local, pipeline_mode, model,
|
target=_run,
|
||||||
),
|
args=(job_id, pdf_path, project_input, text_local, pipeline_mode,
|
||||||
daemon=True).start()
|
vision_model, text_model),
|
||||||
|
daemon=True,
|
||||||
|
).start()
|
||||||
return job_id
|
return job_id
|
||||||
|
|
||||||
|
|
||||||
def _run(job_id: str, pdf_path: str, project_input: Optional[Dict] = None,
|
def _run(
|
||||||
text_local: bool = False, pipeline_mode: str = "classic",
|
job_id: str,
|
||||||
model: Optional[str] = None) -> None:
|
pdf_path: str,
|
||||||
|
project_input: Optional[Dict] = None,
|
||||||
|
text_local: bool = False,
|
||||||
|
pipeline_mode: str = "classic",
|
||||||
|
vision_model: Optional[str] = None,
|
||||||
|
text_model: Optional[str] = None,
|
||||||
|
) -> None:
|
||||||
out_dir = os.path.join(config.OUTPUT_DIR, job_id)
|
out_dir = os.path.join(config.OUTPUT_DIR, job_id)
|
||||||
|
log_path = os.path.join(out_dir, "job.log")
|
||||||
try:
|
try:
|
||||||
_set(job_id, status="running")
|
_set(job_id, status="running")
|
||||||
# Keep a copy of the source PDF so its sheets can be viewed later.
|
|
||||||
os.makedirs(out_dir, exist_ok=True)
|
os.makedirs(out_dir, exist_ok=True)
|
||||||
|
# Truncate any leftover log if job_id somehow collided (shouldn't).
|
||||||
|
with open(log_path, "w", encoding="utf-8"):
|
||||||
|
pass
|
||||||
header = (f"=== Job {job_id} | {pipeline_mode} | {_jobs[job_id].get('source')} | "
|
header = (f"=== Job {job_id} | {pipeline_mode} | {_jobs[job_id].get('source')} | "
|
||||||
f"model={model or 'default'} | "
|
f"vision={vision_model or 'default'} text={text_model or 'default'} | "
|
||||||
f"started {time.strftime('%Y-%m-%d %H:%M:%S %Z', time.gmtime())} UTC ===")
|
f"started {time.strftime('%Y-%m-%d %H:%M:%S %Z', time.gmtime())} UTC ===")
|
||||||
with _tee_log(os.path.join(out_dir, "job.log"), header):
|
_append_log(job_id, header, log_path)
|
||||||
|
|
||||||
|
def on_line(raw: str) -> None:
|
||||||
|
_append_log(job_id, raw, log_path)
|
||||||
|
|
||||||
|
with capture_stdio(on_line):
|
||||||
_run_pipeline(job_id, pdf_path, out_dir, project_input, text_local,
|
_run_pipeline(job_id, pdf_path, out_dir, project_input, text_local,
|
||||||
pipeline_mode, model)
|
pipeline_mode, vision_model, text_model)
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
|
# Land the failure AND its traceback in the job log so failed runs can
|
||||||
|
# be diagnosed from the log alone (the stdio tee is already torn down).
|
||||||
|
try:
|
||||||
|
_append_log(job_id, f"[Jobs] Job {job_id} failed: {e}", log_path)
|
||||||
|
for ln in traceback.format_exc().rstrip().splitlines():
|
||||||
|
_append_log(job_id, ln, log_path)
|
||||||
|
# Cost-so-far for the failed run (counters are reset per job).
|
||||||
|
cost = llm.get_cost()
|
||||||
|
_append_log(job_id,
|
||||||
|
f"[Jobs] Estimated LLM cost before failure: "
|
||||||
|
f"${cost['usd']:.4f} over {cost['calls']} live calls "
|
||||||
|
f"({cost.get('cached', 0)} cached)", log_path)
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
print(f"[Jobs] Job {job_id} failed: {e}")
|
print(f"[Jobs] Job {job_id} failed: {e}")
|
||||||
_set(job_id, status="error", error=str(e), finished_at=time.time())
|
_set(job_id, status="error", error=str(e), finished_at=time.time())
|
||||||
_notify_error(job_id)
|
_notify_error(job_id)
|
||||||
@@ -132,7 +160,8 @@ def _run(job_id: str, pdf_path: str, project_input: Optional[Dict] = None,
|
|||||||
|
|
||||||
def _run_pipeline(job_id: str, pdf_path: str, out_dir: str,
|
def _run_pipeline(job_id: str, pdf_path: str, out_dir: str,
|
||||||
project_input: Optional[Dict], text_local: bool,
|
project_input: Optional[Dict], text_local: bool,
|
||||||
pipeline_mode: str, model: Optional[str]) -> None:
|
pipeline_mode: str, vision_model: Optional[str],
|
||||||
|
text_model: Optional[str]) -> None:
|
||||||
"""The body of a job run; executes inside the job's tee'd log capture."""
|
"""The body of a job run; executes inside the job's tee'd log capture."""
|
||||||
# Persist minimal job metadata so the disk fallback in get_job can
|
# Persist minimal job metadata so the disk fallback in get_job can
|
||||||
# recover the recipient email / pipeline mode after a server restart
|
# recover the recipient email / pipeline mode after a server restart
|
||||||
@@ -143,8 +172,14 @@ def _run_pipeline(job_id: str, pdf_path: str, out_dir: str,
|
|||||||
"email": _jobs[job_id].get("email"),
|
"email": _jobs[job_id].get("email"),
|
||||||
"pipeline_mode": pipeline_mode,
|
"pipeline_mode": pipeline_mode,
|
||||||
"source": _jobs[job_id].get("source"),
|
"source": _jobs[job_id].get("source"),
|
||||||
|
"vision_model": vision_model,
|
||||||
|
"text_model": text_model,
|
||||||
}, f, indent=2)
|
}, f, indent=2)
|
||||||
|
# Keep a copy of the source PDF so its sheets can be viewed later.
|
||||||
shutil.copy2(pdf_path, os.path.join(out_dir, "source.pdf"))
|
shutil.copy2(pdf_path, os.path.join(out_dir, "source.pdf"))
|
||||||
|
# Raw per-call LLM request/response dumps land in <out_dir>/llm_raw/.
|
||||||
|
if config.LLM_RAW_DUMP:
|
||||||
|
llm.set_raw_dump_dir(os.path.join(out_dir, "llm_raw"))
|
||||||
runner = run_agent_pipeline if pipeline_mode == "agent" else run_pipeline
|
runner = run_agent_pipeline if pipeline_mode == "agent" else run_pipeline
|
||||||
runner_kwargs = {
|
runner_kwargs = {
|
||||||
"out_dir": out_dir,
|
"out_dir": out_dir,
|
||||||
@@ -153,18 +188,25 @@ def _run_pipeline(job_id: str, pdf_path: str, out_dir: str,
|
|||||||
"source_name": _jobs[job_id].get("source"),
|
"source_name": _jobs[job_id].get("source"),
|
||||||
}
|
}
|
||||||
if pipeline_mode == "classic":
|
if pipeline_mode == "classic":
|
||||||
|
# run_pipeline takes the picks as params and clears them in finally.
|
||||||
runner_kwargs["text_local"] = text_local
|
runner_kwargs["text_local"] = text_local
|
||||||
|
runner_kwargs["vision_model"] = vision_model
|
||||||
|
runner_kwargs["text_model"] = text_model
|
||||||
else:
|
else:
|
||||||
runner_kwargs["require_review"] = config.AGENT_REQUIRE_REVIEW
|
runner_kwargs["require_review"] = config.AGENT_REQUIRE_REVIEW
|
||||||
if model:
|
# The agent runner has no override params; set them module-level.
|
||||||
print(f"[Jobs] Model override for this run: {model}")
|
if vision_model or text_model:
|
||||||
llm.set_model_override(model)
|
print(f"[Jobs] Model overrides for this run: "
|
||||||
|
f"vision={vision_model or '(default)'} text={text_model or '(default)'}")
|
||||||
|
llm.set_model_overrides(vision_model, text_model)
|
||||||
try:
|
try:
|
||||||
report = runner(pdf_path, **runner_kwargs)
|
report = runner(pdf_path, **runner_kwargs)
|
||||||
finally:
|
finally:
|
||||||
if model:
|
llm.set_raw_dump_dir(None)
|
||||||
llm.set_model_override(None)
|
if pipeline_mode == "agent":
|
||||||
|
llm.set_model_overrides(None, None)
|
||||||
report.setdefault("summary", {})["pipeline_mode"] = pipeline_mode
|
report.setdefault("summary", {})["pipeline_mode"] = pipeline_mode
|
||||||
|
_log_cost_summary(report.get("summary", {}))
|
||||||
if report["summary"].get("agent_status") == "needs_review":
|
if report["summary"].get("agent_status") == "needs_review":
|
||||||
# Human-review gate: hold the job, don't email the unreviewed report.
|
# Human-review gate: hold the job, don't email the unreviewed report.
|
||||||
_set(job_id, status="needs_review", report=report,
|
_set(job_id, status="needs_review", report=report,
|
||||||
@@ -178,6 +220,55 @@ def _run_pipeline(job_id: str, pdf_path: str, out_dir: str,
|
|||||||
_notify(job_id, report, out_dir)
|
_notify(job_id, report, out_dir)
|
||||||
|
|
||||||
|
|
||||||
|
def _log_cost_summary(summary: Dict, label: str = "this run") -> None:
|
||||||
|
"""
|
||||||
|
End-of-log estimated LLM cost block, printed inside the job's stdio tee so
|
||||||
|
it lands at the tail of job.log. Both runners populate the same summary
|
||||||
|
fields (cost_usd / llm_calls / cached_calls / cost_by_stage / models_used).
|
||||||
|
Hybrid note: local text calls carry no usage accounting, so the dollar
|
||||||
|
figure covers OpenRouter calls only (local call counts still appear).
|
||||||
|
"""
|
||||||
|
if summary.get("cost_usd") is None and not summary.get("llm_calls"):
|
||||||
|
return
|
||||||
|
print(f"\n=== Estimated LLM cost ({label}) ===")
|
||||||
|
print(f" Total: ${summary.get('cost_usd', 0.0):.4f} across "
|
||||||
|
f"{summary.get('llm_calls', 0)} live calls "
|
||||||
|
f"({summary.get('cached_calls', 0)} cached at $0)")
|
||||||
|
for name, bucket in (summary.get("cost_by_stage") or {}).items():
|
||||||
|
print(f" {name}: ${bucket.get('usd', 0.0):.4f} "
|
||||||
|
f"({bucket.get('calls', 0)} live, {bucket.get('cached', 0)} cached)")
|
||||||
|
mu = summary.get("models_used") or {}
|
||||||
|
if mu.get("vision"):
|
||||||
|
print(f" Vision models: {', '.join(mu['vision'])}")
|
||||||
|
if mu.get("text_local"):
|
||||||
|
print(f" Text models (local): {', '.join(mu['text_local'])}")
|
||||||
|
if mu.get("text_cloud"):
|
||||||
|
print(f" Text models (cloud): {', '.join(mu['text_cloud'])}")
|
||||||
|
if mu.get("fallback_count"):
|
||||||
|
print(f" Local->cloud fallbacks: {mu['fallback_count']}")
|
||||||
|
if mu.get("text_local"):
|
||||||
|
print(" Note: local calls have no cost accounting; "
|
||||||
|
"the dollar total covers OpenRouter usage only.")
|
||||||
|
|
||||||
|
|
||||||
|
@contextmanager
|
||||||
|
def capture_job_output(job_id: str, out_dir: str) -> Iterator[None]:
|
||||||
|
"""
|
||||||
|
Re-open the stdio tee + raw LLM dump dir for post-run work that still
|
||||||
|
belongs to this job (review finalization): lines append to job.log and
|
||||||
|
the in-memory log, raw dumps resume under <out_dir>/llm_raw/. The tee is
|
||||||
|
process-global — same overlapping-job caveat as the main run.
|
||||||
|
"""
|
||||||
|
log_path = os.path.join(out_dir, "job.log")
|
||||||
|
if config.LLM_RAW_DUMP:
|
||||||
|
llm.set_raw_dump_dir(os.path.join(out_dir, "llm_raw"))
|
||||||
|
try:
|
||||||
|
with capture_stdio(lambda raw: _append_log(job_id, raw, log_path)):
|
||||||
|
yield
|
||||||
|
finally:
|
||||||
|
llm.set_raw_dump_dir(None)
|
||||||
|
|
||||||
|
|
||||||
def _notify(job_id: str, report: Dict, out_dir: str) -> None:
|
def _notify(job_id: str, report: Dict, out_dir: str) -> None:
|
||||||
email = _jobs[job_id].get("email")
|
email = _jobs[job_id].get("email")
|
||||||
if not email:
|
if not email:
|
||||||
@@ -197,11 +288,6 @@ def _notify_error(job_id: str) -> None:
|
|||||||
email = job.get("email")
|
email = job.get("email")
|
||||||
if not email:
|
if not email:
|
||||||
return
|
return
|
||||||
# Reuse the report mailer with a minimal error-shaped payload.
|
|
||||||
err_report = {
|
|
||||||
"source": job.get("source", ""),
|
|
||||||
"summary": {"conflicts_found": 0, "by_severity": {}, "disciplines": []},
|
|
||||||
}
|
|
||||||
try:
|
try:
|
||||||
from backend.email_sender import _smtp_ready, _send
|
from backend.email_sender import _smtp_ready, _send
|
||||||
from email.message import EmailMessage
|
from email.message import EmailMessage
|
||||||
@@ -216,6 +302,7 @@ def _notify_error(job_id: str) -> None:
|
|||||||
"Your conflict check did not complete.\n\n"
|
"Your conflict check did not complete.\n\n"
|
||||||
f"Drawing set: {job.get('source','')}\n"
|
f"Drawing set: {job.get('source','')}\n"
|
||||||
f"Error: {job.get('error','unknown')}\n\n"
|
f"Error: {job.get('error','unknown')}\n\n"
|
||||||
|
f"Review the run log at: {config.APP_BASE_URL.rstrip('/')}/?job={job_id}\n\n"
|
||||||
"Generated by Conflict Checker"
|
"Generated by Conflict Checker"
|
||||||
)
|
)
|
||||||
_send(msg)
|
_send(msg)
|
||||||
@@ -223,25 +310,60 @@ def _notify_error(job_id: str) -> None:
|
|||||||
print(f"[Email] Failed to send error notice: {e}")
|
print(f"[Email] Failed to send error notice: {e}")
|
||||||
|
|
||||||
|
|
||||||
|
def _log_from_disk(job_id: str) -> List[str]:
|
||||||
|
return read_log_file(os.path.join(config.OUTPUT_DIR, job_id, "job.log"))
|
||||||
|
|
||||||
|
|
||||||
|
def get_job_log(job_id: str) -> Optional[List[str]]:
|
||||||
|
"""Full job log lines, from memory or disk. None if job unknown."""
|
||||||
|
with _lock:
|
||||||
|
job = _jobs.get(job_id)
|
||||||
|
if job is not None:
|
||||||
|
return list(job.get("log") or [])
|
||||||
|
|
||||||
|
log = _log_from_disk(job_id)
|
||||||
|
# Job exists on disk if we have a log or a report artifact.
|
||||||
|
report_path = os.path.join(config.OUTPUT_DIR, job_id, "conflicts.json")
|
||||||
|
if log or os.path.isfile(report_path):
|
||||||
|
return log
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
def get_job(job_id: str) -> Optional[Dict]:
|
def get_job(job_id: str) -> Optional[Dict]:
|
||||||
"""Public job view. Includes the full report only when done.
|
"""Public job view. Includes the full report only when done.
|
||||||
|
|
||||||
Falls back to the on-disk conflicts.json when the job isn't in the
|
Falls back to the on-disk artifacts (conflicts.json / job.log) when the
|
||||||
in-memory registry (e.g. after a server restart).
|
job isn't in the in-memory registry (e.g. after a server restart).
|
||||||
"""
|
"""
|
||||||
with _lock:
|
with _lock:
|
||||||
job = _jobs.get(job_id)
|
job = _jobs.get(job_id)
|
||||||
if job:
|
if job:
|
||||||
return dict(job)
|
out = dict(job)
|
||||||
|
log = list(job.get("log") or [])
|
||||||
|
out["log_tail"] = log[-_LOG_TAIL:]
|
||||||
|
# Full log on terminal states so the UI can show it without a
|
||||||
|
# second fetch; keep polls light while running.
|
||||||
|
if out.get("status") in _TERMINAL_STATES:
|
||||||
|
out["log"] = log
|
||||||
|
else:
|
||||||
|
out.pop("log", None)
|
||||||
|
return out
|
||||||
|
|
||||||
# Try loading from disk
|
# Try loading from disk
|
||||||
report_path = os.path.join(config.OUTPUT_DIR, job_id, "conflicts.json")
|
report_path = os.path.join(config.OUTPUT_DIR, job_id, "conflicts.json")
|
||||||
if not os.path.isfile(report_path):
|
log = _log_from_disk(job_id)
|
||||||
|
if not os.path.isfile(report_path) and not log:
|
||||||
return None
|
return None
|
||||||
try:
|
try:
|
||||||
|
report = None
|
||||||
|
if os.path.isfile(report_path):
|
||||||
with open(report_path, encoding="utf-8") as f:
|
with open(report_path, encoding="utf-8") as f:
|
||||||
report = json.load(f)
|
report = json.load(f)
|
||||||
summary = report.get("summary", {})
|
summary = (report or {}).get("summary", {})
|
||||||
|
if report is None:
|
||||||
|
# Crashed before writing a report; the log is the only artifact.
|
||||||
|
status = "error"
|
||||||
|
else:
|
||||||
# Recover the job's real state: a job that stopped at the review gate
|
# Recover the job's real state: a job that stopped at the review gate
|
||||||
# must come back as needs_review (not done) or it can never finalize.
|
# must come back as needs_review (not done) or it can never finalize.
|
||||||
status = "needs_review" if summary.get("agent_status") == "needs_review" else "done"
|
status = "needs_review" if summary.get("agent_status") == "needs_review" else "done"
|
||||||
@@ -261,16 +383,20 @@ def get_job(job_id: str) -> Optional[Dict]:
|
|||||||
job = {
|
job = {
|
||||||
"job_id": job_id,
|
"job_id": job_id,
|
||||||
"status": status,
|
"status": status,
|
||||||
"source": meta.get("source") or report.get("source", os.path.basename(report_path)),
|
"source": meta.get("source") or (report or {}).get("source", os.path.basename(report_path)),
|
||||||
"email": meta.get("email"),
|
"email": meta.get("email"),
|
||||||
"project_input": report.get("project_input", {}),
|
"project_input": (report or {}).get("project_input", {}),
|
||||||
"text_local": summary.get("text_backend") == "local",
|
"text_local": summary.get("text_backend") == "local",
|
||||||
"pipeline_mode": meta.get("pipeline_mode") or summary.get("pipeline_mode", "classic"),
|
"pipeline_mode": meta.get("pipeline_mode") or summary.get("pipeline_mode", "classic"),
|
||||||
|
"vision_model": meta.get("vision_model"),
|
||||||
|
"text_model": meta.get("text_model"),
|
||||||
"stage": None,
|
"stage": None,
|
||||||
"created_at": os.path.getmtime(source_pdf) if os.path.isfile(source_pdf) else None,
|
"created_at": os.path.getmtime(source_pdf) if os.path.isfile(source_pdf) else None,
|
||||||
"finished_at": os.path.getmtime(report_path),
|
"finished_at": os.path.getmtime(report_path) if os.path.isfile(report_path) else None,
|
||||||
"report": report,
|
"report": report,
|
||||||
"error": None,
|
"error": None if report is not None else "Report missing; see job log",
|
||||||
|
"log": log,
|
||||||
|
"log_tail": log[-_LOG_TAIL:],
|
||||||
}
|
}
|
||||||
# Hydrate the in-memory registry so _set(...) transitions (reviewing,
|
# Hydrate the in-memory registry so _set(...) transitions (reviewing,
|
||||||
# finalizing, done) work for restart-recovered jobs.
|
# finalizing, done) work for restart-recovered jobs.
|
||||||
|
|||||||
+186
-17
@@ -8,6 +8,7 @@ with an image) and conflict-reasoning (Stage 3, with images) calls.
|
|||||||
"""
|
"""
|
||||||
|
|
||||||
import os
|
import os
|
||||||
|
import re
|
||||||
import json
|
import json
|
||||||
import hashlib
|
import hashlib
|
||||||
import threading
|
import threading
|
||||||
@@ -23,16 +24,115 @@ _clients: Dict[str, OpenAI] = {}
|
|||||||
# call_json when routing a no-image (text) call. Module-global mirrors the
|
# call_json when routing a no-image (text) call. Module-global mirrors the
|
||||||
# set_stage/cost pattern (single-user tool).
|
# set_stage/cost pattern (single-user tool).
|
||||||
_text_local = False
|
_text_local = False
|
||||||
|
# Per-job model overrides (user picked models in the UI). Same module-global
|
||||||
# Per-job model override (user picked a model in the UI). Same module-global
|
|
||||||
# pattern: set by the job runner before the pipeline starts, cleared after.
|
# pattern: set by the job runner before the pipeline starts, cleared after.
|
||||||
_model_override: Optional[str] = None
|
# Vision applies to image calls, text to no-image calls on OpenRouter (and to
|
||||||
|
# the local->cloud fallback). The LOCAL endpoint's model name is never taken
|
||||||
|
# from these overrides - hybrid local keeps LOCAL_TEXT_MODEL.
|
||||||
|
_vision_model_override: Optional[str] = None
|
||||||
|
_text_model_override: Optional[str] = None
|
||||||
|
|
||||||
|
# Per-job raw request/response dumps (missed/hallucinated-item debugging).
|
||||||
|
# Set by the job runner to <out_dir>/llm_raw at job start, cleared after.
|
||||||
|
# Same module-global pattern as the model overrides (single-user tool).
|
||||||
|
_raw_dump_dir: Optional[str] = None
|
||||||
|
_seq_lock = threading.Lock()
|
||||||
|
_call_seq = 0
|
||||||
|
|
||||||
|
|
||||||
def set_model_override(model: Optional[str]) -> None:
|
def set_raw_dump_dir(path: Optional[str]) -> None:
|
||||||
"""Override the model for all OpenRouter calls (vision + text), or None to clear."""
|
"""Point raw LLM request/response dumps at a directory. None disables."""
|
||||||
global _model_override
|
global _raw_dump_dir, _call_seq
|
||||||
_model_override = (model or "").strip() or None
|
with _seq_lock:
|
||||||
|
_raw_dump_dir = path
|
||||||
|
_call_seq = 0
|
||||||
|
|
||||||
|
|
||||||
|
def _next_seq() -> int:
|
||||||
|
global _call_seq
|
||||||
|
with _seq_lock:
|
||||||
|
_call_seq += 1
|
||||||
|
return _call_seq
|
||||||
|
|
||||||
|
|
||||||
|
def _summarize_parsed(parsed: Any) -> str:
|
||||||
|
"""
|
||||||
|
Compact digest of a parsed response for the job log. List values become
|
||||||
|
item counts (e.g. conflicts[3]) so a stage that returned nothing (miss)
|
||||||
|
or invented items (hallucination) is visible without opening the raw dump.
|
||||||
|
"""
|
||||||
|
if isinstance(parsed, list):
|
||||||
|
return f"list[{len(parsed)}]"
|
||||||
|
if not isinstance(parsed, dict):
|
||||||
|
return type(parsed).__name__
|
||||||
|
parts = []
|
||||||
|
for k, v in parsed.items():
|
||||||
|
if isinstance(v, list):
|
||||||
|
parts.append(f"{k}[{len(v)}]")
|
||||||
|
elif isinstance(v, dict):
|
||||||
|
parts.append(f"{k}{{{len(v)}}}")
|
||||||
|
else:
|
||||||
|
s = str(v)
|
||||||
|
parts.append(f"{k}={s[:40]!r}{'...' if len(s) > 40 else ''}")
|
||||||
|
out = ", ".join(parts)
|
||||||
|
return out[:300] + ("..." if len(out) > 300 else "")
|
||||||
|
|
||||||
|
|
||||||
|
def _dump_raw(seq: int, be: Dict[str, Any], usage_stage: str,
|
||||||
|
system_prompt: str, user_text: str,
|
||||||
|
images_b64: Optional[List[str]], max_tokens: int,
|
||||||
|
raw: str, parsed: Any,
|
||||||
|
finish_reason: Optional[str] = None) -> None:
|
||||||
|
"""Write the full request/response for one call to the job's llm_raw dir."""
|
||||||
|
if not _raw_dump_dir:
|
||||||
|
return
|
||||||
|
try:
|
||||||
|
os.makedirs(_raw_dump_dir, exist_ok=True)
|
||||||
|
safe_stage = re.sub(r"[^A-Za-z0-9_.-]+", "_", usage_stage)[:40]
|
||||||
|
safe_model = re.sub(r"[^A-Za-z0-9_.-]+", "_", be["model"])
|
||||||
|
payload = {
|
||||||
|
"seq": seq,
|
||||||
|
"stage": usage_stage,
|
||||||
|
"model": be["model"],
|
||||||
|
"backend": "local" if be.get("local") else "cloud",
|
||||||
|
"max_tokens": max_tokens,
|
||||||
|
# base64 image payloads deliberately excluded (multi-MB each);
|
||||||
|
# the count + source.pdf in the job dir identify what was sent.
|
||||||
|
"n_images": len(images_b64 or []),
|
||||||
|
"system_prompt": system_prompt,
|
||||||
|
"user_text": user_text,
|
||||||
|
"finish_reason": finish_reason,
|
||||||
|
"raw_response": raw,
|
||||||
|
"parsed": parsed,
|
||||||
|
}
|
||||||
|
path = os.path.join(_raw_dump_dir, f"{seq:04d}_{safe_stage}_{safe_model}.json")
|
||||||
|
tmp = f"{path}.{threading.get_ident()}.tmp"
|
||||||
|
with open(tmp, "w", encoding="utf-8") as f:
|
||||||
|
json.dump(payload, f, indent=2)
|
||||||
|
os.replace(tmp, path) # atomic so thread-pooled stages can't tear it
|
||||||
|
except OSError as e:
|
||||||
|
print(f"[LLM] raw dump failed: {e}")
|
||||||
|
|
||||||
|
|
||||||
|
def _log_call(seq: int, be: Dict[str, Any], usage_stage: str,
|
||||||
|
user_text: str, images_b64: Optional[List[str]],
|
||||||
|
raw: str, parsed: Any, usd: Optional[float],
|
||||||
|
cached: bool = False,
|
||||||
|
reasoning_tokens: Optional[int] = None) -> None:
|
||||||
|
"""One verbose per-call line for the job log (tee'd by job_log.py)."""
|
||||||
|
if not config.LLM_VERBOSE:
|
||||||
|
return
|
||||||
|
backend = "local" if be.get("local") else "cloud"
|
||||||
|
if cached:
|
||||||
|
cost_str = "cache hit"
|
||||||
|
elif usd is not None:
|
||||||
|
cost_str = f"${usd:.4f}"
|
||||||
|
else:
|
||||||
|
cost_str = "cost n/a"
|
||||||
|
think_str = f"think {reasoning_tokens}tk | " if reasoning_tokens is not None else ""
|
||||||
|
print(f"[LLM] #{seq:04d} {usage_stage} | {be['model']} ({backend}) | "
|
||||||
|
f"in {len(user_text)}ch+{len(images_b64 or [])}img | "
|
||||||
|
f"out {len(raw)}ch | {think_str}{cost_str} | {_summarize_parsed(parsed)}")
|
||||||
|
|
||||||
|
|
||||||
def set_text_backend(local: bool) -> None:
|
def set_text_backend(local: bool) -> None:
|
||||||
@@ -40,6 +140,13 @@ def set_text_backend(local: bool) -> None:
|
|||||||
global _text_local
|
global _text_local
|
||||||
_text_local = bool(local)
|
_text_local = bool(local)
|
||||||
|
|
||||||
|
|
||||||
|
def set_model_overrides(vision: Optional[str] = None, text: Optional[str] = None) -> None:
|
||||||
|
"""Per-run OpenRouter vision/text model picks. None/blank clears to defaults."""
|
||||||
|
global _vision_model_override, _text_model_override
|
||||||
|
_vision_model_override = (vision or "").strip() or None
|
||||||
|
_text_model_override = (text or "").strip() or None
|
||||||
|
|
||||||
# --- per-job cost accounting -------------------------------------------------
|
# --- per-job cost accounting -------------------------------------------------
|
||||||
# OpenRouter returns the real USD cost of each call when we request usage
|
# OpenRouter returns the real USD cost of each call when we request usage
|
||||||
# accounting. We accumulate it in a module-level counter; the runner resets it
|
# accounting. We accumulate it in a module-level counter; the runner resets it
|
||||||
@@ -125,9 +232,11 @@ def _add_cached() -> None:
|
|||||||
|
|
||||||
def _cache_key(model: str, system_prompt: str, user_text: str,
|
def _cache_key(model: str, system_prompt: str, user_text: str,
|
||||||
images_b64: Optional[List[str]], max_tokens: int,
|
images_b64: Optional[List[str]], max_tokens: int,
|
||||||
json_mode: bool) -> str:
|
json_mode: bool, reasoning_effort: Optional[str] = None,
|
||||||
|
reasoning_max_tokens: Optional[int] = None) -> str:
|
||||||
h = hashlib.sha256()
|
h = hashlib.sha256()
|
||||||
parts = [model, str(max_tokens), str(json_mode), system_prompt, user_text]
|
parts = [model, str(max_tokens), str(json_mode), str(reasoning_effort),
|
||||||
|
str(reasoning_max_tokens), system_prompt, user_text]
|
||||||
for b in (images_b64 or []):
|
for b in (images_b64 or []):
|
||||||
parts.append(b)
|
parts.append(b)
|
||||||
for p in parts:
|
for p in parts:
|
||||||
@@ -177,17 +286,22 @@ def _resolve_backend(has_images: bool, model_override: Optional[str]) -> Dict[st
|
|||||||
return {
|
return {
|
||||||
"base_url": config.LOCAL_BASE_URL,
|
"base_url": config.LOCAL_BASE_URL,
|
||||||
"api_key": config.LOCAL_API_KEY,
|
"api_key": config.LOCAL_API_KEY,
|
||||||
|
# Local model name comes from per-call args or LOCAL_TEXT_MODEL —
|
||||||
|
# never the UI's OpenRouter picks, which a local server won't serve.
|
||||||
"model": model_override or config.LOCAL_TEXT_MODEL or config.TEXT_MODEL,
|
"model": model_override or config.LOCAL_TEXT_MODEL or config.TEXT_MODEL,
|
||||||
"usage": False, # local has no OpenRouter usage accounting
|
"usage": False, # local has no OpenRouter usage accounting
|
||||||
"local": True,
|
"local": True,
|
||||||
}
|
}
|
||||||
# Vision, or text-on-OpenRouter (default / fallback). A per-job override
|
# Vision, or text-on-OpenRouter (default / fallback). A per-job override
|
||||||
# (user's UI model pick) wins over per-call and env defaults.
|
# (user's UI model pick) wins over per-call and env defaults.
|
||||||
default_model = config.MODEL if has_images else config.TEXT_MODEL
|
if has_images:
|
||||||
|
model = _vision_model_override or model_override or config.MODEL
|
||||||
|
else:
|
||||||
|
model = _text_model_override or model_override or config.TEXT_MODEL
|
||||||
return {
|
return {
|
||||||
"base_url": config.AI_BASE_URL,
|
"base_url": config.AI_BASE_URL,
|
||||||
"api_key": config.AI_API_KEY,
|
"api_key": config.AI_API_KEY,
|
||||||
"model": _model_override or model_override or default_model,
|
"model": model,
|
||||||
"usage": True,
|
"usage": True,
|
||||||
"local": False,
|
"local": False,
|
||||||
}
|
}
|
||||||
@@ -261,6 +375,21 @@ def _response_cost(response) -> Optional[float]:
|
|||||||
return float(cost) if isinstance(cost, (int, float)) else None
|
return float(cost) if isinstance(cost, (int, float)) else None
|
||||||
|
|
||||||
|
|
||||||
|
def _reasoning_tokens(response) -> Optional[int]:
|
||||||
|
"""Hidden thinking tokens for this call (OpenRouter usage details).
|
||||||
|
|
||||||
|
Direct evidence of whether the reasoning knob is working - without it,
|
||||||
|
thinking burn can only be inferred from char counts vs the token cap."""
|
||||||
|
try:
|
||||||
|
dump = response.model_dump()
|
||||||
|
except Exception:
|
||||||
|
return None
|
||||||
|
usage = dump.get("usage") or {}
|
||||||
|
details = usage.get("completion_tokens_details") or {}
|
||||||
|
n = details.get("reasoning_tokens")
|
||||||
|
return int(n) if isinstance(n, (int, float)) else None
|
||||||
|
|
||||||
|
|
||||||
def _parse(raw: str) -> Optional[Dict[str, Any]]:
|
def _parse(raw: str) -> Optional[Dict[str, Any]]:
|
||||||
"""Parse model JSON, falling back to truncation repair."""
|
"""Parse model JSON, falling back to truncation repair."""
|
||||||
try:
|
try:
|
||||||
@@ -280,6 +409,8 @@ def call_json(
|
|||||||
model: Optional[str] = None,
|
model: Optional[str] = None,
|
||||||
usage_tracker: Optional[Any] = None,
|
usage_tracker: Optional[Any] = None,
|
||||||
usage_stage: str = "?",
|
usage_stage: str = "?",
|
||||||
|
reasoning_effort: Optional[str] = None,
|
||||||
|
reasoning_max_tokens: Optional[int] = None,
|
||||||
) -> Optional[Dict[str, Any]]:
|
) -> Optional[Dict[str, Any]]:
|
||||||
"""
|
"""
|
||||||
Send one chat completion expecting a JSON object back.
|
Send one chat completion expecting a JSON object back.
|
||||||
@@ -287,15 +418,23 @@ def call_json(
|
|||||||
Uses the provider's JSON mode (response_format) so the model returns a bare
|
Uses the provider's JSON mode (response_format) so the model returns a bare
|
||||||
JSON object instead of prose/empty text, and recovers from max_tokens
|
JSON object instead of prose/empty text, and recovers from max_tokens
|
||||||
truncation. images_b64: optional base64 JPEGs attached as high-detail image
|
truncation. images_b64: optional base64 JPEGs attached as high-detail image
|
||||||
parts. Returns the parsed dict, or None on a hard failure (caller degrades).
|
parts. reasoning_effort: optional OpenRouter reasoning knob ("low"/"medium"/
|
||||||
|
"high") - keeps thinking models from burning the output budget on hidden
|
||||||
|
reasoning. reasoning_max_tokens: optional hard thinking-token budget
|
||||||
|
(OpenRouter reasoning max_tokens -> Gemini thinking_budget); stronger than
|
||||||
|
effort, and takes precedence when both are given. Returns the parsed dict,
|
||||||
|
or None on a hard failure (caller degrades).
|
||||||
"""
|
"""
|
||||||
has_images = bool(images_b64)
|
has_images = bool(images_b64)
|
||||||
be = _resolve_backend(has_images, model)
|
be = _resolve_backend(has_images, model)
|
||||||
|
seq = _next_seq()
|
||||||
|
|
||||||
cache_key = None
|
cache_key = None
|
||||||
if config.LLM_CACHE:
|
if config.LLM_CACHE:
|
||||||
cache_key = _cache_key(be["model"], system_prompt, user_text,
|
cache_key = _cache_key(be["model"], system_prompt, user_text,
|
||||||
images_b64, max_tokens, json_mode=True)
|
images_b64, max_tokens, json_mode=True,
|
||||||
|
reasoning_effort=reasoning_effort,
|
||||||
|
reasoning_max_tokens=reasoning_max_tokens)
|
||||||
hit = _cache_get(cache_key)
|
hit = _cache_get(cache_key)
|
||||||
if hit is not None:
|
if hit is not None:
|
||||||
_add_cached()
|
_add_cached()
|
||||||
@@ -303,6 +442,8 @@ def call_json(
|
|||||||
usage_tracker.record(
|
usage_tracker.record(
|
||||||
usage_stage, be["model"], cached=True, has_images=has_images
|
usage_stage, be["model"], cached=True, has_images=has_images
|
||||||
)
|
)
|
||||||
|
_log_call(seq, be, usage_stage, user_text, images_b64,
|
||||||
|
"", hit, None, cached=True)
|
||||||
return hit
|
return hit
|
||||||
|
|
||||||
content: List[Dict[str, Any]] = []
|
content: List[Dict[str, Any]] = []
|
||||||
@@ -323,10 +464,21 @@ def call_json(
|
|||||||
for attempt in range(2):
|
for attempt in range(2):
|
||||||
try:
|
try:
|
||||||
client = get_client(be["base_url"], be["api_key"])
|
client = get_client(be["base_url"], be["api_key"])
|
||||||
kwargs = dict(model=be["model"], messages=messages,
|
kwargs: Dict[str, Any] = dict(model=be["model"], messages=messages,
|
||||||
max_tokens=max_tokens, timeout=config.LLM_TIMEOUT)
|
max_tokens=max_tokens, timeout=config.LLM_TIMEOUT)
|
||||||
|
extra_body: Dict[str, Any] = {}
|
||||||
if be["usage"]:
|
if be["usage"]:
|
||||||
kwargs["extra_body"] = {"usage": {"include": True}}
|
extra_body["usage"] = {"include": True}
|
||||||
|
# OpenRouter reasoning knob; only sent to cloud backends (local
|
||||||
|
# servers reject unknown fields). A hard thinking budget wins over
|
||||||
|
# the vaguer effort tier - OpenRouter treats them as exclusive.
|
||||||
|
if not be.get("local"):
|
||||||
|
if reasoning_max_tokens:
|
||||||
|
extra_body["reasoning"] = {"max_tokens": int(reasoning_max_tokens)}
|
||||||
|
elif reasoning_effort:
|
||||||
|
extra_body["reasoning"] = {"effort": reasoning_effort}
|
||||||
|
if extra_body:
|
||||||
|
kwargs["extra_body"] = extra_body
|
||||||
if use_json_mode:
|
if use_json_mode:
|
||||||
kwargs["response_format"] = {"type": "json_object"}
|
kwargs["response_format"] = {"type": "json_object"}
|
||||||
response = client.chat.completions.create(**kwargs)
|
response = client.chat.completions.create(**kwargs)
|
||||||
@@ -337,12 +489,28 @@ def call_json(
|
|||||||
usage_tracker.record(
|
usage_tracker.record(
|
||||||
usage_stage, be["model"], usd=usd or 0.0, has_images=has_images
|
usage_stage, be["model"], usd=usd or 0.0, has_images=has_images
|
||||||
)
|
)
|
||||||
raw = _strip_fences(response.choices[0].message.content or "")
|
choice = response.choices[0]
|
||||||
|
finish_reason = getattr(choice, "finish_reason", None)
|
||||||
|
raw = _strip_fences(choice.message.content or "")
|
||||||
|
reasoning_tokens = _reasoning_tokens(response)
|
||||||
|
if finish_reason == "length":
|
||||||
|
# Hit max_tokens (thinking tokens included on reasoning models).
|
||||||
|
# Logged explicitly so silent truncation isn't mistaken for a
|
||||||
|
# parse problem; _parse below still salvages what it can.
|
||||||
|
print(f"[LLM] output hit max_tokens (finish_reason=length, "
|
||||||
|
f"{len(raw)}ch returned, "
|
||||||
|
f"thinking={reasoning_tokens}tk)")
|
||||||
parsed = _parse(raw)
|
parsed = _parse(raw)
|
||||||
if parsed is not None:
|
if parsed is not None:
|
||||||
_record_model(be, has_images, fell_back)
|
_record_model(be, has_images, fell_back)
|
||||||
if cache_key:
|
if cache_key:
|
||||||
_cache_set(cache_key, parsed)
|
_cache_set(cache_key, parsed)
|
||||||
|
_log_call(seq, be, usage_stage, user_text, images_b64,
|
||||||
|
raw, parsed, usd, reasoning_tokens=reasoning_tokens)
|
||||||
|
if config.LLM_RAW_DUMP:
|
||||||
|
_dump_raw(seq, be, usage_stage, system_prompt, user_text,
|
||||||
|
images_b64, max_tokens, raw, parsed,
|
||||||
|
finish_reason=finish_reason)
|
||||||
return parsed
|
return parsed
|
||||||
if attempt == 0:
|
if attempt == 0:
|
||||||
print("[LLM] JSON parse error (retrying)")
|
print("[LLM] JSON parse error (retrying)")
|
||||||
@@ -362,7 +530,8 @@ def call_json(
|
|||||||
_models["text_local"].add(be["model"])
|
_models["text_local"].add(be["model"])
|
||||||
_models["fallback_count"] += 1
|
_models["fallback_count"] += 1
|
||||||
be = {"base_url": config.AI_BASE_URL, "api_key": config.AI_API_KEY,
|
be = {"base_url": config.AI_BASE_URL, "api_key": config.AI_API_KEY,
|
||||||
"model": config.TEXT_MODEL, "usage": True, "local": False}
|
"model": _text_model_override or config.TEXT_MODEL,
|
||||||
|
"usage": True, "local": False}
|
||||||
cache_key = None # don't cache fallback under the local-model key
|
cache_key = None # don't cache fallback under the local-model key
|
||||||
fell_back = True
|
fell_back = True
|
||||||
continue
|
continue
|
||||||
|
|||||||
+91
-6
@@ -20,9 +20,10 @@ from fastapi.responses import HTMLResponse, JSONResponse, Response
|
|||||||
from fastapi.staticfiles import StaticFiles
|
from fastapi.staticfiles import StaticFiles
|
||||||
|
|
||||||
import backend.jobs
|
import backend.jobs
|
||||||
from backend import config
|
from backend import config, llm
|
||||||
from backend.jobs import PIPELINE_MODES, create_job, get_job, _set
|
from backend.jobs import PIPELINE_MODES, create_job, get_job, _set
|
||||||
from backend.pipeline.pdf_processor import render_page_jpeg
|
from backend.pipeline.pdf_processor import render_page_jpeg
|
||||||
|
from backend.review import chat as review_chat
|
||||||
from backend.review.feedback import decision_to_label, write_label
|
from backend.review.feedback import decision_to_label, write_label
|
||||||
from backend.review.finalizer import finalize_review
|
from backend.review.finalizer import finalize_review
|
||||||
from backend.review.store import ReviewStore
|
from backend.review.store import ReviewStore
|
||||||
@@ -35,6 +36,7 @@ _FRONTEND_DIR = os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", "
|
|||||||
@app.get("/health")
|
@app.get("/health")
|
||||||
def health():
|
def health():
|
||||||
return {"status": "ok", "model": config.MODEL,
|
return {"status": "ok", "model": config.MODEL,
|
||||||
|
"text_model": config.TEXT_MODEL,
|
||||||
"version": config.APP_VERSION,
|
"version": config.APP_VERSION,
|
||||||
"build": config.APP_BUILD,
|
"build": config.APP_BUILD,
|
||||||
"key_configured": bool(config.AI_API_KEY),
|
"key_configured": bool(config.AI_API_KEY),
|
||||||
@@ -43,13 +45,15 @@ def health():
|
|||||||
|
|
||||||
@app.get("/models")
|
@app.get("/models")
|
||||||
def list_models():
|
def list_models():
|
||||||
"""Available OpenRouter models with per-1M-token pricing for the UI picker."""
|
"""Vision/text OpenRouter model lists with pricing for the UI dropdowns."""
|
||||||
from backend.models import fetch_models
|
from backend.models import fetch_models, split_vision_text
|
||||||
models = fetch_models()
|
models = fetch_models()
|
||||||
if models is None:
|
if models is None:
|
||||||
raise HTTPException(status_code=502,
|
raise HTTPException(status_code=502,
|
||||||
detail="Could not fetch the model list from OpenRouter")
|
detail="Could not fetch the model list from OpenRouter")
|
||||||
return {"models": models, "default": config.MODEL, "default_text": config.TEXT_MODEL}
|
vision, text = split_vision_text(models)
|
||||||
|
return {"vision": vision, "text": text,
|
||||||
|
"defaults": {"vision": config.MODEL, "text": config.TEXT_MODEL}}
|
||||||
|
|
||||||
|
|
||||||
@app.get("/jobs/{job_id}/log")
|
@app.get("/jobs/{job_id}/log")
|
||||||
@@ -72,7 +76,8 @@ async def check(
|
|||||||
work_type: Optional[str] = Form(None),
|
work_type: Optional[str] = Form(None),
|
||||||
text_local: bool = Form(False),
|
text_local: bool = Form(False),
|
||||||
pipeline_mode: str = Form("classic"),
|
pipeline_mode: str = Form("classic"),
|
||||||
model: Optional[str] = Form(None),
|
vision_model: Optional[str] = Form(None),
|
||||||
|
text_model: Optional[str] = Form(None),
|
||||||
):
|
):
|
||||||
"""
|
"""
|
||||||
Accept a PDF, start a background conflict check, and return a job_id
|
Accept a PDF, start a background conflict check, and return a job_id
|
||||||
@@ -81,6 +86,9 @@ async def check(
|
|||||||
Optional intake fields (project_name/address/occupancy/work_type) feed the
|
Optional intake fields (project_name/address/occupancy/work_type) feed the
|
||||||
Stage 0 jurisdiction profile; anything left blank is derived from the cover
|
Stage 0 jurisdiction profile; anything left blank is derived from the cover
|
||||||
sheet.
|
sheet.
|
||||||
|
|
||||||
|
vision_model / text_model override the configured defaults for this run
|
||||||
|
(vision always OpenRouter; text follows the OpenRouter vs hybrid choice).
|
||||||
"""
|
"""
|
||||||
if not file.filename.lower().endswith(".pdf"):
|
if not file.filename.lower().endswith(".pdf"):
|
||||||
raise HTTPException(status_code=400, detail="Please upload a PDF.")
|
raise HTTPException(status_code=400, detail="Please upload a PDF.")
|
||||||
@@ -106,9 +114,12 @@ async def check(
|
|||||||
}.items()
|
}.items()
|
||||||
if v and v.strip()
|
if v and v.strip()
|
||||||
}
|
}
|
||||||
|
v_model = (vision_model or "").strip() or None
|
||||||
|
t_model = (text_model or "").strip() or None
|
||||||
job_id = create_job(tmp_path, source_filename=file.filename, email=email,
|
job_id = create_job(tmp_path, source_filename=file.filename, email=email,
|
||||||
project_input=project_input, text_local=text_local,
|
project_input=project_input, text_local=text_local,
|
||||||
pipeline_mode=pipeline_mode, model=model)
|
pipeline_mode=pipeline_mode, vision_model=v_model,
|
||||||
|
text_model=t_model)
|
||||||
return JSONResponse({
|
return JSONResponse({
|
||||||
"job_id": job_id,
|
"job_id": job_id,
|
||||||
"status": "queued",
|
"status": "queued",
|
||||||
@@ -176,7 +187,20 @@ def save_review_decisions(job_id: str, payload: dict):
|
|||||||
def _finalize_job(job_id: str, out_dir: str) -> None:
|
def _finalize_job(job_id: str, out_dir: str) -> None:
|
||||||
"""Background finalization: the ONE place the final report email may fire."""
|
"""Background finalization: the ONE place the final report email may fire."""
|
||||||
try:
|
try:
|
||||||
|
# Re-open the job's log tee + raw dump dir so the finalization LLM
|
||||||
|
# calls (clarification reruns, RFI drafting) land in job.log / llm_raw.
|
||||||
|
with backend.jobs.capture_job_output(job_id, out_dir):
|
||||||
|
print("\n=== Review finalization ===")
|
||||||
|
llm.reset_cost() # finalization-only cost attribution
|
||||||
report = finalize_review(job_id, out_dir)
|
report = finalize_review(job_id, out_dir)
|
||||||
|
cost = llm.get_cost()
|
||||||
|
backend.jobs._log_cost_summary({
|
||||||
|
"cost_usd": round(cost["usd"], 4),
|
||||||
|
"llm_calls": cost["calls"],
|
||||||
|
"cached_calls": cost.get("cached", 0),
|
||||||
|
"cost_by_stage": cost.get("by_stage", {}),
|
||||||
|
"models_used": cost.get("models", {}),
|
||||||
|
}, label="finalization")
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
try:
|
try:
|
||||||
_set(job_id, status="finalization_error", error=str(e),
|
_set(job_id, status="finalization_error", error=str(e),
|
||||||
@@ -229,6 +253,67 @@ def finalize_review_endpoint(job_id: str):
|
|||||||
return {"status": "finalizing"}
|
return {"status": "finalizing"}
|
||||||
|
|
||||||
|
|
||||||
|
# Statuses in which the review chat may be used. The chat is read-only, so it
|
||||||
|
# stays available after finalization - a reviewer often asks "why did it say
|
||||||
|
# that?" about a report they have already sent.
|
||||||
|
_CHAT_STATES = ("needs_review", "reviewing", "finalizing", "done", "finalization_error")
|
||||||
|
|
||||||
|
|
||||||
|
def _chat_out_dir(job_id: str) -> str:
|
||||||
|
"""Resolve a job's output dir for a chat request, or raise an HTTP error."""
|
||||||
|
job = get_job(job_id)
|
||||||
|
if not job:
|
||||||
|
raise HTTPException(status_code=404, detail="Job not found")
|
||||||
|
if job.get("status") not in _CHAT_STATES:
|
||||||
|
raise HTTPException(status_code=409, detail={
|
||||||
|
"detail": f"review chat is not available for a job in status {job.get('status')}",
|
||||||
|
})
|
||||||
|
return job.get("out_dir") or os.path.join(config.OUTPUT_DIR, job_id)
|
||||||
|
|
||||||
|
|
||||||
|
@app.post("/jobs/{job_id}/review-chat")
|
||||||
|
def review_chat_ask(job_id: str, payload: dict):
|
||||||
|
"""Ask one question about a finding, or about the run as a whole.
|
||||||
|
|
||||||
|
Read-only: this answers from the job's artifacts and appends to the chat
|
||||||
|
log. It never changes a finding, a decision, or the report.
|
||||||
|
"""
|
||||||
|
out_dir = _chat_out_dir(job_id)
|
||||||
|
store = ReviewStore(out_dir, create=False)
|
||||||
|
try:
|
||||||
|
turn = review_chat.ask(
|
||||||
|
job_id=job_id,
|
||||||
|
out_dir=out_dir,
|
||||||
|
question=payload.get("question"),
|
||||||
|
review_item_id=payload.get("review_item_id") or None,
|
||||||
|
queue=store.read_queue(),
|
||||||
|
decisions=store.read_decisions(),
|
||||||
|
)
|
||||||
|
except review_chat.ChatError as e:
|
||||||
|
raise HTTPException(status_code=422, detail=str(e))
|
||||||
|
except Exception as e:
|
||||||
|
# A failed model call is an upstream problem, not a bad request; the
|
||||||
|
# review screen shows it inline and the reviewer can retry.
|
||||||
|
raise HTTPException(status_code=502, detail=f"review chat failed: {e}")
|
||||||
|
return {"turn": turn}
|
||||||
|
|
||||||
|
|
||||||
|
@app.get("/jobs/{job_id}/review-chat")
|
||||||
|
def review_chat_history(job_id: str, review_item_id: Optional[str] = None):
|
||||||
|
"""Logged chat turns, oldest first. Without review_item_id, all threads."""
|
||||||
|
out_dir = _chat_out_dir(job_id)
|
||||||
|
turns = review_chat.read_log(out_dir, review_item_id=review_item_id)
|
||||||
|
return {"turns": turns, "enabled": config.ENABLE_REVIEW_CHAT}
|
||||||
|
|
||||||
|
|
||||||
|
@app.get("/jobs/{job_id}/review-chat/log")
|
||||||
|
def review_chat_log(job_id: str):
|
||||||
|
"""The chat log as a readable transcript: issue, questions, findings."""
|
||||||
|
out_dir = _chat_out_dir(job_id)
|
||||||
|
markdown = review_chat.render_log_markdown(review_chat.read_log(out_dir))
|
||||||
|
return Response(content=markdown, media_type="text/markdown; charset=utf-8")
|
||||||
|
|
||||||
|
|
||||||
@app.get("/jobs/{job_id}/sheet-image/{page}")
|
@app.get("/jobs/{job_id}/sheet-image/{page}")
|
||||||
def sheet_image(job_id: str, page: int):
|
def sheet_image(job_id: str, page: int):
|
||||||
"""Render one page of a completed job's source PDF as JPEG (sheet viewer)."""
|
"""Render one page of a completed job's source PDF as JPEG (sheet viewer)."""
|
||||||
|
|||||||
+27
-2
@@ -3,11 +3,13 @@ models.py - Fetch the available OpenRouter model list with pricing (cached).
|
|||||||
|
|
||||||
The /models endpoint is public (no API key needed). Results are normalized to
|
The /models endpoint is public (no API key needed). Results are normalized to
|
||||||
per-1M-token USD costs for display and cached in memory for an hour; callers
|
per-1M-token USD costs for display and cached in memory for an hour; callers
|
||||||
degrade gracefully when OpenRouter is unreachable.
|
degrade gracefully when OpenRouter is unreachable. Each entry also carries a
|
||||||
|
vision flag (accepts image input) so the UI can offer separate vision/text
|
||||||
|
model dropdowns.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
import time
|
import time
|
||||||
from typing import List, Optional
|
from typing import List, Optional, Tuple
|
||||||
|
|
||||||
import httpx
|
import httpx
|
||||||
|
|
||||||
@@ -25,6 +27,18 @@ def _per_mtok(rate) -> float:
|
|||||||
return 0.0
|
return 0.0
|
||||||
|
|
||||||
|
|
||||||
|
def _is_vision(item: dict) -> bool:
|
||||||
|
"""True when the model accepts image input and produces text output."""
|
||||||
|
arch = item.get("architecture") or {}
|
||||||
|
inputs = arch.get("input_modalities") or []
|
||||||
|
outputs = arch.get("output_modalities") or []
|
||||||
|
# Legacy string form: "text+image->text"
|
||||||
|
modality = (arch.get("modality") or "").lower()
|
||||||
|
has_image_in = ("image" in inputs) or ("image" in modality.split("->")[0])
|
||||||
|
has_text_out = ("text" in outputs) or ("->text" in modality) or (not outputs and not modality)
|
||||||
|
return has_image_in and has_text_out
|
||||||
|
|
||||||
|
|
||||||
def _fetch_openrouter_models() -> Optional[List[dict]]:
|
def _fetch_openrouter_models() -> Optional[List[dict]]:
|
||||||
"""Raw GET of the OpenRouter model list; None on any failure."""
|
"""Raw GET of the OpenRouter model list; None on any failure."""
|
||||||
try:
|
try:
|
||||||
@@ -55,6 +69,7 @@ def fetch_models(force: bool = False) -> Optional[List[dict]]:
|
|||||||
"prompt_usd_per_mtok": _per_mtok((item.get("pricing") or {}).get("prompt")),
|
"prompt_usd_per_mtok": _per_mtok((item.get("pricing") or {}).get("prompt")),
|
||||||
"completion_usd_per_mtok": _per_mtok((item.get("pricing") or {}).get("completion")),
|
"completion_usd_per_mtok": _per_mtok((item.get("pricing") or {}).get("completion")),
|
||||||
"context_length": item.get("context_length"),
|
"context_length": item.get("context_length"),
|
||||||
|
"vision": _is_vision(item),
|
||||||
}
|
}
|
||||||
for item in data
|
for item in data
|
||||||
if item.get("id")
|
if item.get("id")
|
||||||
@@ -63,3 +78,13 @@ def fetch_models(force: bool = False) -> Optional[List[dict]]:
|
|||||||
_cache["models"] = models
|
_cache["models"] = models
|
||||||
_cache["at"] = time.time()
|
_cache["at"] = time.time()
|
||||||
return models
|
return models
|
||||||
|
|
||||||
|
|
||||||
|
def split_vision_text(models: List[dict]) -> Tuple[List[dict], List[dict]]:
|
||||||
|
"""Partition the normalized catalog into (vision, text) lists for the UI.
|
||||||
|
|
||||||
|
Every catalog model takes text in/out, so vision models appear in both
|
||||||
|
lists (same dicts, pricing included).
|
||||||
|
"""
|
||||||
|
vision = [m for m in models if m.get("vision")]
|
||||||
|
return vision, list(models)
|
||||||
@@ -54,6 +54,8 @@ def slim_clusters(clusters: List[Dict]) -> List[Dict]:
|
|||||||
"location": c.get("location"),
|
"location": c.get("location"),
|
||||||
"disciplines": c.get("disciplines"),
|
"disciplines": c.get("disciplines"),
|
||||||
"kind": c.get("kind"),
|
"kind": c.get("kind"),
|
||||||
|
**({"disputed_attributes": c["disputed_attributes"]}
|
||||||
|
if c.get("disputed_attributes") else {}),
|
||||||
"assertions": [slim_assertion(a) for a in c.get("assertions", [])],
|
"assertions": [slim_assertion(a) for a in c.get("assertions", [])],
|
||||||
}
|
}
|
||||||
for c in clusters
|
for c in clusters
|
||||||
|
|||||||
@@ -32,6 +32,16 @@ def _evidence_block(cluster: Dict) -> str:
|
|||||||
f"{a.get('attribute','')} = {a.get('value','')} | "
|
f"{a.get('attribute','')} = {a.get('value','')} | "
|
||||||
f"\"{a.get('source_text','')}\""
|
f"\"{a.get('source_text','')}\""
|
||||||
)
|
)
|
||||||
|
disputes = cluster.get("disputed_attributes") or []
|
||||||
|
if disputes:
|
||||||
|
lines.append("")
|
||||||
|
for d in disputes:
|
||||||
|
lines.append(
|
||||||
|
"DISPUTED VALUE (possible extraction misread): "
|
||||||
|
f"attribute={d.get('attribute','')} "
|
||||||
|
f"values={' | '.join(d.get('values') or [])} "
|
||||||
|
f"(assertions {', '.join(d.get('assertion_ids') or [])})"
|
||||||
|
)
|
||||||
return "\n".join(lines)
|
return "\n".join(lines)
|
||||||
|
|
||||||
|
|
||||||
@@ -83,6 +93,8 @@ def _check_one(cluster: Dict, page_to_b64: Dict[int, str]) -> List[Dict]:
|
|||||||
user_text=user_text,
|
user_text=user_text,
|
||||||
images_b64=_images_for(cluster, page_to_b64),
|
images_b64=_images_for(cluster, page_to_b64),
|
||||||
max_tokens=config.REASON_MAX_TOKENS,
|
max_tokens=config.REASON_MAX_TOKENS,
|
||||||
|
reasoning_effort=config.EXTRACT_REASONING_EFFORT or None,
|
||||||
|
reasoning_max_tokens=config.EXTRACT_REASONING_MAX_TOKENS or None,
|
||||||
)
|
)
|
||||||
if isinstance(parsed, list):
|
if isinstance(parsed, list):
|
||||||
candidates = parsed
|
candidates = parsed
|
||||||
|
|||||||
@@ -27,6 +27,10 @@ def constructability_review(sheets: List[Dict], clusters: List[Dict],
|
|||||||
"assertions": dumps(slim_sheets(sheets)),
|
"assertions": dumps(slim_sheets(sheets)),
|
||||||
"clusters": dumps(slim_clusters(clusters)),
|
"clusters": dumps(slim_clusters(clusters)),
|
||||||
"conflicts": dumps(conflicts),
|
"conflicts": dumps(conflicts),
|
||||||
|
"disputes": dumps([
|
||||||
|
d for cluster in clusters
|
||||||
|
for d in (cluster.get("disputed_attributes") or [])
|
||||||
|
]),
|
||||||
},
|
},
|
||||||
max_tokens=config.CONSTRUCT_MAX_TOKENS,
|
max_tokens=config.CONSTRUCT_MAX_TOKENS,
|
||||||
)
|
)
|
||||||
|
|||||||
@@ -0,0 +1,88 @@
|
|||||||
|
"""
|
||||||
|
drawing_integrity.py - Per-sheet Drawing Integrity QA (LLM, classic pipeline).
|
||||||
|
|
||||||
|
The drawing-focused pass: reads ONE sheet's own extracted objects + sheet image
|
||||||
|
+ deterministic text layer and flags defects internal to that single sheet
|
||||||
|
(dangling detail/callout/keynote references, schedule-vs-plan/legend
|
||||||
|
disagreements on the same sheet, dimension strings that do not sum, missing
|
||||||
|
title-block/scale essentials, duplicate/inconsistent tags). It complements the
|
||||||
|
cross-sheet conflict checker; it never does code/ADA or cross-sheet
|
||||||
|
coordination. Emits the canonical issue schema. Returns [] on failure.
|
||||||
|
|
||||||
|
Gated by config.ENABLE_DRAWING_INTEGRITY. Runs sheets concurrently, one call
|
||||||
|
per sheet, skipping sheets below INTEGRITY_MIN_ASSERTIONS.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from concurrent.futures import ThreadPoolExecutor
|
||||||
|
from typing import Dict, List
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
from backend.agents.prompts import (
|
||||||
|
DRAWING_INTEGRITY_SYSTEM_PROMPT,
|
||||||
|
DRAWING_INTEGRITY_USER_PROMPT,
|
||||||
|
)
|
||||||
|
from backend.pipeline._serialize import dumps
|
||||||
|
from backend.pipeline._stage import call_stage, collect_list, validate_issue
|
||||||
|
|
||||||
|
|
||||||
|
def _sheet_meta(sheet: Dict) -> Dict:
|
||||||
|
return {
|
||||||
|
"sheet_number": sheet.get("sheet_number"),
|
||||||
|
"sheet_title": sheet.get("sheet_title"),
|
||||||
|
"discipline": sheet.get("discipline"),
|
||||||
|
"drawing_type": sheet.get("drawing_type"),
|
||||||
|
"level": sheet.get("level"),
|
||||||
|
"scale": sheet.get("scale"),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _review_sheet(sheet: Dict, page_to_b64: Dict, page_to_text: Dict) -> List[Dict]:
|
||||||
|
page_number = sheet.get("page_number")
|
||||||
|
assertions = (sheet.get("assertions") or [])[
|
||||||
|
:config.AGENT_INTEGRITY_MAX_ASSERTIONS
|
||||||
|
]
|
||||||
|
text_layer = (page_to_text.get(page_number) or "")[
|
||||||
|
:config.TEXT_LAYER_MAX_CHARS
|
||||||
|
]
|
||||||
|
image = page_to_b64.get(page_number)
|
||||||
|
images = [image][:config.AGENT_INTEGRITY_MAX_IMAGES] if image else []
|
||||||
|
parsed = call_stage(
|
||||||
|
DRAWING_INTEGRITY_SYSTEM_PROMPT,
|
||||||
|
DRAWING_INTEGRITY_USER_PROMPT,
|
||||||
|
subs={
|
||||||
|
"sheet_meta": dumps(_sheet_meta(sheet)),
|
||||||
|
"assertions": dumps(assertions),
|
||||||
|
"text_layer": text_layer,
|
||||||
|
},
|
||||||
|
images_b64=images,
|
||||||
|
max_tokens=config.INTEGRITY_MAX_TOKENS,
|
||||||
|
)
|
||||||
|
issues = collect_list(
|
||||||
|
parsed, "issues", lambda c: validate_issue(c, "drawing_integrity")
|
||||||
|
)
|
||||||
|
sheet_number = sheet.get("sheet_number")
|
||||||
|
for issue in issues:
|
||||||
|
if not issue.get("sheets") and sheet_number:
|
||||||
|
issue["sheets"] = [sheet_number]
|
||||||
|
return issues
|
||||||
|
|
||||||
|
|
||||||
|
def drawing_integrity_review(sheets: List[Dict], pages: List[Dict]) -> List[Dict]:
|
||||||
|
"""One LLM call per non-sparse sheet, run concurrently."""
|
||||||
|
if not config.ENABLE_DRAWING_INTEGRITY:
|
||||||
|
return []
|
||||||
|
page_to_b64 = {p["page_number"]: p.get("base64") for p in pages}
|
||||||
|
page_to_text = {p["page_number"]: p.get("text_layer") for p in pages}
|
||||||
|
targets = [
|
||||||
|
s for s in sheets
|
||||||
|
if len(s.get("assertions") or []) >= config.INTEGRITY_MIN_ASSERTIONS
|
||||||
|
]
|
||||||
|
issues: List[Dict] = []
|
||||||
|
if targets:
|
||||||
|
with ThreadPoolExecutor(max_workers=config.AGENT_INTEGRITY_CONCURRENCY) as pool:
|
||||||
|
for res in pool.map(
|
||||||
|
lambda s: _review_sheet(s, page_to_b64, page_to_text), targets
|
||||||
|
):
|
||||||
|
issues.extend(res)
|
||||||
|
print(f"[DrawingIntegrity] {len(issues)} issue(s) across {len(targets)} sheet(s)")
|
||||||
|
return issues
|
||||||
@@ -21,9 +21,18 @@ from backend.llm import call_json
|
|||||||
from backend.prompts import (
|
from backend.prompts import (
|
||||||
EXTRACTOR_SYSTEM_PROMPT,
|
EXTRACTOR_SYSTEM_PROMPT,
|
||||||
EXTRACTOR_USER_INSTRUCTION,
|
EXTRACTOR_USER_INSTRUCTION,
|
||||||
|
TEXT_STRUCTURING_SYSTEM_PROMPT,
|
||||||
|
TEXT_STRUCTURING_USER_INSTRUCTION,
|
||||||
DISCIPLINE_PREFIXES,
|
DISCIPLINE_PREFIXES,
|
||||||
ATTRIBUTE_VOCAB,
|
ATTRIBUTE_VOCAB,
|
||||||
)
|
)
|
||||||
|
from backend.text_coverage import (
|
||||||
|
_norm,
|
||||||
|
fallback_objects,
|
||||||
|
merge_objects,
|
||||||
|
recover_sheet_number,
|
||||||
|
text_coverage,
|
||||||
|
)
|
||||||
|
|
||||||
# prefix (upper) -> discipline, longest-prefix-first for greedy matching
|
# prefix (upper) -> discipline, longest-prefix-first for greedy matching
|
||||||
_PREFIX_TO_DISCIPLINE = sorted(
|
_PREFIX_TO_DISCIPLINE = sorted(
|
||||||
@@ -64,7 +73,8 @@ def discipline_from_sheet_number(sheet_number: Optional[str]) -> Optional[str]:
|
|||||||
return None
|
return None
|
||||||
|
|
||||||
|
|
||||||
def _is_grounded(value: str, source_text: str, graphical_basis: str = "") -> bool:
|
def _is_grounded(value: str, source_text: str, graphical_basis: str = "",
|
||||||
|
page_text: Optional[str] = None) -> bool:
|
||||||
"""
|
"""
|
||||||
Keep an object only if its primary value is supported by its source_text,
|
Keep an object only if its primary value is supported by its source_text,
|
||||||
OR it is a graphical object (has graphical_basis with no text to quote).
|
OR it is a graphical object (has graphical_basis with no text to quote).
|
||||||
@@ -72,6 +82,9 @@ def _is_grounded(value: str, source_text: str, graphical_basis: str = "") -> boo
|
|||||||
- If graphical_basis is set and source_text is absent, the object is valid.
|
- If graphical_basis is set and source_text is absent, the object is valid.
|
||||||
- If the value contains digits, every distinct digit-run must appear in
|
- If the value contains digits, every distinct digit-run must appear in
|
||||||
source_text (catches invented dimensions/counts/elevations).
|
source_text (catches invented dimensions/counts/elevations).
|
||||||
|
- Rescue tier: when page_text (the deterministic text layer) is given,
|
||||||
|
digit-runs absent from source_text but present in the page text are
|
||||||
|
still grounded - vision quoted imperfectly but the value is real.
|
||||||
- If the value has no digits, require some alphabetic-token overlap.
|
- If the value has no digits, require some alphabetic-token overlap.
|
||||||
"""
|
"""
|
||||||
# Graphical objects (no readable text on sheet) are always allowed through.
|
# Graphical objects (no readable text on sheet) are always allowed through.
|
||||||
@@ -85,7 +98,11 @@ def _is_grounded(value: str, source_text: str, graphical_basis: str = "") -> boo
|
|||||||
val_digits = set(_DIGITS_RE.findall(value))
|
val_digits = set(_DIGITS_RE.findall(value))
|
||||||
if val_digits:
|
if val_digits:
|
||||||
src_digits = set(_DIGITS_RE.findall(source_text))
|
src_digits = set(_DIGITS_RE.findall(source_text))
|
||||||
return val_digits.issubset(src_digits)
|
if val_digits.issubset(src_digits):
|
||||||
|
return True
|
||||||
|
if page_text:
|
||||||
|
return val_digits.issubset(set(_DIGITS_RE.findall(page_text)))
|
||||||
|
return False
|
||||||
|
|
||||||
# No digits: text-based grounding.
|
# No digits: text-based grounding.
|
||||||
val_norm = re.sub(r"[^a-z0-9]+", " ", value.lower()).strip()
|
val_norm = re.sub(r"[^a-z0-9]+", " ", value.lower()).strip()
|
||||||
@@ -109,7 +126,24 @@ def _primary_value(obj: Dict) -> str:
|
|||||||
or obj.get("name") or obj.get("tag") or "")
|
or obj.get("name") or obj.get("tag") or "")
|
||||||
|
|
||||||
|
|
||||||
def _normalize_sheet(parsed: Dict, page_number: int) -> Dict:
|
def _grounding_stamp(value: str, source_text: str,
|
||||||
|
page_text: Optional[str]) -> Optional[str]:
|
||||||
|
"""\"text_layer\" when the object survived only via the text-layer rescue
|
||||||
|
tier (digits absent from source_text but present in the page text)."""
|
||||||
|
if not page_text:
|
||||||
|
return None
|
||||||
|
val_digits = set(_DIGITS_RE.findall(str(value)))
|
||||||
|
if not val_digits:
|
||||||
|
return None
|
||||||
|
if val_digits.issubset(set(_DIGITS_RE.findall(source_text))):
|
||||||
|
return None
|
||||||
|
if val_digits.issubset(set(_DIGITS_RE.findall(page_text))):
|
||||||
|
return "text_layer"
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_sheet(parsed: Dict, page_number: int,
|
||||||
|
page_text: Optional[str] = None) -> Dict:
|
||||||
"""
|
"""
|
||||||
Validate + clean one parsed sheet result, attaching page_number and ids.
|
Validate + clean one parsed sheet result, attaching page_number and ids.
|
||||||
|
|
||||||
@@ -138,6 +172,9 @@ def _normalize_sheet(parsed: Dict, page_number: int) -> Dict:
|
|||||||
raw_objects = parsed.get("objects") or parsed.get("assertions") or []
|
raw_objects = parsed.get("objects") or parsed.get("assertions") or []
|
||||||
clean: List[Dict] = []
|
clean: List[Dict] = []
|
||||||
dropped = 0
|
dropped = 0
|
||||||
|
rescued = 0
|
||||||
|
unverified = 0
|
||||||
|
page_norm = _norm(page_text) if page_text else ""
|
||||||
|
|
||||||
for idx, obj in enumerate(raw_objects):
|
for idx, obj in enumerate(raw_objects):
|
||||||
if not isinstance(obj, dict):
|
if not isinstance(obj, dict):
|
||||||
@@ -149,9 +186,23 @@ def _normalize_sheet(parsed: Dict, page_number: int) -> Dict:
|
|||||||
# Derive a primary value for the grounding check
|
# Derive a primary value for the grounding check
|
||||||
primary_val = _primary_value(obj)
|
primary_val = _primary_value(obj)
|
||||||
|
|
||||||
if not _is_grounded(primary_val, source_text, graphical_basis):
|
if not _is_grounded(primary_val, source_text, graphical_basis,
|
||||||
|
page_text=page_text):
|
||||||
dropped += 1
|
dropped += 1
|
||||||
continue
|
continue
|
||||||
|
# Pre-set stamps (fallback/merge rungs) win; otherwise compute the
|
||||||
|
# text-layer rescue stamp.
|
||||||
|
grounding = obj.get("grounding") or _grounding_stamp(
|
||||||
|
primary_val, source_text, page_text)
|
||||||
|
if grounding == "text_layer":
|
||||||
|
rescued += 1
|
||||||
|
if not grounding and page_text and source_text:
|
||||||
|
# Vision-unverified: survived the digit guard, but the quoted
|
||||||
|
# source_text is not present in the deterministic text layer.
|
||||||
|
# Kept and stamped - the wave-5b verifier prioritizes these.
|
||||||
|
if _norm(str(source_text)) not in page_norm:
|
||||||
|
grounding = "vision_unverified"
|
||||||
|
unverified += 1
|
||||||
|
|
||||||
# --- location_key: new schema is richer; map to legacy shape + extras ---
|
# --- location_key: new schema is richer; map to legacy shape + extras ---
|
||||||
lk = obj.get("location_key")
|
lk = obj.get("location_key")
|
||||||
@@ -203,10 +254,14 @@ def _normalize_sheet(parsed: Dict, page_number: int) -> Dict:
|
|||||||
"object_attributes": attrs,
|
"object_attributes": attrs,
|
||||||
"graphical_basis": graphical_basis or None,
|
"graphical_basis": graphical_basis or None,
|
||||||
"review_uses": obj.get("review_uses") or [],
|
"review_uses": obj.get("review_uses") or [],
|
||||||
|
**({"grounding": grounding} if grounding else {}),
|
||||||
})
|
})
|
||||||
|
|
||||||
if dropped:
|
if dropped or rescued or unverified:
|
||||||
print(f"[Extract] Page {page_number} ({sheet_number}): dropped {dropped} ungrounded object(s)")
|
print(f"[Extract] Page {page_number} ({sheet_number}): "
|
||||||
|
f"dropped {dropped} ungrounded object(s)"
|
||||||
|
+ (f", rescued {rescued} via text layer" if rescued else "")
|
||||||
|
+ (f", {unverified} vision-unverified" if unverified else ""))
|
||||||
|
|
||||||
unresolved = parsed.get("unresolved_items") or []
|
unresolved = parsed.get("unresolved_items") or []
|
||||||
|
|
||||||
@@ -223,8 +278,41 @@ def _normalize_sheet(parsed: Dict, page_number: int) -> Dict:
|
|||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _text_layer_block(page: Dict) -> str:
|
||||||
|
"""
|
||||||
|
The TEXT LAYER block appended to the extractor instruction at call sites
|
||||||
|
(NOT a template placeholder - render() silently leaves missing keys as
|
||||||
|
literals). Empty string when the page has no usable text layer.
|
||||||
|
"""
|
||||||
|
text = (page.get("text_layer") or "").strip()
|
||||||
|
if not text:
|
||||||
|
return ""
|
||||||
|
return ("\n\nTEXT LAYER (authoritative for alphanumeric content — trust it "
|
||||||
|
"over the image for numbers, tags, and note text):\n"
|
||||||
|
+ text[:config.TEXT_LAYER_MAX_CHARS])
|
||||||
|
|
||||||
|
|
||||||
|
def _text_structuring_extract(page: Dict, sheet_hint: str = ""):
|
||||||
|
"""Rung 2 of the extraction ladder: text-only structuring call (no
|
||||||
|
image). The text layer is authoritative for alphanumeric content - the
|
||||||
|
model segments it instead of transcribing pixels, so vision misreads
|
||||||
|
are impossible on this rung."""
|
||||||
|
instruction = (TEXT_STRUCTURING_USER_INSTRUCTION
|
||||||
|
.replace("{sheet_hint}", sheet_hint or "")
|
||||||
|
.replace("{text_layer}",
|
||||||
|
(page.get("text_layer") or "")
|
||||||
|
[:config.TEXT_LAYER_MAX_CHARS]))
|
||||||
|
return call_json(
|
||||||
|
system_prompt=TEXT_STRUCTURING_SYSTEM_PROMPT,
|
||||||
|
user_text=instruction,
|
||||||
|
max_tokens=config.EXTRACT_MAX_TOKENS,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def _extract_one(page: Dict, sheet_hint: str = "") -> Dict:
|
def _extract_one(page: Dict, sheet_hint: str = "") -> Dict:
|
||||||
user_text = EXTRACTOR_USER_INSTRUCTION.replace("{sheet_hint}", sheet_hint)
|
page_text = page.get("text_layer")
|
||||||
|
user_text = (EXTRACTOR_USER_INSTRUCTION.replace("{sheet_hint}", sheet_hint)
|
||||||
|
+ _text_layer_block(page))
|
||||||
parsed = call_json(
|
parsed = call_json(
|
||||||
system_prompt=EXTRACTOR_SYSTEM_PROMPT,
|
system_prompt=EXTRACTOR_SYSTEM_PROMPT,
|
||||||
user_text=user_text,
|
user_text=user_text,
|
||||||
@@ -232,6 +320,8 @@ def _extract_one(page: Dict, sheet_hint: str = "") -> Dict:
|
|||||||
max_tokens=config.EXTRACT_MAX_TOKENS,
|
max_tokens=config.EXTRACT_MAX_TOKENS,
|
||||||
)
|
)
|
||||||
if not isinstance(parsed, dict):
|
if not isinstance(parsed, dict):
|
||||||
|
if not page_text:
|
||||||
|
# Scanned/raster page: vision-only, keep the legacy failure shape.
|
||||||
return {
|
return {
|
||||||
"page_number": page["page_number"],
|
"page_number": page["page_number"],
|
||||||
"sheet_number": None,
|
"sheet_number": None,
|
||||||
@@ -241,7 +331,63 @@ def _extract_one(page: Dict, sheet_hint: str = "") -> Dict:
|
|||||||
"scale": None,
|
"scale": None,
|
||||||
"assertions": [],
|
"assertions": [],
|
||||||
}
|
}
|
||||||
return _normalize_sheet(parsed, page["page_number"])
|
# Text-bearing page: climb the ladder instead of going dark.
|
||||||
|
parsed = {"sheet": {}, "objects": []}
|
||||||
|
|
||||||
|
sheet = _normalize_sheet(parsed, page["page_number"], page_text=page_text)
|
||||||
|
cov = text_coverage(page_text or "", sheet["assertions"])
|
||||||
|
sheet["coverage"] = cov
|
||||||
|
|
||||||
|
# Rung 2: text-only structuring when coverage is below floor. MERGE,
|
||||||
|
# never replace - vision objects (graphical_basis content exists only
|
||||||
|
# in the image) are kept; the text pass fills what vision missed.
|
||||||
|
if (page_text and config.EXTRACT_TEXT_RETRY_ENABLED
|
||||||
|
and cov["ratio"] < config.EXTRACT_COVERAGE_FLOOR):
|
||||||
|
print(f"[Extract] Page {page['page_number']}: coverage "
|
||||||
|
f"{cov['ratio']:.0%} < floor - text-only structuring pass")
|
||||||
|
parsed2 = _text_structuring_extract(page, sheet_hint)
|
||||||
|
if isinstance(parsed2, dict):
|
||||||
|
sheet2 = _normalize_sheet(parsed2, page["page_number"],
|
||||||
|
page_text=page_text)
|
||||||
|
before = len(sheet["assertions"])
|
||||||
|
sheet["assertions"] = merge_objects(sheet["assertions"],
|
||||||
|
sheet2["assertions"])
|
||||||
|
for key in ("sheet_number", "sheet_title", "discipline",
|
||||||
|
"level", "scale", "drawing_type"):
|
||||||
|
if not sheet.get(key) and sheet2.get(key):
|
||||||
|
sheet[key] = sheet2[key]
|
||||||
|
cov = text_coverage(page_text, sheet["assertions"])
|
||||||
|
sheet["coverage"] = cov
|
||||||
|
print(f"[Extract] Page {page['page_number']}: merged "
|
||||||
|
f"{len(sheet['assertions']) - before} text-structured "
|
||||||
|
f"object(s), coverage now {cov['ratio']:.0%}")
|
||||||
|
|
||||||
|
# Rung 3: deterministic fallback - a dark text-bearing sheet is
|
||||||
|
# impossible. Stubs are deduped against earlier rungs.
|
||||||
|
if (page_text and config.EXTRACT_FALLBACK_ENABLED
|
||||||
|
and cov["ratio"] < config.EXTRACT_COVERAGE_FLOOR):
|
||||||
|
stubs = fallback_objects(page_text, page["page_number"],
|
||||||
|
config.EXTRACT_FALLBACK_MAX_OBJECTS)
|
||||||
|
stubs = _normalize_sheet({"sheet": {}, "objects": stubs},
|
||||||
|
page["page_number"],
|
||||||
|
page_text=page_text)["assertions"]
|
||||||
|
before = len(sheet["assertions"])
|
||||||
|
sheet["assertions"] = merge_objects(sheet["assertions"], stubs)
|
||||||
|
print(f"[Extract] Page {page['page_number']}: fallback merged "
|
||||||
|
f"{len(sheet['assertions']) - before} text-layer stub(s)")
|
||||||
|
sheet["coverage"] = text_coverage(page_text, sheet["assertions"])
|
||||||
|
|
||||||
|
# Identity recovery: never leave a text-bearing page sheet-less.
|
||||||
|
if not sheet.get("sheet_number") and page_text:
|
||||||
|
recovered = recover_sheet_number(page_text)
|
||||||
|
if recovered:
|
||||||
|
sheet["sheet_number"] = recovered
|
||||||
|
sheet["discipline"] = (discipline_from_sheet_number(recovered)
|
||||||
|
or sheet.get("discipline") or "Unknown")
|
||||||
|
print(f"[Extract] Page {page['page_number']}: sheet number "
|
||||||
|
f"recovered from text layer -> {recovered}")
|
||||||
|
|
||||||
|
return sheet
|
||||||
|
|
||||||
|
|
||||||
def extract_assertions(pages: List[Dict], on_progress=None) -> List[Dict]:
|
def extract_assertions(pages: List[Dict], on_progress=None) -> List[Dict]:
|
||||||
|
|||||||
@@ -1,5 +1,4 @@
|
|||||||
"""
|
"""report.py - Stage 4: assemble the final report.
|
||||||
report.py - Stage 4: assemble the final report.
|
|
||||||
|
|
||||||
Produces a single JSON object (also the web API payload) and a human-readable
|
Produces a single JSON object (also the web API payload) and a human-readable
|
||||||
Markdown summary grouped by severity.
|
Markdown summary grouped by severity.
|
||||||
@@ -8,6 +7,37 @@ Markdown summary grouped by severity.
|
|||||||
from typing import List, Dict
|
from typing import List, Dict
|
||||||
from datetime import datetime, timezone
|
from datetime import datetime, timezone
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
|
||||||
|
|
||||||
|
def _extraction_coverage(sheets: List[Dict]) -> Dict | None:
|
||||||
|
"""Summarize per-sheet extraction coverage for the report summary.
|
||||||
|
|
||||||
|
Sheets may carry a ``coverage`` dict (``total_lines``/``covered_lines``/
|
||||||
|
``ratio``) attached during wave-1 extraction. Older paths and scanned
|
||||||
|
pages have none; when no sheet is measured, return None so callers can
|
||||||
|
omit the key entirely.
|
||||||
|
"""
|
||||||
|
measured = [s for s in sheets if isinstance(s.get("coverage"), dict)]
|
||||||
|
if not measured:
|
||||||
|
return None
|
||||||
|
floor = getattr(config, "EXTRACT_COVERAGE_FLOOR", 0.6)
|
||||||
|
return {
|
||||||
|
"pages_measured": len(measured),
|
||||||
|
"pages_below_floor": [
|
||||||
|
s.get("page_number") for s in measured
|
||||||
|
if s["coverage"].get("ratio", 0.0) < floor
|
||||||
|
],
|
||||||
|
"fallback_pages": [
|
||||||
|
s.get("page_number") for s in measured
|
||||||
|
if any(a.get("grounding") == "text_layer_fallback"
|
||||||
|
for a in s.get("assertions", []))
|
||||||
|
],
|
||||||
|
"mean_ratio": round(
|
||||||
|
sum(s["coverage"].get("ratio", 0.0) for s in measured)
|
||||||
|
/ len(measured), 3),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
def build_report(conflicts: List[Dict], sheets: List[Dict], clusters: List[Dict],
|
def build_report(conflicts: List[Dict], sheets: List[Dict], clusters: List[Dict],
|
||||||
source: str = "") -> Dict:
|
source: str = "") -> Dict:
|
||||||
@@ -19,10 +49,7 @@ def build_report(conflicts: List[Dict], sheets: List[Dict], clusters: List[Dict]
|
|||||||
by_cat[c["category"]] = by_cat.get(c["category"], 0) + 1
|
by_cat[c["category"]] = by_cat.get(c["category"], 0) + 1
|
||||||
|
|
||||||
disciplines = sorted({s["discipline"] for s in sheets if s.get("discipline")})
|
disciplines = sorted({s["discipline"] for s in sheets if s.get("discipline")})
|
||||||
return {
|
summary = {
|
||||||
"source": source,
|
|
||||||
"generated_at": datetime.now(timezone.utc).isoformat(),
|
|
||||||
"summary": {
|
|
||||||
"sheets_analyzed": len(sheets),
|
"sheets_analyzed": len(sheets),
|
||||||
"disciplines": disciplines,
|
"disciplines": disciplines,
|
||||||
"assertions_extracted": sum(len(s.get("assertions", [])) for s in sheets),
|
"assertions_extracted": sum(len(s.get("assertions", [])) for s in sheets),
|
||||||
@@ -30,7 +57,14 @@ def build_report(conflicts: List[Dict], sheets: List[Dict], clusters: List[Dict]
|
|||||||
"conflicts_found": len(conflicts),
|
"conflicts_found": len(conflicts),
|
||||||
"by_severity": by_sev,
|
"by_severity": by_sev,
|
||||||
"by_category": by_cat,
|
"by_category": by_cat,
|
||||||
},
|
}
|
||||||
|
coverage = _extraction_coverage(sheets)
|
||||||
|
if coverage is not None:
|
||||||
|
summary["extraction_coverage"] = coverage
|
||||||
|
return {
|
||||||
|
"source": source,
|
||||||
|
"generated_at": datetime.now(timezone.utc).isoformat(),
|
||||||
|
"summary": summary,
|
||||||
"conflicts": conflicts,
|
"conflicts": conflicts,
|
||||||
"sheets": [
|
"sheets": [
|
||||||
{
|
{
|
||||||
|
|||||||
@@ -27,22 +27,28 @@ from typing import Dict, Optional, Callable
|
|||||||
|
|
||||||
from backend.pipeline.pdf_processor import convert_pdf_to_images
|
from backend.pipeline.pdf_processor import convert_pdf_to_images
|
||||||
from backend.pipeline.extractor import extract_assertions
|
from backend.pipeline.extractor import extract_assertions
|
||||||
|
from backend.sheet_reconcile import declared_sheet_list, reconcile_sheets
|
||||||
|
from backend.text_layer import attach_text_layers, coverage_gaps
|
||||||
from backend.pipeline.sheet_index import classify_sheets, derive_project_meta_from_cover
|
from backend.pipeline.sheet_index import classify_sheets, derive_project_meta_from_cover
|
||||||
from backend.pipeline.jurisdiction import run_jurisdiction
|
from backend.pipeline.jurisdiction import run_jurisdiction
|
||||||
from backend.pipeline.normalizer import normalize_assertions, build_project_intelligence
|
from backend.pipeline.normalizer import normalize_assertions, build_project_intelligence
|
||||||
from backend.pipeline.clusterer import cluster_by_location
|
from backend.pipeline.clusterer import cluster_by_location
|
||||||
from backend.pipeline.llm_clusterer import cluster_by_location_llm
|
from backend.pipeline.llm_clusterer import cluster_by_location_llm
|
||||||
from backend import config
|
from backend import config
|
||||||
|
from backend.agents.disputes import annotate_clusters
|
||||||
from backend.pipeline.conflict_checker import check_conflicts
|
from backend.pipeline.conflict_checker import check_conflicts
|
||||||
from backend.pipeline.qaqc_review import senior_review
|
from backend.pipeline.qaqc_review import senior_review
|
||||||
from backend.pipeline.code_review import code_review
|
from backend.pipeline.code_review import code_review
|
||||||
|
from backend.pipeline.drawing_integrity import drawing_integrity_review
|
||||||
from backend.pipeline.constructability import constructability_review
|
from backend.pipeline.constructability import constructability_review
|
||||||
from backend.pipeline.validator import dedup_validate
|
from backend.pipeline.validator import dedup_validate
|
||||||
from backend.pipeline.risk import score_and_prioritize
|
from backend.pipeline.risk import score_and_prioritize
|
||||||
from backend.pipeline.rfi import generate_rfis
|
from backend.pipeline.rfi import generate_rfis
|
||||||
from backend.pipeline.report import build_report, to_markdown
|
from backend.pipeline.report import build_report, to_markdown
|
||||||
from backend.pipeline._stage import validate_issue
|
from backend.pipeline._stage import validate_issue
|
||||||
from backend.llm import reset_cost, get_cost, set_stage, set_text_backend
|
from backend.llm import (
|
||||||
|
reset_cost, get_cost, set_stage, set_text_backend, set_model_overrides,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def run_pipeline(
|
def run_pipeline(
|
||||||
@@ -52,6 +58,8 @@ def run_pipeline(
|
|||||||
project_input: Optional[Dict] = None,
|
project_input: Optional[Dict] = None,
|
||||||
source_name: Optional[str] = None,
|
source_name: Optional[str] = None,
|
||||||
text_local: bool = False,
|
text_local: bool = False,
|
||||||
|
vision_model: Optional[str] = None,
|
||||||
|
text_model: Optional[str] = None,
|
||||||
) -> Dict:
|
) -> Dict:
|
||||||
"""
|
"""
|
||||||
Run the full QAQC pipeline on one PDF and return the report dict.
|
Run the full QAQC pipeline on one PDF and return the report dict.
|
||||||
@@ -59,6 +67,9 @@ def run_pipeline(
|
|||||||
project_input: optional intake fields (project_name, address, occupancy,
|
project_input: optional intake fields (project_name, address, occupancy,
|
||||||
work_type). Cover-sheet-derived values fill any gaps; intake fields win.
|
work_type). Cover-sheet-derived values fill any gaps; intake fields win.
|
||||||
|
|
||||||
|
vision_model / text_model: optional per-run OpenRouter (or local text)
|
||||||
|
model overrides from the UI. Blank/None keeps config defaults.
|
||||||
|
|
||||||
If out_dir is given, writes conflicts.json, report.md, and the intermediate
|
If out_dir is given, writes conflicts.json, report.md, and the intermediate
|
||||||
artifacts (assertions.json, clusters.json, and one json per QAQC stage).
|
artifacts (assertions.json, clusters.json, and one json per QAQC stage).
|
||||||
"""
|
"""
|
||||||
@@ -70,12 +81,50 @@ def run_pipeline(
|
|||||||
|
|
||||||
reset_cost()
|
reset_cost()
|
||||||
set_text_backend(text_local)
|
set_text_backend(text_local)
|
||||||
|
set_model_overrides(vision_model, text_model)
|
||||||
|
if vision_model or text_model:
|
||||||
|
print(f"[Runner] model overrides: vision={vision_model or '(default)'} "
|
||||||
|
f"text={text_model or '(default)'}")
|
||||||
|
|
||||||
|
try:
|
||||||
|
return _run_stages(
|
||||||
|
pdf_path, out_dir, stage, project_input, source_name, text_local,
|
||||||
|
)
|
||||||
|
finally:
|
||||||
|
# Don't leak per-run picks into a later overlapping/CLI call.
|
||||||
|
set_model_overrides(None, None)
|
||||||
|
set_text_backend(False)
|
||||||
|
|
||||||
|
|
||||||
|
def _run_stages(
|
||||||
|
pdf_path: str,
|
||||||
|
out_dir: Optional[str],
|
||||||
|
stage: Callable[[str], None],
|
||||||
|
project_input: Optional[Dict],
|
||||||
|
source_name: Optional[str],
|
||||||
|
text_local: bool,
|
||||||
|
) -> Dict:
|
||||||
stage("PDF -> images")
|
stage("PDF -> images")
|
||||||
pages = convert_pdf_to_images(pdf_path)
|
pages = convert_pdf_to_images(pdf_path)
|
||||||
|
text_dir = os.path.join(out_dir, "text") if out_dir else None
|
||||||
|
attach_text_layers(pdf_path, pages, text_dir=text_dir)
|
||||||
|
|
||||||
stage("Extract assertions")
|
stage("Extract assertions")
|
||||||
sheets = extract_assertions(pages)
|
sheets = extract_assertions(pages)
|
||||||
|
coverage_gaps(pages, sheets) # classic: log-only recall signal
|
||||||
|
|
||||||
|
# Deterministic reconciliation: cover-sheet index vs identified sheets.
|
||||||
|
page_to_text = {p["page_number"]: p.get("text_layer") for p in pages}
|
||||||
|
sheet_recon = reconcile_sheets(sheets, declared_sheet_list(page_to_text))
|
||||||
|
if sheet_recon["declared_total"]:
|
||||||
|
print(f"[SheetIndex] cover declares {sheet_recon['declared_total']} "
|
||||||
|
f"sheets; {sheet_recon['found_total']} identified in set")
|
||||||
|
if sheet_recon["declared_not_in_set"]:
|
||||||
|
print(f"[SheetIndex] declared but not in set: "
|
||||||
|
f"{', '.join(sheet_recon['declared_not_in_set'][:20])}")
|
||||||
|
if sheet_recon["in_set_not_declared"]:
|
||||||
|
print(f"[SheetIndex] in set but not declared: "
|
||||||
|
f"{', '.join(sheet_recon['in_set_not_declared'][:20])}")
|
||||||
|
|
||||||
stage("Classify sheet index")
|
stage("Classify sheet index")
|
||||||
sheet_index = classify_sheets(sheets)
|
sheet_index = classify_sheets(sheets)
|
||||||
@@ -100,21 +149,33 @@ def run_pipeline(
|
|||||||
else:
|
else:
|
||||||
clusters = cluster_by_location(sheets)
|
clusters = cluster_by_location(sheets)
|
||||||
|
|
||||||
|
disputed_count = annotate_clusters(clusters)
|
||||||
|
if disputed_count:
|
||||||
|
print(f"[Cluster] {disputed_count} cluster(s) carry disputed extracted values")
|
||||||
|
|
||||||
stage("Reason over clusters (conflicts)")
|
stage("Reason over clusters (conflicts)")
|
||||||
conflicts = check_conflicts(clusters, pages)
|
conflicts = check_conflicts(clusters, pages)
|
||||||
|
|
||||||
stage("Full-set QAQC review")
|
stage("Full-set QAQC review")
|
||||||
qaqc_issues = senior_review(sheets, clusters, conflicts, sheet_index)
|
qaqc_issues = senior_review(sheets, clusters, conflicts, sheet_index)
|
||||||
|
|
||||||
|
if config.ENABLE_CODE_REVIEW:
|
||||||
stage("Code / ADA review")
|
stage("Code / ADA review")
|
||||||
code_issues = code_review(jurisdiction, sheets, sheet_index)
|
code_issues = code_review(jurisdiction, sheets, sheet_index)
|
||||||
|
else:
|
||||||
|
print("[Code] code/ADA review disabled (ENABLE_CODE_REVIEW=0)")
|
||||||
|
code_issues = []
|
||||||
|
|
||||||
|
stage("Drawing integrity (per-sheet QA)")
|
||||||
|
integrity_issues = drawing_integrity_review(sheets, pages)
|
||||||
|
|
||||||
stage("Constructability review")
|
stage("Constructability review")
|
||||||
construct_issues = constructability_review(sheets, clusters, conflicts)
|
construct_issues = constructability_review(sheets, clusters, conflicts)
|
||||||
|
|
||||||
stage("Validate & deduplicate")
|
stage("Validate & deduplicate")
|
||||||
conflict_issues = [v for v in (validate_issue(c, "conflict") for c in conflicts) if v]
|
conflict_issues = [v for v in (validate_issue(c, "conflict") for c in conflicts) if v]
|
||||||
all_issues = conflict_issues + qaqc_issues + code_issues + construct_issues
|
all_issues = (conflict_issues + integrity_issues + qaqc_issues
|
||||||
|
+ code_issues + construct_issues)
|
||||||
validated = dedup_validate(all_issues)
|
validated = dedup_validate(all_issues)
|
||||||
|
|
||||||
stage("Risk scoring & prioritization")
|
stage("Risk scoring & prioritization")
|
||||||
@@ -129,6 +190,7 @@ def run_pipeline(
|
|||||||
report["project_input"] = merged_input
|
report["project_input"] = merged_input
|
||||||
report["jurisdiction"] = jurisdiction
|
report["jurisdiction"] = jurisdiction
|
||||||
report["sheet_index"] = sheet_index
|
report["sheet_index"] = sheet_index
|
||||||
|
report["sheet_reconciliation"] = sheet_recon
|
||||||
report["project_intelligence"] = project_intel
|
report["project_intelligence"] = project_intel
|
||||||
report["validated_issues"] = prioritized
|
report["validated_issues"] = prioritized
|
||||||
report["rfis"] = rfis
|
report["rfis"] = rfis
|
||||||
@@ -136,6 +198,7 @@ def run_pipeline(
|
|||||||
"conflicts": len(conflicts),
|
"conflicts": len(conflicts),
|
||||||
"qaqc": len(qaqc_issues),
|
"qaqc": len(qaqc_issues),
|
||||||
"code": len(code_issues),
|
"code": len(code_issues),
|
||||||
|
"drawing_integrity": len(integrity_issues),
|
||||||
"constructability": len(construct_issues),
|
"constructability": len(construct_issues),
|
||||||
"validated": len(validated),
|
"validated": len(validated),
|
||||||
"rfis": len(rfis),
|
"rfis": len(rfis),
|
||||||
@@ -160,6 +223,7 @@ def run_pipeline(
|
|||||||
_dump(out_dir, "project_intelligence.json", project_intel)
|
_dump(out_dir, "project_intelligence.json", project_intel)
|
||||||
_dump(out_dir, "qaqc_issues.json", qaqc_issues)
|
_dump(out_dir, "qaqc_issues.json", qaqc_issues)
|
||||||
_dump(out_dir, "code_issues.json", code_issues)
|
_dump(out_dir, "code_issues.json", code_issues)
|
||||||
|
_dump(out_dir, "drawing_integrity.json", integrity_issues)
|
||||||
_dump(out_dir, "constructability.json", construct_issues)
|
_dump(out_dir, "constructability.json", construct_issues)
|
||||||
_dump(out_dir, "validated_issues.json", prioritized)
|
_dump(out_dir, "validated_issues.json", prioritized)
|
||||||
_dump(out_dir, "rfis.json", rfis)
|
_dump(out_dir, "rfis.json", rfis)
|
||||||
|
|||||||
+68
-11
@@ -230,6 +230,7 @@ Rules you must never break:
|
|||||||
- Every object must include source_text copied verbatim from the sheet whenever text is available.
|
- Every object must include source_text copied verbatim from the sheet whenever text is available.
|
||||||
- If the object is graphical and has no text, describe it visually and mark confidence low or medium.
|
- If the object is graphical and has no text, describe it visually and mark confidence low or medium.
|
||||||
- Preserve tags, marks, room numbers, sheet numbers, detail references, and abbreviations exactly as shown.
|
- Preserve tags, marks, room numbers, sheet numbers, detail references, and abbreviations exactly as shown.
|
||||||
|
TEXT LAYER GROUNDING: when a TEXT LAYER block is present in the user message, it is the sheet's deterministic PDF text layer and is authoritative for alphanumeric content (counts, dimensions, member tags, note text). Trust it over your reading of the image for numbers, tags, and note text; quote source_text from it verbatim. Use the image for geometry, symbols, linework, and anything absent from the text layer.
|
||||||
- Use null when information is not determinable.
|
- Use null when information is not determinable.
|
||||||
- Keep objects atomic.
|
- Keep objects atomic.
|
||||||
- Use plain ASCII only.
|
- Use plain ASCII only.
|
||||||
@@ -280,6 +281,30 @@ If the sheet has no extractable objects, return an empty objects array.
|
|||||||
Optional sheet hint: {sheet_hint}"""
|
Optional sheet hint: {sheet_hint}"""
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Stage 2b - text-only structuring (extraction retry ladder, rung 2)
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
TEXT_STRUCTURING_SYSTEM_PROMPT = """You are a construction document structuring engine.
|
||||||
|
You receive the deterministic text layer extracted from one drawing sheet. It is complete and authoritative.
|
||||||
|
Your ONLY job is to segment it into structured objects. You are NOT reading an image. You must NOT invent, complete, or correct any text.
|
||||||
|
Rules:
|
||||||
|
- Every numbered note, schedule row, callout, tag, legend entry, and title-block field becomes its own object.
|
||||||
|
- source_text must be copied VERBATIM from the input, character-for-character. Never paraphrase.
|
||||||
|
- Cover the ENTIRE input. Omitting a note is a failure. When unsure of an object's type, use general_note with confidence low.
|
||||||
|
- Numbers, model numbers, dimensions, and tags must appear in source_text exactly as in the input.
|
||||||
|
Respond only with valid JSON."""
|
||||||
|
|
||||||
|
TEXT_STRUCTURING_USER_INSTRUCTION = """Segment this sheet's text layer into structured construction objects.
|
||||||
|
Every note, schedule row, callout, tag, and title-block field in the text layer must become an object - omit nothing.
|
||||||
|
Respond ONLY with a valid JSON object - no markdown fences:
|
||||||
|
{ "sheet": { "sheet_number": "string or null", "sheet_title": "string or null", "discipline": "string or null", "drawing_type": "string or null", "level": "string or null", "scale": "string or null" }, "objects": [ { "object_id": "string", "object_type": "room | door | window | wall | finish | ceiling | dimension | grid | callout | keynote | general_note | equipment | plumbing_fixture | mechanical_equipment | electrical_device | lighting_fixture | structural_element | schedule_reference | symbol | abbreviation", "category": "architectural | structural | mechanical | electrical | plumbing | code | general", "tag": "string or null", "name": "string or null", "description": "string or null", "attributes": { "attribute_name": "attribute_value" }, "location_key": { "room_number": "string or null", "grid": "string or null", "detail_reference": "string or null" }, "source_text": "VERBATIM text copied from the input", "graphical_basis": null, "review_uses": [ "schedule_comparison", "cross_discipline_coordination", "code_review", "constructability_review" ], "confidence": "high | medium | low" } ], "unresolved_items": [] }
|
||||||
|
Optional sheet hint: {sheet_hint}
|
||||||
|
|
||||||
|
TEXT LAYER (segment ALL of it):
|
||||||
|
{text_layer}"""
|
||||||
|
|
||||||
|
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
# Stage 3a - assertion normalization (WIRED: normalizer.py)
|
# Stage 3a - assertion normalization (WIRED: normalizer.py)
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
@@ -385,24 +410,24 @@ Normalized assertions: {normalized_assertions}"""
|
|||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
CONFLICT_SYSTEM_PROMPT = """You are a Senior Architect and construction-drawing coordination reviewer doing a back-check of a drawing set BEFORE it is issued for bid, permit, or construction.
|
CONFLICT_SYSTEM_PROMPT = """You are a Senior Architect and construction-drawing coordination reviewer doing a back-check of a drawing set BEFORE it is issued for bid, permit, or construction.
|
||||||
You are given clustered facts that multiple disciplines have asserted about the same location or element.
|
You are given clustered facts asserted about the same location or element. Those facts may come from MULTIPLE disciplines, from a SINGLE discipline across several sheets, or from ONE sheet (plan vs schedule vs detail vs keynote on that sheet).
|
||||||
Decide whether these disciplines GENUINELY CONTRADICT each other - the kind of issue a human coordinator would issue as a QAQC comment or RFI before the set goes out.
|
Decide whether these facts GENUINELY CONTRADICT each other - the kind of issue a human coordinator would issue as a QAQC comment or RFI before the set goes out. A contradiction between two facts is a conflict whether or not the two facts come from different disciplines.
|
||||||
You are NOT performing code review in this stage. You are NOT checking ADA in this stage. You are NOT estimating cost or scope. You are NOT rewriting the drawings.
|
You are NOT performing code review in this stage. You are NOT checking ADA in this stage. You are NOT estimating cost or scope. You are NOT rewriting the drawings.
|
||||||
What IS a conflict:
|
What IS a conflict:
|
||||||
- Two disciplines state different values for the same physical quantity at the same place.
|
- Two facts state different values for the same physical quantity at the same place (across disciplines, across sheets of one discipline, or on the same sheet).
|
||||||
- An element is shown in different locations by different disciplines.
|
- An element is shown in different locations by different facts.
|
||||||
- A schedule disagrees with what is drawn on the plan.
|
- A schedule disagrees with what is drawn on the plan (even on the same sheet).
|
||||||
- A detail disagrees with the plan.
|
- A detail disagrees with the plan.
|
||||||
- A keynote disagrees with a schedule, plan, or detail.
|
- A keynote or general note disagrees with a schedule, plan, detail, or legend - including a keynote/legend mismatch on a single sheet.
|
||||||
|
- A callout, detail reference, section marker, or tag references something that does not exist (a dangling reference).
|
||||||
|
- The same room, door, equipment, wall, or utility is labeled or dimensioned inconsistently across sheets or within one sheet.
|
||||||
- An element required by one discipline has no counterpart where another discipline should show it.
|
- An element required by one discipline has no counterpart where another discipline should show it.
|
||||||
- A duct, pipe, conduit, or piece of equipment conflicts with structure, ceiling height, rated wall, or required clearance.
|
- A duct, pipe, conduit, or piece of equipment conflicts with structure, ceiling height, rated wall, or required clearance.
|
||||||
- Equipment shown by one discipline lacks required power, plumbing, ventilation, access, or support in another discipline.
|
- Equipment shown by one discipline lacks required power, plumbing, ventilation, access, or support in another discipline.
|
||||||
- Demolition drawings remove something that new work drawings keep without explanation.
|
- Demolition drawings remove something that new work drawings keep without explanation.
|
||||||
- A callout, keynote, or tag references something that does not exist.
|
|
||||||
- The same room, door, equipment, wall, or utility is labeled inconsistently across sheets.
|
|
||||||
What is NOT a conflict:
|
What is NOT a conflict:
|
||||||
- Two disciplines describing different, compatible aspects of the same place.
|
- Two facts describing different, compatible aspects of the same place.
|
||||||
- A value shown on one discipline and simply not repeated on another, unless that discipline is expected to show it.
|
- A value shown once and simply not repeated elsewhere, unless another sheet or discipline is expected to show it.
|
||||||
- Rounding or representation differences that resolve to the same real value.
|
- Rounding or representation differences that resolve to the same real value.
|
||||||
- A possible code issue.
|
- A possible code issue.
|
||||||
- A design preference.
|
- A design preference.
|
||||||
@@ -410,6 +435,9 @@ What is NOT a conflict:
|
|||||||
- Anything not supported with drawing evidence.
|
- Anything not supported with drawing evidence.
|
||||||
Be conservative:
|
Be conservative:
|
||||||
- Only flag genuine disagreements.
|
- Only flag genuine disagreements.
|
||||||
|
- When a value is marked DISPUTED (possible extraction misread), verify it against
|
||||||
|
the sheet images before relying on either reading; if the images do not resolve
|
||||||
|
it, do not assert a conflict from one reading alone.
|
||||||
- A clean cluster with no contradiction must return an empty conflicts array.
|
- A clean cluster with no contradiction must return an empty conflicts array.
|
||||||
- missing_element requires evidence that another discipline would reasonably be expected to show the missing item.
|
- missing_element requires evidence that another discipline would reasonably be expected to show the missing item.
|
||||||
For each conflict:
|
For each conflict:
|
||||||
@@ -447,6 +475,28 @@ Clustered assertions (evidence):
|
|||||||
{evidence}"""
|
{evidence}"""
|
||||||
|
|
||||||
|
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
# Wave 5b - evidence verification (vision fact-check of cited sheet text)
|
||||||
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
|
VERIFY_SYSTEM_PROMPT = """You are a meticulous construction document checker verifying machine-extracted evidence against the actual drawing sheet images.
|
||||||
|
For each evidence item you are given the sheet it was extracted from and the verbatim text the extractor claims appears there.
|
||||||
|
Judge each item against the images:
|
||||||
|
- confirmed: the text (or an obvious equivalent) appears on the cited sheet and means what the finding claims.
|
||||||
|
- corrected: the sheet shows a DIFFERENT value than the extracted text. Give the actual verbatim text.
|
||||||
|
- not_found: nothing like the extracted text appears on the cited sheet.
|
||||||
|
Be strict about numbers, quantities, and member sizes: "(2) 2x6" and "(5) 2x6" are different values. HSS16x4 and HSS16x16 are different values.
|
||||||
|
Use plain ASCII only.
|
||||||
|
Respond only with valid JSON."""
|
||||||
|
|
||||||
|
VERIFY_USER_INSTRUCTION = """Verify this finding's evidence against the attached sheet images.
|
||||||
|
Respond ONLY with a valid JSON object - no markdown fences, no explanation:
|
||||||
|
{ "verdicts": [ { "sheet": "string", "source_text": "the evidence text judged", "verdict": "confirmed | corrected | not_found", "actual_text": "verbatim sheet text when corrected, else null", "notes": "string or null" } ] }
|
||||||
|
Finding: {finding}
|
||||||
|
TEXT LAYER (deterministic page text extracted from the PDF - an oracle for alphanumeric content such as counts, dimensions, and member tags; when it disagrees with the extracted evidence, trust it and cite it as actual_text):
|
||||||
|
{text_layer}"""
|
||||||
|
|
||||||
|
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
# Stage 6 - senior architect full-set QAQC review (NOT WIRED YET)
|
# Stage 6 - senior architect full-set QAQC review (NOT WIRED YET)
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
@@ -545,6 +595,12 @@ Flag:
|
|||||||
Rules:
|
Rules:
|
||||||
- Only flag issues supported by drawing evidence.
|
- Only flag issues supported by drawing evidence.
|
||||||
- Be specific about the location and why it is a constructability risk.
|
- Be specific about the location and why it is a constructability risk.
|
||||||
|
- Assertions are machine-extracted from sheet images and may contain misread values,
|
||||||
|
especially quantities and member sizes (e.g. "(2) 2x6" vs "(5) 2x6").
|
||||||
|
- When the cluster lists disputed_attributes, or two evidence items disagree on a
|
||||||
|
numeric value, do NOT assert a buildability conclusion from one reading. Report the
|
||||||
|
ambiguity itself (category "detail_gap", confidence "low") and state that the value
|
||||||
|
needs verification against the sheet.
|
||||||
- Use plain ASCII only.
|
- Use plain ASCII only.
|
||||||
Respond only with valid JSON."""
|
Respond only with valid JSON."""
|
||||||
|
|
||||||
@@ -554,7 +610,8 @@ Respond ONLY with a valid JSON object - no markdown fences, no explanation:
|
|||||||
If no constructability issues are found, return: { "issues": [] }
|
If no constructability issues are found, return: { "issues": [] }
|
||||||
Extracted assertions: {assertions}
|
Extracted assertions: {assertions}
|
||||||
Clusters: {clusters}
|
Clusters: {clusters}
|
||||||
Cross-discipline conflicts already found: {conflicts}"""
|
Cross-discipline conflicts already found: {conflicts}
|
||||||
|
Disputed extracted values in this cluster (possible vision misreads - treat as unverified): {disputes}"""
|
||||||
|
|
||||||
|
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
|
|||||||
@@ -0,0 +1,368 @@
|
|||||||
|
"""Review-screen chat: ask the run why it concluded something.
|
||||||
|
|
||||||
|
Read-only by construction. The chat reads job artifacts, calls one LLM, and
|
||||||
|
appends a log record; it never mutates findings, review decisions, or the
|
||||||
|
report, and the prompt forbids it from emitting code or config changes.
|
||||||
|
|
||||||
|
Every turn is logged twice, on purpose:
|
||||||
|
|
||||||
|
- ``<out_dir>/review/chat_log.jsonl`` - job-local, the auditable record of what
|
||||||
|
was asked about which finding and what came back.
|
||||||
|
- ``REVIEW_FEEDBACK_DIR/chat_turns.jsonl`` - cross-job, append-only, so the
|
||||||
|
corrections a reviewer makes in conversation ("that is not a floor drain, it
|
||||||
|
is a power floor box") accumulate somewhere a future run can be primed from.
|
||||||
|
Nothing reads this yet; writing it is what makes that possible later.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import uuid
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from typing import Any, Dict, List, Optional
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
from backend.llm import call_json
|
||||||
|
from backend.pipeline._serialize import dumps
|
||||||
|
from backend.review.chat_context import build_context
|
||||||
|
from backend.review.chat_prompts import (
|
||||||
|
REVIEW_CHAT_SYSTEM_PROMPT,
|
||||||
|
REVIEW_CHAT_USER_PROMPT,
|
||||||
|
)
|
||||||
|
from backend.review.feedback import append_shared_feedback
|
||||||
|
|
||||||
|
_ANSWERABLE = {"yes", "partial", "no"}
|
||||||
|
_ASSESSMENTS = {"looks_supported", "looks_unsupported", "cannot_tell", "not_applicable"}
|
||||||
|
_CONFIDENCE = {"high", "medium", "low"}
|
||||||
|
_MAX_FINDINGS = 12
|
||||||
|
_MAX_EVIDENCE = 12
|
||||||
|
|
||||||
|
|
||||||
|
class ChatError(Exception):
|
||||||
|
"""Raised for a caller-fixable problem (bad question, chat disabled)."""
|
||||||
|
|
||||||
|
|
||||||
|
def _now() -> str:
|
||||||
|
return datetime.now(timezone.utc).isoformat()
|
||||||
|
|
||||||
|
|
||||||
|
def _one_of(value: Any, allowed: set, default: str) -> str:
|
||||||
|
text = str(value or "").strip().lower()
|
||||||
|
return text if text in allowed else default
|
||||||
|
|
||||||
|
|
||||||
|
def _clean_question(raw: Any) -> str:
|
||||||
|
question = str(raw or "").strip()
|
||||||
|
if not question:
|
||||||
|
raise ChatError("question is required")
|
||||||
|
if len(question) > config.REVIEW_CHAT_MAX_QUESTION_CHARS:
|
||||||
|
raise ChatError(
|
||||||
|
f"question is too long (max {config.REVIEW_CHAT_MAX_QUESTION_CHARS} characters)")
|
||||||
|
return question
|
||||||
|
|
||||||
|
|
||||||
|
def _log_path(out_dir: str) -> str:
|
||||||
|
return os.path.join(out_dir, "review", "chat_log.jsonl")
|
||||||
|
|
||||||
|
|
||||||
|
def read_log(out_dir: str, review_item_id: Optional[str] = None,
|
||||||
|
scope_only: bool = False) -> List[Dict[str, Any]]:
|
||||||
|
"""Chat turns for this job, oldest first.
|
||||||
|
|
||||||
|
``review_item_id`` filters to one finding's thread; with ``scope_only`` and
|
||||||
|
no id, returns only the run-scope turns. Corrupt lines are skipped rather
|
||||||
|
than failing the read - a truncated log must not hide the rest.
|
||||||
|
"""
|
||||||
|
path = _log_path(out_dir)
|
||||||
|
if not os.path.isfile(path):
|
||||||
|
return []
|
||||||
|
turns: List[Dict[str, Any]] = []
|
||||||
|
try:
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
for line in f:
|
||||||
|
line = line.strip()
|
||||||
|
if not line:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
turn = json.loads(line)
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
continue
|
||||||
|
if not isinstance(turn, dict):
|
||||||
|
continue
|
||||||
|
if review_item_id is not None:
|
||||||
|
if turn.get("review_item_id") != review_item_id:
|
||||||
|
continue
|
||||||
|
elif scope_only and turn.get("review_item_id") is not None:
|
||||||
|
continue
|
||||||
|
turns.append(turn)
|
||||||
|
except OSError:
|
||||||
|
return []
|
||||||
|
return turns
|
||||||
|
|
||||||
|
|
||||||
|
def _append_log(out_dir: str, turn: Dict[str, Any]) -> None:
|
||||||
|
"""Append one turn as a JSON line; never raises on I/O failure."""
|
||||||
|
try:
|
||||||
|
os.makedirs(os.path.join(out_dir, "review"), exist_ok=True)
|
||||||
|
with open(_log_path(out_dir), "a", encoding="utf-8") as f:
|
||||||
|
f.write(json.dumps(turn) + "\n")
|
||||||
|
except OSError as e:
|
||||||
|
print(f"[ReviewChat] chat log write failed: {e}")
|
||||||
|
|
||||||
|
|
||||||
|
def _issue_snapshot(item: Optional[Dict]) -> Optional[Dict[str, Any]]:
|
||||||
|
"""The issue as it stood when asked about - the log's 'issue in question'.
|
||||||
|
|
||||||
|
Copied rather than referenced by id so the log stays readable after
|
||||||
|
finalization renumbers or suppresses the finding.
|
||||||
|
"""
|
||||||
|
if not item:
|
||||||
|
return None
|
||||||
|
payload = item.get("payload") or {}
|
||||||
|
if item.get("kind") == "clean_cluster":
|
||||||
|
return {
|
||||||
|
"review_item_id": item.get("review_item_id"),
|
||||||
|
"kind": item.get("kind"),
|
||||||
|
"cluster_key": payload.get("key"),
|
||||||
|
"location": payload.get("location"),
|
||||||
|
"disciplines": payload.get("disciplines"),
|
||||||
|
}
|
||||||
|
return {
|
||||||
|
"review_item_id": item.get("review_item_id"),
|
||||||
|
"kind": item.get("kind"),
|
||||||
|
"issue_id": payload.get("issue_id"),
|
||||||
|
"source_stage": payload.get("source_stage"),
|
||||||
|
"category": payload.get("category"),
|
||||||
|
"severity": payload.get("severity"),
|
||||||
|
"confidence": payload.get("confidence"),
|
||||||
|
"location": payload.get("location"),
|
||||||
|
"disciplines": payload.get("disciplines"),
|
||||||
|
"sheets": payload.get("sheets"),
|
||||||
|
"description": payload.get("description"),
|
||||||
|
"blocking": item.get("blocking"),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _history_block(turns: List[Dict[str, Any]]) -> str:
|
||||||
|
if not turns:
|
||||||
|
return ""
|
||||||
|
recent = turns[-config.REVIEW_CHAT_HISTORY_TURNS:]
|
||||||
|
lines = ["Earlier turns in this thread (oldest first):"]
|
||||||
|
for turn in recent:
|
||||||
|
lines.append(f"Reviewer: {turn.get('question') or ''}")
|
||||||
|
lines.append(f"You: {turn.get('answer') or ''}")
|
||||||
|
lines.append("")
|
||||||
|
return "\n".join(lines)
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_answer(parsed: Optional[Dict]) -> Optional[Dict[str, Any]]:
|
||||||
|
"""Coerce the model's JSON into the log/API shape, or None if unusable."""
|
||||||
|
if not isinstance(parsed, dict):
|
||||||
|
return None
|
||||||
|
answer = str(parsed.get("answer") or "").strip()
|
||||||
|
if not answer:
|
||||||
|
return None
|
||||||
|
findings = [
|
||||||
|
str(item).strip()
|
||||||
|
for item in (parsed.get("findings") or [])
|
||||||
|
if isinstance(item, (str, int, float)) and str(item).strip()
|
||||||
|
][:_MAX_FINDINGS]
|
||||||
|
evidence = []
|
||||||
|
for item in (parsed.get("evidence_cited") or [])[:_MAX_EVIDENCE]:
|
||||||
|
if not isinstance(item, dict):
|
||||||
|
continue
|
||||||
|
evidence.append({
|
||||||
|
"artifact": str(item.get("artifact") or "").strip() or None,
|
||||||
|
"sheet": item.get("sheet"),
|
||||||
|
"quote": str(item.get("quote") or "").strip() or None,
|
||||||
|
"why_it_matters": str(item.get("why_it_matters") or "").strip() or None,
|
||||||
|
})
|
||||||
|
correction = parsed.get("suggested_category_correction")
|
||||||
|
correction = str(correction).strip() if correction else ""
|
||||||
|
missing = parsed.get("missing_information")
|
||||||
|
return {
|
||||||
|
"answer": answer,
|
||||||
|
"findings": findings,
|
||||||
|
"evidence_cited": evidence,
|
||||||
|
"answerable": _one_of(parsed.get("answerable"), _ANSWERABLE, "partial"),
|
||||||
|
"missing_information": str(missing).strip() if missing else None,
|
||||||
|
"assessment_of_finding": _one_of(parsed.get("assessment_of_finding"),
|
||||||
|
_ASSESSMENTS, "cannot_tell"),
|
||||||
|
# The feedback signal: a reviewer correcting a misidentification in
|
||||||
|
# conversation ("that is a power floor box") lands here as structured
|
||||||
|
# data instead of dying in free text.
|
||||||
|
"suggested_category_correction": correction or None,
|
||||||
|
"confidence": _one_of(parsed.get("confidence"), _CONFIDENCE, "low"),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _feedback_record(turn: Dict[str, Any]) -> Dict[str, Any]:
|
||||||
|
"""Cross-job roll-up of one turn: metadata + the correction signal.
|
||||||
|
|
||||||
|
Mirrors the privacy stance of the decision labels - no images, no raw sheet
|
||||||
|
dumps. The question and answer ARE carried, because a chat turn without its
|
||||||
|
question is not usable as feedback; keep this store job-internal.
|
||||||
|
"""
|
||||||
|
issue = turn.get("issue") or {}
|
||||||
|
return {
|
||||||
|
"kind": "review_chat_turn",
|
||||||
|
"turn_id": turn.get("turn_id"),
|
||||||
|
"job_id": turn.get("job_id"),
|
||||||
|
"created_at": turn.get("created_at"),
|
||||||
|
"review_item_id": turn.get("review_item_id"),
|
||||||
|
"scope": turn.get("scope"),
|
||||||
|
"issue_id": issue.get("issue_id"),
|
||||||
|
"source_stage": issue.get("source_stage"),
|
||||||
|
"category": issue.get("category"),
|
||||||
|
"severity": issue.get("severity"),
|
||||||
|
"confidence": issue.get("confidence"),
|
||||||
|
"sheets": issue.get("sheets"),
|
||||||
|
"question": turn.get("question"),
|
||||||
|
"answer": turn.get("answer"),
|
||||||
|
"findings": turn.get("findings"),
|
||||||
|
"assessment_of_finding": turn.get("assessment_of_finding"),
|
||||||
|
"suggested_category_correction": turn.get("suggested_category_correction"),
|
||||||
|
"answerable": turn.get("answerable"),
|
||||||
|
"model": turn.get("model"),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def ask(job_id: str, out_dir: str, question: str,
|
||||||
|
review_item_id: Optional[str] = None,
|
||||||
|
queue: Optional[List[Dict]] = None,
|
||||||
|
decisions: Optional[Dict[str, Dict]] = None) -> Dict[str, Any]:
|
||||||
|
"""Answer one reviewer question and log the turn.
|
||||||
|
|
||||||
|
Returns the logged turn. Raises ChatError for a bad question or a disabled
|
||||||
|
chat, and RuntimeError when the model call fails outright (the caller maps
|
||||||
|
both to HTTP status codes).
|
||||||
|
"""
|
||||||
|
if not config.ENABLE_REVIEW_CHAT:
|
||||||
|
raise ChatError("review chat is disabled (ENABLE_REVIEW_CHAT=false)")
|
||||||
|
question = _clean_question(question)
|
||||||
|
queue = queue or []
|
||||||
|
|
||||||
|
item = next((candidate for candidate in queue
|
||||||
|
if candidate.get("review_item_id") == review_item_id), None)
|
||||||
|
if review_item_id and item is None:
|
||||||
|
raise ChatError(f"unknown review_item_id {review_item_id!r}")
|
||||||
|
|
||||||
|
context = build_context(out_dir, review_item_id, queue, decisions, question)
|
||||||
|
history = read_log(out_dir, review_item_id=review_item_id) if review_item_id \
|
||||||
|
else read_log(out_dir, scope_only=True)
|
||||||
|
|
||||||
|
scope_line = (
|
||||||
|
f"Scope: this question is about review item {review_item_id}."
|
||||||
|
if item else
|
||||||
|
"Scope: this question is about the run as a whole, not one finding."
|
||||||
|
)
|
||||||
|
user_text = (REVIEW_CHAT_USER_PROMPT
|
||||||
|
.replace("{scope_line}", scope_line)
|
||||||
|
.replace("{question}", question)
|
||||||
|
.replace("{history_block}", _history_block(history))
|
||||||
|
.replace("{context}", dumps(context)))
|
||||||
|
|
||||||
|
parsed = call_json(
|
||||||
|
system_prompt=REVIEW_CHAT_SYSTEM_PROMPT,
|
||||||
|
user_text=user_text,
|
||||||
|
max_tokens=config.REVIEW_CHAT_MAX_TOKENS,
|
||||||
|
model=config.REVIEW_CHAT_MODEL,
|
||||||
|
usage_stage="review.chat",
|
||||||
|
)
|
||||||
|
answer = _normalize_answer(parsed)
|
||||||
|
if answer is None:
|
||||||
|
raise RuntimeError("the model did not return a usable answer")
|
||||||
|
|
||||||
|
turn = {
|
||||||
|
"turn_id": uuid.uuid4().hex[:12],
|
||||||
|
"job_id": job_id,
|
||||||
|
"created_at": _now(),
|
||||||
|
"review_item_id": review_item_id,
|
||||||
|
"scope": context.get("scope"),
|
||||||
|
"issue": _issue_snapshot(item),
|
||||||
|
"question": question,
|
||||||
|
"reviewer_decision_at_time": context.get("reviewer_decision_so_far"),
|
||||||
|
"artifacts_consulted": sorted(
|
||||||
|
name for name, present
|
||||||
|
in (context.get("artifacts_available") or {}).items() if present
|
||||||
|
),
|
||||||
|
"model": config.REVIEW_CHAT_MODEL,
|
||||||
|
**answer,
|
||||||
|
}
|
||||||
|
_append_log(out_dir, turn)
|
||||||
|
append_shared_feedback(_feedback_record(turn))
|
||||||
|
return turn
|
||||||
|
|
||||||
|
|
||||||
|
def render_log_markdown(turns: List[Dict[str, Any]]) -> str:
|
||||||
|
"""Human-readable transcript: the issue, the questions, the findings.
|
||||||
|
|
||||||
|
Grouped by review item so one finding's whole thread reads together, with
|
||||||
|
run-scope questions last under their own heading.
|
||||||
|
"""
|
||||||
|
by_item: Dict[str, List[Dict[str, Any]]] = {}
|
||||||
|
for turn in turns:
|
||||||
|
by_item.setdefault(turn.get("review_item_id") or "", []).append(turn)
|
||||||
|
|
||||||
|
lines = ["# Review chat log", ""]
|
||||||
|
if not turns:
|
||||||
|
lines.append("_No questions have been asked about this run._")
|
||||||
|
return "\n".join(lines) + "\n"
|
||||||
|
lines.append(f"{len(turns)} turn(s) across {len(by_item)} thread(s).")
|
||||||
|
lines.append("")
|
||||||
|
|
||||||
|
for item_id in sorted(by_item, key=lambda key: (key == "", key)):
|
||||||
|
item_turns = by_item[item_id]
|
||||||
|
issue = next((turn.get("issue") for turn in item_turns if turn.get("issue")), None)
|
||||||
|
if not item_id:
|
||||||
|
lines += ["## Run-scope questions", "",
|
||||||
|
"_Not about a single finding._", ""]
|
||||||
|
elif issue:
|
||||||
|
lines.append(f"## {issue.get('issue_id') or item_id}")
|
||||||
|
lines.append("")
|
||||||
|
meta = [
|
||||||
|
("Category", issue.get("category")),
|
||||||
|
("Severity", issue.get("severity")),
|
||||||
|
("Run confidence", issue.get("confidence")),
|
||||||
|
("Location", issue.get("location")),
|
||||||
|
("Sheets", ", ".join(str(s) for s in issue.get("sheets") or []) or None),
|
||||||
|
("Stage", issue.get("source_stage")),
|
||||||
|
]
|
||||||
|
for label, value in meta:
|
||||||
|
if value:
|
||||||
|
lines.append(f"- **{label}:** {value}")
|
||||||
|
if issue.get("description"):
|
||||||
|
lines += ["", f"> {issue['description']}"]
|
||||||
|
lines.append("")
|
||||||
|
else:
|
||||||
|
lines += [f"## {item_id}", ""]
|
||||||
|
|
||||||
|
for turn in item_turns:
|
||||||
|
lines.append(f"### Q ({turn.get('created_at') or ''})")
|
||||||
|
lines += ["", turn.get("question") or "", "", "**Answer**", "",
|
||||||
|
turn.get("answer") or "", ""]
|
||||||
|
if turn.get("findings"):
|
||||||
|
lines.append("**Findings**")
|
||||||
|
lines.append("")
|
||||||
|
lines += [f"- {finding}" for finding in turn["findings"]]
|
||||||
|
lines.append("")
|
||||||
|
if turn.get("evidence_cited"):
|
||||||
|
lines += ["**Evidence cited**", ""]
|
||||||
|
for item in turn["evidence_cited"]:
|
||||||
|
where = item.get("artifact") or "?"
|
||||||
|
sheet = f" ({item['sheet']})" if item.get("sheet") else ""
|
||||||
|
quote = item.get("quote") or ""
|
||||||
|
lines.append(f"- `{where}`{sheet}: \"{quote}\"")
|
||||||
|
if item.get("why_it_matters"):
|
||||||
|
lines.append(f" - {item['why_it_matters']}")
|
||||||
|
lines.append("")
|
||||||
|
tail = [
|
||||||
|
("Answerable", turn.get("answerable")),
|
||||||
|
("Assessment", turn.get("assessment_of_finding")),
|
||||||
|
("Confidence", turn.get("confidence")),
|
||||||
|
("Missing", turn.get("missing_information")),
|
||||||
|
("Suggested correction", turn.get("suggested_category_correction")),
|
||||||
|
("Model", turn.get("model")),
|
||||||
|
]
|
||||||
|
lines.append(" | ".join(f"{label}: {value}" for label, value in tail if value))
|
||||||
|
lines.append("")
|
||||||
|
return "\n".join(lines) + "\n"
|
||||||
@@ -0,0 +1,315 @@
|
|||||||
|
"""Evidence bundles for the review chat.
|
||||||
|
|
||||||
|
The chat is an explainer, not an investigator: it may only answer from what the
|
||||||
|
run actually produced. This module assembles that material from the job's own
|
||||||
|
artifacts and hands the model a bounded, slimmed view.
|
||||||
|
|
||||||
|
Two shapes, matching the two kinds of question a reviewer asks:
|
||||||
|
|
||||||
|
- item scope ("why does it think the AC unit is on the ground?") - the finding,
|
||||||
|
its evidence, the cluster the finding came from, the sheets those assertions
|
||||||
|
were extracted from, any wave-5b verification verdict, the Brain's merge/drop
|
||||||
|
decision, and the reviewer's own saved decision.
|
||||||
|
- run scope ("why didn't it pick up the Civil set?") - the sheet index by
|
||||||
|
discipline, the deterministic cover-index reconciliation, per-stage counts,
|
||||||
|
what got suppressed and why, and matching job.log lines.
|
||||||
|
|
||||||
|
Everything here is read-only and degrades to empty on a missing or corrupt
|
||||||
|
artifact; a chat request must never be the thing that breaks a review screen.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
from typing import Any, Dict, List, Optional
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
|
||||||
|
# Assertion/evidence text is quoted back verbatim so the reviewer can check the
|
||||||
|
# answer against the sheet, but a whole cluster of them would swamp the prompt.
|
||||||
|
_MAX_CLUSTER_ASSERTIONS = 40
|
||||||
|
_MAX_SHEET_ASSERTIONS = 25
|
||||||
|
_MAX_SOURCE_TEXT_CHARS = 400
|
||||||
|
_MAX_SHEETS_IN_ROSTER = 400
|
||||||
|
_MAX_SUPPRESSED = 25
|
||||||
|
_MAX_LOG_LINE_CHARS = 400
|
||||||
|
|
||||||
|
|
||||||
|
def _read_json(path: str, default):
|
||||||
|
try:
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
return json.load(f)
|
||||||
|
except (OSError, json.JSONDecodeError):
|
||||||
|
return default
|
||||||
|
|
||||||
|
|
||||||
|
def _truncate(value: Any, limit: int = _MAX_SOURCE_TEXT_CHARS) -> Any:
|
||||||
|
if not isinstance(value, str) or len(value) <= limit:
|
||||||
|
return value
|
||||||
|
return value[:limit] + "..."
|
||||||
|
|
||||||
|
|
||||||
|
def _slim_assertion(assertion: Dict) -> Dict:
|
||||||
|
"""Drop base64/bookkeeping; keep what explains where a value came from."""
|
||||||
|
out = {
|
||||||
|
"sheet_number": assertion.get("sheet_number"),
|
||||||
|
"discipline": assertion.get("discipline"),
|
||||||
|
"attribute": assertion.get("attribute"),
|
||||||
|
"value": assertion.get("value"),
|
||||||
|
"source_text": _truncate(assertion.get("source_text")),
|
||||||
|
"location_key": assertion.get("location_key"),
|
||||||
|
"normalized_value": assertion.get("normalized_value"),
|
||||||
|
"disputed": assertion.get("disputed"),
|
||||||
|
}
|
||||||
|
return {key: value for key, value in out.items() if value is not None}
|
||||||
|
|
||||||
|
|
||||||
|
def _slim_finding(finding: Dict) -> Dict:
|
||||||
|
"""The finding as the run recorded it, including how it was checked."""
|
||||||
|
out = {
|
||||||
|
"issue_id": finding.get("issue_id"),
|
||||||
|
"source_stage": finding.get("source_stage"),
|
||||||
|
"agent": finding.get("agent"),
|
||||||
|
"category": finding.get("category"),
|
||||||
|
"severity": finding.get("severity"),
|
||||||
|
"confidence": finding.get("confidence"),
|
||||||
|
"location": finding.get("location"),
|
||||||
|
"disciplines": finding.get("disciplines"),
|
||||||
|
"sheets": finding.get("sheets"),
|
||||||
|
"description": _truncate(finding.get("description"), 1200),
|
||||||
|
"recommended_resolution": finding.get("recommended_resolution"),
|
||||||
|
"code_reference": finding.get("code_reference"),
|
||||||
|
"risk_score": finding.get("risk_score"),
|
||||||
|
"recommended_priority": finding.get("recommended_priority"),
|
||||||
|
"scope_id": finding.get("scope_id"),
|
||||||
|
"evidence": [
|
||||||
|
{
|
||||||
|
"discipline": item.get("discipline"),
|
||||||
|
"sheet": item.get("sheet"),
|
||||||
|
"source_text": _truncate(item.get("source_text")),
|
||||||
|
"asserted_value": item.get("asserted_value"),
|
||||||
|
}
|
||||||
|
for item in (finding.get("evidence") or [])
|
||||||
|
if isinstance(item, dict)
|
||||||
|
],
|
||||||
|
# Wave 5b / Brain-clarify re-checked some findings against fresh sheet
|
||||||
|
# images + the text layer. When present this is the single best answer
|
||||||
|
# to "did it actually look again?", so it is never dropped.
|
||||||
|
"verification": finding.get("verification"),
|
||||||
|
"clarification_of": finding.get("clarification_of"),
|
||||||
|
}
|
||||||
|
return {key: value for key, value in out.items() if value is not None}
|
||||||
|
|
||||||
|
|
||||||
|
def _slim_sheet(sheet: Dict, limit: int = _MAX_SHEET_ASSERTIONS) -> Dict:
|
||||||
|
assertions = sheet.get("assertions") or []
|
||||||
|
out = {
|
||||||
|
"sheet_number": sheet.get("sheet_number"),
|
||||||
|
"sheet_title": sheet.get("sheet_title"),
|
||||||
|
"discipline": sheet.get("discipline"),
|
||||||
|
"level": sheet.get("level"),
|
||||||
|
"page_number": sheet.get("page_number"),
|
||||||
|
"assertion_count": len(assertions),
|
||||||
|
"assertions": [_slim_assertion(item) for item in assertions[:limit]],
|
||||||
|
}
|
||||||
|
if len(assertions) > limit:
|
||||||
|
out["assertions_omitted"] = len(assertions) - limit
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def _discipline_roster(sheet_index: Dict, sheets: List[Dict]) -> Dict[str, List[str]]:
|
||||||
|
"""Sheet numbers grouped by discipline - the 'is Civil in here?' answer.
|
||||||
|
|
||||||
|
Built from the classified sheet index when there is one, falling back to
|
||||||
|
raw extraction, so an empty/failed index stage does not read as "no sheets".
|
||||||
|
"""
|
||||||
|
entries = (sheet_index or {}).get("sheet_index") or []
|
||||||
|
if not entries:
|
||||||
|
entries = [
|
||||||
|
{"sheet_number": sheet.get("sheet_number"),
|
||||||
|
"discipline": sheet.get("discipline")}
|
||||||
|
for sheet in sheets or []
|
||||||
|
]
|
||||||
|
roster: Dict[str, List[str]] = {}
|
||||||
|
for entry in entries:
|
||||||
|
if not isinstance(entry, dict):
|
||||||
|
continue
|
||||||
|
discipline = str(entry.get("discipline") or "unknown")
|
||||||
|
number = entry.get("sheet_number") or entry.get("sheet_id") or "?"
|
||||||
|
bucket = roster.setdefault(discipline, [])
|
||||||
|
if len(bucket) < _MAX_SHEETS_IN_ROSTER and number not in bucket:
|
||||||
|
bucket.append(str(number))
|
||||||
|
return roster
|
||||||
|
|
||||||
|
|
||||||
|
def _log_excerpt(out_dir: str, terms: List[str], limit: int) -> List[str]:
|
||||||
|
"""job.log lines mentioning any search term, newest last.
|
||||||
|
|
||||||
|
The run log is where stage skips, retries, and coverage decisions are
|
||||||
|
recorded ("[Code] gated off", "[Extract] page 12 empty"), which is often
|
||||||
|
the literal answer to "why didn't it look at X".
|
||||||
|
"""
|
||||||
|
path = os.path.join(out_dir, "job.log")
|
||||||
|
needles = [term.lower() for term in terms if term and len(str(term)) >= 2]
|
||||||
|
if not needles or not os.path.isfile(path):
|
||||||
|
return []
|
||||||
|
hits: List[str] = []
|
||||||
|
try:
|
||||||
|
with open(path, encoding="utf-8", errors="replace") as f:
|
||||||
|
for line in f:
|
||||||
|
lowered = line.lower()
|
||||||
|
if any(needle in lowered for needle in needles):
|
||||||
|
hits.append(_truncate(line.rstrip("\n"), _MAX_LOG_LINE_CHARS))
|
||||||
|
except OSError:
|
||||||
|
return []
|
||||||
|
return hits[-limit:]
|
||||||
|
|
||||||
|
|
||||||
|
def _stage_terms(question: str) -> List[str]:
|
||||||
|
"""Search terms for the log: quoted sheet-ish tokens plus long words.
|
||||||
|
|
||||||
|
Deliberately crude - this only decides which log lines get shown, and an
|
||||||
|
over-broad match is bounded by REVIEW_CHAT_LOG_LINES anyway.
|
||||||
|
"""
|
||||||
|
tokens = re.findall(r"[A-Za-z][A-Za-z0-9.\-]{2,}", question or "")
|
||||||
|
stop = {"the", "why", "did", "not", "and", "for", "was", "were", "does",
|
||||||
|
"this", "that", "with", "from", "what", "how", "you", "its",
|
||||||
|
"it's", "there", "when", "have", "has", "any", "are", "but"}
|
||||||
|
return [token for token in tokens if token.lower() not in stop][:12]
|
||||||
|
|
||||||
|
|
||||||
|
def build_context(out_dir: str, review_item_id: Optional[str],
|
||||||
|
queue: Optional[List[Dict]] = None,
|
||||||
|
decisions: Optional[Dict[str, Dict]] = None,
|
||||||
|
question: str = "") -> Dict[str, Any]:
|
||||||
|
"""Assemble the evidence bundle for one chat turn.
|
||||||
|
|
||||||
|
``review_item_id`` selects item scope; None (or an id not in the queue)
|
||||||
|
gives run scope. Missing artifacts degrade to empty sections rather than
|
||||||
|
raising - the model is told what is missing via ``artifacts_available``.
|
||||||
|
"""
|
||||||
|
report = _read_json(os.path.join(out_dir, "conflicts.json"), {}) or {}
|
||||||
|
snapshot = _read_json(os.path.join(out_dir, "agent", "memory.json"), {}) or {}
|
||||||
|
summary = report.get("summary") or {}
|
||||||
|
sheets = snapshot.get("sheets") or []
|
||||||
|
sheet_index = report.get("sheet_index") or snapshot.get("sheet_index") or {}
|
||||||
|
|
||||||
|
context: Dict[str, Any] = {
|
||||||
|
"scope": "run",
|
||||||
|
"run": {
|
||||||
|
"source": report.get("source"),
|
||||||
|
"pipeline_mode": summary.get("pipeline_mode"),
|
||||||
|
"agent_status": summary.get("agent_status"),
|
||||||
|
"sheets_analyzed": summary.get("sheets_analyzed") or len(sheets),
|
||||||
|
"by_stage": summary.get("by_stage"),
|
||||||
|
"conflicts_found": summary.get("conflicts_found"),
|
||||||
|
"by_severity": summary.get("by_severity"),
|
||||||
|
"models_used": summary.get("models_used"),
|
||||||
|
# Stage gating is the answer to a whole class of "why didn't it
|
||||||
|
# check X" questions, so it is stated rather than left implied.
|
||||||
|
"code_review_enabled": config.ENABLE_CODE_REVIEW,
|
||||||
|
},
|
||||||
|
"sheets_by_discipline": _discipline_roster(sheet_index, sheets),
|
||||||
|
"sheet_reconciliation": report.get("sheet_reconciliation"),
|
||||||
|
"missing_expected_sheets": (sheet_index or {}).get("missing_expected_sheets"),
|
||||||
|
"suppressed_by_the_run": [
|
||||||
|
{
|
||||||
|
"issue_id": item.get("issue_id"),
|
||||||
|
"category": item.get("category"),
|
||||||
|
"description": _truncate(item.get("description"), 300),
|
||||||
|
"verification": item.get("verification"),
|
||||||
|
}
|
||||||
|
for item in (snapshot.get("suppressed") or [])[:_MAX_SUPPRESSED]
|
||||||
|
if isinstance(item, dict)
|
||||||
|
],
|
||||||
|
"artifacts_available": {
|
||||||
|
"conflicts.json": bool(report),
|
||||||
|
"agent/memory.json": bool(snapshot),
|
||||||
|
"job.log": os.path.isfile(os.path.join(out_dir, "job.log")),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
item = None
|
||||||
|
for candidate in queue or []:
|
||||||
|
if candidate.get("review_item_id") == review_item_id:
|
||||||
|
item = candidate
|
||||||
|
break
|
||||||
|
if item is None:
|
||||||
|
context["log_excerpt"] = _log_excerpt(
|
||||||
|
out_dir, _stage_terms(question), config.REVIEW_CHAT_LOG_LINES)
|
||||||
|
return context
|
||||||
|
|
||||||
|
context["scope"] = "item"
|
||||||
|
payload = item.get("payload") or {}
|
||||||
|
context["review_item"] = {
|
||||||
|
"review_item_id": item.get("review_item_id"),
|
||||||
|
"kind": item.get("kind"),
|
||||||
|
"blocking": item.get("blocking"),
|
||||||
|
"review_triggers": item.get("reasons"),
|
||||||
|
}
|
||||||
|
saved = (decisions or {}).get(review_item_id) or {}
|
||||||
|
if saved:
|
||||||
|
context["reviewer_decision_so_far"] = {
|
||||||
|
"decision": saved.get("decision"),
|
||||||
|
"reason_code": saved.get("reason_code"),
|
||||||
|
"comment": _truncate(saved.get("comment")),
|
||||||
|
}
|
||||||
|
|
||||||
|
if item.get("kind") == "clean_cluster":
|
||||||
|
context["cluster"] = {
|
||||||
|
"key": payload.get("key"),
|
||||||
|
"location": payload.get("location"),
|
||||||
|
"disciplines": payload.get("disciplines"),
|
||||||
|
"kind": payload.get("kind"),
|
||||||
|
"assertions": [_slim_assertion(a)
|
||||||
|
for a in (payload.get("assertions") or [])[:_MAX_CLUSTER_ASSERTIONS]],
|
||||||
|
}
|
||||||
|
cited_sheets = [a.get("sheet_number") for a in payload.get("assertions") or []]
|
||||||
|
else:
|
||||||
|
context["finding"] = _slim_finding(payload)
|
||||||
|
cited_sheets = list(payload.get("sheets") or [])
|
||||||
|
cited_sheets += [e.get("sheet") for e in payload.get("evidence") or []
|
||||||
|
if isinstance(e, dict)]
|
||||||
|
scope_id = str(payload.get("scope_id") or "")
|
||||||
|
if scope_id.startswith("conflict:"):
|
||||||
|
cluster_key = scope_id.split(":", 1)[1]
|
||||||
|
cluster = next((c for c in snapshot.get("clusters") or []
|
||||||
|
if c.get("key") == cluster_key), None)
|
||||||
|
if cluster is not None:
|
||||||
|
assertions = cluster.get("assertions") or []
|
||||||
|
context["originating_cluster"] = {
|
||||||
|
"key": cluster.get("key"),
|
||||||
|
"location": cluster.get("location"),
|
||||||
|
"disciplines": cluster.get("disciplines"),
|
||||||
|
"kind": cluster.get("kind"),
|
||||||
|
"disputed_attributes": cluster.get("disputed_attributes"),
|
||||||
|
"assertion_count": len(assertions),
|
||||||
|
"assertions": [_slim_assertion(a)
|
||||||
|
for a in assertions[:_MAX_CLUSTER_ASSERTIONS]],
|
||||||
|
}
|
||||||
|
cited_sheets += [a.get("sheet_number") for a in assertions]
|
||||||
|
issue_id = payload.get("issue_id")
|
||||||
|
brain_decisions = [
|
||||||
|
decision for decision in snapshot.get("decisions") or []
|
||||||
|
if isinstance(decision, dict) and (
|
||||||
|
decision.get("kept_issue_id") == issue_id
|
||||||
|
or issue_id in (decision.get("finding_refs") or []))
|
||||||
|
]
|
||||||
|
if brain_decisions:
|
||||||
|
context["brain_decisions"] = brain_decisions[:10]
|
||||||
|
|
||||||
|
# The sheets the finding actually rests on, with their raw extraction -
|
||||||
|
# this is what lets the model say "it read 'MOUNTED ON GRADE' off M2.1".
|
||||||
|
wanted = {str(number) for number in cited_sheets if number}
|
||||||
|
if wanted:
|
||||||
|
context["source_sheets"] = [
|
||||||
|
_slim_sheet(sheet) for sheet in sheets
|
||||||
|
if str(sheet.get("sheet_number") or "") in wanted
|
||||||
|
]
|
||||||
|
|
||||||
|
context["log_excerpt"] = _log_excerpt(
|
||||||
|
out_dir,
|
||||||
|
_stage_terms(question) + sorted(wanted),
|
||||||
|
config.REVIEW_CHAT_LOG_LINES,
|
||||||
|
)
|
||||||
|
return context
|
||||||
@@ -0,0 +1,36 @@
|
|||||||
|
"""Prompts for the review-screen chat (read-only run explainer)."""
|
||||||
|
|
||||||
|
REVIEW_CHAT_SYSTEM_PROMPT = """You are the explainer for a completed automated construction-drawing review run. A human reviewer is working through the review queue and is asking you why the run reached a particular conclusion.
|
||||||
|
|
||||||
|
Your ONLY job is to explain what the run did and why, using the run's own artifacts, which are supplied to you as a JSON context bundle. You are a witness to the run, not a participant in it.
|
||||||
|
|
||||||
|
HARD RULES - never break these:
|
||||||
|
- You do NOT write, propose, suggest, or output code, patches, diffs, file edits, configuration changes, prompt changes, or shell commands. If the reviewer asks for any of those, say that this chat only explains findings, and answer the underlying question in construction-review terms instead.
|
||||||
|
- You do NOT change, re-decide, confirm, reject, or re-score any finding. The reviewer owns that decision; the radio buttons on their screen are the only thing that changes a finding. You may explain what the evidence supports, and you may say plainly that a finding looks wrong, but you never state that a finding "has been" changed.
|
||||||
|
- You answer ONLY from the supplied context bundle. You have no access to the PDF, to sheets that were not extracted, or to anything outside the bundle. Never invent a sheet number, a quotation, a dimension, or a stage that is not in the bundle.
|
||||||
|
- Separate what the run RECORDED from what you INFER. Attribute recorded facts to the artifact they came from ("the extractor recorded ... on M2.1"). Mark reasoning of your own as inference.
|
||||||
|
- When the bundle does not contain the answer, say so directly and name what is missing and which artifact would have held it. "The Civil sheets were never extracted, so there are no Civil assertions to compare" is a good answer. Guessing is not.
|
||||||
|
|
||||||
|
HOW TO ANSWER "why does it think X":
|
||||||
|
Trace the chain backwards through the bundle and quote it: the finding's evidence, the assertions in the originating cluster, the source_text the extractor pulled off each sheet, any verification verdict from the re-check pass, and the Brain's merge or drop decision. If a value is marked disputed, or the verification status is refuted or unverified, say so - that is usually the real answer.
|
||||||
|
|
||||||
|
HOW TO ANSWER "why didn't it pick up X":
|
||||||
|
Work through the bundle's coverage material in this order and report which one explains it: (1) sheets_by_discipline - was the discipline in the set at all? (2) sheet_reconciliation - did the cover sheet's own index declare sheets that were never identified (declared_not_in_set)? (3) run.by_stage and run.code_review_enabled - was the responsible stage gated off or did it produce nothing? (4) suppressed_by_the_run - was something found and then dropped? (5) log_excerpt - did the run log record a skip, a retry, or an empty page? Name the specific reason. If several are possible, say which is best supported and what would confirm it.
|
||||||
|
|
||||||
|
Be direct and concrete. Quote verbatim source_text when it carries the answer. A short, specific, evidence-anchored answer is worth more than a thorough hedge. Use plain ASCII. Respond only with valid JSON."""
|
||||||
|
|
||||||
|
REVIEW_CHAT_USER_PROMPT = """A reviewer is asking about this run. Answer from the context bundle only.
|
||||||
|
|
||||||
|
Respond ONLY with a valid JSON object - no markdown fences, no prose outside the JSON:
|
||||||
|
{"answer":"your direct explanation to the reviewer, plain text, no markdown headings","findings":["one short factual determination per item - what you established about this question, each standing on its own"],"evidence_cited":[{"artifact":"which part of the bundle, e.g. 'finding.evidence' or 'source_sheets[M2.1]' or 'log_excerpt'","sheet":"sheet number or null","quote":"verbatim text from the bundle","why_it_matters":"one sentence"}],"answerable":"yes | partial | no","missing_information":"what the bundle would need to answer fully, or null if fully answered","assessment_of_finding":"looks_supported | looks_unsupported | cannot_tell | not_applicable","suggested_category_correction":"if the reviewer is telling you the run misidentified an object, the object they say it actually is, e.g. 'power floor box'; otherwise null","confidence":"high | medium | low"}
|
||||||
|
|
||||||
|
Set assessment_of_finding to not_applicable for run-scope questions that are not about one finding. Set suggested_category_correction to null unless the reviewer is asserting a correction - do not invent one.
|
||||||
|
|
||||||
|
{scope_line}
|
||||||
|
|
||||||
|
Reviewer's question:
|
||||||
|
{question}
|
||||||
|
|
||||||
|
{history_block}
|
||||||
|
Context bundle (the complete set of artifacts you may reason from):
|
||||||
|
{context}"""
|
||||||
@@ -1,9 +1,19 @@
|
|||||||
"""Feedback labels: one label artifact per human-review decision, for metrics."""
|
"""Feedback labels: one label artifact per human-review decision, for metrics.
|
||||||
|
|
||||||
|
Labels are written twice: job-locally under ``<out_dir>/review/`` (the
|
||||||
|
auditable record for that run) and, via ``append_shared_feedback``, to the
|
||||||
|
cross-job store at ``config.REVIEW_FEEDBACK_DIR``. The shared store is
|
||||||
|
append-only and nothing reads it yet - it exists so that a later pass can prime
|
||||||
|
a run with what reviewers corrected on previous sets without having to walk
|
||||||
|
every job directory.
|
||||||
|
"""
|
||||||
|
|
||||||
import json
|
import json
|
||||||
import os
|
import os
|
||||||
from datetime import datetime, timezone
|
from datetime import datetime, timezone
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
|
||||||
|
|
||||||
def _as_dict(value) -> dict:
|
def _as_dict(value) -> dict:
|
||||||
return value if isinstance(value, dict) else {}
|
return value if isinstance(value, dict) else {}
|
||||||
@@ -21,6 +31,7 @@ def decision_to_label(queue_item: dict, decision: dict, job: dict) -> dict:
|
|||||||
payload = _as_dict(queue_item.get("payload"))
|
payload = _as_dict(queue_item.get("payload"))
|
||||||
summary = _as_dict(_as_dict(job.get("report")).get("summary"))
|
summary = _as_dict(_as_dict(job.get("report")).get("summary"))
|
||||||
return {
|
return {
|
||||||
|
"kind": "review_decision",
|
||||||
"review_item_id": queue_item.get("review_item_id"),
|
"review_item_id": queue_item.get("review_item_id"),
|
||||||
"job_id": job.get("job_id"),
|
"job_id": job.get("job_id"),
|
||||||
"pipeline_mode": job.get("pipeline_mode"),
|
"pipeline_mode": job.get("pipeline_mode"),
|
||||||
@@ -30,6 +41,11 @@ def decision_to_label(queue_item: dict, decision: dict, job: dict) -> dict:
|
|||||||
"confidence": payload.get("confidence"),
|
"confidence": payload.get("confidence"),
|
||||||
"decision": decision.get("decision"),
|
"decision": decision.get("decision"),
|
||||||
"reason_code": decision.get("reason_code"),
|
"reason_code": decision.get("reason_code"),
|
||||||
|
# The reviewer's structured corrections. Carried here (and into the
|
||||||
|
# cross-job store) so "wrong category" survives as data rather than
|
||||||
|
# only as free text on the suppressed issue.
|
||||||
|
"category_correction": decision.get("category_correction"),
|
||||||
|
"severity_correction": decision.get("severity_correction"),
|
||||||
"location": payload.get("location"),
|
"location": payload.get("location"),
|
||||||
"disciplines": payload.get("disciplines"),
|
"disciplines": payload.get("disciplines"),
|
||||||
"sheets": payload.get("sheets"),
|
"sheets": payload.get("sheets"),
|
||||||
@@ -40,7 +56,12 @@ def decision_to_label(queue_item: dict, decision: dict, job: dict) -> dict:
|
|||||||
|
|
||||||
|
|
||||||
def write_label(out_dir: str, label: dict) -> None:
|
def write_label(out_dir: str, label: dict) -> None:
|
||||||
"""Append one label as a JSON line; never raises on I/O failure."""
|
"""Append one label job-locally and to the cross-job store.
|
||||||
|
|
||||||
|
Never raises on I/O failure: a lost label must not fail the save that
|
||||||
|
produced it.
|
||||||
|
"""
|
||||||
|
append_shared_feedback(label)
|
||||||
try:
|
try:
|
||||||
review_dir = os.path.join(out_dir, "review")
|
review_dir = os.path.join(out_dir, "review")
|
||||||
os.makedirs(review_dir, exist_ok=True)
|
os.makedirs(review_dir, exist_ok=True)
|
||||||
@@ -49,3 +70,49 @@ def write_label(out_dir: str, label: dict) -> None:
|
|||||||
f.write(json.dumps(label) + "\n")
|
f.write(json.dumps(label) + "\n")
|
||||||
except OSError as e:
|
except OSError as e:
|
||||||
print(f"[Review] feedback label write failed: {e}")
|
print(f"[Review] feedback label write failed: {e}")
|
||||||
|
|
||||||
|
|
||||||
|
def append_shared_feedback(record: dict) -> None:
|
||||||
|
"""Append one record to the cross-job feedback store; never raises.
|
||||||
|
|
||||||
|
One JSONL file per record ``kind`` so a reader can pick up decisions and
|
||||||
|
chat turns independently. Failure here is logged and swallowed: the
|
||||||
|
cross-job roll-up is a convenience, and losing a line must never fail the
|
||||||
|
review action that produced it.
|
||||||
|
"""
|
||||||
|
try:
|
||||||
|
kind = str(record.get("kind") or "misc")
|
||||||
|
os.makedirs(config.REVIEW_FEEDBACK_DIR, exist_ok=True)
|
||||||
|
name = "chat_turns.jsonl" if kind == "review_chat_turn" else "decisions.jsonl"
|
||||||
|
path = os.path.join(config.REVIEW_FEEDBACK_DIR, name)
|
||||||
|
with open(path, "a", encoding="utf-8") as f:
|
||||||
|
f.write(json.dumps(record) + "\n")
|
||||||
|
except OSError as e:
|
||||||
|
print(f"[Review] shared feedback write failed: {e}")
|
||||||
|
|
||||||
|
|
||||||
|
def read_shared_feedback(kind: str = "review_decision") -> list:
|
||||||
|
"""Read the cross-job store for one record kind, oldest first.
|
||||||
|
|
||||||
|
Corrupt lines are skipped so a partial write cannot hide the rest.
|
||||||
|
"""
|
||||||
|
name = "chat_turns.jsonl" if kind == "review_chat_turn" else "decisions.jsonl"
|
||||||
|
path = os.path.join(config.REVIEW_FEEDBACK_DIR, name)
|
||||||
|
if not os.path.isfile(path):
|
||||||
|
return []
|
||||||
|
records = []
|
||||||
|
try:
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
for line in f:
|
||||||
|
line = line.strip()
|
||||||
|
if not line:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
value = json.loads(line)
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
continue
|
||||||
|
if isinstance(value, dict):
|
||||||
|
records.append(value)
|
||||||
|
except OSError:
|
||||||
|
return []
|
||||||
|
return records
|
||||||
@@ -216,7 +216,7 @@ def finalize_review(job_id: str, out_dir: str) -> dict:
|
|||||||
rfis = _draft_rfis(kept)
|
rfis = _draft_rfis(kept)
|
||||||
|
|
||||||
report["validated_issues"] = kept
|
report["validated_issues"] = kept
|
||||||
report["suppressed_issues"] = suppressed
|
report["suppressed_issues"] = (report.get("suppressed_issues") or []) + suppressed
|
||||||
report["rfis"] = rfis
|
report["rfis"] = rfis
|
||||||
summary = report.setdefault("summary", {})
|
summary = report.setdefault("summary", {})
|
||||||
summary["agent_status"] = "complete"
|
summary["agent_status"] = "complete"
|
||||||
|
|||||||
@@ -0,0 +1,93 @@
|
|||||||
|
"""sheet_reconcile.py - deterministic sheet-list reconciliation (no LLM).
|
||||||
|
|
||||||
|
The cover sheet's own sheet index (SHEET LIST / DRAWING INDEX) declares which
|
||||||
|
sheets the set is SUPPOSED to contain. Comparing that declaration against the
|
||||||
|
sheets wave-1 actually identified answers two early questions:
|
||||||
|
|
||||||
|
- declared_not_in_set: sheets the index lists but we didn't identify - dark
|
||||||
|
pages, misidentification, or disciplines genuinely absent from this PDF.
|
||||||
|
- in_set_not_declared: sheet numbers we extracted that the index doesn't
|
||||||
|
list - misread title blocks or unlisted sheets.
|
||||||
|
|
||||||
|
Deterministic complement to the LLM sheet_index stage, which can only infer
|
||||||
|
from what extraction already found.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import re
|
||||||
|
from typing import Dict, List, Optional
|
||||||
|
|
||||||
|
# Markers that introduce the drawing set's own sheet index on a cover page.
|
||||||
|
_INDEX_MARKERS = (
|
||||||
|
"SHEET LIST",
|
||||||
|
"DRAWING INDEX",
|
||||||
|
"SHEET INDEX",
|
||||||
|
"DRAWING LIST",
|
||||||
|
"INDEX OF DRAWINGS",
|
||||||
|
)
|
||||||
|
|
||||||
|
# Sheet ids: 1-2 letters, optional hyphen, 2-3 digits, optional decimal suffix.
|
||||||
|
# Covers S301, A102, LS101, C-001, C-001.1; excludes dates/project numbers
|
||||||
|
# (pure digits) and member marks (W12X26 - letter after digits).
|
||||||
|
_SHEET_TOKEN_RE = re.compile(r"\b([A-Z]{1,2}-?\d{2,3}(?:\.\d+)?)\b")
|
||||||
|
|
||||||
|
# Only cover-front pages carry the set index.
|
||||||
|
_MAX_INDEX_PAGE = 5
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_id(sheet_id: str) -> str:
|
||||||
|
return (sheet_id or "").upper().replace("-", "").strip()
|
||||||
|
|
||||||
|
|
||||||
|
def declared_sheet_list(page_texts: Dict[int, Optional[str]]) -> List[str]:
|
||||||
|
"""Scrape the declared sheet list off the cover page's text layer.
|
||||||
|
|
||||||
|
page_texts: {page_number: text_layer_or_None}. Returns the ordered,
|
||||||
|
deduped list of declared sheet ids, or [] when no index marker exists.
|
||||||
|
Only the FIRST page containing a marker is parsed (later 'sheet list'
|
||||||
|
echoes in legends/schedules are ignored).
|
||||||
|
"""
|
||||||
|
for page_number in sorted(page_texts):
|
||||||
|
if page_number > _MAX_INDEX_PAGE:
|
||||||
|
break
|
||||||
|
text = page_texts.get(page_number) or ""
|
||||||
|
upper = text.upper()
|
||||||
|
marker_at = -1
|
||||||
|
for marker in _INDEX_MARKERS:
|
||||||
|
marker_at = upper.find(marker)
|
||||||
|
if marker_at >= 0:
|
||||||
|
break
|
||||||
|
if marker_at < 0:
|
||||||
|
continue
|
||||||
|
section = text[marker_at:]
|
||||||
|
declared: List[str] = []
|
||||||
|
for token in _SHEET_TOKEN_RE.findall(section):
|
||||||
|
if token not in declared:
|
||||||
|
declared.append(token)
|
||||||
|
return declared
|
||||||
|
return []
|
||||||
|
|
||||||
|
|
||||||
|
def reconcile_sheets(sheets: List[Dict], declared: List[str]) -> Dict:
|
||||||
|
"""Compare extracted sheet_numbers against the declared index.
|
||||||
|
|
||||||
|
Comparison is hyphen/case-normalized; output lists keep the declared /
|
||||||
|
extracted originals.
|
||||||
|
"""
|
||||||
|
found: List[str] = [str(s["sheet_number"]) for s in sheets or []
|
||||||
|
if s.get("sheet_number")]
|
||||||
|
found_norm = {_normalize_id(n) for n in found}
|
||||||
|
declared_norm = {_normalize_id(n) for n in declared}
|
||||||
|
|
||||||
|
declared_not_in_set = [n for n in declared if _normalize_id(n) not in found_norm]
|
||||||
|
# Preserve extraction order, dedupe, keep originals.
|
||||||
|
in_set_not_declared: List[str] = []
|
||||||
|
for n in found:
|
||||||
|
if _normalize_id(n) not in declared_norm and n not in in_set_not_declared:
|
||||||
|
in_set_not_declared.append(n)
|
||||||
|
|
||||||
|
return {
|
||||||
|
"declared_total": len(declared),
|
||||||
|
"found_total": len(found),
|
||||||
|
"declared_not_in_set": declared_not_in_set,
|
||||||
|
"in_set_not_declared": in_set_not_declared,
|
||||||
|
}
|
||||||
@@ -0,0 +1,140 @@
|
|||||||
|
"""text_coverage.py - deterministic extraction-coverage measurement.
|
||||||
|
|
||||||
|
The coverage guarantee: for any page with a usable text layer, measure how
|
||||||
|
much of that layer ended up represented in extracted objects. Pages below
|
||||||
|
the floor route into the extraction retry ladder (agents/extractors.py and
|
||||||
|
pipeline/extractor.py). fallback_objects() is the last rung: stub objects
|
||||||
|
segmented straight from the text layer so no text-bearing page goes dark.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import re
|
||||||
|
from typing import Dict, List, Optional
|
||||||
|
|
||||||
|
MIN_LINE_CHARS = 12
|
||||||
|
_TICK_RE = re.compile(r"^[\d\s'\"/.,-]+$")
|
||||||
|
_WORD_RE = re.compile(r"[a-z0-9]+")
|
||||||
|
|
||||||
|
|
||||||
|
def _meaningful_lines(text: str) -> List[str]:
|
||||||
|
lines = []
|
||||||
|
for raw in (text or "").splitlines():
|
||||||
|
line = " ".join(raw.split())
|
||||||
|
if len(line) < MIN_LINE_CHARS or _TICK_RE.match(line):
|
||||||
|
continue
|
||||||
|
lines.append(line)
|
||||||
|
return lines
|
||||||
|
|
||||||
|
|
||||||
|
def _norm(text: str) -> str:
|
||||||
|
return " ".join(_WORD_RE.findall((text or "").lower()))
|
||||||
|
|
||||||
|
|
||||||
|
def text_coverage(page_text: str, objects: List[Dict]) -> Dict:
|
||||||
|
"""Fraction of meaningful text-layer lines whose normalized form appears
|
||||||
|
in the concatenated normalized source_text of extracted objects."""
|
||||||
|
lines = _meaningful_lines(page_text)
|
||||||
|
if not lines:
|
||||||
|
return {"total_lines": 0, "covered_lines": 0, "ratio": 1.0}
|
||||||
|
haystack = " ".join(
|
||||||
|
_norm(str(o.get("source_text") or o.get("object_description")
|
||||||
|
or o.get("value") or ""))
|
||||||
|
for o in objects if isinstance(o, dict)
|
||||||
|
)
|
||||||
|
covered = sum(1 for ln in lines if _norm(ln) and _norm(ln) in haystack)
|
||||||
|
return {
|
||||||
|
"total_lines": len(lines),
|
||||||
|
"covered_lines": covered,
|
||||||
|
"ratio": covered / len(lines) if lines else 1.0,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def segment_text_layer(text: str) -> List[str]:
|
||||||
|
"""Segment a page text layer into note-sized blocks."""
|
||||||
|
segments: List[str] = []
|
||||||
|
buf: List[str] = []
|
||||||
|
number_re = re.compile(r"^(\d{1,2}[.)]?|[A-Z]\d{0,2}[.)]?)\s*$")
|
||||||
|
|
||||||
|
def flush():
|
||||||
|
joined = " ".join(buf).strip()
|
||||||
|
if len(joined) >= MIN_LINE_CHARS:
|
||||||
|
segments.append(joined)
|
||||||
|
buf.clear()
|
||||||
|
|
||||||
|
for raw in (text or "").splitlines():
|
||||||
|
line = raw.strip()
|
||||||
|
if not line:
|
||||||
|
flush()
|
||||||
|
continue
|
||||||
|
if number_re.match(line):
|
||||||
|
flush()
|
||||||
|
buf.append(line.rstrip(".)"))
|
||||||
|
continue
|
||||||
|
buf.append(line)
|
||||||
|
if line.endswith(".") and len(" ".join(buf)) > 120:
|
||||||
|
flush()
|
||||||
|
flush()
|
||||||
|
return segments
|
||||||
|
|
||||||
|
|
||||||
|
def fallback_objects(page_text: str, page_number: int,
|
||||||
|
max_objects: int = 200) -> List[Dict]:
|
||||||
|
"""Last-rung deterministic extraction: one stub object per text segment,
|
||||||
|
source_text verbatim from the text layer."""
|
||||||
|
objs = []
|
||||||
|
for idx, seg in enumerate(segment_text_layer(page_text)[:max_objects]):
|
||||||
|
objs.append({
|
||||||
|
"object_id": f"p{page_number}-tl{idx}",
|
||||||
|
"object_type": "general_note",
|
||||||
|
"category": "general",
|
||||||
|
"tag": None,
|
||||||
|
"name": seg[:80],
|
||||||
|
"description": seg,
|
||||||
|
"attributes": {},
|
||||||
|
"location_key": {},
|
||||||
|
"source_text": seg,
|
||||||
|
"graphical_basis": None,
|
||||||
|
"review_uses": ["code_review", "constructability_review"],
|
||||||
|
"confidence": "low",
|
||||||
|
"grounding": "text_layer_fallback",
|
||||||
|
})
|
||||||
|
return objs
|
||||||
|
|
||||||
|
|
||||||
|
def merge_objects(vision_objs: List[Dict], text_objs: List[Dict]) -> List[Dict]:
|
||||||
|
"""Union of vision and text-structured objects. Vision results come first
|
||||||
|
and are never dropped. Text objects are appended unless their normalized
|
||||||
|
source_text is already represented."""
|
||||||
|
merged = list(vision_objs or [])
|
||||||
|
seen = {_norm(str(o.get("source_text") or ""))
|
||||||
|
for o in merged if isinstance(o, dict)}
|
||||||
|
seen.discard("")
|
||||||
|
for obj in text_objs or []:
|
||||||
|
if not isinstance(obj, dict):
|
||||||
|
continue
|
||||||
|
key = _norm(str(obj.get("source_text") or ""))
|
||||||
|
if key and key in seen:
|
||||||
|
continue
|
||||||
|
seen.add(key)
|
||||||
|
merged.append(obj)
|
||||||
|
return merged
|
||||||
|
|
||||||
|
|
||||||
|
# Sheet ids: 1-2 letters, OPTIONAL HYPHEN, 2-3 digits, optional decimal suffix.
|
||||||
|
# The hyphen matters: civil/landscape sets number sheets C-001 / L-101, and a
|
||||||
|
# regex without it leaves those pages sheet_number=None, which then shows up as
|
||||||
|
# a false "declared but not in set" in sheet_reconcile. Kept in sync with
|
||||||
|
# sheet_reconcile._SHEET_TOKEN_RE.
|
||||||
|
_SHEET_ID_RE = re.compile(r"\b([A-Z]{1,2}-?\d{2,3}(?:\.\d+)?)\b")
|
||||||
|
|
||||||
|
|
||||||
|
def recover_sheet_number(page_text: str) -> Optional[str]:
|
||||||
|
"""Deterministic sheet id from the text layer: prefer candidates in the
|
||||||
|
last ~15% of the page (title block lives at the drawing edge)."""
|
||||||
|
text = page_text or ""
|
||||||
|
cands = _SHEET_ID_RE.findall(text)
|
||||||
|
if not cands:
|
||||||
|
return None
|
||||||
|
tail = text[int(len(text) * 0.85):]
|
||||||
|
for cand in reversed(_SHEET_ID_RE.findall(tail)):
|
||||||
|
return cand
|
||||||
|
return cands[0]
|
||||||
@@ -0,0 +1,230 @@
|
|||||||
|
"""
|
||||||
|
text_layer.py - deterministic PDF text-layer extraction (PyMuPDF, no LLM).
|
||||||
|
|
||||||
|
Most CAD-produced drawing sets carry a real vector text layer. We extract it
|
||||||
|
once per job and feed it to the extractor (grounding), the grounding guard
|
||||||
|
(rescue tier), and the wave-5b verifier (text oracle + high-DPI evidence
|
||||||
|
crops). Pages below TEXT_LAYER_MIN_CHARS of text are treated as having no
|
||||||
|
text layer (scanned/raster sheets stay vision-only).
|
||||||
|
|
||||||
|
If PyMuPDF is unavailable the module degrades gracefully: every public
|
||||||
|
function returns empty/None, equivalent to TEXT_LAYER_ENABLED=false.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import re
|
||||||
|
from typing import Dict, List, Optional, Tuple
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
|
||||||
|
try: # PyMuPDF >= 1.24 prefers the pymupdf name; fitz works everywhere.
|
||||||
|
import pymupdf as fitz
|
||||||
|
except ImportError: # pragma: no cover - older PyMuPDF
|
||||||
|
try:
|
||||||
|
import fitz
|
||||||
|
except ImportError: # pragma: no cover - PyMuPDF not installed
|
||||||
|
fitz = None
|
||||||
|
|
||||||
|
_warned_unavailable = False
|
||||||
|
|
||||||
|
# Word token normalization for evidence matching: lowercase alphanumeric only.
|
||||||
|
_TOKEN_RE = re.compile(r"[^a-z0-9]+")
|
||||||
|
# Fuzzy match floor: fraction of needle tokens that must align with the page's
|
||||||
|
# word sequence for a bbox to count as a confident evidence location.
|
||||||
|
_FUZZY_MIN_RATIO = 0.6
|
||||||
|
|
||||||
|
|
||||||
|
def _fitz_or_none():
|
||||||
|
"""Return the fitz module, logging once if PyMuPDF is missing."""
|
||||||
|
global _warned_unavailable
|
||||||
|
if fitz is None and not _warned_unavailable:
|
||||||
|
print("[TextLayer] PyMuPDF not available - text-layer grounding disabled")
|
||||||
|
_warned_unavailable = True
|
||||||
|
return fitz
|
||||||
|
|
||||||
|
|
||||||
|
def extract_text_layers(pdf_path: str) -> Dict[int, Dict]:
|
||||||
|
"""
|
||||||
|
Extract the text layer of every page. Returns {1-based page_number:
|
||||||
|
{"text": str, "words": [{"text", "bbox": (x0,y0,x1,y1)}, ...],
|
||||||
|
"has_text_layer": bool}}. Returns {} when disabled or unavailable.
|
||||||
|
"""
|
||||||
|
if not config.TEXT_LAYER_ENABLED:
|
||||||
|
return {}
|
||||||
|
f = _fitz_or_none()
|
||||||
|
if f is None:
|
||||||
|
return {}
|
||||||
|
try:
|
||||||
|
doc = f.open(pdf_path)
|
||||||
|
except Exception as exc:
|
||||||
|
print(f"[TextLayer] could not open {pdf_path}: {exc}")
|
||||||
|
return {}
|
||||||
|
layers: Dict[int, Dict] = {}
|
||||||
|
try:
|
||||||
|
for index in range(doc.page_count):
|
||||||
|
page = doc[index]
|
||||||
|
text = page.get_text("text") or ""
|
||||||
|
words = [
|
||||||
|
{"text": w[4], "bbox": (w[0], w[1], w[2], w[3])}
|
||||||
|
for w in (page.get_text("words") or [])
|
||||||
|
]
|
||||||
|
has_text_layer = len(text.strip()) >= config.TEXT_LAYER_MIN_CHARS
|
||||||
|
if not has_text_layer:
|
||||||
|
print(f"[TextLayer] Page {index + 1}: {len(text.strip())} chars "
|
||||||
|
f"(< TEXT_LAYER_MIN_CHARS={config.TEXT_LAYER_MIN_CHARS}) - "
|
||||||
|
f"vision-only")
|
||||||
|
layers[index + 1] = {
|
||||||
|
"text": text,
|
||||||
|
"words": words,
|
||||||
|
"has_text_layer": has_text_layer,
|
||||||
|
}
|
||||||
|
finally:
|
||||||
|
doc.close()
|
||||||
|
return layers
|
||||||
|
|
||||||
|
|
||||||
|
def attach_text_layers(
|
||||||
|
pdf_path: str,
|
||||||
|
pages: List[Dict],
|
||||||
|
text_dir: Optional[str] = None,
|
||||||
|
) -> Dict[int, List[Dict]]:
|
||||||
|
"""
|
||||||
|
Attach page["text_layer"] (text or None) to each converted page dict and
|
||||||
|
return the runner-local {page_number: words} map (kept off page dicts -
|
||||||
|
those get serialized). When text_dir is set, dump one .txt per page there
|
||||||
|
(plain file writes; ProjectMemory is a closed registry).
|
||||||
|
"""
|
||||||
|
layers = extract_text_layers(pdf_path)
|
||||||
|
page_words: Dict[int, List[Dict]] = {}
|
||||||
|
for page in pages:
|
||||||
|
layer = layers.get(page["page_number"]) or {}
|
||||||
|
page["text_layer"] = layer.get("text") if layer.get("has_text_layer") else None
|
||||||
|
page_words[page["page_number"]] = layer.get("words") or []
|
||||||
|
if text_dir and layers:
|
||||||
|
import os
|
||||||
|
os.makedirs(text_dir, exist_ok=True)
|
||||||
|
for page_number, layer in layers.items():
|
||||||
|
if not layer.get("has_text_layer"):
|
||||||
|
continue
|
||||||
|
with open(os.path.join(text_dir, f"page-{page_number:03d}.txt"),
|
||||||
|
"w", encoding="utf-8") as fh:
|
||||||
|
fh.write(layer.get("text") or "")
|
||||||
|
return page_words
|
||||||
|
|
||||||
|
|
||||||
|
def _tokens(text: str) -> List[str]:
|
||||||
|
return [t for t in _TOKEN_RE.split(text.lower()) if t]
|
||||||
|
|
||||||
|
|
||||||
|
def _union_bbox(boxes: List[Tuple[float, float, float, float]]):
|
||||||
|
return (
|
||||||
|
min(b[0] for b in boxes),
|
||||||
|
min(b[1] for b in boxes),
|
||||||
|
max(b[2] for b in boxes),
|
||||||
|
max(b[3] for b in boxes),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def find_evidence_bbox(
|
||||||
|
words: List[Dict],
|
||||||
|
needle: str,
|
||||||
|
) -> Optional[Tuple[float, float, float, float]]:
|
||||||
|
"""
|
||||||
|
Best-effort fuzzy substring match of an evidence source_text against the
|
||||||
|
page's word sequence. Returns the union bbox of the matched words, or
|
||||||
|
None when nothing aligns confidently.
|
||||||
|
|
||||||
|
Exact contiguous token runs win; otherwise the best-scoring window with
|
||||||
|
>= _FUZZY_MIN_RATIO token alignment is accepted (vision quotes imperfectly
|
||||||
|
but the value is real page text).
|
||||||
|
"""
|
||||||
|
if not words or not needle:
|
||||||
|
return None
|
||||||
|
needle_tokens = _tokens(str(needle))
|
||||||
|
if not needle_tokens:
|
||||||
|
return None
|
||||||
|
page_tokens = [_tokens(w.get("text") or "") for w in words]
|
||||||
|
# Flatten multi-token words, remembering which word each token came from.
|
||||||
|
flat: List[Tuple[str, int]] = []
|
||||||
|
for word_index, parts in enumerate(page_tokens):
|
||||||
|
for part in parts:
|
||||||
|
flat.append((part, word_index))
|
||||||
|
if not flat:
|
||||||
|
return None
|
||||||
|
|
||||||
|
n = len(needle_tokens)
|
||||||
|
best_span = None
|
||||||
|
best_score = 0.0
|
||||||
|
for start in range(0, len(flat)):
|
||||||
|
window = flat[start:start + n]
|
||||||
|
if not window:
|
||||||
|
break
|
||||||
|
score = sum(1 for i, tok in enumerate(needle_tokens)
|
||||||
|
if i < len(window) and window[i][0] == tok) / n
|
||||||
|
if score > best_score:
|
||||||
|
best_score = score
|
||||||
|
best_span = window
|
||||||
|
if best_score == 1.0:
|
||||||
|
break
|
||||||
|
if best_span is None or best_score < _FUZZY_MIN_RATIO:
|
||||||
|
return None
|
||||||
|
word_indexes = {word_index for _, word_index in best_span}
|
||||||
|
return _union_bbox([words[i]["bbox"] for i in sorted(word_indexes)])
|
||||||
|
|
||||||
|
|
||||||
|
def render_crop(
|
||||||
|
pdf_path: str,
|
||||||
|
page_number: int,
|
||||||
|
bbox: Tuple[float, float, float, float],
|
||||||
|
dpi: Optional[int] = None,
|
||||||
|
margin_pts: Optional[float] = None,
|
||||||
|
) -> Optional[bytes]:
|
||||||
|
"""
|
||||||
|
Render a clip of one page around bbox (+ margin, clamped to the page) at
|
||||||
|
the given DPI and return JPEG bytes, or None on any failure.
|
||||||
|
"""
|
||||||
|
f = _fitz_or_none()
|
||||||
|
if f is None:
|
||||||
|
return None
|
||||||
|
dpi = dpi or config.VERIFY_CROP_DPI
|
||||||
|
margin_pts = config.VERIFY_CROP_MARGIN_PTS if margin_pts is None else margin_pts
|
||||||
|
try:
|
||||||
|
doc = f.open(pdf_path)
|
||||||
|
try:
|
||||||
|
page = doc[page_number - 1]
|
||||||
|
rect = f.Rect(
|
||||||
|
bbox[0] - margin_pts,
|
||||||
|
bbox[1] - margin_pts,
|
||||||
|
bbox[2] + margin_pts,
|
||||||
|
bbox[3] + margin_pts,
|
||||||
|
) & page.rect
|
||||||
|
if rect.is_empty:
|
||||||
|
return None
|
||||||
|
pix = page.get_pixmap(clip=rect, dpi=dpi)
|
||||||
|
return pix.tobytes("jpeg")
|
||||||
|
finally:
|
||||||
|
doc.close()
|
||||||
|
except Exception as exc:
|
||||||
|
print(f"[TextLayer] render_crop failed on page {page_number}: {exc}")
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def coverage_gaps(pages: List[Dict], sheets: List[Dict]) -> List[int]:
|
||||||
|
"""
|
||||||
|
Page numbers that have a text layer but whose extraction failed or
|
||||||
|
returned 0 objects - the silent extraction-loss signal. Logs one
|
||||||
|
[TextLayer] line per gap.
|
||||||
|
"""
|
||||||
|
by_page = {s.get("page_number"): s for s in sheets or []}
|
||||||
|
gaps: List[int] = []
|
||||||
|
for page in pages:
|
||||||
|
text = page.get("text_layer")
|
||||||
|
if not text:
|
||||||
|
continue
|
||||||
|
sheet = by_page.get(page["page_number"])
|
||||||
|
extracted = len(sheet.get("assertions") or []) if sheet else 0
|
||||||
|
if extracted == 0:
|
||||||
|
gaps.append(page["page_number"])
|
||||||
|
print(f"[TextLayer] Page {page['page_number']}: text layer present "
|
||||||
|
f"({len(text)} chars) but no objects extracted — possible "
|
||||||
|
f"extraction gap")
|
||||||
|
return gaps
|
||||||
@@ -0,0 +1,190 @@
|
|||||||
|
<!DOCTYPE html>
|
||||||
|
<html lang="en">
|
||||||
|
<head>
|
||||||
|
<meta charset="UTF-8">
|
||||||
|
<title>Conflict Checker — How Your Plans Get Reviewed</title>
|
||||||
|
<style>
|
||||||
|
* { box-sizing: border-box; margin: 0; padding: 0; }
|
||||||
|
body {
|
||||||
|
font-family: "Segoe UI", "Helvetica Neue", Arial, sans-serif;
|
||||||
|
background: #f4f7fb;
|
||||||
|
color: #1f2d3d;
|
||||||
|
width: 1280px;
|
||||||
|
padding: 40px 48px;
|
||||||
|
}
|
||||||
|
header { text-align: center; margin-bottom: 10px; }
|
||||||
|
h1 { font-size: 34px; color: #123c6e; letter-spacing: 0.5px; }
|
||||||
|
.subtitle { font-size: 17px; color: #5a6b7f; margin-top: 8px; }
|
||||||
|
|
||||||
|
.blueprint {
|
||||||
|
background: #ffffff;
|
||||||
|
border: 2px solid #d5e3f2;
|
||||||
|
border-radius: 18px;
|
||||||
|
padding: 32px 36px;
|
||||||
|
margin-top: 24px;
|
||||||
|
background-image:
|
||||||
|
linear-gradient(#eef4fb 1px, transparent 1px),
|
||||||
|
linear-gradient(90deg, #eef4fb 1px, transparent 1px);
|
||||||
|
background-size: 28px 28px;
|
||||||
|
}
|
||||||
|
|
||||||
|
.row { display: flex; justify-content: center; align-items: stretch; gap: 0; }
|
||||||
|
.row + .connector-down { margin: 0; }
|
||||||
|
|
||||||
|
.card {
|
||||||
|
background: #ffffff;
|
||||||
|
border-radius: 14px;
|
||||||
|
border: 2px solid #cfdcec;
|
||||||
|
box-shadow: 0 3px 8px rgba(18,60,110,0.08);
|
||||||
|
width: 250px;
|
||||||
|
padding: 16px 16px 14px;
|
||||||
|
position: relative;
|
||||||
|
flex-shrink: 0;
|
||||||
|
}
|
||||||
|
.card .num {
|
||||||
|
position: absolute; top: -16px; left: -14px;
|
||||||
|
width: 36px; height: 36px; border-radius: 50%;
|
||||||
|
background: #123c6e; color: #fff;
|
||||||
|
font-weight: 700; font-size: 18px;
|
||||||
|
display: flex; align-items: center; justify-content: center;
|
||||||
|
box-shadow: 0 2px 5px rgba(0,0,0,0.2);
|
||||||
|
}
|
||||||
|
.card .icon { font-size: 34px; text-align: center; margin: 4px 0 6px; }
|
||||||
|
.card h2 { font-size: 17px; color: #123c6e; text-align: center; margin-bottom: 6px; }
|
||||||
|
.card p { font-size: 13px; line-height: 1.4; color: #42536a; text-align: center; }
|
||||||
|
|
||||||
|
.card.scan { border-color: #7fb3e0; background: #f0f7ff; }
|
||||||
|
.card.read { border-color: #7fb3e0; background: #f0f7ff; }
|
||||||
|
.card.lib { border-color: #8fd0a8; background: #f1faf4; }
|
||||||
|
.card.link { border-color: #8fd0a8; background: #f1faf4; }
|
||||||
|
.card.det { border-color: #f2b879; background: #fff8ef; }
|
||||||
|
.card.spec { border-color: #f2b879; background: #fff8ef; }
|
||||||
|
.card.brain { border-color: #c39bd3; background: #f9f3fc; }
|
||||||
|
.card.human { border-color: #e58f8f; background: #fdf1f1; }
|
||||||
|
|
||||||
|
.arrow {
|
||||||
|
display: flex; align-items: center; justify-content: center;
|
||||||
|
color: #123c6e; font-size: 30px; font-weight: bold;
|
||||||
|
width: 44px; flex-shrink: 0;
|
||||||
|
}
|
||||||
|
.connector-down {
|
||||||
|
text-align: center; color: #123c6e; font-size: 30px;
|
||||||
|
font-weight: bold; line-height: 1; padding: 6px 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
.finish {
|
||||||
|
margin: 22px auto 0;
|
||||||
|
width: 560px;
|
||||||
|
background: #123c6e; color: #ffffff;
|
||||||
|
border-radius: 14px; padding: 18px 24px; text-align: center;
|
||||||
|
box-shadow: 0 4px 10px rgba(18,60,110,0.3);
|
||||||
|
}
|
||||||
|
.finish .big { font-size: 20px; font-weight: 700; }
|
||||||
|
.finish .small { font-size: 14px; margin-top: 6px; color: #cfe0f4; }
|
||||||
|
|
||||||
|
footer {
|
||||||
|
margin-top: 26px; text-align: center;
|
||||||
|
font-size: 13px; color: #7a8aa0;
|
||||||
|
}
|
||||||
|
.note {
|
||||||
|
margin: 18px auto 0; width: 900px; font-size: 13.5px; color: #42536a;
|
||||||
|
background: #ffffff; border-left: 4px solid #7fb3e0; border-radius: 6px;
|
||||||
|
padding: 10px 16px; line-height: 1.5;
|
||||||
|
}
|
||||||
|
</style>
|
||||||
|
</head>
|
||||||
|
<body>
|
||||||
|
|
||||||
|
<header>
|
||||||
|
<h1>🔍 CONFLICT CHECKER</h1>
|
||||||
|
<div class="subtitle">How your construction plans get reviewed — a team of AI assistants, each with one job, passing notes down the line.</div>
|
||||||
|
</header>
|
||||||
|
|
||||||
|
<div class="blueprint">
|
||||||
|
|
||||||
|
<!-- Row 1 -->
|
||||||
|
<div class="row">
|
||||||
|
<div class="card scan">
|
||||||
|
<div class="num">1</div>
|
||||||
|
<div class="icon">📄</div>
|
||||||
|
<h2>The Scanner</h2>
|
||||||
|
<p>Turns every page of your PDF blueprints into a picture the AI can read.</p>
|
||||||
|
</div>
|
||||||
|
<div class="arrow">→</div>
|
||||||
|
<div class="card read">
|
||||||
|
<div class="num">2</div>
|
||||||
|
<div class="icon">👓</div>
|
||||||
|
<h2>The Readers</h2>
|
||||||
|
<p>One assistant per page, all working at once. Each writes down every fact: dimensions, notes, materials, callouts.</p>
|
||||||
|
</div>
|
||||||
|
<div class="arrow">→</div>
|
||||||
|
<div class="card lib">
|
||||||
|
<div class="num">3</div>
|
||||||
|
<div class="icon">📚</div>
|
||||||
|
<h2>Librarian & Code Scout</h2>
|
||||||
|
<p>Builds the table of contents (electrical, plumbing, structural…) and figures out <b>where</b> the project is, so the right building codes apply.</p>
|
||||||
|
</div>
|
||||||
|
<div class="arrow">→</div>
|
||||||
|
<div class="card link">
|
||||||
|
<div class="num">4</div>
|
||||||
|
<div class="icon">🔗</div>
|
||||||
|
<h2>The Connector</h2>
|
||||||
|
<p>Connects the dots across sheets — "this water heater on the plumbing sheet is the same one on the electrical sheet" — and sorts facts into topic piles.</p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="connector-down">↓</div>
|
||||||
|
|
||||||
|
<!-- Row 2 -->
|
||||||
|
<div class="row">
|
||||||
|
<div class="card det">
|
||||||
|
<div class="num">5</div>
|
||||||
|
<div class="icon">🕵️</div>
|
||||||
|
<h2>The Detectives</h2>
|
||||||
|
<p>One per topic pile. Hunts for contradictions between sheets that should agree — <i>and</i> mistakes within a single sheet: "wall shown here, but not on the structural plan."</p>
|
||||||
|
</div>
|
||||||
|
<div class="arrow">→</div>
|
||||||
|
<div class="card spec">
|
||||||
|
<div class="num">6</div>
|
||||||
|
<div class="icon">📐</div>
|
||||||
|
<h2>The Specialists</h2>
|
||||||
|
<p>Working at once: a <b>drawing proofreader</b> (dangling callouts, a schedule vs its own plan, dimensions that don't add up), a veteran <b>builder</b> ("can this be built?"), and a <b>checklist keeper</b> ("is anything missing?").</p>
|
||||||
|
</div>
|
||||||
|
<div class="arrow">→</div>
|
||||||
|
<div class="card brain">
|
||||||
|
<div class="num">7</div>
|
||||||
|
<div class="icon">🧠</div>
|
||||||
|
<h2>The Brain</h2>
|
||||||
|
<p>The senior reviewer. Merges duplicates, discards weak findings, ranks the rest — then sends the ones it doubts back for a <b>zoomed-in second look</b> and drops any that don't hold up.</p>
|
||||||
|
</div>
|
||||||
|
<div class="arrow">→</div>
|
||||||
|
<div class="card human">
|
||||||
|
<div class="num">8</div>
|
||||||
|
<div class="icon">✅</div>
|
||||||
|
<h2>Human Review</h2>
|
||||||
|
<p>The important and uncertain findings land on <b>your</b> desk. You confirm, reject, or mark unsure — nothing goes out unapproved.</p>
|
||||||
|
</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="connector-down">↓</div>
|
||||||
|
|
||||||
|
<div class="finish">
|
||||||
|
<div class="big">📋 Final Report + ✉️ Draft RFIs</div>
|
||||||
|
<div class="small">A prioritized list of every problem found — plus ready-to-send "please clarify" letters (Requests For Information) for the design team.</div>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="note">
|
||||||
|
<b>Good to know:</b> the review is focused on the <b>drawings themselves</b> —
|
||||||
|
contradictions, single-sheet mistakes, buildability, and missing pieces
|
||||||
|
(building-code checks are built in but turned off by default). Everyone shares one
|
||||||
|
notebook, so each step builds on the last. If one page can't be read, the team keeps
|
||||||
|
going and that page is flagged as a gap instead of stopping the whole review. Every
|
||||||
|
finding links back to the sheet it came from.
|
||||||
|
</div>
|
||||||
|
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<footer>Conflict Checker · conchecker.scoutitsystems.com · Review your plans before they cost you money in the field.</footer>
|
||||||
|
|
||||||
|
</body>
|
||||||
|
</html>
|
||||||
Binary file not shown.
|
After Width: | Height: | Size: 492 KiB |
@@ -0,0 +1,213 @@
|
|||||||
|
# Conflict Checker — How It Works (Plain Language)
|
||||||
|
|
||||||
|
**What it does:** You upload a set of construction drawings (a PDF of blueprints).
|
||||||
|
A team of AI assistants reads every page, compares everything against everything
|
||||||
|
else, and hands you a list of problems — contradictions between sheets, mistakes
|
||||||
|
within a single sheet, missing information, and things that would be hard to
|
||||||
|
build — before they cost you money in the field.
|
||||||
|
|
||||||
|
Think of it like hiring a room full of specialist consultants to review your
|
||||||
|
plans overnight. Each one has a specific job, they pass their notes down the
|
||||||
|
table, and a senior reviewer at the end sorts it all into one clean report.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The Big Picture (one sentence per step)
|
||||||
|
|
||||||
|
```
|
||||||
|
YOUR PDF OF BLUEPRINTS
|
||||||
|
|
|
||||||
|
v
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
| 0. SCANNER |
|
||||||
|
| Turns every PDF page into a picture the AI can read |
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
|
|
||||||
|
v
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
| 1. READERS (one assistant per page, all at once) |
|
||||||
|
| Reads each sheet and writes down every fact: |
|
||||||
|
| dimensions, notes, materials, room names, callouts |
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
|
|
||||||
|
v
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
| 2. LIBRARIAN + LOCAL-CODE SCOUT (work side by side) |
|
||||||
|
| Librarian: builds the table of contents — which |
|
||||||
|
| sheets exist (electrical, plumbing, structural...) |
|
||||||
|
| Scout: figures out WHERE the project is, so we know |
|
||||||
|
| which building codes apply |
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
|
|
||||||
|
v
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
| 3. CONNECTOR |
|
||||||
|
| Connects the dots across sheets — e.g. "the water |
|
||||||
|
| heater on the plumbing sheet is the same one on the |
|
||||||
|
| electrical sheet" — and groups related facts into |
|
||||||
|
| topic piles (clusters) |
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
|
|
||||||
|
v
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
| 4. CONFLICT DETECTIVES (one per topic pile) |
|
||||||
|
| Compares sheets that should agree and looks for |
|
||||||
|
| contradictions: "Wall shown here on A-201 but not |
|
||||||
|
| on S-101", "Pipe runs through the duct". Now also |
|
||||||
|
| catches contradictions WITHIN a single sheet |
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
|
|
||||||
|
v
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
| 5. THREE SPECIALISTS (work side by side) |
|
||||||
|
| * Drawing Checker — problems on a sheet BY ITSELF: |
|
||||||
|
| a callout pointing to a detail that isn't there, |
|
||||||
|
| a schedule that disagrees with its own plan, |
|
||||||
|
| dimensions that don't add up, missing scale |
|
||||||
|
| * Builder — can this actually be built as |
|
||||||
|
| drawn? (access, clearances, sequencing) |
|
||||||
|
| * Completeness Checker — is anything MISSING from |
|
||||||
|
| the set? (sheets, schedules, required details) |
|
||||||
|
| (A Code Inspector also lives here but is turned OFF |
|
||||||
|
| by default — the focus is the drawings themselves) |
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
|
|
||||||
|
v
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
| 6. THE BRAIN (senior reviewer) |
|
||||||
|
| Collects EVERY finding from everyone, merges the |
|
||||||
|
| duplicates, throws out the weak ones, and ranks the |
|
||||||
|
| rest by how much trouble they'd cause |
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
|
|
||||||
|
v
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
| 6.5 THE BRAIN DOUBLE-CHECKS (asks for a second look) |
|
||||||
|
| For the findings it's unsure about, the Brain sends |
|
||||||
|
| them back to a fact-checker that re-reads the actual |
|
||||||
|
| sheet (zoomed-in image + the page's real text) to |
|
||||||
|
| confirm or debunk. Debunked findings are dropped |
|
||||||
|
| before they ever reach you |
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
|
|
||||||
|
v
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
| 7. HUMAN REVIEW GATE |
|
||||||
|
| The important/uncertain findings are queued for a |
|
||||||
|
| real person to Confirm / Reject / mark Unsure. |
|
||||||
|
| You can also ASK the run why it concluded any of |
|
||||||
|
| them, or why it never looked at something |
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
|
|
||||||
|
v
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
| 8. LETTER WRITER |
|
||||||
|
| Drafts a formal RFI (Request For Information — the |
|
||||||
|
| official "please clarify this" letter) for each |
|
||||||
|
| confirmed issue, ready to send to the design team |
|
||||||
|
+----------------------------------------------------------+
|
||||||
|
|
|
||||||
|
v
|
||||||
|
FINAL REPORT + DRAFT RFIs
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Who's Who (the "agents")
|
||||||
|
|
||||||
|
| # | Name | Analogy | What it actually does |
|
||||||
|
|---|------|---------|-----------------------|
|
||||||
|
| 0 | PDF Scanner | Photocopier | Converts each PDF page into an image the AI can "see" |
|
||||||
|
| 1 | Sheet Extractor | Speed-reader | Reads one page, writes structured notes (every page gets its own reader, in parallel) |
|
||||||
|
| 2 | Sheet Indexer | Librarian | Builds the table of contents of the drawing set |
|
||||||
|
| 2 | Jurisdiction Scout | Local guide | Identifies the project's location so the right building codes are used |
|
||||||
|
| 3 | Linker | Connector | Groups related facts from different sheets into topic clusters |
|
||||||
|
| 4 | Conflict Critic | Detective | Examines each cluster for contradictions — between disciplines, across a discipline's own sheets, or within one sheet |
|
||||||
|
| 5 | Drawing Integrity Agent | Proofreader | Checks each sheet on its own: dangling callouts, a schedule vs its own plan, dimensions that don't sum, missing scale/north/title-block |
|
||||||
|
| 5 | Constructability Agent | Veteran builder | Flags things that are drawn fine but can't be built practically |
|
||||||
|
| 5 | Completeness Agent | Checklist keeper | Flags missing sheets, missing details, gaps in the set |
|
||||||
|
| 5 | Code Agent *(off by default)* | Code inspector | Building-code/ADA checks — kept in the codebase but disabled so the review focuses on the drawings; one flag turns it back on |
|
||||||
|
| 6 | Brain | Chief estimator | Deduplicates, judges, and prioritizes all findings |
|
||||||
|
| 6.5 | Brain (clarification) | Second opinion | For findings it distrusts, sends them back to the fact-checker to re-read the sheet; debunked findings are dropped |
|
||||||
|
| 7 | Review Gate | Your desk | Presents the findings a human should approve before anything goes out |
|
||||||
|
| 7 | Review Chat | The analyst you can question | Answers "why did it decide that?" and "why didn't it check that?" from the run's own records — it explains, it never changes anything |
|
||||||
|
| 8 | RFI Writer | Secretary | Writes the formal clarification letters for confirmed issues |
|
||||||
|
|
||||||
|
Everything the assistants learn is kept in a shared notebook (the "project
|
||||||
|
memory"), so each step builds on the last. If one reader fails on one page, the
|
||||||
|
rest of the team keeps going — that page is noted as a gap instead of crashing
|
||||||
|
the whole review.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Where Improvements Could Be Made
|
||||||
|
|
||||||
|
*(Several items from earlier versions have since shipped — noted below.)*
|
||||||
|
|
||||||
|
### 1. Coverage — "make sure every page actually got read" ✅ *largely shipped*
|
||||||
|
- Failed pages used to quietly disappear, and the Completeness Checker would
|
||||||
|
then report a sheet as "missing" when it was really just unread.
|
||||||
|
- **Done:** a retry ladder now re-reads a page (text-only pass, then a
|
||||||
|
deterministic text-layer fallback) so no text-bearing page goes dark, and the
|
||||||
|
report separates "sheet doesn't exist" from "sheet couldn't be read."
|
||||||
|
- **Still open:** try a different backup model on the hardest pages.
|
||||||
|
|
||||||
|
### 2. Speed — "the team waits in line more than it needs to"
|
||||||
|
- The steps run strictly one after another, but some could start earlier. The
|
||||||
|
Jurisdiction Scout only needs the cover page — it could run while the other
|
||||||
|
Readers are still working. The Letter Writer could start on high-confidence
|
||||||
|
findings instead of waiting for all human review.
|
||||||
|
- **Improvement:** overlap independent steps; start drafting letters for
|
||||||
|
confirmed/high-confidence findings sooner.
|
||||||
|
|
||||||
|
### 3. Cost — "smarter reading, fewer wasted words"
|
||||||
|
- Every page is read by a large, expensive AI model, and that model's
|
||||||
|
"thinking time" counts against its answer budget — we've seen it spend its
|
||||||
|
whole budget thinking and return a cut-off answer.
|
||||||
|
- **Improvement:** use cheaper models for simple pages (schedules, title
|
||||||
|
sheets), save the expensive model for dense drawings; keep tuning the
|
||||||
|
thinking budget knobs; reuse cached answers when the same plan set is
|
||||||
|
re-run. *(The truncation bug itself is now fixed.)*
|
||||||
|
|
||||||
|
### 4. Smarter detective work — "catch conflicts that span piles" ✅ *shipped*
|
||||||
|
- The Detectives only saw one topic pile at a time, so a contradiction spanning
|
||||||
|
two piles — or a mistake on a single sheet — could slip through.
|
||||||
|
- **Done:** the Detectives now also flag contradictions *within* a single sheet,
|
||||||
|
a new Drawing Checker proofreads every sheet on its own, and after the Brain
|
||||||
|
sorts everything it can send doubtful findings back for a zoomed-in second
|
||||||
|
look (wave 6.5) and drop the ones that don't hold up.
|
||||||
|
- **Still open:** revisit the hard cap on how many topic piles are kept on very
|
||||||
|
large sets.
|
||||||
|
|
||||||
|
### 5. Human time — "review less, but review what matters"
|
||||||
|
- Today the review queue is built from rules about severity and confidence.
|
||||||
|
- **Improvement:** learn from your past Confirm/Reject decisions to sort the
|
||||||
|
queue better — the system already records your feedback, so it can get
|
||||||
|
smarter over time about what actually needs your eyes.
|
||||||
|
|
||||||
|
### 5b. Explaining itself — "why did you think that?" ✅ *shipped*
|
||||||
|
- **Done:** every finding on the review screen has a chat panel, plus one for
|
||||||
|
the run as a whole. Ask why a unit was read as ground-mounted, or why a whole
|
||||||
|
discipline never got looked at, and it traces the answer back through what it
|
||||||
|
actually recorded — quoting the note it read off the sheet, or naming the
|
||||||
|
stage that skipped the pages. It cannot change a finding; that stays yours.
|
||||||
|
- **Done:** when you correct it in conversation ("that's not a floor drain,
|
||||||
|
it's a power floor box"), the correction is filed as structured data rather
|
||||||
|
than a free-text comment.
|
||||||
|
- **Still open:** nothing reads those filed corrections back yet. The next step
|
||||||
|
is priming a new run with what reviewers corrected on previous sets, so the
|
||||||
|
same misread does not come back on the next job.
|
||||||
|
|
||||||
|
### 6. Trust — "show the receipts" ✅ *partially shipped*
|
||||||
|
- **Done:** the fact-checker already pulls a zoomed-in crop of the exact spot on
|
||||||
|
the sheet when it re-reads a finding.
|
||||||
|
- **Still open:** attach that crop to the finding in the final report so a
|
||||||
|
non-technical reader can verify it in seconds without opening the PDF.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
*Technical reference for the curious: the pipeline lives in
|
||||||
|
`backend/agents/runner.py` (the waves above are the "Agent wave N" stages), the
|
||||||
|
team's shared notebook is `backend/agents/memory.py`, the review queue is
|
||||||
|
`backend/review/gate.py` + `backend/review/finalizer.py`, and the review chat is
|
||||||
|
`backend/review/chat.py` + `backend/review/chat_context.py`.*
|
||||||
@@ -0,0 +1,166 @@
|
|||||||
|
# Text-Layer Grounding — Design Spec
|
||||||
|
|
||||||
|
**Date:** 2026-08-12 · **Branch:** `agent-mode` · **Status:** approved by user (2026-08-12)
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
The pipeline is vision-only for extraction, but most CAD-produced drawing sets
|
||||||
|
carry a real vector text layer. Two worst documented failure modes are text
|
||||||
|
problems being solved with pixels:
|
||||||
|
|
||||||
|
1. **Wave-1 text misreads propagate immutably** — e.g. job `959e16407573`:
|
||||||
|
vision read "(2) 2x6 STUD PACK" where the sheet says "(5)"; text-only
|
||||||
|
downstream specialists treated the misread as ground truth → confident
|
||||||
|
false-positive findings.
|
||||||
|
2. **Silent extraction loss** — failed/under-extracted pages are invisible
|
||||||
|
(job `475a6f184dd1`: 42% extraction loss), producing false
|
||||||
|
`missing_expected_sheets` warnings and missed conflicts.
|
||||||
|
|
||||||
|
Priority (user, 2026-08-12): reduce false positives **and** missed items;
|
||||||
|
more accurate conflicts.
|
||||||
|
|
||||||
|
## Approach
|
||||||
|
|
||||||
|
Extract the PDF text layer deterministically (PyMuPDF) once per job, and make
|
||||||
|
it a first-class citizen at three points: extractor grounding, the grounding
|
||||||
|
guard, and the wave-5b verifier (as text oracle + high-DPI evidence crops).
|
||||||
|
|
||||||
|
Inspired by `hamzaabduljabbar/construction-drawing-analyzer` (patterns only —
|
||||||
|
its license is source-available/no-resale; all code here is original).
|
||||||
|
|
||||||
|
## Components
|
||||||
|
|
||||||
|
### 1. New module `backend/text_layer.py` (deterministic, no LLM)
|
||||||
|
|
||||||
|
- `extract_text_layers(pdf_path) -> Dict[int, dict]` — per 1-based page:
|
||||||
|
`{"text": str, "words": [{"text", "bbox": (x0,y0,x1,y1)}, ...],
|
||||||
|
"has_text_layer": bool}`. Pages with < `TEXT_LAYER_MIN_CHARS` of text are
|
||||||
|
`has_text_layer=False` (scanned/raster sheets stay vision-only; logged).
|
||||||
|
- `find_evidence_bbox(words, needle) -> bbox | None` — best-effort fuzzy
|
||||||
|
substring match of an evidence `source_text` against word sequence; returns
|
||||||
|
union rect of matched words.
|
||||||
|
- `render_crop(pdf_path, page_number, bbox, dpi, margin_pts) -> bytes` —
|
||||||
|
PyMuPDF `page.get_pixmap(clip=rect, dpi=dpi)` → JPEG bytes.
|
||||||
|
|
||||||
|
Both runners call `extract_text_layers` right after `convert_pdf_to_images`
|
||||||
|
and attach `page["text_layer"] = <text or None>` to each page dict. Word
|
||||||
|
positions stay in a separate `page_words: Dict[int, list]` runner-local map
|
||||||
|
(not attached to page dicts — they get serialized).
|
||||||
|
|
||||||
|
### 2. Extractor grounding (both pipelines)
|
||||||
|
|
||||||
|
- Static paragraph added to `_EXTRACTOR_SYSTEM_TEMPLATE` in
|
||||||
|
`backend/prompts.py` (no new placeholder): when a TEXT LAYER block is
|
||||||
|
present in the user message it is **authoritative for alphanumeric content**
|
||||||
|
(counts, dimensions, member tags, notes); the image is for geometry,
|
||||||
|
symbols, linework, and anything absent from the text layer.
|
||||||
|
- Text-layer content is **appended programmatically** at each extractor call
|
||||||
|
site (classic `extractor.py::_extract_one`, agent
|
||||||
|
`extractors.py::SheetExtractorAgent.run`) — NOT a new `{placeholder}` in the
|
||||||
|
shared template (two-render-path trap: `render()` silently leaves missing
|
||||||
|
keys as literals). Block capped at `TEXT_LAYER_MAX_CHARS`.
|
||||||
|
Format: `\n\nTEXT LAYER (authoritative for alphanumeric content — trust it
|
||||||
|
over the image for numbers, tags, and note text):\n<text>`
|
||||||
|
|
||||||
|
### 3. Grounding guard rescue tier (`pipeline/extractor.py::_normalize_sheet`)
|
||||||
|
|
||||||
|
Current guard drops an object when its primary value's digit-runs aren't in
|
||||||
|
its own `source_text`. New tier, only when a text layer exists for the page:
|
||||||
|
|
||||||
|
- digits ⊆ source_text → keep (unchanged)
|
||||||
|
- digits ⊆ page text layer but ⊄ source_text → keep, stamp
|
||||||
|
`grounding: "text_layer"` on the assertion (recall rescue — vision quoted
|
||||||
|
imperfectly but the value is real page text)
|
||||||
|
- otherwise → drop (unchanged)
|
||||||
|
|
||||||
|
`_is_grounded` gains an optional `page_text` param; existing callers/tests
|
||||||
|
unaffected. Dropped/ rescued counts logged per page.
|
||||||
|
|
||||||
|
### 4. Verifier: text oracle + high-DPI crops (wave 5b)
|
||||||
|
|
||||||
|
Wherever verify scopes are built (agent runner confirmed; classic runner to be
|
||||||
|
checked — integrate at both if present):
|
||||||
|
|
||||||
|
- Scope payload gains `text_layer_excerpt`: concatenated text of the finding's
|
||||||
|
cited sheets, capped at `VERIFY_TEXT_MAX_CHARS`. `VERIFY_USER_INSTRUCTION`
|
||||||
|
gains a `{text_layer}` placeholder with instructions to treat it as
|
||||||
|
deterministic page text (verdicts may cite it as `actual_text`). **Both
|
||||||
|
render sites** (agent verifier + any classic-path render) must substitute it
|
||||||
|
— grep the template name across `backend/agents/` and `backend/pipeline/`.
|
||||||
|
- When `VERIFY_HI_DPI_CROPS` and the page has words: for each evidence item,
|
||||||
|
`find_evidence_bbox` on the cited page's words; on hit, `render_crop` at
|
||||||
|
`VERIFY_CROP_DPI` with margin → crop images replace full-page images (up to
|
||||||
|
`AGENT_CONFLICT_MAX_IMAGES`). On any miss/failure → fall back to the current
|
||||||
|
full-page image. Zero-resolved-images ⇒ scope skipped (I2 guard preserved).
|
||||||
|
|
||||||
|
### 5. Coverage signal (recall)
|
||||||
|
|
||||||
|
After extraction in both runners: for each page with `has_text_layer=True`
|
||||||
|
whose extraction failed or returned 0 objects, log
|
||||||
|
`[TextLayer] Page N: text layer present (M chars) but no objects extracted —
|
||||||
|
possible extraction gap` and add the page to the existing gap-finding path
|
||||||
|
(agent: `orchestrator.stats.failed_scopes`-style finding; classic: log only).
|
||||||
|
|
||||||
|
## Config knobs (`backend/config.py`, env-overridable, documented in `.env.example`)
|
||||||
|
|
||||||
|
| Key | Default | Effect |
|
||||||
|
|-----|---------|--------|
|
||||||
|
| `TEXT_LAYER_ENABLED` | `true` | Master switch |
|
||||||
|
| `TEXT_LAYER_MIN_CHARS` | `20` | Below this per page → `has_text_layer=False` |
|
||||||
|
| `TEXT_LAYER_MAX_CHARS` | `12000` | Cap per sheet injected into extractor prompt |
|
||||||
|
| `VERIFY_TEXT_MAX_CHARS` | `8000` | Cap of text-layer excerpt in verify scope |
|
||||||
|
| `VERIFY_HI_DPI_CROPS` | `true` | Evidence-located crops in verifier |
|
||||||
|
| `VERIFY_CROP_DPI` | `300` | Crop render DPI |
|
||||||
|
| `VERIFY_CROP_MARGIN_PTS` | `36` | Padding around evidence bbox (PDF points) |
|
||||||
|
|
||||||
|
## Known traps (from project history — designed around)
|
||||||
|
|
||||||
|
- **Two render paths:** no new `{placeholder}` in extractor templates; the one
|
||||||
|
new placeholder (`{text_layer}` in VERIFY_USER_INSTRUCTION) substituted at
|
||||||
|
every render site; a render test asserts no `{...}` literals remain.
|
||||||
|
- **ProjectMemory closed registry:** no new memory keys. Text artifacts dump
|
||||||
|
via plain file writes under `outputs/<job>/text/` (agent: under `agent/`).
|
||||||
|
- **`slim_clusters`:** no new cluster fields — unchanged.
|
||||||
|
- **I2 zero-image path:** crops replace full-page images only on confident
|
||||||
|
bbox match; never reduce image count to zero.
|
||||||
|
- **Base64 hygiene:** page dicts already carry base64; `text_layer` strings
|
||||||
|
must not leak into `clusters.json` dumps — reuse `_without_base64` pattern
|
||||||
|
if assertions ever carry page refs (they don't today).
|
||||||
|
|
||||||
|
## Dependencies
|
||||||
|
|
||||||
|
`PyMuPDF>=1.23` added to `requirements.txt` (Docker image rebuild picks it up;
|
||||||
|
pdf2image/poppler unchanged).
|
||||||
|
|
||||||
|
## Testing
|
||||||
|
|
||||||
|
- `tests/test_text_layer.py` — build tiny PDFs with PyMuPDF in-test:
|
||||||
|
extraction, `has_text_layer` thresholds, `find_evidence_bbox` hit/miss,
|
||||||
|
`render_crop` dimensions.
|
||||||
|
- Extractor guard: rescue-tier unit tests (keep-with-flag, still-drop,
|
||||||
|
unchanged behavior without text layer).
|
||||||
|
- Prompt render test: extractor + verify instructions fully substituted at
|
||||||
|
every site (both pipelines).
|
||||||
|
- Runner-level (pattern from `tests/agents/test_wave5b_suppression.py`):
|
||||||
|
stubbed waves, assert text layer reaches extract scopes and verify scopes
|
||||||
|
(excerpt present, crop fallback on no-match), full `run_agent_pipeline`.
|
||||||
|
- Full `pytest tests/` green before push.
|
||||||
|
|
||||||
|
## Validation (post-deploy)
|
||||||
|
|
||||||
|
Re-run the Cypress set (source PDF persists at
|
||||||
|
`/app/backend/outputs/959e16407573/source.pdf` on sits-docker) per the
|
||||||
|
documented re-run workflow. Success criteria:
|
||||||
|
|
||||||
|
1. The "(2) vs (5)"-class findings are not generated, or are verifier-refuted
|
||||||
|
with text-layer evidence cited.
|
||||||
|
2. Coverage-gap log lines appear for any page with text but no objects.
|
||||||
|
3. No new `finish_reason=length` in waves 1/4; cost delta reported vs
|
||||||
|
baseline job.
|
||||||
|
|
||||||
|
## Out of scope (future PRs)
|
||||||
|
|
||||||
|
- Legend/symbol-library wave injected into extractor + critic prompts.
|
||||||
|
- Deterministic schedule-row recall pass (text-layer tables → assertions).
|
||||||
|
- pdf-markup export for the review UI.
|
||||||
|
- Takeoff/polygon geometry (belongs to AI_Takeoffs, not this product).
|
||||||
+316
-38
@@ -3,6 +3,9 @@
|
|||||||
<head>
|
<head>
|
||||||
<meta charset="utf-8" />
|
<meta charset="utf-8" />
|
||||||
<meta name="viewport" content="width=device-width, initial-scale=1" />
|
<meta name="viewport" content="width=device-width, initial-scale=1" />
|
||||||
|
<meta http-equiv="Cache-Control" content="no-cache, no-store, must-revalidate" />
|
||||||
|
<meta http-equiv="Pragma" content="no-cache" />
|
||||||
|
<meta http-equiv="Expires" content="0" />
|
||||||
<title>Conflict Checker</title>
|
<title>Conflict Checker</title>
|
||||||
<style>
|
<style>
|
||||||
:root {
|
:root {
|
||||||
@@ -15,6 +18,7 @@
|
|||||||
header { padding:24px 28px; border-bottom:1px solid var(--line); }
|
header { padding:24px 28px; border-bottom:1px solid var(--line); }
|
||||||
h1 { margin:0; font-size:20px; letter-spacing:.2px; }
|
h1 { margin:0; font-size:20px; letter-spacing:.2px; }
|
||||||
.sub { color:var(--muted); font-size:13px; margin-top:4px; }
|
.sub { color:var(--muted); font-size:13px; margin-top:4px; }
|
||||||
|
#buildTag { display:inline-block; font-size:11px; background:rgba(91,140,255,.15); color:var(--accent); padding:2px 8px; border-radius:12px; margin-left:8px; vertical-align:middle; text-transform:none; letter-spacing:.3px; }
|
||||||
main { max-width:920px; margin:0 auto; padding:28px; }
|
main { max-width:920px; margin:0 auto; padding:28px; }
|
||||||
.drop { border:1.5px dashed var(--line); border-radius:12px; padding:36px; text-align:center;
|
.drop { border:1.5px dashed var(--line); border-radius:12px; padding:36px; text-align:center;
|
||||||
background:var(--panel); transition:border-color .15s; cursor:pointer; }
|
background:var(--panel); transition:border-color .15s; cursor:pointer; }
|
||||||
@@ -60,7 +64,15 @@
|
|||||||
border:1px solid var(--line); background:#0c0e13; color:var(--text); font-size:16px; outline:none; }
|
border:1px solid var(--line); background:#0c0e13; color:var(--text); font-size:16px; outline:none; }
|
||||||
.email-card input[type=email]::placeholder { color:#6b7280; }
|
.email-card input[type=email]::placeholder { color:#6b7280; }
|
||||||
.email-card input[type=email]:focus { border-color:var(--accent); box-shadow:0 0 0 3px rgba(91,140,255,.18); }
|
.email-card input[type=email]:focus { border-color:var(--accent); box-shadow:0 0 0 3px rgba(91,140,255,.18); }
|
||||||
|
.email-card select { width:100%; padding:10px 12px; border-radius:8px; margin-top:6px;
|
||||||
|
border:1px solid var(--line); background:#0c0e13; color:var(--text); font-size:14px; }
|
||||||
|
.email-card .field { margin-top:12px; }
|
||||||
|
.email-card .field > span { display:block; font-size:13px; color:var(--muted); margin-bottom:2px; }
|
||||||
.btn.full { width:100%; padding:14px; font-size:15px; margin-top:0; }
|
.btn.full { width:100%; padding:14px; font-size:15px; margin-top:0; }
|
||||||
|
.logbox { background:#0c0e13; border:1px solid var(--line); border-radius:8px; padding:12px 14px;
|
||||||
|
margin-top:10px; max-height:320px; overflow:auto; font:12px/1.45 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace;
|
||||||
|
color:#c6cdd8; white-space:pre-wrap; word-break:break-word; }
|
||||||
|
.logbox .empty-log { color:var(--muted); }
|
||||||
.note { background:var(--panel); border:1px solid var(--line); border-radius:10px;
|
.note { background:var(--panel); border:1px solid var(--line); border-radius:10px;
|
||||||
padding:16px 18px; margin:18px 0; }
|
padding:16px 18px; margin:18px 0; }
|
||||||
.note b { color:var(--text); }
|
.note b { color:var(--text); }
|
||||||
@@ -70,6 +82,29 @@
|
|||||||
border:1px solid var(--line); border-radius:6px; padding:6px 8px; font-size:13px; margin-top:6px; }
|
border:1px solid var(--line); border-radius:6px; padding:6px 8px; font-size:13px; margin-top:6px; }
|
||||||
.review-controls input[type=text] { width:100%; }
|
.review-controls input[type=text] { width:100%; }
|
||||||
.review-controls .hidden { display:none; }
|
.review-controls .hidden { display:none; }
|
||||||
|
.chat { margin-top:10px; padding-top:10px; border-top:1px solid var(--line); font-size:13px; }
|
||||||
|
.chat-toggle { background:none; border:0; color:var(--accent); cursor:pointer; padding:0;
|
||||||
|
font-size:13px; text-decoration:underline dotted; }
|
||||||
|
.chat-body { margin-top:10px; }
|
||||||
|
.chat-body.hidden, .chat .hidden { display:none; }
|
||||||
|
.chat-turns { max-height:340px; overflow-y:auto; margin-bottom:8px; }
|
||||||
|
.chat-q, .chat-a { border-radius:8px; padding:8px 10px; margin:6px 0; }
|
||||||
|
.chat-q { background:#161b25; }
|
||||||
|
.chat-q b { color:var(--accent); }
|
||||||
|
.chat-a { background:#11141a; }
|
||||||
|
.chat-a ul { margin:6px 0 0; padding-left:18px; }
|
||||||
|
.chat-a li { margin:2px 0; }
|
||||||
|
.chat-cite { color:var(--muted); font-size:12px; margin-top:6px; }
|
||||||
|
.chat-cite code { color:var(--accent); }
|
||||||
|
.chat-tags { color:var(--muted); font-size:12px; margin-top:6px; }
|
||||||
|
.chat-row { display:flex; gap:8px; align-items:flex-start; }
|
||||||
|
.chat-row textarea { flex:1; background:#0c0e13; color:var(--text); border:1px solid var(--line);
|
||||||
|
border-radius:6px; padding:8px; font-size:13px; font-family:inherit; resize:vertical;
|
||||||
|
min-height:38px; }
|
||||||
|
.chat-row button { white-space:nowrap; }
|
||||||
|
.chat-hint { color:var(--muted); font-size:12px; margin-top:6px; }
|
||||||
|
.chat-err { color:var(--hi); font-size:12px; margin-top:6px; }
|
||||||
|
.btn.sm { padding:8px 14px; font-size:13px; margin-top:0; }
|
||||||
.pill.blocking { background:rgba(255,93,87,.15); color:var(--hi); }
|
.pill.blocking { background:rgba(255,93,87,.15); color:var(--hi); }
|
||||||
.pill.audit { background:rgba(91,140,255,.15); color:var(--accent); }
|
.pill.audit { background:rgba(91,140,255,.15); color:var(--accent); }
|
||||||
.pill.critical { background:rgba(255,93,87,.28); color:#fff; }
|
.pill.critical { background:rgba(255,93,87,.28); color:#fff; }
|
||||||
@@ -87,7 +122,7 @@
|
|||||||
<body>
|
<body>
|
||||||
<header>
|
<header>
|
||||||
<h1>Conflict Checker</h1>
|
<h1>Conflict Checker</h1>
|
||||||
<div class="sub">Cross-discipline design contradiction review for construction drawing sets<span id="buildTag"></span></div>
|
<div class="sub">Cross-discipline design contradiction review for construction drawing sets<span id="buildTag">build ...</span></div>
|
||||||
</header>
|
</header>
|
||||||
<main>
|
<main>
|
||||||
<div class="drop" id="drop">
|
<div class="drop" id="drop">
|
||||||
@@ -99,7 +134,7 @@
|
|||||||
<input type="email" id="email" placeholder="you@firm.com" />
|
<input type="email" id="email" placeholder="you@firm.com" />
|
||||||
</div>
|
</div>
|
||||||
<details class="email-card" id="intake">
|
<details class="email-card" id="intake">
|
||||||
<summary style="cursor:pointer">Project details <span class="opt">(optional — improves code/ADA review)</span></summary>
|
<summary style="cursor:pointer">Project details <span class="opt">(optional)</span></summary>
|
||||||
<input type="text" id="project_name" placeholder="Project name" style="width:100%;margin-top:8px" />
|
<input type="text" id="project_name" placeholder="Project name" style="width:100%;margin-top:8px" />
|
||||||
<input type="text" id="address" placeholder="Project address" style="width:100%;margin-top:8px" />
|
<input type="text" id="address" placeholder="Project address" style="width:100%;margin-top:8px" />
|
||||||
<input type="text" id="occupancy" placeholder="Occupancy (e.g. Business, Assembly)" style="width:100%;margin-top:8px" />
|
<input type="text" id="occupancy" placeholder="Occupancy (e.g. Business, Assembly)" style="width:100%;margin-top:8px" />
|
||||||
@@ -118,15 +153,23 @@
|
|||||||
<label>⚙️ Compute <span class="opt">(text stages; vision always runs on the API)</span></label>
|
<label>⚙️ Compute <span class="opt">(text stages; vision always runs on the API)</span></label>
|
||||||
<label style="display:block;font-weight:400;margin-top:6px">
|
<label style="display:block;font-weight:400;margin-top:6px">
|
||||||
<input type="radio" name="compute" value="openrouter" checked> OpenRouter — all stages (fastest, paid)</label>
|
<input type="radio" name="compute" value="openrouter" checked> OpenRouter — all stages (fastest, paid)</label>
|
||||||
<label style="display:block;font-weight:400;margin-top:6px">
|
|
||||||
<input type="radio" name="compute" value="local"> Hybrid — text stages on local LLM (cheaper, slower)</label>
|
|
||||||
<div id="modelPick" style="margin-top:10px">
|
<div id="modelPick" style="margin-top:10px">
|
||||||
<label for="model" style="font-weight:400">Model <span class="opt" id="modelNote">loading...</span></label>
|
<div class="field">
|
||||||
<select id="model" style="width:100%;margin-top:6px;padding:10px;border-radius:8px;border:1px solid var(--line);background:#0c0e13;color:var(--text)"></select>
|
<span>Vision model <span class="opt">(image stages)</span> <span class="opt" id="modelNote">loading...</span></span>
|
||||||
|
<select id="vision_model" disabled><option value="">Loading models…</option></select>
|
||||||
|
</div>
|
||||||
|
<div class="field">
|
||||||
|
<span>Text model <span class="opt">(non-image stages)</span></span>
|
||||||
|
<select id="text_model" disabled><option value="">Loading models…</option></select>
|
||||||
|
</div>
|
||||||
</div>
|
</div>
|
||||||
</div>
|
</div>
|
||||||
<button class="btn full" id="run" disabled>Run conflict check</button>
|
<button class="btn full" id="run" disabled>Run conflict check</button>
|
||||||
<div class="status" id="status"></div>
|
<div class="status" id="status"></div>
|
||||||
|
<div id="liveLog" style="display:none" class="note">
|
||||||
|
<b>Run log</b> <span class="opt" id="logHint">(updates live)</span>
|
||||||
|
<pre class="logbox" id="logBox"><span class="empty-log">Waiting for output…</span></pre>
|
||||||
|
</div>
|
||||||
<div id="results"></div>
|
<div id="results"></div>
|
||||||
</main>
|
</main>
|
||||||
<div id="viewer">
|
<div id="viewer">
|
||||||
@@ -144,8 +187,106 @@
|
|||||||
const drop=document.getElementById('drop'), fileInput=document.getElementById('file'),
|
const drop=document.getElementById('drop'), fileInput=document.getElementById('file'),
|
||||||
runBtn=document.getElementById('run'), statusEl=document.getElementById('status'),
|
runBtn=document.getElementById('run'), statusEl=document.getElementById('status'),
|
||||||
results=document.getElementById('results'), dropLabel=document.getElementById('dropLabel'),
|
results=document.getElementById('results'), dropLabel=document.getElementById('dropLabel'),
|
||||||
emailEl=document.getElementById('email');
|
emailEl=document.getElementById('email'),
|
||||||
|
visionSel=document.getElementById('vision_model'),
|
||||||
|
textSel=document.getElementById('text_model'),
|
||||||
|
liveLog=document.getElementById('liveLog'),
|
||||||
|
logBox=document.getElementById('logBox'),
|
||||||
|
logHint=document.getElementById('logHint');
|
||||||
let chosen=null, polling=null, currentJobId=null, sheetPage={}, viewerZoom=1, reviewDirty=false;
|
let chosen=null, polling=null, currentJobId=null, sheetPage={}, viewerZoom=1, reviewDirty=false;
|
||||||
|
let defaultVisionModel='', defaultTextModel='';
|
||||||
|
|
||||||
|
function modelLabel(m){
|
||||||
|
// Include per-1M-token pricing when the catalog provides it.
|
||||||
|
let s=m.name||m.id;
|
||||||
|
if(m.prompt_usd_per_mtok!=null)
|
||||||
|
s+=' — $'+m.prompt_usd_per_mtok+' / $'+m.completion_usd_per_mtok+' per 1M tok';
|
||||||
|
return s;
|
||||||
|
}
|
||||||
|
|
||||||
|
function fillSelect(sel, items, preferred){
|
||||||
|
sel.innerHTML='';
|
||||||
|
(items||[]).forEach(m=>{
|
||||||
|
const opt=document.createElement('option');
|
||||||
|
opt.value=m.id; opt.textContent=modelLabel(m);
|
||||||
|
if(m.id===preferred) opt.selected=true;
|
||||||
|
sel.appendChild(opt);
|
||||||
|
});
|
||||||
|
if(!sel.options.length){
|
||||||
|
const opt=document.createElement('option');
|
||||||
|
opt.value=preferred||''; opt.textContent=preferred||'(no models)';
|
||||||
|
sel.appendChild(opt);
|
||||||
|
}
|
||||||
|
sel.disabled=false;
|
||||||
|
}
|
||||||
|
|
||||||
|
let modelsLoaded=false;
|
||||||
|
const MODELS_CACHE_KEY='cc_models_v1';
|
||||||
|
const MODELS_CACHE_TTL=24*60*60*1000;
|
||||||
|
function loadModelsCache(){
|
||||||
|
try{
|
||||||
|
const raw=localStorage.getItem(MODELS_CACHE_KEY);
|
||||||
|
if(!raw) return null;
|
||||||
|
const parsed=JSON.parse(raw);
|
||||||
|
if(!parsed.ts || Date.now()-parsed.ts > MODELS_CACHE_TTL) return null;
|
||||||
|
return parsed.data||null;
|
||||||
|
}catch(e){ return null; }
|
||||||
|
}
|
||||||
|
function saveModelsCache(data){
|
||||||
|
try{ localStorage.setItem(MODELS_CACHE_KEY, JSON.stringify({ts:Date.now(), data})); }
|
||||||
|
catch(e){}
|
||||||
|
}
|
||||||
|
function fetchWithTimeout(url, ms){
|
||||||
|
return Promise.race([
|
||||||
|
fetch(url, {cache:'no-store'}),
|
||||||
|
new Promise((_,reject)=>setTimeout(()=>reject(new Error('timeout')), ms))
|
||||||
|
]);
|
||||||
|
}
|
||||||
|
|
||||||
|
async function loadModels(){
|
||||||
|
const note=document.getElementById('modelNote');
|
||||||
|
const cached=loadModelsCache();
|
||||||
|
if(cached){
|
||||||
|
const defs=cached.defaults||{};
|
||||||
|
fillSelect(visionSel, cached.vision, defs.vision);
|
||||||
|
fillSelect(textSel, cached.text, defs.text);
|
||||||
|
modelsLoaded=true;
|
||||||
|
note.textContent='('+(cached.text||[]).length+' text / '+(cached.vision||[]).length+' vision cached)';
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
try{
|
||||||
|
note.textContent='fetching models...';
|
||||||
|
const res=await fetchWithTimeout('/models', 5000);
|
||||||
|
if(!res.ok) throw new Error('models HTTP '+res.status);
|
||||||
|
const data=await res.json();
|
||||||
|
saveModelsCache(data);
|
||||||
|
const defs=data.defaults||{};
|
||||||
|
fillSelect(visionSel, data.vision, defs.vision);
|
||||||
|
fillSelect(textSel, data.text, defs.text);
|
||||||
|
modelsLoaded=true;
|
||||||
|
note.textContent='('+(data.text||[]).length+' text / '+(data.vision||[]).length+' vision available)';
|
||||||
|
}catch(err){
|
||||||
|
const fallback=[{id:defaultVisionModel||'', name:defaultVisionModel||'(default)', prompt_usd_per_mtok:null, completion_usd_per_mtok:null}];
|
||||||
|
fillSelect(visionSel, fallback.filter(m=>m.id), defaultVisionModel);
|
||||||
|
const fallbackText=[{id:defaultTextModel||'', name:defaultTextModel||'(default)', prompt_usd_per_mtok:null, completion_usd_per_mtok:null}];
|
||||||
|
fillSelect(textSel, fallbackText.filter(m=>m.id), defaultTextModel);
|
||||||
|
visionSel.disabled=false; textSel.disabled=false;
|
||||||
|
note.textContent='using configured defaults ('+(err.message==='timeout'?'fetch timed out':'list unavailable')+')';
|
||||||
|
console.warn('Could not load models:', err);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function showLog(lines, live){
|
||||||
|
liveLog.style.display='block';
|
||||||
|
logHint.textContent=live?'(updates live)':'(saved with this job)';
|
||||||
|
const arr=lines||[];
|
||||||
|
if(!arr.length){
|
||||||
|
logBox.innerHTML='<span class="empty-log">No log lines yet…</span>';
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
logBox.textContent=arr.join('\n');
|
||||||
|
logBox.scrollTop=logBox.scrollHeight;
|
||||||
|
}
|
||||||
|
|
||||||
function setFile(f){ chosen=f; dropLabel.textContent=f?('Selected: '+f.name):'Drop a PDF drawing set here, or click to choose';
|
function setFile(f){ chosen=f; dropLabel.textContent=f?('Selected: '+f.name):'Drop a PDF drawing set here, or click to choose';
|
||||||
runBtn.disabled=!f; }
|
runBtn.disabled=!f; }
|
||||||
@@ -160,6 +301,7 @@ runBtn.addEventListener('click',async e=>{
|
|||||||
if(!chosen) return;
|
if(!chosen) return;
|
||||||
runBtn.disabled=true; results.innerHTML='';
|
runBtn.disabled=true; results.innerHTML='';
|
||||||
statusEl.innerHTML='<span class="spinner"></span>Uploading...';
|
statusEl.innerHTML='<span class="spinner"></span>Uploading...';
|
||||||
|
showLog([], true);
|
||||||
const fd=new FormData(); fd.append('file',chosen);
|
const fd=new FormData(); fd.append('file',chosen);
|
||||||
const email=(emailEl.value||'').trim(); if(email) fd.append('notification_email',email);
|
const email=(emailEl.value||'').trim(); if(email) fd.append('notification_email',email);
|
||||||
['project_name','address','occupancy','work_type'].forEach(id=>{
|
['project_name','address','occupancy','work_type'].forEach(id=>{
|
||||||
@@ -168,8 +310,9 @@ runBtn.addEventListener('click',async e=>{
|
|||||||
const compute=(document.querySelector('input[name="compute"]:checked')||{}).value;
|
const compute=(document.querySelector('input[name="compute"]:checked')||{}).value;
|
||||||
fd.append('text_local', compute==='local' ? 'true' : 'false');
|
fd.append('text_local', compute==='local' ? 'true' : 'false');
|
||||||
if(compute==='openrouter'){
|
if(compute==='openrouter'){
|
||||||
const modelSel=document.getElementById('model');
|
// Model picks only apply to OpenRouter compute; hybrid keeps its local model.
|
||||||
if(modelSel.value) fd.append('model', modelSel.value);
|
if(visionSel.value) fd.append('vision_model', visionSel.value);
|
||||||
|
if(textSel.value) fd.append('text_model', textSel.value);
|
||||||
}
|
}
|
||||||
const pipelineMode=(document.querySelector('input[name="pipeline_mode"]:checked')||{}).value||'classic';
|
const pipelineMode=(document.querySelector('input[name="pipeline_mode"]:checked')||{}).value||'classic';
|
||||||
fd.append('pipeline_mode',pipelineMode);
|
fd.append('pipeline_mode',pipelineMode);
|
||||||
@@ -196,21 +339,30 @@ function poll(jobId){
|
|||||||
const res=await fetch('/jobs/'+jobId);
|
const res=await fetch('/jobs/'+jobId);
|
||||||
if(!res.ok) throw new Error('job not found');
|
if(!res.ok) throw new Error('job not found');
|
||||||
const job=await res.json();
|
const job=await res.json();
|
||||||
|
const live=['running','queued','finalizing'].includes(job.status);
|
||||||
|
if(job.log_tail && job.log_tail.length) showLog(job.log_tail, live);
|
||||||
if(job.status==='running'||job.status==='queued'){
|
if(job.status==='running'||job.status==='queued'){
|
||||||
statusEl.innerHTML='<span class="spinner"></span>'+esc(job.stage||'Working...')+
|
statusEl.innerHTML='<span class="spinner"></span>'+esc(job.stage||'Working...')+
|
||||||
' · you can leave this page';
|
' · you can leave this page';
|
||||||
} else if(job.status==='done'){
|
} else if(job.status==='done'){
|
||||||
clearInterval(polling); polling=null; runBtn.disabled=false; render(job.report);
|
clearInterval(polling); polling=null; runBtn.disabled=false;
|
||||||
|
if(job.log && job.log.length) showLog(job.log, false);
|
||||||
|
render(job.report);
|
||||||
} else if(job.status==='needs_review'||job.status==='reviewing'){
|
} else if(job.status==='needs_review'||job.status==='reviewing'){
|
||||||
clearInterval(polling); polling=null; runBtn.disabled=false; renderReview(job);
|
clearInterval(polling); polling=null; runBtn.disabled=false;
|
||||||
|
if(job.log && job.log.length) showLog(job.log, false);
|
||||||
|
renderReview(job);
|
||||||
} else if(job.status==='finalizing'){
|
} else if(job.status==='finalizing'){
|
||||||
statusEl.innerHTML='<span class="spinner"></span>Finalizing reviewed report...';
|
statusEl.innerHTML='<span class="spinner"></span>Finalizing reviewed report...';
|
||||||
} else if(job.status==='finalization_error'){
|
} else if(job.status==='finalization_error'){
|
||||||
clearInterval(polling); polling=null; runBtn.disabled=false;
|
clearInterval(polling); polling=null; runBtn.disabled=false;
|
||||||
statusEl.textContent='Finalization failed: '+(job.error||'unknown error');
|
statusEl.textContent='Finalization failed: '+(job.error||'unknown error');
|
||||||
|
if(job.log && job.log.length) showLog(job.log, false);
|
||||||
} else if(job.status==='error'){
|
} else if(job.status==='error'){
|
||||||
clearInterval(polling); polling=null; runBtn.disabled=false;
|
clearInterval(polling); polling=null; runBtn.disabled=false;
|
||||||
statusEl.textContent='Run failed: '+(job.error||'unknown error');
|
statusEl.textContent='Run failed: '+(job.error||'unknown error');
|
||||||
|
if(job.log && job.log.length) showLog(job.log, false);
|
||||||
|
else if(job.log_tail && job.log_tail.length) showLog(job.log_tail, false);
|
||||||
}
|
}
|
||||||
}catch(err){ clearInterval(polling); polling=null; runBtn.disabled=false;
|
}catch(err){ clearInterval(polling); polling=null; runBtn.disabled=false;
|
||||||
statusEl.textContent='Error: '+err.message; }
|
statusEl.textContent='Error: '+err.message; }
|
||||||
@@ -224,35 +376,19 @@ function escAttr(s){ return esc(s).replace(/"/g,'"'); }
|
|||||||
function syncPipelineOptions(){
|
function syncPipelineOptions(){
|
||||||
const agent=(document.querySelector('input[name="pipeline_mode"]:checked')||{}).value==='agent';
|
const agent=(document.querySelector('input[name="pipeline_mode"]:checked')||{}).value==='agent';
|
||||||
const local=document.querySelector('input[name="compute"][value="local"]');
|
const local=document.querySelector('input[name="compute"][value="local"]');
|
||||||
|
if(local){
|
||||||
local.disabled=agent;
|
local.disabled=agent;
|
||||||
if(agent&&local.checked) document.querySelector('input[name="compute"][value="openrouter"]').checked=true;
|
if(agent&&local.checked) document.querySelector('input[name="compute"][value="openrouter"]').checked=true;
|
||||||
|
}
|
||||||
}
|
}
|
||||||
document.querySelectorAll('input[name="pipeline_mode"]').forEach(el=>el.addEventListener('change',syncPipelineOptions));
|
document.querySelectorAll('input[name="pipeline_mode"]').forEach(el=>el.addEventListener('change',syncPipelineOptions));
|
||||||
syncPipelineOptions();
|
syncPipelineOptions();
|
||||||
|
|
||||||
// --- model picker (OpenRouter compute only) ---
|
// --- model pickers (OpenRouter compute only) ---
|
||||||
let modelList=null;
|
|
||||||
async function loadModels(){
|
|
||||||
const note=document.getElementById('modelNote'), sel=document.getElementById('model');
|
|
||||||
try{
|
|
||||||
const res=await fetch('/models');
|
|
||||||
if(!res.ok) throw new Error('list unavailable');
|
|
||||||
const data=await res.json();
|
|
||||||
modelList=data.models||[];
|
|
||||||
sel.innerHTML=modelList.map(m=>
|
|
||||||
'<option value="'+escAttr(m.id)+'"'+(m.id===data.default?' selected':'')+'>'+
|
|
||||||
esc(m.name||m.id)+' — $'+esc(m.prompt_usd_per_mtok)+' / $'+esc(m.completion_usd_per_mtok)+
|
|
||||||
' per 1M tok</option>').join('');
|
|
||||||
note.textContent='('+modelList.length+' available)';
|
|
||||||
}catch(e){
|
|
||||||
sel.innerHTML='';
|
|
||||||
note.textContent='using configured default (list unavailable)';
|
|
||||||
}
|
|
||||||
}
|
|
||||||
function syncCompute(){
|
function syncCompute(){
|
||||||
const openrouter=(document.querySelector('input[name="compute"]:checked')||{}).value==='openrouter';
|
const openrouter=(document.querySelector('input[name="compute"]:checked')||{}).value==='openrouter';
|
||||||
document.getElementById('modelPick').style.display=openrouter?'block':'none';
|
document.getElementById('modelPick').style.display=openrouter?'block':'none';
|
||||||
if(openrouter&&!modelList) loadModels();
|
if(openrouter&&!modelsLoaded) loadModels();
|
||||||
}
|
}
|
||||||
document.querySelectorAll('input[name="compute"]').forEach(el=>el.addEventListener('change',syncCompute));
|
document.querySelectorAll('input[name="compute"]').forEach(el=>el.addEventListener('change',syncCompute));
|
||||||
syncCompute();
|
syncCompute();
|
||||||
@@ -360,7 +496,7 @@ function render(rep){
|
|||||||
const issues=rep.validated_issues||[];
|
const issues=rep.validated_issues||[];
|
||||||
if(issues.length){
|
if(issues.length){
|
||||||
html+='<details open style="margin-top:24px"><summary><b>QAQC issues ('+issues.length+')</b> '+
|
html+='<details open style="margin-top:24px"><summary><b>QAQC issues ('+issues.length+')</b> '+
|
||||||
'<span class="opt">conflicts + full-set + code/ADA + constructability, deduplicated</span></summary>';
|
'<span class="opt">conflicts + drawing integrity + full-set + constructability, deduplicated</span></summary>';
|
||||||
for(const c of issues){
|
for(const c of issues){
|
||||||
const rs=c.review_state;
|
const rs=c.review_state;
|
||||||
html+='<div class="conflict '+esc(c.severity)+'">'+
|
html+='<div class="conflict '+esc(c.severity)+'">'+
|
||||||
@@ -394,6 +530,10 @@ function render(rep){
|
|||||||
html+='</details>';
|
html+='</details>';
|
||||||
}
|
}
|
||||||
|
|
||||||
|
html+='<div class="meta" style="margin-top:14px">Full run log: <a href="/jobs/'+
|
||||||
|
esc(currentJobId)+'/log?plain=1" target="_blank" rel="noopener">/jobs/'+
|
||||||
|
esc(currentJobId)+'/log</a> (also saved as job.log on the server)</div>';
|
||||||
|
|
||||||
results.innerHTML=html;
|
results.innerHTML=html;
|
||||||
}
|
}
|
||||||
function stat(v,l){ return '<div class="stat"><b>'+esc(v)+'</b><span>'+esc(l)+'</span></div>'; }
|
function stat(v,l){ return '<div class="stat"><b>'+esc(v)+'</b><span>'+esc(l)+'</span></div>'; }
|
||||||
@@ -419,6 +559,12 @@ async function renderReview(job){
|
|||||||
esc(prog.completed||0)+' of '+esc(prog.required||0)+' required items decided.'+
|
esc(prog.completed||0)+' of '+esc(prog.required||0)+' required items decided.'+
|
||||||
((prog.remaining||0)>0?' Decide all blocking items, save, then finalize.':
|
((prog.remaining||0)>0?' Decide all blocking items, save, then finalize.':
|
||||||
' All required items decided \u2014 you can finalize.')+'</div>';
|
' All required items decided \u2014 you can finalize.')+'</div>';
|
||||||
|
html+='<div class="conflict" style="border-left-color:var(--accent)">'+
|
||||||
|
'<div class="row"><span class="cat">Ask about this run</span></div>'+
|
||||||
|
'<div class="meta">Questions about coverage or about the run as a whole — '+
|
||||||
|
'e.g. "why didn\'t it pick up the Civil set?". Answers are explanations only; '+
|
||||||
|
'they never change a finding.</div>'+
|
||||||
|
chatPanelHtml('','Ask a question about this run',true)+'</div>';
|
||||||
const blocking=queue.filter(i=>i.blocking), audit=queue.filter(i=>!i.blocking);
|
const blocking=queue.filter(i=>i.blocking), audit=queue.filter(i=>!i.blocking);
|
||||||
blocking.forEach((item,i)=>{ html+=reviewItemHtml(item,'b'+i,prior[item.review_item_id]); });
|
blocking.forEach((item,i)=>{ html+=reviewItemHtml(item,'b'+i,prior[item.review_item_id]); });
|
||||||
if(audit.length){
|
if(audit.length){
|
||||||
@@ -431,7 +577,11 @@ async function renderReview(job){
|
|||||||
'<button class="btn" id="saveReviewBtn">Save decisions</button> '+
|
'<button class="btn" id="saveReviewBtn">Save decisions</button> '+
|
||||||
'<button class="btn" id="finalizeBtn"'+((prog.remaining||0)===0?'':' disabled')+
|
'<button class="btn" id="finalizeBtn"'+((prog.remaining||0)===0?'':' disabled')+
|
||||||
'>Finalize & send report</button></div>'+
|
'>Finalize & send report</button></div>'+
|
||||||
'<div class="status" id="reviewMsg"></div>';
|
'<div class="status" id="reviewMsg"></div>'+
|
||||||
|
'<div class="meta" style="margin-top:10px">Chat transcript: '+
|
||||||
|
'<a href="/jobs/'+esc(jobId)+'/review-chat/log" target="_blank" rel="noopener">'+
|
||||||
|
'/jobs/'+esc(jobId)+'/review-chat/log</a> '+
|
||||||
|
'(also saved as review/chat_log.jsonl on the server)</div>';
|
||||||
results.innerHTML=html;
|
results.innerHTML=html;
|
||||||
results.querySelectorAll('.review-item input[type=radio]').forEach(r=>{
|
results.querySelectorAll('.review-item input[type=radio]').forEach(r=>{
|
||||||
r.addEventListener('change',()=>syncReviewControls(r.closest('.review-item')));
|
r.addEventListener('change',()=>syncReviewControls(r.closest('.review-item')));
|
||||||
@@ -442,6 +592,8 @@ async function renderReview(job){
|
|||||||
});
|
});
|
||||||
document.getElementById('saveReviewBtn').addEventListener('click',saveReviewDecisions);
|
document.getElementById('saveReviewBtn').addEventListener('click',saveReviewDecisions);
|
||||||
document.getElementById('finalizeBtn').addEventListener('click',finalizeReview);
|
document.getElementById('finalizeBtn').addEventListener('click',finalizeReview);
|
||||||
|
wireChatPanels();
|
||||||
|
loadChatHistory();
|
||||||
}
|
}
|
||||||
|
|
||||||
function reviewItemHtml(item,uid,prev){
|
function reviewItemHtml(item,uid,prev){
|
||||||
@@ -467,6 +619,9 @@ function reviewItemHtml(item,uid,prev){
|
|||||||
esc(e.source_text)+'"</div>').join('')+'</div>';
|
esc(e.source_text)+'"</div>').join('')+'</div>';
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
if((p.sheets||[]).length){
|
||||||
|
html+='<div class="meta">Sheets: '+sheetList(p.sheets)+'</div>';
|
||||||
|
}
|
||||||
if((item.reasons||[]).length){
|
if((item.reasons||[]).length){
|
||||||
html+='<div class="meta">Review triggers: '+esc(item.reasons.join(', '))+'</div>';
|
html+='<div class="meta">Review triggers: '+esc(item.reasons.join(', '))+'</div>';
|
||||||
}
|
}
|
||||||
@@ -480,10 +635,126 @@ function reviewItemHtml(item,uid,prev){
|
|||||||
'<input type="text" class="comment" placeholder="Comment (optional)" value="'+escAttr(prev.comment||'')+'">'+
|
'<input type="text" class="comment" placeholder="Comment (optional)" value="'+escAttr(prev.comment||'')+'">'+
|
||||||
'<input type="text" class="clar'+(prev.decision==='needs_clarification'?'':' hidden')+
|
'<input type="text" class="clar'+(prev.decision==='needs_clarification'?'':' hidden')+
|
||||||
'" placeholder="Clarification answer" value="'+escAttr(prev.clarification_answer||'')+'">'+
|
'" placeholder="Clarification answer" value="'+escAttr(prev.clarification_answer||'')+'">'+
|
||||||
'</div></div>';
|
'</div>'+
|
||||||
|
chatPanelHtml(item.review_item_id,'Ask about this finding')+
|
||||||
|
'</div>';
|
||||||
return html;
|
return html;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- review chat (read-only run explainer) ---
|
||||||
|
// One panel per review item plus one run-scope panel. Panels are keyed by
|
||||||
|
// review_item_id ('' = run scope); history for every panel is fetched once.
|
||||||
|
|
||||||
|
function chatPanelHtml(key,label,openByDefault){
|
||||||
|
const open=!!openByDefault;
|
||||||
|
return '<div class="chat" data-chat="'+escAttr(key||'')+'">'+
|
||||||
|
'<button type="button" class="chat-toggle">'+esc(label)+'</button>'+
|
||||||
|
'<div class="chat-body'+(open?'':' hidden')+'">'+
|
||||||
|
'<div class="chat-turns"></div>'+
|
||||||
|
'<div class="chat-row">'+
|
||||||
|
'<textarea rows="2" placeholder="Why did it conclude that?"></textarea>'+
|
||||||
|
'<button type="button" class="btn sm chat-send">Ask</button>'+
|
||||||
|
'</div>'+
|
||||||
|
'<div class="chat-hint">Explains what the run did, from its own artifacts. '+
|
||||||
|
'It cannot change this finding or your decision.</div>'+
|
||||||
|
'<div class="chat-err"></div>'+
|
||||||
|
'</div></div>';
|
||||||
|
}
|
||||||
|
|
||||||
|
function chatTurnHtml(turn){
|
||||||
|
let html='<div class="chat-q"><b>You:</b> '+esc(turn.question||'')+'</div>'+
|
||||||
|
'<div class="chat-a">'+esc(turn.answer||'');
|
||||||
|
if((turn.findings||[]).length){
|
||||||
|
html+='<ul>'+turn.findings.map(f=>'<li>'+esc(f)+'</li>').join('')+'</ul>';
|
||||||
|
}
|
||||||
|
(turn.evidence_cited||[]).forEach(c=>{
|
||||||
|
html+='<div class="chat-cite"><code>'+esc(c.artifact||'?')+'</code>'+
|
||||||
|
(c.sheet?' ('+esc(c.sheet)+')':'')+
|
||||||
|
(c.quote?': "'+esc(c.quote)+'"':'')+
|
||||||
|
(c.why_it_matters?' — '+esc(c.why_it_matters):'')+'</div>';
|
||||||
|
});
|
||||||
|
const tags=[];
|
||||||
|
if(turn.answerable&&turn.answerable!=='yes') tags.push('answerable: '+turn.answerable);
|
||||||
|
if(turn.assessment_of_finding&&turn.assessment_of_finding!=='not_applicable')
|
||||||
|
tags.push(turn.assessment_of_finding.replace(/_/g,' '));
|
||||||
|
if(turn.confidence) tags.push('confidence: '+turn.confidence);
|
||||||
|
if(turn.missing_information) tags.push('missing: '+turn.missing_information);
|
||||||
|
if(turn.suggested_category_correction)
|
||||||
|
tags.push('correction noted: '+turn.suggested_category_correction);
|
||||||
|
if(tags.length) html+='<div class="chat-tags">'+esc(tags.join(' \u00b7 '))+'</div>';
|
||||||
|
return html+'</div>';
|
||||||
|
}
|
||||||
|
|
||||||
|
function renderChatTurns(panel,turns){
|
||||||
|
const box=panel.querySelector('.chat-turns');
|
||||||
|
box.innerHTML=turns.length?turns.map(chatTurnHtml).join('')
|
||||||
|
:'<div class="meta">No questions asked yet.</div>';
|
||||||
|
box.scrollTop=box.scrollHeight;
|
||||||
|
}
|
||||||
|
|
||||||
|
async function loadChatHistory(){
|
||||||
|
let data;
|
||||||
|
try{
|
||||||
|
const res=await fetch('/jobs/'+currentJobId+'/review-chat');
|
||||||
|
if(!res.ok) return; // chat unavailable for this job: leave panels empty
|
||||||
|
data=await res.json();
|
||||||
|
}catch(err){ return; }
|
||||||
|
const byKey={};
|
||||||
|
(data.turns||[]).forEach(t=>{
|
||||||
|
const k=t.review_item_id||'';
|
||||||
|
(byKey[k]=byKey[k]||[]).push(t);
|
||||||
|
});
|
||||||
|
results.querySelectorAll('.chat').forEach(panel=>{
|
||||||
|
const key=panel.getAttribute('data-chat')||'';
|
||||||
|
renderChatTurns(panel,byKey[key]||[]);
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
function wireChatPanels(){
|
||||||
|
results.querySelectorAll('.chat').forEach(panel=>{
|
||||||
|
panel.querySelector('.chat-toggle').addEventListener('click',()=>{
|
||||||
|
panel.querySelector('.chat-body').classList.toggle('hidden');
|
||||||
|
});
|
||||||
|
const send=panel.querySelector('.chat-send');
|
||||||
|
const box=panel.querySelector('textarea');
|
||||||
|
send.addEventListener('click',()=>askChat(panel));
|
||||||
|
// Enter sends, Shift+Enter newlines - the questions are usually one line.
|
||||||
|
box.addEventListener('keydown',e=>{
|
||||||
|
if(e.key==='Enter'&&!e.shiftKey){ e.preventDefault(); askChat(panel); }
|
||||||
|
});
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async function askChat(panel){
|
||||||
|
const box=panel.querySelector('textarea');
|
||||||
|
const send=panel.querySelector('.chat-send');
|
||||||
|
const err=panel.querySelector('.chat-err');
|
||||||
|
const question=box.value.trim();
|
||||||
|
err.textContent='';
|
||||||
|
if(!question) return;
|
||||||
|
const key=panel.getAttribute('data-chat')||'';
|
||||||
|
const body={question:question};
|
||||||
|
if(key) body.review_item_id=key;
|
||||||
|
send.disabled=true; send.textContent='Asking...';
|
||||||
|
try{
|
||||||
|
const res=await fetch('/jobs/'+currentJobId+'/review-chat',
|
||||||
|
{method:'POST',headers:{'Content-Type':'application/json'},
|
||||||
|
body:JSON.stringify(body)});
|
||||||
|
if(!res.ok){
|
||||||
|
const e=await res.json().catch(()=>({detail:res.statusText}));
|
||||||
|
const detail=typeof e.detail==='string'?e.detail:JSON.stringify(e.detail);
|
||||||
|
throw new Error(detail||'Request failed');
|
||||||
|
}
|
||||||
|
const data=await res.json();
|
||||||
|
const turnsBox=panel.querySelector('.chat-turns');
|
||||||
|
if(turnsBox.querySelector('.meta')) turnsBox.innerHTML='';
|
||||||
|
turnsBox.insertAdjacentHTML('beforeend',chatTurnHtml(data.turn||{}));
|
||||||
|
turnsBox.scrollTop=turnsBox.scrollHeight;
|
||||||
|
box.value='';
|
||||||
|
}catch(e){ err.textContent=e.message; }
|
||||||
|
finally{ send.disabled=false; send.textContent='Ask'; }
|
||||||
|
}
|
||||||
|
|
||||||
function syncReviewControls(el){
|
function syncReviewControls(el){
|
||||||
const sel=el.querySelector('input[type=radio]:checked');
|
const sel=el.querySelector('input[type=radio]:checked');
|
||||||
const v=sel?sel.value:'';
|
const v=sel?sel.value:'';
|
||||||
@@ -566,10 +837,17 @@ async function finalizeReview(){
|
|||||||
|
|
||||||
// If opened from an email link (/?job=<id>), load that job's results directly.
|
// If opened from an email link (/?job=<id>), load that job's results directly.
|
||||||
(function init(){
|
(function init(){
|
||||||
// Show the deployed build in the header so it's obvious which version is up.
|
// Fetch /health first: build tag for the header and default models as a
|
||||||
fetch('/health').then(r=>r.ok?r.json():null).then(h=>{
|
// fallback if the larger /models catalog fails or times out.
|
||||||
if(h&&h.build) document.getElementById('buildTag').textContent=' · build '+h.build;
|
fetch('/health', {cache:'no-store'}).then(r=>r.ok?r.json():null).then(h=>{
|
||||||
}).catch(()=>{});
|
if(h){
|
||||||
|
if(h.build) document.getElementById('buildTag').textContent=' · build '+h.build;
|
||||||
|
if(h.model) defaultVisionModel=h.model;
|
||||||
|
if(h.text_model) defaultTextModel=h.text_model;
|
||||||
|
}
|
||||||
|
}).catch(()=>{}).finally(()=>{
|
||||||
|
loadModels();
|
||||||
|
});
|
||||||
const jobId=new URLSearchParams(location.search).get('job');
|
const jobId=new URLSearchParams(location.search).get('job');
|
||||||
if(jobId){ statusEl.innerHTML='<span class="spinner"></span>Loading job '+esc(jobId)+'...'; poll(jobId); }
|
if(jobId){ statusEl.innerHTML='<span class="spinner"></span>Loading job '+esc(jobId)+'...'; poll(jobId); }
|
||||||
})();
|
})();
|
||||||
|
|||||||
@@ -2,6 +2,7 @@ fastapi==0.115.0
|
|||||||
uvicorn[standard]==0.30.6
|
uvicorn[standard]==0.30.6
|
||||||
python-multipart==0.0.12
|
python-multipart==0.0.12
|
||||||
pdf2image==1.17.0
|
pdf2image==1.17.0
|
||||||
|
PyMuPDF>=1.23.0 # deterministic text-layer extraction (extractor grounding, verifier crops)
|
||||||
Pillow==10.4.0
|
Pillow==10.4.0
|
||||||
openai==1.51.0
|
openai==1.51.0
|
||||||
httpx==0.27.2 # openai 1.51 passes proxies= to httpx; >=0.28 dropped it
|
httpx==0.27.2 # openai 1.51 passes proxies= to httpx; >=0.28 dropped it
|
||||||
|
|||||||
@@ -0,0 +1,198 @@
|
|||||||
|
"""Wave 6.5 Brain-directed clarification: planning unit + runner integration."""
|
||||||
|
|
||||||
|
import backend.agents.brain as brain_mod
|
||||||
|
import backend.agents.runner as runner_mod
|
||||||
|
from backend import config
|
||||||
|
from backend.agents.base import AgentResult
|
||||||
|
from backend.agents.base import AgentUsage
|
||||||
|
from backend.agents.brain import BrainAgent
|
||||||
|
from backend.agents.runner import run_agent_pipeline
|
||||||
|
|
||||||
|
|
||||||
|
# --- plan_clarifications unit tests ---------------------------------------
|
||||||
|
|
||||||
|
def _finding(issue_id, **kw):
|
||||||
|
base = {"issue_id": issue_id, "severity": "medium", "confidence": "low",
|
||||||
|
"source_stage": "conflict", "description": "d",
|
||||||
|
"evidence": [{"sheet": "A1", "source_text": "x"}]}
|
||||||
|
base.update(kw)
|
||||||
|
return base
|
||||||
|
|
||||||
|
|
||||||
|
def test_plan_caps_and_filters_unknown_ids(monkeypatch):
|
||||||
|
monkeypatch.setattr(config, "BRAIN_CLARIFY_MAX_REQUESTS", 2)
|
||||||
|
monkeypatch.setattr(brain_mod, "call_json", lambda **k: {"requests": [
|
||||||
|
{"issue_id": "A", "request_type": "verify_evidence", "reason": "thin"},
|
||||||
|
{"issue_id": "GHOST", "request_type": "verify_evidence", "reason": "x"},
|
||||||
|
{"issue_id": "B", "request_type": "verify_evidence", "reason": "amb"},
|
||||||
|
{"issue_id": "C", "request_type": "verify_evidence", "reason": "over cap"},
|
||||||
|
]})
|
||||||
|
prioritized = [_finding("A"), _finding("B"), _finding("C")]
|
||||||
|
reqs = BrainAgent(AgentUsage()).plan_clarifications(prioritized)
|
||||||
|
ids = [r["issue_id"] for r in reqs]
|
||||||
|
assert ids == ["A", "B"] # GHOST filtered, capped at 2
|
||||||
|
|
||||||
|
|
||||||
|
def test_plan_skips_already_verified(monkeypatch):
|
||||||
|
monkeypatch.setattr(config, "BRAIN_CLARIFY_MAX_REQUESTS", 8)
|
||||||
|
captured = {}
|
||||||
|
|
||||||
|
def fake(**kwargs):
|
||||||
|
captured["user_text"] = kwargs["user_text"]
|
||||||
|
return {"requests": [
|
||||||
|
{"issue_id": "A", "request_type": "verify_evidence", "reason": "y"},
|
||||||
|
]}
|
||||||
|
|
||||||
|
monkeypatch.setattr(brain_mod, "call_json", fake)
|
||||||
|
prioritized = [
|
||||||
|
_finding("A"),
|
||||||
|
_finding("V", verification={"status": "confirmed", "verdicts": []}),
|
||||||
|
]
|
||||||
|
reqs = BrainAgent(AgentUsage()).plan_clarifications(prioritized)
|
||||||
|
assert [r["issue_id"] for r in reqs] == ["A"]
|
||||||
|
# The already-verified finding must not even be offered to the model.
|
||||||
|
assert '"V"' not in captured["user_text"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_plan_empty_on_call_failure(monkeypatch):
|
||||||
|
monkeypatch.setattr(config, "BRAIN_CLARIFY_MAX_REQUESTS", 8)
|
||||||
|
|
||||||
|
def boom(**k):
|
||||||
|
raise RuntimeError("brain down")
|
||||||
|
monkeypatch.setattr(brain_mod, "call_json", boom)
|
||||||
|
assert BrainAgent(AgentUsage()).plan_clarifications([_finding("A")]) == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_plan_no_requests_returns_empty(monkeypatch):
|
||||||
|
monkeypatch.setattr(config, "BRAIN_CLARIFY_MAX_REQUESTS", 8)
|
||||||
|
monkeypatch.setattr(brain_mod, "call_json", lambda **k: {"requests": []})
|
||||||
|
assert BrainAgent(AgentUsage()).plan_clarifications([_finding("A")]) == []
|
||||||
|
|
||||||
|
|
||||||
|
# --- runner-level integration ---------------------------------------------
|
||||||
|
|
||||||
|
def _stub_agent(artifacts):
|
||||||
|
return lambda usage: type("S", (), {
|
||||||
|
"name": "stub",
|
||||||
|
"run": lambda self, scope: AgentResult(
|
||||||
|
scope_id=scope.scope_id, artifacts=list(artifacts)),
|
||||||
|
})()
|
||||||
|
|
||||||
|
|
||||||
|
def _patch_pipeline(monkeypatch, brain_finding, plan_requests):
|
||||||
|
monkeypatch.setattr(
|
||||||
|
runner_mod, "convert_pdf_to_images",
|
||||||
|
lambda path: [{"page_number": 1, "base64": "QUJD"}])
|
||||||
|
monkeypatch.setattr(runner_mod, "SheetExtractorAgent", _stub_agent([
|
||||||
|
{"sheet_number": "S401", "page_number": 1, "level": "roof",
|
||||||
|
"discipline": "S", "assertions": [
|
||||||
|
{"text": "(2) 2x6 STUD PACK", "object_type": "framing"},
|
||||||
|
{"text": "HSS16X4 beam", "object_type": "framing"},
|
||||||
|
]},
|
||||||
|
]))
|
||||||
|
monkeypatch.setattr(runner_mod, "SheetIndexAgent", _stub_agent([{}]))
|
||||||
|
monkeypatch.setattr(runner_mod, "JurisdictionAgent", _stub_agent([{}]))
|
||||||
|
monkeypatch.setattr(runner_mod, "LinkerAgent", _stub_agent([
|
||||||
|
{"key": "c1", "location": "roof beam pocket", "assertions": []},
|
||||||
|
]))
|
||||||
|
# A filler conflict finding so memory["findings"] is non-empty and wave 6
|
||||||
|
# actually invokes Brain.run (which our stub replaces with brain_finding).
|
||||||
|
monkeypatch.setattr(runner_mod, "ConflictCriticAgent", _stub_agent([
|
||||||
|
{"issue_id": "FILLER", "severity": "low", "confidence": "low",
|
||||||
|
"source_stage": "conflict", "sheets": [], "description": "filler",
|
||||||
|
"evidence": []},
|
||||||
|
]))
|
||||||
|
monkeypatch.setattr(runner_mod, "CodeAgent", _stub_agent([]))
|
||||||
|
monkeypatch.setattr(runner_mod, "ConstructabilityAgent", _stub_agent([]))
|
||||||
|
monkeypatch.setattr(runner_mod, "CompletenessAgent", _stub_agent([]))
|
||||||
|
monkeypatch.setattr(runner_mod, "DrawingIntegrityAgent", _stub_agent([]))
|
||||||
|
# Brain.run returns our finding; plan_clarifications returns the requests.
|
||||||
|
monkeypatch.setattr(
|
||||||
|
runner_mod, "BrainAgent",
|
||||||
|
lambda usage: type("B", (), {
|
||||||
|
"run": lambda self, findings, si, ju: ([dict(brain_finding)], []),
|
||||||
|
"plan_clarifications": lambda self, prioritized: list(plan_requests),
|
||||||
|
})())
|
||||||
|
|
||||||
|
|
||||||
|
def test_brain_clarify_refutes_and_suppresses(monkeypatch, tmp_path):
|
||||||
|
"""Brain flags a MEDIUM finding wave-5b's severity gate skipped; the
|
||||||
|
clarification verifier refutes it, so it moves to suppressed_issues."""
|
||||||
|
monkeypatch.setattr(config, "ENABLE_BRAIN_CLARIFY", True)
|
||||||
|
finding = {
|
||||||
|
"issue_id": "M1", "severity": "medium", "confidence": "low",
|
||||||
|
"source_stage": "conflict", "sheets": ["S401"],
|
||||||
|
"description": "beam bears on (2) 2x6 stud pack",
|
||||||
|
"evidence": [{"sheet": "S401", "source_text": "(2) 2x6 STUD PACK"}],
|
||||||
|
}
|
||||||
|
_patch_pipeline(monkeypatch, finding, [
|
||||||
|
{"issue_id": "M1", "request_type": "verify_evidence", "reason": "misread?"},
|
||||||
|
])
|
||||||
|
# The clarification verifier returns a 'corrected' verdict -> refuted.
|
||||||
|
monkeypatch.setattr(
|
||||||
|
"backend.agents.verifier.call_json",
|
||||||
|
lambda **kwargs: {"verdicts": [
|
||||||
|
{"sheet": "S401", "source_text": "(2) 2x6 STUD PACK",
|
||||||
|
"verdict": "corrected", "actual_text": "(5) 2x6 STUD PACK",
|
||||||
|
"notes": "reads (5)"},
|
||||||
|
]})
|
||||||
|
pdf = tmp_path / "d.pdf"
|
||||||
|
pdf.write_bytes(b"%PDF-1.4\n")
|
||||||
|
report = run_agent_pipeline(str(pdf), out_dir=str(tmp_path),
|
||||||
|
require_review=False)
|
||||||
|
validated = report.get("validated_issues") or []
|
||||||
|
assert all(f.get("issue_id") != "M1" for f in validated) # dropped
|
||||||
|
assert [f["issue_id"] for f in report["suppressed_issues"]] == ["M1"]
|
||||||
|
assert report["suppressed_issues"][0]["verification"]["status"] == "refuted"
|
||||||
|
|
||||||
|
|
||||||
|
def test_brain_clarify_confirms_keeps_finding(monkeypatch, tmp_path):
|
||||||
|
monkeypatch.setattr(config, "ENABLE_BRAIN_CLARIFY", True)
|
||||||
|
finding = {
|
||||||
|
"issue_id": "M2", "severity": "medium", "confidence": "low",
|
||||||
|
"source_stage": "conflict", "sheets": ["S401"],
|
||||||
|
"description": "beam bears on (5) 2x6 stud pack",
|
||||||
|
"evidence": [{"sheet": "S401", "source_text": "(5) 2x6 STUD PACK"}],
|
||||||
|
}
|
||||||
|
_patch_pipeline(monkeypatch, finding, [
|
||||||
|
{"issue_id": "M2", "request_type": "verify_evidence", "reason": "check"},
|
||||||
|
])
|
||||||
|
monkeypatch.setattr(
|
||||||
|
"backend.agents.verifier.call_json",
|
||||||
|
lambda **kwargs: {"verdicts": [
|
||||||
|
{"sheet": "S401", "source_text": "(5) 2x6 STUD PACK",
|
||||||
|
"verdict": "confirmed", "actual_text": None, "notes": None},
|
||||||
|
]})
|
||||||
|
pdf = tmp_path / "d.pdf"
|
||||||
|
pdf.write_bytes(b"%PDF-1.4\n")
|
||||||
|
report = run_agent_pipeline(str(pdf), out_dir=str(tmp_path),
|
||||||
|
require_review=False)
|
||||||
|
validated = report.get("validated_issues") or []
|
||||||
|
kept = [f for f in validated if f.get("issue_id") == "M2"]
|
||||||
|
assert len(kept) == 1
|
||||||
|
assert kept[0]["verification"]["status"] == "confirmed"
|
||||||
|
assert report["suppressed_issues"] == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_brain_clarify_disabled_is_noop(monkeypatch, tmp_path):
|
||||||
|
monkeypatch.setattr(config, "ENABLE_BRAIN_CLARIFY", False)
|
||||||
|
finding = {
|
||||||
|
"issue_id": "M3", "severity": "medium", "confidence": "low",
|
||||||
|
"source_stage": "conflict", "sheets": ["S401"],
|
||||||
|
"description": "d", "evidence": [{"sheet": "S401", "source_text": "t"}],
|
||||||
|
}
|
||||||
|
# plan_clarifications should never be consulted; give it a bomb to prove it.
|
||||||
|
def _bomb(self, prioritized):
|
||||||
|
raise AssertionError("plan_clarifications must not run when disabled")
|
||||||
|
_patch_pipeline(monkeypatch, finding, [])
|
||||||
|
monkeypatch.setattr(
|
||||||
|
runner_mod, "BrainAgent",
|
||||||
|
lambda usage: type("B", (), {
|
||||||
|
"run": lambda self, findings, si, ju: ([dict(finding)], []),
|
||||||
|
"plan_clarifications": _bomb})())
|
||||||
|
pdf = tmp_path / "d.pdf"
|
||||||
|
pdf.write_bytes(b"%PDF-1.4\n")
|
||||||
|
report = run_agent_pipeline(str(pdf), out_dir=str(tmp_path),
|
||||||
|
require_review=False)
|
||||||
|
validated = report.get("validated_issues") or []
|
||||||
|
assert any(f.get("issue_id") == "M3" for f in validated)
|
||||||
@@ -0,0 +1,51 @@
|
|||||||
|
"""Classic pipeline path must also satisfy the {disputes} placeholder added to
|
||||||
|
CONSTRUCTABILITY_USER_INSTRUCTION (agent path substitutes it in construct_agent.py;
|
||||||
|
the classic stage builds its own subs dict)."""
|
||||||
|
|
||||||
|
from unittest.mock import patch
|
||||||
|
|
||||||
|
from backend.pipeline._stage import render
|
||||||
|
from backend.pipeline.constructability import constructability_review
|
||||||
|
from backend.prompts import CONSTRUCTABILITY_USER_INSTRUCTION
|
||||||
|
|
||||||
|
|
||||||
|
def _cluster_with_dispute():
|
||||||
|
return {
|
||||||
|
"key": "c1",
|
||||||
|
"assertions": [],
|
||||||
|
"disputed_attributes": [{
|
||||||
|
"attribute": "stud_pack_size",
|
||||||
|
"values": ["(2) 2x6 STUD PACK", "(5) 2x6 STUD PACK"],
|
||||||
|
"assertion_ids": ["a1", "a2"],
|
||||||
|
}],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def test_classic_constructability_supplies_disputes_sub():
|
||||||
|
captured = {}
|
||||||
|
|
||||||
|
def fake_call_stage(system_prompt, user_instruction, subs=None, **kwargs):
|
||||||
|
captured["subs"] = subs or {}
|
||||||
|
return {"issues": []}
|
||||||
|
|
||||||
|
with patch("backend.pipeline.constructability.call_stage", fake_call_stage):
|
||||||
|
constructability_review([], [_cluster_with_dispute()], [])
|
||||||
|
|
||||||
|
assert "disputes" in captured["subs"], "classic path must substitute {disputes}"
|
||||||
|
rendered = render(CONSTRUCTABILITY_USER_INSTRUCTION, captured["subs"])
|
||||||
|
assert "{disputes}" not in rendered
|
||||||
|
assert "(5) 2x6 STUD PACK" in rendered
|
||||||
|
|
||||||
|
|
||||||
|
def test_classic_constructability_disputes_defaults_empty():
|
||||||
|
captured = {}
|
||||||
|
|
||||||
|
def fake_call_stage(system_prompt, user_instruction, subs=None, **kwargs):
|
||||||
|
captured["subs"] = subs or {}
|
||||||
|
return {"issues": []}
|
||||||
|
|
||||||
|
with patch("backend.pipeline.constructability.call_stage", fake_call_stage):
|
||||||
|
constructability_review([], [{"key": "c2", "assertions": []}], [])
|
||||||
|
|
||||||
|
rendered = render(CONSTRUCTABILITY_USER_INSTRUCTION, captured["subs"])
|
||||||
|
assert "{disputes}" not in rendered
|
||||||
@@ -0,0 +1,69 @@
|
|||||||
|
from backend.agents.disputes import annotate_clusters, find_disputes
|
||||||
|
|
||||||
|
|
||||||
|
def _a(id_, attribute, value):
|
||||||
|
return {"id": id_, "attribute": attribute, "value": value,
|
||||||
|
"source_text": value}
|
||||||
|
|
||||||
|
|
||||||
|
def test_find_disputes_flags_same_attribute_different_values():
|
||||||
|
assertions = [
|
||||||
|
_a("a1", "stud_pack_size", "(2) 2x6 STUD PACK"),
|
||||||
|
_a("a2", "stud_pack_size", "(5) 2x6 STUD PACK"),
|
||||||
|
_a("a3", "beam_size", "HSS16X4X5/8"),
|
||||||
|
]
|
||||||
|
disputes = find_disputes(assertions)
|
||||||
|
assert len(disputes) == 1
|
||||||
|
assert disputes[0]["attribute"] == "stud_pack_size"
|
||||||
|
assert disputes[0]["values"] == ["(2) 2x6 STUD PACK", "(5) 2x6 STUD PACK"]
|
||||||
|
assert disputes[0]["assertion_ids"] == ["a1", "a2"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_find_disputes_ignores_agreeing_values_and_blanks():
|
||||||
|
assertions = [
|
||||||
|
_a("a1", "beam_size", "HSS16X4X5/8"),
|
||||||
|
_a("a2", "beam_size", " hss16x4x5/8 "), # same after normalize
|
||||||
|
_a("a3", "", "orphan"), # no attribute -> skipped
|
||||||
|
_a("a4", "beam_size", ""), # no value -> skipped
|
||||||
|
]
|
||||||
|
assert find_disputes(assertions) == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_annotate_clusters_writes_disputed_attributes():
|
||||||
|
clusters = [
|
||||||
|
{"key": "c1", "assertions": [
|
||||||
|
_a("a1", "stud_pack_size", "(2) 2x6"),
|
||||||
|
_a("a2", "stud_pack_size", "(5) 2x6"),
|
||||||
|
]},
|
||||||
|
{"key": "c2", "assertions": [_a("a3", "x", "1"), _a("a4", "x", "1")]},
|
||||||
|
]
|
||||||
|
assert annotate_clusters(clusters) == 1
|
||||||
|
assert clusters[0]["disputed_attributes"][0]["attribute"] == "stud_pack_size"
|
||||||
|
assert "disputed_attributes" not in clusters[1]
|
||||||
|
|
||||||
|
|
||||||
|
def test_slim_clusters_preserves_disputed_attributes():
|
||||||
|
from backend.pipeline._serialize import slim_clusters
|
||||||
|
cluster = {"key": "c1", "assertions": [],
|
||||||
|
"disputed_attributes": [{"attribute": "a", "values": ["1", "2"],
|
||||||
|
"assertion_ids": ["x", "y"]}]}
|
||||||
|
slim = slim_clusters([cluster])[0]
|
||||||
|
assert slim["disputed_attributes"][0]["values"] == ["1", "2"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_find_disputes_handles_none_and_zero_values():
|
||||||
|
# None value/attribute -> skipped; numeric 0 is a real value, not blank
|
||||||
|
assertions = [
|
||||||
|
{"id": "a1", "attribute": "count", "value": 0},
|
||||||
|
{"id": "a2", "attribute": "count", "value": 1},
|
||||||
|
{"id": "a3", "attribute": None, "value": "x"},
|
||||||
|
{"id": "a4", "attribute": "count", "value": None},
|
||||||
|
]
|
||||||
|
disputes = find_disputes(assertions)
|
||||||
|
assert len(disputes) == 1
|
||||||
|
assert disputes[0]["values"] == ["0", "1"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_find_disputes_empty_input():
|
||||||
|
assert find_disputes([]) == []
|
||||||
|
assert annotate_clusters([]) == 0
|
||||||
@@ -0,0 +1,95 @@
|
|||||||
|
"""Unit tests for the per-sheet DrawingIntegrityAgent and scope builder."""
|
||||||
|
|
||||||
|
import backend.agents.integrity_agent as integ
|
||||||
|
from backend.agents.base import AgentScope, AgentUsage
|
||||||
|
from backend.agents.integrity_agent import (
|
||||||
|
DrawingIntegrityAgent, build_integrity_scopes,
|
||||||
|
)
|
||||||
|
from backend import config
|
||||||
|
|
||||||
|
|
||||||
|
def _sheet(page, sheet_number, n_assertions):
|
||||||
|
return {
|
||||||
|
"sheet_number": sheet_number,
|
||||||
|
"page_number": page,
|
||||||
|
"sheet_title": f"Sheet {sheet_number}",
|
||||||
|
"discipline": "Architectural",
|
||||||
|
"assertions": [
|
||||||
|
{"object_type": "note", "source_text": f"note {i}"}
|
||||||
|
for i in range(n_assertions)
|
||||||
|
],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def test_build_scopes_skips_sparse_sheets(monkeypatch):
|
||||||
|
monkeypatch.setattr(config, "INTEGRITY_MIN_ASSERTIONS", 3)
|
||||||
|
sheets = [
|
||||||
|
_sheet(1, "A101", 5), # kept
|
||||||
|
_sheet(2, "A102", 2), # skipped (too sparse)
|
||||||
|
_sheet(3, "A103", 3), # kept (== floor)
|
||||||
|
]
|
||||||
|
page_to_b64 = {1: "IMG1", 2: "IMG2", 3: "IMG3"}
|
||||||
|
page_to_text = {1: "text one", 2: "text two", 3: "text three"}
|
||||||
|
scopes = build_integrity_scopes(sheets, page_to_b64, page_to_text)
|
||||||
|
ids = sorted(s.scope_id for s in scopes)
|
||||||
|
assert ids == ["integrity:1", "integrity:3"]
|
||||||
|
# Each scope carries its own page image + text layer.
|
||||||
|
by_id = {s.scope_id: s for s in scopes}
|
||||||
|
assert by_id["integrity:1"].payload["image_b64"] == "IMG1"
|
||||||
|
assert by_id["integrity:1"].payload["text_layer"] == "text one"
|
||||||
|
|
||||||
|
|
||||||
|
def test_agent_parses_and_anchors_sheet(monkeypatch):
|
||||||
|
"""Findings with blank sheets get anchored to the scope's sheet number."""
|
||||||
|
captured = {}
|
||||||
|
|
||||||
|
def fake_call_json(**kwargs):
|
||||||
|
captured.update(kwargs)
|
||||||
|
return {"issues": [
|
||||||
|
{"issue_id": "DI-1", "severity": "high", "confidence": "high",
|
||||||
|
"category": "dangling_reference", "sheets": [],
|
||||||
|
"description": "Detail callout 5/A101 has no detail 5 on this sheet",
|
||||||
|
"evidence": [{"sheet": "A101", "source_text": "5/A101"}]},
|
||||||
|
]}
|
||||||
|
|
||||||
|
monkeypatch.setattr(integ, "call_json", fake_call_json)
|
||||||
|
sheet = _sheet(1, "A101", 5)
|
||||||
|
scope = AgentScope("integrity:1", {
|
||||||
|
"sheet": sheet, "page_number": 1,
|
||||||
|
"image_b64": "IMG1", "text_layer": "the deterministic text layer",
|
||||||
|
})
|
||||||
|
result = DrawingIntegrityAgent(AgentUsage()).run(scope)
|
||||||
|
assert result.error == ""
|
||||||
|
assert len(result.artifacts) == 1
|
||||||
|
finding = result.artifacts[0]
|
||||||
|
assert finding["source_stage"] == "drawing_integrity"
|
||||||
|
assert finding["sheets"] == ["A101"] # anchored
|
||||||
|
assert finding["agent"] == "drawing_integrity"
|
||||||
|
assert finding["scope_id"] == "integrity:1"
|
||||||
|
# The image + text layer reached the model.
|
||||||
|
assert captured["images_b64"] == ["IMG1"]
|
||||||
|
assert "the deterministic text layer" in captured["user_text"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_agent_empty_issues_is_clean(monkeypatch):
|
||||||
|
monkeypatch.setattr(integ, "call_json", lambda **k: {"issues": []})
|
||||||
|
scope = AgentScope("integrity:1", {
|
||||||
|
"sheet": _sheet(1, "A101", 5), "page_number": 1,
|
||||||
|
"image_b64": "IMG1", "text_layer": "t",
|
||||||
|
})
|
||||||
|
result = DrawingIntegrityAgent(AgentUsage()).run(scope)
|
||||||
|
assert result.error == ""
|
||||||
|
assert result.artifacts == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_agent_survives_call_failure(monkeypatch):
|
||||||
|
def boom(**kwargs):
|
||||||
|
raise RuntimeError("model exploded")
|
||||||
|
monkeypatch.setattr(integ, "call_json", boom)
|
||||||
|
scope = AgentScope("integrity:1", {
|
||||||
|
"sheet": _sheet(1, "A101", 5), "page_number": 1,
|
||||||
|
"image_b64": "IMG1", "text_layer": "t",
|
||||||
|
})
|
||||||
|
result = DrawingIntegrityAgent(AgentUsage()).run(scope)
|
||||||
|
assert "model exploded" in result.error
|
||||||
|
assert result.artifacts == []
|
||||||
@@ -0,0 +1,86 @@
|
|||||||
|
from backend import config
|
||||||
|
from backend.agents.base import AgentScope, AgentUsage
|
||||||
|
from backend.agents.extractors import SheetExtractorAgent
|
||||||
|
from backend.prompts import TEXT_STRUCTURING_SYSTEM_PROMPT, TEXT_STRUCTURING_USER_INSTRUCTION
|
||||||
|
|
||||||
|
def test_text_structuring_prompt_demands_verbatim_and_completeness():
|
||||||
|
assert "verbatim" in TEXT_STRUCTURING_USER_INSTRUCTION.lower()
|
||||||
|
assert "every" in TEXT_STRUCTURING_USER_INSTRUCTION.lower()
|
||||||
|
assert "{text_layer}" in TEXT_STRUCTURING_USER_INSTRUCTION
|
||||||
|
|
||||||
|
|
||||||
|
def _page(n=8, text="1. \nALL SAWN LUMBER IN CONTACT WITH SOIL TO BE SOUTHERN PINE, PRESSURE TREATED.\n2. \nROOF SHEATHING: 5/8\" PLYWOOD, C-D GRADE, STRUCTURAL I."):
|
||||||
|
return {"page_number": n, "base64": "AAAA", "text_layer": text}
|
||||||
|
|
||||||
|
|
||||||
|
def _run(agent, page, hint=""):
|
||||||
|
scope = AgentScope(scope_id=f"sheet:{page['page_number']}",
|
||||||
|
payload={"page": page, "sheet_hint": hint})
|
||||||
|
result = agent.run(scope)
|
||||||
|
assert not result.error, result.error
|
||||||
|
return result.artifacts[0]
|
||||||
|
|
||||||
|
|
||||||
|
def test_ladder_falls_back_when_vision_returns_nothing(monkeypatch):
|
||||||
|
# vision pass returns 1 summary object that the guard drops;
|
||||||
|
# text-structuring disabled to exercise the deterministic rung
|
||||||
|
monkeypatch.setattr("backend.agents.extractors.call_json",
|
||||||
|
lambda **kw: [{"name": "general notes", "value": "notes"}])
|
||||||
|
monkeypatch.setattr("backend.config.EXTRACT_TEXT_RETRY_ENABLED", False)
|
||||||
|
agent = SheetExtractorAgent(AgentUsage())
|
||||||
|
sheet = _run(agent, _page())
|
||||||
|
assert sheet["assertions"], "dark sheet must be impossible with fallback enabled"
|
||||||
|
assert all(a.get("grounding") == "text_layer_fallback" for a in sheet["assertions"])
|
||||||
|
assert sheet["coverage"]["ratio"] >= 0.6
|
||||||
|
|
||||||
|
|
||||||
|
def test_ladder_merge_preserves_graphical_objects(monkeypatch):
|
||||||
|
# vision finds a graphical symbol; text rung adds notes.
|
||||||
|
# The graphical object MUST survive the merge.
|
||||||
|
calls = {"n": 0}
|
||||||
|
def fake_call_json(**kw):
|
||||||
|
calls["n"] += 1
|
||||||
|
if kw.get("images_b64"): # vision pass
|
||||||
|
return {"sheet": {}, "objects": [
|
||||||
|
{"object_id": "g1", "object_type": "lighting_fixture",
|
||||||
|
"name": "pendant at grid C-4", "source_text": None,
|
||||||
|
"graphical_basis": "16in pendant symbol at grid C-4"}]}
|
||||||
|
return {"sheet": {}, "objects": [ # text-structuring pass
|
||||||
|
{"object_id": "t1", "object_type": "general_note",
|
||||||
|
"source_text": "ALL SAWN LUMBER IN CONTACT WITH SOIL TO BE SOUTHERN PINE, PRESSURE TREATED.",
|
||||||
|
"name": "lumber note"}]}
|
||||||
|
monkeypatch.setattr("backend.agents.extractors.call_json", fake_call_json)
|
||||||
|
agent = SheetExtractorAgent(AgentUsage())
|
||||||
|
sheet = _run(agent, _page())
|
||||||
|
assert calls["n"] >= 2, "text-structuring rung should have fired"
|
||||||
|
assert any(a.get("graphical_basis") for a in sheet["assertions"])
|
||||||
|
assert any("SAWN LUMBER" in (a.get("source_text") or "") for a in sheet["assertions"])
|
||||||
|
|
||||||
|
|
||||||
|
def test_ladder_recovers_sheet_number_from_text_layer(monkeypatch):
|
||||||
|
monkeypatch.setattr(
|
||||||
|
"backend.agents.extractors.call_json",
|
||||||
|
lambda **kw: {"sheet": {}, "objects": [
|
||||||
|
{"object_id": "o1", "name": "RCP note",
|
||||||
|
"source_text": "GYP. BD. CEILING 8'-11 3/8\" A.F.F. TYP. FOR ALL STOREFRONT",
|
||||||
|
"attributes": {"height": "8'-11 3/8\""}}]})
|
||||||
|
agent = SheetExtractorAgent(AgentUsage())
|
||||||
|
sheet = _run(agent, _page(18, "REFLECTED CEILING PLAN\nGYP. BD. CEILING 8'-11 3/8\" A.F.F. TYP. FOR ALL STOREFRONT\nA102"))
|
||||||
|
assert sheet["sheet_number"] == "A102"
|
||||||
|
|
||||||
|
|
||||||
|
def test_ladder_skips_retry_when_coverage_healthy(monkeypatch):
|
||||||
|
# vision covers every meaningful text-layer line -> no rung 2/3 calls
|
||||||
|
calls = {"n": 0}
|
||||||
|
def fake_call_json(**kw):
|
||||||
|
calls["n"] += 1
|
||||||
|
return {"sheet": {"sheet_number": "A101"}, "objects": [
|
||||||
|
{"object_id": "o1", "object_type": "general_note", "name": "lumber note",
|
||||||
|
"source_text": "ALL SAWN LUMBER IN CONTACT WITH SOIL TO BE SOUTHERN PINE, PRESSURE TREATED."},
|
||||||
|
{"object_id": "o2", "object_type": "general_note", "name": "sheathing note",
|
||||||
|
"source_text": "ROOF SHEATHING: 5/8\" PLYWOOD, C-D GRADE, STRUCTURAL I."}]}
|
||||||
|
monkeypatch.setattr("backend.agents.extractors.call_json", fake_call_json)
|
||||||
|
agent = SheetExtractorAgent(AgentUsage())
|
||||||
|
sheet = _run(agent, _page())
|
||||||
|
assert sheet["coverage"]["ratio"] >= config.EXTRACT_COVERAGE_FLOOR
|
||||||
|
assert calls["n"] == 1
|
||||||
@@ -0,0 +1,74 @@
|
|||||||
|
from backend.agents.base import AgentScope
|
||||||
|
from backend.agents.linker import build_link_scopes
|
||||||
|
|
||||||
|
|
||||||
|
def _sheet(number, page, level, assertions):
|
||||||
|
return {"sheet_number": number, "page_number": page,
|
||||||
|
"discipline": "Structural", "level": level,
|
||||||
|
"assertions": assertions}
|
||||||
|
|
||||||
|
|
||||||
|
def _assertion(id_, ref=None, tag=None, level=None):
|
||||||
|
return {"id": id_, "attribute": "stud_pack_size", "value": "(5) 2x6",
|
||||||
|
"source_text": "(5) 2x6 STUD PACK",
|
||||||
|
"location_key": {"detail_reference": ref, "tag": tag,
|
||||||
|
"level": level}}
|
||||||
|
|
||||||
|
|
||||||
|
def test_xref_scope_joins_same_detail_reference_across_levels():
|
||||||
|
sheets = [
|
||||||
|
_sheet("S101", 10, "foundation", [_assertion("a1", ref="A/S205")]),
|
||||||
|
_sheet("S205", 20, "roof", [_assertion("a2", ref="A/S205")]),
|
||||||
|
_sheet("S401", 30, "roof", [_assertion("a3", ref="A/S205")]),
|
||||||
|
]
|
||||||
|
scopes = build_link_scopes(sheets)
|
||||||
|
xref = [s for s in scopes if s.scope_id.startswith("xref:")]
|
||||||
|
assert xref, "expected a cross-level detail-reference scope"
|
||||||
|
ids = {a["id"] for s in xref for a in s.payload["assertions"]}
|
||||||
|
assert ids == {"a1", "a2", "a3"}
|
||||||
|
|
||||||
|
|
||||||
|
def test_xref_scope_requires_two_distinct_sheets():
|
||||||
|
sheets = [
|
||||||
|
_sheet("S401", 30, "roof", [_assertion("a1", ref="A/S205"),
|
||||||
|
_assertion("a2", ref="A/S205")]),
|
||||||
|
]
|
||||||
|
scopes = build_link_scopes(sheets)
|
||||||
|
assert not [s for s in scopes if s.scope_id.startswith("xref:")]
|
||||||
|
|
||||||
|
|
||||||
|
def test_xref_scope_joins_shared_member_tag():
|
||||||
|
sheets = [
|
||||||
|
_sheet("S102", 5, "roof", [_assertion("a1", tag="HSS16X4X5/8")]),
|
||||||
|
_sheet("S401", 30, "unknown", [_assertion("a2", tag="HSS16X4X5/8")]),
|
||||||
|
]
|
||||||
|
scopes = build_link_scopes(sheets)
|
||||||
|
xref = [s for s in scopes if s.scope_id.startswith("xref:")]
|
||||||
|
assert xref
|
||||||
|
|
||||||
|
|
||||||
|
def test_xref_scope_joins_single_letter_member_mark():
|
||||||
|
# W-shapes (W12X26) are the most common steel marks and have one leading letter
|
||||||
|
sheets = [
|
||||||
|
_sheet("S102", 5, "roof", [_assertion("a1", tag="W12X26")]),
|
||||||
|
_sheet("S401", 30, "unknown", [_assertion("a2", tag="W12X26")]),
|
||||||
|
]
|
||||||
|
scopes = build_link_scopes(sheets)
|
||||||
|
xref = [s for s in scopes if s.scope_id.startswith("xref:")]
|
||||||
|
assert xref, "single-letter member marks (W12X26) must join xref scopes"
|
||||||
|
|
||||||
|
|
||||||
|
def test_xref_scope_rechecks_sheet_diversity_after_cap(monkeypatch):
|
||||||
|
from backend import config
|
||||||
|
monkeypatch.setattr(config, "AGENT_LINK_MAX_ASSERTIONS", 2)
|
||||||
|
sheets = [
|
||||||
|
_sheet("S401", 30, "roof", [_assertion("a1", ref="A/S205"),
|
||||||
|
_assertion("a2", ref="A/S205")]),
|
||||||
|
_sheet("S205", 20, "roof", [_assertion("a3", ref="A/S205")]),
|
||||||
|
]
|
||||||
|
scopes = build_link_scopes(sheets)
|
||||||
|
xref = [s for s in scopes if s.scope_id.startswith("xref:")]
|
||||||
|
for scope in xref:
|
||||||
|
sheets_in_scope = {a["sheet_number"] for a in scope.payload["assertions"]}
|
||||||
|
assert len(sheets_in_scope) >= 2, \
|
||||||
|
"capped xref scope must still span two sheets"
|
||||||
@@ -0,0 +1,102 @@
|
|||||||
|
"""Flag-gating tests: ENABLE_CODE_REVIEW off skips code, drawing_integrity runs.
|
||||||
|
|
||||||
|
Runner-level smoke tests using stubbed agents (same pattern as
|
||||||
|
test_wave5b_suppression). Verifies the code/ADA wave is skipped when
|
||||||
|
ENABLE_CODE_REVIEW is false and the Drawing Integrity wave feeds findings
|
||||||
|
into the report by_stage counters.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import backend.agents.runner as runner_mod
|
||||||
|
from backend import config
|
||||||
|
from backend.agents.base import AgentResult
|
||||||
|
from backend.agents.runner import run_agent_pipeline
|
||||||
|
|
||||||
|
|
||||||
|
def _stub_agent(artifacts):
|
||||||
|
return lambda usage: type("S", (), {
|
||||||
|
"name": "stub",
|
||||||
|
"run": lambda self, scope: AgentResult(
|
||||||
|
scope_id=scope.scope_id, artifacts=list(artifacts)),
|
||||||
|
})()
|
||||||
|
|
||||||
|
|
||||||
|
def _integrity_finding():
|
||||||
|
return {
|
||||||
|
"issue_id": "DI-1", "severity": "high", "confidence": "high",
|
||||||
|
"source_stage": "drawing_integrity", "sheets": ["A101"],
|
||||||
|
"category": "dangling_reference",
|
||||||
|
"description": "Detail callout 5/A101 has no detail 5 on this sheet",
|
||||||
|
"evidence": [{"sheet": "A101", "source_text": "5/A101"}],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _patch(monkeypatch, code_should_raise):
|
||||||
|
monkeypatch.setattr(
|
||||||
|
runner_mod, "convert_pdf_to_images",
|
||||||
|
lambda path: [{"page_number": 1, "base64": "QUJD"}])
|
||||||
|
# Sheet has 3+ assertions so the integrity wave does NOT skip it.
|
||||||
|
monkeypatch.setattr(runner_mod, "SheetExtractorAgent", _stub_agent([
|
||||||
|
{"sheet_number": "A101", "page_number": 1, "level": "1",
|
||||||
|
"discipline": "A", "assertions": [
|
||||||
|
{"text": "5/A101", "object_type": "detail_marker"},
|
||||||
|
{"text": "ROOM 101", "object_type": "room"},
|
||||||
|
{"text": "DOOR 101A", "object_type": "door"},
|
||||||
|
]},
|
||||||
|
]))
|
||||||
|
monkeypatch.setattr(runner_mod, "SheetIndexAgent", _stub_agent([{}]))
|
||||||
|
monkeypatch.setattr(runner_mod, "JurisdictionAgent", _stub_agent([{}]))
|
||||||
|
monkeypatch.setattr(runner_mod, "LinkerAgent", _stub_agent([]))
|
||||||
|
monkeypatch.setattr(runner_mod, "ConflictCriticAgent", _stub_agent([]))
|
||||||
|
|
||||||
|
def code_boom(usage):
|
||||||
|
if code_should_raise:
|
||||||
|
raise AssertionError("CodeAgent must not run when gated off")
|
||||||
|
return _stub_agent([])(usage)
|
||||||
|
monkeypatch.setattr(runner_mod, "CodeAgent", code_boom)
|
||||||
|
|
||||||
|
monkeypatch.setattr(runner_mod, "DrawingIntegrityAgent",
|
||||||
|
_stub_agent([_integrity_finding()]))
|
||||||
|
monkeypatch.setattr(runner_mod, "ConstructabilityAgent", _stub_agent([]))
|
||||||
|
monkeypatch.setattr(runner_mod, "CompletenessAgent", _stub_agent([]))
|
||||||
|
monkeypatch.setattr(
|
||||||
|
runner_mod, "BrainAgent",
|
||||||
|
lambda usage: type("B", (), {
|
||||||
|
"run": lambda self, findings, si, ju: (list(findings), []),
|
||||||
|
"plan_clarifications": lambda self, prioritized: []})())
|
||||||
|
# Stub the wave-5b verifier so the high-severity integrity finding is
|
||||||
|
# confirmed (never a live network call).
|
||||||
|
monkeypatch.setattr(
|
||||||
|
"backend.agents.verifier.call_json",
|
||||||
|
lambda **kwargs: {"verdicts": [
|
||||||
|
{"sheet": "A101", "source_text": "5/A101",
|
||||||
|
"verdict": "confirmed", "actual_text": None, "notes": None},
|
||||||
|
]})
|
||||||
|
|
||||||
|
|
||||||
|
def test_code_gated_off_integrity_on(monkeypatch, tmp_path):
|
||||||
|
monkeypatch.setattr(config, "ENABLE_CODE_REVIEW", False)
|
||||||
|
monkeypatch.setattr(config, "ENABLE_DRAWING_INTEGRITY", True)
|
||||||
|
_patch(monkeypatch, code_should_raise=True)
|
||||||
|
pdf = tmp_path / "d.pdf"
|
||||||
|
pdf.write_bytes(b"%PDF-1.4\n")
|
||||||
|
report = run_agent_pipeline(str(pdf), out_dir=str(tmp_path),
|
||||||
|
require_review=False)
|
||||||
|
by_stage = report["summary"]["by_stage"]
|
||||||
|
assert by_stage["code"] == 0
|
||||||
|
assert by_stage["drawing_integrity"] == 1
|
||||||
|
# The integrity finding survived into the validated set.
|
||||||
|
assert any(f.get("issue_id") == "DI-1"
|
||||||
|
for f in report.get("validated_issues") or [])
|
||||||
|
|
||||||
|
|
||||||
|
def test_code_enabled_runs(monkeypatch, tmp_path):
|
||||||
|
monkeypatch.setattr(config, "ENABLE_CODE_REVIEW", True)
|
||||||
|
monkeypatch.setattr(config, "ENABLE_DRAWING_INTEGRITY", True)
|
||||||
|
_patch(monkeypatch, code_should_raise=False)
|
||||||
|
# build_code_scopes runs on the real sheet; CodeAgent is stubbed to []
|
||||||
|
pdf = tmp_path / "d.pdf"
|
||||||
|
pdf.write_bytes(b"%PDF-1.4\n")
|
||||||
|
report = run_agent_pipeline(str(pdf), out_dir=str(tmp_path),
|
||||||
|
require_review=False)
|
||||||
|
# No crash; integrity still reported.
|
||||||
|
assert report["summary"]["by_stage"]["drawing_integrity"] == 1
|
||||||
@@ -4,7 +4,7 @@ from backend.agents.runner import run_agent_pipeline
|
|||||||
|
|
||||||
def _patch_brain(monkeypatch):
|
def _patch_brain(monkeypatch):
|
||||||
monkeypatch.setattr("backend.agents.runner.convert_pdf_to_images", lambda path: [{"page_number": 1, "base64": "x"}])
|
monkeypatch.setattr("backend.agents.runner.convert_pdf_to_images", lambda path: [{"page_number": 1, "base64": "x"}])
|
||||||
monkeypatch.setattr("backend.agents.runner.BrainAgent", lambda usage: type("B", (), {"run": lambda self, findings, sheet_index, jurisdiction: ([{"issue_id": "AGENT-0001", "severity": "high", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"}], [])})())
|
monkeypatch.setattr("backend.agents.runner.BrainAgent", lambda usage: type("B", (), {"run": lambda self, findings, sheet_index, jurisdiction: ([{"issue_id": "AGENT-0001", "severity": "high", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"}], []), "plan_clarifications": lambda self, prioritized: []})())
|
||||||
|
|
||||||
|
|
||||||
def test_agent_runner_can_enter_review_mode(monkeypatch, tmp_path):
|
def test_agent_runner_can_enter_review_mode(monkeypatch, tmp_path):
|
||||||
|
|||||||
@@ -0,0 +1,79 @@
|
|||||||
|
"""SheetExtractorAgent fallback ladder tests (bare-list wrap + compact retry)."""
|
||||||
|
|
||||||
|
from unittest.mock import patch
|
||||||
|
|
||||||
|
from backend.agents.base import AgentScope, AgentUsage
|
||||||
|
from backend.agents.extractors import SheetExtractorAgent, _wrap_bare_list
|
||||||
|
|
||||||
|
|
||||||
|
def _scope():
|
||||||
|
return AgentScope(
|
||||||
|
scope_id="sheet:4",
|
||||||
|
payload={"page": {"page_number": 4, "base64": "QUJD"}},
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _objects(n=2):
|
||||||
|
return [
|
||||||
|
{
|
||||||
|
"object_id": f"obj-{i}",
|
||||||
|
"object_type": "equipment",
|
||||||
|
"category": "mechanical",
|
||||||
|
"name": f"RTU-{i}",
|
||||||
|
"source_text": f"RTU-{i}",
|
||||||
|
"confidence": "high",
|
||||||
|
}
|
||||||
|
for i in range(n)
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def test_wrap_bare_list_builds_sheet_envelope():
|
||||||
|
wrapped = _wrap_bare_list(_objects(3), page_number=4)
|
||||||
|
assert wrapped["sheet"] == {}
|
||||||
|
assert len(wrapped["objects"]) == 3
|
||||||
|
|
||||||
|
|
||||||
|
def test_wrap_bare_list_passes_dicts_and_none_through():
|
||||||
|
assert _wrap_bare_list({"sheet": {}, "objects": []}, 1) == {"sheet": {}, "objects": []}
|
||||||
|
assert _wrap_bare_list(None, 1) is None
|
||||||
|
|
||||||
|
|
||||||
|
def test_run_accepts_bare_list_response():
|
||||||
|
agent = SheetExtractorAgent(usage=AgentUsage())
|
||||||
|
with patch("backend.agents.extractors.call_json",
|
||||||
|
return_value=_objects(5)) as mock_call:
|
||||||
|
result = agent.run(_scope())
|
||||||
|
assert not result.error
|
||||||
|
assert len(result.artifacts) == 1
|
||||||
|
sheet = result.artifacts[0]
|
||||||
|
assert sheet["page_number"] == 4
|
||||||
|
assert len(sheet["assertions"]) == 5
|
||||||
|
# No compact retry needed when the first call yields data.
|
||||||
|
assert mock_call.call_count == 1
|
||||||
|
# Reasoning knobs are forwarded (None when config is blank in tests).
|
||||||
|
assert "reasoning_effort" in mock_call.call_args.kwargs
|
||||||
|
assert "reasoning_max_tokens" in mock_call.call_args.kwargs
|
||||||
|
|
||||||
|
|
||||||
|
def test_run_compact_retry_after_hard_failure():
|
||||||
|
agent = SheetExtractorAgent(usage=AgentUsage())
|
||||||
|
with patch("backend.agents.extractors.call_json",
|
||||||
|
side_effect=[None, {"sheet": {"sheet_number": "A102"},
|
||||||
|
"objects": _objects(2)}]) as mock_call:
|
||||||
|
result = agent.run(_scope())
|
||||||
|
assert not result.error
|
||||||
|
assert result.artifacts[0]["sheet_number"] == "A102"
|
||||||
|
assert mock_call.call_count == 2
|
||||||
|
# Second call carried the compact suffix.
|
||||||
|
assert "COMPACT RETRY" in mock_call.call_args_list[1].kwargs["user_text"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_run_returns_empty_sheet_after_both_attempts_miss():
|
||||||
|
agent = SheetExtractorAgent(usage=AgentUsage())
|
||||||
|
with patch("backend.agents.extractors.call_json", return_value=None) as mock_call:
|
||||||
|
result = agent.run(_scope())
|
||||||
|
# Coverage ladder: no text layer to rescue the page -> empty sheet,
|
||||||
|
# but no hard failure (the ladder replaced the old raise).
|
||||||
|
assert not result.error
|
||||||
|
assert result.artifacts[0]["assertions"] == []
|
||||||
|
assert mock_call.call_count == 2
|
||||||
@@ -0,0 +1,129 @@
|
|||||||
|
"""Runner-level text-layer flow: excerpt into verify scopes, hi-DPI crop
|
||||||
|
replacement with full-page fallback, and coverage-gap findings."""
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
fitz = pytest.importorskip("pymupdf")
|
||||||
|
|
||||||
|
import backend.agents.runner as runner_mod
|
||||||
|
from backend.agents.base import AgentResult
|
||||||
|
from backend.agents.runner import run_agent_pipeline
|
||||||
|
|
||||||
|
PAGE_TEXT = "(5) 2X6 STUD PACK AT BEARING"
|
||||||
|
|
||||||
|
|
||||||
|
def _make_pdf(path):
|
||||||
|
doc = fitz.open()
|
||||||
|
page = doc.new_page(width=612, height=792)
|
||||||
|
page.insert_text((72, 72), PAGE_TEXT, fontsize=11)
|
||||||
|
doc.save(str(path))
|
||||||
|
doc.close()
|
||||||
|
return str(path)
|
||||||
|
|
||||||
|
|
||||||
|
def _finding(sheets, evidence_text):
|
||||||
|
return {
|
||||||
|
"issue_id": "C1", "severity": "critical", "confidence": "high",
|
||||||
|
"source_stage": "constructability", "sheets": sheets,
|
||||||
|
"description": "stud pack conflict",
|
||||||
|
"evidence": [{"sheet": sheets[0], "source_text": evidence_text}],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _stub_agent(artifacts):
|
||||||
|
return lambda usage: type("S", (), {
|
||||||
|
"name": "stub",
|
||||||
|
"run": lambda self, scope: AgentResult(
|
||||||
|
scope_id=scope.scope_id, artifacts=list(artifacts)),
|
||||||
|
})()
|
||||||
|
|
||||||
|
|
||||||
|
def _patch_pipeline(monkeypatch, finding, verify_sink):
|
||||||
|
monkeypatch.setattr(
|
||||||
|
runner_mod, "convert_pdf_to_images",
|
||||||
|
lambda path: [{"page_number": 1, "base64": "QUJD"}])
|
||||||
|
monkeypatch.setattr(runner_mod, "SheetExtractorAgent", _stub_agent([
|
||||||
|
{"sheet_number": "S401", "page_number": 1, "level": "roof",
|
||||||
|
"discipline": "S", "assertions": [
|
||||||
|
{"text": "(5) 2X6 STUD PACK", "object_type": "framing"},
|
||||||
|
{"text": "HSS16X4 beam", "object_type": "framing"},
|
||||||
|
]},
|
||||||
|
]))
|
||||||
|
monkeypatch.setattr(runner_mod, "SheetIndexAgent", _stub_agent([{}]))
|
||||||
|
monkeypatch.setattr(runner_mod, "JurisdictionAgent", _stub_agent([{}]))
|
||||||
|
monkeypatch.setattr(runner_mod, "LinkerAgent", _stub_agent([
|
||||||
|
{"key": "c1", "location": "roof beam pocket", "assertions": []},
|
||||||
|
]))
|
||||||
|
monkeypatch.setattr(runner_mod, "ConflictCriticAgent", _stub_agent([]))
|
||||||
|
monkeypatch.setattr(runner_mod, "CodeAgent", _stub_agent([]))
|
||||||
|
monkeypatch.setattr(runner_mod, "ConstructabilityAgent",
|
||||||
|
_stub_agent([finding]))
|
||||||
|
monkeypatch.setattr(runner_mod, "CompletenessAgent", _stub_agent([]))
|
||||||
|
monkeypatch.setattr(
|
||||||
|
runner_mod, "BrainAgent",
|
||||||
|
lambda usage: type("B", (), {
|
||||||
|
"run": lambda self, findings, sheet_index, jurisdiction:
|
||||||
|
(list(findings), []),
|
||||||
|
"plan_clarifications": lambda self, prioritized: []})())
|
||||||
|
|
||||||
|
class _RecordingVerifier:
|
||||||
|
name = "verify"
|
||||||
|
|
||||||
|
def __init__(self, usage):
|
||||||
|
pass
|
||||||
|
|
||||||
|
def run(self, scope):
|
||||||
|
verify_sink.append(scope.payload)
|
||||||
|
return AgentResult(scope_id=scope.scope_id, artifacts=[{
|
||||||
|
"finding_index": scope.payload["finding_index"],
|
||||||
|
"status": "confirmed",
|
||||||
|
"verdicts": [],
|
||||||
|
}])
|
||||||
|
|
||||||
|
monkeypatch.setattr(runner_mod, "EvidenceVerifierAgent",
|
||||||
|
lambda usage: _RecordingVerifier(usage))
|
||||||
|
|
||||||
|
|
||||||
|
def test_verify_scope_carries_text_excerpt_and_crop(monkeypatch, tmp_path):
|
||||||
|
"""Evidence text matches the page text layer -> excerpt present and the
|
||||||
|
full-page image is replaced by a hi-DPI crop."""
|
||||||
|
sink = []
|
||||||
|
_patch_pipeline(monkeypatch,
|
||||||
|
_finding(["S401"], "(5) 2X6 STUD PACK AT BEARING"), sink)
|
||||||
|
pdf = _make_pdf(tmp_path / "set.pdf")
|
||||||
|
run_agent_pipeline(pdf, out_dir=str(tmp_path), require_review=False)
|
||||||
|
assert len(sink) == 1
|
||||||
|
payload = sink[0]
|
||||||
|
assert "2X6 STUD PACK" in payload["text_layer_excerpt"]
|
||||||
|
assert payload["images_b64"], "crop must never drop all images"
|
||||||
|
assert payload["images_b64"][0] != "QUJD", "expected crop, not full page"
|
||||||
|
|
||||||
|
|
||||||
|
def test_verify_scope_falls_back_to_full_page(monkeypatch, tmp_path):
|
||||||
|
"""Evidence text not in the text layer -> keep the full-page image."""
|
||||||
|
sink = []
|
||||||
|
_patch_pipeline(monkeypatch,
|
||||||
|
_finding(["S401"], "PENTHOUSE EXHAUST FAN EF-9"), sink)
|
||||||
|
pdf = _make_pdf(tmp_path / "set.pdf")
|
||||||
|
run_agent_pipeline(pdf, out_dir=str(tmp_path), require_review=False)
|
||||||
|
assert len(sink) == 1
|
||||||
|
assert sink[0]["images_b64"] == ["QUJD"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_coverage_gap_becomes_gap_finding(monkeypatch, tmp_path):
|
||||||
|
"""Text layer present but zero objects extracted -> failed-scope gap
|
||||||
|
finding survives into the report."""
|
||||||
|
sink = []
|
||||||
|
_patch_pipeline(monkeypatch, _finding(["S401"], PAGE_TEXT), sink)
|
||||||
|
# Extractor returns a sheet with NO objects despite a real text layer.
|
||||||
|
monkeypatch.setattr(runner_mod, "SheetExtractorAgent", _stub_agent([
|
||||||
|
{"sheet_number": "S401", "page_number": 1, "level": "roof",
|
||||||
|
"discipline": "S", "assertions": []},
|
||||||
|
]))
|
||||||
|
pdf = _make_pdf(tmp_path / "set.pdf")
|
||||||
|
report = run_agent_pipeline(pdf, out_dir=str(tmp_path),
|
||||||
|
require_review=False)
|
||||||
|
gaps = [f for f in (report.get("validated_issues") or [])
|
||||||
|
if f.get("category") == "analysis_gap"]
|
||||||
|
assert any("extraction gap" in (g.get("description") or "")
|
||||||
|
for g in gaps)
|
||||||
@@ -0,0 +1,68 @@
|
|||||||
|
from unittest.mock import patch
|
||||||
|
|
||||||
|
from backend.agents.base import AgentScope, AgentUsage
|
||||||
|
|
||||||
|
from backend.agents.verifier import (
|
||||||
|
EvidenceVerifierAgent, apply_verdicts, select_findings,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _finding(sev="critical", issue_id="i1", sheets=("S401",), cluster_key=None):
|
||||||
|
f = {"issue_id": issue_id, "severity": sev, "confidence": "high",
|
||||||
|
"source_stage": "constructability", "sheets": list(sheets),
|
||||||
|
"description": "HSS16x4 on (2) 2x6 STUD PACK is unbuildable",
|
||||||
|
"evidence": [{"sheet": "S401", "source_text": "(2) 2x6 STUD PACK",
|
||||||
|
"asserted_value": "3-inch width"}]}
|
||||||
|
if cluster_key:
|
||||||
|
f["cluster_key"] = cluster_key
|
||||||
|
return f
|
||||||
|
|
||||||
|
|
||||||
|
def test_select_findings_by_severity_and_dispute():
|
||||||
|
findings = [_finding("critical"), _finding("low", "i2"),
|
||||||
|
_finding("medium", "i3", cluster_key="c9")]
|
||||||
|
clusters = [{"key": "c9", "disputed_attributes": [{"attribute": "a"}]}]
|
||||||
|
selected = select_findings(findings, clusters, max_checks=20,
|
||||||
|
severities={"critical", "high"})
|
||||||
|
assert [f["issue_id"] for f in selected] == ["i1", "i3"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_select_findings_respects_cap():
|
||||||
|
findings = [_finding("critical", f"i{n}") for n in range(30)]
|
||||||
|
selected = select_findings(findings, [], max_checks=5,
|
||||||
|
severities={"critical"})
|
||||||
|
assert len(selected) == 5
|
||||||
|
|
||||||
|
|
||||||
|
def test_run_attaches_verdicts_and_marks_refuted():
|
||||||
|
agent = EvidenceVerifierAgent(usage=AgentUsage())
|
||||||
|
scope = AgentScope(scope_id="verify:0", payload={
|
||||||
|
"finding_index": 0,
|
||||||
|
"finding": _finding(),
|
||||||
|
"images_b64": ["QUJD"],
|
||||||
|
})
|
||||||
|
verdicts = {"verdicts": [
|
||||||
|
{"sheet": "S401", "source_text": "(2) 2x6 STUD PACK",
|
||||||
|
"verdict": "corrected", "actual_text": "(5) 2x6 STUD PACK",
|
||||||
|
"notes": "callout reads (5)"},
|
||||||
|
]}
|
||||||
|
with patch("backend.agents.verifier.call_json", return_value=verdicts):
|
||||||
|
result = agent.run(scope)
|
||||||
|
assert not result.error
|
||||||
|
artifact = result.artifacts[0]
|
||||||
|
assert artifact["finding_index"] == 0
|
||||||
|
assert artifact["status"] == "refuted" # no evidence confirmed
|
||||||
|
assert artifact["verdicts"][0]["actual_text"] == "(5) 2x6 STUD PACK"
|
||||||
|
|
||||||
|
|
||||||
|
def test_apply_verdicts_annotates_and_suppresses():
|
||||||
|
from backend.agents.base import AgentResult
|
||||||
|
findings = [_finding("critical", "i1"), _finding("high", "i2")]
|
||||||
|
results = [AgentResult(scope_id="verify:0", artifacts=[
|
||||||
|
{"finding_index": 0, "status": "refuted", "verdicts": []},
|
||||||
|
{"finding_index": 1, "status": "confirmed", "verdicts": []},
|
||||||
|
])]
|
||||||
|
suppressed = apply_verdicts(findings, results)
|
||||||
|
assert suppressed == [findings[0]]
|
||||||
|
assert findings[0]["verification"]["status"] == "refuted"
|
||||||
|
assert findings[1]["verification"]["status"] == "confirmed"
|
||||||
@@ -0,0 +1,87 @@
|
|||||||
|
"""Runner-level wave-5b tests: suppression path and zero-image guard."""
|
||||||
|
|
||||||
|
import backend.agents.runner as runner_mod
|
||||||
|
from backend.agents.base import AgentResult
|
||||||
|
from backend.agents.runner import run_agent_pipeline
|
||||||
|
|
||||||
|
|
||||||
|
def _finding(sheets):
|
||||||
|
return {
|
||||||
|
"issue_id": "C1", "severity": "critical", "confidence": "high",
|
||||||
|
"source_stage": "constructability", "sheets": sheets,
|
||||||
|
"description": "HSS16x4 on (2) 2x6 STUD PACK is unbuildable",
|
||||||
|
"evidence": [{"sheet": sheets[0], "source_text": "(2) 2x6 STUD PACK"}],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _stub_agent(artifacts):
|
||||||
|
return lambda usage: type("S", (), {
|
||||||
|
"name": "stub",
|
||||||
|
"run": lambda self, scope: AgentResult(
|
||||||
|
scope_id=scope.scope_id, artifacts=list(artifacts)),
|
||||||
|
})()
|
||||||
|
|
||||||
|
|
||||||
|
def _patch_pipeline(monkeypatch, finding):
|
||||||
|
monkeypatch.setattr(
|
||||||
|
runner_mod, "convert_pdf_to_images",
|
||||||
|
lambda path: [{"page_number": 1, "base64": "QUJD"}])
|
||||||
|
monkeypatch.setattr(runner_mod, "SheetExtractorAgent", _stub_agent([
|
||||||
|
{"sheet_number": "S401", "page_number": 1, "level": "roof",
|
||||||
|
"discipline": "S", "assertions": [
|
||||||
|
{"text": "(2) 2x6 STUD PACK", "object_type": "framing"},
|
||||||
|
{"text": "HSS16X4 beam", "object_type": "framing"},
|
||||||
|
]},
|
||||||
|
]))
|
||||||
|
monkeypatch.setattr(runner_mod, "SheetIndexAgent", _stub_agent([{}]))
|
||||||
|
monkeypatch.setattr(runner_mod, "JurisdictionAgent", _stub_agent([{}]))
|
||||||
|
monkeypatch.setattr(runner_mod, "LinkerAgent", _stub_agent([
|
||||||
|
{"key": "c1", "location": "roof beam pocket", "assertions": []},
|
||||||
|
]))
|
||||||
|
monkeypatch.setattr(runner_mod, "ConflictCriticAgent", _stub_agent([]))
|
||||||
|
monkeypatch.setattr(runner_mod, "CodeAgent", _stub_agent([]))
|
||||||
|
monkeypatch.setattr(runner_mod, "ConstructabilityAgent", _stub_agent([finding]))
|
||||||
|
monkeypatch.setattr(runner_mod, "CompletenessAgent", _stub_agent([]))
|
||||||
|
monkeypatch.setattr(
|
||||||
|
runner_mod, "BrainAgent",
|
||||||
|
lambda usage: type("B", (), {
|
||||||
|
"run": lambda self, findings, sheet_index, jurisdiction:
|
||||||
|
(list(findings), []),
|
||||||
|
"plan_clarifications": lambda self, prioritized: []})())
|
||||||
|
|
||||||
|
|
||||||
|
def test_refuted_finding_is_suppressed_not_crash(monkeypatch, tmp_path):
|
||||||
|
"""Regression: memory.replace("suppressed", ...) must not KeyError."""
|
||||||
|
_patch_pipeline(monkeypatch, _finding(["S401"]))
|
||||||
|
monkeypatch.setattr(
|
||||||
|
"backend.agents.verifier.call_json",
|
||||||
|
lambda **kwargs: {"verdicts": [
|
||||||
|
{"sheet": "S401", "source_text": "(2) 2x6 STUD PACK",
|
||||||
|
"verdict": "corrected", "actual_text": "(5) 2x6 STUD PACK",
|
||||||
|
"notes": "callout reads (5)"},
|
||||||
|
]})
|
||||||
|
pdf = tmp_path / "dummy.pdf"
|
||||||
|
pdf.write_bytes(b"%PDF-1.4\n")
|
||||||
|
report = run_agent_pipeline(str(pdf), out_dir=str(tmp_path),
|
||||||
|
require_review=False)
|
||||||
|
assert [f["issue_id"] for f in report["suppressed_issues"]] == ["C1"]
|
||||||
|
assert report["suppressed_issues"][0]["verification"]["status"] == "refuted"
|
||||||
|
|
||||||
|
|
||||||
|
def test_zero_image_finding_is_not_suppressed(monkeypatch, tmp_path):
|
||||||
|
"""A finding whose sheets resolve to no page images must not be judged
|
||||||
|
(and must never be refuted) without pixels."""
|
||||||
|
_patch_pipeline(monkeypatch, _finding(["S999"])) # no such sheet
|
||||||
|
monkeypatch.setattr(
|
||||||
|
"backend.agents.verifier.call_json",
|
||||||
|
lambda **kwargs: {"verdicts": [
|
||||||
|
{"sheet": "S999", "source_text": "(2) 2x6 STUD PACK",
|
||||||
|
"verdict": "not_found", "actual_text": None, "notes": None},
|
||||||
|
]})
|
||||||
|
pdf = tmp_path / "dummy.pdf"
|
||||||
|
pdf.write_bytes(b"%PDF-1.4\n")
|
||||||
|
report = run_agent_pipeline(str(pdf), out_dir=str(tmp_path),
|
||||||
|
require_review=False)
|
||||||
|
assert report["suppressed_issues"] == []
|
||||||
|
validated = report.get("validated_issues") or []
|
||||||
|
assert any(f.get("issue_id") == "C1" for f in validated)
|
||||||
@@ -60,18 +60,56 @@ def test_job_log_endpoint_serves_log_and_404s(job_env, monkeypatch):
|
|||||||
assert client.get("/jobs/nope/log").status_code == 404
|
assert client.get("/jobs/nope/log").status_code == 404
|
||||||
|
|
||||||
|
|
||||||
def test_model_override_set_and_cleared_around_run(job_env, monkeypatch):
|
def test_model_overrides_passed_to_classic_runner(job_env, monkeypatch):
|
||||||
from backend import llm
|
"""Classic mode: per-run picks travel as run_pipeline kwargs (the runner
|
||||||
|
sets and clears llm.set_model_overrides itself)."""
|
||||||
seen = {}
|
seen = {}
|
||||||
|
|
||||||
def fake_runner(pdf_path, **kwargs):
|
def fake_runner(pdf_path, **kwargs):
|
||||||
seen["override"] = llm._model_override
|
seen.update(kwargs)
|
||||||
return {"source": "set.pdf", "summary": {}}
|
return {"source": "set.pdf", "summary": {}}
|
||||||
|
|
||||||
monkeypatch.setattr("backend.jobs.run_pipeline", fake_runner)
|
monkeypatch.setattr("backend.jobs.run_pipeline", fake_runner)
|
||||||
jobs.create_job(str(job_env / "set.pdf"), "set.pdf",
|
jobs.create_job(str(job_env / "set.pdf"), "set.pdf",
|
||||||
pipeline_mode="classic", model="openai/gpt-4o")
|
pipeline_mode="classic",
|
||||||
|
vision_model="openai/gpt-4o", text_model="openai/gpt-4o-mini")
|
||||||
|
|
||||||
assert seen["override"] == "openai/gpt-4o"
|
assert seen["vision_model"] == "openai/gpt-4o"
|
||||||
assert llm._model_override is None # cleared after the run
|
assert seen["text_model"] == "openai/gpt-4o-mini"
|
||||||
|
|
||||||
|
|
||||||
|
def test_model_overrides_set_and_cleared_around_agent_run(job_env, monkeypatch):
|
||||||
|
"""Agent mode: the agent runner has no override params, so jobs.py sets
|
||||||
|
them module-level for the duration of the run."""
|
||||||
|
from backend import llm
|
||||||
|
|
||||||
|
seen = {}
|
||||||
|
|
||||||
|
def fake_agent_runner(pdf_path, **kwargs):
|
||||||
|
seen["vision"] = llm._vision_model_override
|
||||||
|
seen["text"] = llm._text_model_override
|
||||||
|
return {"source": "set.pdf", "summary": {}}
|
||||||
|
|
||||||
|
monkeypatch.setattr("backend.jobs.run_agent_pipeline", fake_agent_runner)
|
||||||
|
jobs.create_job(str(job_env / "set.pdf"), "set.pdf",
|
||||||
|
pipeline_mode="agent",
|
||||||
|
vision_model="openai/gpt-4o", text_model="openai/gpt-4o-mini")
|
||||||
|
|
||||||
|
assert seen["vision"] == "openai/gpt-4o"
|
||||||
|
assert seen["text"] == "openai/gpt-4o-mini"
|
||||||
|
assert llm._vision_model_override is None # cleared after the run
|
||||||
|
assert llm._text_model_override is None
|
||||||
|
|
||||||
|
|
||||||
|
def test_failed_run_logs_traceback(job_env, monkeypatch):
|
||||||
|
"""A crashed job must leave the traceback in job.log, not just str(e)."""
|
||||||
|
def boom(pdf_path, **kwargs):
|
||||||
|
raise RuntimeError("kaboom-stage-failure")
|
||||||
|
|
||||||
|
monkeypatch.setattr("backend.jobs.run_pipeline", boom)
|
||||||
|
job_id = jobs.create_job(str(job_env / "set.pdf"), "set.pdf", pipeline_mode="classic")
|
||||||
|
|
||||||
|
assert jobs._jobs[job_id]["status"] == "error"
|
||||||
|
content = (job_env / job_id / "job.log").read_text()
|
||||||
|
assert "Traceback (most recent call last)" in content
|
||||||
|
assert "RuntimeError: kaboom-stage-failure" in content
|
||||||
@@ -11,12 +11,23 @@ _PAYLOAD = {
|
|||||||
"name": "GPT-4o",
|
"name": "GPT-4o",
|
||||||
"pricing": {"prompt": "0.0000025", "completion": "0.00001"},
|
"pricing": {"prompt": "0.0000025", "completion": "0.00001"},
|
||||||
"context_length": 128000,
|
"context_length": 128000,
|
||||||
|
"architecture": {"input_modalities": ["text", "image"],
|
||||||
|
"output_modalities": ["text"]},
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"id": "google/gemini-2.5-pro",
|
"id": "google/gemini-2.5-pro",
|
||||||
"name": "Gemini 2.5 Pro",
|
"name": "Gemini 2.5 Pro",
|
||||||
"pricing": {"prompt": "0.00000125", "completion": "0.00001"},
|
"pricing": {"prompt": "0.00000125", "completion": "0.00001"},
|
||||||
"context_length": 1000000,
|
"context_length": 1000000,
|
||||||
|
"architecture": {"modality": "text+image->text"},
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "meta-llama/llama-3.1-70b-instruct",
|
||||||
|
"name": "Llama 3.1 70B Instruct",
|
||||||
|
"pricing": {"prompt": "0.0000005", "completion": "0.0000008"},
|
||||||
|
"context_length": 131072,
|
||||||
|
"architecture": {"input_modalities": ["text"],
|
||||||
|
"output_modalities": ["text"]},
|
||||||
},
|
},
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
@@ -34,14 +45,27 @@ def test_models_endpoint_normalizes_pricing(monkeypatch):
|
|||||||
response = client.get("/models")
|
response = client.get("/models")
|
||||||
assert response.status_code == 200
|
assert response.status_code == 200
|
||||||
body = response.json()
|
body = response.json()
|
||||||
assert body["default"] == config.MODEL
|
assert body["defaults"] == {"vision": config.MODEL, "text": config.TEXT_MODEL}
|
||||||
assert body["default_text"] == config.TEXT_MODEL
|
by_id = {m["id"]: m for m in body["text"]}
|
||||||
by_id = {m["id"]: m for m in body["models"]}
|
|
||||||
assert by_id["openai/gpt-4o"]["prompt_usd_per_mtok"] == 2.5
|
assert by_id["openai/gpt-4o"]["prompt_usd_per_mtok"] == 2.5
|
||||||
assert by_id["openai/gpt-4o"]["completion_usd_per_mtok"] == 10.0
|
assert by_id["openai/gpt-4o"]["completion_usd_per_mtok"] == 10.0
|
||||||
assert by_id["openai/gpt-4o"]["context_length"] == 128000
|
assert by_id["openai/gpt-4o"]["context_length"] == 128000
|
||||||
|
|
||||||
|
|
||||||
|
def test_models_endpoint_splits_vision_and_text(monkeypatch):
|
||||||
|
_reset_cache()
|
||||||
|
monkeypatch.setattr(models, "_fetch_openrouter_models", lambda: _PAYLOAD["data"])
|
||||||
|
client = TestClient(app)
|
||||||
|
body = client.get("/models").json()
|
||||||
|
vision_ids = {m["id"] for m in body["vision"]}
|
||||||
|
text_ids = {m["id"] for m in body["text"]}
|
||||||
|
# Both modality shapes (structured and legacy string) are recognized.
|
||||||
|
assert vision_ids == {"openai/gpt-4o", "google/gemini-2.5-pro"}
|
||||||
|
# Text list is the full catalog; vision models appear in both.
|
||||||
|
assert text_ids == {"openai/gpt-4o", "google/gemini-2.5-pro",
|
||||||
|
"meta-llama/llama-3.1-70b-instruct"}
|
||||||
|
|
||||||
|
|
||||||
def test_models_endpoint_caches(monkeypatch):
|
def test_models_endpoint_caches(monkeypatch):
|
||||||
_reset_cache()
|
_reset_cache()
|
||||||
calls = []
|
calls = []
|
||||||
|
|||||||
@@ -0,0 +1,139 @@
|
|||||||
|
"""API tests for the review-chat endpoints."""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
from fastapi.testclient import TestClient
|
||||||
|
|
||||||
|
from backend.main import app
|
||||||
|
from backend.review.store import ReviewStore
|
||||||
|
|
||||||
|
|
||||||
|
def _queue_item() -> dict:
|
||||||
|
return {"review_item_id": "finding:AGENT-0007", "kind": "finding",
|
||||||
|
"blocking": True, "reasons": ["severity_high"],
|
||||||
|
"payload": {"issue_id": "AGENT-0007", "category": "elevation_disagreement",
|
||||||
|
"severity": "high", "sheets": ["M2.1"],
|
||||||
|
"description": "AC-1 at grade vs roof.",
|
||||||
|
"scope_id": "conflict:roof-ac1"}}
|
||||||
|
|
||||||
|
|
||||||
|
def _reply() -> dict:
|
||||||
|
return {"answer": "It read 'AC-1 MOUNTED ON GRADE' off M2.1.",
|
||||||
|
"findings": ["The grade value came from M2.1."],
|
||||||
|
"evidence_cited": [], "answerable": "yes",
|
||||||
|
"assessment_of_finding": "looks_supported", "confidence": "high"}
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
def job(monkeypatch, tmp_path):
|
||||||
|
"""A finished, review-gated job with one queued finding and a stub model."""
|
||||||
|
store = ReviewStore(str(tmp_path))
|
||||||
|
store.write_queue([_queue_item()])
|
||||||
|
monkeypatch.setattr("backend.main.get_job", lambda job_id: {
|
||||||
|
"job_id": job_id, "status": "needs_review", "out_dir": str(tmp_path)})
|
||||||
|
monkeypatch.setattr("backend.review.chat.call_json", lambda **kw: _reply())
|
||||||
|
return str(tmp_path)
|
||||||
|
|
||||||
|
|
||||||
|
def test_ask_about_a_finding_returns_and_logs_a_turn(job):
|
||||||
|
client = TestClient(app)
|
||||||
|
response = client.post("/jobs/job1/review-chat", json={
|
||||||
|
"question": "Why does it think AC-1 is at grade?",
|
||||||
|
"review_item_id": "finding:AGENT-0007"})
|
||||||
|
assert response.status_code == 200
|
||||||
|
turn = response.json()["turn"]
|
||||||
|
assert turn["issue"]["issue_id"] == "AGENT-0007"
|
||||||
|
assert turn["findings"] == ["The grade value came from M2.1."]
|
||||||
|
with open(os.path.join(job, "review", "chat_log.jsonl"), encoding="utf-8") as f:
|
||||||
|
assert len([line for line in f if line.strip()]) == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_ask_about_the_run_needs_no_item(job):
|
||||||
|
client = TestClient(app)
|
||||||
|
response = client.post("/jobs/job1/review-chat",
|
||||||
|
json={"question": "Why didn't it pick up the Civil set?"})
|
||||||
|
assert response.status_code == 200
|
||||||
|
assert response.json()["turn"]["scope"] == "run"
|
||||||
|
|
||||||
|
|
||||||
|
def test_history_endpoint_filters_by_item(job):
|
||||||
|
client = TestClient(app)
|
||||||
|
client.post("/jobs/job1/review-chat", json={
|
||||||
|
"question": "Why grade?", "review_item_id": "finding:AGENT-0007"})
|
||||||
|
client.post("/jobs/job1/review-chat", json={"question": "Why no Civil?"})
|
||||||
|
assert len(client.get("/jobs/job1/review-chat").json()["turns"]) == 2
|
||||||
|
filtered = client.get("/jobs/job1/review-chat",
|
||||||
|
params={"review_item_id": "finding:AGENT-0007"}).json()
|
||||||
|
assert len(filtered["turns"]) == 1
|
||||||
|
assert filtered["turns"][0]["question"] == "Why grade?"
|
||||||
|
|
||||||
|
|
||||||
|
def test_transcript_endpoint_renders_markdown(job):
|
||||||
|
client = TestClient(app)
|
||||||
|
client.post("/jobs/job1/review-chat", json={
|
||||||
|
"question": "Why grade?", "review_item_id": "finding:AGENT-0007"})
|
||||||
|
response = client.get("/jobs/job1/review-chat/log")
|
||||||
|
assert response.status_code == 200
|
||||||
|
assert response.headers["content-type"].startswith("text/markdown")
|
||||||
|
assert "## AGENT-0007" in response.text
|
||||||
|
assert "Why grade?" in response.text
|
||||||
|
|
||||||
|
|
||||||
|
def test_blank_question_is_422(job):
|
||||||
|
client = TestClient(app)
|
||||||
|
assert client.post("/jobs/job1/review-chat", json={"question": " "}).status_code == 422
|
||||||
|
|
||||||
|
|
||||||
|
def test_unknown_item_is_422(job):
|
||||||
|
client = TestClient(app)
|
||||||
|
response = client.post("/jobs/job1/review-chat",
|
||||||
|
json={"question": "why?", "review_item_id": "finding:NOPE"})
|
||||||
|
assert response.status_code == 422
|
||||||
|
|
||||||
|
|
||||||
|
def test_model_failure_is_502_not_500(monkeypatch, job):
|
||||||
|
monkeypatch.setattr("backend.review.chat.call_json", lambda **kw: None)
|
||||||
|
client = TestClient(app)
|
||||||
|
response = client.post("/jobs/job1/review-chat", json={"question": "why?"})
|
||||||
|
assert response.status_code == 502
|
||||||
|
|
||||||
|
|
||||||
|
def test_chat_stays_available_after_the_job_is_done(monkeypatch, tmp_path):
|
||||||
|
"""The chat is read-only, so a finalized report can still be questioned."""
|
||||||
|
ReviewStore(str(tmp_path)).write_queue([_queue_item()])
|
||||||
|
monkeypatch.setattr("backend.main.get_job", lambda job_id: {
|
||||||
|
"job_id": job_id, "status": "done", "out_dir": str(tmp_path)})
|
||||||
|
monkeypatch.setattr("backend.review.chat.call_json", lambda **kw: _reply())
|
||||||
|
client = TestClient(app)
|
||||||
|
assert client.post("/jobs/job1/review-chat",
|
||||||
|
json={"question": "why?"}).status_code == 200
|
||||||
|
|
||||||
|
|
||||||
|
def test_chat_is_409_while_the_job_is_still_running(monkeypatch, tmp_path):
|
||||||
|
monkeypatch.setattr("backend.main.get_job", lambda job_id: {
|
||||||
|
"job_id": job_id, "status": "running", "out_dir": str(tmp_path)})
|
||||||
|
client = TestClient(app)
|
||||||
|
assert client.post("/jobs/job1/review-chat",
|
||||||
|
json={"question": "why?"}).status_code == 409
|
||||||
|
|
||||||
|
|
||||||
|
def test_chat_404s_for_unknown_job(monkeypatch, tmp_path):
|
||||||
|
monkeypatch.setattr("backend.main.get_job", lambda job_id: None)
|
||||||
|
client = TestClient(app)
|
||||||
|
assert client.post("/jobs/nope/review-chat",
|
||||||
|
json={"question": "why?"}).status_code == 404
|
||||||
|
|
||||||
|
|
||||||
|
def test_chat_never_mutates_review_decisions(job):
|
||||||
|
"""The whole point: asking questions cannot change the review state."""
|
||||||
|
client = TestClient(app)
|
||||||
|
before = ReviewStore(job, create=False).read_decisions()
|
||||||
|
client.post("/jobs/job1/review-chat", json={
|
||||||
|
"question": "This is wrong, reject it.",
|
||||||
|
"review_item_id": "finding:AGENT-0007"})
|
||||||
|
after = ReviewStore(job, create=False).read_decisions()
|
||||||
|
assert before == after == {}
|
||||||
|
with open(os.path.join(job, "review", "review_queue.json"), encoding="utf-8") as f:
|
||||||
|
assert json.load(f) == [_queue_item()]
|
||||||
@@ -0,0 +1,15 @@
|
|||||||
|
"""Shared test fixtures."""
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture(autouse=True)
|
||||||
|
def isolated_feedback_store(tmp_path, monkeypatch):
|
||||||
|
"""Keep the cross-job feedback store out of the real outputs directory.
|
||||||
|
|
||||||
|
Saving a review decision or asking a chat question appends to
|
||||||
|
config.REVIEW_FEEDBACK_DIR, which is process-wide rather than job-local.
|
||||||
|
Without this, running the suite would accumulate junk in backend/outputs.
|
||||||
|
"""
|
||||||
|
monkeypatch.setattr("backend.config.REVIEW_FEEDBACK_DIR",
|
||||||
|
str(tmp_path / "_feedback"))
|
||||||
@@ -0,0 +1,172 @@
|
|||||||
|
"""Review chat: answer normalization, logging, and the feedback roll-up."""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
from backend.review import chat
|
||||||
|
|
||||||
|
|
||||||
|
def _queue() -> list:
|
||||||
|
return [{
|
||||||
|
"review_item_id": "finding:AGENT-0007", "kind": "finding", "blocking": True,
|
||||||
|
"reasons": ["severity_high"],
|
||||||
|
"payload": {"issue_id": "AGENT-0007", "source_stage": "conflict",
|
||||||
|
"category": "elevation_disagreement", "severity": "high",
|
||||||
|
"confidence": "medium", "location": "Roof / AC-1",
|
||||||
|
"sheets": ["M2.1"], "description": "AC-1 at grade vs roof.",
|
||||||
|
"evidence": [], "scope_id": "conflict:roof-ac1"},
|
||||||
|
}]
|
||||||
|
|
||||||
|
|
||||||
|
def _model_reply(**overrides) -> dict:
|
||||||
|
reply = {
|
||||||
|
"answer": "The extractor read 'AC-1 MOUNTED ON GRADE' off M2.1.",
|
||||||
|
"findings": ["The grade reading came from M2.1's text layer."],
|
||||||
|
"evidence_cited": [{"artifact": "source_sheets[M2.1]", "sheet": "M2.1",
|
||||||
|
"quote": "AC-1 MOUNTED ON GRADE",
|
||||||
|
"why_it_matters": "It is the sole basis for 'grade'."}],
|
||||||
|
"answerable": "yes",
|
||||||
|
"missing_information": None,
|
||||||
|
"assessment_of_finding": "looks_supported",
|
||||||
|
"suggested_category_correction": None,
|
||||||
|
"confidence": "high",
|
||||||
|
}
|
||||||
|
reply.update(overrides)
|
||||||
|
return reply
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
def fake_llm(monkeypatch):
|
||||||
|
"""Stub the model; the chat must never need a network to be tested."""
|
||||||
|
calls = []
|
||||||
|
|
||||||
|
def _call(**kwargs):
|
||||||
|
calls.append(kwargs)
|
||||||
|
return calls_reply[0]
|
||||||
|
|
||||||
|
calls_reply = [_model_reply()]
|
||||||
|
monkeypatch.setattr("backend.review.chat.call_json", lambda **kw: _call(**kw))
|
||||||
|
return calls, calls_reply
|
||||||
|
|
||||||
|
|
||||||
|
def test_ask_logs_issue_question_and_findings(tmp_path, fake_llm):
|
||||||
|
turn = chat.ask("job1", str(tmp_path), "Why is AC-1 at grade?",
|
||||||
|
review_item_id="finding:AGENT-0007", queue=_queue())
|
||||||
|
assert turn["question"] == "Why is AC-1 at grade?"
|
||||||
|
assert turn["findings"] == ["The grade reading came from M2.1's text layer."]
|
||||||
|
# The log's "issue in question" is a snapshot, not a bare id.
|
||||||
|
assert turn["issue"]["issue_id"] == "AGENT-0007"
|
||||||
|
assert turn["issue"]["severity"] == "high"
|
||||||
|
path = os.path.join(str(tmp_path), "review", "chat_log.jsonl")
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
logged = [json.loads(line) for line in f if line.strip()]
|
||||||
|
assert len(logged) == 1
|
||||||
|
assert logged[0]["turn_id"] == turn["turn_id"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_ask_appends_to_the_cross_job_feedback_store(tmp_path, fake_llm):
|
||||||
|
chat.ask("job1", str(tmp_path), "Is this really a floor drain?",
|
||||||
|
review_item_id="finding:AGENT-0007", queue=_queue())
|
||||||
|
path = os.path.join(config.REVIEW_FEEDBACK_DIR, "chat_turns.jsonl")
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
records = [json.loads(line) for line in f if line.strip()]
|
||||||
|
assert records[0]["kind"] == "review_chat_turn"
|
||||||
|
assert records[0]["issue_id"] == "AGENT-0007"
|
||||||
|
assert records[0]["job_id"] == "job1"
|
||||||
|
|
||||||
|
|
||||||
|
def test_correction_signal_is_captured_as_structured_data(tmp_path, fake_llm):
|
||||||
|
"""A misidentification correction survives as a field, not free text."""
|
||||||
|
_, reply = fake_llm
|
||||||
|
reply[0] = _model_reply(suggested_category_correction="power floor box",
|
||||||
|
assessment_of_finding="looks_unsupported")
|
||||||
|
turn = chat.ask("job1", str(tmp_path), "That is not a floor drain.",
|
||||||
|
review_item_id="finding:AGENT-0007", queue=_queue())
|
||||||
|
assert turn["suggested_category_correction"] == "power floor box"
|
||||||
|
path = os.path.join(config.REVIEW_FEEDBACK_DIR, "chat_turns.jsonl")
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
record = json.loads(f.readline())
|
||||||
|
assert record["suggested_category_correction"] == "power floor box"
|
||||||
|
assert record["assessment_of_finding"] == "looks_unsupported"
|
||||||
|
|
||||||
|
|
||||||
|
def test_run_scope_question_needs_no_item(tmp_path, fake_llm):
|
||||||
|
turn = chat.ask("job1", str(tmp_path), "Why didn't it pick up the Civil set?")
|
||||||
|
assert turn["review_item_id"] is None
|
||||||
|
assert turn["scope"] == "run"
|
||||||
|
assert turn["issue"] is None
|
||||||
|
|
||||||
|
|
||||||
|
def test_history_is_replayed_for_the_same_thread(tmp_path, fake_llm):
|
||||||
|
calls, _ = fake_llm
|
||||||
|
chat.ask("job1", str(tmp_path), "First question?",
|
||||||
|
review_item_id="finding:AGENT-0007", queue=_queue())
|
||||||
|
chat.ask("job1", str(tmp_path), "Follow-up?",
|
||||||
|
review_item_id="finding:AGENT-0007", queue=_queue())
|
||||||
|
assert "First question?" in calls[1]["user_text"]
|
||||||
|
# A run-scope turn must not inherit an item thread's history.
|
||||||
|
chat.ask("job1", str(tmp_path), "Unrelated run question?")
|
||||||
|
assert "First question?" not in calls[2]["user_text"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_blank_and_oversized_questions_are_rejected(tmp_path, fake_llm):
|
||||||
|
with pytest.raises(chat.ChatError):
|
||||||
|
chat.ask("job1", str(tmp_path), " ")
|
||||||
|
with pytest.raises(chat.ChatError):
|
||||||
|
chat.ask("job1", str(tmp_path),
|
||||||
|
"x" * (config.REVIEW_CHAT_MAX_QUESTION_CHARS + 1))
|
||||||
|
|
||||||
|
|
||||||
|
def test_unknown_review_item_is_rejected(tmp_path, fake_llm):
|
||||||
|
with pytest.raises(chat.ChatError):
|
||||||
|
chat.ask("job1", str(tmp_path), "why?", review_item_id="finding:NOPE",
|
||||||
|
queue=_queue())
|
||||||
|
|
||||||
|
|
||||||
|
def test_unusable_model_reply_raises_and_logs_nothing(tmp_path, monkeypatch):
|
||||||
|
monkeypatch.setattr("backend.review.chat.call_json", lambda **kw: None)
|
||||||
|
with pytest.raises(RuntimeError):
|
||||||
|
chat.ask("job1", str(tmp_path), "why?")
|
||||||
|
assert not os.path.exists(os.path.join(str(tmp_path), "review", "chat_log.jsonl"))
|
||||||
|
|
||||||
|
|
||||||
|
def test_bad_enum_values_fall_back_instead_of_failing(tmp_path, fake_llm):
|
||||||
|
_, reply = fake_llm
|
||||||
|
reply[0] = _model_reply(answerable="probably", confidence="",
|
||||||
|
assessment_of_finding="made_up")
|
||||||
|
turn = chat.ask("job1", str(tmp_path), "why?")
|
||||||
|
assert turn["answerable"] == "partial"
|
||||||
|
assert turn["confidence"] == "low"
|
||||||
|
assert turn["assessment_of_finding"] == "cannot_tell"
|
||||||
|
|
||||||
|
|
||||||
|
def test_disabled_chat_refuses(tmp_path, monkeypatch, fake_llm):
|
||||||
|
monkeypatch.setattr("backend.config.ENABLE_REVIEW_CHAT", False)
|
||||||
|
with pytest.raises(chat.ChatError):
|
||||||
|
chat.ask("job1", str(tmp_path), "why?")
|
||||||
|
|
||||||
|
|
||||||
|
def test_read_log_skips_corrupt_lines(tmp_path, fake_llm):
|
||||||
|
chat.ask("job1", str(tmp_path), "why?")
|
||||||
|
path = os.path.join(str(tmp_path), "review", "chat_log.jsonl")
|
||||||
|
with open(path, "a", encoding="utf-8") as f:
|
||||||
|
f.write("{not json\n")
|
||||||
|
assert len(chat.read_log(str(tmp_path))) == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_markdown_transcript_groups_by_issue(tmp_path, fake_llm):
|
||||||
|
chat.ask("job1", str(tmp_path), "Why is AC-1 at grade?",
|
||||||
|
review_item_id="finding:AGENT-0007", queue=_queue())
|
||||||
|
chat.ask("job1", str(tmp_path), "Why no Civil?")
|
||||||
|
markdown = chat.render_log_markdown(chat.read_log(str(tmp_path)))
|
||||||
|
assert "## AGENT-0007" in markdown
|
||||||
|
assert "## Run-scope questions" in markdown
|
||||||
|
assert "Why is AC-1 at grade?" in markdown
|
||||||
|
assert "**Findings**" in markdown
|
||||||
|
|
||||||
|
|
||||||
|
def test_markdown_transcript_handles_empty_log():
|
||||||
|
assert "No questions" in chat.render_log_markdown([])
|
||||||
@@ -0,0 +1,137 @@
|
|||||||
|
"""Context bundles for the review chat: what the model is allowed to see."""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
|
||||||
|
from backend.review.chat_context import build_context
|
||||||
|
|
||||||
|
|
||||||
|
def _write(out_dir: str, name: str, value) -> None:
|
||||||
|
path = os.path.join(out_dir, name)
|
||||||
|
os.makedirs(os.path.dirname(path), exist_ok=True)
|
||||||
|
with open(path, "w", encoding="utf-8") as f:
|
||||||
|
json.dump(value, f)
|
||||||
|
|
||||||
|
|
||||||
|
def _finding() -> dict:
|
||||||
|
return {
|
||||||
|
"issue_id": "AGENT-0007",
|
||||||
|
"source_stage": "conflict",
|
||||||
|
"category": "elevation_disagreement",
|
||||||
|
"severity": "high",
|
||||||
|
"confidence": "medium",
|
||||||
|
"location": "Roof / AC-1",
|
||||||
|
"disciplines": ["Mechanical"],
|
||||||
|
"sheets": ["M2.1"],
|
||||||
|
"description": "AC-1 shown at grade on M2.1 but on the roof elsewhere.",
|
||||||
|
"evidence": [{"discipline": "Mechanical", "sheet": "M2.1",
|
||||||
|
"source_text": "AC-1 MOUNTED ON GRADE", "asserted_value": "grade"}],
|
||||||
|
"scope_id": "conflict:roof-ac1",
|
||||||
|
"verification": {"status": "unverified", "verdicts": []},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _queue() -> list:
|
||||||
|
return [{"review_item_id": "finding:AGENT-0007", "kind": "finding",
|
||||||
|
"blocking": True, "reasons": ["severity_high"], "payload": _finding()}]
|
||||||
|
|
||||||
|
|
||||||
|
def _job_dir(tmp_path) -> str:
|
||||||
|
out_dir = str(tmp_path)
|
||||||
|
_write(out_dir, "conflicts.json", {
|
||||||
|
"source": "set.pdf",
|
||||||
|
"summary": {"pipeline_mode": "agent", "agent_status": "needs_review",
|
||||||
|
"by_stage": {"conflicts": 3}, "conflicts_found": 3},
|
||||||
|
"sheet_index": {"sheet_index": [
|
||||||
|
{"sheet_number": "M2.1", "discipline": "Mechanical"},
|
||||||
|
{"sheet_number": "A1.1", "discipline": "Architectural"},
|
||||||
|
]},
|
||||||
|
"sheet_reconciliation": {"declared_total": 4, "found_total": 2,
|
||||||
|
"declared_not_in_set": ["C-001", "C-101"],
|
||||||
|
"in_set_not_declared": []},
|
||||||
|
})
|
||||||
|
_write(out_dir, "agent/memory.json", {
|
||||||
|
"sheets": [{
|
||||||
|
"sheet_number": "M2.1", "discipline": "Mechanical", "page_number": 7,
|
||||||
|
"assertions": [{"attribute": "mounting", "value": "grade",
|
||||||
|
"source_text": "AC-1 MOUNTED ON GRADE",
|
||||||
|
"base64": "SHOULD-NOT-APPEAR"}],
|
||||||
|
}],
|
||||||
|
"clusters": [{"key": "roof-ac1", "location": "Roof / AC-1",
|
||||||
|
"disciplines": ["Mechanical", "Architectural"],
|
||||||
|
"assertions": [{"sheet_number": "M2.1", "attribute": "mounting",
|
||||||
|
"value": "grade", "base64": "SHOULD-NOT-APPEAR"}]}],
|
||||||
|
"decisions": [{"finding_refs": ["AGENT-0007"], "action": "kept",
|
||||||
|
"reason": "supported", "kept_issue_id": "AGENT-0007"}],
|
||||||
|
"suppressed": [],
|
||||||
|
})
|
||||||
|
return out_dir
|
||||||
|
|
||||||
|
|
||||||
|
def test_item_scope_carries_the_reasoning_chain(tmp_path):
|
||||||
|
context = build_context(_job_dir(tmp_path), "finding:AGENT-0007", _queue())
|
||||||
|
assert context["scope"] == "item"
|
||||||
|
assert context["finding"]["issue_id"] == "AGENT-0007"
|
||||||
|
# The chain a "why does it think X" answer has to walk.
|
||||||
|
assert context["originating_cluster"]["key"] == "roof-ac1"
|
||||||
|
assert context["source_sheets"][0]["sheet_number"] == "M2.1"
|
||||||
|
assert context["brain_decisions"][0]["action"] == "kept"
|
||||||
|
assert context["finding"]["verification"]["status"] == "unverified"
|
||||||
|
|
||||||
|
|
||||||
|
def test_context_never_leaks_base64(tmp_path):
|
||||||
|
"""Page images blow up the prompt and are useless as quotable evidence."""
|
||||||
|
context = build_context(_job_dir(tmp_path), "finding:AGENT-0007", _queue())
|
||||||
|
assert "SHOULD-NOT-APPEAR" not in json.dumps(context)
|
||||||
|
|
||||||
|
|
||||||
|
def test_run_scope_carries_coverage_material(tmp_path):
|
||||||
|
"""The 'why didn't it pick up the Civil set' inputs are all present."""
|
||||||
|
context = build_context(_job_dir(tmp_path), None, _queue(),
|
||||||
|
question="why didn't it pick up the Civil set?")
|
||||||
|
assert context["scope"] == "run"
|
||||||
|
assert set(context["sheets_by_discipline"]) == {"Mechanical", "Architectural"}
|
||||||
|
assert context["sheet_reconciliation"]["declared_not_in_set"] == ["C-001", "C-101"]
|
||||||
|
assert "finding" not in context
|
||||||
|
assert context["run"]["code_review_enabled"] in (True, False)
|
||||||
|
|
||||||
|
|
||||||
|
def test_unknown_item_falls_back_to_run_scope(tmp_path):
|
||||||
|
context = build_context(_job_dir(tmp_path), "finding:NOPE", _queue())
|
||||||
|
assert context["scope"] == "run"
|
||||||
|
|
||||||
|
|
||||||
|
def test_missing_artifacts_degrade_to_empty(tmp_path):
|
||||||
|
context = build_context(str(tmp_path), None, [])
|
||||||
|
assert context["scope"] == "run"
|
||||||
|
assert context["artifacts_available"] == {
|
||||||
|
"conflicts.json": False, "agent/memory.json": False, "job.log": False}
|
||||||
|
|
||||||
|
|
||||||
|
def test_log_excerpt_matches_question_terms(tmp_path):
|
||||||
|
out_dir = _job_dir(tmp_path)
|
||||||
|
with open(os.path.join(out_dir, "job.log"), "w", encoding="utf-8") as f:
|
||||||
|
f.write("[Extract] page 3 Civil sheet unreadable, skipped\n")
|
||||||
|
f.write("[Brain] merged 2 findings\n")
|
||||||
|
context = build_context(out_dir, None, [], question="why no Civil sheets?")
|
||||||
|
assert any("Civil" in line for line in context["log_excerpt"])
|
||||||
|
assert context["artifacts_available"]["job.log"] is True
|
||||||
|
|
||||||
|
|
||||||
|
def test_reviewer_decision_so_far_is_included(tmp_path):
|
||||||
|
decisions = {"finding:AGENT-0007": {"decision": "reject",
|
||||||
|
"reason_code": "extraction_misread",
|
||||||
|
"comment": "that is a power floor box"}}
|
||||||
|
context = build_context(_job_dir(tmp_path), "finding:AGENT-0007", _queue(), decisions)
|
||||||
|
assert context["reviewer_decision_so_far"]["reason_code"] == "extraction_misread"
|
||||||
|
|
||||||
|
|
||||||
|
def test_clean_cluster_item_uses_cluster_scope(tmp_path):
|
||||||
|
queue = [{"review_item_id": "clean_cluster:roof-ac1", "kind": "clean_cluster",
|
||||||
|
"blocking": False, "reasons": ["audit_sample"],
|
||||||
|
"payload": {"key": "roof-ac1", "location": "Roof / AC-1",
|
||||||
|
"assertions": [{"sheet_number": "M2.1", "value": "grade"}]}}]
|
||||||
|
context = build_context(_job_dir(tmp_path), "clean_cluster:roof-ac1", queue)
|
||||||
|
assert context["scope"] == "item"
|
||||||
|
assert context["cluster"]["key"] == "roof-ac1"
|
||||||
|
assert "finding" not in context
|
||||||
@@ -85,3 +85,34 @@ def test_write_label_appends_json_lines(tmp_path):
|
|||||||
with open(path, encoding="utf-8") as f:
|
with open(path, encoding="utf-8") as f:
|
||||||
lines = [json.loads(line) for line in f if line.strip()]
|
lines = [json.loads(line) for line in f if line.strip()]
|
||||||
assert lines == [label1, label2]
|
assert lines == [label1, label2]
|
||||||
|
|
||||||
|
|
||||||
|
def test_decision_label_carries_reviewer_corrections():
|
||||||
|
"""category/severity corrections reach the label instead of being dropped."""
|
||||||
|
decision = {"decision": "reject", "reason_code": "extraction_misread",
|
||||||
|
"category_correction": "power floor box",
|
||||||
|
"severity_correction": "low"}
|
||||||
|
label = decision_to_label(_queue_item(), decision, {"job_id": "abc123", "pipeline_mode": "agent",
|
||||||
|
"report": {"summary": {}}})
|
||||||
|
assert label["category_correction"] == "power floor box"
|
||||||
|
assert label["severity_correction"] == "low"
|
||||||
|
assert label["kind"] == "review_decision"
|
||||||
|
|
||||||
|
|
||||||
|
def test_write_label_also_lands_in_the_cross_job_store(tmp_path):
|
||||||
|
from backend import config
|
||||||
|
from backend.review.feedback import read_shared_feedback
|
||||||
|
label = decision_to_label(_queue_item(), {"decision": "reject",
|
||||||
|
"reason_code": "extraction_misread"},
|
||||||
|
{"job_id": "abc123", "pipeline_mode": "agent",
|
||||||
|
"report": {"summary": {}}})
|
||||||
|
write_label(str(tmp_path), label)
|
||||||
|
assert os.path.isfile(os.path.join(config.REVIEW_FEEDBACK_DIR, "decisions.jsonl"))
|
||||||
|
records = read_shared_feedback("review_decision")
|
||||||
|
assert len(records) == 1
|
||||||
|
assert records[0]["reason_code"] == "extraction_misread"
|
||||||
|
|
||||||
|
|
||||||
|
def test_shared_feedback_read_is_empty_when_nothing_written():
|
||||||
|
from backend.review.feedback import read_shared_feedback
|
||||||
|
assert read_shared_feedback("review_decision") == []
|
||||||
@@ -85,6 +85,29 @@ def test_finalize_confirm_keeps_confirmed(monkeypatch, tmp_path):
|
|||||||
assert report["summary"]["agent_status"] == "complete"
|
assert report["summary"]["agent_status"] == "complete"
|
||||||
|
|
||||||
|
|
||||||
|
def test_finalize_preserves_verifier_suppressed(monkeypatch, tmp_path):
|
||||||
|
"""Wave-5b (verifier) suppressions must survive review finalization and
|
||||||
|
merge with review-rejected suppressions."""
|
||||||
|
monkeypatch.setattr("backend.review.finalizer._draft_rfis", lambda kept: [])
|
||||||
|
_write_job(
|
||||||
|
str(tmp_path),
|
||||||
|
prioritized=[{"issue_id": "AGENT-0001", "severity": "high"}],
|
||||||
|
queue=[_blocking_item("AGENT-0001")],
|
||||||
|
decisions=[{"review_item_id": "finding:AGENT-0001",
|
||||||
|
"decision": "reject", "reason_code": "not_a_contradiction"}],
|
||||||
|
)
|
||||||
|
path = os.path.join(str(tmp_path), "conflicts.json")
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
report = json.load(f)
|
||||||
|
report["suppressed_issues"] = [
|
||||||
|
{"issue_id": "C1", "verification": {"status": "refuted"}}]
|
||||||
|
with open(path, "w", encoding="utf-8") as f:
|
||||||
|
json.dump(report, f)
|
||||||
|
final = finalize_review("job1", str(tmp_path))
|
||||||
|
ids = [f["issue_id"] for f in final["suppressed_issues"]]
|
||||||
|
assert ids == ["C1", "AGENT-0001"]
|
||||||
|
|
||||||
|
|
||||||
def test_finalize_no_decision_keeps_unreviewed(monkeypatch, tmp_path):
|
def test_finalize_no_decision_keeps_unreviewed(monkeypatch, tmp_path):
|
||||||
"""Non-blocking (audit) items don't need a decision; issue stays unreviewed."""
|
"""Non-blocking (audit) items don't need a decision; issue stays unreviewed."""
|
||||||
monkeypatch.setattr("backend.review.finalizer._draft_rfis", lambda kept: [])
|
monkeypatch.setattr("backend.review.finalizer._draft_rfis", lambda kept: [])
|
||||||
|
|||||||
@@ -0,0 +1,84 @@
|
|||||||
|
"""Unit tests for the extraction_coverage block in the report summary.
|
||||||
|
|
||||||
|
Both the classic pipeline and the agent runner build their summary via
|
||||||
|
backend.pipeline.report.build_report, so unit tests on that function cover
|
||||||
|
every summary-producing path.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from backend.pipeline.report import build_report
|
||||||
|
|
||||||
|
|
||||||
|
def _sheet(page, coverage=None, assertions=None):
|
||||||
|
sheet = {
|
||||||
|
"page_number": page,
|
||||||
|
"sheet_number": f"S{page:03d}",
|
||||||
|
"discipline": "S",
|
||||||
|
"assertions": assertions if assertions is not None else [
|
||||||
|
{"text": "NOTE ALPHA", "object_type": "note"},
|
||||||
|
{"text": "NOTE BETA", "object_type": "note"},
|
||||||
|
],
|
||||||
|
}
|
||||||
|
if coverage is not None:
|
||||||
|
sheet["coverage"] = coverage
|
||||||
|
return sheet
|
||||||
|
|
||||||
|
|
||||||
|
def _cov(total, covered):
|
||||||
|
return {
|
||||||
|
"total_lines": total,
|
||||||
|
"covered_lines": covered,
|
||||||
|
"ratio": covered / total if total else 0.0,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def test_extraction_coverage_omitted_without_coverage_data():
|
||||||
|
report = build_report(conflicts=[], sheets=[_sheet(1), _sheet(2)], clusters=[])
|
||||||
|
assert "extraction_coverage" not in report["summary"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_extraction_coverage_healthy():
|
||||||
|
sheets = [_sheet(1, _cov(100, 95)), _sheet(2, _cov(80, 76))]
|
||||||
|
report = build_report(conflicts=[], sheets=sheets, clusters=[])
|
||||||
|
cov = report["summary"]["extraction_coverage"]
|
||||||
|
assert cov["pages_measured"] == 2
|
||||||
|
assert cov["pages_below_floor"] == []
|
||||||
|
assert cov["fallback_pages"] == []
|
||||||
|
assert cov["mean_ratio"] == round((0.95 + 0.95) / 2, 3)
|
||||||
|
|
||||||
|
|
||||||
|
def test_extraction_coverage_flags_below_floor():
|
||||||
|
sheets = [
|
||||||
|
_sheet(1, _cov(100, 95)),
|
||||||
|
_sheet(2, _cov(100, 40)), # ratio 0.4 < 0.6 floor
|
||||||
|
_sheet(3, _cov(100, 59)), # ratio 0.59 < 0.6 floor
|
||||||
|
]
|
||||||
|
report = build_report(conflicts=[], sheets=sheets, clusters=[])
|
||||||
|
cov = report["summary"]["extraction_coverage"]
|
||||||
|
assert cov["pages_measured"] == 3
|
||||||
|
assert cov["pages_below_floor"] == [2, 3]
|
||||||
|
assert cov["mean_ratio"] == round((0.95 + 0.4 + 0.59) / 3, 3)
|
||||||
|
|
||||||
|
|
||||||
|
def test_extraction_coverage_flags_fallback_pages():
|
||||||
|
fallback_assertions = [
|
||||||
|
{"text": "NOTE ALPHA", "object_type": "note"},
|
||||||
|
{"text": "NOTE BETA", "object_type": "note",
|
||||||
|
"grounding": "text_layer_fallback"},
|
||||||
|
]
|
||||||
|
sheets = [
|
||||||
|
_sheet(1, _cov(100, 90), assertions=fallback_assertions),
|
||||||
|
_sheet(2, _cov(100, 90)),
|
||||||
|
]
|
||||||
|
report = build_report(conflicts=[], sheets=sheets, clusters=[])
|
||||||
|
cov = report["summary"]["extraction_coverage"]
|
||||||
|
assert cov["fallback_pages"] == [1]
|
||||||
|
|
||||||
|
|
||||||
|
def test_extraction_coverage_mixed_sheets_only_counts_measured():
|
||||||
|
# Sheet 2 has no coverage dict (e.g. scanned page / older path).
|
||||||
|
sheets = [_sheet(1, _cov(100, 50)), _sheet(2)]
|
||||||
|
report = build_report(conflicts=[], sheets=sheets, clusters=[])
|
||||||
|
cov = report["summary"]["extraction_coverage"]
|
||||||
|
assert cov["pages_measured"] == 1
|
||||||
|
assert cov["pages_below_floor"] == [1]
|
||||||
|
assert cov["mean_ratio"] == 0.5
|
||||||
@@ -0,0 +1,52 @@
|
|||||||
|
"""Tests for the classic-path drawing_integrity_review stage + gating."""
|
||||||
|
|
||||||
|
import backend.pipeline.drawing_integrity as di
|
||||||
|
from backend import config
|
||||||
|
from backend.pipeline.drawing_integrity import drawing_integrity_review
|
||||||
|
|
||||||
|
|
||||||
|
def _sheet(page, sheet_number, n):
|
||||||
|
return {
|
||||||
|
"sheet_number": sheet_number,
|
||||||
|
"page_number": page,
|
||||||
|
"discipline": "Architectural",
|
||||||
|
"assertions": [{"source_text": f"n{i}"} for i in range(n)],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _pages(pages):
|
||||||
|
return [{"page_number": p, "base64": f"IMG{p}", "text_layer": f"txt{p}"}
|
||||||
|
for p in pages]
|
||||||
|
|
||||||
|
|
||||||
|
def test_disabled_returns_empty(monkeypatch):
|
||||||
|
monkeypatch.setattr(config, "ENABLE_DRAWING_INTEGRITY", False)
|
||||||
|
called = []
|
||||||
|
monkeypatch.setattr(di, "call_stage", lambda *a, **k: called.append(1) or {})
|
||||||
|
out = drawing_integrity_review([_sheet(1, "A101", 5)], _pages([1]))
|
||||||
|
assert out == []
|
||||||
|
assert called == [] # no LLM calls when disabled
|
||||||
|
|
||||||
|
|
||||||
|
def test_reviews_only_dense_sheets(monkeypatch):
|
||||||
|
monkeypatch.setattr(config, "ENABLE_DRAWING_INTEGRITY", True)
|
||||||
|
monkeypatch.setattr(config, "INTEGRITY_MIN_ASSERTIONS", 3)
|
||||||
|
seen_sheets = []
|
||||||
|
|
||||||
|
def fake_call_stage(system, user, subs=None, images_b64=None, **k):
|
||||||
|
# Record which sheet_meta reached the model.
|
||||||
|
seen_sheets.append(subs["sheet_meta"])
|
||||||
|
return {"issues": [
|
||||||
|
{"severity": "medium", "confidence": "high",
|
||||||
|
"category": "on_sheet_contradiction", "sheets": [],
|
||||||
|
"description": "plan disagrees with same-sheet schedule"},
|
||||||
|
]}
|
||||||
|
|
||||||
|
monkeypatch.setattr(di, "call_stage", fake_call_stage)
|
||||||
|
sheets = [_sheet(1, "A101", 5), _sheet(2, "A102", 1), _sheet(3, "A103", 4)]
|
||||||
|
out = drawing_integrity_review(sheets, _pages([1, 2, 3]))
|
||||||
|
# Two dense sheets reviewed, one skipped; each produced one anchored finding.
|
||||||
|
assert len(out) == 2
|
||||||
|
assert all(f["source_stage"] == "drawing_integrity" for f in out)
|
||||||
|
assert {tuple(f["sheets"]) for f in out} == {("A101",), ("A103",)}
|
||||||
|
assert len(seen_sheets) == 2
|
||||||
@@ -0,0 +1,98 @@
|
|||||||
|
"""Coverage-driven extraction retry ladder — classic path (pipeline/extractor.py)."""
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
from backend.pipeline import extractor
|
||||||
|
|
||||||
|
PAGE_TEXT = ("1. \nALL SAWN LUMBER IN CONTACT WITH SOIL TO BE SOUTHERN PINE, "
|
||||||
|
"PRESSURE TREATED.\n2. \nROOF SHEATHING: 5/8\" PLYWOOD, C-D GRADE, "
|
||||||
|
"STRUCTURAL I.")
|
||||||
|
|
||||||
|
|
||||||
|
def _page(n=8, text=PAGE_TEXT):
|
||||||
|
return {"page_number": n, "base64": "AAAA", "text_layer": text}
|
||||||
|
|
||||||
|
|
||||||
|
def test_classic_fallback_when_vision_returns_nothing(monkeypatch):
|
||||||
|
"""Vision pass returns an unusable bare list; text retry disabled ->
|
||||||
|
deterministic fallback stubs make a dark sheet impossible."""
|
||||||
|
monkeypatch.setattr(extractor, "call_json",
|
||||||
|
lambda **kw: [{"name": "general notes", "value": "notes"}])
|
||||||
|
monkeypatch.setattr(config, "EXTRACT_TEXT_RETRY_ENABLED", False)
|
||||||
|
sheet = extractor._extract_one(_page())
|
||||||
|
assert sheet["assertions"], "dark sheet must be impossible with fallback enabled"
|
||||||
|
assert all(a.get("grounding") == "text_layer_fallback"
|
||||||
|
for a in sheet["assertions"])
|
||||||
|
assert sheet["coverage"]["ratio"] >= 0.6
|
||||||
|
|
||||||
|
|
||||||
|
def test_classic_merge_preserves_graphical_objects(monkeypatch):
|
||||||
|
"""Rung-2 merge must never drop vision-only graphical objects."""
|
||||||
|
def fake(**kw):
|
||||||
|
if kw.get("images_b64"):
|
||||||
|
return {"sheet": {}, "objects": [
|
||||||
|
{"object_id": "g1", "object_type": "lighting_fixture",
|
||||||
|
"name": "pendant at grid C-4", "source_text": None,
|
||||||
|
"graphical_basis": "16in pendant symbol at grid C-4"}]}
|
||||||
|
return {"sheet": {}, "objects": [
|
||||||
|
{"object_id": "t1", "object_type": "general_note",
|
||||||
|
"source_text": "ALL SAWN LUMBER IN CONTACT WITH SOIL TO BE "
|
||||||
|
"SOUTHERN PINE, PRESSURE TREATED.",
|
||||||
|
"name": "lumber note"}]}
|
||||||
|
|
||||||
|
monkeypatch.setattr(extractor, "call_json", fake)
|
||||||
|
sheet = extractor._extract_one(_page())
|
||||||
|
assert any(a.get("graphical_basis") for a in sheet["assertions"])
|
||||||
|
assert any("SAWN LUMBER" in (a.get("source_text") or "")
|
||||||
|
for a in sheet["assertions"])
|
||||||
|
|
||||||
|
|
||||||
|
def test_classic_recovers_sheet_number(monkeypatch):
|
||||||
|
text = ("REFLECTED CEILING PLAN\n"
|
||||||
|
"GYP. BD. CEILING 8'-11 3/8\" A.F.F. TYP. FOR ALL STOREFRONT\n"
|
||||||
|
"LED TAPE LIGHT. SEE ELEC. SCONCE 8'-0\" A.F.F., SEE ELEC.\n"
|
||||||
|
"A102")
|
||||||
|
|
||||||
|
def fake(**kw):
|
||||||
|
if kw.get("images_b64"):
|
||||||
|
return {"sheet": {}, "objects": [
|
||||||
|
{"object_id": "o1", "name": "RCP ceiling note",
|
||||||
|
"source_text": "GYP. BD. CEILING 8'-11 3/8\" A.F.F. TYP. "
|
||||||
|
"FOR ALL STOREFRONT",
|
||||||
|
"attributes": {"height": "8'-11 3/8\""}}]}
|
||||||
|
return {"sheet": {}, "objects": []}
|
||||||
|
|
||||||
|
monkeypatch.setattr(extractor, "call_json", fake)
|
||||||
|
sheet = extractor._extract_one(_page(18, text))
|
||||||
|
assert sheet["sheet_number"] == "A102"
|
||||||
|
|
||||||
|
|
||||||
|
def test_classic_skips_retry_when_coverage_healthy(monkeypatch):
|
||||||
|
calls = []
|
||||||
|
|
||||||
|
def fake(**kw):
|
||||||
|
calls.append(kw)
|
||||||
|
return {"sheet": {"sheet_number": "S202"}, "objects": [
|
||||||
|
{"object_id": "o1", "name": "lumber note",
|
||||||
|
"source_text": "ALL SAWN LUMBER IN CONTACT WITH SOIL TO BE "
|
||||||
|
"SOUTHERN PINE, PRESSURE TREATED.",
|
||||||
|
"attributes": {"species": "southern pine"}},
|
||||||
|
{"object_id": "o2", "name": "sheathing note",
|
||||||
|
"source_text": "ROOF SHEATHING: 5/8\" PLYWOOD, C-D GRADE, "
|
||||||
|
"STRUCTURAL I.",
|
||||||
|
"attributes": {"sheathing": "5/8 plywood"}}]}
|
||||||
|
|
||||||
|
monkeypatch.setattr(extractor, "call_json", fake)
|
||||||
|
sheet = extractor._extract_one(_page())
|
||||||
|
assert len(calls) == 1, "healthy coverage must not trigger the text-only rung"
|
||||||
|
assert sheet["coverage"]["ratio"] == 1.0
|
||||||
|
assert sheet["sheet_number"] == "S202"
|
||||||
|
|
||||||
|
|
||||||
|
def test_classic_scanned_page_keeps_failed_sheet_shape(monkeypatch):
|
||||||
|
"""No text layer (scanned page): total parse failure keeps the existing
|
||||||
|
'extraction failed' empty-sheet return — ladder is text-layer-only."""
|
||||||
|
monkeypatch.setattr(extractor, "call_json", lambda **kw: None)
|
||||||
|
sheet = extractor._extract_one({"page_number": 4, "base64": "AAAA",
|
||||||
|
"text_layer": None})
|
||||||
|
assert sheet["assertions"] == []
|
||||||
|
assert "extraction failed" in (sheet.get("sheet_title") or "")
|
||||||
@@ -0,0 +1,122 @@
|
|||||||
|
"""Grounding-guard rescue tier, text-layer prompt block, and render hygiene."""
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
from backend.pipeline._stage import render
|
||||||
|
from backend.pipeline.extractor import (
|
||||||
|
_is_grounded,
|
||||||
|
_normalize_sheet,
|
||||||
|
_text_layer_block,
|
||||||
|
)
|
||||||
|
from backend.prompts import EXTRACTOR_USER_INSTRUCTION, VERIFY_USER_INSTRUCTION
|
||||||
|
|
||||||
|
PAGE_TEXT = "NOTES: (5) 2X6 STUD PACK AT BEARING. HSS16X4 BEAM. 7'-0\" AFF."
|
||||||
|
|
||||||
|
|
||||||
|
def _parsed(value, source_text):
|
||||||
|
return {
|
||||||
|
"sheet": {"sheet_number": "S401"},
|
||||||
|
"objects": [{
|
||||||
|
"object_id": "o1",
|
||||||
|
"object_type": "framing",
|
||||||
|
"name": "stud pack",
|
||||||
|
"attributes": {"count": value},
|
||||||
|
"source_text": source_text,
|
||||||
|
}],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def test_rescue_tier_keeps_and_stamps():
|
||||||
|
"""Digits absent from source_text but present in the page text layer:
|
||||||
|
kept, stamped grounding=text_layer (vision quoted imperfectly)."""
|
||||||
|
sheet = _normalize_sheet(_parsed("(2)", "(2) 2x6 STUD PACK"), 1,
|
||||||
|
page_text=PAGE_TEXT)
|
||||||
|
# "(2)" is not grounded by its own source_text alone? it is - use a value
|
||||||
|
# whose digits differ from the quote to exercise the rescue path.
|
||||||
|
sheet = _normalize_sheet(_parsed("5", "(2) 2x6 STUD PACK"), 1,
|
||||||
|
page_text=PAGE_TEXT)
|
||||||
|
assert len(sheet["assertions"]) == 1
|
||||||
|
assert sheet["assertions"][0]["grounding"] == "text_layer"
|
||||||
|
|
||||||
|
|
||||||
|
def test_no_rescue_without_page_text():
|
||||||
|
sheet = _normalize_sheet(_parsed("5", "(2) 2x6 STUD PACK"), 1)
|
||||||
|
assert sheet["assertions"] == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_still_dropped_when_digits_nowhere():
|
||||||
|
sheet = _normalize_sheet(_parsed("99", "(2) 2x6 STUD PACK"), 1,
|
||||||
|
page_text=PAGE_TEXT)
|
||||||
|
assert sheet["assertions"] == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_is_grounded_backward_compatible():
|
||||||
|
assert _is_grounded("(5)", "(5) 2x6 STUD PACK") is True
|
||||||
|
# Digit-run guard is a set check: "(3)" has no support anywhere.
|
||||||
|
assert _is_grounded("(3)", "(5) 2x6 STUD PACK") is False
|
||||||
|
assert _is_grounded("(3)", "(5) 2x6 STUD PACK",
|
||||||
|
page_text="(3) 2x6 STUD PACK") is True
|
||||||
|
|
||||||
|
|
||||||
|
def test_text_layer_block_empty_without_layer():
|
||||||
|
assert _text_layer_block({"page_number": 1}) == ""
|
||||||
|
assert _text_layer_block({"page_number": 1, "text_layer": None}) == ""
|
||||||
|
|
||||||
|
|
||||||
|
def test_text_layer_block_appends_and_caps(monkeypatch):
|
||||||
|
block = _text_layer_block({"page_number": 1, "text_layer": PAGE_TEXT})
|
||||||
|
assert "TEXT LAYER" in block and "STUD PACK" in block
|
||||||
|
monkeypatch.setattr(config, "TEXT_LAYER_MAX_CHARS", 50)
|
||||||
|
block = _text_layer_block({"page_number": 1, "text_layer": "x" * 500})
|
||||||
|
assert len(block.split(":\n", 1)[1]) == 50
|
||||||
|
|
||||||
|
|
||||||
|
def test_verify_instruction_fully_rendered():
|
||||||
|
"""render() silently leaves missing keys as literals - both placeholders
|
||||||
|
must be substituted at the (single) verify render site."""
|
||||||
|
out = render(VERIFY_USER_INSTRUCTION,
|
||||||
|
{"finding": "FINDING_JSON", "text_layer": "PAGE_TEXT"})
|
||||||
|
assert "{finding}" not in out and "{text_layer}" not in out
|
||||||
|
assert "FINDING_JSON" in out and "PAGE_TEXT" in out
|
||||||
|
|
||||||
|
|
||||||
|
def test_extractor_instruction_fully_substituted():
|
||||||
|
page = {"page_number": 1, "text_layer": PAGE_TEXT}
|
||||||
|
out = (EXTRACTOR_USER_INSTRUCTION.replace("{sheet_hint}", "")
|
||||||
|
+ _text_layer_block(page))
|
||||||
|
assert "{sheet_hint}" not in out
|
||||||
|
|
||||||
|
|
||||||
|
def test_vision_unverified_stamp_when_source_text_not_in_text_layer():
|
||||||
|
"""Digits ground the object against the page text, but its quoted
|
||||||
|
source_text is not actually present in the text layer: kept, stamped
|
||||||
|
vision_unverified (the wave-5b verifier consumes grounding stamps)."""
|
||||||
|
page_text = "WALL: 2X6 WD STUD @ 16\" O.C. WITH R-13 BATT INSULATION"
|
||||||
|
parsed = {"sheet": {}, "objects": [
|
||||||
|
{"object_id": "x1", "name": "stud pack",
|
||||||
|
"source_text": "(5) 2X6 STUD PACK AT JAMB", # NOT in page text
|
||||||
|
"attributes": {"count": "5"}}]}
|
||||||
|
sheet = _normalize_sheet(parsed, 1, page_text=page_text)
|
||||||
|
assert len(sheet["assertions"]) == 1
|
||||||
|
assert sheet["assertions"][0]["grounding"] == "vision_unverified"
|
||||||
|
|
||||||
|
|
||||||
|
def test_no_unverified_stamp_when_source_text_in_text_layer():
|
||||||
|
page_text = "WALL: 2X6 WD STUD @ 16\" O.C. WITH R-13 BATT INSULATION"
|
||||||
|
parsed = {"sheet": {}, "objects": [
|
||||||
|
{"object_id": "x1", "name": "stud note",
|
||||||
|
"source_text": "2X6 WD STUD @ 16\" O.C.",
|
||||||
|
"attributes": {"size": "2x6"}}]}
|
||||||
|
sheet = _normalize_sheet(parsed, 1, page_text=page_text)
|
||||||
|
assert len(sheet["assertions"]) == 1
|
||||||
|
assert "grounding" not in sheet["assertions"][0]
|
||||||
|
|
||||||
|
|
||||||
|
def test_preset_grounding_stamp_survives_normalization():
|
||||||
|
"""Fallback/merge rungs stamp grounding upstream; normalization must
|
||||||
|
preserve a pre-set stamp instead of recomputing it away."""
|
||||||
|
parsed = {"sheet": {}, "objects": [
|
||||||
|
{"object_id": "f1", "name": "lumber note",
|
||||||
|
"source_text": "ALL LUMBER SOUTHERN PINE",
|
||||||
|
"grounding": "text_layer_fallback"}]}
|
||||||
|
sheet = _normalize_sheet(parsed, 1, page_text="ALL LUMBER SOUTHERN PINE")
|
||||||
|
assert sheet["assertions"][0]["grounding"] == "text_layer_fallback"
|
||||||
@@ -1,26 +1,44 @@
|
|||||||
from backend import config
|
from backend import config
|
||||||
from backend.llm import _resolve_backend, set_model_override
|
from backend.llm import _resolve_backend, set_model_overrides, set_text_backend
|
||||||
|
|
||||||
|
|
||||||
def test_override_wins_for_vision_and_text():
|
def teardown_function():
|
||||||
set_model_override("openai/gpt-4o")
|
set_model_overrides(None, None)
|
||||||
try:
|
set_text_backend(False)
|
||||||
|
|
||||||
|
|
||||||
|
def test_vision_override_wins_for_vision_only():
|
||||||
|
set_model_overrides(vision="openai/gpt-4o", text=None)
|
||||||
assert _resolve_backend(has_images=True, model_override=None)["model"] == "openai/gpt-4o"
|
assert _resolve_backend(has_images=True, model_override=None)["model"] == "openai/gpt-4o"
|
||||||
assert _resolve_backend(has_images=False, model_override=None)["model"] == "openai/gpt-4o"
|
assert _resolve_backend(has_images=False, model_override=None)["model"] == config.TEXT_MODEL
|
||||||
finally:
|
|
||||||
set_model_override(None)
|
|
||||||
|
def test_text_override_wins_for_text_only():
|
||||||
|
set_model_overrides(vision=None, text="anthropic/claude-sonnet-4")
|
||||||
|
assert _resolve_backend(has_images=False, model_override=None)["model"] == "anthropic/claude-sonnet-4"
|
||||||
|
assert _resolve_backend(has_images=True, model_override=None)["model"] == config.MODEL
|
||||||
|
|
||||||
|
|
||||||
def test_override_beats_per_call_model_arg():
|
def test_override_beats_per_call_model_arg():
|
||||||
set_model_override("openai/gpt-4o")
|
set_model_overrides(vision="openai/gpt-4o", text="openai/gpt-4o-mini")
|
||||||
try:
|
|
||||||
# Agents pass their AGENT_*_MODEL per call; the user's job pick wins.
|
# Agents pass their AGENT_*_MODEL per call; the user's job pick wins.
|
||||||
assert _resolve_backend(has_images=False, model_override="other/model")["model"] == "openai/gpt-4o"
|
assert _resolve_backend(has_images=True, model_override="other/model")["model"] == "openai/gpt-4o"
|
||||||
finally:
|
assert _resolve_backend(has_images=False, model_override="other/model")["model"] == "openai/gpt-4o-mini"
|
||||||
set_model_override(None)
|
|
||||||
|
|
||||||
|
|
||||||
def test_no_override_keeps_defaults():
|
def test_no_override_keeps_defaults():
|
||||||
set_model_override(None)
|
set_model_overrides(None, None)
|
||||||
assert _resolve_backend(has_images=True, model_override=None)["model"] == config.MODEL
|
assert _resolve_backend(has_images=True, model_override=None)["model"] == config.MODEL
|
||||||
assert _resolve_backend(has_images=False, model_override=None)["model"] == config.TEXT_MODEL
|
assert _resolve_backend(has_images=False, model_override=None)["model"] == config.TEXT_MODEL
|
||||||
|
|
||||||
|
|
||||||
|
def test_ui_picks_never_name_the_local_model(monkeypatch):
|
||||||
|
"""Hybrid runs keep LOCAL_TEXT_MODEL; OpenRouter picks must not leak into
|
||||||
|
the local endpoint (a vLLM server won't serve OpenRouter model ids)."""
|
||||||
|
monkeypatch.setattr(config, "LOCAL_BASE_URL", "http://localhost:8000/v1")
|
||||||
|
monkeypatch.setattr(config, "LOCAL_TEXT_MODEL", "qwen/local-instruct")
|
||||||
|
set_text_backend(True)
|
||||||
|
set_model_overrides(vision="openai/gpt-4o", text="anthropic/claude-sonnet-4")
|
||||||
|
be = _resolve_backend(has_images=False, model_override=None)
|
||||||
|
assert be["local"] is True
|
||||||
|
assert be["model"] == "qwen/local-instruct"
|
||||||
@@ -0,0 +1,69 @@
|
|||||||
|
"""Regression tests for defects found in the Aug 2026 agent-mode code review.
|
||||||
|
|
||||||
|
R2 - text_coverage._SHEET_ID_RE could not match hyphenated sheet ids (C-001),
|
||||||
|
leaving civil/landscape pages sheet_number=None and producing false
|
||||||
|
"declared but not in set" reconciliation warnings.
|
||||||
|
R3 - config bool knobs mixed `== "true"` with the 1/true/yes set, so setting
|
||||||
|
EXTRACT_TEXT_RETRY_ENABLED=1 silently DISABLED the retry ladder.
|
||||||
|
|
||||||
|
See also tests/agents/test_verifier.py for the verifier's "corrected"
|
||||||
|
semantics, which are intentional and pinned there.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import importlib
|
||||||
|
import os
|
||||||
|
from unittest import mock
|
||||||
|
|
||||||
|
from backend.sheet_reconcile import declared_sheet_list, reconcile_sheets
|
||||||
|
from backend.text_coverage import recover_sheet_number
|
||||||
|
|
||||||
|
|
||||||
|
# --- R2: hyphenated sheet ids ----------------------------------------------
|
||||||
|
|
||||||
|
def test_recover_sheet_number_handles_hyphenated_civil_id():
|
||||||
|
page_text = ("GENERAL NOTES\n" * 40) + "PROJECT NO 2024-118\nSHEET\nC-001\n"
|
||||||
|
assert recover_sheet_number(page_text) == "C-001"
|
||||||
|
|
||||||
|
|
||||||
|
def test_recover_sheet_number_still_handles_plain_ids():
|
||||||
|
page_text = ("NOTES\n" * 40) + "SHEET\nS302\n"
|
||||||
|
assert recover_sheet_number(page_text) == "S302"
|
||||||
|
|
||||||
|
|
||||||
|
def test_recovered_hyphenated_id_reconciles_against_declared_index():
|
||||||
|
"""The whole point: a recovered C-001 must not read as a missing sheet."""
|
||||||
|
declared = declared_sheet_list({1: "SHEET LIST\nC-001 CIVIL\nA102 PLAN\n"})
|
||||||
|
assert declared == ["C-001", "A102"]
|
||||||
|
recovered = recover_sheet_number(("X\n" * 40) + "SHEET\nC-001\n")
|
||||||
|
recon = reconcile_sheets(
|
||||||
|
[{"sheet_number": recovered}, {"sheet_number": "A102"}], declared)
|
||||||
|
assert recon["declared_not_in_set"] == []
|
||||||
|
assert recon["in_set_not_declared"] == []
|
||||||
|
|
||||||
|
|
||||||
|
# --- R3: boolean env knob parsing ------------------------------------------
|
||||||
|
|
||||||
|
def test_numeric_one_enables_ladder_knobs():
|
||||||
|
with mock.patch.dict(os.environ, {
|
||||||
|
"EXTRACT_TEXT_RETRY_ENABLED": "1",
|
||||||
|
"EXTRACT_FALLBACK_ENABLED": "yes",
|
||||||
|
}):
|
||||||
|
cfg = importlib.reload(importlib.import_module("backend.config"))
|
||||||
|
try:
|
||||||
|
assert cfg.EXTRACT_TEXT_RETRY_ENABLED is True
|
||||||
|
assert cfg.EXTRACT_FALLBACK_ENABLED is True
|
||||||
|
finally:
|
||||||
|
importlib.reload(cfg)
|
||||||
|
|
||||||
|
|
||||||
|
def test_false_values_still_disable_ladder_knobs():
|
||||||
|
with mock.patch.dict(os.environ, {
|
||||||
|
"EXTRACT_TEXT_RETRY_ENABLED": "false",
|
||||||
|
"EXTRACT_FALLBACK_ENABLED": "0",
|
||||||
|
}):
|
||||||
|
cfg = importlib.reload(importlib.import_module("backend.config"))
|
||||||
|
try:
|
||||||
|
assert cfg.EXTRACT_TEXT_RETRY_ENABLED is False
|
||||||
|
assert cfg.EXTRACT_FALLBACK_ENABLED is False
|
||||||
|
finally:
|
||||||
|
importlib.reload(cfg)
|
||||||
@@ -0,0 +1,86 @@
|
|||||||
|
"""Deterministic sheet-list reconciliation: cover index vs extracted sheets."""
|
||||||
|
|
||||||
|
from backend.sheet_reconcile import declared_sheet_list, reconcile_sheets
|
||||||
|
|
||||||
|
COVER_TEXT = """VERIZON CYPRESS
|
||||||
|
SHEET LIST
|
||||||
|
SHEET NUMBER
|
||||||
|
SHEET NAME
|
||||||
|
G000
|
||||||
|
COVER
|
||||||
|
G001
|
||||||
|
GENERAL INFO
|
||||||
|
C-001
|
||||||
|
CIVIL COVER
|
||||||
|
C-001.1
|
||||||
|
ALTA SURVEY
|
||||||
|
L-101
|
||||||
|
LANDSCAPE PLAN
|
||||||
|
S101
|
||||||
|
FOUNDATION PLAN
|
||||||
|
S301
|
||||||
|
WALL SECTIONS
|
||||||
|
S401
|
||||||
|
PERSPECTIVE VIEW
|
||||||
|
A101
|
||||||
|
FLOOR PLAN
|
||||||
|
A102
|
||||||
|
REFLECTED CEILING PLAN
|
||||||
|
E400
|
||||||
|
ELECTRICAL SITE PLAN
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
def test_declared_sheet_list_from_cover():
|
||||||
|
declared = declared_sheet_list({1: COVER_TEXT, 2: "symbols legend"})
|
||||||
|
assert declared[0] == "G000"
|
||||||
|
assert "C-001" in declared and "C-001.1" in declared # hyphenated ids kept
|
||||||
|
assert "L-101" in declared
|
||||||
|
assert "A102" in declared
|
||||||
|
assert declared.count("G000") == 1
|
||||||
|
assert len(declared) == 11
|
||||||
|
|
||||||
|
|
||||||
|
def test_declared_sheet_list_uses_first_index_page_only():
|
||||||
|
texts = {1: "no index here", 2: COVER_TEXT, 3: "SHEET LIST\nXX999\nBOGUS"}
|
||||||
|
declared = declared_sheet_list(texts)
|
||||||
|
assert "XX999" not in declared # only the first marker page is parsed
|
||||||
|
|
||||||
|
|
||||||
|
def test_declared_sheet_list_none_when_no_marker():
|
||||||
|
assert declared_sheet_list({1: "just notes", 2: "floor plan stuff"}) == []
|
||||||
|
|
||||||
|
|
||||||
|
def _sheets(*nums):
|
||||||
|
return [{"page_number": i + 1, "sheet_number": n}
|
||||||
|
for i, n in enumerate(nums)]
|
||||||
|
|
||||||
|
|
||||||
|
def test_reconcile_both_directions():
|
||||||
|
declared = declared_sheet_list({1: COVER_TEXT})
|
||||||
|
rec = reconcile_sheets(_sheets("G000", "G001", "S101", "S301", "S302", "A101"),
|
||||||
|
declared)
|
||||||
|
# declared but not extracted (civil/landscape not in this PDF + missing)
|
||||||
|
assert "C-001" in rec["declared_not_in_set"]
|
||||||
|
assert "A102" in rec["declared_not_in_set"]
|
||||||
|
assert "E400" in rec["declared_not_in_set"]
|
||||||
|
# extracted but not on the cover index (misread or unlisted sheet)
|
||||||
|
assert rec["in_set_not_declared"] == ["S302"]
|
||||||
|
assert rec["declared_total"] == 11
|
||||||
|
assert rec["found_total"] == 6
|
||||||
|
|
||||||
|
|
||||||
|
def test_reconcile_normalizes_hyphens():
|
||||||
|
declared = ["C-001", "S301"]
|
||||||
|
rec = reconcile_sheets(_sheets("C001", "S301"), declared)
|
||||||
|
assert rec["declared_not_in_set"] == []
|
||||||
|
assert rec["in_set_not_declared"] == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_reconcile_ignores_unidentified_sheets():
|
||||||
|
rec = reconcile_sheets(
|
||||||
|
[{"page_number": 8, "sheet_number": None},
|
||||||
|
{"page_number": 9, "sheet_number": "S301"}],
|
||||||
|
["S301", "A102"])
|
||||||
|
assert rec["found_total"] == 1
|
||||||
|
assert rec["declared_not_in_set"] == ["A102"]
|
||||||
@@ -0,0 +1,76 @@
|
|||||||
|
from backend.text_coverage import (text_coverage, segment_text_layer,
|
||||||
|
fallback_objects, merge_objects,
|
||||||
|
recover_sheet_number)
|
||||||
|
|
||||||
|
|
||||||
|
def test_coverage_full():
|
||||||
|
text = "NOTE 1\nALL LUMBER NO. 2 SOUTHERN PINE\nNOTE 2\nUSE 5/8\" PLYWOOD"
|
||||||
|
objects = [{"source_text": "ALL LUMBER NO. 2 SOUTHERN PINE"},
|
||||||
|
{"source_text": "USE 5/8\" PLYWOOD"}]
|
||||||
|
cov = text_coverage(text, objects)
|
||||||
|
assert cov["covered_lines"] == 2
|
||||||
|
assert cov["total_lines"] == 2
|
||||||
|
assert cov["ratio"] == 1.0
|
||||||
|
|
||||||
|
|
||||||
|
def test_coverage_zero_on_empty_objects():
|
||||||
|
cov = text_coverage("LINE ALPHA CONTENT\nLINE BETA CONTENT\nLINE GAMMA CONTENT", [])
|
||||||
|
assert cov["ratio"] == 0.0 and cov["total_lines"] == 3
|
||||||
|
|
||||||
|
|
||||||
|
def test_coverage_ignores_short_and_numeric_noise_lines():
|
||||||
|
text = "15\"\n19\"\nA\nB\nREAL NOTE ABOUT FRAMING HERE"
|
||||||
|
cov = text_coverage(text, [{"source_text": "REAL NOTE ABOUT FRAMING HERE"}])
|
||||||
|
assert cov["total_lines"] == 1 and cov["ratio"] == 1.0
|
||||||
|
|
||||||
|
|
||||||
|
def test_segment_notes_and_rows():
|
||||||
|
text = "WOOD CONSTRUCTION\n1. \nALL SAWN LUMBER TO BE SOUTHERN PINE.\n2. \nROOF SHEATHING 5/8\" PLYWOOD."
|
||||||
|
segs = segment_text_layer(text)
|
||||||
|
assert any("ALL SAWN LUMBER" in s for s in segs)
|
||||||
|
assert any("ROOF SHEATHING" in s for s in segs)
|
||||||
|
|
||||||
|
|
||||||
|
def test_fallback_objects_verbatim_and_stamped():
|
||||||
|
objs = fallback_objects("1. \nALL SAWN LUMBER TO BE SOUTHERN PINE.", page_number=8)
|
||||||
|
assert len(objs) == 1
|
||||||
|
assert objs[0]["source_text"] == "1 ALL SAWN LUMBER TO BE SOUTHERN PINE."
|
||||||
|
assert objs[0]["grounding"] == "text_layer_fallback"
|
||||||
|
assert objs[0]["confidence"] == "low"
|
||||||
|
|
||||||
|
|
||||||
|
def test_merge_objects_keeps_vision_and_unions_text():
|
||||||
|
vision = [
|
||||||
|
{"source_text": "2X6 WD STUD @ 16\" O.C.", "object_type": "wall"},
|
||||||
|
{"source_text": None, "graphical_basis": "light fixture symbol, grid C-4",
|
||||||
|
"object_type": "lighting_fixture"},
|
||||||
|
]
|
||||||
|
text = [
|
||||||
|
{"source_text": "2X6 WD STUD @ 16\" O.C.", "object_type": "wall"},
|
||||||
|
{"source_text": "ALL LUMBER NO. 2 SOUTHERN PINE", "object_type": "general_note"},
|
||||||
|
]
|
||||||
|
merged = merge_objects(vision, text)
|
||||||
|
assert len(merged) == 3
|
||||||
|
assert any(o.get("graphical_basis") for o in merged)
|
||||||
|
assert merged[0]["object_type"] == "wall"
|
||||||
|
|
||||||
|
|
||||||
|
def test_merge_objects_dedupes_by_normalized_text():
|
||||||
|
a = [{"source_text": "RTU-1: 5 TON, 1600 CFM"}]
|
||||||
|
b = [{"source_text": "rtu 1 5 ton 1600 cfm"}]
|
||||||
|
assert len(merge_objects(a, b)) == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_recover_sheet_number_from_title_block():
|
||||||
|
text = ("WALL SECTIONS\n...\nSheet Information\nS301\n"
|
||||||
|
"Issue Date 05.29.26\nProject Number 25177")
|
||||||
|
assert recover_sheet_number(text) == "S301"
|
||||||
|
|
||||||
|
|
||||||
|
def test_recover_sheet_number_none_when_absent():
|
||||||
|
assert recover_sheet_number("just some notes about lumber") is None
|
||||||
|
|
||||||
|
|
||||||
|
def test_recover_prefers_discipline_pattern_over_dates():
|
||||||
|
text = "Issue Date 05.29.26\nProject Number 25177\nA102 REFLECTED CEILING PLAN"
|
||||||
|
assert recover_sheet_number(text) == "A102"
|
||||||
@@ -0,0 +1,111 @@
|
|||||||
|
"""Text-layer extraction, evidence bbox matching, and crop rendering."""
|
||||||
|
|
||||||
|
import os
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
fitz = pytest.importorskip("pymupdf")
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
from backend.text_layer import (
|
||||||
|
attach_text_layers,
|
||||||
|
coverage_gaps,
|
||||||
|
extract_text_layers,
|
||||||
|
find_evidence_bbox,
|
||||||
|
render_crop,
|
||||||
|
)
|
||||||
|
|
||||||
|
EVIDENCE = "(5) 2X6 STUD PACK @ 16 IN O.C."
|
||||||
|
|
||||||
|
|
||||||
|
def _make_pdf(path, pages):
|
||||||
|
"""pages: list of str ('' = effectively blank page)."""
|
||||||
|
doc = fitz.open()
|
||||||
|
for text in pages:
|
||||||
|
page = doc.new_page(width=612, height=792)
|
||||||
|
if text:
|
||||||
|
page.insert_text((72, 72), text, fontsize=11)
|
||||||
|
doc.save(str(path))
|
||||||
|
doc.close()
|
||||||
|
return str(path)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
def text_pdf(tmp_path):
|
||||||
|
return _make_pdf(tmp_path / "set.pdf", [EVIDENCE, ""])
|
||||||
|
|
||||||
|
|
||||||
|
def test_extract_text_layers(text_pdf):
|
||||||
|
layers = extract_text_layers(text_pdf)
|
||||||
|
assert set(layers) == {1, 2}
|
||||||
|
assert layers[1]["has_text_layer"] is True
|
||||||
|
assert "2X6 STUD PACK" in layers[1]["text"]
|
||||||
|
assert layers[1]["words"], "expected word-level bboxes"
|
||||||
|
assert all("bbox" in w and len(w["bbox"]) == 4 for w in layers[1]["words"])
|
||||||
|
|
||||||
|
|
||||||
|
def test_blank_page_below_min_chars(text_pdf):
|
||||||
|
layers = extract_text_layers(text_pdf)
|
||||||
|
assert layers[2]["has_text_layer"] is False
|
||||||
|
|
||||||
|
|
||||||
|
def test_disabled_returns_empty(text_pdf, monkeypatch):
|
||||||
|
monkeypatch.setattr(config, "TEXT_LAYER_ENABLED", False)
|
||||||
|
assert extract_text_layers(text_pdf) == {}
|
||||||
|
|
||||||
|
|
||||||
|
def test_attach_text_layers(text_pdf, tmp_path):
|
||||||
|
pages = [{"page_number": 1}, {"page_number": 2}]
|
||||||
|
words = attach_text_layers(text_pdf, pages,
|
||||||
|
text_dir=str(tmp_path / "text"))
|
||||||
|
assert pages[0]["text_layer"] and "STUD PACK" in pages[0]["text_layer"]
|
||||||
|
assert pages[1]["text_layer"] is None
|
||||||
|
assert words[1] and not words[2]
|
||||||
|
assert os.path.isfile(tmp_path / "text" / "page-001.txt")
|
||||||
|
assert not os.path.exists(tmp_path / "text" / "page-002.txt")
|
||||||
|
|
||||||
|
|
||||||
|
def test_find_evidence_bbox_exact(text_pdf):
|
||||||
|
words = extract_text_layers(text_pdf)[1]["words"]
|
||||||
|
bbox = find_evidence_bbox(words, EVIDENCE)
|
||||||
|
assert bbox is not None
|
||||||
|
assert bbox[2] > bbox[0] and bbox[3] > bbox[1]
|
||||||
|
|
||||||
|
|
||||||
|
def test_find_evidence_bbox_fuzzy(text_pdf):
|
||||||
|
# Vision quotes imperfectly: wrong count token, rest exact.
|
||||||
|
words = extract_text_layers(text_pdf)[1]["words"]
|
||||||
|
bbox = find_evidence_bbox(words, "(2) 2X6 STUD PACK @ 16 IN O.C.")
|
||||||
|
assert bbox is not None
|
||||||
|
|
||||||
|
|
||||||
|
def test_find_evidence_bbox_miss(text_pdf):
|
||||||
|
words = extract_text_layers(text_pdf)[1]["words"]
|
||||||
|
assert find_evidence_bbox(words, "PENTHOUSE EXHAUST FAN EF-9") is None
|
||||||
|
assert find_evidence_bbox([], EVIDENCE) is None
|
||||||
|
assert find_evidence_bbox(words, "") is None
|
||||||
|
|
||||||
|
|
||||||
|
def test_render_crop(text_pdf):
|
||||||
|
words = extract_text_layers(text_pdf)[1]["words"]
|
||||||
|
bbox = find_evidence_bbox(words, EVIDENCE)
|
||||||
|
crop = render_crop(text_pdf, 1, bbox)
|
||||||
|
assert crop is not None
|
||||||
|
# Decodes as an image of plausible size (margin around the text line).
|
||||||
|
doc = fitz.open(stream=crop, filetype="jpeg")
|
||||||
|
pix = doc[0].get_pixmap()
|
||||||
|
assert pix.width > 100 and pix.height > 20
|
||||||
|
doc.close()
|
||||||
|
|
||||||
|
|
||||||
|
def test_render_crop_bad_page(text_pdf):
|
||||||
|
assert render_crop(text_pdf, 99, (0, 0, 10, 10)) is None
|
||||||
|
|
||||||
|
|
||||||
|
def test_coverage_gaps():
|
||||||
|
pages = [{"page_number": 1, "text_layer": "some real text"},
|
||||||
|
{"page_number": 2, "text_layer": "more text"},
|
||||||
|
{"page_number": 3, "text_layer": None}]
|
||||||
|
sheets = [{"page_number": 1, "assertions": [{"id": "a"}]},
|
||||||
|
{"page_number": 2, "assertions": []}]
|
||||||
|
assert coverage_gaps(pages, sheets) == [2]
|
||||||
Reference in new issue
Block a user