Compare commits

...
6 Commits
Author SHA1 Message Date
woogi f7e1b6bb7c Merge main: dual model dropdowns + richer job logs, adapted for agent-mode.
Docker Release / build-and-push (push) Successful in 1m0s
Docker Release / release (push) Skipped
- llm.py: set_model_overrides(vision, text) replaces the single job override;
  UI picks still beat per-call agent model args, but never name the hybrid
  local model (avoids main's hybrid footgun); local->cloud fallback uses the
  text pick.
- jobs.py: timestamped line-split tee (job_log.py), in-memory log + log_tail
  polls, full log on terminal states (done/error/needs_review/finalization_error),
  log-only disk recovery, error email links to the run log, and failed runs now
  append the full traceback to job.log. Keeps pipeline_mode, job.json, and the
  review gate.
- models.py: vision/text split via architecture modalities, pricing kept;
  /models returns {vision, text, defaults}; /check takes vision_model/text_model
  (replacing model); /health adds text_model. models_catalog.py dropped.
- UI: two priced dropdowns (OpenRouter compute only) + live run-log panel.
- Tests updated for dual overrides and the /models shape; new coverage for
  traceback capture and local-model immunity.
2026-08-02 09:55:09 -05:00
John Wilganowski afa1089311 Add job run logs, OpenRouter model picker, and discipline grouping.
Docker Release / build-and-push (push) Successful in 55s
Docker Release / release (push) Skipped
- Job logs: each job's stdout/stderr is teed into outputs/<id>/job.log
  (survives restarts) and served at GET /jobs/{id}/log as text/plain, so
  full run logs can be shared for debugging and refinement.
- Model picker: GET /models proxies OpenRouter's public model list with
  per-1M-token pricing (1h cache, 502 on failure); the UI shows a model
  dropdown with costs when OpenRouter compute is selected, and the pick
  overrides vision+text models for that job (Classic and Agent modes).
- Conflicts in the report view are grouped by discipline pair
  (collapsible sections, severity-ordered within groups) instead of one
  flat severity-only list.
2026-07-28 21:23:42 +00:00
John Wilganowski 4ecc7c5cef Point email links at conchecker.scoutitsystems.com and show build SHA in header.
Docker Release / build-and-push (push) Successful in 58s
Docker Release / release (push) Skipped
APP_BASE_URL default (config, .env.example, both compose files) is now
https://conchecker.scoutitsystems.com with no port, so review-required
and final-report email links use the public site. CI bakes the short
commit SHA into the image as APP_BUILD via a Docker build-arg; /health
returns version+build and the site header shows the build so it's easy
to confirm which image is deployed. Local runs default to 'dev'.
2026-07-28 20:34:33 +00:00
John Wilganowski 1c1d2ff21b Add required human review gate to the Agent pipeline.
Docker Release / build-and-push (push) Successful in 1m10s
Docker Release / release (push) Skipped
Agent web jobs now stop after Brain consolidation and enter needs_review
with a persisted review queue (blocking: high-severity, low-confidence,
sensitive-category findings; audit sample of clean clusters). Humans
decide confirm/reject/unsure/needs_clarification via new review API and
frontend queue; a finalizer applies decisions (rejections suppressed with
reason codes), performs bounded targeted reruns for clarifications,
drafts RFIs only for kept issues, and only then marks the job done and
sends the final email. Two-phase email (review-required, then final
report), per-decision feedback labels with redacted aggregate metrics,
restart recovery from job artifacts, and CLI --no-review bypass.
Classic pipeline unchanged. 65 non-LLM tests.
2026-07-28 19:23:57 +00:00
woogiandCursor ac328d34fd Build agent-mode branch image in CI under a branch tag.
Docker Release / build-and-push (push) Successful in 3m38s
Docker Release / release (push) Has been skipped
Publish agent-mode pushes as :agent-mode (and :sha-<commit>) so the experimental fork is available from the registry without overwriting the Classic :latest image.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-18 14:42:13 +00:00
woogiandCursor 82a48d99cf Add scoped Agent-mode pipeline as experimental Classic fork.
Wire specialist waves, Brain consolidation, and Classic-compatible reports so Agent mode can run end-to-end via OpenRouter without changing the default Classic path.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-18 14:31:13 +00:00
53 changed files with 5450 additions and 242 deletions
+11 -2
View File
@@ -2,7 +2,7 @@ name: Docker Release
on:
push:
branches: [main]
branches: [main, agent-mode]
tags: ["v*"]
env:
@@ -25,12 +25,19 @@ jobs:
{
echo "tags<<EOF"
echo "${IMAGE}:latest"
echo "${IMAGE}:sha-${short_sha}"
if [ "${GITHUB_REF_TYPE}" = "tag" ]; then
ref_name="${GITHUB_REF_NAME}"
echo "${IMAGE}:${ref_name}"
echo "${IMAGE}:${ref_name#v}"
elif [ "${GITHUB_REF_NAME}" = "main" ]; then
# Only the stable Classic line publishes :latest.
echo "${IMAGE}:latest"
else
# Feature branches (e.g. agent-mode) publish under a branch tag
# so they never overwrite the default :latest image.
branch_tag="${GITHUB_REF_NAME//\//-}"
echo "${IMAGE}:${branch_tag}"
fi
echo "EOF"
} >> "$GITHUB_OUTPUT"
@@ -56,6 +63,8 @@ jobs:
context: .
push: true
tags: ${{ steps.meta.outputs.tags }}
build-args: |
APP_BUILD=sha-${{ steps.meta.outputs.short_sha }}
release:
needs: build-and-push
+7
View File
@@ -1,5 +1,8 @@
FROM python:3.12-slim-bookworm
LABEL org.opencontainers.image.title="Conflict Checker" \
org.opencontainers.image.description="Classic and experimental scoped Agent pipelines"
RUN apt-get update \
&& apt-get install -y --no-install-recommends poppler-utils \
&& rm -rf /var/lib/apt/lists/*
@@ -16,6 +19,10 @@ COPY cli cli
RUN mkdir -p backend/uploads backend/outputs backend/.llm_cache
ENV PYTHONUNBUFFERED=1
# Build identifier baked in by CI (sha-<short_sha>, matches the image tag);
# defaults to "dev" for local builds. Surfaced in /health and the site header.
ARG APP_BUILD=dev
ENV APP_BUILD=${APP_BUILD}
EXPOSE 8099
HEALTHCHECK --interval=30s --timeout=5s --start-period=10s --retries=3 \
+36 -23
View File
@@ -2,7 +2,7 @@
Orientation for a new coding session. Setup and Docker details live in [README.md](README.md). This file tracks what the code actually does and what tends to waste time.
**Last updated:** 2026-07-31 · tip `a6b0c8f` on Gitea `main`
**Last updated:** 2026-08-02 · `agent-mode` branch (merged `main` tip `bf508bf`)
## What this is
@@ -18,11 +18,13 @@ Source of truth: Scout IT Gitea — `gitea.scoutitsystems.com/woogi/Conflict_Che
|-------|--------|
| API | Python 3.12, FastAPI, Uvicorn ([backend/main.py](backend/main.py)) |
| UI | Single static file [frontend/index.html](frontend/index.html), served by FastAPI |
| Pipeline | Shared by web + CLI: [backend/pipeline/runner.py](backend/pipeline/runner.py) |
| LLM | OpenRouter via `openai` SDK; default `google/gemini-2.5-pro`. Vision always OpenRouter; text stages can use local vLLM |
| Pipelines | **Classic:** [backend/pipeline/runner.py](backend/pipeline/runner.py) (shared by web + CLI). **Agent:** [backend/agents/runner.py](backend/agents/runner.py) — experimental scoped specialist agents, selected per job (`pipeline_mode`) |
| Review gate | Agent jobs stop at `needs_review` for human decisions before the report emails ([backend/review/](backend/review/)) |
| LLM | OpenRouter via `openai` SDK; default `google/gemini-2.5-pro`. Vision always OpenRouter; text stages can use local vLLM (classic/hybrid only) |
| PDF | `pdf2image` + system `poppler-utils` → JPEG page images |
| Jobs | In-memory threads ([backend/jobs.py](backend/jobs.py)) — no Redis/DB |
| Deploy | Docker Compose; app on port **8099** |
| Tests | `pytest tests/` (~82 tests; see Quick start) |
| Deploy | Docker Compose; app on port **8099**; public URL `https://conchecker.scoutitsystems.com` |
## Live pipeline (authoritative)
@@ -53,15 +55,20 @@ PDF → images → extract → sheet index → jurisdiction
Design notes for Stage 2/3 engines also live under `Changes/*.docx`.
**Agent pipeline** (`pipeline_mode=agent`): scoped specialist agents in [backend/agents/](backend/agents/) (extractors, linker, brain, critics, RFI writer) run through `Orchestrator` + `ProjectMemory`; artifacts under `outputs/<job_id>/agent/`. Agent mode is OpenRouter-only (no hybrid) and, when `AGENT_REQUIRE_REVIEW=true`, stops at `needs_review` until a human saves decisions and finalizes via the review endpoints. Design docs: `docs/superpowers/`.
## HTTP API (current)
| Method | Path | Purpose |
|--------|------|---------|
| GET | `/health` | Liveness + default `model` / `text_model` + key/email flags |
| GET | `/models` | OpenRouter catalog split into `vision[]` / `text[]` + `defaults` (cached ~1h) |
| GET | `/health` | Liveness + `model` / `text_model` + `version` / `build` + key/email flags |
| GET | `/models` | OpenRouter catalog split into `vision[]` / `text[]` + `defaults`, with per-1M-token pricing (cached ~1h; **502** when OpenRouter is unreachable) |
| POST | `/check` | Upload PDF; returns `{job_id}` immediately |
| GET | `/jobs/{id}` | Status poll. Running: `stage` + `log_tail`. Done/error: `report` and/or `error` + full `log` |
| GET | `/jobs/{id}/log` | Full run log JSON (`lines`, `text`); `?plain=1` for text/plain |
| GET | `/jobs/{id}` | Status poll. Running: `stage` + `log_tail`. Done/error/needs_review/finalization_error: `report` and/or `error` + full `log` |
| GET | `/jobs/{id}/log` | Full run log as `text/plain` (404 when no log file) |
| GET | `/jobs/{id}/review` | Review queue + progress + saved decisions (agent jobs) |
| POST | `/jobs/{id}/review-decisions` | Save reviewer decisions (409 outside needs_review/reviewing) |
| POST | `/jobs/{id}/finalize-review` | Background finalize + send report (409 unless review gate passed) |
| GET | `/jobs/{id}/sheet-image/{page}` | JPEG of source PDF page for the sheet viewer |
| GET | `/` | Serves `frontend/index.html` |
@@ -69,7 +76,8 @@ Design notes for Stage 2/3 engines also live under `Changes/*.docx`.
- Required: `file` (PDF)
- Optional: `notification_email`, `project_name`, `address`, `occupancy`, `work_type`
- Compute: `text_local` (`true` = hybrid local text)
- Pipeline: `pipeline_mode` (`classic` default, or `agent`)
- Compute: `text_local` (`true` = hybrid local text; forced off for agent mode)
- Models: `vision_model`, `text_model` (OpenRouter ids; blank = config defaults)
## Job logs
@@ -77,10 +85,12 @@ Design notes for Stage 2/3 engines also live under `Changes/*.docx`.
Pipeline `print()` is teed for the job thread ([backend/job_log.py](backend/job_log.py)):
- Live: `GET /jobs/{id}``log_tail` (last 80 lines)
- Done/error: same payload includes full `log`
- Done/error/needs_review/finalization_error: same payload includes full `log`
- Disk: `backend/outputs/<job_id>/job.log` (survives restart; status registry does not)
- API: `GET /jobs/{id}/log` or `?plain=1`
- UI: “Run log” panel updates while running; stays visible after finish/fail
- API: `GET /jobs/{id}/log` `text/plain`
- UI: “Run log” panel updates while running; stays visible after finish/fail/review
- Failures append the **full traceback** to the log; each run starts with a header line (job id, mode, models, start time)
- `outputs/<job_id>/job.json` (written at start) carries email/mode/models so the disk fallback can rebuild a job after restart
## Vision vs text models
@@ -91,7 +101,7 @@ Two models, not one:
| Vision | `MODEL` | Extract, conflict reason (images) | Always OpenRouter |
| Text | `TEXT_MODEL` (falls back to `MODEL`) | Sheet index, jurisdiction, normalize, cluster(LLM), QAQC, code, construct, validate, risk, RFI | OpenRouter, or local when hybrid |
UI: two dropdowns filled from `GET /models` ([backend/models_catalog.py](backend/models_catalog.py)). Per-run picks go through `set_model_overrides()` in [backend/llm.py](backend/llm.py); runner clears them in `finally`. Hybrid: text dropdown also names the local model override and OpenRouter fallback.
UI: two dropdowns (with per-1M pricing) filled from `GET /models` ([backend/models.py](backend/models.py)), shown only for OpenRouter compute. Per-run picks go through `set_model_overrides(vision, text)` in [backend/llm.py](backend/llm.py): classic runs pass them as `run_pipeline` kwargs (runner clears in `finally`); agent runs set them module-level around `run_agent_pipeline`. UI picks beat per-call agent `AGENT_*_MODEL` args but **never name the hybrid local model** — local stays on `LOCAL_TEXT_MODEL`; the text pick only covers the cloud fallback.
## Where to change what
@@ -103,7 +113,9 @@ UI: two dropdowns filled from `GET /models` ([backend/models_catalog.py](backend
| UI (upload, models, live log, results) | [frontend/index.html](frontend/index.html) |
| CLI tuning loop | [cli/run_check.py](cli/run_check.py) |
| LLM client, cache, cost, model overrides | [backend/llm.py](backend/llm.py) |
| OpenRouter vision/text model lists | [backend/models_catalog.py](backend/models_catalog.py) |
| OpenRouter catalog + pricing + vision/text split | [backend/models.py](backend/models.py) |
| Agent-mode pipeline | [backend/agents/](backend/agents/) (`runner.py` entry; agents call `llm.call_json` with per-call model args) |
| Human review gate (queue, decisions, finalize) | [backend/review/](backend/review/) |
| Job registry + stdout tee log | [backend/jobs.py](backend/jobs.py), [backend/job_log.py](backend/job_log.py) |
| Stage helpers (prompt render, issue validate) | [backend/pipeline/_stage.py](backend/pipeline/_stage.py) |
| Code text corpus (Stage 7) | [backend/code_corpus/](backend/code_corpus/) |
@@ -115,7 +127,7 @@ Older prompt snapshot: `backend/prompts.py.v1`.
1. **Prompt placeholders** — Use `str.replace` via `_stage.render`, never `str.format`. Prompts contain literal `{` JSON braces.
2. **Clustering** — Default is LLM (`CLUSTERER=llm`); empty LLM result falls back to deterministic. Deterministic clusters need ≥2 disciplines (or schedule-vs-plan); single-discipline “missing” gaps are a known limit.
3. **Jobs are in-memory** — Process restart clears job status; `outputs/<job_id>/` (report + `job.log` + `source.pdf`) still reload via disk fallback.
3. **Jobs are in-memory** — Process restart clears job status; `outputs/<job_id>/` (report + `job.log` + `job.json` + `source.pdf`) still reload via disk fallback, including `needs_review` recovery.
4. **Dependency pin**`httpx==0.27.2` with `openai==1.51.0`. httpx ≥0.28 breaks openais `proxies=` kwarg.
5. **Code corpus licensing** — Only `ada_2010.txt` is shipped. Do not paste IBC/IFC/IECC without a license (see `backend/code_corpus/README.md`).
6. **Dual assertion schema** — Newer `{sheet, objects[]}` is mapped to legacy `{assertions[]}` with `attribute`/`value` for older stages.
@@ -124,23 +136,24 @@ Older prompt snapshot: `backend/prompts.py.v1`.
9. **Samples**`samples/*.pdf` are gitignored; drop PDFs locally for CLI runs.
10. **README drift** — Treat README for setup/CI; treat this file + `runner.py` for pipeline truth. `prompts.py` header may still say some prompts are unwired — they are wired through the runner.
11. **Git identity** — This box has no `user.name` / `user.email`; commits need `GIT_AUTHOR_*` / `GIT_COMMITTER_*` env vars (do not `git config`). Remote push to Gitea works.
12. **No local Python deps on host** — App is meant to run in Docker; bare `python3` imports may miss `dotenv` / `httpx`. Prefer `docker compose`.
12. **No local Python deps on host** — App is meant to run in Docker; bare `python3` imports may miss `dotenv` / `httpx`. Prefer `docker compose`. For tests on this box: venv + requirements, but unpin Pillow (`Pillow>=11`) — 10.4.0 doesn't build on Python 3.14 (Docker uses 3.12, where the pin is fine).
13. **Agent mode constraints** — OpenRouter-only (hybrid disabled in UI and forced off server-side); review gate statuses are `needs_review → reviewing → finalizing → done` (`finalization_error` on finalize failure); only terminal states include the full `log` in polls.
## Quick start pointers
- Full setup: [README.md](README.md) (`docker compose up -d --build` → http://localhost:8099).
- Local CLI: `python cli/run_check.py samples/your_set.pdf --out out/your_set`.
- Prompt iteration: set `LLM_CACHE=true` in `backend/.env` so unchanged stages replay for free; clear with `rm -rf backend/.llm_cache`.
- Artifacts per job: `assertions.json`, `clusters.json`, per-stage JSON, `conflicts.json`, `report.md`, `job.log`, `source.pdf` under `backend/outputs/<job_id>/`.
- No automated test suite; validate via CLI dumps and golden-set diffs (described in README).
- Artifacts per job: `assertions.json`, `clusters.json`, per-stage JSON, `conflicts.json`, `report.md`, `job.log`, `job.json`, `source.pdf` under `backend/outputs/<job_id>/`.
- Tests: `python -m pytest tests/` (needs the deps from `requirements.txt` + `pytest`; on this box use a venv, see gotcha #12).
## Recent work (2026-07-31)
## Recent work (2026-08-02, agent-mode)
Shipped on `main` as `a6b0c8f`:
Merged `main` tip (`a6b0c8f` + `bf508bf`) into `agent-mode`, reconciling with this branch's own earlier implementations:
- Per-job run log (tee + disk + API + UI)
- Separate vision/text model dropdowns backed by OpenRouter `/models`
- Session notes file (this doc)
- **Two model dropdowns** — main's vision/text split ported onto this branch's priced catalog (`models.py`); pickers stay OpenRouter-compute-only, and UI picks never override `LOCAL_TEXT_MODEL` (main's hybrid footgun avoided).
- **Better run logs** — main's timestamped line-splitting tee, `log_tail` polls, terminal-state full log, and log-only disk recovery merged with this branch's header line, `job.json` metadata, and review-gate states. Failed runs now also append the traceback to `job.log`.
- `backend/models_catalog.py` (main's unpriced catalog) intentionally dropped in favor of `models.py`.
## Conflict categories (taxonomy)
+75
View File
@@ -147,6 +147,76 @@ python cli/run_check.py samples/your_set.pdf --out out/your_set
# -> out/your_set/report.md + conflicts.json
```
The Classic pipeline remains the recommended default. The experimental Agent
fork runs in the same image and can be selected in the web UI or from the CLI:
```bash
python cli/run_check.py samples/your_set.pdf --mode agent --out out/agent-run
```
Agent mode uses OpenRouter for every model call and runs bounded specialist
waves: one-sheet extraction, sheet-index/jurisdiction orientation, semantic
linkers partitioned by level and object family, per-cluster conflict critics,
batched code review, cluster-scoped constructability, summary-only completeness,
central Brain consolidation, and one-finding RFI writers. It returns the same
`conflicts`, `validated_issues`, `rfis`, and `summary` fields as Classic.
Agent artifacts are also written under `<output>/agent/`, including wave
snapshots and the final Project Memory. `summary.agent_stats`,
`summary.cost_by_stage`, and `summary.models_used` are job-local, so concurrent
Agent jobs do not share accounting.
Optional `AGENT_*_MODEL` variables select an OpenRouter model per specialist.
The `AGENT_*_CONCURRENCY` and scope-cap variables in `backend/.env.example`
bound fan-out and prompt size. Agent mode intentionally ignores the hybrid/local
text option in v1.
### Agent mode: required human review
By default (`AGENT_REQUIRE_REVIEW=true`) an Agent run **stops after the Brain
consolidation wave** and waits for a human before anything ships:
```
Brain merge -> needs_review -> review UI (/?job=<id>) -> finalize -> final report
```
The job lifecycle adds review states: `needs_review` (queue built, waiting),
`reviewing` (decisions submitted), `finalizing` (targeted reruns + RFI writers
running), then `done` — or `finalization_error` if finalization fails. Open the
job in the web UI to work the queue: blocking items (high/critical severity,
low confidence, sensitive categories) must be decided; clean-cluster items are
non-blocking spot-checks.
Email is **two-phase**: a "review required" notice goes out when the job enters
`needs_review` (with a link to the review UI); the final conflict report email
is only sent after finalization completes. The unreviewed report never leaves
the server.
**Privacy boundary:** all review artifacts (queue, decisions, final report) are
job-local under `outputs/<job_id>/review/`. Cross-job review-feedback
aggregation, when built, excludes verbatim `source_text`, images, and comments
unless `REVIEW_AGGREGATE_INCLUDE_TEXT=true`.
Config knobs (see `backend/.env.example`):
| Key | Default | Effect |
|-----|---------|--------|
| `AGENT_REQUIRE_REVIEW` | `true` | `false` = Agent jobs skip the gate entirely (old behavior: RFIs, final report, one email) |
| `AGENT_REVIEW_AUDIT_SAMPLE` | `5` | Max clean clusters added to the queue as spot-checks |
| `REVIEW_AGGREGATE_INCLUDE_TEXT` | `false` | Allow future aggregate feedback to include source text/images/comments |
From the CLI, `--no-review` bypasses the gate for that run (it overrides
`AGENT_REQUIRE_REVIEW=true`):
```bash
python cli/run_check.py samples/your_set.pdf --mode agent --no-review --out out/agent-run
```
**Deployment note:** the review endpoints (`/jobs/{id}/review-decisions`,
`/jobs/{id}/finalize-review`) are **state-changing and sensitive** — they accept
human decisions that alter the final report. Do **not** expose the UI/API
publicly without reverse-proxy auth or a shared access token in front of it.
Web UI (upload + view):
```bash
@@ -155,6 +225,11 @@ uvicorn backend.main:app --reload --port 8099 # open http://127.0.0.1:8099
Or use Docker: `docker compose up -d` (see **Setup** above).
The standard Docker image contains both pipelines; no additional queue,
database, or model service is required. Set `AI_API_KEY` in `backend/.env` and
choose Agent mode per request. Treat Agent output as experimental and compare it
against a reviewed golden set before using it for issuance decisions.
## Conflict categories
`dimensional_disagreement`, `elevation_disagreement`, `location_mismatch`,
+32 -1
View File
@@ -4,6 +4,36 @@ AI_BASE_URL=https://openrouter.ai/api/v1
AI_API_KEY=sk-or-...
MODEL=google/gemini-2.5-pro
# Optional Agent-mode OpenRouter model overrides (inherit MODEL/TEXT_MODEL when blank)
AGENT_EXTRACT_MODEL=
AGENT_INDEX_MODEL=
AGENT_JURISDICTION_MODEL=
AGENT_LINKER_MODEL=
AGENT_CONFLICT_MODEL=
AGENT_CODE_MODEL=
AGENT_CONSTRUCT_MODEL=
AGENT_COMPLETENESS_MODEL=
AGENT_BRAIN_MODEL=
AGENT_RFI_MODEL=
# Agent-mode hard scope limits / concurrency
AGENT_LINK_MAX_ASSERTIONS=60
AGENT_CLUSTER_MAX_ASSERTIONS=24
AGENT_CONFLICT_MAX_IMAGES=6
AGENT_CODE_BATCH_SIZE=60
AGENT_BRAIN_MAX_TOKENS=16384
AGENT_LINK_CONCURRENCY=4
AGENT_CONFLICT_CONCURRENCY=4
AGENT_SPECIALIST_CONCURRENCY=4
AGENT_RFI_CONCURRENCY=4
# Agent-mode human-review gate (pipeline stops after Brain until a human reviews)
AGENT_REQUIRE_REVIEW=true
# Max clean clusters added to the review queue as non-blocking spot-checks
AGENT_REVIEW_AUDIT_SAMPLE=5
# Allow future cross-job review-feedback aggregation to include source_text/images/comments
REVIEW_AGGREGATE_INCLUDE_TEXT=false
# Pipeline tuning
PDF_DPI=100
MAX_PAGES=60
@@ -15,7 +45,8 @@ EXTRACT_CONCURRENCY=4
REASON_CONCURRENCY=4
# Public URL users reach this server on (used for the link in result emails)
APP_BASE_URL=http://localhost:8099
APP_BASE_URL=https://conchecker.scoutitsystems.com
# APP_BUILD is set by CI at image build time (sha-<short_sha>) - do not set manually.
# Email notifications (optional). Leave SMTP_HOST blank to disable.
# Examples:
+10
View File
@@ -0,0 +1,10 @@
"""Parallel, specialist-agent pipeline isolated from the Classic runner."""
def run_agent_pipeline(*args, **kwargs):
"""Lazy package-level entry point that avoids importing optional runtime deps."""
from backend.agents.runner import run_agent_pipeline as _run
return _run(*args, **kwargs)
__all__ = ["run_agent_pipeline"]
+89
View File
@@ -0,0 +1,89 @@
"""Shared contracts and job-local accounting for Agent-mode workers."""
import threading
from dataclasses import dataclass, field
from typing import Any, Dict, List, Optional, Protocol
@dataclass(frozen=True)
class AgentScope:
"""A bounded work package passed to exactly one specialist agent."""
scope_id: str
payload: Dict[str, Any] = field(default_factory=dict)
@dataclass
class AgentResult:
"""Artifacts returned by a specialist for collection by the orchestrator."""
scope_id: str
artifacts: List[Dict[str, Any]] = field(default_factory=list)
error: str = ""
@dataclass
class AgentUsage:
"""Thread-safe usage accounting owned by one Agent pipeline run."""
usd: float = 0.0
calls: int = 0
cached: int = 0
by_stage: Dict[str, Dict[str, Any]] = field(default_factory=dict)
models: Dict[str, set] = field(default_factory=lambda: {
"vision": set(),
"text_cloud": set(),
})
_lock: threading.Lock = field(default_factory=threading.Lock, repr=False)
def record(
self,
stage: str,
model: str,
usd: float = 0.0,
cached: bool = False,
has_images: bool = False,
) -> None:
with self._lock:
bucket = self.by_stage.setdefault(
stage, {"usd": 0.0, "calls": 0, "cached": 0}
)
if cached:
self.cached += 1
bucket["cached"] += 1
else:
self.calls += 1
self.usd += usd
bucket["calls"] += 1
bucket["usd"] += usd
family = "vision" if has_images else "text_cloud"
self.models[family].add(model)
def snapshot(self) -> Dict[str, Any]:
with self._lock:
return {
"usd": self.usd,
"calls": self.calls,
"cached": self.cached,
"by_stage": {k: dict(v) for k, v in self.by_stage.items()},
"models": {
"vision": sorted(self.models["vision"]),
"text_local": [],
"text_cloud": sorted(self.models["text_cloud"]),
"fallback_count": 0,
},
}
class ScopedAgent(Protocol):
"""Protocol implemented by each future specialist agent."""
name: str
def run(self, scope: AgentScope) -> AgentResult:
...
def failure(scope: AgentScope, error: Exception) -> AgentResult:
"""Convert a worker exception into a non-fatal scoped result."""
return AgentResult(scope_id=scope.scope_id, error=str(error))
+129
View File
@@ -0,0 +1,129 @@
"""Central merge, judge, and prioritization agent."""
import json
import re
from typing import Dict, List, Tuple
from backend import config
from backend.agents.base import AgentUsage
from backend.agents.prompts import BRAIN_SYSTEM_PROMPT, BRAIN_USER_PROMPT
from backend.llm import call_json
from backend.pipeline._stage import collect_list, validate_issue
def _finding_ref(finding: Dict, index: int) -> str:
return (
finding.get("issue_id")
or f"{finding.get('agent', 'agent')}:{finding.get('scope_id', '?')}:{index + 1}"
)
def _signature(finding: Dict) -> Tuple[str, str, str]:
norm = lambda value: re.sub(r"[^a-z0-9]+", " ", str(value).lower()).strip()
description = " ".join(norm(finding.get("description")).split()[:12])
return (
norm(finding.get("category")),
norm(finding.get("location")),
description,
)
def _fallback(findings: List[Dict]) -> Tuple[List[Dict], List[Dict]]:
"""Conservative local consolidation when the Brain call fails."""
kept: Dict[Tuple[str, str, str], Dict] = {}
refs: Dict[Tuple[str, str, str], List[str]] = {}
decisions: List[Dict] = []
severity_rank = {"critical": 4, "high": 3, "medium": 2, "low": 1}
for index, finding in enumerate(findings):
ref = _finding_ref(finding, index)
supported = bool(finding.get("evidence")) or finding.get("agent") == "completeness"
if not supported or not finding.get("description"):
decisions.append({
"finding_refs": [ref],
"action": "dropped",
"reason": "missing actionable support",
"kept_issue_id": None,
})
continue
signature = _signature(finding)
if signature not in kept:
kept[signature] = dict(finding)
refs[signature] = [ref]
else:
refs[signature].append(ref)
existing = kept[signature]
if severity_rank.get(finding.get("severity"), 2) > severity_rank.get(
existing.get("severity"), 2
):
existing["severity"] = finding.get("severity")
existing["evidence"] = (
existing.get("evidence") or []
) + (finding.get("evidence") or [])
issues = list(kept.values())
for index, (signature, issue) in enumerate(kept.items()):
issue["issue_id"] = issue.get("issue_id") or f"AGENT-{index + 1:04d}"
issue["risk_score"] = {
"critical": 95, "high": 75, "medium": 50, "low": 25
}.get(issue.get("severity"), 50)
issue["recommended_priority"] = {
"critical": "immediate",
"high": "before_bid",
"medium": "before_construction",
"low": "track_only",
}.get(issue.get("severity"), "before_construction")
decisions.append({
"finding_refs": refs[signature],
"action": "merged" if len(refs[signature]) > 1 else "kept",
"reason": "conservative deterministic fallback",
"kept_issue_id": issue["issue_id"],
})
issues.sort(key=lambda item: -int(item.get("risk_score") or 0))
return issues, decisions
class BrainAgent:
name = "brain"
def __init__(self, usage: AgentUsage) -> None:
self.usage = usage
def run(
self,
findings: List[Dict],
sheet_index: Dict,
jurisdiction: Dict,
) -> Tuple[List[Dict], List[Dict]]:
instruction = BRAIN_USER_PROMPT
for key, value in {
"sheet_index": sheet_index,
"jurisdiction": jurisdiction,
"findings": findings,
}.items():
instruction = instruction.replace(
"{" + key + "}", json.dumps(value, ensure_ascii=True)
)
parsed = call_json(
system_prompt=BRAIN_SYSTEM_PROMPT,
user_text=instruction,
max_tokens=config.AGENT_BRAIN_MAX_TOKENS,
model=config.AGENT_BRAIN_MODEL,
usage_tracker=self.usage,
usage_stage="agent.brain",
)
issues = collect_list(
parsed, "issues", lambda item: validate_issue(item, item.get("source_stage", ""))
)
if not issues:
return _fallback(findings)
raw_issues = parsed.get("issues") if isinstance(parsed, dict) else []
for index, issue in enumerate(issues):
raw = raw_issues[index] if index < len(raw_issues) else {}
issue["issue_id"] = issue.get("issue_id") or f"AGENT-{index + 1:04d}"
issue["risk_score"] = raw.get("risk_score") or issue.get("risk_score") or 50
issue["recommended_priority"] = (
raw.get("recommended_priority") or "before_construction"
)
issues.sort(key=lambda item: -int(item.get("risk_score") or 0))
decisions = parsed.get("decisions") or []
return issues, [item for item in decisions if isinstance(item, dict)]
+95
View File
@@ -0,0 +1,95 @@
"""Scoped code/accessibility review agents."""
from typing import Dict, List
from backend import config
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
from backend.llm import call_json
from backend.pipeline import code_refs
from backend.pipeline._serialize import dumps, slim_sheets
from backend.pipeline._stage import collect_list, validate_issue
from backend.pipeline.jurisdiction import active_review_paths
from backend.prompts import CODE_REVIEW_SYSTEM_PROMPT, CODE_REVIEW_USER_INSTRUCTION
def build_code_scopes(
sheets: List[Dict], jurisdiction: Dict, sheet_index: Dict
) -> List[AgentScope]:
cap = max(1, config.AGENT_CODE_BATCH_SIZE)
fragments: List[Dict] = []
for sheet in sheets:
assertions = sheet.get("assertions") or []
if not assertions:
continue
for offset in range(0, len(assertions), cap):
fragments.append({
**sheet,
"assertions": assertions[offset:offset + cap],
})
batches: List[List[Dict]] = []
current: List[Dict] = []
count = 0
for fragment in fragments:
size = len(fragment["assertions"])
if current and count + size > cap:
batches.append(current)
current, count = [], 0
current.append(fragment)
count += size
if current:
batches.append(current)
return [
AgentScope(
scope_id=f"code:{index + 1}",
payload={
"sheets": batch,
"jurisdiction": jurisdiction,
"sheet_index": sheet_index,
},
)
for index, batch in enumerate(batches)
]
class CodeAgent:
name = "code_reviewer"
def __init__(self, usage: AgentUsage) -> None:
self.usage = usage
def run(self, scope: AgentScope) -> AgentResult:
try:
sheets = scope.payload.get("sheets") or []
jurisdiction = scope.payload.get("jurisdiction") or {}
sheet_index = scope.payload.get("sheet_index") or {}
assertions = [
assertion for sheet in sheets
for assertion in sheet.get("assertions", [])
]
excerpts = code_refs.retrieve(
active_review_paths(jurisdiction), assertions
)
instruction = CODE_REVIEW_USER_INSTRUCTION
for key, value in {
"jurisdiction": dumps(jurisdiction),
"sheet_index": dumps(sheet_index),
"assertions": dumps(slim_sheets(sheets)),
"code_references": code_refs.format_excerpts(excerpts),
}.items():
instruction = instruction.replace("{" + key + "}", value)
parsed = call_json(
system_prompt=CODE_REVIEW_SYSTEM_PROMPT,
user_text=instruction,
max_tokens=config.CODE_MAX_TOKENS,
model=config.AGENT_CODE_MODEL,
usage_tracker=self.usage,
usage_stage="agent.code",
)
findings = collect_list(
parsed, "issues", lambda item: validate_issue(item, "code")
)
for finding in findings:
finding.update(agent=self.name, scope_id=scope.scope_id)
return AgentResult(scope_id=scope.scope_id, artifacts=findings)
except Exception as exc:
return failure(scope, exc)
+55
View File
@@ -0,0 +1,55 @@
"""Summary-only drawing-set completeness agent."""
import json
from typing import Dict, List
from backend import config
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
from backend.agents.prompts import COMPLETENESS_SYSTEM_PROMPT, COMPLETENESS_USER_PROMPT
from backend.llm import call_json
from backend.pipeline._stage import collect_list, validate_issue
def build_sheet_summaries(sheets: List[Dict]) -> List[Dict]:
"""Return counts and classifications only; never raw assertions."""
return [{
"sheet_number": sheet.get("sheet_number"),
"sheet_title": sheet.get("sheet_title"),
"discipline": sheet.get("discipline"),
"drawing_type": sheet.get("drawing_type"),
"level": sheet.get("level"),
"assertion_count": len(sheet.get("assertions") or []),
"unresolved_count": len(sheet.get("unresolved_items") or []),
} for sheet in sheets]
class CompletenessAgent:
name = "completeness"
def __init__(self, usage: AgentUsage) -> None:
self.usage = usage
def run(self, scope: AgentScope) -> AgentResult:
try:
instruction = COMPLETENESS_USER_PROMPT
for key in ("sheet_index", "sheet_summaries", "cluster_summary"):
instruction = instruction.replace(
"{" + key + "}",
json.dumps(scope.payload.get(key) or {}, ensure_ascii=True),
)
parsed = call_json(
system_prompt=COMPLETENESS_SYSTEM_PROMPT,
user_text=instruction,
max_tokens=config.QAQC_MAX_TOKENS,
model=config.AGENT_COMPLETENESS_MODEL,
usage_tracker=self.usage,
usage_stage="agent.completeness",
)
findings = collect_list(
parsed, "issues", lambda item: validate_issue(item, "qaqc")
)
for finding in findings:
finding.update(agent=self.name, scope_id=scope.scope_id)
return AgentResult(scope_id=scope.scope_id, artifacts=findings)
except Exception as exc:
return failure(scope, exc)
+74
View File
@@ -0,0 +1,74 @@
"""Per-cluster conflict critics with hard evidence and image caps."""
from typing import Dict, List
from backend import config
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
from backend.llm import call_json
from backend.pipeline.conflict_checker import _evidence_block, _valid_conflict
from backend.prompts import CONFLICT_SYSTEM_PROMPT, CONFLICT_USER_INSTRUCTION
def _as_finding(conflict: Dict, scope_id: str) -> Dict:
return {
"issue_id": conflict.get("conflict_id") or "",
"source_stage": "conflict",
"category": conflict.get("category") or "uncategorized",
"severity": conflict.get("severity") or "medium",
"confidence": conflict.get("confidence") or "medium",
"location": conflict.get("location") or "",
"disciplines": conflict.get("disciplines") or [],
"sheets": conflict.get("sheets") or [],
"description": conflict.get("description") or "",
"evidence": conflict.get("evidence") or [],
"recommended_resolution": conflict.get("recommended_resolution") or "",
"code_reference": None,
"agent": "conflict_critic",
"scope_id": scope_id,
}
class ConflictCriticAgent:
name = "conflict_critic"
def __init__(self, usage: AgentUsage) -> None:
self.usage = usage
def run(self, scope: AgentScope) -> AgentResult:
try:
cluster = dict(scope.payload["cluster"])
cluster["assertions"] = (
cluster.get("assertions") or []
)[:config.AGENT_CLUSTER_MAX_ASSERTIONS]
page_to_b64: Dict[int, str] = scope.payload.get("page_to_b64") or {}
images: List[str] = []
for page_number in (
cluster.get("page_numbers") or []
)[:config.AGENT_CONFLICT_MAX_IMAGES]:
if page_to_b64.get(page_number):
images.append(page_to_b64[page_number])
instruction = (
CONFLICT_USER_INSTRUCTION
.replace("{location}", cluster.get("location") or "")
.replace("{evidence}", _evidence_block(cluster))
)
parsed = call_json(
system_prompt=CONFLICT_SYSTEM_PROMPT,
user_text=instruction,
images_b64=images,
max_tokens=config.REASON_MAX_TOKENS,
model=config.AGENT_CONFLICT_MODEL,
usage_tracker=self.usage,
usage_stage="agent.conflict",
)
candidates = parsed if isinstance(parsed, list) else (
parsed.get("conflicts") if isinstance(parsed, dict) else []
)
findings = []
for candidate in candidates or []:
conflict = _valid_conflict(candidate, cluster)
if conflict:
findings.append(_as_finding(conflict, scope.scope_id))
return AgentResult(scope_id=scope.scope_id, artifacts=findings)
except Exception as exc:
return failure(scope, exc)
+70
View File
@@ -0,0 +1,70 @@
"""Zone/cluster-scoped constructability agents."""
from typing import Dict, List
from backend import config
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
from backend.llm import call_json
from backend.pipeline._serialize import dumps, slim_clusters
from backend.pipeline._stage import collect_list, validate_issue
from backend.prompts import (
CONSTRUCTABILITY_SYSTEM_PROMPT,
CONSTRUCTABILITY_USER_INSTRUCTION,
)
def build_construct_scopes(
clusters: List[Dict], conflict_findings: List[Dict]
) -> List[AgentScope]:
scopes = []
for index, cluster in enumerate(clusters):
related = [
finding for finding in conflict_findings
if finding.get("scope_id") == f"conflict:{cluster.get('key')}"
or finding.get("location") == cluster.get("location")
]
scopes.append(AgentScope(
scope_id=f"construct:{index + 1}",
payload={"cluster": cluster, "conflicts": related},
))
return scopes
class ConstructabilityAgent:
name = "constructability"
def __init__(self, usage: AgentUsage) -> None:
self.usage = usage
def run(self, scope: AgentScope) -> AgentResult:
try:
cluster = dict(scope.payload["cluster"])
cluster["assertions"] = (
cluster.get("assertions") or []
)[:config.AGENT_CLUSTER_MAX_ASSERTIONS]
instruction = CONSTRUCTABILITY_USER_INSTRUCTION
substitutions = {
"assertions": dumps(cluster["assertions"]),
"clusters": dumps(slim_clusters([cluster])),
"conflicts": dumps(scope.payload.get("conflicts") or []),
}
for key, value in substitutions.items():
instruction = instruction.replace("{" + key + "}", value)
parsed = call_json(
system_prompt=CONSTRUCTABILITY_SYSTEM_PROMPT,
user_text=instruction,
max_tokens=config.CONSTRUCT_MAX_TOKENS,
model=config.AGENT_CONSTRUCT_MODEL,
usage_tracker=self.usage,
usage_stage="agent.constructability",
)
findings = collect_list(
parsed,
"issues",
lambda item: validate_issue(item, "constructability"),
)
for finding in findings:
finding.update(agent=self.name, scope_id=scope.scope_id)
return AgentResult(scope_id=scope.scope_id, artifacts=findings)
except Exception as exc:
return failure(scope, exc)
+106
View File
@@ -0,0 +1,106 @@
"""Scoped extraction and orientation agents."""
import json
from typing import Dict
from backend import config
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
from backend.llm import call_json
from backend.pipeline.extractor import _normalize_sheet
from backend.pipeline.sheet_index import _index_input
from backend.prompts import (
EXTRACTOR_SYSTEM_PROMPT,
EXTRACTOR_USER_INSTRUCTION,
JURISDICTION_SYSTEM_PROMPT,
JURISDICTION_USER_INSTRUCTION,
SHEET_INDEX_SYSTEM_PROMPT,
SHEET_INDEX_USER_INSTRUCTION,
)
class SheetExtractorAgent:
name = "sheet_extractor"
def __init__(self, usage: AgentUsage) -> None:
self.usage = usage
def run(self, scope: AgentScope) -> AgentResult:
try:
page = scope.payload["page"]
instruction = EXTRACTOR_USER_INSTRUCTION.replace(
"{sheet_hint}", str(scope.payload.get("sheet_hint") or "")
)
parsed = call_json(
system_prompt=EXTRACTOR_SYSTEM_PROMPT,
user_text=instruction,
images_b64=[page["base64"]],
max_tokens=config.EXTRACT_MAX_TOKENS,
model=config.AGENT_EXTRACT_MODEL,
usage_tracker=self.usage,
usage_stage="agent.extract",
)
if not isinstance(parsed, dict):
raise ValueError("no structured extraction returned")
sheet = _normalize_sheet(parsed, page["page_number"])
return AgentResult(scope_id=scope.scope_id, artifacts=[sheet])
except Exception as exc:
return failure(scope, exc)
class SheetIndexAgent:
name = "sheet_index"
def __init__(self, usage: AgentUsage) -> None:
self.usage = usage
def run(self, scope: AgentScope) -> AgentResult:
try:
sheets = scope.payload.get("sheets") or []
instruction = SHEET_INDEX_USER_INSTRUCTION.replace(
"{sheet_index_input}",
json.dumps(_index_input(sheets), ensure_ascii=True),
)
parsed = call_json(
system_prompt=SHEET_INDEX_SYSTEM_PROMPT,
user_text=instruction,
max_tokens=config.SHEET_INDEX_MAX_TOKENS,
model=config.AGENT_INDEX_MODEL,
usage_tracker=self.usage,
usage_stage="agent.sheet_index",
)
if isinstance(parsed, list):
parsed = {"sheet_index": parsed, "missing_expected_sheets": []}
if not isinstance(parsed, dict):
raise ValueError("no sheet index returned")
return AgentResult(scope_id=scope.scope_id, artifacts=[parsed])
except Exception as exc:
return failure(scope, exc)
class JurisdictionAgent:
name = "jurisdiction"
def __init__(self, usage: AgentUsage) -> None:
self.usage = usage
def run(self, scope: AgentScope) -> AgentResult:
try:
project_input: Dict = scope.payload.get("project_input") or {}
instruction = JURISDICTION_USER_INSTRUCTION.replace(
"{project_input}", json.dumps(project_input, ensure_ascii=True)
)
parsed = call_json(
system_prompt=JURISDICTION_SYSTEM_PROMPT,
user_text=instruction,
max_tokens=config.JURISDICTION_MAX_TOKENS,
model=config.AGENT_JURISDICTION_MODEL,
usage_tracker=self.usage,
usage_stage="agent.jurisdiction",
)
if not isinstance(parsed, dict):
raise ValueError("no jurisdiction profile returned")
profile = parsed.get("project_code_profile")
artifact = profile if isinstance(profile, dict) else parsed
return AgentResult(scope_id=scope.scope_id, artifacts=[artifact])
except Exception as exc:
return failure(scope, exc)
+193
View File
@@ -0,0 +1,193 @@
"""Bounded semantic linkers that build coordination clusters."""
import json
import re
from collections import defaultdict
from typing import Dict, Iterable, List, Tuple
from backend import config
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
from backend.llm import call_json
from backend.pipeline.clusterer import cluster_by_location
from backend.pipeline.llm_clusterer import _location
from backend.prompts import CLUSTER_SYSTEM_PROMPT, CLUSTER_USER_INSTRUCTION
def _family(assertion: Dict) -> str:
location = assertion.get("location_key") or {}
if location.get("room"):
return "room"
if location.get("grid"):
return "grid"
if location.get("detail_reference"):
return "detail"
tag = str(location.get("tag") or "")
match = re.match(r"[A-Za-z]+", tag)
return (
(match.group(0).lower() if match else "")
or (assertion.get("object_type") or "").lower()
or "general"
)
def build_link_scopes(sheets: List[Dict]) -> List[AgentScope]:
"""Partition facts by level and object/tag family, then enforce a hard cap."""
buckets: Dict[Tuple[str, str], List[Dict]] = defaultdict(list)
for sheet in sheets:
for assertion in sheet.get("assertions", []):
enriched = {
**assertion,
"discipline": sheet.get("discipline") or "Unknown",
"sheet_number": sheet.get("sheet_number"),
"page_number": sheet.get("page_number"),
}
level = str((assertion.get("location_key") or {}).get("level")
or sheet.get("level") or "unknown").lower()
buckets[(level, _family(assertion))].append(enriched)
scopes: List[AgentScope] = []
cap = max(2, config.AGENT_LINK_MAX_ASSERTIONS)
for (level, family), assertions in sorted(buckets.items()):
for offset in range(0, len(assertions), cap):
chunk = assertions[offset:offset + cap]
if len(chunk) < 2:
continue
scopes.append(AgentScope(
scope_id=f"{level}:{family}:{offset // cap + 1}",
payload={"assertions": chunk, "level": level, "family": family},
))
return scopes
def _payload(assertions: Iterable[Dict]) -> List[Dict]:
return [{
"assertion_id": item.get("id"),
"discipline": item.get("discipline"),
"sheet_number": item.get("sheet_number"),
"attribute": item.get("attribute"),
"value": item.get("value"),
"location_key": item.get("location_key"),
"source_text": item.get("source_text"),
} for item in assertions]
def _fallback(assertions: List[Dict]) -> List[Dict]:
"""Use the deterministic linker within this scope when semantic linking fails."""
by_sheet: Dict[Tuple, Dict] = {}
for item in assertions:
key = (item.get("sheet_number"), item.get("page_number"))
sheet = by_sheet.setdefault(key, {
"sheet_number": item.get("sheet_number"),
"page_number": item.get("page_number"),
"discipline": item.get("discipline"),
"assertions": [],
})
sheet["assertions"].append(item)
return cluster_by_location(list(by_sheet.values()))
class LinkerAgent:
name = "linker"
def __init__(self, usage: AgentUsage) -> None:
self.usage = usage
def run(self, scope: AgentScope) -> AgentResult:
try:
assertions = scope.payload.get("assertions") or []
by_id = {item.get("id"): item for item in assertions if item.get("id")}
instruction = CLUSTER_USER_INSTRUCTION.replace(
"{normalized_assertions}",
json.dumps(_payload(assertions), ensure_ascii=True),
)
parsed = call_json(
system_prompt=CLUSTER_SYSTEM_PROMPT,
user_text=instruction,
max_tokens=config.CLUSTER_MAX_TOKENS,
model=config.AGENT_LINKER_MODEL,
usage_tracker=self.usage,
usage_stage="agent.link",
)
raw = parsed if isinstance(parsed, list) else (
parsed.get("clusters") if isinstance(parsed, dict) else []
)
clusters: List[Dict] = []
for candidate in raw or []:
if not isinstance(candidate, dict):
continue
members = [
by_id[item_id]
for item_id in candidate.get("assertion_ids") or []
if item_id in by_id
]
if len(members) < 2:
continue
allowed_sheets = []
for member in members:
sheet = member.get("sheet_number")
if sheet not in allowed_sheets:
allowed_sheets.append(sheet)
allowed_sheets = allowed_sheets[:config.AGENT_CONFLICT_MAX_IMAGES]
members = [
member for member in members
if member.get("sheet_number") in allowed_sheets
][:config.AGENT_CLUSTER_MAX_ASSERTIONS]
primary = candidate.get("primary_location_key") or {}
clusters.append({
"key": f"{scope.scope_id}:{candidate.get('cluster_id') or len(clusters) + 1}",
"location": _location(primary),
"disciplines": sorted({
member.get("discipline") or "Unknown" for member in members
}),
"page_numbers": sorted({
member["page_number"] for member in members
if member.get("page_number")
}),
"sheets": sorted({
member["sheet_number"] for member in members
if member.get("sheet_number")
}),
"assertions": members,
"kind": "agent_semantic",
"scope_id": scope.scope_id,
})
if not clusters:
clusters = _fallback(assertions)
for cluster in clusters:
cluster["scope_id"] = scope.scope_id
cluster["kind"] = "agent_deterministic"
return AgentResult(scope_id=scope.scope_id, artifacts=clusters)
except Exception as exc:
return failure(scope, exc)
def build_object_graph(clusters: List[Dict]) -> Dict:
"""Build a deterministic graph view from linker output."""
nodes = []
edges = []
seen = set()
for cluster in clusters:
cluster_id = cluster.get("key")
nodes.append({
"id": cluster_id,
"type": "cluster",
"location": cluster.get("location"),
"sheets": cluster.get("sheets") or [],
})
for assertion in cluster.get("assertions") or []:
assertion_id = assertion.get("id")
if not assertion_id:
continue
if assertion_id not in seen:
seen.add(assertion_id)
nodes.append({
"id": assertion_id,
"type": assertion.get("object_type") or "assertion",
"sheet": assertion.get("sheet_number"),
})
edges.append({
"source": assertion_id,
"target": cluster_id,
"relationship": "member_of",
})
return {"nodes": nodes, "edges": edges}
+67
View File
@@ -0,0 +1,67 @@
"""Thread-safe per-job blackboard for the Agent pipeline."""
import copy
import json
import os
import threading
from typing import Any, Dict, Iterable, Optional
_COLLECTION_KEYS = {"sheets", "clusters", "findings", "decisions", "rfis"}
_MAPPING_KEYS = {"sheet_index", "jurisdiction", "object_graph"}
_MEMORY_KEYS = _COLLECTION_KEYS | _MAPPING_KEYS
class ProjectMemory:
"""Owns intermediate Agent-mode state and optional debug snapshots."""
def __init__(self, artifact_dir: Optional[str] = None) -> None:
self.artifact_dir = artifact_dir
self._lock = threading.RLock()
self._data: Dict[str, Any] = {
**{key: [] for key in _COLLECTION_KEYS},
**{key: {} for key in _MAPPING_KEYS},
}
if artifact_dir:
os.makedirs(artifact_dir, exist_ok=True)
def replace(self, key: str, value: Any) -> None:
"""Replace one named memory section."""
self._validate_key(key)
with self._lock:
self._data[key] = copy.deepcopy(value)
def append(self, key: str, value: Dict[str, Any]) -> None:
"""Append one artifact to a list-backed memory section."""
if key not in _COLLECTION_KEYS:
raise KeyError(f"{key!r} is not an appendable memory section")
with self._lock:
self._data[key].append(copy.deepcopy(value))
def extend(self, key: str, values: Iterable[Dict[str, Any]]) -> None:
"""Append several artifacts under one lock."""
if key not in _COLLECTION_KEYS:
raise KeyError(f"{key!r} is not an appendable memory section")
with self._lock:
self._data[key].extend(copy.deepcopy(list(values)))
def snapshot(self) -> Dict[str, Any]:
"""Return a detached, JSON-serializable view of current state."""
with self._lock:
return copy.deepcopy(self._data)
def dump(self, filename: str = "memory.json") -> Optional[str]:
"""Persist a snapshot when this job has an artifact directory."""
if not self.artifact_dir:
return None
path = os.path.join(self.artifact_dir, filename)
temp_path = f"{path}.tmp"
with open(temp_path, "w", encoding="utf-8") as f:
json.dump(self.snapshot(), f, indent=2)
os.replace(temp_path, path)
return path
@staticmethod
def _validate_key(key: str) -> None:
if key not in _MEMORY_KEYS:
raise KeyError(f"Unknown project memory section: {key!r}")
+78
View File
@@ -0,0 +1,78 @@
"""Wave scheduler for the Agent pipeline."""
from concurrent.futures import ThreadPoolExecutor, as_completed
from dataclasses import dataclass, field
from typing import Callable, Dict, Iterable, List, Optional
from backend.agents.base import AgentResult, AgentScope, ScopedAgent
from backend.agents.memory import ProjectMemory
@dataclass
class AgentStats:
"""Job-local accounting; never shared across concurrent jobs."""
calls: int = 0
scopes: int = 0
merges: int = 0
failed_scopes: List[str] = field(default_factory=list)
def as_dict(self) -> Dict:
return {
"calls": self.calls,
"scopes": self.scopes,
"merges": self.merges,
"failed_scopes": list(self.failed_scopes),
}
class Orchestrator:
"""Coordinates bounded fan-out/fan-in waves against one ProjectMemory."""
def __init__(
self,
memory: ProjectMemory,
on_stage: Optional[Callable[[str], None]] = None,
) -> None:
self.memory = memory
self.on_stage = on_stage
self.stats = AgentStats()
def stage(self, name: str) -> None:
print(f"\n=== {name} ===")
if self.on_stage:
self.on_stage(name)
def initialize(self) -> Dict:
"""Initialize the job-local artifact store."""
self.stage("Initialize agent pipeline")
self.memory.dump()
return self.stats.as_dict()
def run_scopes(
self,
agent: ScopedAgent,
scopes: Iterable[AgentScope],
concurrency: int,
) -> List[AgentResult]:
"""Run independent scopes; one failure never aborts the wave."""
scope_list = list(scopes)
if not scope_list:
return []
results: List[AgentResult] = []
with ThreadPoolExecutor(max_workers=max(1, concurrency)) as pool:
futures = {pool.submit(agent.run, scope): scope for scope in scope_list}
for future in as_completed(futures):
scope = futures[future]
self.stats.scopes += 1
try:
result = future.result()
except Exception as exc:
result = AgentResult(scope_id=scope.scope_id, error=str(exc))
if result.error:
self.stats.failed_scopes.append(
f"{agent.name}:{scope.scope_id}: {result.error}"
)
results.append(result)
results.sort(key=lambda result: result.scope_id)
return results
+25
View File
@@ -0,0 +1,25 @@
"""Prompts unique to the scoped Agent pipeline."""
COMPLETENESS_SYSTEM_PROMPT = """You are a senior construction-document completeness reviewer.
Review only the supplied sheet index and aggregate counts. Identify missing sheets,
schedules, details, or clearly incomplete coverage. Do not infer drawing facts and do
not report direct design conflicts. A missing-information finding may use the sheet
index itself as evidence. Respond only with valid JSON."""
COMPLETENESS_USER_PROMPT = """Review this summarized drawing set for completeness.
Return {"issues":[{"issue_id":"string","source_stage":"qaqc","category":"missing_sheet | missing_schedule | missing_detail | missing_information | bid_readiness | permit_readiness | other","severity":"critical | high | medium | low","confidence":"high | medium | low","location":"sheet or drawing set","disciplines":["string"],"sheets":["string"],"description":"string","evidence":[],"recommended_resolution":"string","code_reference":null}]}.
Sheet index: {sheet_index}
Aggregate sheet summaries: {sheet_summaries}
Cluster summary: {cluster_summary}"""
BRAIN_SYSTEM_PROMPT = """You are the central decision layer for a construction drawing
review. Merge duplicate specialist findings, reject vague or unsupported findings,
preserve verbatim evidence, and prioritize the kept issues. Do not create new issues.
Conflicts need drawing evidence; completeness findings may instead cite an explicit
missing item from the sheet index. Return only valid JSON."""
BRAIN_USER_PROMPT = """Judge and consolidate these scoped specialist findings.
Return {"issues":[{"issue_id":"string","source_stage":"conflict | qaqc | code | constructability","category":"string","severity":"critical | high | medium | low","confidence":"high | medium | low","location":"string","disciplines":["string"],"sheets":["string"],"description":"string","evidence":[{"discipline":"string","sheet":"string","source_text":"string","asserted_value":"string"}],"recommended_resolution":"string","code_reference":"string or null","risk_score":1,"recommended_priority":"immediate | before_bid | before_permit | before_construction | track_only"}],"decisions":[{"finding_refs":["string"],"action":"kept | merged | dropped","reason":"string","kept_issue_id":"string or null"}]}.
Sheet index: {sheet_index}
Jurisdiction summary: {jurisdiction}
Specialist findings: {findings}"""
+42
View File
@@ -0,0 +1,42 @@
"""One-finding-per-call RFI writers."""
from backend import config
from backend.agents.base import AgentResult, AgentScope, AgentUsage, failure
from backend.llm import call_json
from backend.pipeline._serialize import dumps
from backend.pipeline.rfi import _valid_rfi
from backend.prompts import RFI_SYSTEM_PROMPT, RFI_USER_INSTRUCTION
class RFIWriterAgent:
name = "rfi_writer"
def __init__(self, usage: AgentUsage) -> None:
self.usage = usage
def run(self, scope: AgentScope) -> AgentResult:
try:
finding = scope.payload["finding"]
instruction = RFI_USER_INSTRUCTION.replace(
"{prioritized_issues}", dumps([finding])
)
parsed = call_json(
system_prompt=RFI_SYSTEM_PROMPT,
user_text=instruction,
max_tokens=config.RFI_MAX_TOKENS,
model=config.AGENT_RFI_MODEL,
usage_tracker=self.usage,
usage_stage="agent.rfi",
)
candidates = parsed if isinstance(parsed, list) else (
parsed.get("rfi_comments") if isinstance(parsed, dict) else []
)
rfis = []
for candidate in candidates or []:
rfi = _valid_rfi(candidate)
if rfi:
rfi["issue_id"] = rfi.get("issue_id") or finding.get("issue_id")
rfis.append(rfi)
return AgentResult(scope_id=scope.scope_id, artifacts=rfis[:1])
except Exception as exc:
return failure(scope, exc)
+379
View File
@@ -0,0 +1,379 @@
"""Public entry point for the scoped Agent-mode pipeline."""
import json
import os
from typing import Callable, Dict, Optional
from backend import config
from backend.agents.base import AgentScope, AgentUsage
from backend.agents.brain import BrainAgent
from backend.agents.code_agent import CodeAgent, build_code_scopes
from backend.agents.completeness import CompletenessAgent, build_sheet_summaries
from backend.agents.conflict_critic import ConflictCriticAgent
from backend.agents.construct_agent import ConstructabilityAgent, build_construct_scopes
from backend.agents.extractors import (
JurisdictionAgent,
SheetExtractorAgent,
SheetIndexAgent,
)
from backend.agents.linker import LinkerAgent, build_link_scopes, build_object_graph
from backend.agents.memory import ProjectMemory
from backend.agents.orchestrator import Orchestrator
from backend.agents.rfi_writer import RFIWriterAgent
from backend.pipeline.pdf_processor import convert_pdf_to_images
from backend.pipeline.report import build_report, to_markdown
from backend.pipeline.sheet_index import derive_project_meta_from_cover
from backend.review.gate import build_review_queue
from backend.review.store import ReviewStore
def run_agent_pipeline(
pdf_path: str,
out_dir: Optional[str] = None,
on_stage: Optional[Callable[[str], None]] = None,
project_input: Optional[Dict] = None,
source_name: Optional[str] = None,
require_review: bool = True,
) -> Dict:
"""Run all scoped specialist waves and return a Classic-compatible report."""
if not os.path.isfile(pdf_path):
raise FileNotFoundError(pdf_path)
agent_dir = os.path.join(out_dir, "agent") if out_dir else None
memory = ProjectMemory(artifact_dir=agent_dir)
orchestrator = Orchestrator(memory=memory, on_stage=on_stage)
usage = AgentUsage()
orchestrator.initialize()
orchestrator.stage("Agent ingest: PDF -> images")
pages = convert_pdf_to_images(pdf_path)
page_to_b64 = {page["page_number"]: page["base64"] for page in pages}
orchestrator.stage("Agent wave 1: extract sheets")
extract_scopes = [
AgentScope(
scope_id=f"sheet:{page['page_number']}",
payload={"page": page},
)
for page in pages
]
extract_results = orchestrator.run_scopes(
SheetExtractorAgent(usage), extract_scopes, config.EXTRACT_CONCURRENCY
)
sheets = [
artifact
for result in extract_results
for artifact in result.artifacts
]
sheets.sort(key=lambda sheet: sheet.get("page_number") or 0)
memory.replace("sheets", sheets)
memory.dump("01-extract.json")
cover_meta = derive_project_meta_from_cover(
sheets, source_name or os.path.basename(pdf_path)
)
merged_input = {**cover_meta, **(project_input or {})}
orchestrator.stage("Agent wave 2: sheet index and jurisdiction")
index_results = orchestrator.run_scopes(
SheetIndexAgent(usage),
[AgentScope("sheet-index", {"sheets": sheets})],
1,
)
jurisdiction_results = orchestrator.run_scopes(
JurisdictionAgent(usage),
[AgentScope("jurisdiction", {"project_input": merged_input})],
1,
)
sheet_index = (
index_results[0].artifacts[0]
if index_results and index_results[0].artifacts else {}
)
jurisdiction = (
jurisdiction_results[0].artifacts[0]
if jurisdiction_results and jurisdiction_results[0].artifacts else {}
)
memory.replace("sheet_index", sheet_index)
memory.replace("jurisdiction", jurisdiction)
memory.dump("02-orient.json")
orchestrator.stage("Agent wave 3: scoped semantic linking")
link_results = orchestrator.run_scopes(
LinkerAgent(usage),
build_link_scopes(sheets),
config.AGENT_LINK_CONCURRENCY,
)
clusters = [
artifact
for result in link_results
for artifact in result.artifacts
][:config.CLUSTER_MAX]
object_graph = build_object_graph(clusters)
memory.replace("clusters", clusters)
memory.replace("object_graph", object_graph)
memory.dump("03-link.json")
orchestrator.stage("Agent wave 4: per-cluster conflict critics")
conflict_scopes = [
AgentScope(
scope_id=f"conflict:{cluster.get('key')}",
payload={"cluster": cluster, "page_to_b64": page_to_b64},
)
for cluster in clusters
]
conflict_results = orchestrator.run_scopes(
ConflictCriticAgent(usage),
conflict_scopes,
config.AGENT_CONFLICT_CONCURRENCY,
)
conflict_findings = [
artifact
for result in conflict_results
for artifact in result.artifacts
]
memory.extend("findings", conflict_findings)
orchestrator.stage("Agent wave 5: scoped specialists")
code_results = orchestrator.run_scopes(
CodeAgent(usage),
build_code_scopes(sheets, jurisdiction, sheet_index),
config.AGENT_SPECIALIST_CONCURRENCY,
)
construct_results = orchestrator.run_scopes(
ConstructabilityAgent(usage),
build_construct_scopes(clusters, conflict_findings),
config.AGENT_SPECIALIST_CONCURRENCY,
)
completeness_scope = AgentScope("completeness", {
"sheet_index": sheet_index,
"sheet_summaries": build_sheet_summaries(sheets),
"cluster_summary": {
"count": len(clusters),
"by_kind": _counts(clusters, "kind"),
},
})
completeness_results = orchestrator.run_scopes(
CompletenessAgent(usage), [completeness_scope], 1
)
specialist_findings = [
artifact
for result in code_results + construct_results + completeness_results
for artifact in result.artifacts
]
memory.extend("findings", specialist_findings)
gap_findings = [
{
"issue_id": f"AGENT-GAP-{index + 1:03d}",
"source_stage": "qaqc",
"category": "analysis_gap",
"severity": "low",
"confidence": "high",
"location": failed_scope.split(":", 2)[1] if ":" in failed_scope else "",
"disciplines": [],
"sheets": [],
"description": f"Agent analysis scope did not complete: {failed_scope}",
"evidence": [],
"recommended_resolution": "Review this scope manually or rerun the job.",
"code_reference": None,
"agent": "completeness",
"scope_id": "failed-scopes",
}
for index, failed_scope in enumerate(orchestrator.stats.failed_scopes)
]
memory.extend("findings", gap_findings)
memory.dump("05-specialists.json")
orchestrator.stage("Agent wave 6: Brain merge, judge, prioritize")
all_findings = memory.snapshot()["findings"]
if all_findings:
prioritized, decisions = BrainAgent(usage).run(
all_findings, sheet_index, jurisdiction
)
else:
prioritized, decisions = [], []
memory.extend("decisions", decisions)
orchestrator.stats.merges = sum(
1 for decision in decisions if decision.get("action") == "merged"
)
if require_review:
orchestrator.stage("Agent review gate: build human-review queue")
memory_snapshot = memory.snapshot()
queue = build_review_queue(memory_snapshot, prioritized, decisions,
limit=config.AGENT_REVIEW_AUDIT_SAMPLE)
store = ReviewStore(out_dir)
store.write_queue(queue)
candidate_conflicts = [_finding_as_conflict(item) for item in conflict_findings]
report = build_report(
conflicts=candidate_conflicts,
sheets=sheets,
clusters=clusters,
source=source_name or os.path.basename(pdf_path),
)
report.update({
"project_input": merged_input,
"jurisdiction": jurisdiction,
"sheet_index": sheet_index,
"project_intelligence": object_graph,
"validated_issues": prioritized,
"rfis": [],
"suppressed_issues": [],
})
progress = store.progress(queue)
# Same usage/stats summary block as the wave-7 path (rfis: 0 — they
# are drafted only after human review finalizes the run).
cost = usage.snapshot()
orchestrator.stats.calls = cost["calls"]
stats = orchestrator.stats.as_dict()
report["summary"].update({
"pipeline_mode": "agent",
"agent_status": "needs_review",
"review": progress,
"agent_stats": stats,
"by_stage": {
"conflicts": len(conflict_findings),
"qaqc": sum(
1 for item in specialist_findings
if item.get("source_stage") == "qaqc"
),
"code": sum(
1 for item in specialist_findings
if item.get("source_stage") == "code"
),
"constructability": sum(
1 for item in specialist_findings
if item.get("source_stage") == "constructability"
),
"validated": len(prioritized),
"rfis": 0,
},
"cost_usd": round(cost["usd"], 4),
"llm_calls": cost["calls"],
"cached_calls": cost["cached"],
"cost_by_stage": cost["by_stage"],
"text_backend": "openrouter",
"models_used": cost["models"],
})
if out_dir:
_dump(out_dir, "conflicts.json", report)
_dump(out_dir, "validated_issues.json", prioritized)
# Snapshot for the review finalizer's targeted clarification reruns.
memory.dump("memory.json")
return report
orchestrator.stage("Agent wave 7: per-finding RFI writers")
rfi_scopes = [
AgentScope(
scope_id=f"rfi:{finding.get('issue_id') or index + 1}",
payload={"finding": finding},
)
for index, finding in enumerate(prioritized)
]
rfi_results = orchestrator.run_scopes(
RFIWriterAgent(usage), rfi_scopes, config.AGENT_RFI_CONCURRENCY
)
rfis = [
artifact for result in rfi_results for artifact in result.artifacts
]
memory.extend("rfis", rfis)
memory.dump("memory.json")
orchestrator.stage("Build agent report")
conflicts = [_finding_as_conflict(item) for item in conflict_findings]
report = build_report(
conflicts=conflicts,
sheets=sheets,
clusters=clusters,
source=source_name or os.path.basename(pdf_path),
)
report.update({
"project_input": merged_input,
"jurisdiction": jurisdiction,
"sheet_index": sheet_index,
"project_intelligence": object_graph,
"validated_issues": prioritized,
"rfis": rfis,
})
cost = usage.snapshot()
orchestrator.stats.calls = cost["calls"]
stats = orchestrator.stats.as_dict()
report["summary"].update({
"pipeline_mode": "agent",
"agent_status": "complete",
"agent_stats": stats,
"by_stage": {
"conflicts": len(conflict_findings),
"qaqc": sum(
1 for item in specialist_findings
if item.get("source_stage") == "qaqc"
),
"code": sum(
1 for item in specialist_findings
if item.get("source_stage") == "code"
),
"constructability": sum(
1 for item in specialist_findings
if item.get("source_stage") == "constructability"
),
"validated": len(prioritized),
"rfis": len(rfis),
},
"cost_usd": round(cost["usd"], 4),
"llm_calls": cost["calls"],
"cached_calls": cost["cached"],
"cost_by_stage": cost["by_stage"],
"text_backend": "openrouter",
"models_used": cost["models"],
})
if out_dir:
os.makedirs(out_dir, exist_ok=True)
_dump(out_dir, "assertions.json", sheets)
_dump(out_dir, "clusters.json", [_without_base64(item) for item in clusters])
_dump(out_dir, "sheet_index.json", sheet_index)
_dump(out_dir, "jurisdiction.json", jurisdiction)
_dump(out_dir, "project_intelligence.json", object_graph)
_dump(out_dir, "validated_issues.json", prioritized)
_dump(out_dir, "rfis.json", rfis)
_dump(out_dir, "conflicts.json", report)
with open(os.path.join(out_dir, "report.md"), "w", encoding="utf-8") as f:
f.write(to_markdown(report))
return report
def _dump(out_dir: str, name: str, value) -> None:
with open(os.path.join(out_dir, name), "w", encoding="utf-8") as f:
json.dump(value, f, indent=2)
def _counts(items, key: str) -> Dict[str, int]:
counts: Dict[str, int] = {}
for item in items:
value = str(item.get(key) or "unknown")
counts[value] = counts.get(value, 0) + 1
return counts
def _finding_as_conflict(finding: Dict) -> Dict:
return {
"category": finding.get("category") or "uncategorized",
"severity": finding.get("severity") or "medium",
"disciplines": finding.get("disciplines") or [],
"location": finding.get("location") or "",
"sheets": finding.get("sheets") or [],
"description": finding.get("description") or "",
"evidence": finding.get("evidence") or [],
"recommended_resolution": finding.get("recommended_resolution") or "",
"confidence": finding.get("confidence") or "medium",
"cluster_key": finding.get("scope_id"),
}
def _without_base64(cluster: Dict) -> Dict:
return {
**cluster,
"assertions": [
{key: value for key, value in assertion.items() if key != "base64"}
for assertion in cluster.get("assertions") or []
],
}
+40 -1
View File
@@ -22,6 +22,42 @@ MODEL = os.getenv("MODEL", "google/gemini-2.5-pro")
# MODEL when unset.
TEXT_MODEL = os.getenv("TEXT_MODEL", "") or MODEL
# Agent-mode model overrides (OpenRouter IDs). Empty values inherit the
# matching general-purpose model so the skeleton requires no extra config.
AGENT_EXTRACT_MODEL = os.getenv("AGENT_EXTRACT_MODEL", "") or MODEL
AGENT_INDEX_MODEL = os.getenv("AGENT_INDEX_MODEL", "") or TEXT_MODEL
AGENT_JURISDICTION_MODEL = os.getenv("AGENT_JURISDICTION_MODEL", "") or TEXT_MODEL
AGENT_LINKER_MODEL = os.getenv("AGENT_LINKER_MODEL", "") or TEXT_MODEL
AGENT_CONFLICT_MODEL = os.getenv("AGENT_CONFLICT_MODEL", "") or MODEL
AGENT_CODE_MODEL = os.getenv("AGENT_CODE_MODEL", "") or TEXT_MODEL
AGENT_CONSTRUCT_MODEL = os.getenv("AGENT_CONSTRUCT_MODEL", "") or TEXT_MODEL
AGENT_COMPLETENESS_MODEL = os.getenv("AGENT_COMPLETENESS_MODEL", "") or TEXT_MODEL
AGENT_BRAIN_MODEL = os.getenv("AGENT_BRAIN_MODEL", "") or TEXT_MODEL
AGENT_RFI_MODEL = os.getenv("AGENT_RFI_MODEL", "") or TEXT_MODEL
# Agent-mode hard scope limits. These are intentionally independent of Classic
# batching so Agent workers can never grow into whole-set reasoning calls.
AGENT_LINK_MAX_ASSERTIONS = int(os.getenv("AGENT_LINK_MAX_ASSERTIONS", "60"))
AGENT_CLUSTER_MAX_ASSERTIONS = int(os.getenv("AGENT_CLUSTER_MAX_ASSERTIONS", "24"))
AGENT_CONFLICT_MAX_IMAGES = int(os.getenv("AGENT_CONFLICT_MAX_IMAGES", "6"))
AGENT_CODE_BATCH_SIZE = int(os.getenv("AGENT_CODE_BATCH_SIZE", "60"))
AGENT_BRAIN_MAX_TOKENS = int(os.getenv("AGENT_BRAIN_MAX_TOKENS", "16384"))
AGENT_LINK_CONCURRENCY = int(os.getenv("AGENT_LINK_CONCURRENCY", "4"))
AGENT_CONFLICT_CONCURRENCY = int(os.getenv("AGENT_CONFLICT_CONCURRENCY", "4"))
AGENT_SPECIALIST_CONCURRENCY = int(os.getenv("AGENT_SPECIALIST_CONCURRENCY", "4"))
AGENT_RFI_CONCURRENCY = int(os.getenv("AGENT_RFI_CONCURRENCY", "4"))
# Agent-mode human-review gate. When on (default), Agent runs stop after the
# Brain merge and wait for human decisions before RFIs/final report/email go
# out. AGENT_REVIEW_AUDIT_SAMPLE caps how many clean clusters get added to the
# queue as non-blocking spot-checks. REVIEW_AGGREGATE_INCLUDE_TEXT controls
# whether future cross-job review feedback aggregation may include verbatim
# source_text/images/comments (off by default = privacy-preserving).
AGENT_REQUIRE_REVIEW = os.getenv("AGENT_REQUIRE_REVIEW", "true").strip().lower() in ("1", "true", "yes")
AGENT_REVIEW_AUDIT_SAMPLE = int(os.getenv("AGENT_REVIEW_AUDIT_SAMPLE", "5"))
# NOTE: currently unwired - reserved for future cross-job aggregation tooling.
REVIEW_AGGREGATE_INCLUDE_TEXT = os.getenv("REVIEW_AGGREGATE_INCLUDE_TEXT", "false").strip().lower() in ("1", "true", "yes")
# -- Hybrid (local text LLM) ----------------------------------------
# Optional OpenAI-compatible local endpoint (e.g. a vLLM box) for the text-only
# QAQC stages. Vision stages ALWAYS use OpenRouter. The user picks hybrid per
@@ -87,7 +123,10 @@ APP_VERSION = "0.1.0"
# Public base URL used to build the "view results" link in notification
# emails. Set to whatever address users reach this server on (e.g. the
# Tailscale/LAN URL) so the link in the email actually resolves.
APP_BASE_URL = os.getenv("APP_BASE_URL", "http://localhost:8099")
APP_BASE_URL = os.getenv("APP_BASE_URL", "https://conchecker.scoutitsystems.com")
# Build identifier baked into the Docker image by CI (sha-<short_sha>, matching
# the image tag). Shown in the site header and /health. "dev" for local runs.
APP_BUILD = os.getenv("APP_BUILD", "dev")
# -- Email / SMTP (optional notification on completion) -------------
# If unset, the app still works; it just logs "SMTP not configured" and
+16
View File
@@ -38,6 +38,22 @@ def _send(msg: EmailMessage) -> bool:
return False
def send_review_required(recipient_email: str, report: dict, review_url: str) -> bool:
if not recipient_email or not _smtp_ready():
return False
msg = EmailMessage()
msg["Subject"] = f"Conflict Checker - review required - {report.get('source', 'drawing set')}"
msg["From"] = config.SMTP_FROM or config.SMTP_USER
msg["To"] = recipient_email
review = report.get("summary", {}).get("review", {})
msg.set_content(
"Agent analysis is complete and waiting for human review.\n\n"
f"Required review items: {review.get('required', 0)}\n"
f"Review URL: {review_url}\n"
)
return _send(msg)
def send_conflict_report(
recipient_email: str,
report: Dict,
+129 -39
View File
@@ -14,21 +14,28 @@ A teed stdout/stderr log is kept in memory and written to outputs/<job_id>/job.l
so failed or suspicious runs can be reviewed after the fact.
"""
import json
import os
import time
import traceback
import uuid
import shutil
import threading
from typing import Dict, List, Optional
from backend import config
from backend import llm
from backend.job_log import capture_stdio, read_log_file, stamp_line
from backend.agents.runner import run_agent_pipeline
from backend.pipeline.runner import run_pipeline
from backend.email_sender import send_conflict_report
from backend.email_sender import send_conflict_report, send_review_required
_jobs: Dict[str, Dict] = {}
_lock = threading.Lock()
_LOG_TAIL = 80
PIPELINE_MODES = {"classic", "agent"}
# States where the job will produce no more log output; polls get the full log.
_TERMINAL_STATES = {"done", "error", "needs_review", "finalization_error"}
def _set(job_id: str, **fields) -> None:
@@ -59,19 +66,26 @@ def create_job(
email: Optional[str] = None,
project_input: Optional[Dict] = None,
text_local: bool = False,
pipeline_mode: str = "classic",
vision_model: Optional[str] = None,
text_model: Optional[str] = None,
) -> str:
"""Register a job and kick off its background thread. Returns the job_id."""
pipeline_mode = pipeline_mode.strip().lower()
if pipeline_mode not in PIPELINE_MODES:
raise ValueError(f"Unsupported pipeline mode: {pipeline_mode!r}")
# Agent mode v1 is OpenRouter-only.
text_local = bool(text_local and pipeline_mode == "classic")
job_id = uuid.uuid4().hex[:12]
with _lock:
_jobs[job_id] = {
"job_id": job_id,
"status": "queued", # queued -> running -> done | error
"status": "queued", # queued -> running -> done | needs_review | error
"source": source_filename,
"email": email or None,
"project_input": project_input or {},
"text_local": text_local,
"pipeline_mode": pipeline_mode,
"vision_model": (vision_model or "").strip() or None,
"text_model": (text_model or "").strip() or None,
"stage": None,
@@ -83,7 +97,8 @@ def create_job(
}
threading.Thread(
target=_run,
args=(job_id, pdf_path, project_input, text_local, vision_model, text_model),
args=(job_id, pdf_path, project_input, text_local, pipeline_mode,
vision_model, text_model),
daemon=True,
).start()
return job_id
@@ -94,6 +109,7 @@ def _run(
pdf_path: str,
project_input: Optional[Dict] = None,
text_local: bool = False,
pipeline_mode: str = "classic",
vision_model: Optional[str] = None,
text_model: Optional[str] = None,
) -> None:
@@ -101,33 +117,28 @@ def _run(
log_path = os.path.join(out_dir, "job.log")
try:
_set(job_id, status="running")
# Keep a copy of the source PDF so its sheets can be viewed later.
os.makedirs(out_dir, exist_ok=True)
# Truncate any leftover log if job_id somehow collided (shouldn't).
with open(log_path, "w", encoding="utf-8"):
pass
shutil.copy2(pdf_path, os.path.join(out_dir, "source.pdf"))
header = (f"=== Job {job_id} | {pipeline_mode} | {_jobs[job_id].get('source')} | "
f"vision={vision_model or 'default'} text={text_model or 'default'} | "
f"started {time.strftime('%Y-%m-%d %H:%M:%S %Z', time.gmtime())} UTC ===")
_append_log(job_id, header, log_path)
def on_line(raw: str) -> None:
_append_log(job_id, raw, log_path)
with capture_stdio(on_line):
report = run_pipeline(
pdf_path,
out_dir=out_dir,
on_stage=lambda name: _set(job_id, stage=name),
project_input=project_input,
source_name=_jobs[job_id].get("source"),
text_local=text_local,
vision_model=vision_model,
text_model=text_model,
)
_set(job_id, status="done", report=report, finished_at=time.time(), stage=None)
_notify(job_id, report, out_dir)
_run_pipeline(job_id, pdf_path, out_dir, project_input, text_local,
pipeline_mode, vision_model, text_model)
except Exception as e:
# Also land in the job log via print under the tee when possible.
# Land the failure AND its traceback in the job log so failed runs can
# be diagnosed from the log alone (the stdio tee is already torn down).
try:
_append_log(job_id, f"[Jobs] Job {job_id} failed: {e}", log_path)
for ln in traceback.format_exc().rstrip().splitlines():
_append_log(job_id, ln, log_path)
except Exception:
pass
print(f"[Jobs] Job {job_id} failed: {e}")
@@ -140,6 +151,63 @@ def _run(
pass
def _run_pipeline(job_id: str, pdf_path: str, out_dir: str,
project_input: Optional[Dict], text_local: bool,
pipeline_mode: str, vision_model: Optional[str],
text_model: Optional[str]) -> None:
"""The body of a job run; executes inside the job's tee'd log capture."""
# Persist minimal job metadata so the disk fallback in get_job can
# recover the recipient email / pipeline mode after a server restart
# (plain json.dump, matching the _dump style used elsewhere).
with open(os.path.join(out_dir, "job.json"), "w", encoding="utf-8") as f:
json.dump({
"job_id": job_id,
"email": _jobs[job_id].get("email"),
"pipeline_mode": pipeline_mode,
"source": _jobs[job_id].get("source"),
"vision_model": vision_model,
"text_model": text_model,
}, f, indent=2)
# Keep a copy of the source PDF so its sheets can be viewed later.
shutil.copy2(pdf_path, os.path.join(out_dir, "source.pdf"))
runner = run_agent_pipeline if pipeline_mode == "agent" else run_pipeline
runner_kwargs = {
"out_dir": out_dir,
"on_stage": lambda name: _set(job_id, stage=name),
"project_input": project_input,
"source_name": _jobs[job_id].get("source"),
}
if pipeline_mode == "classic":
# run_pipeline takes the picks as params and clears them in finally.
runner_kwargs["text_local"] = text_local
runner_kwargs["vision_model"] = vision_model
runner_kwargs["text_model"] = text_model
else:
runner_kwargs["require_review"] = config.AGENT_REQUIRE_REVIEW
# The agent runner has no override params; set them module-level.
if vision_model or text_model:
print(f"[Jobs] Model overrides for this run: "
f"vision={vision_model or '(default)'} text={text_model or '(default)'}")
llm.set_model_overrides(vision_model, text_model)
try:
report = runner(pdf_path, **runner_kwargs)
finally:
if pipeline_mode == "agent":
llm.set_model_overrides(None, None)
report.setdefault("summary", {})["pipeline_mode"] = pipeline_mode
if report["summary"].get("agent_status") == "needs_review":
# Human-review gate: hold the job, don't email the unreviewed report.
_set(job_id, status="needs_review", report=report,
finished_at=time.time(), stage=None)
email = _jobs[job_id].get("email")
if email:
review_url = f"{config.APP_BASE_URL.rstrip('/')}/?job={job_id}"
send_review_required(email, report, review_url)
else:
_set(job_id, status="done", report=report, finished_at=time.time(), stage=None)
_notify(job_id, report, out_dir)
def _notify(job_id: str, report: Dict, out_dir: str) -> None:
email = _jobs[job_id].get("email")
if not email:
@@ -159,7 +227,6 @@ def _notify_error(job_id: str) -> None:
email = job.get("email")
if not email:
return
# Reuse the report mailer with a minimal error-shaped payload.
try:
from backend.email_sender import _smtp_ready, _send
from email.message import EmailMessage
@@ -204,8 +271,8 @@ def get_job_log(job_id: str) -> Optional[List[str]]:
def get_job(job_id: str) -> Optional[Dict]:
"""Public job view. Includes the full report only when done.
Falls back to the on-disk conflicts.json when the job isn't in the
in-memory registry (e.g. after a server restart).
Falls back to the on-disk artifacts (conflicts.json / job.log) when the
job isn't in the in-memory registry (e.g. after a server restart).
"""
with _lock:
job = _jobs.get(job_id)
@@ -215,7 +282,7 @@ def get_job(job_id: str) -> Optional[Dict]:
out["log_tail"] = log[-_LOG_TAIL:]
# Full log on terminal states so the UI can show it without a
# second fetch; keep polls light while running.
if out.get("status") in ("done", "error"):
if out.get("status") in _TERMINAL_STATES:
out["log"] = log
else:
out.pop("log", None)
@@ -227,30 +294,53 @@ def get_job(job_id: str) -> Optional[Dict]:
if not os.path.isfile(report_path) and not log:
return None
try:
import json
report = None
if os.path.isfile(report_path):
with open(report_path, encoding="utf-8") as f:
report = json.load(f)
summary = (report or {}).get("summary", {})
if report is None:
# Crashed before writing a report; the log is the only artifact.
status = "error"
else:
# Recover the job's real state: a job that stopped at the review gate
# must come back as needs_review (not done) or it can never finalize.
status = "needs_review" if summary.get("agent_status") == "needs_review" else "done"
# job.json (written at job start) carries the recipient email and
# pipeline mode so the final notification still fires after a restart.
# Missing/corrupt job.json degrades to the previous derivations.
meta: Dict = {}
meta_path = os.path.join(config.OUTPUT_DIR, job_id, "job.json")
try:
with open(meta_path, encoding="utf-8") as f:
loaded = json.load(f)
if isinstance(loaded, dict):
meta = loaded
except (OSError, json.JSONDecodeError):
pass
source_pdf = os.path.join(config.OUTPUT_DIR, job_id, "source.pdf")
status = "done" if report is not None else "error"
return {
"job_id": job_id,
"status": status,
"source": (report or {}).get("source", os.path.basename(report_path)),
"email": None,
job = {
"job_id": job_id,
"status": status,
"source": meta.get("source") or (report or {}).get("source", os.path.basename(report_path)),
"email": meta.get("email"),
"project_input": (report or {}).get("project_input", {}),
"text_local": (report or {}).get("summary", {}).get("text_backend") == "local",
"vision_model": None,
"text_model": None,
"stage": None,
"created_at": os.path.getmtime(source_pdf) if os.path.isfile(source_pdf) else None,
"finished_at": os.path.getmtime(report_path) if os.path.isfile(report_path) else None,
"report": report,
"error": None if report is not None else "Report missing; see job log",
"log": log,
"log_tail": log[-_LOG_TAIL:],
"text_local": summary.get("text_backend") == "local",
"pipeline_mode": meta.get("pipeline_mode") or summary.get("pipeline_mode", "classic"),
"vision_model": meta.get("vision_model"),
"text_model": meta.get("text_model"),
"stage": None,
"created_at": os.path.getmtime(source_pdf) if os.path.isfile(source_pdf) else None,
"finished_at": os.path.getmtime(report_path) if os.path.isfile(report_path) else None,
"report": report,
"error": None if report is not None else "Report missing; see job log",
"log": log,
"log_tail": log[-_LOG_TAIL:],
}
# Hydrate the in-memory registry so _set(...) transitions (reviewing,
# finalizing, done) work for restart-recovered jobs.
with _lock:
return dict(_jobs.setdefault(job_id, job))
except Exception as e:
print(f"[Jobs] Failed to load job {job_id} from disk: {e}")
return None
+30 -13
View File
@@ -23,7 +23,11 @@ _clients: Dict[str, OpenAI] = {}
# call_json when routing a no-image (text) call. Module-global mirrors the
# set_stage/cost pattern (single-user tool).
_text_local = False
# Optional per-run model overrides from the UI (empty = use config defaults).
# Per-job model overrides (user picked models in the UI). Same module-global
# pattern: set by the job runner before the pipeline starts, cleared after.
# Vision applies to image calls, text to no-image calls on OpenRouter (and to
# the local->cloud fallback). The LOCAL endpoint's model name is never taken
# from these overrides - hybrid local keeps LOCAL_TEXT_MODEL.
_vision_model_override: Optional[str] = None
_text_model_override: Optional[str] = None
@@ -35,7 +39,7 @@ def set_text_backend(local: bool) -> None:
def set_model_overrides(vision: Optional[str] = None, text: Optional[str] = None) -> None:
"""Per-run OpenRouter model picks. None/blank clears back to config defaults."""
"""Per-run OpenRouter vision/text model picks. None/blank clears to defaults."""
global _vision_model_override, _text_model_override
_vision_model_override = (vision or "").strip() or None
_text_model_override = (text or "").strip() or None
@@ -177,20 +181,22 @@ def _resolve_backend(has_images: bool, model_override: Optional[str]) -> Dict[st
return {
"base_url": config.LOCAL_BASE_URL,
"api_key": config.LOCAL_API_KEY,
"model": (model_override or _text_model_override
or config.LOCAL_TEXT_MODEL or config.TEXT_MODEL),
# Local model name comes from per-call args or LOCAL_TEXT_MODEL —
# never the UI's OpenRouter picks, which a local server won't serve.
"model": model_override or config.LOCAL_TEXT_MODEL or config.TEXT_MODEL,
"usage": False, # local has no OpenRouter usage accounting
"local": True,
}
# Vision, or text-on-OpenRouter (default / fallback).
# Vision, or text-on-OpenRouter (default / fallback). A per-job override
# (user's UI model pick) wins over per-call and env defaults.
if has_images:
default_model = _vision_model_override or config.MODEL
model = _vision_model_override or model_override or config.MODEL
else:
default_model = _text_model_override or config.TEXT_MODEL
model = _text_model_override or model_override or config.TEXT_MODEL
return {
"base_url": config.AI_BASE_URL,
"api_key": config.AI_API_KEY,
"model": model_override or default_model,
"model": model,
"usage": True,
"local": False,
}
@@ -251,18 +257,17 @@ def _repair_truncated(raw: str) -> Optional[Dict[str, Any]]:
return None
def _record_cost(response) -> None:
def _response_cost(response) -> Optional[float]:
"""Pull OpenRouter's per-call USD cost out of the usage object, if present."""
try:
dump = response.model_dump()
except Exception:
return
return None
usage = dump.get("usage") or {}
cost = usage.get("cost")
if cost is None:
cost = (usage.get("cost_details") or {}).get("upstream_inference_cost")
if isinstance(cost, (int, float)):
_add_cost(float(cost))
return float(cost) if isinstance(cost, (int, float)) else None
def _parse(raw: str) -> Optional[Dict[str, Any]]:
@@ -282,6 +287,8 @@ def call_json(
images_b64: Optional[List[str]] = None,
max_tokens: int = 4096,
model: Optional[str] = None,
usage_tracker: Optional[Any] = None,
usage_stage: str = "?",
) -> Optional[Dict[str, Any]]:
"""
Send one chat completion expecting a JSON object back.
@@ -301,6 +308,10 @@ def call_json(
hit = _cache_get(cache_key)
if hit is not None:
_add_cached()
if usage_tracker:
usage_tracker.record(
usage_stage, be["model"], cached=True, has_images=has_images
)
return hit
content: List[Dict[str, Any]] = []
@@ -328,7 +339,13 @@ def call_json(
if use_json_mode:
kwargs["response_format"] = {"type": "json_object"}
response = client.chat.completions.create(**kwargs)
_record_cost(response)
usd = _response_cost(response)
if usd is not None:
_add_cost(usd)
if usage_tracker:
usage_tracker.record(
usage_stage, be["model"], usd=usd or 0.0, has_images=has_images
)
raw = _strip_fences(response.choices[0].message.content or "")
parsed = _parse(raw)
if parsed is not None:
+148 -24
View File
@@ -11,16 +11,21 @@ ever needs concurrency.
import os
import tempfile
import threading
import time
from typing import Optional
from fastapi import FastAPI, UploadFile, File, Form, HTTPException
from fastapi.responses import HTMLResponse, JSONResponse, Response, PlainTextResponse
from fastapi.responses import HTMLResponse, JSONResponse, Response
from fastapi.staticfiles import StaticFiles
import backend.jobs
from backend import config
from backend.jobs import create_job, get_job, get_job_log
from backend.models_catalog import list_models
from backend.jobs import PIPELINE_MODES, create_job, get_job, _set
from backend.pipeline.pdf_processor import render_page_jpeg
from backend.review.feedback import decision_to_label, write_label
from backend.review.finalizer import finalize_review
from backend.review.store import ReviewStore
app = FastAPI(title=config.APP_TITLE, version=config.APP_VERSION)
@@ -31,14 +36,33 @@ _FRONTEND_DIR = os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", "
def health():
return {"status": "ok", "model": config.MODEL,
"text_model": config.TEXT_MODEL,
"version": config.APP_VERSION,
"build": config.APP_BUILD,
"key_configured": bool(config.AI_API_KEY),
"email_configured": bool(config.SMTP_HOST and config.SMTP_USER and config.SMTP_PASSWORD)}
@app.get("/models")
def models():
"""Vision vs text OpenRouter model lists for the UI dropdowns."""
return JSONResponse(list_models())
def list_models():
"""Vision/text OpenRouter model lists with pricing for the UI dropdowns."""
from backend.models import fetch_models, split_vision_text
models = fetch_models()
if models is None:
raise HTTPException(status_code=502,
detail="Could not fetch the model list from OpenRouter")
vision, text = split_vision_text(models)
return {"vision": vision, "text": text,
"defaults": {"vision": config.MODEL, "text": config.TEXT_MODEL}}
@app.get("/jobs/{job_id}/log")
def job_log(job_id: str):
"""The full captured stdout/stderr log of a job run (persists on disk)."""
path = os.path.join(config.OUTPUT_DIR, job_id, "job.log")
if not os.path.isfile(path):
raise HTTPException(status_code=404, detail="Log not found for this job")
with open(path, encoding="utf-8", errors="replace") as f:
return Response(content=f.read(), media_type="text/plain")
@app.post("/check")
@@ -50,6 +74,7 @@ async def check(
occupancy: Optional[str] = Form(None),
work_type: Optional[str] = Form(None),
text_local: bool = Form(False),
pipeline_mode: str = Form("classic"),
vision_model: Optional[str] = Form(None),
text_model: Optional[str] = Form(None),
):
@@ -66,6 +91,12 @@ async def check(
"""
if not file.filename.lower().endswith(".pdf"):
raise HTTPException(status_code=400, detail="Please upload a PDF.")
pipeline_mode = pipeline_mode.strip().lower()
if pipeline_mode not in PIPELINE_MODES:
raise HTTPException(
status_code=400,
detail=f"pipeline_mode must be one of: {', '.join(sorted(PIPELINE_MODES))}",
)
os.makedirs(config.UPLOAD_DIR, exist_ok=True)
suffix = "_" + os.path.basename(file.filename)
@@ -84,16 +115,16 @@ async def check(
}
v_model = (vision_model or "").strip() or None
t_model = (text_model or "").strip() or None
job_id = create_job(
tmp_path,
source_filename=file.filename,
email=email,
project_input=project_input,
text_local=text_local,
vision_model=v_model,
text_model=t_model,
)
return JSONResponse({"job_id": job_id, "status": "queued", "email": email})
job_id = create_job(tmp_path, source_filename=file.filename, email=email,
project_input=project_input, text_local=text_local,
pipeline_mode=pipeline_mode, vision_model=v_model,
text_model=t_model)
return JSONResponse({
"job_id": job_id,
"status": "queued",
"email": email,
"pipeline_mode": pipeline_mode,
})
@app.get("/jobs/{job_id}")
@@ -104,15 +135,108 @@ def job_status(job_id: str):
return JSONResponse(job)
@app.get("/jobs/{job_id}/log")
def job_log(job_id: str, plain: bool = False):
"""Full captured run log (also on disk as outputs/<job_id>/job.log)."""
lines = get_job_log(job_id)
if lines is None:
@app.get("/jobs/{job_id}/review")
def review_queue(job_id: str):
job = get_job(job_id)
if not job:
raise HTTPException(status_code=404, detail="Job not found")
if plain:
return PlainTextResponse("\n".join(lines) + ("\n" if lines else ""))
return JSONResponse({"job_id": job_id, "lines": lines, "text": "\n".join(lines)})
out_dir = job.get("out_dir") or os.path.join(config.OUTPUT_DIR, job_id)
# Read-only endpoint: don't create review/ dirs just by looking at them
# (readers already degrade to empty on missing files).
store = ReviewStore(out_dir, create=False)
queue = store.read_queue()
return {"queue": queue, "progress": store.progress(queue),
"decisions": store.read_decisions()}
@app.post("/jobs/{job_id}/review-decisions")
def save_review_decisions(job_id: str, payload: dict):
job = get_job(job_id)
if not job:
raise HTTPException(status_code=404, detail="Job not found")
if job.get("status") not in ("needs_review", "reviewing"):
# Positive state guard, mirroring the finalize endpoint: only jobs
# sitting at (or working through) the review gate accept decisions.
raise HTTPException(status_code=409, detail={
"detail": f"cannot save review decisions for a job in status {job.get('status')}",
})
out_dir = job.get("out_dir") or os.path.join(config.OUTPUT_DIR, job_id)
store = ReviewStore(out_dir)
queue = store.read_queue()
items_by_id = {item.get("review_item_id"): item for item in queue}
saved = 0
try:
for decision in payload.get("decisions") or []:
store.append_decision(decision)
queue_item = items_by_id.get(decision.get("review_item_id"))
if queue_item is not None:
write_label(out_dir, decision_to_label(queue_item, decision, job))
saved += 1
except ValueError as e:
raise HTTPException(status_code=422, detail=str(e))
progress = store.progress(queue)
if job.get("status") == "needs_review" and saved > 0 and progress["remaining"] > 0:
try:
_set(job_id, status="reviewing")
except KeyError:
pass # job not in the in-memory registry (e.g. loaded from disk)
return {"progress": progress}
def _finalize_job(job_id: str, out_dir: str) -> None:
"""Background finalization: the ONE place the final report email may fire."""
try:
report = finalize_review(job_id, out_dir)
except Exception as e:
try:
_set(job_id, status="finalization_error", error=str(e),
finished_at=time.time(), stage=None)
except KeyError:
pass # job not in the in-memory registry
return
try:
_set(job_id, status="done", report=report,
finished_at=time.time(), stage=None)
except KeyError:
pass
try:
backend.jobs._notify(job_id, report, out_dir)
except Exception as e:
print(f"[Jobs] Final notification for {job_id} failed: {e}")
@app.post("/jobs/{job_id}/finalize-review")
def finalize_review_endpoint(job_id: str):
job = get_job(job_id)
if not job:
raise HTTPException(status_code=404, detail="Job not found")
if job.get("status") in ("done", "finalizing"):
raise HTTPException(status_code=409, detail={
"detail": f"job is already {job['status']}",
})
if job.get("status") not in ("needs_review", "reviewing", "finalization_error"):
# Positive state-machine guard: finalization (and the final email) is
# only reachable after the job has passed through the review gate.
raise HTTPException(status_code=409, detail={
"detail": f"cannot finalize a job in status {job.get('status')}",
})
out_dir = job.get("out_dir") or os.path.join(config.OUTPUT_DIR, job_id)
store = ReviewStore(out_dir)
queue = store.read_queue()
decisions = store.read_decisions()
if any(item.get("blocking") and item.get("review_item_id") not in decisions
for item in queue):
# 409 detail shape: {"detail": <message>, "progress": <store.progress()>}
raise HTTPException(status_code=409, detail={
"detail": "incomplete review",
"progress": store.progress(queue),
})
try:
_set(job_id, status="finalizing")
except KeyError:
pass # job not in the in-memory registry (e.g. loaded from disk)
threading.Thread(target=_finalize_job, args=(job_id, out_dir), daemon=True).start()
return {"status": "finalizing"}
@app.get("/jobs/{job_id}/sheet-image/{page}")
+90
View File
@@ -0,0 +1,90 @@
"""
models.py - Fetch the available OpenRouter model list with pricing (cached).
The /models endpoint is public (no API key needed). Results are normalized to
per-1M-token USD costs for display and cached in memory for an hour; callers
degrade gracefully when OpenRouter is unreachable. Each entry also carries a
vision flag (accepts image input) so the UI can offer separate vision/text
model dropdowns.
"""
import time
from typing import List, Optional, Tuple
import httpx
from backend import config
_CACHE_TTL_SECONDS = 3600
_cache = {"at": 0.0, "models": None}
def _per_mtok(rate) -> float:
"""OpenRouter pricing is USD per token (as a string); display is per 1M."""
try:
return round(float(rate) * 1_000_000, 4)
except (TypeError, ValueError):
return 0.0
def _is_vision(item: dict) -> bool:
"""True when the model accepts image input and produces text output."""
arch = item.get("architecture") or {}
inputs = arch.get("input_modalities") or []
outputs = arch.get("output_modalities") or []
# Legacy string form: "text+image->text"
modality = (arch.get("modality") or "").lower()
has_image_in = ("image" in inputs) or ("image" in modality.split("->")[0])
has_text_out = ("text" in outputs) or ("->text" in modality) or (not outputs and not modality)
return has_image_in and has_text_out
def _fetch_openrouter_models() -> Optional[List[dict]]:
"""Raw GET of the OpenRouter model list; None on any failure."""
try:
response = httpx.get(f"{config.AI_BASE_URL.rstrip('/')}/models", timeout=10)
response.raise_for_status()
data = response.json().get("data")
return data if isinstance(data, list) else None
except Exception as e:
print(f"[Models] OpenRouter /models fetch failed: {e}")
return None
def fetch_models(force: bool = False) -> Optional[List[dict]]:
"""Normalized model list for the UI picker, or None when unavailable."""
if (
not force
and _cache["models"] is not None
and time.time() - _cache["at"] < _CACHE_TTL_SECONDS
):
return _cache["models"]
data = _fetch_openrouter_models()
if data is None:
return None
models = [
{
"id": item.get("id") or "",
"name": item.get("name") or item.get("id") or "",
"prompt_usd_per_mtok": _per_mtok((item.get("pricing") or {}).get("prompt")),
"completion_usd_per_mtok": _per_mtok((item.get("pricing") or {}).get("completion")),
"context_length": item.get("context_length"),
"vision": _is_vision(item),
}
for item in data
if item.get("id")
]
models.sort(key=lambda m: m["id"])
_cache["models"] = models
_cache["at"] = time.time()
return models
def split_vision_text(models: List[dict]) -> Tuple[List[dict], List[dict]]:
"""Partition the normalized catalog into (vision, text) lists for the UI.
Every catalog model takes text in/out, so vision models appear in both
lists (same dicts, pricing included).
"""
vision = [m for m in models if m.get("vision")]
return vision, list(models)
-101
View File
@@ -1,101 +0,0 @@
"""
models_catalog.py - OpenRouter model list for the UI dropdowns.
Fetches https://openrouter.ai/api/v1/models (cached ~1h) and splits into:
- vision: accepts image input and returns text
- text: chat models that return text (may also be multimodal)
"""
import time
from typing import Any, Dict, List
from backend import config
_TTL_SEC = 3600
_cache: Dict[str, Any] = {"at": 0.0, "payload": None}
def _entry(m: Dict[str, Any]) -> Dict[str, str]:
return {
"id": m.get("id") or "",
"name": m.get("name") or m.get("id") or "",
}
def _ensure_default(items: List[Dict[str, str]], model_id: str) -> List[Dict[str, str]]:
"""Prepend the configured default if OpenRouter didn't return it."""
if not model_id:
return items
if any(x["id"] == model_id for x in items):
return items
return [{"id": model_id, "name": model_id}] + items
def _fetch_raw() -> List[Dict[str, Any]]:
import httpx # local import so the app can start without httpx in odd envs
headers = {"Accept": "application/json"}
if config.AI_API_KEY:
headers["Authorization"] = f"Bearer {config.AI_API_KEY}"
url = f"{config.AI_BASE_URL.rstrip('/')}/models"
# Ask for text-output chat models (includes multimodal). "all" is huge.
with httpx.Client(timeout=30.0) as client:
r = client.get(url, headers=headers, params={"output_modalities": "text"})
r.raise_for_status()
data = r.json()
return data.get("data") or []
def list_models() -> Dict[str, Any]:
"""Return {vision, text, defaults} for the frontend selects."""
now = time.time()
if _cache["payload"] and (now - _cache["at"]) < _TTL_SEC:
return _cache["payload"]
try:
raw = _fetch_raw()
except Exception as e:
# Degrade to configured defaults so the UI still works offline.
print(f"[Models] OpenRouter catalog fetch failed: {e}")
vision = _ensure_default([], config.MODEL)
text = _ensure_default([], config.TEXT_MODEL)
payload = {
"vision": vision,
"text": text,
"defaults": {"vision": config.MODEL, "text": config.TEXT_MODEL},
"error": str(e),
}
return payload
vision: List[Dict[str, str]] = []
text: List[Dict[str, str]] = []
for m in raw:
mid = m.get("id") or ""
if not mid:
continue
arch = m.get("architecture") or {}
inputs = arch.get("input_modalities") or []
outputs = arch.get("output_modalities") or []
# Legacy string form: "text+image->text"
modality = (arch.get("modality") or "").lower()
has_image_in = ("image" in inputs) or ("image" in modality.split("->")[0])
has_text_out = ("text" in outputs) or ("->text" in modality) or (not outputs and not modality)
has_text_in = ("text" in inputs) or ("text" in modality) or not inputs
if has_image_in and has_text_out:
vision.append(_entry(m))
if has_text_in and has_text_out:
text.append(_entry(m))
vision.sort(key=lambda x: x["name"].lower())
text.sort(key=lambda x: x["name"].lower())
vision = _ensure_default(vision, config.MODEL)
text = _ensure_default(text, config.TEXT_MODEL)
payload = {
"vision": vision,
"text": text,
"defaults": {"vision": config.MODEL, "text": config.TEXT_MODEL},
}
_cache["at"] = now
_cache["payload"] = payload
return payload
+1
View File
@@ -176,6 +176,7 @@ def _run_stages(
report["summary"]["cost_by_stage"] = cost.get("by_stage", {})
report["summary"]["text_backend"] = "local" if text_local else "openrouter"
report["summary"]["models_used"] = cost.get("models", {})
report["summary"]["pipeline_mode"] = "classic"
print(f"[Runner] LLM cost: ${cost['usd']:.4f} over {cost['calls']} live calls"
f" ({cost.get('cached', 0)} cached)")
+1
View File
@@ -0,0 +1 @@
"""Human-review gate: decision schemas and review-trigger policy."""
+51
View File
@@ -0,0 +1,51 @@
"""Feedback labels: one label artifact per human-review decision, for metrics."""
import json
import os
from datetime import datetime, timezone
def _as_dict(value) -> dict:
return value if isinstance(value, dict) else {}
def decision_to_label(queue_item: dict, decision: dict, job: dict) -> dict:
"""Build one feedback label from a queue item, its decision, and the job.
All field access is defensive: missing fields degrade to None (or [] for
models_used) rather than raising.
"""
queue_item = _as_dict(queue_item)
decision = _as_dict(decision)
job = _as_dict(job)
payload = _as_dict(queue_item.get("payload"))
summary = _as_dict(_as_dict(job.get("report")).get("summary"))
return {
"review_item_id": queue_item.get("review_item_id"),
"job_id": job.get("job_id"),
"pipeline_mode": job.get("pipeline_mode"),
"source_stage": payload.get("source_stage"),
"category": payload.get("category"),
"severity": payload.get("severity"),
"confidence": payload.get("confidence"),
"decision": decision.get("decision"),
"reason_code": decision.get("reason_code"),
"location": payload.get("location"),
"disciplines": payload.get("disciplines"),
"sheets": payload.get("sheets"),
"drawing_type": payload.get("drawing_type"),
"models_used": summary.get("models_used") or [],
"created_at": datetime.now(timezone.utc).isoformat(),
}
def write_label(out_dir: str, label: dict) -> None:
"""Append one label as a JSON line; never raises on I/O failure."""
try:
review_dir = os.path.join(out_dir, "review")
os.makedirs(review_dir, exist_ok=True)
path = os.path.join(review_dir, "feedback_labels.jsonl")
with open(path, "a", encoding="utf-8") as f:
f.write(json.dumps(label) + "\n")
except OSError as e:
print(f"[Review] feedback label write failed: {e}")
+255
View File
@@ -0,0 +1,255 @@
"""ReviewFinalizer: apply human decisions, targeted reruns, RFIs, final artifacts.
All LLM-touching helpers degrade gracefully: a failed or empty targeted rerun
becomes a visible ``analysis_gap`` finding instead of raising, and RFI drafting
returns whatever was produced (possibly []). Finalization never crashes the job
on a single bad scope.
"""
import json
import os
from typing import Dict, List, Optional, Tuple
from backend import config
from backend.agents.base import AgentScope, AgentUsage
from backend.agents.conflict_critic import ConflictCriticAgent
from backend.agents.memory import ProjectMemory
from backend.agents.orchestrator import Orchestrator
from backend.agents.rfi_writer import RFIWriterAgent
# Private import, acceptable here: the runner's _finding_as_conflict is the
# canonical finding -> report["conflicts"] mapping; reusing it keeps the
# finalized report's conflicts in exactly the shape build_report produces.
from backend.agents.runner import _finding_as_conflict
from backend.pipeline.report import to_markdown
from backend.review.store import ReviewStore
def apply_decisions(prioritized: List[dict], decisions: Dict[str, dict]) -> Tuple[List[dict], List[dict]]:
kept: List[dict] = []
suppressed: List[dict] = []
for issue in prioritized:
review_id = f"finding:{issue.get('issue_id')}"
decision = decisions.get(review_id) or {}
action = decision.get("decision")
if action == "reject":
suppressed.append({
**issue,
"review_state": "rejected",
"reason_code": decision.get("reason_code"),
"review_comment": decision.get("comment") or "",
})
elif action == "unsure":
kept.append({**issue, "review_state": "unsure"})
else:
kept.append({**issue, "review_state": "confirmed" if action == "confirm" else "unreviewed"})
return kept, suppressed
def _gap_finding(index: int, scope_id: str, description: str) -> dict:
"""Same shape as the runner's gap_findings: low severity, high confidence."""
return {
"issue_id": f"AGENT-GAP-CLARIFY-{index + 1:03d}",
"source_stage": "qaqc",
"category": "analysis_gap",
"severity": "low",
"confidence": "high",
"location": scope_id.split(":", 2)[1] if ":" in scope_id else "",
"disciplines": [],
"sheets": [],
"description": description,
"evidence": [],
"recommended_resolution": "Review this scope manually or rerun the job.",
"code_reference": None,
"agent": "completeness",
"scope_id": scope_id,
}
def rerun_clarified_scopes(
memory_snapshot: dict,
decisions: Dict[str, dict],
prioritized: Optional[List[dict]] = None,
) -> List[dict]:
"""Bounded targeted reruns for ``needs_clarification`` decisions.
v1 reruns conflict scopes only: at most ONE ConflictCriticAgent scope per
clarified finding. The clarification answer is injected as a pseudo
"Reviewer" assertion prepended to the cluster's assertions so it reaches
the critic's evidence block (and survives front-truncation to
AGENT_CLUSTER_MAX_ASSERTIONS); ``page_to_b64`` is empty (cluster
assertions may carry their own
base64). Every per-scope failure degrades to an ``analysis_gap`` finding
and never raises. Non-conflict scopes are NOT rerun; they produce an
``analysis_gap`` noting the scope is not rerunnable in v1.
"""
findings_pool = list(prioritized or []) + list(memory_snapshot.get("findings") or [])
clusters = memory_snapshot.get("clusters") or []
out: List[dict] = []
for item_id, decision in (decisions or {}).items():
if (decision or {}).get("decision") != "needs_clarification":
continue
answer = str(decision.get("clarification_answer") or "").strip()
if not answer:
continue
issue_id = item_id.split(":", 1)[1] if item_id.startswith("finding:") else item_id
finding = next((f for f in findings_pool if f.get("issue_id") == issue_id), None)
scope_id = str((finding or {}).get("scope_id") or "")
if not scope_id.startswith("conflict:"):
out.append(_gap_finding(
len(out), scope_id or item_id,
f"Clarification rerun not supported in v1 for non-conflict scope "
f"{scope_id or item_id!r} (finding {issue_id}).",
))
continue
cluster_key = scope_id.split(":", 1)[1]
cluster = next((c for c in clusters if c.get("key") == cluster_key), None)
if cluster is None:
out.append(_gap_finding(
len(out), scope_id,
f"Clarification rerun failed: cluster {cluster_key!r} not found "
f"for scope {scope_id} (finding {issue_id}).",
))
continue
rerun_cluster = {
**cluster,
# Prepend: ConflictCriticAgent truncates assertions from the front
# (AGENT_CLUSTER_MAX_ASSERTIONS), so the clarification must come
# first or a full cluster would silently drop it.
"assertions": [{
"discipline": "Reviewer",
"sheet_number": "REVIEW",
"attribute": "clarification",
"value": answer,
"source_text": answer,
}] + list(cluster.get("assertions") or []),
}
scope = AgentScope(
scope_id=scope_id,
payload={"cluster": rerun_cluster, "page_to_b64": {}},
)
result = ConflictCriticAgent(AgentUsage()).run(scope)
if result.error or not result.artifacts:
out.append(_gap_finding(
len(out), scope_id,
f"Clarification rerun did not complete for scope {scope_id} "
f"(finding {issue_id}): {result.error or 'no findings produced'}.",
))
continue
for rerun_finding in result.artifacts:
rerun_finding["clarification_of"] = issue_id
out.append(rerun_finding)
return out
def _draft_rfis(kept: List[dict]) -> List[dict]:
"""Draft RFIs for kept issues only, mirroring the runner's wave 7."""
orchestrator = Orchestrator(ProjectMemory())
scopes = [
AgentScope(
scope_id=f"rfi:{finding.get('issue_id') or index + 1}",
payload={"finding": finding},
)
for index, finding in enumerate(kept)
]
try:
results = orchestrator.run_scopes(
RFIWriterAgent(AgentUsage()), scopes, config.AGENT_RFI_CONCURRENCY
)
except Exception:
return []
return [artifact for result in results for artifact in result.artifacts]
def _read_json(path: str, default):
try:
with open(path, encoding="utf-8") as f:
return json.load(f)
except (OSError, json.JSONDecodeError):
return default
def _dump(out_dir: str, name: str, value) -> None:
with open(os.path.join(out_dir, name), "w", encoding="utf-8") as f:
json.dump(value, f, indent=2)
def finalize_review(job_id: str, out_dir: str) -> dict:
"""Apply review decisions and write the final report artifacts.
Raises ValueError("incomplete review") if any blocking queue item lacks a
decision. Never raises for rerun/RFI degradation.
"""
store = ReviewStore(out_dir)
queue = store.read_queue()
decisions = store.read_decisions()
for item in queue:
if item.get("blocking") and item.get("review_item_id") not in decisions:
raise ValueError("incomplete review")
report = _read_json(os.path.join(out_dir, "conflicts.json"), {}) or {}
snapshot = _read_json(os.path.join(out_dir, "agent", "memory.json"), {}) or {}
prioritized = list(report.get("validated_issues") or [])
rerun_findings = rerun_clarified_scopes(snapshot, decisions, prioritized)
replacements: Dict[str, List[dict]] = {}
for finding in rerun_findings:
origin = finding.get("clarification_of")
if origin:
replacements.setdefault(origin, []).append(finding)
else:
prioritized.append(finding) # gap findings stay as additions
for origin, new_findings in replacements.items():
for index, issue in enumerate(prioritized):
if issue.get("issue_id") == origin:
prioritized[index:index + 1] = new_findings
break
kept, suppressed = apply_decisions(prioritized, decisions)
for issue in kept:
if issue.get("clarification_of"):
issue["review_state"] = "clarified"
else:
decision = decisions.get(f"finding:{issue.get('issue_id')}") or {}
if decision.get("decision") == "needs_clarification":
issue["review_state"] = "clarification_failed"
rfis = _draft_rfis(kept)
report["validated_issues"] = kept
report["suppressed_issues"] = suppressed
report["rfis"] = rfis
summary = report.setdefault("summary", {})
summary["agent_status"] = "complete"
by_stage = summary.get("by_stage")
if isinstance(by_stage, dict):
if "validated" in by_stage:
by_stage["validated"] = len(kept)
if "rfis" in by_stage:
by_stage["rfis"] = len(rfis)
# Rebuild the conflicts view and headline counts from the KEPT
# conflict-stage findings so rejected findings no longer appear as
# conflicts in report.md / the UI (mirrors pipeline.report.build_report).
conflicts = [
_finding_as_conflict(finding)
for finding in kept
if finding.get("source_stage") == "conflict"
]
report["conflicts"] = conflicts
by_severity = {"high": 0, "medium": 0, "low": 0}
by_category: Dict[str, int] = {}
for conflict in conflicts:
by_severity[conflict["severity"]] = by_severity.get(conflict["severity"], 0) + 1
by_category[conflict["category"]] = by_category.get(conflict["category"], 0) + 1
summary["conflicts_found"] = len(conflicts)
summary["by_severity"] = by_severity
summary["by_category"] = by_category
summary["review"] = store.progress(queue)
os.makedirs(out_dir, exist_ok=True)
_dump(out_dir, "conflicts.json", report)
_dump(out_dir, "validated_issues.json", kept)
_dump(out_dir, "suppressed_issues.json", suppressed)
_dump(out_dir, "rfis.json", rfis)
with open(os.path.join(out_dir, "report.md"), "w", encoding="utf-8") as f:
f.write(to_markdown(report))
return report
+31
View File
@@ -0,0 +1,31 @@
"""ReviewGate: build the human-review queue from prioritized findings."""
from typing import Dict, List, Optional
from backend.review.policy import build_audit_sample, requires_review
def _finding_item(issue: Dict, blocking: bool, reasons: List[str], kind: str) -> Dict:
issue_id = issue.get("issue_id") or "unknown"
return {
"review_item_id": f"finding:{issue_id}",
"kind": kind,
"blocking": blocking,
"reasons": reasons,
"payload": issue,
}
def build_review_queue(memory_snapshot: Dict, prioritized: List[Dict], decisions: List[Dict],
limit: Optional[int] = None) -> List[Dict]:
queue: List[Dict] = []
for issue in prioritized:
reasons = requires_review(issue)
queue.append(_finding_item(issue, bool(reasons), reasons, "finding" if reasons else "audit_finding"))
if limit is None:
for item in build_audit_sample(memory_snapshot, prioritized):
queue.append(item)
else:
for item in build_audit_sample(memory_snapshot, prioritized, limit=limit):
queue.append(item)
return queue
+25
View File
@@ -0,0 +1,25 @@
"""Aggregate metrics over feedback labels.
Default aggregates exclude source_text, images, raw sheet content, and
reviewer free-text comments; include_text=True is the only path that embeds
the raw labels.
"""
from collections import Counter
from typing import Dict, List
def aggregate_labels(labels: List[dict], include_text: bool = False) -> Dict:
decisions = Counter(label.get("decision") or "unknown" for label in labels)
reasons = Counter(label.get("reason_code") or "none" for label in labels if label.get("decision") == "reject")
summary = {
"total": len(labels),
"decisions": dict(decisions),
"reject_reasons": dict(reasons),
}
for label in labels:
decision = label.get("decision") or "unknown"
summary[decision] = summary.get(decision, 0) + 1
if include_text:
summary["labels"] = labels
return summary
+70
View File
@@ -0,0 +1,70 @@
"""Review-trigger policy: which findings block on human review."""
from typing import Any, Dict, List
_SENSITIVE_CATEGORIES = {
"missing_element",
"ada",
"tas_tdlr",
"egress",
"fire_separation",
"occupancy",
"spatial_clash",
"clearance_conflict",
"penetration_conflict",
}
def requires_review(issue: Dict) -> List[str]:
"""Return trigger reasons that require human review for one issue."""
reasons: List[str] = []
severity = str(issue.get("severity") or "").lower()
confidence = str(issue.get("confidence") or "").lower()
category = str(issue.get("category") or "").lower()
if severity in {"critical", "high"}:
reasons.append("severity_high")
if confidence == "low":
reasons.append("confidence_low")
if category in _SENSITIVE_CATEGORIES or issue.get("source_stage") == "code":
reasons.append("sensitive_category")
return reasons
def build_audit_sample(
memory_snapshot: Dict,
prioritized: List[Dict],
limit: int = 5,
) -> List[Dict[str, Any]]:
"""Build non-blocking spot-check items for clean (finding-free) clusters."""
implicated = {
str(finding.get("scope_id") or "")
for finding in (memory_snapshot.get("findings") or []) + list(prioritized)
}
items: List[Dict[str, Any]] = []
for cluster in memory_snapshot.get("clusters") or []:
if len(items) >= limit:
break
assertions = cluster.get("assertions") or []
if len(assertions) < 2:
continue
cluster_key = cluster.get("key") or "unknown"
if any(cluster_key in scope_id for scope_id in implicated):
continue
items.append({
"review_item_id": f"clean_cluster:{cluster_key}",
"kind": "clean_cluster",
"blocking": False,
"reasons": ["audit_sample"],
"payload": _without_base64(cluster),
})
return items
def _without_base64(cluster: Dict) -> Dict:
return {
**cluster,
"assertions": [
{key: value for key, value in assertion.items() if key != "base64"}
for assertion in cluster.get("assertions") or []
],
}
+45
View File
@@ -0,0 +1,45 @@
"""Human-review decision schema and validation."""
from typing import Optional
DECISIONS = {"confirm", "reject", "unsure", "needs_clarification"}
REASON_CODES = {
"wrong_cluster_link",
"same_value_different_representation",
"not_a_contradiction",
"missing_evidence",
"extraction_misread",
"code_path_not_applicable",
"duplicate",
"severity_too_high",
"severity_too_low",
"other",
}
def validate_decision(raw: dict) -> Optional[dict]:
"""Normalize a reviewer decision payload, or return None if invalid."""
if not isinstance(raw, dict):
return None
decision = str(raw.get("decision") or "").strip()
if decision not in DECISIONS:
return None
reason_code = raw.get("reason_code")
if decision == "reject":
reason_code = str(reason_code or "").strip()
if reason_code not in REASON_CODES:
return None
elif reason_code is not None:
reason_code = str(reason_code).strip() or None
if reason_code and reason_code not in REASON_CODES:
return None
return {
"review_item_id": str(raw.get("review_item_id") or "").strip(),
"decision": decision,
"reason_code": reason_code,
"category_correction": raw.get("category_correction"),
"severity_correction": raw.get("severity_correction"),
"comment": str(raw.get("comment") or "").strip(),
"clarification_answer": raw.get("clarification_answer"),
"reviewed_at": raw.get("reviewed_at"),
}
+62
View File
@@ -0,0 +1,62 @@
"""Persistence for human-review queue and decisions within a job output dir."""
import json
import os
from typing import Dict, List
from backend.review.schemas import validate_decision
class ReviewStore:
def __init__(self, job_out_dir: str, create: bool = True) -> None:
self.review_dir = os.path.join(job_out_dir, "review")
if create:
os.makedirs(self.review_dir, exist_ok=True)
def _path(self, name: str) -> str:
return os.path.join(self.review_dir, name)
def _write_json(self, name: str, value) -> None:
path = self._path(name)
tmp = f"{path}.tmp"
with open(tmp, "w", encoding="utf-8") as f:
json.dump(value, f, indent=2)
os.replace(tmp, path)
def write_queue(self, queue: List[dict]) -> None:
self._write_json("review_queue.json", queue)
def read_queue(self) -> List[dict]:
try:
with open(self._path("review_queue.json"), encoding="utf-8") as f:
value = json.load(f)
return value if isinstance(value, list) else []
except (OSError, json.JSONDecodeError):
return []
def append_decision(self, decision: dict) -> None:
valid = validate_decision(decision)
if not valid or not valid["review_item_id"]:
raise ValueError("invalid review decision")
decisions = self.read_decisions()
decisions[valid["review_item_id"]] = valid
self._write_json("review_decisions.json", decisions)
def read_decisions(self) -> Dict[str, dict]:
try:
with open(self._path("review_decisions.json"), encoding="utf-8") as f:
value = json.load(f)
return value if isinstance(value, dict) else {}
except (OSError, json.JSONDecodeError):
return {}
def progress(self, queue: List[dict]) -> dict:
decisions = self.read_decisions()
required = [item for item in queue if item.get("blocking")]
completed = [item for item in required if item.get("review_item_id") in decisions]
return {
"required": len(required),
"completed": len(completed),
"remaining": len(required) - len(completed),
"total": len(queue),
}
+27 -2
View File
@@ -17,6 +17,8 @@ _ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
if _ROOT not in sys.path:
sys.path.insert(0, _ROOT)
from backend import config # noqa: E402
from backend.agents.runner import run_agent_pipeline # noqa: E402
from backend.pipeline.runner import run_pipeline # noqa: E402
@@ -25,11 +27,16 @@ def main() -> int:
parser.add_argument("pdf", help="Path to the PDF drawing set")
parser.add_argument("--out", default=None,
help="Directory for artifacts (default: out/<pdf-stem>)")
parser.add_argument("--mode", choices=("classic", "agent"), default="classic",
help="Pipeline implementation to run (default: classic)")
parser.add_argument("--project-name", default=None)
parser.add_argument("--address", default=None)
parser.add_argument("--occupancy", default=None)
parser.add_argument("--work-type", default=None,
help="new_building | remodel | tenant_improvement | addition | ...")
parser.add_argument("--no-review", action="store_true",
help="Agent mode only: skip the human-review gate and finish the run "
"(overrides AGENT_REQUIRE_REVIEW=true)")
args = parser.parse_args()
if not os.path.isfile(args.pdf):
@@ -43,7 +50,21 @@ def main() -> int:
}.items() if v
}
out_dir = args.out or os.path.join("out", os.path.splitext(os.path.basename(args.pdf))[0])
report = run_pipeline(args.pdf, out_dir=out_dir, project_input=project_input or None)
if args.mode == "agent":
report = run_agent_pipeline(
args.pdf,
out_dir=out_dir,
project_input=project_input or None,
source_name=os.path.basename(args.pdf),
require_review=config.AGENT_REQUIRE_REVIEW and not args.no_review,
)
else:
report = run_pipeline(
args.pdf,
out_dir=out_dir,
project_input=project_input or None,
source_name=os.path.basename(args.pdf),
)
s = report["summary"]
print("\n" + "=" * 60)
@@ -51,7 +72,11 @@ def main() -> int:
f"(high {s['by_severity']['high']}, "
f"medium {s['by_severity']['medium']}, "
f"low {s['by_severity']['low']})")
print(f" Report: {os.path.join(out_dir, 'report.md')}")
if s.get("agent_status") == "needs_review":
print(" Stopped for human review - finalize via the web UI, "
"or rerun with --no-review.")
else:
print(f" Report: {os.path.join(out_dir, 'report.md')}")
print("=" * 60)
return 0
+1 -1
View File
@@ -13,7 +13,7 @@ services:
env_file:
- backend/.env
environment:
APP_BASE_URL: ${APP_BASE_URL:-http://localhost:8099}
APP_BASE_URL: ${APP_BASE_URL:-https://conchecker.scoutitsystems.com}
volumes:
- uploads:/app/backend/uploads
- outputs:/app/backend/outputs
+1 -1
View File
@@ -7,7 +7,7 @@ services:
- backend/.env
environment:
# Override in backend/.env for production (email links, etc.)
APP_BASE_URL: ${APP_BASE_URL:-http://localhost:8099}
APP_BASE_URL: ${APP_BASE_URL:-https://conchecker.scoutitsystems.com}
volumes:
- uploads:/app/backend/uploads
- outputs:/app/backend/outputs
@@ -0,0 +1,867 @@
# Agent Human Review Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Add required human review to the Agent pipeline so findings are confirmed, rejected, clarified, and measured before final RFIs/reports are issued.
**Architecture:** Keep the existing Agent pipeline through Brain, then insert a ReviewGate that writes a persistent review queue and moves the job to `needs_review`. A ReviewFinalizer applies human decisions, performs bounded targeted reruns for clarification, drafts RFIs only for kept issues, and only then marks the job done and sends final email.
**Tech Stack:** Python 3, FastAPI, pytest, vanilla JS frontend, JSON file artifacts under `backend/outputs/<job_id>/`.
## Global Constraints
- Do not change Classic pipeline behavior.
- Agent mode remains OpenRouter-only in v1.
- No final email before human review finalization.
- No raw `source_text`, sheet images, or drawing content in aggregate metrics by default.
- All new review logic must have non-LLM tests.
- Follow existing patterns: small modules, graceful degradation, JSON artifacts under job output dir.
- Review endpoints are state-changing and must be treated as sensitive in docs and deployment notes.
---
### Task 1: Review schemas and policy
**Files:**
- Create: `backend/review/__init__.py`
- Create: `backend/review/schemas.py`
- Create: `backend/review/policy.py`
- Test: `tests/review/test_policy.py`
**Interfaces:**
- Consumes: nothing from earlier tasks.
- Produces:
- `DECISIONS = {"confirm", "reject", "unsure", "needs_clarification"}`
- `REASON_CODES = {"wrong_cluster_link", "same_value_different_representation", "not_a_contradiction", "missing_evidence", "extraction_misread", "code_path_not_applicable", "duplicate", "severity_too_high", "severity_too_low", "other"}`
- `validate_decision(raw: dict) -> dict | None`
- `requires_review(issue: dict) -> list[str]`
- `build_audit_sample(memory_snapshot: dict, prioritized: list[dict], limit: int = 5) -> list[dict]`
- [ ] **Step 1: Write failing policy tests**
```python
from backend.review.policy import requires_review
def test_high_severity_requires_review():
issue = {"severity": "high", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"}
assert "severity_high" in requires_review(issue)
def test_low_confidence_requires_review():
issue = {"severity": "low", "confidence": "low", "category": "note_or_spec_contradiction", "source_stage": "conflict"}
assert "confidence_low" in requires_review(issue)
def test_sensitive_code_category_requires_review():
issue = {"severity": "medium", "confidence": "high", "category": "egress", "source_stage": "code"}
assert "sensitive_category" in requires_review(issue)
def test_medium_high_confidence_note_does_not_require_review():
issue = {"severity": "medium", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"}
assert requires_review(issue) == []
```
- [ ] **Step 2: Run tests to verify they fail**
Run: `pytest tests/review/test_policy.py -v`
Expected: FAIL with `ModuleNotFoundError: No module named 'backend.review'`
- [ ] **Step 3: Implement schemas and policy**
```python
# backend/review/schemas.py
from typing import Optional
DECISIONS = {"confirm", "reject", "unsure", "needs_clarification"}
REASON_CODES = {
"wrong_cluster_link",
"same_value_different_representation",
"not_a_contradiction",
"missing_evidence",
"extraction_misread",
"code_path_not_applicable",
"duplicate",
"severity_too_high",
"severity_too_low",
"other",
}
def validate_decision(raw: dict) -> Optional[dict]:
if not isinstance(raw, dict):
return None
decision = str(raw.get("decision") or "").strip()
if decision not in DECISIONS:
return None
reason_code = raw.get("reason_code")
if decision == "reject":
reason_code = str(reason_code or "").strip()
if reason_code not in REASON_CODES:
return None
elif reason_code is not None:
reason_code = str(reason_code).strip() or None
if reason_code and reason_code not in REASON_CODES:
return None
return {
"review_item_id": str(raw.get("review_item_id") or "").strip(),
"decision": decision,
"reason_code": reason_code,
"category_correction": raw.get("category_correction"),
"severity_correction": raw.get("severity_correction"),
"comment": str(raw.get("comment") or "").strip(),
"clarification_answer": raw.get("clarification_answer"),
"reviewed_at": raw.get("reviewed_at"),
}
```
```python
# backend/review/policy.py
from typing import Dict, List
_SENSITIVE_CATEGORIES = {
"missing_element",
"ada",
"tas_tdlr",
"egress",
"fire_separation",
"occupancy",
"spatial_clash",
"clearance_conflict",
"penetration_conflict",
}
def requires_review(issue: Dict) -> List[str]:
reasons: List[str] = []
severity = str(issue.get("severity") or "").lower()
confidence = str(issue.get("confidence") or "").lower()
category = str(issue.get("category") or "").lower()
if severity in {"critical", "high"}:
reasons.append("severity_high")
if confidence == "low":
reasons.append("confidence_low")
if category in _SENSITIVE_CATEGORIES or issue.get("source_stage") == "code":
reasons.append("sensitive_category")
return reasons
```
- [ ] **Step 4: Run tests to verify they pass**
Run: `pytest tests/review/test_policy.py -v`
Expected: PASS
- [ ] **Step 5: Commit**
```bash
git add backend/review tests/review/test_policy.py
git commit -m "Add review decision schema and trigger policy"
```
---
### Task 2: Review persistence
**Files:**
- Create: `backend/review/store.py`
- Test: `tests/review/test_store.py`
**Interfaces:**
- Consumes: `validate_decision` from Task 1.
- Produces:
- `ReviewStore(job_out_dir: str)`
- `.write_queue(queue: list[dict]) -> None`
- `.read_queue() -> list[dict]`
- `.append_decision(decision: dict) -> None`
- `.read_decisions() -> dict[str, dict]`
- `.progress(queue: list[dict]) -> dict`
- [ ] **Step 1: Write failing persistence tests**
```python
import json
from backend.review.store import ReviewStore
def test_queue_and_decisions_round_trip(tmp_path):
store = ReviewStore(str(tmp_path))
queue = [{"review_item_id": "finding:1", "blocking": True}]
store.write_queue(queue)
assert store.read_queue() == queue
store.append_decision({"review_item_id": "finding:1", "decision": "confirm"})
assert store.read_decisions()["finding:1"]["decision"] == "confirm"
def test_progress_counts_required_items(tmp_path):
store = ReviewStore(str(tmp_path))
queue = [
{"review_item_id": "a", "blocking": True},
{"review_item_id": "b", "blocking": False},
]
store.write_queue(queue)
store.append_decision({"review_item_id": "a", "decision": "confirm"})
progress = store.progress(queue)
assert progress["required"] == 1
assert progress["completed"] == 1
```
- [ ] **Step 2: Run tests to verify they fail**
Run: `pytest tests/review/test_store.py -v`
Expected: FAIL with `ModuleNotFoundError: No module named 'backend.review.store'`
- [ ] **Step 3: Implement ReviewStore**
```python
import json
import os
from typing import Dict, List
from backend.review.schemas import validate_decision
class ReviewStore:
def __init__(self, job_out_dir: str) -> None:
self.review_dir = os.path.join(job_out_dir, "review")
os.makedirs(self.review_dir, exist_ok=True)
def _path(self, name: str) -> str:
return os.path.join(self.review_dir, name)
def _write_json(self, name: str, value) -> None:
path = self._path(name)
tmp = f"{path}.tmp"
with open(tmp, "w", encoding="utf-8") as f:
json.dump(value, f, indent=2)
os.replace(tmp, path)
def write_queue(self, queue: List[dict]) -> None:
self._write_json("review_queue.json", queue)
def read_queue(self) -> List[dict]:
try:
with open(self._path("review_queue.json"), encoding="utf-8") as f:
value = json.load(f)
return value if isinstance(value, list) else []
except (OSError, json.JSONDecodeError):
return []
def append_decision(self, decision: dict) -> None:
valid = validate_decision(decision)
if not valid or not valid["review_item_id"]:
raise ValueError("invalid review decision")
decisions = self.read_decisions()
decisions[valid["review_item_id"]] = valid
self._write_json("review_decisions.json", decisions)
def read_decisions(self) -> Dict[str, dict]:
try:
with open(self._path("review_decisions.json"), encoding="utf-8") as f:
value = json.load(f)
return value if isinstance(value, dict) else {}
except (OSError, json.JSONDecodeError):
return {}
def progress(self, queue: List[dict]) -> dict:
decisions = self.read_decisions()
required = [item for item in queue if item.get("blocking")]
completed = [item for item in required if item.get("review_item_id") in decisions]
return {
"required": len(required),
"completed": len(completed),
"remaining": len(required) - len(completed),
"total": len(queue),
}
```
- [ ] **Step 4: Run tests to verify they pass**
Run: `pytest tests/review/test_store.py -v`
Expected: PASS
- [ ] **Step 5: Commit**
```bash
git add backend/review/store.py tests/review/test_store.py
git commit -m "Add persistent review store"
```
---
### Task 3: ReviewGate queue builder
**Files:**
- Create: `backend/review/gate.py`
- Test: `tests/review/test_gate.py`
**Interfaces:**
- Consumes: `requires_review`, `build_audit_sample` from Task 1.
- Produces:
- `build_review_queue(memory_snapshot: dict, prioritized: list[dict], decisions: list[dict]) -> list[dict]`
- queue item shape: `{ "review_item_id": str, "kind": "finding|audit_finding|clean_cluster", "blocking": bool, "reasons": list[str], "payload": dict }`
- [ ] **Step 1: Write failing gate tests**
```python
from backend.review.gate import build_review_queue
def test_gate_marks_blocking_and_audit_items():
memory = {"clusters": [{"key": "room:101", "location": "Room 101", "assertions": [{"id": "a1"}, {"id": "a2"}]}], "findings": []}
prioritized = [
{"issue_id": "AGENT-0001", "severity": "high", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"},
{"issue_id": "AGENT-0002", "severity": "low", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"},
]
queue = build_review_queue(memory, prioritized, [])
by_id = {item["review_item_id"]: item for item in queue}
assert by_id["finding:AGENT-0001"]["blocking"] is True
assert by_id["finding:AGENT-0002"]["blocking"] is False
assert any(item["kind"] == "clean_cluster" for item in queue)
```
- [ ] **Step 2: Run tests to verify they fail**
Run: `pytest tests/review/test_gate.py -v`
Expected: FAIL with `ModuleNotFoundError: No module named 'backend.review.gate'`
- [ ] **Step 3: Implement ReviewGate**
```python
from typing import Dict, List
from backend.review.policy import build_audit_sample, requires_review
def _finding_item(issue: Dict, blocking: bool, reasons: List[str], kind: str) -> Dict:
issue_id = issue.get("issue_id") or "unknown"
return {
"review_item_id": f"finding:{issue_id}",
"kind": kind,
"blocking": blocking,
"reasons": reasons,
"payload": issue,
}
def build_review_queue(memory_snapshot: Dict, prioritized: List[Dict], decisions: List[Dict]) -> List[Dict]:
queue: List[Dict] = []
for issue in prioritized:
reasons = requires_review(issue)
queue.append(_finding_item(issue, bool(reasons), reasons, "finding" if reasons else "audit_finding"))
for item in build_audit_sample(memory_snapshot, prioritized):
queue.append(item)
return queue
```
- [ ] **Step 4: Run tests to verify they pass**
Run: `pytest tests/review/test_gate.py -v`
Expected: PASS
- [ ] **Step 5: Commit**
```bash
git add backend/review/gate.py tests/review/test_gate.py
git commit -m "Add review gate queue builder"
```
---
### Task 4: Agent runner stops after Brain
**Files:**
- Modify: `backend/agents/runner.py`
- Modify: `cli/run_check.py`
- Test: `tests/agents/test_runner_review_gate.py`
**Interfaces:**
- Consumes: `build_review_queue`, `ReviewStore`.
- Produces:
- `run_agent_pipeline(..., require_review: bool = True) -> dict`
- candidate report contains `summary.agent_status = "needs_review"` and `summary.review = {"required": int, "completed": 0, "blocking": int}` when review is required.
- [ ] **Step 1: Write failing runner gate test**
```python
from backend.agents.runner import run_agent_pipeline
def test_agent_runner_can_enter_review_mode(monkeypatch, tmp_path):
monkeypatch.setattr("backend.agents.runner.convert_pdf_to_images", lambda path: [{"page_number": 1, "base64": "x"}])
monkeypatch.setattr("backend.agents.runner.BrainAgent", lambda usage: type("B", (), {"run": lambda self, findings, sheet_index, jurisdiction: ([{"issue_id": "AGENT-0001", "severity": "high", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"}], [])})())
report = run_agent_pipeline("dummy.pdf", out_dir=str(tmp_path), require_review=True)
assert report["summary"]["agent_status"] == "needs_review"
assert report["summary"]["review"]["required"] == 1
```
- [ ] **Step 2: Run test to verify it fails**
Run: `pytest tests/agents/test_runner_review_gate.py -v`
Expected: FAIL because `require_review` is not a supported argument.
- [ ] **Step 3: Implement review-mode branch in runner**
```python
from backend.review.gate import build_review_queue
from backend.review.store import ReviewStore
def run_agent_pipeline(..., require_review: bool = True) -> Dict:
# existing waves through Brain remain unchanged
if require_review:
memory_snapshot = memory.snapshot()
queue = build_review_queue(memory_snapshot, prioritized, decisions)
store = ReviewStore(out_dir)
store.write_queue(queue)
candidate_conflicts = [_finding_as_conflict(item) for item in conflict_findings]
report = build_report(
conflicts=candidate_conflicts,
sheets=sheets,
clusters=clusters,
source=source_name or os.path.basename(pdf_path),
)
report.update({
"project_input": merged_input,
"jurisdiction": jurisdiction,
"sheet_index": sheet_index,
"project_intelligence": object_graph,
"validated_issues": prioritized,
"rfis": [],
"suppressed_issues": [],
})
progress = store.progress(queue)
report["summary"].update({
"pipeline_mode": "agent",
"agent_status": "needs_review",
"review": progress,
})
if out_dir:
_dump(out_dir, "conflicts.json", report)
_dump(out_dir, "validated_issues.json", prioritized)
return report
# existing RFI/report path remains for require_review=False
```
- [ ] **Step 4: Run test to verify it passes**
Run: `pytest tests/agents/test_runner_review_gate.py -v`
Expected: PASS
- [ ] **Step 5: Commit**
```bash
git add backend/agents/runner.py cli/run_check.py tests/agents/test_runner_review_gate.py
git commit -m "Gate agent runs behind required human review"
```
---
### Task 5: Job states and review API
**Files:**
- Modify: `backend/jobs.py`
- Modify: `backend/main.py`
- Test: `tests/api/test_review_api.py`
**Interfaces:**
- Consumes: `ReviewStore`, `validate_decision`.
- Produces:
- statuses: `needs_review`, `reviewing`, `finalizing`, `finalization_error`
- `GET /jobs/{job_id}/review -> {"queue": list[dict], "progress": dict}`
- `POST /jobs/{job_id}/review-decisions`
- [ ] **Step 1: Write failing API tests**
```python
from fastapi.testclient import TestClient
from backend.main import app
def test_review_queue_and_decision_save(monkeypatch, tmp_path):
client = TestClient(app)
monkeypatch.setattr("backend.main.get_job", lambda job_id: {"job_id": job_id, "status": "needs_review", "report": {"summary": {}}, "out_dir": str(tmp_path)})
queue_response = client.get("/jobs/job1/review")
assert queue_response.status_code == 200
decision_response = client.post("/jobs/job1/review-decisions", json={"decisions": [{"review_item_id": "finding:AGENT-0001", "decision": "confirm"}]})
assert decision_response.status_code == 200
```
- [ ] **Step 2: Run tests to verify they fail**
Run: `pytest tests/api/test_review_api.py -v`
Expected: FAIL with 404 because review endpoints do not exist.
- [ ] **Step 3: Implement job status and endpoints**
```python
# backend/main.py
from backend.review.store import ReviewStore
@app.get("/jobs/{job_id}/review")
def review_queue(job_id: str):
job = get_job(job_id)
if not job:
raise HTTPException(status_code=404, detail="Job not found")
out_dir = job.get("out_dir") or os.path.join(config.OUTPUT_DIR, job_id)
store = ReviewStore(out_dir)
queue = store.read_queue()
return {"queue": queue, "progress": store.progress(queue)}
@app.post("/jobs/{job_id}/review-decisions")
def save_review_decisions(job_id: str, payload: dict):
job = get_job(job_id)
if not job:
raise HTTPException(status_code=404, detail="Job not found")
out_dir = job.get("out_dir") or os.path.join(config.OUTPUT_DIR, job_id)
store = ReviewStore(out_dir)
for decision in payload.get("decisions") or []:
store.append_decision(decision)
return {"progress": store.progress(store.read_queue())}
```
- [ ] **Step 4: Run tests to verify they pass**
Run: `pytest tests/api/test_review_api.py -v`
Expected: PASS
- [ ] **Step 5: Commit**
```bash
git add backend/jobs.py backend/main.py tests/api/test_review_api.py
git commit -m "Add review job states and API endpoints"
```
---
### Task 6: Review finalizer and targeted rerun
**Files:**
- Create: `backend/review/finalizer.py`
- Modify: `backend/agents/runner.py`
- Modify: `backend/main.py`
- Test: `tests/review/test_finalizer.py`
**Interfaces:**
- Consumes: `ReviewStore`, queue items from Task 3, Agent runner helpers.
- Produces:
- `finalize_review(job_id: str, out_dir: str) -> dict`
- `apply_decisions(prioritized: list[dict], decisions: dict[str, dict]) -> tuple[list[dict], list[dict]]`
- `rerun_clarified_scopes(memory_snapshot: dict, decisions: dict[str, dict]) -> list[dict]`
- `POST /jobs/{job_id}/finalize-review` returns `409` until blocking decisions are complete
- [ ] **Step 1: Write failing finalizer tests**
```python
from backend.review.finalizer import apply_decisions
def test_reject_suppresses_with_reason():
prioritized = [{"issue_id": "AGENT-0001", "severity": "high"}]
decisions = {"finding:AGENT-0001": {"decision": "reject", "reason_code": "duplicate"}}
kept, suppressed = apply_decisions(prioritized, decisions)
assert kept == []
assert suppressed[0]["review_state"] == "rejected"
assert suppressed[0]["reason_code"] == "duplicate"
def test_unsure_is_kept_but_flagged():
prioritized = [{"issue_id": "AGENT-0002", "severity": "medium"}]
decisions = {"finding:AGENT-0002": {"decision": "unsure"}}
kept, suppressed = apply_decisions(prioritized, decisions)
assert kept[0]["review_state"] == "unsure"
assert suppressed == []
```
- [ ] **Step 2: Run tests to verify they fail**
Run: `pytest tests/review/test_finalizer.py -v`
Expected: FAIL with `ModuleNotFoundError: No module named 'backend.review.finalizer'`
- [ ] **Step 3: Implement finalizer decision application**
```python
from typing import Dict, List, Tuple
def apply_decisions(prioritized: List[dict], decisions: Dict[str, dict]) -> Tuple[List[dict], List[dict]]:
kept: List[dict] = []
suppressed: List[dict] = []
for issue in prioritized:
review_id = f"finding:{issue.get('issue_id')}"
decision = decisions.get(review_id) or {}
action = decision.get("decision")
if action == "reject":
suppressed.append({
**issue,
"review_state": "rejected",
"reason_code": decision.get("reason_code"),
"review_comment": decision.get("comment") or "",
})
elif action == "unsure":
kept.append({**issue, "review_state": "unsure"})
else:
kept.append({**issue, "review_state": "confirmed" if action == "confirm" else "unreviewed"})
return kept, suppressed
```
- [ ] **Step 4: Run tests to verify they pass**
Run: `pytest tests/review/test_finalizer.py -v`
Expected: PASS
- [ ] **Step 5: Commit**
```bash
git add backend/review/finalizer.py backend/agents/runner.py tests/review/test_finalizer.py
git commit -m "Finalize reviewed agent findings"
```
---
### Task 7: Feedback labels and metrics
**Files:**
- Create: `backend/review/feedback.py`
- Create: `backend/review/metrics.py`
- Test: `tests/review/test_feedback.py`
**Interfaces:**
- Consumes: queue items and validated decisions.
- Produces:
- `decision_to_label(queue_item: dict, decision: dict, job: dict) -> dict`
- `write_label(out_dir: str, label: dict) -> None`
- `aggregate_labels(labels: list[dict], include_text: bool = False) -> dict`
- [ ] **Step 1: Write failing feedback tests**
```python
from backend.review.metrics import aggregate_labels
def test_aggregate_redacts_text_by_default():
labels = [{"decision": "reject", "reason_code": "missing_evidence", "comment": "secret", "payload": {"evidence": [{"source_text": "secret"}]}}]
summary = aggregate_labels(labels)
assert summary["reject"] == 1
assert "secret" not in str(summary)
```
- [ ] **Step 2: Run tests to verify they fail**
Run: `pytest tests/review/test_feedback.py -v`
Expected: FAIL with `ModuleNotFoundError: No module named 'backend.review.metrics'`
- [ ] **Step 3: Implement label writing and aggregation**
```python
from collections import Counter
from typing import Dict, List
def aggregate_labels(labels: List[dict], include_text: bool = False) -> Dict:
decisions = Counter(label.get("decision") or "unknown" for label in labels)
reasons = Counter(label.get("reason_code") or "none" for label in labels if label.get("decision") == "reject")
summary = {
"total": len(labels),
"decisions": dict(decisions),
"reject_reasons": dict(reasons),
}
for label in labels:
decision = label.get("decision") or "unknown"
summary[decision] = summary.get(decision, 0) + 1
if include_text:
summary["labels"] = labels
return summary
```
- [ ] **Step 4: Run tests to verify they pass**
Run: `pytest tests/review/test_feedback.py -v`
Expected: PASS
- [ ] **Step 5: Commit**
```bash
git add backend/review/feedback.py backend/review/metrics.py tests/review/test_feedback.py
git commit -m "Add review feedback labels and aggregate metrics"
```
---
### Task 8: Two-phase email
**Files:**
- Modify: `backend/email_sender.py`
- Modify: `backend/jobs.py`
- Test: `tests/api/test_review_email_flow.py`
**Interfaces:**
- Consumes: existing `_smtp_ready` and `_send` helpers.
- Produces:
- `send_review_required(recipient_email: str, report: dict, review_url: str) -> bool`
- [ ] **Step 1: Write failing email flow test**
```python
from backend.email_sender import send_review_required
def test_review_required_email_skips_without_smtp(monkeypatch):
monkeypatch.setattr("backend.email_sender._smtp_ready", lambda: False)
assert send_review_required("user@example.com", {"source": "set.pdf", "summary": {}}, "http://localhost:8099/?job=abc") is False
```
- [ ] **Step 2: Run tests to verify they fail**
Run: `pytest tests/api/test_review_email_flow.py -v`
Expected: FAIL with `ImportError: cannot import name 'send_review_required'`
- [ ] **Step 3: Implement review-required email**
```python
def send_review_required(recipient_email: str, report: dict, review_url: str) -> bool:
if not recipient_email or not _smtp_ready():
return False
msg = EmailMessage()
msg["Subject"] = f"Conflict Checker - review required - {report.get('source', 'drawing set')}"
msg["From"] = config.SMTP_FROM or config.SMTP_USER
msg["To"] = recipient_email
review = report.get("summary", {}).get("review", {})
msg.set_content(
"Agent analysis is complete and waiting for human review.\n\n"
f"Required review items: {review.get('required', 0)}\n"
f"Review URL: {review_url}\n"
)
return _send(msg)
```
- [ ] **Step 4: Run tests to verify they pass**
Run: `pytest tests/api/test_review_email_flow.py -v`
Expected: PASS
- [ ] **Step 5: Commit**
```bash
git add backend/email_sender.py backend/jobs.py tests/api/test_review_email_flow.py
git commit -m "Send review-required email before final report"
```
---
### Task 9: Frontend review queue
**Files:**
- Modify: `frontend/index.html`
- Test: `tests/api/test_review_api.py` plus manual browser check
**Interfaces:**
- Consumes: `GET /jobs/{id}`, `GET /jobs/{id}/review`, `POST /jobs/{id}/review-decisions`, `POST /jobs/{id}/finalize-review`.
- Produces: browser flow for `needs_review` jobs.
- [ ] **Step 1: Add failing API expectation for review progress field**
```python
def test_job_includes_review_progress(monkeypatch):
# Extend tests/api/test_review_api.py to assert get_job returns report.summary.review.
assert "review" in {"summary": {"review": {"required": 1, "completed": 0}}}["summary"]
```
- [ ] **Step 2: Run tests to verify current behavior**
Run: `pytest tests/api/test_review_api.py -v`
Expected: PASS for API fields added in Task 5.
- [ ] **Step 3: Implement minimal review UI**
Add a `renderReview(job)` path in `frontend/index.html` that:
- fetches `/jobs/${jobId}/review`,
- renders blocking items first,
- shows `payload.description`, `payload.location`, `payload.category`, `payload.severity`, `payload.confidence`, and `payload.evidence`,
- requires a reason code when `reject` is selected,
- posts decisions to `/jobs/${jobId}/review-decisions`,
- calls `/jobs/${jobId}/finalize-review` only when `progress.remaining === 0`.
- [ ] **Step 4: Manual browser check**
Run: `uvicorn backend.main:app --reload --port 8099`
Expected: a synthetic `needs_review` job shows the queue, decisions persist across refresh, and finalize is blocked until required items are decided.
- [ ] **Step 5: Commit**
```bash
git add frontend/index.html tests/api/test_review_api.py
git commit -m "Add frontend human review queue"
```
---
### Task 10: Config, docs, and rollout
**Files:**
- Modify: `backend/config.py`
- Modify: `backend/.env.example`
- Modify: `README.md`
- Test: `tests/review/test_policy.py`, `tests/review/test_store.py`, `tests/review/test_gate.py`, `tests/agents/test_runner_review_gate.py`, `tests/api/test_review_api.py`, `tests/review/test_finalizer.py`, `tests/review/test_feedback.py`, `tests/api/test_review_email_flow.py`
**Interfaces:**
- Consumes: all previous tasks.
- Produces:
- `AGENT_REQUIRE_REVIEW = true`
- `AGENT_REVIEW_AUDIT_SAMPLE = 5`
- `REVIEW_AGGREGATE_INCLUDE_TEXT = false`
- [ ] **Step 1: Add config assertions to existing policy test file**
```python
from backend import config
def test_review_defaults():
assert config.AGENT_REQUIRE_REVIEW is True
assert config.AGENT_REVIEW_AUDIT_SAMPLE == 5
assert config.REVIEW_AGGREGATE_INCLUDE_TEXT is False
```
- [ ] **Step 2: Run tests to verify they fail**
Run: `pytest tests/review/test_policy.py::test_review_defaults -v`
Expected: FAIL with `AttributeError` for missing config values.
- [ ] **Step 3: Implement config and docs**
Add to `backend/config.py`:
```python
AGENT_REQUIRE_REVIEW = os.getenv("AGENT_REQUIRE_REVIEW", "true").strip().lower() in ("1", "true", "yes")
AGENT_REVIEW_AUDIT_SAMPLE = int(os.getenv("AGENT_REVIEW_AUDIT_SAMPLE", "5"))
REVIEW_AGGREGATE_INCLUDE_TEXT = os.getenv("REVIEW_AGGREGATE_INCLUDE_TEXT", "false").strip().lower() in ("1", "true", "yes")
```
Add the same keys to `backend/.env.example` and document the two-email flow and privacy boundary in `README.md`.
- [ ] **Step 4: Run full test suite**
Run: `pytest -v`
Expected: PASS
- [ ] **Step 5: Commit**
```bash
git add backend/config.py backend/.env.example README.md tests
git commit -m "Configure required agent human review"
```
---
## Execution Handoff
Plan complete and saved to `docs/superpowers/plans/2026-07-28-agent-human-review.md`. Two execution options:
**1. Subagent-Driven (recommended)** - Dispatch a fresh subagent per task, review between tasks, fast iteration.
**2. Inline Execution** - Execute tasks in this session using executing-plans, batch execution with checkpoints.
Which approach?
@@ -0,0 +1,245 @@
# Required Human Review for Agent Pipeline Design
**Date:** 2026-07-28
**Status:** Approved
**Owner:** Conflict Checker Agent pipeline
## Goal
Make Agent mode produce higher-quality findings by requiring structured human review before final RFIs/reports are issued, and by turning review decisions into usable feedback for future prompt, rule, threshold, and evaluation improvements.
## Background
Agent mode is intended to replace Classic mode. Its advantage is the holistic project picture: sheet extraction, sheet index, jurisdiction, semantic linking, specialist findings, Brain consolidation, and RFI generation. The main quality risks are missed real conflicts, false positives, weak or unsupported findings, and silent stage/scope degradation.
There are not enough known-good golden sets to rely only on golden-set regression. Human review becomes the feedback mechanism. The human is not expected to review every raw extraction; the human reviews a curated queue after Brain consolidation and before final report/RFI issuance.
## Requirements
### Functional requirements
1. Agent web jobs must not reach `done` until required human review is complete.
2. The Agent pipeline runs through Brain, then enters `needs_review`.
3. RFI generation happens only after review finalization.
4. Required review items include:
- all critical/high severity findings,
- all low-confidence findings,
- sensitive categories: missing element, code/ADA/egress/fire separation, spatial clash/clearance,
- a small audit sample of medium/low findings and clean/no-finding clusters.
5. Review decisions support `confirm`, `reject`, `unsure`, and `needs_clarification`.
6. Rejections require a reason code.
7. Review progress persists to disk and survives server restart.
8. Rejected findings are suppressed, not deleted.
9. Clarifications are stored as first-class artifacts.
10. Where practical, clarification triggers targeted rerun of only the affected scope.
11. Aggregate feedback must not contain raw drawing text/images by default.
12. Classic mode remains unchanged.
### Non-functional requirements
- No automatic prompt mutation from human labels.
- No final email before review completion.
- Review endpoints must be treated as state-changing and sensitive.
- Review logic must be testable without LLM calls, PDFs, OpenRouter, or network access.
- Targeted reruns must degrade gracefully and must not crash finalization.
## Architecture
Add three small components.
### ReviewGate
Runs after Brain and before RFI/report finalization.
Consumes:
- `ProjectMemory` snapshot
- Brain prioritized issues
- Brain decisions
- review policy
Produces:
- `review/review_queue.json`
- candidate report with `summary.agent_status = "needs_review"`
- job transition to `needs_review`
### ReviewStore
Owns review persistence under the job output directory.
Stores:
- `review/review_queue.json`
- `review/review_decisions.json`
- `review/review_progress.json`
Writes must be atomic using a temporary file plus `os.replace`, matching the existing LLM cache/report artifact style.
### ReviewFinalizer
Runs after required decisions are submitted.
Responsibilities:
- validate completeness,
- apply decisions,
- perform bounded targeted reruns for clarification where supported,
- re-run Brain only for affected findings,
- draft RFIs only for kept/confirmed issues,
- write final artifacts,
- transition to `done`,
- send final email.
## Job lifecycle
Current lifecycle:
`queued -> running -> done -> email`
New Agent lifecycle:
`queued -> running -> needs_review -> reviewing -> finalizing -> done -> email`
Additional failure state:
- `finalization_error`
If the server restarts while a job is in `needs_review` or `reviewing`, the backend rebuilds state from `outputs/<job_id>/conflicts.json`, `outputs/<job_id>/review/review_queue.json`, and `outputs/<job_id>/review/review_decisions.json`.
## Email behavior
If email is enabled, Agent mode sends two emails:
1. **Review required** when the job enters `needs_review`.
2. **Final report** only after review finalization.
If SMTP is not configured, the UI still shows `needs_review` and no email failure crashes the job.
## Review queue policy
Blocking review items are findings that meet any of these rules:
- severity is `critical` or `high`,
- confidence is `low`,
- category is `missing_element`,
- source stage is `code`,
- category is in `ada`, `tas_tdlr`, `egress`, `fire_separation`, `occupancy`, `spatial_clash`, `clearance_conflict`, or `penetration_conflict`.
Audit sample items are selected deterministically from:
- medium/low findings not already blocking,
- clean clusters with no findings,
- no-finding scopes when available.
Default audit sample size is 5 items.
## Review decision schema
```json
{
"review_item_id": "finding:AGENT-0007",
"decision": "reject",
"reason_code": "same_value_different_representation",
"category_correction": null,
"severity_correction": null,
"comment": "9'-0\" AFF and 108 inches are the same value here.",
"clarification_answer": null,
"reviewed_at": "2026-07-28T12:00:00Z"
}
```
Allowed reason codes:
- `wrong_cluster_link`
- `same_value_different_representation`
- `not_a_contradiction`
- `missing_evidence`
- `extraction_misread`
- `code_path_not_applicable`
- `duplicate`
- `severity_too_high`
- `severity_too_low`
- `other`
## Finalization rules
- All blocking review items must have a valid decision before finalization.
- Confirmed findings become final `validated_issues`.
- Unsure findings remain included but are flagged as `review_state = "unsure"`.
- Rejected findings become `suppressed_issues` with reason code and comment.
- Clarification answers are stored and, when the affected scope is rerunnable, trigger a targeted rerun.
- Targeted rerun failure creates an `analysis_gap` finding and does not block finalization unless the reviewer chooses to reject the affected item.
- RFIs are drafted only for final kept issues.
## Feedback labels
Every decision emits a label artifact for metrics:
```json
{
"review_item_id": "finding:AGENT-0007",
"job_id": "abc123",
"pipeline_mode": "agent",
"source_stage": "conflict",
"category": "elevation_disagreement",
"severity": "high",
"confidence": "medium",
"decision": "reject",
"reason_code": "same_value_different_representation",
"location": "Room 204 / Level 2",
"disciplines": ["Architectural", "Mechanical"],
"sheets": ["A2.1", "M2.1"],
"drawing_type": "floor_plan",
"models_used": ["google/gemini-2.5-pro"],
"created_at": "2026-07-28T12:00:00Z"
}
```
Default aggregate metrics exclude `source_text`, images, raw sheet content, and reviewer free-text comments.
## API shape
- `GET /jobs/{job_id}` includes `needs_review`, `reviewing`, `finalizing`, `done`, `error`, or `finalization_error` plus review progress.
- `GET /jobs/{job_id}/review` returns `{ "queue": [...], "progress": {...} }`.
- `POST /jobs/{job_id}/review-decisions` saves one or more decisions.
- `POST /jobs/{job_id}/finalize-review` validates completeness and finalizes the job.
## Security and privacy
Review endpoints are more sensitive than read-only report endpoints because they mutate job state and expose evidence. Before required review is enabled beyond a trusted LAN, the app should have reverse-proxy auth, a shared access token, or explicit deployment documentation stating that the UI/API must not be exposed publicly.
Review artifacts stay job-local by default. Cross-job aggregate metrics use metadata and reason codes only unless richer retention is explicitly enabled later.
## Testing strategy
Tests must not require PDFs, LLMs, OpenRouter, or network access.
Cover:
- required-review trigger policy,
- review queue construction,
- decision validation and reason codes,
- finalization behavior for confirm/reject/unsure/clarification,
- restart recovery from review artifacts,
- targeted rerun failure degradation,
- metrics redaction,
- API state transitions,
- email flow blocking until finalization.
## Rollout
- Classic mode is unchanged.
- Agent web jobs default to required human review.
- CLI supports an explicit bypass flag, `--no-review`, for tuning/debug runs.
- Review state and decisions are always written to job artifacts.
- Aggregate feedback is metadata-only by default.
## Acceptance criteria
- An Agent web job cannot reach `done` or send the final email while required review items are undecided.
- Rejected findings are suppressed with reason codes and remain auditable.
- Review progress survives server restart.
- Clarification failures degrade to visible `analysis_gap`, not job failure.
- Aggregate feedback contains no raw drawing text/images by default.
- New tests cover the review gate without requiring LLM calls.
+304 -34
View File
@@ -72,6 +72,14 @@
.note { background:var(--panel); border:1px solid var(--line); border-radius:10px;
padding:16px 18px; margin:18px 0; }
.note b { color:var(--text); }
.review-controls { margin-top:10px; padding-top:10px; border-top:1px solid var(--line); font-size:13px; }
.review-controls label { margin-right:14px; cursor:pointer; white-space:nowrap; }
.review-controls select, .review-controls input[type=text] { background:#0c0e13; color:var(--text);
border:1px solid var(--line); border-radius:6px; padding:6px 8px; font-size:13px; margin-top:6px; }
.review-controls input[type=text] { width:100%; }
.review-controls .hidden { display:none; }
.pill.blocking { background:rgba(255,93,87,.15); color:var(--hi); }
.pill.audit { background:rgba(91,140,255,.15); color:var(--accent); }
.pill.critical { background:rgba(255,93,87,.28); color:#fff; }
.sheetlink { color:var(--accent); cursor:pointer; text-decoration:underline dotted; }
#viewer { position:fixed; inset:0; background:rgba(0,0,0,.88); display:none;
@@ -87,7 +95,7 @@
<body>
<header>
<h1>Conflict Checker</h1>
<div class="sub">Cross-discipline design contradiction review for construction drawing sets</div>
<div class="sub">Cross-discipline design contradiction review for construction drawing sets<span id="buildTag"></span></div>
</header>
<main>
<div class="drop" id="drop">
@@ -105,19 +113,30 @@
<input type="text" id="occupancy" placeholder="Occupancy (e.g. Business, Assembly)" style="width:100%;margin-top:8px" />
<input type="text" id="work_type" placeholder="Work type (new building, remodel, TI, addition)" style="width:100%;margin-top:8px" />
</details>
<div class="email-card">
<label>&#129302; Pipeline</label>
<label style="display:block;font-weight:400;margin-top:6px">
<input type="radio" name="pipeline_mode" value="classic" checked>
Classic pipeline &mdash; current production workflow</label>
<label style="display:block;font-weight:400;margin-top:6px">
<input type="radio" name="pipeline_mode" value="agent">
Agent pipeline &mdash; experimental specialist-agent workflow</label>
</div>
<div class="email-card">
<label>&#9881;&#65039; Compute <span class="opt">(text stages; vision always runs on the API)</span></label>
<label style="display:block;font-weight:400;margin-top:6px">
<input type="radio" name="compute" value="openrouter" checked> OpenRouter &mdash; all stages (fastest, paid)</label>
<label style="display:block;font-weight:400;margin-top:6px">
<input type="radio" name="compute" value="local"> Hybrid &mdash; text stages on local LLM (cheaper, slower)</label>
<div class="field">
<span>Vision model <span class="opt">(image stages)</span></span>
<select id="vision_model" disabled><option value="">Loading models&hellip;</option></select>
</div>
<div class="field">
<span>Text model <span class="opt">(non-image stages / hybrid fallback)</span></span>
<select id="text_model" disabled><option value="">Loading models&hellip;</option></select>
<div id="modelPick" style="margin-top:10px">
<div class="field">
<span>Vision model <span class="opt">(image stages)</span> <span class="opt" id="modelNote">loading...</span></span>
<select id="vision_model" disabled><option value="">Loading models&hellip;</option></select>
</div>
<div class="field">
<span>Text model <span class="opt">(non-image stages)</span></span>
<select id="text_model" disabled><option value="">Loading models&hellip;</option></select>
</div>
</div>
</div>
<button class="btn full" id="run" disabled>Run conflict check</button>
@@ -149,13 +168,21 @@ const drop=document.getElementById('drop'), fileInput=document.getElementById('f
liveLog=document.getElementById('liveLog'),
logBox=document.getElementById('logBox'),
logHint=document.getElementById('logHint');
let chosen=null, polling=null, currentJobId=null, sheetPage={}, viewerZoom=1;
let chosen=null, polling=null, currentJobId=null, sheetPage={}, viewerZoom=1, reviewDirty=false;
function modelLabel(m){
// Include per-1M-token pricing when the catalog provides it.
let s=m.name||m.id;
if(m.prompt_usd_per_mtok!=null)
s+=' — $'+m.prompt_usd_per_mtok+' / $'+m.completion_usd_per_mtok+' per 1M tok';
return s;
}
function fillSelect(sel, items, preferred){
sel.innerHTML='';
(items||[]).forEach(m=>{
const opt=document.createElement('option');
opt.value=m.id; opt.textContent=m.name||m.id;
opt.value=m.id; opt.textContent=modelLabel(m);
if(m.id===preferred) opt.selected=true;
sel.appendChild(opt);
});
@@ -167,7 +194,9 @@ function fillSelect(sel, items, preferred){
sel.disabled=false;
}
let modelsLoaded=false;
async function loadModels(){
const note=document.getElementById('modelNote');
try{
const res=await fetch('/models');
if(!res.ok) throw new Error('models HTTP '+res.status);
@@ -175,13 +204,13 @@ async function loadModels(){
const defs=data.defaults||{};
fillSelect(visionSel, data.vision, defs.vision);
fillSelect(textSel, data.text, defs.text);
if(data.error){
console.warn('Model catalog degraded:', data.error);
}
modelsLoaded=true;
note.textContent='('+(data.text||[]).length+' text / '+(data.vision||[]).length+' vision available)';
}catch(err){
visionSel.innerHTML='<option value="">(default)</option>';
textSel.innerHTML='<option value="">(default)</option>';
visionSel.disabled=false; textSel.disabled=false;
note.textContent='using configured defaults (list unavailable)';
console.warn('Could not load models:', err);
}
}
@@ -219,8 +248,13 @@ runBtn.addEventListener('click',async e=>{
});
const compute=(document.querySelector('input[name="compute"]:checked')||{}).value;
fd.append('text_local', compute==='local' ? 'true' : 'false');
if(visionSel.value) fd.append('vision_model', visionSel.value);
if(textSel.value) fd.append('text_model', textSel.value);
if(compute==='openrouter'){
// Model picks only apply to OpenRouter compute; hybrid keeps its local model.
if(visionSel.value) fd.append('vision_model', visionSel.value);
if(textSel.value) fd.append('text_model', textSel.value);
}
const pipelineMode=(document.querySelector('input[name="pipeline_mode"]:checked')||{}).value||'classic';
fd.append('pipeline_mode',pipelineMode);
try{
const res=await fetch('/check',{method:'POST',body:fd});
if(!res.ok){ const err=await res.json().catch(()=>({detail:res.statusText}));
@@ -244,7 +278,8 @@ function poll(jobId){
const res=await fetch('/jobs/'+jobId);
if(!res.ok) throw new Error('job not found');
const job=await res.json();
if(job.log_tail && job.log_tail.length) showLog(job.log_tail, job.status==='running'||job.status==='queued');
const live=['running','queued','finalizing'].includes(job.status);
if(job.log_tail && job.log_tail.length) showLog(job.log_tail, live);
if(job.status==='running'||job.status==='queued'){
statusEl.innerHTML='<span class="spinner"></span>'+esc(job.stage||'Working...')+
' &middot; you can leave this page';
@@ -252,6 +287,16 @@ function poll(jobId){
clearInterval(polling); polling=null; runBtn.disabled=false;
if(job.log && job.log.length) showLog(job.log, false);
render(job.report);
} else if(job.status==='needs_review'||job.status==='reviewing'){
clearInterval(polling); polling=null; runBtn.disabled=false;
if(job.log && job.log.length) showLog(job.log, false);
renderReview(job);
} else if(job.status==='finalizing'){
statusEl.innerHTML='<span class="spinner"></span>Finalizing reviewed report...';
} else if(job.status==='finalization_error'){
clearInterval(polling); polling=null; runBtn.disabled=false;
statusEl.textContent='Finalization failed: '+(job.error||'unknown error');
if(job.log && job.log.length) showLog(job.log, false);
} else if(job.status==='error'){
clearInterval(polling); polling=null; runBtn.disabled=false;
statusEl.textContent='Run failed: '+(job.error||'unknown error');
@@ -265,6 +310,25 @@ function poll(jobId){
}
function esc(s){ return (s==null?'':String(s)).replace(/[&<>]/g,c=>({'&':'&amp;','<':'&lt;','>':'&gt;'}[c])); }
function escAttr(s){ return esc(s).replace(/"/g,'&quot;'); }
function syncPipelineOptions(){
const agent=(document.querySelector('input[name="pipeline_mode"]:checked')||{}).value==='agent';
const local=document.querySelector('input[name="compute"][value="local"]');
local.disabled=agent;
if(agent&&local.checked) document.querySelector('input[name="compute"][value="openrouter"]').checked=true;
}
document.querySelectorAll('input[name="pipeline_mode"]').forEach(el=>el.addEventListener('change',syncPipelineOptions));
syncPipelineOptions();
// --- model pickers (OpenRouter compute only) ---
function syncCompute(){
const openrouter=(document.querySelector('input[name="compute"]:checked')||{}).value==='openrouter';
document.getElementById('modelPick').style.display=openrouter?'block':'none';
if(openrouter&&!modelsLoaded) loadModels();
}
document.querySelectorAll('input[name="compute"]').forEach(el=>el.addEventListener('change',syncCompute));
syncCompute();
// --- sheet viewer ---
function pageFor(num){ return sheetPage[num] || sheetPage[(num||'').toUpperCase()] || null; }
@@ -292,6 +356,43 @@ function closeSheet(){ document.getElementById('viewer').classList.remove('open'
document.getElementById('viewer').addEventListener('click',e=>{ if(e.target.id==='viewer') closeSheet(); });
document.addEventListener('keydown',e=>{ if(e.key==='Escape') closeSheet(); });
// --- conflicts grouped by discipline pair ---
const SEV_RANK={critical:0,high:1,medium:2,low:3};
function sevRank(c){ const r=SEV_RANK[(c.severity||'').toLowerCase()]; return r==null?4:r; }
function groupConflicts(conflicts){
// Group key: disciplines sorted alphabetically, joined ' vs ' (order-independent
// pair). Missing disciplines -> 'General'. Groups ordered by their most severe
// conflict, then name; items within a group ordered critical->high->medium->low.
const groups={};
for(const c of conflicts||[]){
const ds=(c.disciplines||[]).map(d=>String(d)).filter(Boolean).sort();
const key=ds.length?ds.join(' vs '):'General';
(groups[key]=groups[key]||[]).push(c);
}
const names=Object.keys(groups).sort((a,b)=>{
const ra=Math.min.apply(null,groups[a].map(sevRank)),
rb=Math.min.apply(null,groups[b].map(sevRank));
return (ra-rb)||a.localeCompare(b);
});
return names.map(name=>({name:name,
items:groups[name].slice().sort((x,y)=>sevRank(x)-sevRank(y))}));
}
function conflictCard(c){
let html='<div class="conflict '+esc(c.severity)+'">'+
'<div class="row"><span class="cat">'+esc(c.category)+'</span>'+
'<span class="pill '+esc(c.severity)+'">'+esc(c.severity)+'</span></div>'+
'<div class="loc">'+esc(c.location)+'</div>'+
'<div class="meta">'+esc((c.disciplines||[]).join(' vs '))+
' &middot; sheets '+sheetList(c.sheets)+'</div>'+
'<div class="desc">'+esc(c.description)+'</div>';
if(c.evidence&&c.evidence.length){
html+='<div class="ev">'+c.evidence.map(e=>
'<div><span class="d">'+esc(e.discipline)+'</span> ('+sheetSpan(e.sheet)+'): "'+esc(e.source_text)+'"</div>').join('')+'</div>';
}
if(c.recommended_resolution){ html+='<div class="reso">Resolution: '+esc(c.recommended_resolution)+'</div>'; }
return html+'</div>';
}
function render(rep){
const s=rep.summary;
if(!currentJobId) currentJobId=new URLSearchParams(location.search).get('job');
@@ -306,7 +407,9 @@ function render(rep){
if(textModel) modelLine+=' · text: '+esc(textModel);
if(fallbacks) modelLine+=' ('+fallbacks+' cloud fallback'+(fallbacks>1?'s':'')+')';
}
statusEl.textContent='Analyzed '+s.sheets_analyzed+' sheets ('+(s.disciplines.join(', ')||'none')+')'+modelLine+'.';
const mode=s.pipeline_mode||'classic';
statusEl.textContent=(mode==='agent'?'Agent':'Classic')+' pipeline analyzed '+s.sheets_analyzed+
' sheets ('+(s.disciplines.join(', ')||'none')+')'+modelLine+'.';
let html='<div class="summary">'+
stat(s.conflicts_found,'conflicts')+
stat(s.by_severity.high,'high')+
@@ -315,21 +418,16 @@ function render(rep){
stat(s.assertions_extracted,'facts')+
stat(s.clusters_checked,'clusters')+
(s.cost_usd!=null?stat('$'+Number(s.cost_usd).toFixed(2),'cost'):'')+'</div>';
if(!rep.conflicts.length){ html+='<div class="empty">No cross-discipline conflicts detected.</div>'; }
for(const c of rep.conflicts){
html+='<div class="conflict '+esc(c.severity)+'">'+
'<div class="row"><span class="cat">'+esc(c.category)+'</span>'+
'<span class="pill '+esc(c.severity)+'">'+esc(c.severity)+'</span></div>'+
'<div class="loc">'+esc(c.location)+'</div>'+
'<div class="meta">'+esc((c.disciplines||[]).join(' vs '))+
' &middot; sheets '+sheetList(c.sheets)+'</div>'+
'<div class="desc">'+esc(c.description)+'</div>';
if(c.evidence&&c.evidence.length){
html+='<div class="ev">'+c.evidence.map(e=>
'<div><span class="d">'+esc(e.discipline)+'</span> ('+sheetSpan(e.sheet)+'): "'+esc(e.source_text)+'"</div>').join('')+'</div>';
}
if(c.recommended_resolution){ html+='<div class="reso">Resolution: '+esc(c.recommended_resolution)+'</div>'; }
html+='</div>';
if(s.agent_status==='skeleton'){
html+='<div class="note"><b>Agent pipeline skeleton:</b> routing and artifacts are active; '+
'specialist analysis is added in the next implementation phases.</div>';
} else if(!rep.conflicts.length){
html+='<div class="empty">No cross-discipline conflicts detected.</div>';
}
for(const g of groupConflicts(rep.conflicts)){
html+='<details open style="margin-top:16px"><summary><b>'+esc(g.name)+' ('+g.items.length+')</b></summary>';
for(const c of g.items){ html+=conflictCard(c); }
html+='</details>';
}
const issues=rep.validated_issues||[];
@@ -337,10 +435,13 @@ function render(rep){
html+='<details open style="margin-top:24px"><summary><b>QAQC issues ('+issues.length+')</b> '+
'<span class="opt">conflicts + full-set + code/ADA + constructability, deduplicated</span></summary>';
for(const c of issues){
const rs=c.review_state;
html+='<div class="conflict '+esc(c.severity)+'">'+
'<div class="row"><span class="cat">'+esc(c.source_stage)+' &middot; '+esc(c.category)+'</span>'+
'<span class="pill '+esc(c.severity)+'">'+esc(c.severity)+
(c.risk_score!=null?(' &middot; risk '+esc(c.risk_score)):'')+'</span></div>'+
(c.risk_score!=null?(' &middot; risk '+esc(c.risk_score)):'')+'</span>'+
(rs&&['unsure','clarified','clarification_failed'].includes(rs)?
' <span class="pill audit">'+esc(rs.replace(/_/g,' '))+'</span>':'')+'</div>'+
'<div class="loc">'+esc(c.location)+'</div>'+
((c.sheets||[]).length?('<div class="meta">Sheets: '+sheetList(c.sheets)+'</div>'):'')+
'<div class="desc">'+esc(c.description)+'</div>';
@@ -374,9 +475,178 @@ function render(rep){
}
function stat(v,l){ return '<div class="stat"><b>'+esc(v)+'</b><span>'+esc(l)+'</span></div>'; }
// --- human review queue (agent pipeline) ---
const REVIEW_REASON_CODES=['wrong_cluster_link','same_value_different_representation',
'not_a_contradiction','missing_evidence','extraction_misread','code_path_not_applicable',
'duplicate','severity_too_high','severity_too_low','other'];
const REVIEW_DECISIONS=['confirm','reject','unsure','needs_clarification'];
async function renderReview(job){
const jobId=job.job_id||currentJobId;
currentJobId=jobId;
statusEl.textContent='Analysis complete \u2014 human review required.';
let data;
try{
const res=await fetch('/jobs/'+jobId+'/review');
if(!res.ok) throw new Error('could not load review queue');
data=await res.json();
}catch(err){ statusEl.textContent='Error: '+err.message; return; }
const queue=data.queue||[], prog=data.progress||{}, prior=data.decisions||{};
let html='<div class="note"><b>Analysis complete \u2014 human review required.</b><br>'+
esc(prog.completed||0)+' of '+esc(prog.required||0)+' required items decided.'+
((prog.remaining||0)>0?' Decide all blocking items, save, then finalize.':
' All required items decided \u2014 you can finalize.')+'</div>';
const blocking=queue.filter(i=>i.blocking), audit=queue.filter(i=>!i.blocking);
blocking.forEach((item,i)=>{ html+=reviewItemHtml(item,'b'+i,prior[item.review_item_id]); });
if(audit.length){
html+='<details style="margin-top:16px"><summary><b>Audit items ('+audit.length+')</b> '+
'<span class="opt">non-blocking &mdash; decisions optional</span></summary>';
audit.forEach((item,i)=>{ html+=reviewItemHtml(item,'a'+i,prior[item.review_item_id]); });
html+='</details>';
}
html+='<div style="margin:18px 0">'+
'<button class="btn" id="saveReviewBtn">Save decisions</button> '+
'<button class="btn" id="finalizeBtn"'+((prog.remaining||0)===0?'':' disabled')+
'>Finalize &amp; send report</button></div>'+
'<div class="status" id="reviewMsg"></div>';
results.innerHTML=html;
results.querySelectorAll('.review-item input[type=radio]').forEach(r=>{
r.addEventListener('change',()=>syncReviewControls(r.closest('.review-item')));
});
reviewDirty=false;
results.querySelectorAll('.review-item input,.review-item select').forEach(el=>{
el.addEventListener('change',()=>{ reviewDirty=true; });
});
document.getElementById('saveReviewBtn').addEventListener('click',saveReviewDecisions);
document.getElementById('finalizeBtn').addEventListener('click',finalizeReview);
}
function reviewItemHtml(item,uid,prev){
prev=prev||{};
const p=item.payload||{};
const sev=p.severity||'medium';
let html='<div class="conflict '+escAttr(sev)+' review-item" data-id="'+escAttr(item.review_item_id)+'">'+
'<div class="row"><span class="cat">'+esc(p.category||item.kind)+'</span>'+
'<span><span class="pill '+(item.blocking?'blocking':'audit')+'">'+
(item.blocking?'blocking':'audit')+'</span> '+
(p.severity?'<span class="pill '+escAttr(sev)+'">'+esc(sev)+'</span>':'')+'</span></div>';
if(item.kind==='clean_cluster'){
html+='<div class="loc">'+esc(p.location||p.key||'(cluster)')+'</div>'+
'<div class="meta">Cluster '+esc(p.key||'')+' &middot; '+
esc((p.assertions||[]).length)+' assertions</div>';
}else{
html+='<div class="loc">'+esc(p.location||'')+'</div>'+
'<div class="desc">'+esc(p.description||'')+'</div>'+
(p.confidence?'<div class="meta">Confidence: '+esc(p.confidence)+'</div>':'');
if(p.evidence&&p.evidence.length){
html+='<div class="ev">'+p.evidence.map(e=>
'<div><span class="d">'+esc(e.discipline)+'</span> ('+sheetSpan(e.sheet)+'): "'+
esc(e.source_text)+'"</div>').join('')+'</div>';
}
}
if((item.reasons||[]).length){
html+='<div class="meta">Review triggers: '+esc(item.reasons.join(', '))+'</div>';
}
html+='<div class="review-controls">'+
REVIEW_DECISIONS.map(d=>'<label><input type="radio" name="dec-'+uid+'" value="'+d+'"'+
(prev.decision===d?' checked':'')+'> '+esc(d.replace(/_/g,' '))+'</label>').join('')+
'<select class="reason'+(prev.decision==='reject'?'':' hidden')+'">'+
'<option value="">Reason code (required for reject)...</option>'+
REVIEW_REASON_CODES.map(c=>'<option value="'+c+'"'+(prev.reason_code===c?' selected':'')+
'>'+esc(c.replace(/_/g,' '))+'</option>').join('')+'</select>'+
'<input type="text" class="comment" placeholder="Comment (optional)" value="'+escAttr(prev.comment||'')+'">'+
'<input type="text" class="clar'+(prev.decision==='needs_clarification'?'':' hidden')+
'" placeholder="Clarification answer" value="'+escAttr(prev.clarification_answer||'')+'">'+
'</div></div>';
return html;
}
function syncReviewControls(el){
const sel=el.querySelector('input[type=radio]:checked');
const v=sel?sel.value:'';
el.querySelector('.reason').classList.toggle('hidden',v!=='reject');
if(v!=='reject') el.querySelector('.reason').value='';
el.querySelector('.clar').classList.toggle('hidden',v!=='needs_clarification');
}
function reviewMsg(m,isErr){
const el=document.getElementById('reviewMsg');
if(el){ el.style.color=isErr?'var(--hi)':'var(--muted)'; el.textContent=m; }
}
function collectReviewDecisions(){
const decisions=[], missingReason=[];
results.querySelectorAll('.review-item').forEach(el=>{
const id=el.getAttribute('data-id');
const sel=el.querySelector('input[type=radio]:checked');
if(!sel) return;
const reason=el.querySelector('.reason').value;
if(sel.value==='reject'&&!reason){ missingReason.push(id); return; }
const d={review_item_id:id, decision:sel.value,
comment:el.querySelector('.comment').value.trim()};
if(sel.value==='reject') d.reason_code=reason;
else if(reason) d.reason_code=reason;
if(sel.value==='needs_clarification')
d.clarification_answer=el.querySelector('.clar').value.trim()||null;
decisions.push(d);
});
return {decisions, missingReason};
}
async function postReviewDecisions(decisions){
const res=await fetch('/jobs/'+currentJobId+'/review-decisions',
{method:'POST',headers:{'Content-Type':'application/json'},
body:JSON.stringify({decisions})});
if(!res.ok){
const err=await res.json().catch(()=>({detail:res.statusText}));
const detail=typeof err.detail==='string'?err.detail:JSON.stringify(err.detail);
throw new Error(detail||'Request failed');
}
}
async function saveReviewDecisions(){
const {decisions,missingReason}=collectReviewDecisions();
if(missingReason.length){
reviewMsg('Reject requires a reason code: '+missingReason.join(', '),true); return;
}
if(!decisions.length){ reviewMsg('No decisions set yet.',true); return; }
try{
await postReviewDecisions(decisions);
renderReview({job_id:currentJobId});
}catch(err){ reviewMsg('Save failed: '+err.message,true); }
}
async function finalizeReview(){
reviewMsg('');
try{
if(reviewDirty){
// Auto-save unsaved control edits so they aren't lost at finalize.
const {decisions,missingReason}=collectReviewDecisions();
if(missingReason.length){
reviewMsg('Reject requires a reason code: '+missingReason.join(', '),true); return;
}
if(decisions.length) await postReviewDecisions(decisions);
reviewDirty=false;
}
const res=await fetch('/jobs/'+currentJobId+'/finalize-review',{method:'POST'});
if(res.status===409){
const err=await res.json().catch(()=>({}));
const d=err.detail||{};
const prog=d.progress?(' ('+(d.progress.remaining||0)+' required items undecided)'):'';
reviewMsg('Cannot finalize: '+(d.detail||'conflict')+prog,true); return;
}
if(!res.ok) throw new Error('Request failed ('+res.status+')');
statusEl.innerHTML='<span class="spinner"></span>Finalizing reviewed report...';
poll(currentJobId);
}catch(err){ reviewMsg('Finalize failed: '+err.message,true); }
}
// If opened from an email link (/?job=<id>), load that job's results directly.
(function init(){
loadModels();
// Show the deployed build in the header so it's obvious which version is up.
fetch('/health').then(r=>r.ok?r.json():null).then(h=>{
if(h&&h.build) document.getElementById('buildTag').textContent=' · build '+h.build;
}).catch(()=>{});
const jobId=new URLSearchParams(location.search).get('job');
if(jobId){ statusEl.innerHTML='<span class="spinner"></span>Loading job '+esc(jobId)+'...'; poll(jobId); }
})();
+65
View File
@@ -0,0 +1,65 @@
from backend.agents.base import AgentResult
from backend.agents.runner import run_agent_pipeline
def _patch_brain(monkeypatch):
monkeypatch.setattr("backend.agents.runner.convert_pdf_to_images", lambda path: [{"page_number": 1, "base64": "x"}])
monkeypatch.setattr("backend.agents.runner.BrainAgent", lambda usage: type("B", (), {"run": lambda self, findings, sheet_index, jurisdiction: ([{"issue_id": "AGENT-0001", "severity": "high", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"}], [])})())
def test_agent_runner_can_enter_review_mode(monkeypatch, tmp_path):
_patch_brain(monkeypatch)
pdf = tmp_path / "dummy.pdf"
pdf.write_bytes(b"%PDF-1.4\n")
report = run_agent_pipeline(str(pdf), out_dir=str(tmp_path), require_review=True)
assert report["summary"]["agent_status"] == "needs_review"
assert report["summary"]["review"]["required"] == 1
def test_review_mode_writes_memory_snapshot(monkeypatch, tmp_path):
"""The finalizer needs agent/memory.json for targeted clarification reruns."""
_patch_brain(monkeypatch)
pdf = tmp_path / "dummy.pdf"
pdf.write_bytes(b"%PDF-1.4\n")
run_agent_pipeline(str(pdf), out_dir=str(tmp_path), require_review=True)
assert (tmp_path / "agent" / "memory.json").is_file()
def test_review_mode_summary_includes_agent_observability(monkeypatch, tmp_path):
"""Review-mode candidate reports must carry the same usage/stats block as
the wave-7 path so finalizer fix-ups and feedback labels have real data."""
_patch_brain(monkeypatch)
pdf = tmp_path / "dummy.pdf"
pdf.write_bytes(b"%PDF-1.4\n")
report = run_agent_pipeline(str(pdf), out_dir=str(tmp_path), require_review=True)
summary = report["summary"]
assert "agent_stats" in summary
assert summary["by_stage"]["rfis"] == 0
assert summary["by_stage"]["validated"] == 1
assert "conflicts" in summary["by_stage"]
assert "cost_usd" in summary
assert "llm_calls" in summary
assert "cached_calls" in summary
assert "cost_by_stage" in summary
assert "models_used" in summary
def test_agent_runner_without_review_still_writes_rfis(monkeypatch, tmp_path):
_patch_brain(monkeypatch)
monkeypatch.setattr(
"backend.agents.runner.RFIWriterAgent",
lambda usage: type("R", (), {
"name": "rfi_writer",
"run": lambda self, scope: AgentResult(
scope_id=scope.scope_id,
artifacts=[{"issue_id": "AGENT-0001", "question": "Confirm intent?"}],
),
})(),
)
pdf = tmp_path / "dummy.pdf"
pdf.write_bytes(b"%PDF-1.4\n")
report = run_agent_pipeline(str(pdf), out_dir=str(tmp_path), require_review=False)
assert report["summary"]["agent_status"] == "complete"
assert "review" not in report["summary"]
assert len(report["rfis"]) == 1
assert report["rfis"][0]["issue_id"] == "AGENT-0001"
+21
View File
@@ -0,0 +1,21 @@
from fastapi.testclient import TestClient
from backend import config
from backend.main import app
def test_health_includes_version_and_build():
client = TestClient(app)
response = client.get("/health")
assert response.status_code == 200
body = response.json()
assert body["version"] == config.APP_VERSION
assert body["build"] == config.APP_BUILD
def test_app_base_url_default_is_public_site():
assert config.APP_BASE_URL == "https://conchecker.scoutitsystems.com"
def test_app_build_defaults_to_dev():
assert config.APP_BUILD == "dev"
+115
View File
@@ -0,0 +1,115 @@
import threading
import pytest
from fastapi.testclient import TestClient
import backend.jobs as jobs
from backend.main import app
class _SyncThread:
"""Drop-in threading.Thread replacement that runs the target inline."""
def __init__(self, target=None, args=(), kwargs=None, **_ignored):
self._target = target
self._args = args
self._kwargs = kwargs or {}
def start(self):
self._target(*self._args, **self._kwargs)
@pytest.fixture
def job_env(monkeypatch, tmp_path):
monkeypatch.setattr("backend.config.OUTPUT_DIR", str(tmp_path))
monkeypatch.setattr(threading, "Thread", _SyncThread)
monkeypatch.setattr("backend.jobs.send_conflict_report", lambda *a, **k: True)
pdf = tmp_path / "set.pdf"
pdf.write_bytes(b"%PDF-1.4\n")
yield tmp_path
jobs._jobs.clear()
def test_job_log_captures_pipeline_output(job_env, monkeypatch):
def fake_runner(pdf_path, **kwargs):
print("STAGE banner: fake wave ran")
return {"source": "set.pdf", "summary": {"conflicts_found": 0}}
monkeypatch.setattr("backend.jobs.run_pipeline", fake_runner)
job_id = jobs.create_job(str(job_env / "set.pdf"), "set.pdf", pipeline_mode="classic")
log_path = job_env / job_id / "job.log"
assert log_path.is_file()
content = log_path.read_text()
assert "STAGE banner: fake wave ran" in content
assert job_id in content # header line
def test_job_log_endpoint_serves_log_and_404s(job_env, monkeypatch):
monkeypatch.setattr(
"backend.jobs.run_pipeline",
lambda pdf_path, **kw: {"source": "s", "summary": {}},
)
job_id = jobs.create_job(str(job_env / "set.pdf"), "set.pdf", pipeline_mode="classic")
client = TestClient(app)
ok = client.get(f"/jobs/{job_id}/log")
assert ok.status_code == 200
assert ok.headers["content-type"].startswith("text/plain")
assert "Job " + job_id in ok.text
assert client.get("/jobs/nope/log").status_code == 404
def test_model_overrides_passed_to_classic_runner(job_env, monkeypatch):
"""Classic mode: per-run picks travel as run_pipeline kwargs (the runner
sets and clears llm.set_model_overrides itself)."""
seen = {}
def fake_runner(pdf_path, **kwargs):
seen.update(kwargs)
return {"source": "set.pdf", "summary": {}}
monkeypatch.setattr("backend.jobs.run_pipeline", fake_runner)
jobs.create_job(str(job_env / "set.pdf"), "set.pdf",
pipeline_mode="classic",
vision_model="openai/gpt-4o", text_model="openai/gpt-4o-mini")
assert seen["vision_model"] == "openai/gpt-4o"
assert seen["text_model"] == "openai/gpt-4o-mini"
def test_model_overrides_set_and_cleared_around_agent_run(job_env, monkeypatch):
"""Agent mode: the agent runner has no override params, so jobs.py sets
them module-level for the duration of the run."""
from backend import llm
seen = {}
def fake_agent_runner(pdf_path, **kwargs):
seen["vision"] = llm._vision_model_override
seen["text"] = llm._text_model_override
return {"source": "set.pdf", "summary": {}}
monkeypatch.setattr("backend.jobs.run_agent_pipeline", fake_agent_runner)
jobs.create_job(str(job_env / "set.pdf"), "set.pdf",
pipeline_mode="agent",
vision_model="openai/gpt-4o", text_model="openai/gpt-4o-mini")
assert seen["vision"] == "openai/gpt-4o"
assert seen["text"] == "openai/gpt-4o-mini"
assert llm._vision_model_override is None # cleared after the run
assert llm._text_model_override is None
def test_failed_run_logs_traceback(job_env, monkeypatch):
"""A crashed job must leave the traceback in job.log, not just str(e)."""
def boom(pdf_path, **kwargs):
raise RuntimeError("kaboom-stage-failure")
monkeypatch.setattr("backend.jobs.run_pipeline", boom)
job_id = jobs.create_job(str(job_env / "set.pdf"), "set.pdf", pipeline_mode="classic")
assert jobs._jobs[job_id]["status"] == "error"
content = (job_env / job_id / "job.log").read_text()
assert "Traceback (most recent call last)" in content
assert "RuntimeError: kaboom-stage-failure" in content
+88
View File
@@ -0,0 +1,88 @@
from fastapi.testclient import TestClient
import backend.models as models
from backend import config
from backend.main import app
_PAYLOAD = {
"data": [
{
"id": "openai/gpt-4o",
"name": "GPT-4o",
"pricing": {"prompt": "0.0000025", "completion": "0.00001"},
"context_length": 128000,
"architecture": {"input_modalities": ["text", "image"],
"output_modalities": ["text"]},
},
{
"id": "google/gemini-2.5-pro",
"name": "Gemini 2.5 Pro",
"pricing": {"prompt": "0.00000125", "completion": "0.00001"},
"context_length": 1000000,
"architecture": {"modality": "text+image->text"},
},
{
"id": "meta-llama/llama-3.1-70b-instruct",
"name": "Llama 3.1 70B Instruct",
"pricing": {"prompt": "0.0000005", "completion": "0.0000008"},
"context_length": 131072,
"architecture": {"input_modalities": ["text"],
"output_modalities": ["text"]},
},
]
}
def _reset_cache():
models._cache["models"] = None
models._cache["at"] = 0.0
def test_models_endpoint_normalizes_pricing(monkeypatch):
_reset_cache()
monkeypatch.setattr(models, "_fetch_openrouter_models", lambda: _PAYLOAD["data"])
client = TestClient(app)
response = client.get("/models")
assert response.status_code == 200
body = response.json()
assert body["defaults"] == {"vision": config.MODEL, "text": config.TEXT_MODEL}
by_id = {m["id"]: m for m in body["text"]}
assert by_id["openai/gpt-4o"]["prompt_usd_per_mtok"] == 2.5
assert by_id["openai/gpt-4o"]["completion_usd_per_mtok"] == 10.0
assert by_id["openai/gpt-4o"]["context_length"] == 128000
def test_models_endpoint_splits_vision_and_text(monkeypatch):
_reset_cache()
monkeypatch.setattr(models, "_fetch_openrouter_models", lambda: _PAYLOAD["data"])
client = TestClient(app)
body = client.get("/models").json()
vision_ids = {m["id"] for m in body["vision"]}
text_ids = {m["id"] for m in body["text"]}
# Both modality shapes (structured and legacy string) are recognized.
assert vision_ids == {"openai/gpt-4o", "google/gemini-2.5-pro"}
# Text list is the full catalog; vision models appear in both.
assert text_ids == {"openai/gpt-4o", "google/gemini-2.5-pro",
"meta-llama/llama-3.1-70b-instruct"}
def test_models_endpoint_caches(monkeypatch):
_reset_cache()
calls = []
def fake_fetch():
calls.append(1)
return _PAYLOAD["data"]
monkeypatch.setattr(models, "_fetch_openrouter_models", fake_fetch)
client = TestClient(app)
assert client.get("/models").status_code == 200
assert client.get("/models").status_code == 200
assert len(calls) == 1
def test_models_endpoint_502_on_fetch_failure(monkeypatch):
_reset_cache()
monkeypatch.setattr(models, "_fetch_openrouter_models", lambda: None)
client = TestClient(app)
assert client.get("/models").status_code == 502
+439
View File
@@ -0,0 +1,439 @@
"""API tests for the human-review endpoints and review-aware job states."""
import json
import os
import threading
from fastapi.testclient import TestClient
import backend.jobs
from backend.main import app
from backend.review.store import ReviewStore
class _SyncThread:
"""Drop-in threading.Thread replacement that runs the target inline."""
def __init__(self, target=None, args=(), **kwargs):
self._target = target
self._args = args
def start(self):
self._target(*self._args)
def _queue_item(item_id: str) -> dict:
return {"review_item_id": item_id, "kind": "finding",
"blocking": True, "reasons": ["high_severity"], "payload": {}}
def test_review_queue_and_decision_save(monkeypatch, tmp_path):
store = ReviewStore(str(tmp_path))
store.write_queue([_queue_item("finding:AGENT-0001")])
client = TestClient(app)
monkeypatch.setattr("backend.main.get_job", lambda job_id: {"job_id": job_id, "status": "needs_review", "report": {"summary": {}}, "out_dir": str(tmp_path)})
queue_response = client.get("/jobs/job1/review")
assert queue_response.status_code == 200
decision_response = client.post("/jobs/job1/review-decisions", json={"decisions": [{"review_item_id": "finding:AGENT-0001", "decision": "confirm"}]})
assert decision_response.status_code == 200
def test_review_decision_emits_feedback_label(monkeypatch, tmp_path):
"""Every saved decision appends one feedback label under review/."""
store = ReviewStore(str(tmp_path))
store.write_queue([_queue_item("finding:AGENT-0001")])
client = TestClient(app)
monkeypatch.setattr("backend.main.get_job", lambda job_id: {"job_id": job_id, "status": "needs_review", "report": {"summary": {}}, "out_dir": str(tmp_path)})
response = client.post("/jobs/job1/review-decisions", json={"decisions": [{"review_item_id": "finding:AGENT-0001", "decision": "confirm"}]})
assert response.status_code == 200
path = os.path.join(str(tmp_path), "review", "feedback_labels.jsonl")
with open(path, encoding="utf-8") as f:
labels = [json.loads(line) for line in f if line.strip()]
assert len(labels) == 1
assert labels[0]["review_item_id"] == "finding:AGENT-0001"
assert labels[0]["decision"] == "confirm"
def test_review_endpoints_404_for_unknown_job(monkeypatch, tmp_path):
monkeypatch.setattr("backend.config.OUTPUT_DIR", str(tmp_path))
client = TestClient(app)
assert client.get("/jobs/nope/review").status_code == 404
assert client.post("/jobs/nope/review-decisions", json={"decisions": []}).status_code == 404
def test_review_decision_invalid_returns_422(monkeypatch, tmp_path):
client = TestClient(app)
monkeypatch.setattr("backend.main.get_job", lambda job_id: {"job_id": job_id, "status": "needs_review", "report": {"summary": {}}, "out_dir": str(tmp_path)})
response = client.post("/jobs/job1/review-decisions", json={"decisions": [{"review_item_id": "finding:AGENT-0001", "decision": "bogus"}]})
assert response.status_code == 422
def test_partial_review_moves_job_to_reviewing(monkeypatch, tmp_path):
"""Saving some but not all required decisions flips needs_review -> reviewing."""
store = ReviewStore(str(tmp_path))
store.write_queue([_queue_item("finding:AGENT-0001"), _queue_item("finding:AGENT-0002")])
job_id = "jobreviewing"
backend.jobs._jobs[job_id] = {
"job_id": job_id,
"status": "needs_review",
"out_dir": str(tmp_path),
"report": {"summary": {}},
}
try:
client = TestClient(app)
response = client.post(f"/jobs/{job_id}/review-decisions", json={
"decisions": [{"review_item_id": "finding:AGENT-0001", "decision": "confirm"}],
})
assert response.status_code == 200
assert response.json()["progress"]["remaining"] == 1
assert backend.jobs.get_job(job_id)["status"] == "reviewing"
finally:
backend.jobs._jobs.pop(job_id, None)
def test_agent_job_needs_review_skips_notify(monkeypatch, tmp_path):
"""Carried finding from Task 4: an agent report needing review must not be emailed."""
sent = []
monkeypatch.setattr("backend.config.OUTPUT_DIR", str(tmp_path))
monkeypatch.setattr("backend.jobs.run_agent_pipeline", lambda pdf_path, **kw: {
"summary": {"agent_status": "needs_review"}, "conflicts": [],
})
monkeypatch.setattr("backend.jobs.send_conflict_report", lambda *a, **kw: sent.append((a, kw)))
monkeypatch.setattr(threading, "Thread", _SyncThread)
pdf = tmp_path / "upload.pdf"
pdf.write_bytes(b"%PDF-1.4 dummy")
job_id = backend.jobs.create_job(str(pdf), source_filename="set.pdf",
email="arch@example.com", pipeline_mode="agent")
try:
job = backend.jobs.get_job(job_id)
assert job["status"] == "needs_review"
assert job["report"]["summary"]["agent_status"] == "needs_review"
assert sent == []
finally:
backend.jobs._jobs.pop(job_id, None)
def test_classic_job_still_completes_and_notifies(monkeypatch, tmp_path):
"""Classic pipeline behavior is unchanged: done status + completion email."""
sent = []
monkeypatch.setattr("backend.config.OUTPUT_DIR", str(tmp_path))
monkeypatch.setattr("backend.jobs.run_pipeline", lambda pdf_path, **kw: {
"summary": {}, "conflicts": [],
})
monkeypatch.setattr("backend.jobs.send_conflict_report", lambda *a, **kw: sent.append((a, kw)))
monkeypatch.setattr(threading, "Thread", _SyncThread)
pdf = tmp_path / "upload.pdf"
pdf.write_bytes(b"%PDF-1.4 dummy")
job_id = backend.jobs.create_job(str(pdf), source_filename="set.pdf",
email="arch@example.com", pipeline_mode="classic")
try:
assert backend.jobs.get_job(job_id)["status"] == "done"
assert len(sent) == 1
finally:
backend.jobs._jobs.pop(job_id, None)
def test_job_status_includes_review_progress(monkeypatch, tmp_path):
"""GET /jobs/{id} surfaces report.summary.review for a needs_review job."""
review = {"required": 1, "completed": 0, "remaining": 1, "total": 1}
client = TestClient(app)
monkeypatch.setattr("backend.main.get_job", lambda job_id: {
"job_id": job_id, "status": "needs_review",
"report": {"summary": {"agent_status": "needs_review", "review": review}},
"out_dir": str(tmp_path),
})
response = client.get("/jobs/job1")
assert response.status_code == 200
body = response.json()
assert body["status"] == "needs_review"
assert body["report"]["summary"]["review"] == review
def test_review_response_includes_saved_decisions(monkeypatch, tmp_path):
"""GET /jobs/{id}/review also returns the decisions map for UI pre-population."""
store = ReviewStore(str(tmp_path))
store.write_queue([_queue_item("finding:AGENT-0001")])
client = TestClient(app)
monkeypatch.setattr("backend.main.get_job", lambda job_id: {"job_id": job_id, "status": "needs_review", "report": {"summary": {}}, "out_dir": str(tmp_path)})
post = client.post("/jobs/job1/review-decisions", json={"decisions": [
{"review_item_id": "finding:AGENT-0001", "decision": "confirm", "comment": "looks right"},
]})
assert post.status_code == 200
get = client.get("/jobs/job1/review")
assert get.status_code == 200
decisions = get.json()["decisions"]
assert decisions["finding:AGENT-0001"]["decision"] == "confirm"
assert decisions["finding:AGENT-0001"]["comment"] == "looks right"
def test_review_flow_via_disk_fallback(monkeypatch, tmp_path):
"""Smoke: a synthetic on-disk needs_review job served by the REAL get_job
(disk fallback), with decisions persisting across review GETs."""
job_id = "jobdisk"
out_dir = os.path.join(str(tmp_path), job_id)
os.makedirs(out_dir)
report = {
"source": "set.pdf",
"summary": {
"agent_status": "needs_review",
"review": {"required": 1, "completed": 0, "remaining": 1, "total": 1},
},
"conflicts": [],
}
with open(os.path.join(out_dir, "conflicts.json"), "w", encoding="utf-8") as f:
json.dump(report, f)
store = ReviewStore(out_dir)
store.write_queue([_queue_item("finding:AGENT-0001")])
monkeypatch.setattr("backend.config.OUTPUT_DIR", str(tmp_path))
client = TestClient(app)
job_response = client.get(f"/jobs/{job_id}")
assert job_response.status_code == 200
assert job_response.json()["report"]["summary"]["review"]["required"] == 1
review_response = client.get(f"/jobs/{job_id}/review")
assert review_response.status_code == 200
body = review_response.json()
assert [i["review_item_id"] for i in body["queue"]] == ["finding:AGENT-0001"]
assert body["progress"]["remaining"] == 1
assert body["decisions"] == {}
post = client.post(f"/jobs/{job_id}/review-decisions", json={"decisions": [
{"review_item_id": "finding:AGENT-0001", "decision": "reject",
"reason_code": "not_a_contradiction"},
]})
assert post.status_code == 200
assert post.json()["progress"]["remaining"] == 0
again = client.get(f"/jobs/{job_id}/review")
assert again.status_code == 200
saved = again.json()["decisions"]["finding:AGENT-0001"]
assert saved["decision"] == "reject"
assert saved["reason_code"] == "not_a_contradiction"
def _write_restart_job(tmp_path, job_id, email=None):
"""On-disk needs_review job artifacts, as a pre-restart run would leave
them: candidate report + memory snapshot + review queue (+ job.json)."""
out_dir = os.path.join(str(tmp_path), job_id)
os.makedirs(os.path.join(out_dir, "agent"))
report = {
"source": "set.pdf",
"generated_at": "2026-07-28T00:00:00+00:00",
"summary": {
"sheets_analyzed": 0, "disciplines": [], "assertions_extracted": 0,
"clusters_checked": 0, "conflicts_found": 0,
"by_severity": {"high": 0, "medium": 0, "low": 0}, "by_category": {},
"pipeline_mode": "agent", "agent_status": "needs_review",
"review": {"required": 1, "completed": 0, "remaining": 1, "total": 1},
},
"conflicts": [], "sheets": [],
"validated_issues": [{"issue_id": "AGENT-0001", "severity": "high"}],
"suppressed_issues": [], "rfis": [],
}
with open(os.path.join(out_dir, "conflicts.json"), "w", encoding="utf-8") as f:
json.dump(report, f)
with open(os.path.join(out_dir, "agent", "memory.json"), "w", encoding="utf-8") as f:
json.dump({}, f)
if email is not None:
with open(os.path.join(out_dir, "job.json"), "w", encoding="utf-8") as f:
json.dump({"job_id": job_id, "email": email,
"pipeline_mode": "agent", "source": "set.pdf"}, f)
store = ReviewStore(out_dir)
store.write_queue([_queue_item("finding:AGENT-0001")])
return out_dir
def test_restart_recovered_needs_review_job_finalizes(monkeypatch, tmp_path):
"""CRITICAL: after a restart, a needs_review job recovered from disk keeps
its status (not "done"), hydrates the in-memory registry, and the whole
decide -> finalize flow completes to done via the real get_job."""
job_id = "jobrestart"
_write_restart_job(tmp_path, job_id)
monkeypatch.setattr("backend.config.OUTPUT_DIR", str(tmp_path))
monkeypatch.setattr("backend.review.finalizer._draft_rfis", lambda kept: [])
monkeypatch.setattr("backend.jobs._notify", lambda *a, **kw: None)
monkeypatch.setattr(threading, "Thread", _SyncThread)
try:
# Real disk fallback: simulates a fresh post-restart process.
job = backend.jobs.get_job(job_id)
assert job["status"] == "needs_review"
assert job_id in backend.jobs._jobs # hydrated for _set() transitions
client = TestClient(app)
post = client.post(f"/jobs/{job_id}/review-decisions", json={"decisions": [
{"review_item_id": "finding:AGENT-0001", "decision": "confirm"},
]})
assert post.status_code == 200
fin = client.post(f"/jobs/{job_id}/finalize-review")
assert fin.status_code == 200
job = backend.jobs.get_job(job_id)
assert job["status"] == "done"
assert job["report"]["summary"]["agent_status"] == "complete"
finally:
backend.jobs._jobs.pop(job_id, None)
def test_restart_recovered_job_final_email_uses_job_json(monkeypatch, tmp_path):
"""CRITICAL: job.json (written at job start) restores the recipient email
after a restart, so finalization still fires the final report email."""
job_id = "jobemail"
_write_restart_job(tmp_path, job_id, email="arch@example.com")
monkeypatch.setattr("backend.config.OUTPUT_DIR", str(tmp_path))
monkeypatch.setattr("backend.review.finalizer._draft_rfis", lambda kept: [])
sent = []
monkeypatch.setattr("backend.jobs.send_conflict_report",
lambda email, report, **kw: sent.append(email))
monkeypatch.setattr(threading, "Thread", _SyncThread)
try:
job = backend.jobs.get_job(job_id)
assert job["status"] == "needs_review"
assert job["email"] == "arch@example.com"
client = TestClient(app)
post = client.post(f"/jobs/{job_id}/review-decisions", json={"decisions": [
{"review_item_id": "finding:AGENT-0001", "decision": "confirm"},
]})
assert post.status_code == 200
fin = client.post(f"/jobs/{job_id}/finalize-review")
assert fin.status_code == 200
assert backend.jobs.get_job(job_id)["status"] == "done"
assert sent == ["arch@example.com"]
finally:
backend.jobs._jobs.pop(job_id, None)
def test_review_decisions_409_for_non_review_job(monkeypatch, tmp_path):
"""Positive state guard: only needs_review/reviewing jobs accept decisions."""
client = TestClient(app)
monkeypatch.setattr("backend.main.get_job", lambda job_id: {
"job_id": job_id, "status": "done", "out_dir": str(tmp_path),
})
response = client.post("/jobs/job1/review-decisions", json={"decisions": [
{"review_item_id": "finding:AGENT-0001", "decision": "confirm"}]})
assert response.status_code == 409
assert "done" in response.json()["detail"]["detail"]
def test_review_queue_get_does_not_create_review_dir(monkeypatch, tmp_path):
"""The read-only GET endpoint must not create review/ dirs on read."""
client = TestClient(app)
monkeypatch.setattr("backend.main.get_job", lambda job_id: {
"job_id": job_id, "status": "needs_review",
"report": {"summary": {}}, "out_dir": str(tmp_path),
})
response = client.get("/jobs/job1/review")
assert response.status_code == 200
assert response.json()["queue"] == []
assert response.json()["decisions"] == {}
assert not os.path.exists(os.path.join(str(tmp_path), "review"))
def _write_finalizable_job(tmp_path, decisions):
"""Minimal review-mode artifacts: candidate report + queue + decisions."""
out_dir = str(tmp_path)
os.makedirs(os.path.join(out_dir, "agent"), exist_ok=True)
report = {
"source": "set.pdf",
"generated_at": "2026-07-28T00:00:00+00:00",
"summary": {
"sheets_analyzed": 0, "disciplines": [], "assertions_extracted": 0,
"clusters_checked": 0, "conflicts_found": 0,
"by_severity": {"high": 0, "medium": 0, "low": 0}, "by_category": {},
"pipeline_mode": "agent", "agent_status": "needs_review",
"review": {"required": 1, "completed": 0, "remaining": 1, "total": 1},
},
"conflicts": [], "sheets": [],
"validated_issues": [{"issue_id": "AGENT-0001", "severity": "high"}],
"suppressed_issues": [], "rfis": [],
}
with open(os.path.join(out_dir, "conflicts.json"), "w", encoding="utf-8") as f:
json.dump(report, f)
with open(os.path.join(out_dir, "agent", "memory.json"), "w", encoding="utf-8") as f:
json.dump({}, f)
store = ReviewStore(out_dir)
store.write_queue([_queue_item("finding:AGENT-0001")])
for decision in decisions:
store.append_decision(decision)
return out_dir
def test_finalize_review_409_while_undecided(monkeypatch, tmp_path):
out_dir = _write_finalizable_job(tmp_path, decisions=[])
client = TestClient(app)
monkeypatch.setattr("backend.main.get_job", lambda job_id: {
"job_id": job_id, "status": "needs_review", "out_dir": out_dir,
})
response = client.post("/jobs/job1/finalize-review")
assert response.status_code == 409
body = response.json()
assert body["detail"]["detail"] == "incomplete review"
assert body["detail"]["progress"]["remaining"] == 1
def test_finalize_review_404_for_unknown_job(monkeypatch, tmp_path):
monkeypatch.setattr("backend.config.OUTPUT_DIR", str(tmp_path))
client = TestClient(app)
assert client.post("/jobs/nope/finalize-review").status_code == 404
def test_finalize_review_409_when_already_done(monkeypatch, tmp_path):
client = TestClient(app)
monkeypatch.setattr("backend.main.get_job", lambda job_id: {
"job_id": job_id, "status": "done", "out_dir": str(tmp_path),
})
assert client.post("/jobs/job1/finalize-review").status_code == 409
def test_finalize_review_409_for_job_not_in_review(monkeypatch, tmp_path):
"""A running (or otherwise non-review) job must not be finalizable: no
finalization thread, no artifact clobbering, no final email."""
threads = []
notified = []
monkeypatch.setattr(threading, "Thread",
lambda *a, **kw: threads.append((a, kw)) or _SyncThread(*a, **kw))
monkeypatch.setattr("backend.jobs._notify",
lambda *a, **kw: notified.append(a))
monkeypatch.setattr("backend.main.get_job", lambda job_id: {
"job_id": job_id, "status": "running", "out_dir": str(tmp_path),
})
client = TestClient(app)
response = client.post("/jobs/job1/finalize-review")
assert response.status_code == 409
assert "running" in response.json()["detail"]["detail"]
assert threads == []
assert notified == []
assert not os.path.exists(os.path.join(str(tmp_path), "conflicts.json"))
def test_finalize_review_happy_path_notifies_once(monkeypatch, tmp_path):
out_dir = _write_finalizable_job(tmp_path, decisions=[{
"review_item_id": "finding:AGENT-0001", "decision": "confirm",
}])
notified = []
monkeypatch.setattr("backend.review.finalizer._draft_rfis", lambda kept: [])
monkeypatch.setattr("backend.jobs._notify",
lambda job_id, report, out_dir: notified.append(job_id))
monkeypatch.setattr(threading, "Thread", _SyncThread)
job_id = "jobfinalize"
backend.jobs._jobs[job_id] = {
"job_id": job_id, "status": "reviewing", "out_dir": out_dir,
"report": None, "email": "arch@example.com",
}
try:
client = TestClient(app)
response = client.post(f"/jobs/{job_id}/finalize-review")
assert response.status_code == 200
assert response.json() == {"status": "finalizing"}
job = backend.jobs.get_job(job_id)
assert job["status"] == "done"
assert job["report"]["summary"]["agent_status"] == "complete"
assert notified == [job_id]
for name in ("conflicts.json", "validated_issues.json",
"suppressed_issues.json", "rfis.json", "report.md"):
assert os.path.isfile(os.path.join(out_dir, name)), name
finally:
backend.jobs._jobs.pop(job_id, None)
+106
View File
@@ -0,0 +1,106 @@
"""Tests for the two-phase email flow: review-required notice, then final report."""
import threading
import backend.jobs
from backend.email_sender import send_review_required
class _SyncThread:
"""Drop-in threading.Thread replacement that runs the target inline."""
def __init__(self, target=None, args=(), **kwargs):
self._target = target
self._args = args
def start(self):
self._target(*self._args)
def test_review_required_email_skips_without_smtp(monkeypatch):
monkeypatch.setattr("backend.email_sender._smtp_ready", lambda: False)
assert send_review_required("user@example.com", {"source": "set.pdf", "summary": {}}, "http://localhost:8099/?job=abc") is False
def test_review_required_email_sends_with_smtp(monkeypatch):
"""With SMTP ready, the message goes out with recipient, review URL, and
the required-item count (0 when the report has no review summary)."""
sent = []
monkeypatch.setattr("backend.email_sender._smtp_ready", lambda: True)
monkeypatch.setattr("backend.email_sender._send",
lambda msg: sent.append(msg) or True)
review_url = "http://localhost:8099/?job=abc"
report = {"source": "set.pdf", "summary": {"review": {"required": 3}}}
assert send_review_required("user@example.com", report, review_url) is True
assert len(sent) == 1
msg = sent[0]
assert msg["To"] == "user@example.com"
assert "review" in msg["Subject"].lower()
body = msg.get_content()
assert "review" in body.lower()
assert review_url in body
assert "3" in body
# Missing review summary -> required count defaults to 0.
sent.clear()
assert send_review_required("user@example.com", {"source": "set.pdf", "summary": {}}, review_url) is True
assert "0" in sent[0].get_content()
def test_agent_needs_review_sends_review_email_not_report(monkeypatch, tmp_path):
"""An agent job entering needs_review emails the review-required notice
exactly once and never sends the final conflict report."""
review_emails = []
report_emails = []
monkeypatch.setattr("backend.config.OUTPUT_DIR", str(tmp_path))
monkeypatch.setattr("backend.jobs.run_agent_pipeline", lambda pdf_path, **kw: {
"summary": {"agent_status": "needs_review"}, "conflicts": [],
})
monkeypatch.setattr("backend.jobs.send_review_required",
lambda *a, **kw: review_emails.append((a, kw)))
monkeypatch.setattr("backend.jobs.send_conflict_report",
lambda *a, **kw: report_emails.append((a, kw)))
monkeypatch.setattr(threading, "Thread", _SyncThread)
pdf = tmp_path / "upload.pdf"
pdf.write_bytes(b"%PDF-1.4 dummy")
job_id = backend.jobs.create_job(str(pdf), source_filename="set.pdf",
email="arch@example.com", pipeline_mode="agent")
try:
job = backend.jobs.get_job(job_id)
assert job["status"] == "needs_review"
assert len(review_emails) == 1
args, _ = review_emails[0]
assert args[0] == "arch@example.com"
assert f"/?job={job_id}" in args[2]
assert report_emails == []
finally:
backend.jobs._jobs.pop(job_id, None)
def test_classic_job_sends_only_conflict_report(monkeypatch, tmp_path):
"""Classic pipeline is untouched: only the final report email fires."""
review_emails = []
report_emails = []
monkeypatch.setattr("backend.config.OUTPUT_DIR", str(tmp_path))
monkeypatch.setattr("backend.jobs.run_pipeline", lambda pdf_path, **kw: {
"summary": {}, "conflicts": [],
})
monkeypatch.setattr("backend.jobs.send_review_required",
lambda *a, **kw: review_emails.append((a, kw)))
monkeypatch.setattr("backend.jobs.send_conflict_report",
lambda *a, **kw: report_emails.append((a, kw)))
monkeypatch.setattr(threading, "Thread", _SyncThread)
pdf = tmp_path / "upload.pdf"
pdf.write_bytes(b"%PDF-1.4 dummy")
job_id = backend.jobs.create_job(str(pdf), source_filename="set.pdf",
email="arch@example.com", pipeline_mode="classic")
try:
assert backend.jobs.get_job(job_id)["status"] == "done"
assert len(report_emails) == 1
assert review_emails == []
finally:
backend.jobs._jobs.pop(job_id, None)
+87
View File
@@ -0,0 +1,87 @@
"""Feedback labels and aggregate metrics for human-review decisions."""
import json
import os
from datetime import datetime
from backend.review.feedback import decision_to_label, write_label
from backend.review.metrics import aggregate_labels
def test_aggregate_redacts_text_by_default():
labels = [{"decision": "reject", "reason_code": "missing_evidence", "comment": "secret", "payload": {"evidence": [{"source_text": "secret"}]}}]
summary = aggregate_labels(labels)
assert summary["reject"] == 1
assert "secret" not in str(summary)
def test_aggregate_include_text_embeds_labels():
labels = [{"decision": "reject", "reason_code": "missing_evidence", "comment": "secret"}]
summary = aggregate_labels(labels, include_text=True)
assert summary["labels"] == labels
def _queue_item() -> dict:
return {
"review_item_id": "finding:AGENT-0007",
"kind": "finding",
"blocking": True,
"reasons": ["high_severity"],
"payload": {
"issue_id": "AGENT-0007",
"source_stage": "conflict",
"category": "elevation_disagreement",
"severity": "high",
"confidence": "medium",
"location": "Room 204 / Level 2",
"disciplines": ["Architectural", "Mechanical"],
"sheets": ["A2.1", "M2.1"],
"drawing_type": "floor_plan",
},
}
def test_decision_to_label_builds_spec_shape():
decision = {"review_item_id": "finding:AGENT-0007", "decision": "reject",
"reason_code": "same_value_different_representation"}
job = {"job_id": "abc123", "pipeline_mode": "agent",
"report": {"summary": {"models_used": ["google/gemini-2.5-pro"]}}}
label = decision_to_label(_queue_item(), decision, job)
assert label["review_item_id"] == "finding:AGENT-0007"
assert label["job_id"] == "abc123"
assert label["pipeline_mode"] == "agent"
assert label["source_stage"] == "conflict"
assert label["category"] == "elevation_disagreement"
assert label["severity"] == "high"
assert label["confidence"] == "medium"
assert label["decision"] == "reject"
assert label["reason_code"] == "same_value_different_representation"
assert label["location"] == "Room 204 / Level 2"
assert label["disciplines"] == ["Architectural", "Mechanical"]
assert label["sheets"] == ["A2.1", "M2.1"]
assert label["drawing_type"] == "floor_plan"
assert label["models_used"] == ["google/gemini-2.5-pro"]
datetime.fromisoformat(label["created_at"])
def test_decision_to_label_degrades_on_missing_fields():
label = decision_to_label({"review_item_id": "finding:AGENT-0001"}, {}, {})
assert label["review_item_id"] == "finding:AGENT-0001"
assert label["job_id"] is None
assert label["decision"] is None
assert label["reason_code"] is None
assert label["category"] is None
assert label["source_stage"] is None
assert label["models_used"] == []
datetime.fromisoformat(label["created_at"])
def test_write_label_appends_json_lines(tmp_path):
label1 = {"review_item_id": "finding:AGENT-0001", "decision": "confirm"}
label2 = {"review_item_id": "finding:AGENT-0002", "decision": "reject"}
write_label(str(tmp_path), label1)
write_label(str(tmp_path), label2)
path = os.path.join(str(tmp_path), "review", "feedback_labels.jsonl")
with open(path, encoding="utf-8") as f:
lines = [json.loads(line) for line in f if line.strip()]
assert lines == [label1, label2]
+291
View File
@@ -0,0 +1,291 @@
"""Non-LLM tests for the review finalizer: decisions, reruns, final artifacts."""
import json
import os
import pytest
from backend.agents.base import AgentResult
from backend.review.finalizer import (
apply_decisions,
finalize_review,
rerun_clarified_scopes,
)
from backend.review.store import ReviewStore
def test_reject_suppresses_with_reason():
prioritized = [{"issue_id": "AGENT-0001", "severity": "high"}]
decisions = {"finding:AGENT-0001": {"decision": "reject", "reason_code": "duplicate"}}
kept, suppressed = apply_decisions(prioritized, decisions)
assert kept == []
assert suppressed[0]["review_state"] == "rejected"
assert suppressed[0]["reason_code"] == "duplicate"
def test_unsure_is_kept_but_flagged():
prioritized = [{"issue_id": "AGENT-0002", "severity": "medium"}]
decisions = {"finding:AGENT-0002": {"decision": "unsure"}}
kept, suppressed = apply_decisions(prioritized, decisions)
assert kept[0]["review_state"] == "unsure"
assert suppressed == []
def _write_job(out_dir, prioritized, queue, decisions=None, memory=None):
"""Hand-written review-mode artifacts (conflicts.json + agent/memory.json)."""
os.makedirs(os.path.join(out_dir, "agent"), exist_ok=True)
report = {
"source": "set.pdf",
"generated_at": "2026-07-28T00:00:00+00:00",
"summary": {
"sheets_analyzed": 0,
"disciplines": [],
"assertions_extracted": 0,
"clusters_checked": 0,
"conflicts_found": 0,
"by_severity": {"high": 0, "medium": 0, "low": 0},
"by_category": {},
"pipeline_mode": "agent",
"agent_status": "needs_review",
"review": {"required": 1, "completed": 0, "remaining": 1, "total": 1},
"by_stage": {"validated": len(prioritized), "rfis": 0},
},
"conflicts": [],
"sheets": [],
"validated_issues": prioritized,
"suppressed_issues": [],
"rfis": [],
}
with open(os.path.join(out_dir, "conflicts.json"), "w", encoding="utf-8") as f:
json.dump(report, f)
with open(os.path.join(out_dir, "agent", "memory.json"), "w", encoding="utf-8") as f:
json.dump(memory or {}, f)
store = ReviewStore(out_dir)
store.write_queue(queue)
for decision in decisions or []:
store.append_decision(decision)
def _blocking_item(issue_id):
return {"review_item_id": f"finding:{issue_id}", "kind": "finding",
"blocking": True, "reasons": ["high_severity"], "payload": {}}
def test_finalize_confirm_keeps_confirmed(monkeypatch, tmp_path):
monkeypatch.setattr("backend.review.finalizer._draft_rfis", lambda kept: [])
_write_job(
str(tmp_path),
prioritized=[{"issue_id": "AGENT-0001", "severity": "high"}],
queue=[_blocking_item("AGENT-0001")],
decisions=[{"review_item_id": "finding:AGENT-0001", "decision": "confirm"}],
)
report = finalize_review("job1", str(tmp_path))
assert report["validated_issues"][0]["review_state"] == "confirmed"
assert report["suppressed_issues"] == []
assert report["summary"]["agent_status"] == "complete"
def test_finalize_no_decision_keeps_unreviewed(monkeypatch, tmp_path):
"""Non-blocking (audit) items don't need a decision; issue stays unreviewed."""
monkeypatch.setattr("backend.review.finalizer._draft_rfis", lambda kept: [])
item = {**_blocking_item("AGENT-0001"), "blocking": False, "kind": "audit_finding"}
_write_job(
str(tmp_path),
prioritized=[{"issue_id": "AGENT-0001", "severity": "medium"}],
queue=[item],
)
report = finalize_review("job1", str(tmp_path))
assert report["validated_issues"][0]["review_state"] == "unreviewed"
def test_finalize_clarification_replacement_marked_clarified(monkeypatch, tmp_path):
monkeypatch.setattr("backend.review.finalizer._draft_rfis", lambda kept: [])
replacement = {"issue_id": "AGENT-0001-R1", "severity": "medium",
"clarification_of": "AGENT-0001"}
monkeypatch.setattr(
"backend.review.finalizer.rerun_clarified_scopes",
lambda snapshot, decisions, prioritized=None: [replacement],
)
_write_job(
str(tmp_path),
prioritized=[{"issue_id": "AGENT-0001", "severity": "high"}],
queue=[_blocking_item("AGENT-0001")],
decisions=[{"review_item_id": "finding:AGENT-0001",
"decision": "needs_clarification",
"clarification_answer": "Ceiling is 9'-0\" AFF."}],
)
report = finalize_review("job1", str(tmp_path))
kept = report["validated_issues"]
assert [issue["issue_id"] for issue in kept] == ["AGENT-0001-R1"]
assert kept[0]["review_state"] == "clarified"
def test_finalize_failed_clarification_flagged(monkeypatch, tmp_path):
monkeypatch.setattr("backend.review.finalizer._draft_rfis", lambda kept: [])
monkeypatch.setattr(
"backend.review.finalizer.rerun_clarified_scopes",
lambda snapshot, decisions, prioritized=None: [],
)
_write_job(
str(tmp_path),
prioritized=[{"issue_id": "AGENT-0001", "severity": "high"}],
queue=[_blocking_item("AGENT-0001")],
decisions=[{"review_item_id": "finding:AGENT-0001",
"decision": "needs_clarification",
"clarification_answer": "Ceiling is 9'-0\" AFF."}],
)
report = finalize_review("job1", str(tmp_path))
assert report["validated_issues"][0]["review_state"] == "clarification_failed"
def test_finalize_incomplete_review_raises(monkeypatch, tmp_path):
monkeypatch.setattr("backend.review.finalizer._draft_rfis", lambda kept: [])
_write_job(
str(tmp_path),
prioritized=[{"issue_id": "AGENT-0001", "severity": "high"}],
queue=[_blocking_item("AGENT-0001")],
)
with pytest.raises(ValueError, match="incomplete review"):
finalize_review("job1", str(tmp_path))
def test_finalize_writes_final_artifacts(monkeypatch, tmp_path):
monkeypatch.setattr("backend.review.finalizer._draft_rfis",
lambda kept: [{"issue_id": kept[0]["issue_id"], "question": "?"}])
_write_job(
str(tmp_path),
prioritized=[{"issue_id": "AGENT-0001", "severity": "high"}],
queue=[_blocking_item("AGENT-0001")],
decisions=[{"review_item_id": "finding:AGENT-0001", "decision": "confirm"}],
)
report = finalize_review("job1", str(tmp_path))
assert report["summary"]["by_stage"]["validated"] == 1
assert report["summary"]["by_stage"]["rfis"] == 1
for name in ("conflicts.json", "validated_issues.json",
"suppressed_issues.json", "rfis.json", "report.md"):
assert os.path.isfile(os.path.join(str(tmp_path), name)), name
with open(os.path.join(str(tmp_path), "validated_issues.json"), encoding="utf-8") as f:
assert json.load(f)[0]["review_state"] == "confirmed"
def test_finalize_reject_rebuilds_conflicts_and_counts(monkeypatch, tmp_path):
"""Rejected conflict-stage findings must not survive into the final
report's conflicts / headline counts; suppressed_issues keeps them."""
monkeypatch.setattr("backend.review.finalizer._draft_rfis", lambda kept: [])
kept_finding = {
"issue_id": "AGENT-0001", "source_stage": "conflict",
"category": "note_or_spec_contradiction", "severity": "high",
"location": "Grid A", "disciplines": ["A", "S"], "sheets": ["A-1"],
"description": "kept finding", "evidence": [],
"recommended_resolution": "fix", "confidence": "high",
}
rejected_finding = {
**kept_finding, "issue_id": "AGENT-0002", "severity": "medium",
"description": "rejected finding",
}
_write_job(
str(tmp_path),
prioritized=[kept_finding, rejected_finding],
queue=[_blocking_item("AGENT-0001"), _blocking_item("AGENT-0002")],
decisions=[
{"review_item_id": "finding:AGENT-0001", "decision": "confirm"},
{"review_item_id": "finding:AGENT-0002", "decision": "reject",
"reason_code": "not_a_contradiction"},
],
)
# Simulate the pre-review candidate values the finalizer must overwrite.
candidate_path = os.path.join(str(tmp_path), "conflicts.json")
with open(candidate_path, encoding="utf-8") as f:
candidate = json.load(f)
candidate["conflicts"] = [{"description": "kept finding", "severity": "high",
"category": "note_or_spec_contradiction"},
{"description": "rejected finding", "severity": "medium",
"category": "note_or_spec_contradiction"}]
candidate["summary"]["conflicts_found"] = 2
candidate["summary"]["by_severity"] = {"high": 1, "medium": 1, "low": 0}
candidate["summary"]["by_category"] = {"note_or_spec_contradiction": 2}
with open(candidate_path, "w", encoding="utf-8") as f:
json.dump(candidate, f)
report = finalize_review("job1", str(tmp_path))
assert [c["description"] for c in report["conflicts"]] == ["kept finding"]
assert report["summary"]["conflicts_found"] == 1
assert report["summary"]["by_severity"] == {"high": 1, "medium": 0, "low": 0}
assert report["summary"]["by_category"] == {"note_or_spec_contradiction": 1}
suppressed = report["suppressed_issues"]
assert [s["issue_id"] for s in suppressed] == ["AGENT-0002"]
assert suppressed[0]["review_state"] == "rejected"
assert suppressed[0]["reason_code"] == "not_a_contradiction"
with open(os.path.join(str(tmp_path), "report.md"), encoding="utf-8") as f:
assert "rejected finding" not in f.read()
def test_rerun_missing_cluster_degrades_to_analysis_gap():
snapshot = {"findings": [{"issue_id": "AGENT-0001", "scope_id": "conflict:link:1"}],
"clusters": []}
decisions = {"finding:AGENT-0001": {
"decision": "needs_clarification", "clarification_answer": "9'-0\" AFF"}}
findings = rerun_clarified_scopes(snapshot, decisions)
assert len(findings) == 1
assert findings[0]["category"] == "analysis_gap"
assert findings[0]["source_stage"] == "qaqc"
assert findings[0]["severity"] == "low"
assert findings[0]["confidence"] == "high"
def test_rerun_non_conflict_scope_noted_as_analysis_gap():
"""v1 only reruns conflict scopes; other scopes get a visible gap, no raise."""
snapshot = {"findings": [{"issue_id": "AGENT-0002", "scope_id": "code: egress"}],
"clusters": []}
decisions = {"finding:AGENT-0002": {
"decision": "needs_clarification", "clarification_answer": "Corridor is 44 in."}}
findings = rerun_clarified_scopes(snapshot, decisions)
assert len(findings) == 1
assert findings[0]["category"] == "analysis_gap"
def test_rerun_successful_scope_prepends_clarification_and_tags(monkeypatch):
"""Real rerun path (non-LLM): cluster lookup, pseudo-assertion injection,
and clarification_of tagging through the real rerun_clarified_scopes."""
captured = {}
class FakeCritic:
name = "conflict_critic"
def __init__(self, usage):
pass
def run(self, scope):
captured["scope"] = scope
return AgentResult(
scope_id=scope.scope_id,
artifacts=[{"issue_id": "AGENT-0001-R1", "severity": "medium"}],
)
monkeypatch.setattr("backend.review.finalizer.ConflictCriticAgent", FakeCritic)
snapshot = {
"findings": [{"issue_id": "AGENT-0001", "scope_id": "conflict:link:1"}],
"clusters": [{"key": "link:1",
"assertions": [{"attribute": "height", "value": "10'-0\""}]}],
}
decisions = {"finding:AGENT-0001": {
"decision": "needs_clarification", "clarification_answer": "9'-0\" AFF"}}
findings = rerun_clarified_scopes(snapshot, decisions)
assert len(findings) == 1
assert findings[0]["issue_id"] == "AGENT-0001-R1"
assert findings[0]["clarification_of"] == "AGENT-0001"
payload = captured["scope"].payload
assert payload["page_to_b64"] == {}
assertions = payload["cluster"]["assertions"]
# Prepended at index 0 so front-truncation can't drop the clarification.
assert assertions[0]["discipline"] == "Reviewer"
assert assertions[0]["attribute"] == "clarification"
assert assertions[0]["value"] == "9'-0\" AFF"
assert assertions[1]["attribute"] == "height"
def test_rerun_ignores_other_decisions():
decisions = {"finding:AGENT-0001": {"decision": "confirm"}}
assert rerun_clarified_scopes({}, decisions) == []
+28
View File
@@ -0,0 +1,28 @@
from backend.review.gate import build_review_queue
def test_gate_marks_blocking_and_audit_items():
memory = {"clusters": [{"key": "room:101", "location": "Room 101", "assertions": [{"id": "a1"}, {"id": "a2"}]}], "findings": []}
prioritized = [
{"issue_id": "AGENT-0001", "severity": "high", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"},
{"issue_id": "AGENT-0002", "severity": "low", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"},
]
queue = build_review_queue(memory, prioritized, [])
by_id = {item["review_item_id"]: item for item in queue}
assert by_id["finding:AGENT-0001"]["blocking"] is True
assert by_id["finding:AGENT-0002"]["blocking"] is False
assert any(item["kind"] == "clean_cluster" for item in queue)
def test_gate_limit_caps_clean_cluster_items():
memory = {
"clusters": [
{"key": f"room:{index}", "assertions": [{"id": "a"}, {"id": "b"}]}
for index in range(3)
],
"findings": [],
}
queue = build_review_queue(memory, [], [], limit=1)
clean_items = [item for item in queue if item["kind"] == "clean_cluster"]
assert len(clean_items) == 1
assert clean_items[0]["review_item_id"] == "clean_cluster:room:0"
+130
View File
@@ -0,0 +1,130 @@
from backend import config
from backend.review.policy import build_audit_sample, requires_review
from backend.review.schemas import validate_decision
def test_high_severity_requires_review():
issue = {"severity": "high", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"}
assert "severity_high" in requires_review(issue)
def test_low_confidence_requires_review():
issue = {"severity": "low", "confidence": "low", "category": "note_or_spec_contradiction", "source_stage": "conflict"}
assert "confidence_low" in requires_review(issue)
def test_sensitive_code_category_requires_review():
issue = {"severity": "medium", "confidence": "high", "category": "egress", "source_stage": "code"}
assert "sensitive_category" in requires_review(issue)
def test_medium_high_confidence_note_does_not_require_review():
issue = {"severity": "medium", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"}
assert requires_review(issue) == []
def test_build_audit_sample_returns_clean_cluster_spot_check():
memory = {
"clusters": [
{
"key": "room:101",
"location": "Room 101",
"assertions": [{"id": "a1"}, {"id": "a2"}],
}
],
"findings": [],
}
prioritized = []
items = build_audit_sample(memory, prioritized)
assert len(items) == 1
item = items[0]
assert item["kind"] == "clean_cluster"
assert item["blocking"] is False
assert item["review_item_id"] == "clean_cluster:room:101"
def test_build_audit_sample_strips_base64_from_assertions():
memory = {
"clusters": [
{
"key": "room:101",
"assertions": [
{"id": "a1", "base64": "AAAA"},
{"id": "a2", "base64": "BBBB"},
],
}
],
"findings": [],
}
items = build_audit_sample(memory, [])
assert len(items) == 1
assertions = items[0]["payload"]["assertions"]
assert assertions == [{"id": "a1"}, {"id": "a2"}]
assert all("base64" not in assertion for assertion in assertions)
def test_build_audit_sample_respects_limit():
memory = {
"clusters": [
{"key": f"room:{index}", "assertions": [{"id": "a"}, {"id": "b"}]}
for index in range(4)
],
"findings": [],
}
items = build_audit_sample(memory, [], limit=2)
assert len(items) == 2
assert [item["review_item_id"] for item in items] == [
"clean_cluster:room:0",
"clean_cluster:room:1",
]
def test_build_audit_sample_excludes_implicated_clusters():
memory = {
"clusters": [
{"key": "room:101", "assertions": [{"id": "a1"}, {"id": "a2"}]},
{"key": "room:102", "assertions": [{"id": "b1"}, {"id": "b2"}]},
],
"findings": [{"scope_id": "conflict:room:101"}],
}
items = build_audit_sample(memory, [])
assert [item["review_item_id"] for item in items] == ["clean_cluster:room:102"]
def test_validate_decision_confirm_without_reason_code():
result = validate_decision({"review_item_id": "x", "decision": "confirm"})
assert result is not None
assert result["decision"] == "confirm"
assert result["reason_code"] is None
def test_validate_decision_reject_with_valid_reason_code():
result = validate_decision({"decision": "reject", "reason_code": "duplicate"})
assert result is not None
assert result["reason_code"] == "duplicate"
def test_validate_decision_reject_with_missing_reason_code_returns_none():
assert validate_decision({"decision": "reject"}) is None
def test_validate_decision_reject_with_invalid_reason_code_returns_none():
assert validate_decision({"decision": "reject", "reason_code": "bogus"}) is None
def test_validate_decision_unknown_decision_returns_none():
assert validate_decision({"decision": "approve"}) is None
def test_validate_decision_non_dict_returns_none():
assert validate_decision("confirm") is None
def test_validate_decision_invalid_reason_code_on_non_reject_returns_none():
assert validate_decision({"decision": "confirm", "reason_code": "bogus"}) is None
def test_review_defaults():
assert config.AGENT_REQUIRE_REVIEW is True
assert config.AGENT_REVIEW_AUDIT_SAMPLE == 5
assert config.REVIEW_AGGREGATE_INCLUDE_TEXT is False
+24
View File
@@ -0,0 +1,24 @@
import json
from backend.review.store import ReviewStore
def test_queue_and_decisions_round_trip(tmp_path):
store = ReviewStore(str(tmp_path))
queue = [{"review_item_id": "finding:1", "blocking": True}]
store.write_queue(queue)
assert store.read_queue() == queue
store.append_decision({"review_item_id": "finding:1", "decision": "confirm"})
assert store.read_decisions()["finding:1"]["decision"] == "confirm"
def test_progress_counts_required_items(tmp_path):
store = ReviewStore(str(tmp_path))
queue = [
{"review_item_id": "a", "blocking": True},
{"review_item_id": "b", "blocking": False},
]
store.write_queue(queue)
store.append_decision({"review_item_id": "a", "decision": "confirm"})
progress = store.progress(queue)
assert progress["required"] == 1
assert progress["completed"] == 1
+44
View File
@@ -0,0 +1,44 @@
from backend import config
from backend.llm import _resolve_backend, set_model_overrides, set_text_backend
def teardown_function():
set_model_overrides(None, None)
set_text_backend(False)
def test_vision_override_wins_for_vision_only():
set_model_overrides(vision="openai/gpt-4o", text=None)
assert _resolve_backend(has_images=True, model_override=None)["model"] == "openai/gpt-4o"
assert _resolve_backend(has_images=False, model_override=None)["model"] == config.TEXT_MODEL
def test_text_override_wins_for_text_only():
set_model_overrides(vision=None, text="anthropic/claude-sonnet-4")
assert _resolve_backend(has_images=False, model_override=None)["model"] == "anthropic/claude-sonnet-4"
assert _resolve_backend(has_images=True, model_override=None)["model"] == config.MODEL
def test_override_beats_per_call_model_arg():
set_model_overrides(vision="openai/gpt-4o", text="openai/gpt-4o-mini")
# Agents pass their AGENT_*_MODEL per call; the user's job pick wins.
assert _resolve_backend(has_images=True, model_override="other/model")["model"] == "openai/gpt-4o"
assert _resolve_backend(has_images=False, model_override="other/model")["model"] == "openai/gpt-4o-mini"
def test_no_override_keeps_defaults():
set_model_overrides(None, None)
assert _resolve_backend(has_images=True, model_override=None)["model"] == config.MODEL
assert _resolve_backend(has_images=False, model_override=None)["model"] == config.TEXT_MODEL
def test_ui_picks_never_name_the_local_model(monkeypatch):
"""Hybrid runs keep LOCAL_TEXT_MODEL; OpenRouter picks must not leak into
the local endpoint (a vLLM server won't serve OpenRouter model ids)."""
monkeypatch.setattr(config, "LOCAL_BASE_URL", "http://localhost:8000/v1")
monkeypatch.setattr(config, "LOCAL_TEXT_MODEL", "qwen/local-instruct")
set_text_backend(True)
set_model_overrides(vision="openai/gpt-4o", text="anthropic/claude-sonnet-4")
be = _resolve_backend(has_images=False, model_override=None)
assert be["local"] is True
assert be["model"] == "qwen/local-instruct"