Add required human review gate to the Agent pipeline #1
No files matched your search
@@ -212,10 +212,54 @@ From the CLI, `--no-review` bypasses the gate for that run (it overrides
|
|||||||
python cli/run_check.py samples/your_set.pdf --mode agent --no-review --out out/agent-run
|
python cli/run_check.py samples/your_set.pdf --mode agent --no-review --out out/agent-run
|
||||||
```
|
```
|
||||||
|
|
||||||
|
### Asking the run why: review chat
|
||||||
|
|
||||||
|
Each item on the review screen has an **Ask about this finding** panel, and the
|
||||||
|
screen carries one **Ask about this run** panel for questions that are not about
|
||||||
|
a single finding. The chat answers from the job's own artifacts — the finding's
|
||||||
|
evidence, the cluster it came from, the raw per-sheet extraction, the
|
||||||
|
verification verdict, the Brain's merge decision, the sheet index, the cover-index
|
||||||
|
reconciliation, and matching `job.log` lines.
|
||||||
|
|
||||||
|
```
|
||||||
|
"why does it think the AC unit is mounted on the ground?" -> item scope
|
||||||
|
"why didn't it pick up on the Civil set?" -> run scope
|
||||||
|
```
|
||||||
|
|
||||||
|
The chat is **read-only**. It cannot change a finding, a severity, a decision,
|
||||||
|
or the report, and the prompt forbids it from proposing code or config changes —
|
||||||
|
the radio buttons remain the only thing that alters review state. When the
|
||||||
|
artifacts do not contain the answer, it says so and names what is missing rather
|
||||||
|
than guessing.
|
||||||
|
|
||||||
|
Every turn is logged twice:
|
||||||
|
|
||||||
|
- `outputs/<job_id>/review/chat_log.jsonl` — the auditable record: the issue as
|
||||||
|
it stood when asked about, the question, the answer, the determinations, and
|
||||||
|
the evidence quoted. Readable as a transcript at
|
||||||
|
`GET /jobs/{id}/review-chat/log`.
|
||||||
|
- `REVIEW_FEEDBACK_DIR/chat_turns.jsonl` — the cross-job roll-up, alongside
|
||||||
|
`decisions.jsonl`. When a reviewer corrects a misidentification in
|
||||||
|
conversation ("that is not a floor drain, it is a power floor box"), the
|
||||||
|
correction is captured as `suggested_category_correction` rather than dying in
|
||||||
|
free text. Nothing reads this store yet; writing it is what makes priming a
|
||||||
|
future run on past corrections possible.
|
||||||
|
|
||||||
|
| Key | Default | Effect |
|
||||||
|
|-----|---------|--------|
|
||||||
|
| `ENABLE_REVIEW_CHAT` | `true` | `false` = the chat endpoints refuse and the panels stay empty |
|
||||||
|
| `REVIEW_CHAT_MODEL` | `TEXT_MODEL` | Model for chat answers |
|
||||||
|
| `REVIEW_CHAT_MAX_TOKENS` | `4096` | Answer budget |
|
||||||
|
| `REVIEW_CHAT_HISTORY_TURNS` | `6` | Prior turns replayed into a thread's prompt |
|
||||||
|
| `REVIEW_CHAT_LOG_LINES` | `40` | Max `job.log` lines pulled into the context bundle |
|
||||||
|
| `REVIEW_FEEDBACK_DIR` | `backend/outputs/_feedback` | Cross-job decision + chat feedback store |
|
||||||
|
|
||||||
**Deployment note:** the review endpoints (`/jobs/{id}/review-decisions`,
|
**Deployment note:** the review endpoints (`/jobs/{id}/review-decisions`,
|
||||||
`/jobs/{id}/finalize-review`) are **state-changing and sensitive** — they accept
|
`/jobs/{id}/finalize-review`) are **state-changing and sensitive** — they accept
|
||||||
human decisions that alter the final report. Do **not** expose the UI/API
|
human decisions that alter the final report. `/jobs/{id}/review-chat` does not
|
||||||
publicly without reverse-proxy auth or a shared access token in front of it.
|
change review state, but it does spend model budget and returns drawing
|
||||||
|
evidence. Do **not** expose the UI/API publicly without reverse-proxy auth or a
|
||||||
|
shared access token in front of it.
|
||||||
|
|
||||||
Web UI (upload + view):
|
Web UI (upload + view):
|
||||||
|
|
||||||
|
|||||||
@@ -51,6 +51,21 @@ AGENT_REVIEW_AUDIT_SAMPLE=5
|
|||||||
# Allow future cross-job review-feedback aggregation to include source_text/images/comments
|
# Allow future cross-job review-feedback aggregation to include source_text/images/comments
|
||||||
REVIEW_AGGREGATE_INCLUDE_TEXT=false
|
REVIEW_AGGREGATE_INCLUDE_TEXT=false
|
||||||
|
|
||||||
|
# Review chat: read-only Q&A about findings and coverage on the review screen.
|
||||||
|
# It explains what the run did from the job's artifacts; it never changes a
|
||||||
|
# finding, a decision, or the report.
|
||||||
|
ENABLE_REVIEW_CHAT=true
|
||||||
|
# Model for chat answers (blank inherits TEXT_MODEL)
|
||||||
|
REVIEW_CHAT_MODEL=
|
||||||
|
REVIEW_CHAT_MAX_TOKENS=4096
|
||||||
|
# Prior turns replayed into a thread's prompt
|
||||||
|
REVIEW_CHAT_HISTORY_TURNS=6
|
||||||
|
# Max job.log lines searched into the chat's context bundle
|
||||||
|
REVIEW_CHAT_LOG_LINES=40
|
||||||
|
REVIEW_CHAT_MAX_QUESTION_CHARS=2000
|
||||||
|
# Cross-job store for review decisions + chat turns (blank = backend/outputs/_feedback)
|
||||||
|
REVIEW_FEEDBACK_DIR=
|
||||||
|
|
||||||
# Pipeline tuning
|
# Pipeline tuning
|
||||||
PDF_DPI=100
|
PDF_DPI=100
|
||||||
MAX_PAGES=60
|
MAX_PAGES=60
|
||||||
|
|||||||
@@ -129,6 +129,26 @@ AGENT_REVIEW_AUDIT_SAMPLE = int(os.getenv("AGENT_REVIEW_AUDIT_SAMPLE", "5"))
|
|||||||
# NOTE: currently unwired - reserved for future cross-job aggregation tooling.
|
# NOTE: currently unwired - reserved for future cross-job aggregation tooling.
|
||||||
REVIEW_AGGREGATE_INCLUDE_TEXT = os.getenv("REVIEW_AGGREGATE_INCLUDE_TEXT", "false").strip().lower() in ("1", "true", "yes")
|
REVIEW_AGGREGATE_INCLUDE_TEXT = os.getenv("REVIEW_AGGREGATE_INCLUDE_TEXT", "false").strip().lower() in ("1", "true", "yes")
|
||||||
|
|
||||||
|
# -- Review chat (ask-the-run Q&A on the review screen) --------------
|
||||||
|
# A read-only explainer: it answers "why did the run decide X?" from the job's
|
||||||
|
# own artifacts and never mutates findings, decisions, or code. Every turn is
|
||||||
|
# appended to <out_dir>/review/chat_log.jsonl AND to the cross-job feedback
|
||||||
|
# store (REVIEW_FEEDBACK_DIR) so answers are available to future prompt priors.
|
||||||
|
# HISTORY_TURNS caps how much of a thread is replayed into the prompt;
|
||||||
|
# LOG_LINES caps how many job.log lines are searched into the context bundle.
|
||||||
|
ENABLE_REVIEW_CHAT = _flag("ENABLE_REVIEW_CHAT", "true")
|
||||||
|
REVIEW_CHAT_MODEL = os.getenv("REVIEW_CHAT_MODEL", "") or TEXT_MODEL
|
||||||
|
REVIEW_CHAT_MAX_TOKENS = int(os.getenv("REVIEW_CHAT_MAX_TOKENS", "4096"))
|
||||||
|
REVIEW_CHAT_HISTORY_TURNS = int(os.getenv("REVIEW_CHAT_HISTORY_TURNS", "6"))
|
||||||
|
REVIEW_CHAT_LOG_LINES = int(os.getenv("REVIEW_CHAT_LOG_LINES", "40"))
|
||||||
|
REVIEW_CHAT_MAX_QUESTION_CHARS = int(os.getenv("REVIEW_CHAT_MAX_QUESTION_CHARS", "2000"))
|
||||||
|
# Cross-job feedback store: where review decisions and chat turns accumulate so
|
||||||
|
# a future run can be primed with "what humans corrected last time". Job-local
|
||||||
|
# artifacts stay the source of truth; this is the append-only roll-up.
|
||||||
|
# (OUTPUT_DIR is defined further down; keep this in sync with it.)
|
||||||
|
REVIEW_FEEDBACK_DIR = os.getenv("REVIEW_FEEDBACK_DIR", "") or os.path.join(
|
||||||
|
_BASE_DIR, "outputs", "_feedback")
|
||||||
|
|
||||||
# -- Hybrid (local text LLM) ----------------------------------------
|
# -- Hybrid (local text LLM) ----------------------------------------
|
||||||
# Optional OpenAI-compatible local endpoint (e.g. a vLLM box) for the text-only
|
# Optional OpenAI-compatible local endpoint (e.g. a vLLM box) for the text-only
|
||||||
# QAQC stages. Vision stages ALWAYS use OpenRouter. The user picks hybrid per
|
# QAQC stages. Vision stages ALWAYS use OpenRouter. The user picks hybrid per
|
||||||
|
|||||||
@@ -23,6 +23,7 @@ import backend.jobs
|
|||||||
from backend import config, llm
|
from backend import config, llm
|
||||||
from backend.jobs import PIPELINE_MODES, create_job, get_job, _set
|
from backend.jobs import PIPELINE_MODES, create_job, get_job, _set
|
||||||
from backend.pipeline.pdf_processor import render_page_jpeg
|
from backend.pipeline.pdf_processor import render_page_jpeg
|
||||||
|
from backend.review import chat as review_chat
|
||||||
from backend.review.feedback import decision_to_label, write_label
|
from backend.review.feedback import decision_to_label, write_label
|
||||||
from backend.review.finalizer import finalize_review
|
from backend.review.finalizer import finalize_review
|
||||||
from backend.review.store import ReviewStore
|
from backend.review.store import ReviewStore
|
||||||
@@ -252,6 +253,67 @@ def finalize_review_endpoint(job_id: str):
|
|||||||
return {"status": "finalizing"}
|
return {"status": "finalizing"}
|
||||||
|
|
||||||
|
|
||||||
|
# Statuses in which the review chat may be used. The chat is read-only, so it
|
||||||
|
# stays available after finalization - a reviewer often asks "why did it say
|
||||||
|
# that?" about a report they have already sent.
|
||||||
|
_CHAT_STATES = ("needs_review", "reviewing", "finalizing", "done", "finalization_error")
|
||||||
|
|
||||||
|
|
||||||
|
def _chat_out_dir(job_id: str) -> str:
|
||||||
|
"""Resolve a job's output dir for a chat request, or raise an HTTP error."""
|
||||||
|
job = get_job(job_id)
|
||||||
|
if not job:
|
||||||
|
raise HTTPException(status_code=404, detail="Job not found")
|
||||||
|
if job.get("status") not in _CHAT_STATES:
|
||||||
|
raise HTTPException(status_code=409, detail={
|
||||||
|
"detail": f"review chat is not available for a job in status {job.get('status')}",
|
||||||
|
})
|
||||||
|
return job.get("out_dir") or os.path.join(config.OUTPUT_DIR, job_id)
|
||||||
|
|
||||||
|
|
||||||
|
@app.post("/jobs/{job_id}/review-chat")
|
||||||
|
def review_chat_ask(job_id: str, payload: dict):
|
||||||
|
"""Ask one question about a finding, or about the run as a whole.
|
||||||
|
|
||||||
|
Read-only: this answers from the job's artifacts and appends to the chat
|
||||||
|
log. It never changes a finding, a decision, or the report.
|
||||||
|
"""
|
||||||
|
out_dir = _chat_out_dir(job_id)
|
||||||
|
store = ReviewStore(out_dir, create=False)
|
||||||
|
try:
|
||||||
|
turn = review_chat.ask(
|
||||||
|
job_id=job_id,
|
||||||
|
out_dir=out_dir,
|
||||||
|
question=payload.get("question"),
|
||||||
|
review_item_id=payload.get("review_item_id") or None,
|
||||||
|
queue=store.read_queue(),
|
||||||
|
decisions=store.read_decisions(),
|
||||||
|
)
|
||||||
|
except review_chat.ChatError as e:
|
||||||
|
raise HTTPException(status_code=422, detail=str(e))
|
||||||
|
except Exception as e:
|
||||||
|
# A failed model call is an upstream problem, not a bad request; the
|
||||||
|
# review screen shows it inline and the reviewer can retry.
|
||||||
|
raise HTTPException(status_code=502, detail=f"review chat failed: {e}")
|
||||||
|
return {"turn": turn}
|
||||||
|
|
||||||
|
|
||||||
|
@app.get("/jobs/{job_id}/review-chat")
|
||||||
|
def review_chat_history(job_id: str, review_item_id: Optional[str] = None):
|
||||||
|
"""Logged chat turns, oldest first. Without review_item_id, all threads."""
|
||||||
|
out_dir = _chat_out_dir(job_id)
|
||||||
|
turns = review_chat.read_log(out_dir, review_item_id=review_item_id)
|
||||||
|
return {"turns": turns, "enabled": config.ENABLE_REVIEW_CHAT}
|
||||||
|
|
||||||
|
|
||||||
|
@app.get("/jobs/{job_id}/review-chat/log")
|
||||||
|
def review_chat_log(job_id: str):
|
||||||
|
"""The chat log as a readable transcript: issue, questions, findings."""
|
||||||
|
out_dir = _chat_out_dir(job_id)
|
||||||
|
markdown = review_chat.render_log_markdown(review_chat.read_log(out_dir))
|
||||||
|
return Response(content=markdown, media_type="text/markdown; charset=utf-8")
|
||||||
|
|
||||||
|
|
||||||
@app.get("/jobs/{job_id}/sheet-image/{page}")
|
@app.get("/jobs/{job_id}/sheet-image/{page}")
|
||||||
def sheet_image(job_id: str, page: int):
|
def sheet_image(job_id: str, page: int):
|
||||||
"""Render one page of a completed job's source PDF as JPEG (sheet viewer)."""
|
"""Render one page of a completed job's source PDF as JPEG (sheet viewer)."""
|
||||||
|
|||||||
@@ -0,0 +1,368 @@
|
|||||||
|
"""Review-screen chat: ask the run why it concluded something.
|
||||||
|
|
||||||
|
Read-only by construction. The chat reads job artifacts, calls one LLM, and
|
||||||
|
appends a log record; it never mutates findings, review decisions, or the
|
||||||
|
report, and the prompt forbids it from emitting code or config changes.
|
||||||
|
|
||||||
|
Every turn is logged twice, on purpose:
|
||||||
|
|
||||||
|
- ``<out_dir>/review/chat_log.jsonl`` - job-local, the auditable record of what
|
||||||
|
was asked about which finding and what came back.
|
||||||
|
- ``REVIEW_FEEDBACK_DIR/chat_turns.jsonl`` - cross-job, append-only, so the
|
||||||
|
corrections a reviewer makes in conversation ("that is not a floor drain, it
|
||||||
|
is a power floor box") accumulate somewhere a future run can be primed from.
|
||||||
|
Nothing reads this yet; writing it is what makes that possible later.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import uuid
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from typing import Any, Dict, List, Optional
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
from backend.llm import call_json
|
||||||
|
from backend.pipeline._serialize import dumps
|
||||||
|
from backend.review.chat_context import build_context
|
||||||
|
from backend.review.chat_prompts import (
|
||||||
|
REVIEW_CHAT_SYSTEM_PROMPT,
|
||||||
|
REVIEW_CHAT_USER_PROMPT,
|
||||||
|
)
|
||||||
|
from backend.review.feedback import append_shared_feedback
|
||||||
|
|
||||||
|
_ANSWERABLE = {"yes", "partial", "no"}
|
||||||
|
_ASSESSMENTS = {"looks_supported", "looks_unsupported", "cannot_tell", "not_applicable"}
|
||||||
|
_CONFIDENCE = {"high", "medium", "low"}
|
||||||
|
_MAX_FINDINGS = 12
|
||||||
|
_MAX_EVIDENCE = 12
|
||||||
|
|
||||||
|
|
||||||
|
class ChatError(Exception):
|
||||||
|
"""Raised for a caller-fixable problem (bad question, chat disabled)."""
|
||||||
|
|
||||||
|
|
||||||
|
def _now() -> str:
|
||||||
|
return datetime.now(timezone.utc).isoformat()
|
||||||
|
|
||||||
|
|
||||||
|
def _one_of(value: Any, allowed: set, default: str) -> str:
|
||||||
|
text = str(value or "").strip().lower()
|
||||||
|
return text if text in allowed else default
|
||||||
|
|
||||||
|
|
||||||
|
def _clean_question(raw: Any) -> str:
|
||||||
|
question = str(raw or "").strip()
|
||||||
|
if not question:
|
||||||
|
raise ChatError("question is required")
|
||||||
|
if len(question) > config.REVIEW_CHAT_MAX_QUESTION_CHARS:
|
||||||
|
raise ChatError(
|
||||||
|
f"question is too long (max {config.REVIEW_CHAT_MAX_QUESTION_CHARS} characters)")
|
||||||
|
return question
|
||||||
|
|
||||||
|
|
||||||
|
def _log_path(out_dir: str) -> str:
|
||||||
|
return os.path.join(out_dir, "review", "chat_log.jsonl")
|
||||||
|
|
||||||
|
|
||||||
|
def read_log(out_dir: str, review_item_id: Optional[str] = None,
|
||||||
|
scope_only: bool = False) -> List[Dict[str, Any]]:
|
||||||
|
"""Chat turns for this job, oldest first.
|
||||||
|
|
||||||
|
``review_item_id`` filters to one finding's thread; with ``scope_only`` and
|
||||||
|
no id, returns only the run-scope turns. Corrupt lines are skipped rather
|
||||||
|
than failing the read - a truncated log must not hide the rest.
|
||||||
|
"""
|
||||||
|
path = _log_path(out_dir)
|
||||||
|
if not os.path.isfile(path):
|
||||||
|
return []
|
||||||
|
turns: List[Dict[str, Any]] = []
|
||||||
|
try:
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
for line in f:
|
||||||
|
line = line.strip()
|
||||||
|
if not line:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
turn = json.loads(line)
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
continue
|
||||||
|
if not isinstance(turn, dict):
|
||||||
|
continue
|
||||||
|
if review_item_id is not None:
|
||||||
|
if turn.get("review_item_id") != review_item_id:
|
||||||
|
continue
|
||||||
|
elif scope_only and turn.get("review_item_id") is not None:
|
||||||
|
continue
|
||||||
|
turns.append(turn)
|
||||||
|
except OSError:
|
||||||
|
return []
|
||||||
|
return turns
|
||||||
|
|
||||||
|
|
||||||
|
def _append_log(out_dir: str, turn: Dict[str, Any]) -> None:
|
||||||
|
"""Append one turn as a JSON line; never raises on I/O failure."""
|
||||||
|
try:
|
||||||
|
os.makedirs(os.path.join(out_dir, "review"), exist_ok=True)
|
||||||
|
with open(_log_path(out_dir), "a", encoding="utf-8") as f:
|
||||||
|
f.write(json.dumps(turn) + "\n")
|
||||||
|
except OSError as e:
|
||||||
|
print(f"[ReviewChat] chat log write failed: {e}")
|
||||||
|
|
||||||
|
|
||||||
|
def _issue_snapshot(item: Optional[Dict]) -> Optional[Dict[str, Any]]:
|
||||||
|
"""The issue as it stood when asked about - the log's 'issue in question'.
|
||||||
|
|
||||||
|
Copied rather than referenced by id so the log stays readable after
|
||||||
|
finalization renumbers or suppresses the finding.
|
||||||
|
"""
|
||||||
|
if not item:
|
||||||
|
return None
|
||||||
|
payload = item.get("payload") or {}
|
||||||
|
if item.get("kind") == "clean_cluster":
|
||||||
|
return {
|
||||||
|
"review_item_id": item.get("review_item_id"),
|
||||||
|
"kind": item.get("kind"),
|
||||||
|
"cluster_key": payload.get("key"),
|
||||||
|
"location": payload.get("location"),
|
||||||
|
"disciplines": payload.get("disciplines"),
|
||||||
|
}
|
||||||
|
return {
|
||||||
|
"review_item_id": item.get("review_item_id"),
|
||||||
|
"kind": item.get("kind"),
|
||||||
|
"issue_id": payload.get("issue_id"),
|
||||||
|
"source_stage": payload.get("source_stage"),
|
||||||
|
"category": payload.get("category"),
|
||||||
|
"severity": payload.get("severity"),
|
||||||
|
"confidence": payload.get("confidence"),
|
||||||
|
"location": payload.get("location"),
|
||||||
|
"disciplines": payload.get("disciplines"),
|
||||||
|
"sheets": payload.get("sheets"),
|
||||||
|
"description": payload.get("description"),
|
||||||
|
"blocking": item.get("blocking"),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _history_block(turns: List[Dict[str, Any]]) -> str:
|
||||||
|
if not turns:
|
||||||
|
return ""
|
||||||
|
recent = turns[-config.REVIEW_CHAT_HISTORY_TURNS:]
|
||||||
|
lines = ["Earlier turns in this thread (oldest first):"]
|
||||||
|
for turn in recent:
|
||||||
|
lines.append(f"Reviewer: {turn.get('question') or ''}")
|
||||||
|
lines.append(f"You: {turn.get('answer') or ''}")
|
||||||
|
lines.append("")
|
||||||
|
return "\n".join(lines)
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_answer(parsed: Optional[Dict]) -> Optional[Dict[str, Any]]:
|
||||||
|
"""Coerce the model's JSON into the log/API shape, or None if unusable."""
|
||||||
|
if not isinstance(parsed, dict):
|
||||||
|
return None
|
||||||
|
answer = str(parsed.get("answer") or "").strip()
|
||||||
|
if not answer:
|
||||||
|
return None
|
||||||
|
findings = [
|
||||||
|
str(item).strip()
|
||||||
|
for item in (parsed.get("findings") or [])
|
||||||
|
if isinstance(item, (str, int, float)) and str(item).strip()
|
||||||
|
][:_MAX_FINDINGS]
|
||||||
|
evidence = []
|
||||||
|
for item in (parsed.get("evidence_cited") or [])[:_MAX_EVIDENCE]:
|
||||||
|
if not isinstance(item, dict):
|
||||||
|
continue
|
||||||
|
evidence.append({
|
||||||
|
"artifact": str(item.get("artifact") or "").strip() or None,
|
||||||
|
"sheet": item.get("sheet"),
|
||||||
|
"quote": str(item.get("quote") or "").strip() or None,
|
||||||
|
"why_it_matters": str(item.get("why_it_matters") or "").strip() or None,
|
||||||
|
})
|
||||||
|
correction = parsed.get("suggested_category_correction")
|
||||||
|
correction = str(correction).strip() if correction else ""
|
||||||
|
missing = parsed.get("missing_information")
|
||||||
|
return {
|
||||||
|
"answer": answer,
|
||||||
|
"findings": findings,
|
||||||
|
"evidence_cited": evidence,
|
||||||
|
"answerable": _one_of(parsed.get("answerable"), _ANSWERABLE, "partial"),
|
||||||
|
"missing_information": str(missing).strip() if missing else None,
|
||||||
|
"assessment_of_finding": _one_of(parsed.get("assessment_of_finding"),
|
||||||
|
_ASSESSMENTS, "cannot_tell"),
|
||||||
|
# The feedback signal: a reviewer correcting a misidentification in
|
||||||
|
# conversation ("that is a power floor box") lands here as structured
|
||||||
|
# data instead of dying in free text.
|
||||||
|
"suggested_category_correction": correction or None,
|
||||||
|
"confidence": _one_of(parsed.get("confidence"), _CONFIDENCE, "low"),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _feedback_record(turn: Dict[str, Any]) -> Dict[str, Any]:
|
||||||
|
"""Cross-job roll-up of one turn: metadata + the correction signal.
|
||||||
|
|
||||||
|
Mirrors the privacy stance of the decision labels - no images, no raw sheet
|
||||||
|
dumps. The question and answer ARE carried, because a chat turn without its
|
||||||
|
question is not usable as feedback; keep this store job-internal.
|
||||||
|
"""
|
||||||
|
issue = turn.get("issue") or {}
|
||||||
|
return {
|
||||||
|
"kind": "review_chat_turn",
|
||||||
|
"turn_id": turn.get("turn_id"),
|
||||||
|
"job_id": turn.get("job_id"),
|
||||||
|
"created_at": turn.get("created_at"),
|
||||||
|
"review_item_id": turn.get("review_item_id"),
|
||||||
|
"scope": turn.get("scope"),
|
||||||
|
"issue_id": issue.get("issue_id"),
|
||||||
|
"source_stage": issue.get("source_stage"),
|
||||||
|
"category": issue.get("category"),
|
||||||
|
"severity": issue.get("severity"),
|
||||||
|
"confidence": issue.get("confidence"),
|
||||||
|
"sheets": issue.get("sheets"),
|
||||||
|
"question": turn.get("question"),
|
||||||
|
"answer": turn.get("answer"),
|
||||||
|
"findings": turn.get("findings"),
|
||||||
|
"assessment_of_finding": turn.get("assessment_of_finding"),
|
||||||
|
"suggested_category_correction": turn.get("suggested_category_correction"),
|
||||||
|
"answerable": turn.get("answerable"),
|
||||||
|
"model": turn.get("model"),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def ask(job_id: str, out_dir: str, question: str,
|
||||||
|
review_item_id: Optional[str] = None,
|
||||||
|
queue: Optional[List[Dict]] = None,
|
||||||
|
decisions: Optional[Dict[str, Dict]] = None) -> Dict[str, Any]:
|
||||||
|
"""Answer one reviewer question and log the turn.
|
||||||
|
|
||||||
|
Returns the logged turn. Raises ChatError for a bad question or a disabled
|
||||||
|
chat, and RuntimeError when the model call fails outright (the caller maps
|
||||||
|
both to HTTP status codes).
|
||||||
|
"""
|
||||||
|
if not config.ENABLE_REVIEW_CHAT:
|
||||||
|
raise ChatError("review chat is disabled (ENABLE_REVIEW_CHAT=false)")
|
||||||
|
question = _clean_question(question)
|
||||||
|
queue = queue or []
|
||||||
|
|
||||||
|
item = next((candidate for candidate in queue
|
||||||
|
if candidate.get("review_item_id") == review_item_id), None)
|
||||||
|
if review_item_id and item is None:
|
||||||
|
raise ChatError(f"unknown review_item_id {review_item_id!r}")
|
||||||
|
|
||||||
|
context = build_context(out_dir, review_item_id, queue, decisions, question)
|
||||||
|
history = read_log(out_dir, review_item_id=review_item_id) if review_item_id \
|
||||||
|
else read_log(out_dir, scope_only=True)
|
||||||
|
|
||||||
|
scope_line = (
|
||||||
|
f"Scope: this question is about review item {review_item_id}."
|
||||||
|
if item else
|
||||||
|
"Scope: this question is about the run as a whole, not one finding."
|
||||||
|
)
|
||||||
|
user_text = (REVIEW_CHAT_USER_PROMPT
|
||||||
|
.replace("{scope_line}", scope_line)
|
||||||
|
.replace("{question}", question)
|
||||||
|
.replace("{history_block}", _history_block(history))
|
||||||
|
.replace("{context}", dumps(context)))
|
||||||
|
|
||||||
|
parsed = call_json(
|
||||||
|
system_prompt=REVIEW_CHAT_SYSTEM_PROMPT,
|
||||||
|
user_text=user_text,
|
||||||
|
max_tokens=config.REVIEW_CHAT_MAX_TOKENS,
|
||||||
|
model=config.REVIEW_CHAT_MODEL,
|
||||||
|
usage_stage="review.chat",
|
||||||
|
)
|
||||||
|
answer = _normalize_answer(parsed)
|
||||||
|
if answer is None:
|
||||||
|
raise RuntimeError("the model did not return a usable answer")
|
||||||
|
|
||||||
|
turn = {
|
||||||
|
"turn_id": uuid.uuid4().hex[:12],
|
||||||
|
"job_id": job_id,
|
||||||
|
"created_at": _now(),
|
||||||
|
"review_item_id": review_item_id,
|
||||||
|
"scope": context.get("scope"),
|
||||||
|
"issue": _issue_snapshot(item),
|
||||||
|
"question": question,
|
||||||
|
"reviewer_decision_at_time": context.get("reviewer_decision_so_far"),
|
||||||
|
"artifacts_consulted": sorted(
|
||||||
|
name for name, present
|
||||||
|
in (context.get("artifacts_available") or {}).items() if present
|
||||||
|
),
|
||||||
|
"model": config.REVIEW_CHAT_MODEL,
|
||||||
|
**answer,
|
||||||
|
}
|
||||||
|
_append_log(out_dir, turn)
|
||||||
|
append_shared_feedback(_feedback_record(turn))
|
||||||
|
return turn
|
||||||
|
|
||||||
|
|
||||||
|
def render_log_markdown(turns: List[Dict[str, Any]]) -> str:
|
||||||
|
"""Human-readable transcript: the issue, the questions, the findings.
|
||||||
|
|
||||||
|
Grouped by review item so one finding's whole thread reads together, with
|
||||||
|
run-scope questions last under their own heading.
|
||||||
|
"""
|
||||||
|
by_item: Dict[str, List[Dict[str, Any]]] = {}
|
||||||
|
for turn in turns:
|
||||||
|
by_item.setdefault(turn.get("review_item_id") or "", []).append(turn)
|
||||||
|
|
||||||
|
lines = ["# Review chat log", ""]
|
||||||
|
if not turns:
|
||||||
|
lines.append("_No questions have been asked about this run._")
|
||||||
|
return "\n".join(lines) + "\n"
|
||||||
|
lines.append(f"{len(turns)} turn(s) across {len(by_item)} thread(s).")
|
||||||
|
lines.append("")
|
||||||
|
|
||||||
|
for item_id in sorted(by_item, key=lambda key: (key == "", key)):
|
||||||
|
item_turns = by_item[item_id]
|
||||||
|
issue = next((turn.get("issue") for turn in item_turns if turn.get("issue")), None)
|
||||||
|
if not item_id:
|
||||||
|
lines += ["## Run-scope questions", "",
|
||||||
|
"_Not about a single finding._", ""]
|
||||||
|
elif issue:
|
||||||
|
lines.append(f"## {issue.get('issue_id') or item_id}")
|
||||||
|
lines.append("")
|
||||||
|
meta = [
|
||||||
|
("Category", issue.get("category")),
|
||||||
|
("Severity", issue.get("severity")),
|
||||||
|
("Run confidence", issue.get("confidence")),
|
||||||
|
("Location", issue.get("location")),
|
||||||
|
("Sheets", ", ".join(str(s) for s in issue.get("sheets") or []) or None),
|
||||||
|
("Stage", issue.get("source_stage")),
|
||||||
|
]
|
||||||
|
for label, value in meta:
|
||||||
|
if value:
|
||||||
|
lines.append(f"- **{label}:** {value}")
|
||||||
|
if issue.get("description"):
|
||||||
|
lines += ["", f"> {issue['description']}"]
|
||||||
|
lines.append("")
|
||||||
|
else:
|
||||||
|
lines += [f"## {item_id}", ""]
|
||||||
|
|
||||||
|
for turn in item_turns:
|
||||||
|
lines.append(f"### Q ({turn.get('created_at') or ''})")
|
||||||
|
lines += ["", turn.get("question") or "", "", "**Answer**", "",
|
||||||
|
turn.get("answer") or "", ""]
|
||||||
|
if turn.get("findings"):
|
||||||
|
lines.append("**Findings**")
|
||||||
|
lines.append("")
|
||||||
|
lines += [f"- {finding}" for finding in turn["findings"]]
|
||||||
|
lines.append("")
|
||||||
|
if turn.get("evidence_cited"):
|
||||||
|
lines += ["**Evidence cited**", ""]
|
||||||
|
for item in turn["evidence_cited"]:
|
||||||
|
where = item.get("artifact") or "?"
|
||||||
|
sheet = f" ({item['sheet']})" if item.get("sheet") else ""
|
||||||
|
quote = item.get("quote") or ""
|
||||||
|
lines.append(f"- `{where}`{sheet}: \"{quote}\"")
|
||||||
|
if item.get("why_it_matters"):
|
||||||
|
lines.append(f" - {item['why_it_matters']}")
|
||||||
|
lines.append("")
|
||||||
|
tail = [
|
||||||
|
("Answerable", turn.get("answerable")),
|
||||||
|
("Assessment", turn.get("assessment_of_finding")),
|
||||||
|
("Confidence", turn.get("confidence")),
|
||||||
|
("Missing", turn.get("missing_information")),
|
||||||
|
("Suggested correction", turn.get("suggested_category_correction")),
|
||||||
|
("Model", turn.get("model")),
|
||||||
|
]
|
||||||
|
lines.append(" | ".join(f"{label}: {value}" for label, value in tail if value))
|
||||||
|
lines.append("")
|
||||||
|
return "\n".join(lines) + "\n"
|
||||||
@@ -0,0 +1,315 @@
|
|||||||
|
"""Evidence bundles for the review chat.
|
||||||
|
|
||||||
|
The chat is an explainer, not an investigator: it may only answer from what the
|
||||||
|
run actually produced. This module assembles that material from the job's own
|
||||||
|
artifacts and hands the model a bounded, slimmed view.
|
||||||
|
|
||||||
|
Two shapes, matching the two kinds of question a reviewer asks:
|
||||||
|
|
||||||
|
- item scope ("why does it think the AC unit is on the ground?") - the finding,
|
||||||
|
its evidence, the cluster the finding came from, the sheets those assertions
|
||||||
|
were extracted from, any wave-5b verification verdict, the Brain's merge/drop
|
||||||
|
decision, and the reviewer's own saved decision.
|
||||||
|
- run scope ("why didn't it pick up the Civil set?") - the sheet index by
|
||||||
|
discipline, the deterministic cover-index reconciliation, per-stage counts,
|
||||||
|
what got suppressed and why, and matching job.log lines.
|
||||||
|
|
||||||
|
Everything here is read-only and degrades to empty on a missing or corrupt
|
||||||
|
artifact; a chat request must never be the thing that breaks a review screen.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
from typing import Any, Dict, List, Optional
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
|
||||||
|
# Assertion/evidence text is quoted back verbatim so the reviewer can check the
|
||||||
|
# answer against the sheet, but a whole cluster of them would swamp the prompt.
|
||||||
|
_MAX_CLUSTER_ASSERTIONS = 40
|
||||||
|
_MAX_SHEET_ASSERTIONS = 25
|
||||||
|
_MAX_SOURCE_TEXT_CHARS = 400
|
||||||
|
_MAX_SHEETS_IN_ROSTER = 400
|
||||||
|
_MAX_SUPPRESSED = 25
|
||||||
|
_MAX_LOG_LINE_CHARS = 400
|
||||||
|
|
||||||
|
|
||||||
|
def _read_json(path: str, default):
|
||||||
|
try:
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
return json.load(f)
|
||||||
|
except (OSError, json.JSONDecodeError):
|
||||||
|
return default
|
||||||
|
|
||||||
|
|
||||||
|
def _truncate(value: Any, limit: int = _MAX_SOURCE_TEXT_CHARS) -> Any:
|
||||||
|
if not isinstance(value, str) or len(value) <= limit:
|
||||||
|
return value
|
||||||
|
return value[:limit] + "..."
|
||||||
|
|
||||||
|
|
||||||
|
def _slim_assertion(assertion: Dict) -> Dict:
|
||||||
|
"""Drop base64/bookkeeping; keep what explains where a value came from."""
|
||||||
|
out = {
|
||||||
|
"sheet_number": assertion.get("sheet_number"),
|
||||||
|
"discipline": assertion.get("discipline"),
|
||||||
|
"attribute": assertion.get("attribute"),
|
||||||
|
"value": assertion.get("value"),
|
||||||
|
"source_text": _truncate(assertion.get("source_text")),
|
||||||
|
"location_key": assertion.get("location_key"),
|
||||||
|
"normalized_value": assertion.get("normalized_value"),
|
||||||
|
"disputed": assertion.get("disputed"),
|
||||||
|
}
|
||||||
|
return {key: value for key, value in out.items() if value is not None}
|
||||||
|
|
||||||
|
|
||||||
|
def _slim_finding(finding: Dict) -> Dict:
|
||||||
|
"""The finding as the run recorded it, including how it was checked."""
|
||||||
|
out = {
|
||||||
|
"issue_id": finding.get("issue_id"),
|
||||||
|
"source_stage": finding.get("source_stage"),
|
||||||
|
"agent": finding.get("agent"),
|
||||||
|
"category": finding.get("category"),
|
||||||
|
"severity": finding.get("severity"),
|
||||||
|
"confidence": finding.get("confidence"),
|
||||||
|
"location": finding.get("location"),
|
||||||
|
"disciplines": finding.get("disciplines"),
|
||||||
|
"sheets": finding.get("sheets"),
|
||||||
|
"description": _truncate(finding.get("description"), 1200),
|
||||||
|
"recommended_resolution": finding.get("recommended_resolution"),
|
||||||
|
"code_reference": finding.get("code_reference"),
|
||||||
|
"risk_score": finding.get("risk_score"),
|
||||||
|
"recommended_priority": finding.get("recommended_priority"),
|
||||||
|
"scope_id": finding.get("scope_id"),
|
||||||
|
"evidence": [
|
||||||
|
{
|
||||||
|
"discipline": item.get("discipline"),
|
||||||
|
"sheet": item.get("sheet"),
|
||||||
|
"source_text": _truncate(item.get("source_text")),
|
||||||
|
"asserted_value": item.get("asserted_value"),
|
||||||
|
}
|
||||||
|
for item in (finding.get("evidence") or [])
|
||||||
|
if isinstance(item, dict)
|
||||||
|
],
|
||||||
|
# Wave 5b / Brain-clarify re-checked some findings against fresh sheet
|
||||||
|
# images + the text layer. When present this is the single best answer
|
||||||
|
# to "did it actually look again?", so it is never dropped.
|
||||||
|
"verification": finding.get("verification"),
|
||||||
|
"clarification_of": finding.get("clarification_of"),
|
||||||
|
}
|
||||||
|
return {key: value for key, value in out.items() if value is not None}
|
||||||
|
|
||||||
|
|
||||||
|
def _slim_sheet(sheet: Dict, limit: int = _MAX_SHEET_ASSERTIONS) -> Dict:
|
||||||
|
assertions = sheet.get("assertions") or []
|
||||||
|
out = {
|
||||||
|
"sheet_number": sheet.get("sheet_number"),
|
||||||
|
"sheet_title": sheet.get("sheet_title"),
|
||||||
|
"discipline": sheet.get("discipline"),
|
||||||
|
"level": sheet.get("level"),
|
||||||
|
"page_number": sheet.get("page_number"),
|
||||||
|
"assertion_count": len(assertions),
|
||||||
|
"assertions": [_slim_assertion(item) for item in assertions[:limit]],
|
||||||
|
}
|
||||||
|
if len(assertions) > limit:
|
||||||
|
out["assertions_omitted"] = len(assertions) - limit
|
||||||
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
def _discipline_roster(sheet_index: Dict, sheets: List[Dict]) -> Dict[str, List[str]]:
|
||||||
|
"""Sheet numbers grouped by discipline - the 'is Civil in here?' answer.
|
||||||
|
|
||||||
|
Built from the classified sheet index when there is one, falling back to
|
||||||
|
raw extraction, so an empty/failed index stage does not read as "no sheets".
|
||||||
|
"""
|
||||||
|
entries = (sheet_index or {}).get("sheet_index") or []
|
||||||
|
if not entries:
|
||||||
|
entries = [
|
||||||
|
{"sheet_number": sheet.get("sheet_number"),
|
||||||
|
"discipline": sheet.get("discipline")}
|
||||||
|
for sheet in sheets or []
|
||||||
|
]
|
||||||
|
roster: Dict[str, List[str]] = {}
|
||||||
|
for entry in entries:
|
||||||
|
if not isinstance(entry, dict):
|
||||||
|
continue
|
||||||
|
discipline = str(entry.get("discipline") or "unknown")
|
||||||
|
number = entry.get("sheet_number") or entry.get("sheet_id") or "?"
|
||||||
|
bucket = roster.setdefault(discipline, [])
|
||||||
|
if len(bucket) < _MAX_SHEETS_IN_ROSTER and number not in bucket:
|
||||||
|
bucket.append(str(number))
|
||||||
|
return roster
|
||||||
|
|
||||||
|
|
||||||
|
def _log_excerpt(out_dir: str, terms: List[str], limit: int) -> List[str]:
|
||||||
|
"""job.log lines mentioning any search term, newest last.
|
||||||
|
|
||||||
|
The run log is where stage skips, retries, and coverage decisions are
|
||||||
|
recorded ("[Code] gated off", "[Extract] page 12 empty"), which is often
|
||||||
|
the literal answer to "why didn't it look at X".
|
||||||
|
"""
|
||||||
|
path = os.path.join(out_dir, "job.log")
|
||||||
|
needles = [term.lower() for term in terms if term and len(str(term)) >= 2]
|
||||||
|
if not needles or not os.path.isfile(path):
|
||||||
|
return []
|
||||||
|
hits: List[str] = []
|
||||||
|
try:
|
||||||
|
with open(path, encoding="utf-8", errors="replace") as f:
|
||||||
|
for line in f:
|
||||||
|
lowered = line.lower()
|
||||||
|
if any(needle in lowered for needle in needles):
|
||||||
|
hits.append(_truncate(line.rstrip("\n"), _MAX_LOG_LINE_CHARS))
|
||||||
|
except OSError:
|
||||||
|
return []
|
||||||
|
return hits[-limit:]
|
||||||
|
|
||||||
|
|
||||||
|
def _stage_terms(question: str) -> List[str]:
|
||||||
|
"""Search terms for the log: quoted sheet-ish tokens plus long words.
|
||||||
|
|
||||||
|
Deliberately crude - this only decides which log lines get shown, and an
|
||||||
|
over-broad match is bounded by REVIEW_CHAT_LOG_LINES anyway.
|
||||||
|
"""
|
||||||
|
tokens = re.findall(r"[A-Za-z][A-Za-z0-9.\-]{2,}", question or "")
|
||||||
|
stop = {"the", "why", "did", "not", "and", "for", "was", "were", "does",
|
||||||
|
"this", "that", "with", "from", "what", "how", "you", "its",
|
||||||
|
"it's", "there", "when", "have", "has", "any", "are", "but"}
|
||||||
|
return [token for token in tokens if token.lower() not in stop][:12]
|
||||||
|
|
||||||
|
|
||||||
|
def build_context(out_dir: str, review_item_id: Optional[str],
|
||||||
|
queue: Optional[List[Dict]] = None,
|
||||||
|
decisions: Optional[Dict[str, Dict]] = None,
|
||||||
|
question: str = "") -> Dict[str, Any]:
|
||||||
|
"""Assemble the evidence bundle for one chat turn.
|
||||||
|
|
||||||
|
``review_item_id`` selects item scope; None (or an id not in the queue)
|
||||||
|
gives run scope. Missing artifacts degrade to empty sections rather than
|
||||||
|
raising - the model is told what is missing via ``artifacts_available``.
|
||||||
|
"""
|
||||||
|
report = _read_json(os.path.join(out_dir, "conflicts.json"), {}) or {}
|
||||||
|
snapshot = _read_json(os.path.join(out_dir, "agent", "memory.json"), {}) or {}
|
||||||
|
summary = report.get("summary") or {}
|
||||||
|
sheets = snapshot.get("sheets") or []
|
||||||
|
sheet_index = report.get("sheet_index") or snapshot.get("sheet_index") or {}
|
||||||
|
|
||||||
|
context: Dict[str, Any] = {
|
||||||
|
"scope": "run",
|
||||||
|
"run": {
|
||||||
|
"source": report.get("source"),
|
||||||
|
"pipeline_mode": summary.get("pipeline_mode"),
|
||||||
|
"agent_status": summary.get("agent_status"),
|
||||||
|
"sheets_analyzed": summary.get("sheets_analyzed") or len(sheets),
|
||||||
|
"by_stage": summary.get("by_stage"),
|
||||||
|
"conflicts_found": summary.get("conflicts_found"),
|
||||||
|
"by_severity": summary.get("by_severity"),
|
||||||
|
"models_used": summary.get("models_used"),
|
||||||
|
# Stage gating is the answer to a whole class of "why didn't it
|
||||||
|
# check X" questions, so it is stated rather than left implied.
|
||||||
|
"code_review_enabled": config.ENABLE_CODE_REVIEW,
|
||||||
|
},
|
||||||
|
"sheets_by_discipline": _discipline_roster(sheet_index, sheets),
|
||||||
|
"sheet_reconciliation": report.get("sheet_reconciliation"),
|
||||||
|
"missing_expected_sheets": (sheet_index or {}).get("missing_expected_sheets"),
|
||||||
|
"suppressed_by_the_run": [
|
||||||
|
{
|
||||||
|
"issue_id": item.get("issue_id"),
|
||||||
|
"category": item.get("category"),
|
||||||
|
"description": _truncate(item.get("description"), 300),
|
||||||
|
"verification": item.get("verification"),
|
||||||
|
}
|
||||||
|
for item in (snapshot.get("suppressed") or [])[:_MAX_SUPPRESSED]
|
||||||
|
if isinstance(item, dict)
|
||||||
|
],
|
||||||
|
"artifacts_available": {
|
||||||
|
"conflicts.json": bool(report),
|
||||||
|
"agent/memory.json": bool(snapshot),
|
||||||
|
"job.log": os.path.isfile(os.path.join(out_dir, "job.log")),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
item = None
|
||||||
|
for candidate in queue or []:
|
||||||
|
if candidate.get("review_item_id") == review_item_id:
|
||||||
|
item = candidate
|
||||||
|
break
|
||||||
|
if item is None:
|
||||||
|
context["log_excerpt"] = _log_excerpt(
|
||||||
|
out_dir, _stage_terms(question), config.REVIEW_CHAT_LOG_LINES)
|
||||||
|
return context
|
||||||
|
|
||||||
|
context["scope"] = "item"
|
||||||
|
payload = item.get("payload") or {}
|
||||||
|
context["review_item"] = {
|
||||||
|
"review_item_id": item.get("review_item_id"),
|
||||||
|
"kind": item.get("kind"),
|
||||||
|
"blocking": item.get("blocking"),
|
||||||
|
"review_triggers": item.get("reasons"),
|
||||||
|
}
|
||||||
|
saved = (decisions or {}).get(review_item_id) or {}
|
||||||
|
if saved:
|
||||||
|
context["reviewer_decision_so_far"] = {
|
||||||
|
"decision": saved.get("decision"),
|
||||||
|
"reason_code": saved.get("reason_code"),
|
||||||
|
"comment": _truncate(saved.get("comment")),
|
||||||
|
}
|
||||||
|
|
||||||
|
if item.get("kind") == "clean_cluster":
|
||||||
|
context["cluster"] = {
|
||||||
|
"key": payload.get("key"),
|
||||||
|
"location": payload.get("location"),
|
||||||
|
"disciplines": payload.get("disciplines"),
|
||||||
|
"kind": payload.get("kind"),
|
||||||
|
"assertions": [_slim_assertion(a)
|
||||||
|
for a in (payload.get("assertions") or [])[:_MAX_CLUSTER_ASSERTIONS]],
|
||||||
|
}
|
||||||
|
cited_sheets = [a.get("sheet_number") for a in payload.get("assertions") or []]
|
||||||
|
else:
|
||||||
|
context["finding"] = _slim_finding(payload)
|
||||||
|
cited_sheets = list(payload.get("sheets") or [])
|
||||||
|
cited_sheets += [e.get("sheet") for e in payload.get("evidence") or []
|
||||||
|
if isinstance(e, dict)]
|
||||||
|
scope_id = str(payload.get("scope_id") or "")
|
||||||
|
if scope_id.startswith("conflict:"):
|
||||||
|
cluster_key = scope_id.split(":", 1)[1]
|
||||||
|
cluster = next((c for c in snapshot.get("clusters") or []
|
||||||
|
if c.get("key") == cluster_key), None)
|
||||||
|
if cluster is not None:
|
||||||
|
assertions = cluster.get("assertions") or []
|
||||||
|
context["originating_cluster"] = {
|
||||||
|
"key": cluster.get("key"),
|
||||||
|
"location": cluster.get("location"),
|
||||||
|
"disciplines": cluster.get("disciplines"),
|
||||||
|
"kind": cluster.get("kind"),
|
||||||
|
"disputed_attributes": cluster.get("disputed_attributes"),
|
||||||
|
"assertion_count": len(assertions),
|
||||||
|
"assertions": [_slim_assertion(a)
|
||||||
|
for a in assertions[:_MAX_CLUSTER_ASSERTIONS]],
|
||||||
|
}
|
||||||
|
cited_sheets += [a.get("sheet_number") for a in assertions]
|
||||||
|
issue_id = payload.get("issue_id")
|
||||||
|
brain_decisions = [
|
||||||
|
decision for decision in snapshot.get("decisions") or []
|
||||||
|
if isinstance(decision, dict) and (
|
||||||
|
decision.get("kept_issue_id") == issue_id
|
||||||
|
or issue_id in (decision.get("finding_refs") or []))
|
||||||
|
]
|
||||||
|
if brain_decisions:
|
||||||
|
context["brain_decisions"] = brain_decisions[:10]
|
||||||
|
|
||||||
|
# The sheets the finding actually rests on, with their raw extraction -
|
||||||
|
# this is what lets the model say "it read 'MOUNTED ON GRADE' off M2.1".
|
||||||
|
wanted = {str(number) for number in cited_sheets if number}
|
||||||
|
if wanted:
|
||||||
|
context["source_sheets"] = [
|
||||||
|
_slim_sheet(sheet) for sheet in sheets
|
||||||
|
if str(sheet.get("sheet_number") or "") in wanted
|
||||||
|
]
|
||||||
|
|
||||||
|
context["log_excerpt"] = _log_excerpt(
|
||||||
|
out_dir,
|
||||||
|
_stage_terms(question) + sorted(wanted),
|
||||||
|
config.REVIEW_CHAT_LOG_LINES,
|
||||||
|
)
|
||||||
|
return context
|
||||||
@@ -0,0 +1,36 @@
|
|||||||
|
"""Prompts for the review-screen chat (read-only run explainer)."""
|
||||||
|
|
||||||
|
REVIEW_CHAT_SYSTEM_PROMPT = """You are the explainer for a completed automated construction-drawing review run. A human reviewer is working through the review queue and is asking you why the run reached a particular conclusion.
|
||||||
|
|
||||||
|
Your ONLY job is to explain what the run did and why, using the run's own artifacts, which are supplied to you as a JSON context bundle. You are a witness to the run, not a participant in it.
|
||||||
|
|
||||||
|
HARD RULES - never break these:
|
||||||
|
- You do NOT write, propose, suggest, or output code, patches, diffs, file edits, configuration changes, prompt changes, or shell commands. If the reviewer asks for any of those, say that this chat only explains findings, and answer the underlying question in construction-review terms instead.
|
||||||
|
- You do NOT change, re-decide, confirm, reject, or re-score any finding. The reviewer owns that decision; the radio buttons on their screen are the only thing that changes a finding. You may explain what the evidence supports, and you may say plainly that a finding looks wrong, but you never state that a finding "has been" changed.
|
||||||
|
- You answer ONLY from the supplied context bundle. You have no access to the PDF, to sheets that were not extracted, or to anything outside the bundle. Never invent a sheet number, a quotation, a dimension, or a stage that is not in the bundle.
|
||||||
|
- Separate what the run RECORDED from what you INFER. Attribute recorded facts to the artifact they came from ("the extractor recorded ... on M2.1"). Mark reasoning of your own as inference.
|
||||||
|
- When the bundle does not contain the answer, say so directly and name what is missing and which artifact would have held it. "The Civil sheets were never extracted, so there are no Civil assertions to compare" is a good answer. Guessing is not.
|
||||||
|
|
||||||
|
HOW TO ANSWER "why does it think X":
|
||||||
|
Trace the chain backwards through the bundle and quote it: the finding's evidence, the assertions in the originating cluster, the source_text the extractor pulled off each sheet, any verification verdict from the re-check pass, and the Brain's merge or drop decision. If a value is marked disputed, or the verification status is refuted or unverified, say so - that is usually the real answer.
|
||||||
|
|
||||||
|
HOW TO ANSWER "why didn't it pick up X":
|
||||||
|
Work through the bundle's coverage material in this order and report which one explains it: (1) sheets_by_discipline - was the discipline in the set at all? (2) sheet_reconciliation - did the cover sheet's own index declare sheets that were never identified (declared_not_in_set)? (3) run.by_stage and run.code_review_enabled - was the responsible stage gated off or did it produce nothing? (4) suppressed_by_the_run - was something found and then dropped? (5) log_excerpt - did the run log record a skip, a retry, or an empty page? Name the specific reason. If several are possible, say which is best supported and what would confirm it.
|
||||||
|
|
||||||
|
Be direct and concrete. Quote verbatim source_text when it carries the answer. A short, specific, evidence-anchored answer is worth more than a thorough hedge. Use plain ASCII. Respond only with valid JSON."""
|
||||||
|
|
||||||
|
REVIEW_CHAT_USER_PROMPT = """A reviewer is asking about this run. Answer from the context bundle only.
|
||||||
|
|
||||||
|
Respond ONLY with a valid JSON object - no markdown fences, no prose outside the JSON:
|
||||||
|
{"answer":"your direct explanation to the reviewer, plain text, no markdown headings","findings":["one short factual determination per item - what you established about this question, each standing on its own"],"evidence_cited":[{"artifact":"which part of the bundle, e.g. 'finding.evidence' or 'source_sheets[M2.1]' or 'log_excerpt'","sheet":"sheet number or null","quote":"verbatim text from the bundle","why_it_matters":"one sentence"}],"answerable":"yes | partial | no","missing_information":"what the bundle would need to answer fully, or null if fully answered","assessment_of_finding":"looks_supported | looks_unsupported | cannot_tell | not_applicable","suggested_category_correction":"if the reviewer is telling you the run misidentified an object, the object they say it actually is, e.g. 'power floor box'; otherwise null","confidence":"high | medium | low"}
|
||||||
|
|
||||||
|
Set assessment_of_finding to not_applicable for run-scope questions that are not about one finding. Set suggested_category_correction to null unless the reviewer is asserting a correction - do not invent one.
|
||||||
|
|
||||||
|
{scope_line}
|
||||||
|
|
||||||
|
Reviewer's question:
|
||||||
|
{question}
|
||||||
|
|
||||||
|
{history_block}
|
||||||
|
Context bundle (the complete set of artifacts you may reason from):
|
||||||
|
{context}"""
|
||||||
@@ -1,9 +1,19 @@
|
|||||||
"""Feedback labels: one label artifact per human-review decision, for metrics."""
|
"""Feedback labels: one label artifact per human-review decision, for metrics.
|
||||||
|
|
||||||
|
Labels are written twice: job-locally under ``<out_dir>/review/`` (the
|
||||||
|
auditable record for that run) and, via ``append_shared_feedback``, to the
|
||||||
|
cross-job store at ``config.REVIEW_FEEDBACK_DIR``. The shared store is
|
||||||
|
append-only and nothing reads it yet - it exists so that a later pass can prime
|
||||||
|
a run with what reviewers corrected on previous sets without having to walk
|
||||||
|
every job directory.
|
||||||
|
"""
|
||||||
|
|
||||||
import json
|
import json
|
||||||
import os
|
import os
|
||||||
from datetime import datetime, timezone
|
from datetime import datetime, timezone
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
|
||||||
|
|
||||||
def _as_dict(value) -> dict:
|
def _as_dict(value) -> dict:
|
||||||
return value if isinstance(value, dict) else {}
|
return value if isinstance(value, dict) else {}
|
||||||
@@ -21,6 +31,7 @@ def decision_to_label(queue_item: dict, decision: dict, job: dict) -> dict:
|
|||||||
payload = _as_dict(queue_item.get("payload"))
|
payload = _as_dict(queue_item.get("payload"))
|
||||||
summary = _as_dict(_as_dict(job.get("report")).get("summary"))
|
summary = _as_dict(_as_dict(job.get("report")).get("summary"))
|
||||||
return {
|
return {
|
||||||
|
"kind": "review_decision",
|
||||||
"review_item_id": queue_item.get("review_item_id"),
|
"review_item_id": queue_item.get("review_item_id"),
|
||||||
"job_id": job.get("job_id"),
|
"job_id": job.get("job_id"),
|
||||||
"pipeline_mode": job.get("pipeline_mode"),
|
"pipeline_mode": job.get("pipeline_mode"),
|
||||||
@@ -30,6 +41,11 @@ def decision_to_label(queue_item: dict, decision: dict, job: dict) -> dict:
|
|||||||
"confidence": payload.get("confidence"),
|
"confidence": payload.get("confidence"),
|
||||||
"decision": decision.get("decision"),
|
"decision": decision.get("decision"),
|
||||||
"reason_code": decision.get("reason_code"),
|
"reason_code": decision.get("reason_code"),
|
||||||
|
# The reviewer's structured corrections. Carried here (and into the
|
||||||
|
# cross-job store) so "wrong category" survives as data rather than
|
||||||
|
# only as free text on the suppressed issue.
|
||||||
|
"category_correction": decision.get("category_correction"),
|
||||||
|
"severity_correction": decision.get("severity_correction"),
|
||||||
"location": payload.get("location"),
|
"location": payload.get("location"),
|
||||||
"disciplines": payload.get("disciplines"),
|
"disciplines": payload.get("disciplines"),
|
||||||
"sheets": payload.get("sheets"),
|
"sheets": payload.get("sheets"),
|
||||||
@@ -40,7 +56,12 @@ def decision_to_label(queue_item: dict, decision: dict, job: dict) -> dict:
|
|||||||
|
|
||||||
|
|
||||||
def write_label(out_dir: str, label: dict) -> None:
|
def write_label(out_dir: str, label: dict) -> None:
|
||||||
"""Append one label as a JSON line; never raises on I/O failure."""
|
"""Append one label job-locally and to the cross-job store.
|
||||||
|
|
||||||
|
Never raises on I/O failure: a lost label must not fail the save that
|
||||||
|
produced it.
|
||||||
|
"""
|
||||||
|
append_shared_feedback(label)
|
||||||
try:
|
try:
|
||||||
review_dir = os.path.join(out_dir, "review")
|
review_dir = os.path.join(out_dir, "review")
|
||||||
os.makedirs(review_dir, exist_ok=True)
|
os.makedirs(review_dir, exist_ok=True)
|
||||||
@@ -49,3 +70,49 @@ def write_label(out_dir: str, label: dict) -> None:
|
|||||||
f.write(json.dumps(label) + "\n")
|
f.write(json.dumps(label) + "\n")
|
||||||
except OSError as e:
|
except OSError as e:
|
||||||
print(f"[Review] feedback label write failed: {e}")
|
print(f"[Review] feedback label write failed: {e}")
|
||||||
|
|
||||||
|
|
||||||
|
def append_shared_feedback(record: dict) -> None:
|
||||||
|
"""Append one record to the cross-job feedback store; never raises.
|
||||||
|
|
||||||
|
One JSONL file per record ``kind`` so a reader can pick up decisions and
|
||||||
|
chat turns independently. Failure here is logged and swallowed: the
|
||||||
|
cross-job roll-up is a convenience, and losing a line must never fail the
|
||||||
|
review action that produced it.
|
||||||
|
"""
|
||||||
|
try:
|
||||||
|
kind = str(record.get("kind") or "misc")
|
||||||
|
os.makedirs(config.REVIEW_FEEDBACK_DIR, exist_ok=True)
|
||||||
|
name = "chat_turns.jsonl" if kind == "review_chat_turn" else "decisions.jsonl"
|
||||||
|
path = os.path.join(config.REVIEW_FEEDBACK_DIR, name)
|
||||||
|
with open(path, "a", encoding="utf-8") as f:
|
||||||
|
f.write(json.dumps(record) + "\n")
|
||||||
|
except OSError as e:
|
||||||
|
print(f"[Review] shared feedback write failed: {e}")
|
||||||
|
|
||||||
|
|
||||||
|
def read_shared_feedback(kind: str = "review_decision") -> list:
|
||||||
|
"""Read the cross-job store for one record kind, oldest first.
|
||||||
|
|
||||||
|
Corrupt lines are skipped so a partial write cannot hide the rest.
|
||||||
|
"""
|
||||||
|
name = "chat_turns.jsonl" if kind == "review_chat_turn" else "decisions.jsonl"
|
||||||
|
path = os.path.join(config.REVIEW_FEEDBACK_DIR, name)
|
||||||
|
if not os.path.isfile(path):
|
||||||
|
return []
|
||||||
|
records = []
|
||||||
|
try:
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
for line in f:
|
||||||
|
line = line.strip()
|
||||||
|
if not line:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
value = json.loads(line)
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
continue
|
||||||
|
if isinstance(value, dict):
|
||||||
|
records.append(value)
|
||||||
|
except OSError:
|
||||||
|
return []
|
||||||
|
return records
|
||||||
@@ -94,7 +94,9 @@ table, and a senior reviewer at the end sorts it all into one clean report.
|
|||||||
+----------------------------------------------------------+
|
+----------------------------------------------------------+
|
||||||
| 7. HUMAN REVIEW GATE |
|
| 7. HUMAN REVIEW GATE |
|
||||||
| The important/uncertain findings are queued for a |
|
| The important/uncertain findings are queued for a |
|
||||||
| real person to Confirm / Reject / mark Unsure |
|
| real person to Confirm / Reject / mark Unsure. |
|
||||||
|
| You can also ASK the run why it concluded any of |
|
||||||
|
| them, or why it never looked at something |
|
||||||
+----------------------------------------------------------+
|
+----------------------------------------------------------+
|
||||||
|
|
|
|
||||||
v
|
v
|
||||||
@@ -128,6 +130,7 @@ table, and a senior reviewer at the end sorts it all into one clean report.
|
|||||||
| 6 | Brain | Chief estimator | Deduplicates, judges, and prioritizes all findings |
|
| 6 | Brain | Chief estimator | Deduplicates, judges, and prioritizes all findings |
|
||||||
| 6.5 | Brain (clarification) | Second opinion | For findings it distrusts, sends them back to the fact-checker to re-read the sheet; debunked findings are dropped |
|
| 6.5 | Brain (clarification) | Second opinion | For findings it distrusts, sends them back to the fact-checker to re-read the sheet; debunked findings are dropped |
|
||||||
| 7 | Review Gate | Your desk | Presents the findings a human should approve before anything goes out |
|
| 7 | Review Gate | Your desk | Presents the findings a human should approve before anything goes out |
|
||||||
|
| 7 | Review Chat | The analyst you can question | Answers "why did it decide that?" and "why didn't it check that?" from the run's own records — it explains, it never changes anything |
|
||||||
| 8 | RFI Writer | Secretary | Writes the formal clarification letters for confirmed issues |
|
| 8 | RFI Writer | Secretary | Writes the formal clarification letters for confirmed issues |
|
||||||
|
|
||||||
Everything the assistants learn is kept in a shared notebook (the "project
|
Everything the assistants learn is kept in a shared notebook (the "project
|
||||||
@@ -182,6 +185,19 @@ the whole review.
|
|||||||
queue better — the system already records your feedback, so it can get
|
queue better — the system already records your feedback, so it can get
|
||||||
smarter over time about what actually needs your eyes.
|
smarter over time about what actually needs your eyes.
|
||||||
|
|
||||||
|
### 5b. Explaining itself — "why did you think that?" ✅ *shipped*
|
||||||
|
- **Done:** every finding on the review screen has a chat panel, plus one for
|
||||||
|
the run as a whole. Ask why a unit was read as ground-mounted, or why a whole
|
||||||
|
discipline never got looked at, and it traces the answer back through what it
|
||||||
|
actually recorded — quoting the note it read off the sheet, or naming the
|
||||||
|
stage that skipped the pages. It cannot change a finding; that stays yours.
|
||||||
|
- **Done:** when you correct it in conversation ("that's not a floor drain,
|
||||||
|
it's a power floor box"), the correction is filed as structured data rather
|
||||||
|
than a free-text comment.
|
||||||
|
- **Still open:** nothing reads those filed corrections back yet. The next step
|
||||||
|
is priming a new run with what reviewers corrected on previous sets, so the
|
||||||
|
same misread does not come back on the next job.
|
||||||
|
|
||||||
### 6. Trust — "show the receipts" ✅ *partially shipped*
|
### 6. Trust — "show the receipts" ✅ *partially shipped*
|
||||||
- **Done:** the fact-checker already pulls a zoomed-in crop of the exact spot on
|
- **Done:** the fact-checker already pulls a zoomed-in crop of the exact spot on
|
||||||
the sheet when it re-reads a finding.
|
the sheet when it re-reads a finding.
|
||||||
@@ -192,5 +208,6 @@ the whole review.
|
|||||||
|
|
||||||
*Technical reference for the curious: the pipeline lives in
|
*Technical reference for the curious: the pipeline lives in
|
||||||
`backend/agents/runner.py` (the waves above are the "Agent wave N" stages), the
|
`backend/agents/runner.py` (the waves above are the "Agent wave N" stages), the
|
||||||
team's shared notebook is `backend/agents/memory.py`, and the review queue is
|
team's shared notebook is `backend/agents/memory.py`, the review queue is
|
||||||
`backend/review/gate.py` + `backend/review/finalizer.py`.*
|
`backend/review/gate.py` + `backend/review/finalizer.py`, and the review chat is
|
||||||
|
`backend/review/chat.py` + `backend/review/chat_context.py`.*
|
||||||
+153
-2
@@ -82,6 +82,29 @@
|
|||||||
border:1px solid var(--line); border-radius:6px; padding:6px 8px; font-size:13px; margin-top:6px; }
|
border:1px solid var(--line); border-radius:6px; padding:6px 8px; font-size:13px; margin-top:6px; }
|
||||||
.review-controls input[type=text] { width:100%; }
|
.review-controls input[type=text] { width:100%; }
|
||||||
.review-controls .hidden { display:none; }
|
.review-controls .hidden { display:none; }
|
||||||
|
.chat { margin-top:10px; padding-top:10px; border-top:1px solid var(--line); font-size:13px; }
|
||||||
|
.chat-toggle { background:none; border:0; color:var(--accent); cursor:pointer; padding:0;
|
||||||
|
font-size:13px; text-decoration:underline dotted; }
|
||||||
|
.chat-body { margin-top:10px; }
|
||||||
|
.chat-body.hidden, .chat .hidden { display:none; }
|
||||||
|
.chat-turns { max-height:340px; overflow-y:auto; margin-bottom:8px; }
|
||||||
|
.chat-q, .chat-a { border-radius:8px; padding:8px 10px; margin:6px 0; }
|
||||||
|
.chat-q { background:#161b25; }
|
||||||
|
.chat-q b { color:var(--accent); }
|
||||||
|
.chat-a { background:#11141a; }
|
||||||
|
.chat-a ul { margin:6px 0 0; padding-left:18px; }
|
||||||
|
.chat-a li { margin:2px 0; }
|
||||||
|
.chat-cite { color:var(--muted); font-size:12px; margin-top:6px; }
|
||||||
|
.chat-cite code { color:var(--accent); }
|
||||||
|
.chat-tags { color:var(--muted); font-size:12px; margin-top:6px; }
|
||||||
|
.chat-row { display:flex; gap:8px; align-items:flex-start; }
|
||||||
|
.chat-row textarea { flex:1; background:#0c0e13; color:var(--text); border:1px solid var(--line);
|
||||||
|
border-radius:6px; padding:8px; font-size:13px; font-family:inherit; resize:vertical;
|
||||||
|
min-height:38px; }
|
||||||
|
.chat-row button { white-space:nowrap; }
|
||||||
|
.chat-hint { color:var(--muted); font-size:12px; margin-top:6px; }
|
||||||
|
.chat-err { color:var(--hi); font-size:12px; margin-top:6px; }
|
||||||
|
.btn.sm { padding:8px 14px; font-size:13px; margin-top:0; }
|
||||||
.pill.blocking { background:rgba(255,93,87,.15); color:var(--hi); }
|
.pill.blocking { background:rgba(255,93,87,.15); color:var(--hi); }
|
||||||
.pill.audit { background:rgba(91,140,255,.15); color:var(--accent); }
|
.pill.audit { background:rgba(91,140,255,.15); color:var(--accent); }
|
||||||
.pill.critical { background:rgba(255,93,87,.28); color:#fff; }
|
.pill.critical { background:rgba(255,93,87,.28); color:#fff; }
|
||||||
@@ -536,6 +559,12 @@ async function renderReview(job){
|
|||||||
esc(prog.completed||0)+' of '+esc(prog.required||0)+' required items decided.'+
|
esc(prog.completed||0)+' of '+esc(prog.required||0)+' required items decided.'+
|
||||||
((prog.remaining||0)>0?' Decide all blocking items, save, then finalize.':
|
((prog.remaining||0)>0?' Decide all blocking items, save, then finalize.':
|
||||||
' All required items decided \u2014 you can finalize.')+'</div>';
|
' All required items decided \u2014 you can finalize.')+'</div>';
|
||||||
|
html+='<div class="conflict" style="border-left-color:var(--accent)">'+
|
||||||
|
'<div class="row"><span class="cat">Ask about this run</span></div>'+
|
||||||
|
'<div class="meta">Questions about coverage or about the run as a whole — '+
|
||||||
|
'e.g. "why didn\'t it pick up the Civil set?". Answers are explanations only; '+
|
||||||
|
'they never change a finding.</div>'+
|
||||||
|
chatPanelHtml('','Ask a question about this run',true)+'</div>';
|
||||||
const blocking=queue.filter(i=>i.blocking), audit=queue.filter(i=>!i.blocking);
|
const blocking=queue.filter(i=>i.blocking), audit=queue.filter(i=>!i.blocking);
|
||||||
blocking.forEach((item,i)=>{ html+=reviewItemHtml(item,'b'+i,prior[item.review_item_id]); });
|
blocking.forEach((item,i)=>{ html+=reviewItemHtml(item,'b'+i,prior[item.review_item_id]); });
|
||||||
if(audit.length){
|
if(audit.length){
|
||||||
@@ -548,7 +577,11 @@ async function renderReview(job){
|
|||||||
'<button class="btn" id="saveReviewBtn">Save decisions</button> '+
|
'<button class="btn" id="saveReviewBtn">Save decisions</button> '+
|
||||||
'<button class="btn" id="finalizeBtn"'+((prog.remaining||0)===0?'':' disabled')+
|
'<button class="btn" id="finalizeBtn"'+((prog.remaining||0)===0?'':' disabled')+
|
||||||
'>Finalize & send report</button></div>'+
|
'>Finalize & send report</button></div>'+
|
||||||
'<div class="status" id="reviewMsg"></div>';
|
'<div class="status" id="reviewMsg"></div>'+
|
||||||
|
'<div class="meta" style="margin-top:10px">Chat transcript: '+
|
||||||
|
'<a href="/jobs/'+esc(jobId)+'/review-chat/log" target="_blank" rel="noopener">'+
|
||||||
|
'/jobs/'+esc(jobId)+'/review-chat/log</a> '+
|
||||||
|
'(also saved as review/chat_log.jsonl on the server)</div>';
|
||||||
results.innerHTML=html;
|
results.innerHTML=html;
|
||||||
results.querySelectorAll('.review-item input[type=radio]').forEach(r=>{
|
results.querySelectorAll('.review-item input[type=radio]').forEach(r=>{
|
||||||
r.addEventListener('change',()=>syncReviewControls(r.closest('.review-item')));
|
r.addEventListener('change',()=>syncReviewControls(r.closest('.review-item')));
|
||||||
@@ -559,6 +592,8 @@ async function renderReview(job){
|
|||||||
});
|
});
|
||||||
document.getElementById('saveReviewBtn').addEventListener('click',saveReviewDecisions);
|
document.getElementById('saveReviewBtn').addEventListener('click',saveReviewDecisions);
|
||||||
document.getElementById('finalizeBtn').addEventListener('click',finalizeReview);
|
document.getElementById('finalizeBtn').addEventListener('click',finalizeReview);
|
||||||
|
wireChatPanels();
|
||||||
|
loadChatHistory();
|
||||||
}
|
}
|
||||||
|
|
||||||
function reviewItemHtml(item,uid,prev){
|
function reviewItemHtml(item,uid,prev){
|
||||||
@@ -600,10 +635,126 @@ function reviewItemHtml(item,uid,prev){
|
|||||||
'<input type="text" class="comment" placeholder="Comment (optional)" value="'+escAttr(prev.comment||'')+'">'+
|
'<input type="text" class="comment" placeholder="Comment (optional)" value="'+escAttr(prev.comment||'')+'">'+
|
||||||
'<input type="text" class="clar'+(prev.decision==='needs_clarification'?'':' hidden')+
|
'<input type="text" class="clar'+(prev.decision==='needs_clarification'?'':' hidden')+
|
||||||
'" placeholder="Clarification answer" value="'+escAttr(prev.clarification_answer||'')+'">'+
|
'" placeholder="Clarification answer" value="'+escAttr(prev.clarification_answer||'')+'">'+
|
||||||
'</div></div>';
|
'</div>'+
|
||||||
|
chatPanelHtml(item.review_item_id,'Ask about this finding')+
|
||||||
|
'</div>';
|
||||||
return html;
|
return html;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// --- review chat (read-only run explainer) ---
|
||||||
|
// One panel per review item plus one run-scope panel. Panels are keyed by
|
||||||
|
// review_item_id ('' = run scope); history for every panel is fetched once.
|
||||||
|
|
||||||
|
function chatPanelHtml(key,label,openByDefault){
|
||||||
|
const open=!!openByDefault;
|
||||||
|
return '<div class="chat" data-chat="'+escAttr(key||'')+'">'+
|
||||||
|
'<button type="button" class="chat-toggle">'+esc(label)+'</button>'+
|
||||||
|
'<div class="chat-body'+(open?'':' hidden')+'">'+
|
||||||
|
'<div class="chat-turns"></div>'+
|
||||||
|
'<div class="chat-row">'+
|
||||||
|
'<textarea rows="2" placeholder="Why did it conclude that?"></textarea>'+
|
||||||
|
'<button type="button" class="btn sm chat-send">Ask</button>'+
|
||||||
|
'</div>'+
|
||||||
|
'<div class="chat-hint">Explains what the run did, from its own artifacts. '+
|
||||||
|
'It cannot change this finding or your decision.</div>'+
|
||||||
|
'<div class="chat-err"></div>'+
|
||||||
|
'</div></div>';
|
||||||
|
}
|
||||||
|
|
||||||
|
function chatTurnHtml(turn){
|
||||||
|
let html='<div class="chat-q"><b>You:</b> '+esc(turn.question||'')+'</div>'+
|
||||||
|
'<div class="chat-a">'+esc(turn.answer||'');
|
||||||
|
if((turn.findings||[]).length){
|
||||||
|
html+='<ul>'+turn.findings.map(f=>'<li>'+esc(f)+'</li>').join('')+'</ul>';
|
||||||
|
}
|
||||||
|
(turn.evidence_cited||[]).forEach(c=>{
|
||||||
|
html+='<div class="chat-cite"><code>'+esc(c.artifact||'?')+'</code>'+
|
||||||
|
(c.sheet?' ('+esc(c.sheet)+')':'')+
|
||||||
|
(c.quote?': "'+esc(c.quote)+'"':'')+
|
||||||
|
(c.why_it_matters?' — '+esc(c.why_it_matters):'')+'</div>';
|
||||||
|
});
|
||||||
|
const tags=[];
|
||||||
|
if(turn.answerable&&turn.answerable!=='yes') tags.push('answerable: '+turn.answerable);
|
||||||
|
if(turn.assessment_of_finding&&turn.assessment_of_finding!=='not_applicable')
|
||||||
|
tags.push(turn.assessment_of_finding.replace(/_/g,' '));
|
||||||
|
if(turn.confidence) tags.push('confidence: '+turn.confidence);
|
||||||
|
if(turn.missing_information) tags.push('missing: '+turn.missing_information);
|
||||||
|
if(turn.suggested_category_correction)
|
||||||
|
tags.push('correction noted: '+turn.suggested_category_correction);
|
||||||
|
if(tags.length) html+='<div class="chat-tags">'+esc(tags.join(' \u00b7 '))+'</div>';
|
||||||
|
return html+'</div>';
|
||||||
|
}
|
||||||
|
|
||||||
|
function renderChatTurns(panel,turns){
|
||||||
|
const box=panel.querySelector('.chat-turns');
|
||||||
|
box.innerHTML=turns.length?turns.map(chatTurnHtml).join('')
|
||||||
|
:'<div class="meta">No questions asked yet.</div>';
|
||||||
|
box.scrollTop=box.scrollHeight;
|
||||||
|
}
|
||||||
|
|
||||||
|
async function loadChatHistory(){
|
||||||
|
let data;
|
||||||
|
try{
|
||||||
|
const res=await fetch('/jobs/'+currentJobId+'/review-chat');
|
||||||
|
if(!res.ok) return; // chat unavailable for this job: leave panels empty
|
||||||
|
data=await res.json();
|
||||||
|
}catch(err){ return; }
|
||||||
|
const byKey={};
|
||||||
|
(data.turns||[]).forEach(t=>{
|
||||||
|
const k=t.review_item_id||'';
|
||||||
|
(byKey[k]=byKey[k]||[]).push(t);
|
||||||
|
});
|
||||||
|
results.querySelectorAll('.chat').forEach(panel=>{
|
||||||
|
const key=panel.getAttribute('data-chat')||'';
|
||||||
|
renderChatTurns(panel,byKey[key]||[]);
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
function wireChatPanels(){
|
||||||
|
results.querySelectorAll('.chat').forEach(panel=>{
|
||||||
|
panel.querySelector('.chat-toggle').addEventListener('click',()=>{
|
||||||
|
panel.querySelector('.chat-body').classList.toggle('hidden');
|
||||||
|
});
|
||||||
|
const send=panel.querySelector('.chat-send');
|
||||||
|
const box=panel.querySelector('textarea');
|
||||||
|
send.addEventListener('click',()=>askChat(panel));
|
||||||
|
// Enter sends, Shift+Enter newlines - the questions are usually one line.
|
||||||
|
box.addEventListener('keydown',e=>{
|
||||||
|
if(e.key==='Enter'&&!e.shiftKey){ e.preventDefault(); askChat(panel); }
|
||||||
|
});
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async function askChat(panel){
|
||||||
|
const box=panel.querySelector('textarea');
|
||||||
|
const send=panel.querySelector('.chat-send');
|
||||||
|
const err=panel.querySelector('.chat-err');
|
||||||
|
const question=box.value.trim();
|
||||||
|
err.textContent='';
|
||||||
|
if(!question) return;
|
||||||
|
const key=panel.getAttribute('data-chat')||'';
|
||||||
|
const body={question:question};
|
||||||
|
if(key) body.review_item_id=key;
|
||||||
|
send.disabled=true; send.textContent='Asking...';
|
||||||
|
try{
|
||||||
|
const res=await fetch('/jobs/'+currentJobId+'/review-chat',
|
||||||
|
{method:'POST',headers:{'Content-Type':'application/json'},
|
||||||
|
body:JSON.stringify(body)});
|
||||||
|
if(!res.ok){
|
||||||
|
const e=await res.json().catch(()=>({detail:res.statusText}));
|
||||||
|
const detail=typeof e.detail==='string'?e.detail:JSON.stringify(e.detail);
|
||||||
|
throw new Error(detail||'Request failed');
|
||||||
|
}
|
||||||
|
const data=await res.json();
|
||||||
|
const turnsBox=panel.querySelector('.chat-turns');
|
||||||
|
if(turnsBox.querySelector('.meta')) turnsBox.innerHTML='';
|
||||||
|
turnsBox.insertAdjacentHTML('beforeend',chatTurnHtml(data.turn||{}));
|
||||||
|
turnsBox.scrollTop=turnsBox.scrollHeight;
|
||||||
|
box.value='';
|
||||||
|
}catch(e){ err.textContent=e.message; }
|
||||||
|
finally{ send.disabled=false; send.textContent='Ask'; }
|
||||||
|
}
|
||||||
|
|
||||||
function syncReviewControls(el){
|
function syncReviewControls(el){
|
||||||
const sel=el.querySelector('input[type=radio]:checked');
|
const sel=el.querySelector('input[type=radio]:checked');
|
||||||
const v=sel?sel.value:'';
|
const v=sel?sel.value:'';
|
||||||
|
|||||||
@@ -0,0 +1,139 @@
|
|||||||
|
"""API tests for the review-chat endpoints."""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
from fastapi.testclient import TestClient
|
||||||
|
|
||||||
|
from backend.main import app
|
||||||
|
from backend.review.store import ReviewStore
|
||||||
|
|
||||||
|
|
||||||
|
def _queue_item() -> dict:
|
||||||
|
return {"review_item_id": "finding:AGENT-0007", "kind": "finding",
|
||||||
|
"blocking": True, "reasons": ["severity_high"],
|
||||||
|
"payload": {"issue_id": "AGENT-0007", "category": "elevation_disagreement",
|
||||||
|
"severity": "high", "sheets": ["M2.1"],
|
||||||
|
"description": "AC-1 at grade vs roof.",
|
||||||
|
"scope_id": "conflict:roof-ac1"}}
|
||||||
|
|
||||||
|
|
||||||
|
def _reply() -> dict:
|
||||||
|
return {"answer": "It read 'AC-1 MOUNTED ON GRADE' off M2.1.",
|
||||||
|
"findings": ["The grade value came from M2.1."],
|
||||||
|
"evidence_cited": [], "answerable": "yes",
|
||||||
|
"assessment_of_finding": "looks_supported", "confidence": "high"}
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
def job(monkeypatch, tmp_path):
|
||||||
|
"""A finished, review-gated job with one queued finding and a stub model."""
|
||||||
|
store = ReviewStore(str(tmp_path))
|
||||||
|
store.write_queue([_queue_item()])
|
||||||
|
monkeypatch.setattr("backend.main.get_job", lambda job_id: {
|
||||||
|
"job_id": job_id, "status": "needs_review", "out_dir": str(tmp_path)})
|
||||||
|
monkeypatch.setattr("backend.review.chat.call_json", lambda **kw: _reply())
|
||||||
|
return str(tmp_path)
|
||||||
|
|
||||||
|
|
||||||
|
def test_ask_about_a_finding_returns_and_logs_a_turn(job):
|
||||||
|
client = TestClient(app)
|
||||||
|
response = client.post("/jobs/job1/review-chat", json={
|
||||||
|
"question": "Why does it think AC-1 is at grade?",
|
||||||
|
"review_item_id": "finding:AGENT-0007"})
|
||||||
|
assert response.status_code == 200
|
||||||
|
turn = response.json()["turn"]
|
||||||
|
assert turn["issue"]["issue_id"] == "AGENT-0007"
|
||||||
|
assert turn["findings"] == ["The grade value came from M2.1."]
|
||||||
|
with open(os.path.join(job, "review", "chat_log.jsonl"), encoding="utf-8") as f:
|
||||||
|
assert len([line for line in f if line.strip()]) == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_ask_about_the_run_needs_no_item(job):
|
||||||
|
client = TestClient(app)
|
||||||
|
response = client.post("/jobs/job1/review-chat",
|
||||||
|
json={"question": "Why didn't it pick up the Civil set?"})
|
||||||
|
assert response.status_code == 200
|
||||||
|
assert response.json()["turn"]["scope"] == "run"
|
||||||
|
|
||||||
|
|
||||||
|
def test_history_endpoint_filters_by_item(job):
|
||||||
|
client = TestClient(app)
|
||||||
|
client.post("/jobs/job1/review-chat", json={
|
||||||
|
"question": "Why grade?", "review_item_id": "finding:AGENT-0007"})
|
||||||
|
client.post("/jobs/job1/review-chat", json={"question": "Why no Civil?"})
|
||||||
|
assert len(client.get("/jobs/job1/review-chat").json()["turns"]) == 2
|
||||||
|
filtered = client.get("/jobs/job1/review-chat",
|
||||||
|
params={"review_item_id": "finding:AGENT-0007"}).json()
|
||||||
|
assert len(filtered["turns"]) == 1
|
||||||
|
assert filtered["turns"][0]["question"] == "Why grade?"
|
||||||
|
|
||||||
|
|
||||||
|
def test_transcript_endpoint_renders_markdown(job):
|
||||||
|
client = TestClient(app)
|
||||||
|
client.post("/jobs/job1/review-chat", json={
|
||||||
|
"question": "Why grade?", "review_item_id": "finding:AGENT-0007"})
|
||||||
|
response = client.get("/jobs/job1/review-chat/log")
|
||||||
|
assert response.status_code == 200
|
||||||
|
assert response.headers["content-type"].startswith("text/markdown")
|
||||||
|
assert "## AGENT-0007" in response.text
|
||||||
|
assert "Why grade?" in response.text
|
||||||
|
|
||||||
|
|
||||||
|
def test_blank_question_is_422(job):
|
||||||
|
client = TestClient(app)
|
||||||
|
assert client.post("/jobs/job1/review-chat", json={"question": " "}).status_code == 422
|
||||||
|
|
||||||
|
|
||||||
|
def test_unknown_item_is_422(job):
|
||||||
|
client = TestClient(app)
|
||||||
|
response = client.post("/jobs/job1/review-chat",
|
||||||
|
json={"question": "why?", "review_item_id": "finding:NOPE"})
|
||||||
|
assert response.status_code == 422
|
||||||
|
|
||||||
|
|
||||||
|
def test_model_failure_is_502_not_500(monkeypatch, job):
|
||||||
|
monkeypatch.setattr("backend.review.chat.call_json", lambda **kw: None)
|
||||||
|
client = TestClient(app)
|
||||||
|
response = client.post("/jobs/job1/review-chat", json={"question": "why?"})
|
||||||
|
assert response.status_code == 502
|
||||||
|
|
||||||
|
|
||||||
|
def test_chat_stays_available_after_the_job_is_done(monkeypatch, tmp_path):
|
||||||
|
"""The chat is read-only, so a finalized report can still be questioned."""
|
||||||
|
ReviewStore(str(tmp_path)).write_queue([_queue_item()])
|
||||||
|
monkeypatch.setattr("backend.main.get_job", lambda job_id: {
|
||||||
|
"job_id": job_id, "status": "done", "out_dir": str(tmp_path)})
|
||||||
|
monkeypatch.setattr("backend.review.chat.call_json", lambda **kw: _reply())
|
||||||
|
client = TestClient(app)
|
||||||
|
assert client.post("/jobs/job1/review-chat",
|
||||||
|
json={"question": "why?"}).status_code == 200
|
||||||
|
|
||||||
|
|
||||||
|
def test_chat_is_409_while_the_job_is_still_running(monkeypatch, tmp_path):
|
||||||
|
monkeypatch.setattr("backend.main.get_job", lambda job_id: {
|
||||||
|
"job_id": job_id, "status": "running", "out_dir": str(tmp_path)})
|
||||||
|
client = TestClient(app)
|
||||||
|
assert client.post("/jobs/job1/review-chat",
|
||||||
|
json={"question": "why?"}).status_code == 409
|
||||||
|
|
||||||
|
|
||||||
|
def test_chat_404s_for_unknown_job(monkeypatch, tmp_path):
|
||||||
|
monkeypatch.setattr("backend.main.get_job", lambda job_id: None)
|
||||||
|
client = TestClient(app)
|
||||||
|
assert client.post("/jobs/nope/review-chat",
|
||||||
|
json={"question": "why?"}).status_code == 404
|
||||||
|
|
||||||
|
|
||||||
|
def test_chat_never_mutates_review_decisions(job):
|
||||||
|
"""The whole point: asking questions cannot change the review state."""
|
||||||
|
client = TestClient(app)
|
||||||
|
before = ReviewStore(job, create=False).read_decisions()
|
||||||
|
client.post("/jobs/job1/review-chat", json={
|
||||||
|
"question": "This is wrong, reject it.",
|
||||||
|
"review_item_id": "finding:AGENT-0007"})
|
||||||
|
after = ReviewStore(job, create=False).read_decisions()
|
||||||
|
assert before == after == {}
|
||||||
|
with open(os.path.join(job, "review", "review_queue.json"), encoding="utf-8") as f:
|
||||||
|
assert json.load(f) == [_queue_item()]
|
||||||
@@ -0,0 +1,15 @@
|
|||||||
|
"""Shared test fixtures."""
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture(autouse=True)
|
||||||
|
def isolated_feedback_store(tmp_path, monkeypatch):
|
||||||
|
"""Keep the cross-job feedback store out of the real outputs directory.
|
||||||
|
|
||||||
|
Saving a review decision or asking a chat question appends to
|
||||||
|
config.REVIEW_FEEDBACK_DIR, which is process-wide rather than job-local.
|
||||||
|
Without this, running the suite would accumulate junk in backend/outputs.
|
||||||
|
"""
|
||||||
|
monkeypatch.setattr("backend.config.REVIEW_FEEDBACK_DIR",
|
||||||
|
str(tmp_path / "_feedback"))
|
||||||
@@ -0,0 +1,172 @@
|
|||||||
|
"""Review chat: answer normalization, logging, and the feedback roll-up."""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from backend import config
|
||||||
|
from backend.review import chat
|
||||||
|
|
||||||
|
|
||||||
|
def _queue() -> list:
|
||||||
|
return [{
|
||||||
|
"review_item_id": "finding:AGENT-0007", "kind": "finding", "blocking": True,
|
||||||
|
"reasons": ["severity_high"],
|
||||||
|
"payload": {"issue_id": "AGENT-0007", "source_stage": "conflict",
|
||||||
|
"category": "elevation_disagreement", "severity": "high",
|
||||||
|
"confidence": "medium", "location": "Roof / AC-1",
|
||||||
|
"sheets": ["M2.1"], "description": "AC-1 at grade vs roof.",
|
||||||
|
"evidence": [], "scope_id": "conflict:roof-ac1"},
|
||||||
|
}]
|
||||||
|
|
||||||
|
|
||||||
|
def _model_reply(**overrides) -> dict:
|
||||||
|
reply = {
|
||||||
|
"answer": "The extractor read 'AC-1 MOUNTED ON GRADE' off M2.1.",
|
||||||
|
"findings": ["The grade reading came from M2.1's text layer."],
|
||||||
|
"evidence_cited": [{"artifact": "source_sheets[M2.1]", "sheet": "M2.1",
|
||||||
|
"quote": "AC-1 MOUNTED ON GRADE",
|
||||||
|
"why_it_matters": "It is the sole basis for 'grade'."}],
|
||||||
|
"answerable": "yes",
|
||||||
|
"missing_information": None,
|
||||||
|
"assessment_of_finding": "looks_supported",
|
||||||
|
"suggested_category_correction": None,
|
||||||
|
"confidence": "high",
|
||||||
|
}
|
||||||
|
reply.update(overrides)
|
||||||
|
return reply
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
def fake_llm(monkeypatch):
|
||||||
|
"""Stub the model; the chat must never need a network to be tested."""
|
||||||
|
calls = []
|
||||||
|
|
||||||
|
def _call(**kwargs):
|
||||||
|
calls.append(kwargs)
|
||||||
|
return calls_reply[0]
|
||||||
|
|
||||||
|
calls_reply = [_model_reply()]
|
||||||
|
monkeypatch.setattr("backend.review.chat.call_json", lambda **kw: _call(**kw))
|
||||||
|
return calls, calls_reply
|
||||||
|
|
||||||
|
|
||||||
|
def test_ask_logs_issue_question_and_findings(tmp_path, fake_llm):
|
||||||
|
turn = chat.ask("job1", str(tmp_path), "Why is AC-1 at grade?",
|
||||||
|
review_item_id="finding:AGENT-0007", queue=_queue())
|
||||||
|
assert turn["question"] == "Why is AC-1 at grade?"
|
||||||
|
assert turn["findings"] == ["The grade reading came from M2.1's text layer."]
|
||||||
|
# The log's "issue in question" is a snapshot, not a bare id.
|
||||||
|
assert turn["issue"]["issue_id"] == "AGENT-0007"
|
||||||
|
assert turn["issue"]["severity"] == "high"
|
||||||
|
path = os.path.join(str(tmp_path), "review", "chat_log.jsonl")
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
logged = [json.loads(line) for line in f if line.strip()]
|
||||||
|
assert len(logged) == 1
|
||||||
|
assert logged[0]["turn_id"] == turn["turn_id"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_ask_appends_to_the_cross_job_feedback_store(tmp_path, fake_llm):
|
||||||
|
chat.ask("job1", str(tmp_path), "Is this really a floor drain?",
|
||||||
|
review_item_id="finding:AGENT-0007", queue=_queue())
|
||||||
|
path = os.path.join(config.REVIEW_FEEDBACK_DIR, "chat_turns.jsonl")
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
records = [json.loads(line) for line in f if line.strip()]
|
||||||
|
assert records[0]["kind"] == "review_chat_turn"
|
||||||
|
assert records[0]["issue_id"] == "AGENT-0007"
|
||||||
|
assert records[0]["job_id"] == "job1"
|
||||||
|
|
||||||
|
|
||||||
|
def test_correction_signal_is_captured_as_structured_data(tmp_path, fake_llm):
|
||||||
|
"""A misidentification correction survives as a field, not free text."""
|
||||||
|
_, reply = fake_llm
|
||||||
|
reply[0] = _model_reply(suggested_category_correction="power floor box",
|
||||||
|
assessment_of_finding="looks_unsupported")
|
||||||
|
turn = chat.ask("job1", str(tmp_path), "That is not a floor drain.",
|
||||||
|
review_item_id="finding:AGENT-0007", queue=_queue())
|
||||||
|
assert turn["suggested_category_correction"] == "power floor box"
|
||||||
|
path = os.path.join(config.REVIEW_FEEDBACK_DIR, "chat_turns.jsonl")
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
record = json.loads(f.readline())
|
||||||
|
assert record["suggested_category_correction"] == "power floor box"
|
||||||
|
assert record["assessment_of_finding"] == "looks_unsupported"
|
||||||
|
|
||||||
|
|
||||||
|
def test_run_scope_question_needs_no_item(tmp_path, fake_llm):
|
||||||
|
turn = chat.ask("job1", str(tmp_path), "Why didn't it pick up the Civil set?")
|
||||||
|
assert turn["review_item_id"] is None
|
||||||
|
assert turn["scope"] == "run"
|
||||||
|
assert turn["issue"] is None
|
||||||
|
|
||||||
|
|
||||||
|
def test_history_is_replayed_for_the_same_thread(tmp_path, fake_llm):
|
||||||
|
calls, _ = fake_llm
|
||||||
|
chat.ask("job1", str(tmp_path), "First question?",
|
||||||
|
review_item_id="finding:AGENT-0007", queue=_queue())
|
||||||
|
chat.ask("job1", str(tmp_path), "Follow-up?",
|
||||||
|
review_item_id="finding:AGENT-0007", queue=_queue())
|
||||||
|
assert "First question?" in calls[1]["user_text"]
|
||||||
|
# A run-scope turn must not inherit an item thread's history.
|
||||||
|
chat.ask("job1", str(tmp_path), "Unrelated run question?")
|
||||||
|
assert "First question?" not in calls[2]["user_text"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_blank_and_oversized_questions_are_rejected(tmp_path, fake_llm):
|
||||||
|
with pytest.raises(chat.ChatError):
|
||||||
|
chat.ask("job1", str(tmp_path), " ")
|
||||||
|
with pytest.raises(chat.ChatError):
|
||||||
|
chat.ask("job1", str(tmp_path),
|
||||||
|
"x" * (config.REVIEW_CHAT_MAX_QUESTION_CHARS + 1))
|
||||||
|
|
||||||
|
|
||||||
|
def test_unknown_review_item_is_rejected(tmp_path, fake_llm):
|
||||||
|
with pytest.raises(chat.ChatError):
|
||||||
|
chat.ask("job1", str(tmp_path), "why?", review_item_id="finding:NOPE",
|
||||||
|
queue=_queue())
|
||||||
|
|
||||||
|
|
||||||
|
def test_unusable_model_reply_raises_and_logs_nothing(tmp_path, monkeypatch):
|
||||||
|
monkeypatch.setattr("backend.review.chat.call_json", lambda **kw: None)
|
||||||
|
with pytest.raises(RuntimeError):
|
||||||
|
chat.ask("job1", str(tmp_path), "why?")
|
||||||
|
assert not os.path.exists(os.path.join(str(tmp_path), "review", "chat_log.jsonl"))
|
||||||
|
|
||||||
|
|
||||||
|
def test_bad_enum_values_fall_back_instead_of_failing(tmp_path, fake_llm):
|
||||||
|
_, reply = fake_llm
|
||||||
|
reply[0] = _model_reply(answerable="probably", confidence="",
|
||||||
|
assessment_of_finding="made_up")
|
||||||
|
turn = chat.ask("job1", str(tmp_path), "why?")
|
||||||
|
assert turn["answerable"] == "partial"
|
||||||
|
assert turn["confidence"] == "low"
|
||||||
|
assert turn["assessment_of_finding"] == "cannot_tell"
|
||||||
|
|
||||||
|
|
||||||
|
def test_disabled_chat_refuses(tmp_path, monkeypatch, fake_llm):
|
||||||
|
monkeypatch.setattr("backend.config.ENABLE_REVIEW_CHAT", False)
|
||||||
|
with pytest.raises(chat.ChatError):
|
||||||
|
chat.ask("job1", str(tmp_path), "why?")
|
||||||
|
|
||||||
|
|
||||||
|
def test_read_log_skips_corrupt_lines(tmp_path, fake_llm):
|
||||||
|
chat.ask("job1", str(tmp_path), "why?")
|
||||||
|
path = os.path.join(str(tmp_path), "review", "chat_log.jsonl")
|
||||||
|
with open(path, "a", encoding="utf-8") as f:
|
||||||
|
f.write("{not json\n")
|
||||||
|
assert len(chat.read_log(str(tmp_path))) == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_markdown_transcript_groups_by_issue(tmp_path, fake_llm):
|
||||||
|
chat.ask("job1", str(tmp_path), "Why is AC-1 at grade?",
|
||||||
|
review_item_id="finding:AGENT-0007", queue=_queue())
|
||||||
|
chat.ask("job1", str(tmp_path), "Why no Civil?")
|
||||||
|
markdown = chat.render_log_markdown(chat.read_log(str(tmp_path)))
|
||||||
|
assert "## AGENT-0007" in markdown
|
||||||
|
assert "## Run-scope questions" in markdown
|
||||||
|
assert "Why is AC-1 at grade?" in markdown
|
||||||
|
assert "**Findings**" in markdown
|
||||||
|
|
||||||
|
|
||||||
|
def test_markdown_transcript_handles_empty_log():
|
||||||
|
assert "No questions" in chat.render_log_markdown([])
|
||||||
@@ -0,0 +1,137 @@
|
|||||||
|
"""Context bundles for the review chat: what the model is allowed to see."""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
|
||||||
|
from backend.review.chat_context import build_context
|
||||||
|
|
||||||
|
|
||||||
|
def _write(out_dir: str, name: str, value) -> None:
|
||||||
|
path = os.path.join(out_dir, name)
|
||||||
|
os.makedirs(os.path.dirname(path), exist_ok=True)
|
||||||
|
with open(path, "w", encoding="utf-8") as f:
|
||||||
|
json.dump(value, f)
|
||||||
|
|
||||||
|
|
||||||
|
def _finding() -> dict:
|
||||||
|
return {
|
||||||
|
"issue_id": "AGENT-0007",
|
||||||
|
"source_stage": "conflict",
|
||||||
|
"category": "elevation_disagreement",
|
||||||
|
"severity": "high",
|
||||||
|
"confidence": "medium",
|
||||||
|
"location": "Roof / AC-1",
|
||||||
|
"disciplines": ["Mechanical"],
|
||||||
|
"sheets": ["M2.1"],
|
||||||
|
"description": "AC-1 shown at grade on M2.1 but on the roof elsewhere.",
|
||||||
|
"evidence": [{"discipline": "Mechanical", "sheet": "M2.1",
|
||||||
|
"source_text": "AC-1 MOUNTED ON GRADE", "asserted_value": "grade"}],
|
||||||
|
"scope_id": "conflict:roof-ac1",
|
||||||
|
"verification": {"status": "unverified", "verdicts": []},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _queue() -> list:
|
||||||
|
return [{"review_item_id": "finding:AGENT-0007", "kind": "finding",
|
||||||
|
"blocking": True, "reasons": ["severity_high"], "payload": _finding()}]
|
||||||
|
|
||||||
|
|
||||||
|
def _job_dir(tmp_path) -> str:
|
||||||
|
out_dir = str(tmp_path)
|
||||||
|
_write(out_dir, "conflicts.json", {
|
||||||
|
"source": "set.pdf",
|
||||||
|
"summary": {"pipeline_mode": "agent", "agent_status": "needs_review",
|
||||||
|
"by_stage": {"conflicts": 3}, "conflicts_found": 3},
|
||||||
|
"sheet_index": {"sheet_index": [
|
||||||
|
{"sheet_number": "M2.1", "discipline": "Mechanical"},
|
||||||
|
{"sheet_number": "A1.1", "discipline": "Architectural"},
|
||||||
|
]},
|
||||||
|
"sheet_reconciliation": {"declared_total": 4, "found_total": 2,
|
||||||
|
"declared_not_in_set": ["C-001", "C-101"],
|
||||||
|
"in_set_not_declared": []},
|
||||||
|
})
|
||||||
|
_write(out_dir, "agent/memory.json", {
|
||||||
|
"sheets": [{
|
||||||
|
"sheet_number": "M2.1", "discipline": "Mechanical", "page_number": 7,
|
||||||
|
"assertions": [{"attribute": "mounting", "value": "grade",
|
||||||
|
"source_text": "AC-1 MOUNTED ON GRADE",
|
||||||
|
"base64": "SHOULD-NOT-APPEAR"}],
|
||||||
|
}],
|
||||||
|
"clusters": [{"key": "roof-ac1", "location": "Roof / AC-1",
|
||||||
|
"disciplines": ["Mechanical", "Architectural"],
|
||||||
|
"assertions": [{"sheet_number": "M2.1", "attribute": "mounting",
|
||||||
|
"value": "grade", "base64": "SHOULD-NOT-APPEAR"}]}],
|
||||||
|
"decisions": [{"finding_refs": ["AGENT-0007"], "action": "kept",
|
||||||
|
"reason": "supported", "kept_issue_id": "AGENT-0007"}],
|
||||||
|
"suppressed": [],
|
||||||
|
})
|
||||||
|
return out_dir
|
||||||
|
|
||||||
|
|
||||||
|
def test_item_scope_carries_the_reasoning_chain(tmp_path):
|
||||||
|
context = build_context(_job_dir(tmp_path), "finding:AGENT-0007", _queue())
|
||||||
|
assert context["scope"] == "item"
|
||||||
|
assert context["finding"]["issue_id"] == "AGENT-0007"
|
||||||
|
# The chain a "why does it think X" answer has to walk.
|
||||||
|
assert context["originating_cluster"]["key"] == "roof-ac1"
|
||||||
|
assert context["source_sheets"][0]["sheet_number"] == "M2.1"
|
||||||
|
assert context["brain_decisions"][0]["action"] == "kept"
|
||||||
|
assert context["finding"]["verification"]["status"] == "unverified"
|
||||||
|
|
||||||
|
|
||||||
|
def test_context_never_leaks_base64(tmp_path):
|
||||||
|
"""Page images blow up the prompt and are useless as quotable evidence."""
|
||||||
|
context = build_context(_job_dir(tmp_path), "finding:AGENT-0007", _queue())
|
||||||
|
assert "SHOULD-NOT-APPEAR" not in json.dumps(context)
|
||||||
|
|
||||||
|
|
||||||
|
def test_run_scope_carries_coverage_material(tmp_path):
|
||||||
|
"""The 'why didn't it pick up the Civil set' inputs are all present."""
|
||||||
|
context = build_context(_job_dir(tmp_path), None, _queue(),
|
||||||
|
question="why didn't it pick up the Civil set?")
|
||||||
|
assert context["scope"] == "run"
|
||||||
|
assert set(context["sheets_by_discipline"]) == {"Mechanical", "Architectural"}
|
||||||
|
assert context["sheet_reconciliation"]["declared_not_in_set"] == ["C-001", "C-101"]
|
||||||
|
assert "finding" not in context
|
||||||
|
assert context["run"]["code_review_enabled"] in (True, False)
|
||||||
|
|
||||||
|
|
||||||
|
def test_unknown_item_falls_back_to_run_scope(tmp_path):
|
||||||
|
context = build_context(_job_dir(tmp_path), "finding:NOPE", _queue())
|
||||||
|
assert context["scope"] == "run"
|
||||||
|
|
||||||
|
|
||||||
|
def test_missing_artifacts_degrade_to_empty(tmp_path):
|
||||||
|
context = build_context(str(tmp_path), None, [])
|
||||||
|
assert context["scope"] == "run"
|
||||||
|
assert context["artifacts_available"] == {
|
||||||
|
"conflicts.json": False, "agent/memory.json": False, "job.log": False}
|
||||||
|
|
||||||
|
|
||||||
|
def test_log_excerpt_matches_question_terms(tmp_path):
|
||||||
|
out_dir = _job_dir(tmp_path)
|
||||||
|
with open(os.path.join(out_dir, "job.log"), "w", encoding="utf-8") as f:
|
||||||
|
f.write("[Extract] page 3 Civil sheet unreadable, skipped\n")
|
||||||
|
f.write("[Brain] merged 2 findings\n")
|
||||||
|
context = build_context(out_dir, None, [], question="why no Civil sheets?")
|
||||||
|
assert any("Civil" in line for line in context["log_excerpt"])
|
||||||
|
assert context["artifacts_available"]["job.log"] is True
|
||||||
|
|
||||||
|
|
||||||
|
def test_reviewer_decision_so_far_is_included(tmp_path):
|
||||||
|
decisions = {"finding:AGENT-0007": {"decision": "reject",
|
||||||
|
"reason_code": "extraction_misread",
|
||||||
|
"comment": "that is a power floor box"}}
|
||||||
|
context = build_context(_job_dir(tmp_path), "finding:AGENT-0007", _queue(), decisions)
|
||||||
|
assert context["reviewer_decision_so_far"]["reason_code"] == "extraction_misread"
|
||||||
|
|
||||||
|
|
||||||
|
def test_clean_cluster_item_uses_cluster_scope(tmp_path):
|
||||||
|
queue = [{"review_item_id": "clean_cluster:roof-ac1", "kind": "clean_cluster",
|
||||||
|
"blocking": False, "reasons": ["audit_sample"],
|
||||||
|
"payload": {"key": "roof-ac1", "location": "Roof / AC-1",
|
||||||
|
"assertions": [{"sheet_number": "M2.1", "value": "grade"}]}}]
|
||||||
|
context = build_context(_job_dir(tmp_path), "clean_cluster:roof-ac1", queue)
|
||||||
|
assert context["scope"] == "item"
|
||||||
|
assert context["cluster"]["key"] == "roof-ac1"
|
||||||
|
assert "finding" not in context
|
||||||
@@ -85,3 +85,34 @@ def test_write_label_appends_json_lines(tmp_path):
|
|||||||
with open(path, encoding="utf-8") as f:
|
with open(path, encoding="utf-8") as f:
|
||||||
lines = [json.loads(line) for line in f if line.strip()]
|
lines = [json.loads(line) for line in f if line.strip()]
|
||||||
assert lines == [label1, label2]
|
assert lines == [label1, label2]
|
||||||
|
|
||||||
|
|
||||||
|
def test_decision_label_carries_reviewer_corrections():
|
||||||
|
"""category/severity corrections reach the label instead of being dropped."""
|
||||||
|
decision = {"decision": "reject", "reason_code": "extraction_misread",
|
||||||
|
"category_correction": "power floor box",
|
||||||
|
"severity_correction": "low"}
|
||||||
|
label = decision_to_label(_queue_item(), decision, {"job_id": "abc123", "pipeline_mode": "agent",
|
||||||
|
"report": {"summary": {}}})
|
||||||
|
assert label["category_correction"] == "power floor box"
|
||||||
|
assert label["severity_correction"] == "low"
|
||||||
|
assert label["kind"] == "review_decision"
|
||||||
|
|
||||||
|
|
||||||
|
def test_write_label_also_lands_in_the_cross_job_store(tmp_path):
|
||||||
|
from backend import config
|
||||||
|
from backend.review.feedback import read_shared_feedback
|
||||||
|
label = decision_to_label(_queue_item(), {"decision": "reject",
|
||||||
|
"reason_code": "extraction_misread"},
|
||||||
|
{"job_id": "abc123", "pipeline_mode": "agent",
|
||||||
|
"report": {"summary": {}}})
|
||||||
|
write_label(str(tmp_path), label)
|
||||||
|
assert os.path.isfile(os.path.join(config.REVIEW_FEEDBACK_DIR, "decisions.jsonl"))
|
||||||
|
records = read_shared_feedback("review_decision")
|
||||||
|
assert len(records) == 1
|
||||||
|
assert records[0]["reason_code"] == "extraction_misread"
|
||||||
|
|
||||||
|
|
||||||
|
def test_shared_feedback_read_is_empty_when_nothing_written():
|
||||||
|
from backend.review.feedback import read_shared_feedback
|
||||||
|
assert read_shared_feedback("review_decision") == []
|
||||||
Reference in new issue
Block a user