feat: review chat — ask the run why it concluded a finding
Docker Release / build-and-push (push) Successful in 1m45s
Docker Release / release (push) Skipped

Read-only Q&A on the review screen, per finding and per run, answered from
the job's own artifacts (evidence, cluster, extraction, verification, Brain
merge, sheet index, cover reconciliation, job.log). It never mutates findings,
decisions, or the report.

Turns are logged job-locally (review/chat_log.jsonl, transcript at
/jobs/{id}/review-chat/log) and to a cross-job feedback store
(REVIEW_FEEDBACK_DIR), which now also receives review decisions with their
category/severity corrections.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0115gGtrSxXE9DKvS9XPFSoT
This commit is contained in:
woogiandClaude Opus 5 committed 2026-09-14 10:35:18 -05:00
1 parent 46db871152
commit 23e19f53b2
15 files changed
+1598 -9

No files matched your search

+46 -2
View File
@@ -212,10 +212,54 @@ From the CLI, `--no-review` bypasses the gate for that run (it overrides
python cli/run_check.py samples/your_set.pdf --mode agent --no-review --out out/agent-run
```
### Asking the run why: review chat
Each item on the review screen has an **Ask about this finding** panel, and the
screen carries one **Ask about this run** panel for questions that are not about
a single finding. The chat answers from the job's own artifacts — the finding's
evidence, the cluster it came from, the raw per-sheet extraction, the
verification verdict, the Brain's merge decision, the sheet index, the cover-index
reconciliation, and matching `job.log` lines.
```
"why does it think the AC unit is mounted on the ground?" -> item scope
"why didn't it pick up on the Civil set?" -> run scope
```
The chat is **read-only**. It cannot change a finding, a severity, a decision,
or the report, and the prompt forbids it from proposing code or config changes —
the radio buttons remain the only thing that alters review state. When the
artifacts do not contain the answer, it says so and names what is missing rather
than guessing.
Every turn is logged twice:
- `outputs/<job_id>/review/chat_log.jsonl` — the auditable record: the issue as
it stood when asked about, the question, the answer, the determinations, and
the evidence quoted. Readable as a transcript at
`GET /jobs/{id}/review-chat/log`.
- `REVIEW_FEEDBACK_DIR/chat_turns.jsonl` — the cross-job roll-up, alongside
`decisions.jsonl`. When a reviewer corrects a misidentification in
conversation ("that is not a floor drain, it is a power floor box"), the
correction is captured as `suggested_category_correction` rather than dying in
free text. Nothing reads this store yet; writing it is what makes priming a
future run on past corrections possible.
| Key | Default | Effect |
|-----|---------|--------|
| `ENABLE_REVIEW_CHAT` | `true` | `false` = the chat endpoints refuse and the panels stay empty |
| `REVIEW_CHAT_MODEL` | `TEXT_MODEL` | Model for chat answers |
| `REVIEW_CHAT_MAX_TOKENS` | `4096` | Answer budget |
| `REVIEW_CHAT_HISTORY_TURNS` | `6` | Prior turns replayed into a thread's prompt |
| `REVIEW_CHAT_LOG_LINES` | `40` | Max `job.log` lines pulled into the context bundle |
| `REVIEW_FEEDBACK_DIR` | `backend/outputs/_feedback` | Cross-job decision + chat feedback store |
**Deployment note:** the review endpoints (`/jobs/{id}/review-decisions`,
`/jobs/{id}/finalize-review`) are **state-changing and sensitive** — they accept
human decisions that alter the final report. Do **not** expose the UI/API
publicly without reverse-proxy auth or a shared access token in front of it.
human decisions that alter the final report. `/jobs/{id}/review-chat` does not
change review state, but it does spend model budget and returns drawing
evidence. Do **not** expose the UI/API publicly without reverse-proxy auth or a
shared access token in front of it.
Web UI (upload + view):