feat: review chat — ask the run why it concluded a finding
Docker Release / build-and-push (push) Successful in 1m45s
Docker Release / release (push) Skipped

Read-only Q&A on the review screen, per finding and per run, answered from
the job's own artifacts (evidence, cluster, extraction, verification, Brain
merge, sheet index, cover reconciliation, job.log). It never mutates findings,
decisions, or the report.

Turns are logged job-locally (review/chat_log.jsonl, transcript at
/jobs/{id}/review-chat/log) and to a cross-job feedback store
(REVIEW_FEEDBACK_DIR), which now also receives review decisions with their
category/severity corrections.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0115gGtrSxXE9DKvS9XPFSoT
This commit is contained in:
woogiandClaude Opus 5 committed 2026-09-14 10:35:18 -05:00
1 parent 46db871152
commit 23e19f53b2
15 files changed
+1598 -9

No files matched your search

+20 -3
View File
@@ -94,7 +94,9 @@ table, and a senior reviewer at the end sorts it all into one clean report.
+----------------------------------------------------------+
| 7. HUMAN REVIEW GATE |
| The important/uncertain findings are queued for a |
| real person to Confirm / Reject / mark Unsure |
| real person to Confirm / Reject / mark Unsure. |
| You can also ASK the run why it concluded any of |
| them, or why it never looked at something |
+----------------------------------------------------------+
|
v
@@ -128,6 +130,7 @@ table, and a senior reviewer at the end sorts it all into one clean report.
| 6 | Brain | Chief estimator | Deduplicates, judges, and prioritizes all findings |
| 6.5 | Brain (clarification) | Second opinion | For findings it distrusts, sends them back to the fact-checker to re-read the sheet; debunked findings are dropped |
| 7 | Review Gate | Your desk | Presents the findings a human should approve before anything goes out |
| 7 | Review Chat | The analyst you can question | Answers "why did it decide that?" and "why didn't it check that?" from the run's own records — it explains, it never changes anything |
| 8 | RFI Writer | Secretary | Writes the formal clarification letters for confirmed issues |
Everything the assistants learn is kept in a shared notebook (the "project
@@ -182,6 +185,19 @@ the whole review.
queue better — the system already records your feedback, so it can get
smarter over time about what actually needs your eyes.
### 5b. Explaining itself — "why did you think that?" ✅ *shipped*
- **Done:** every finding on the review screen has a chat panel, plus one for
the run as a whole. Ask why a unit was read as ground-mounted, or why a whole
discipline never got looked at, and it traces the answer back through what it
actually recorded — quoting the note it read off the sheet, or naming the
stage that skipped the pages. It cannot change a finding; that stays yours.
- **Done:** when you correct it in conversation ("that's not a floor drain,
it's a power floor box"), the correction is filed as structured data rather
than a free-text comment.
- **Still open:** nothing reads those filed corrections back yet. The next step
is priming a new run with what reviewers corrected on previous sets, so the
same misread does not come back on the next job.
### 6. Trust — "show the receipts" ✅ *partially shipped*
- **Done:** the fact-checker already pulls a zoomed-in crop of the exact spot on
the sheet when it re-reads a finding.
@@ -192,5 +208,6 @@ the whole review.
*Technical reference for the curious: the pipeline lives in
`backend/agents/runner.py` (the waves above are the "Agent wave N" stages), the
team's shared notebook is `backend/agents/memory.py`, and the review queue is
`backend/review/gate.py` + `backend/review/finalizer.py`.*
team's shared notebook is `backend/agents/memory.py`, the review queue is
`backend/review/gate.py` + `backend/review/finalizer.py`, and the review chat is
`backend/review/chat.py` + `backend/review/chat_context.py`.*