feat: review chat — ask the run why it concluded a finding
Docker Release / build-and-push (push) Successful in 1m45s
Docker Release / release (push) Skipped

Read-only Q&A on the review screen, per finding and per run, answered from
the job's own artifacts (evidence, cluster, extraction, verification, Brain
merge, sheet index, cover reconciliation, job.log). It never mutates findings,
decisions, or the report.

Turns are logged job-locally (review/chat_log.jsonl, transcript at
/jobs/{id}/review-chat/log) and to a cross-job feedback store
(REVIEW_FEEDBACK_DIR), which now also receives review decisions with their
category/severity corrections.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0115gGtrSxXE9DKvS9XPFSoT
This commit is contained in:
woogiandClaude Opus 5 committed 2026-09-14 10:35:18 -05:00
1 parent 46db871152
commit 23e19f53b2
15 files changed
+1598 -9

No files matched your search

+20
View File
@@ -129,6 +129,26 @@ AGENT_REVIEW_AUDIT_SAMPLE = int(os.getenv("AGENT_REVIEW_AUDIT_SAMPLE", "5"))
# NOTE: currently unwired - reserved for future cross-job aggregation tooling.
REVIEW_AGGREGATE_INCLUDE_TEXT = os.getenv("REVIEW_AGGREGATE_INCLUDE_TEXT", "false").strip().lower() in ("1", "true", "yes")
# -- Review chat (ask-the-run Q&A on the review screen) --------------
# A read-only explainer: it answers "why did the run decide X?" from the job's
# own artifacts and never mutates findings, decisions, or code. Every turn is
# appended to <out_dir>/review/chat_log.jsonl AND to the cross-job feedback
# store (REVIEW_FEEDBACK_DIR) so answers are available to future prompt priors.
# HISTORY_TURNS caps how much of a thread is replayed into the prompt;
# LOG_LINES caps how many job.log lines are searched into the context bundle.
ENABLE_REVIEW_CHAT = _flag("ENABLE_REVIEW_CHAT", "true")
REVIEW_CHAT_MODEL = os.getenv("REVIEW_CHAT_MODEL", "") or TEXT_MODEL
REVIEW_CHAT_MAX_TOKENS = int(os.getenv("REVIEW_CHAT_MAX_TOKENS", "4096"))
REVIEW_CHAT_HISTORY_TURNS = int(os.getenv("REVIEW_CHAT_HISTORY_TURNS", "6"))
REVIEW_CHAT_LOG_LINES = int(os.getenv("REVIEW_CHAT_LOG_LINES", "40"))
REVIEW_CHAT_MAX_QUESTION_CHARS = int(os.getenv("REVIEW_CHAT_MAX_QUESTION_CHARS", "2000"))
# Cross-job feedback store: where review decisions and chat turns accumulate so
# a future run can be primed with "what humans corrected last time". Job-local
# artifacts stay the source of truth; this is the append-only roll-up.
# (OUTPUT_DIR is defined further down; keep this in sync with it.)
REVIEW_FEEDBACK_DIR = os.getenv("REVIEW_FEEDBACK_DIR", "") or os.path.join(
_BASE_DIR, "outputs", "_feedback")
# -- Hybrid (local text LLM) ----------------------------------------
# Optional OpenAI-compatible local endpoint (e.g. a vLLM box) for the text-only
# QAQC stages. Vision stages ALWAYS use OpenRouter. The user picks hybrid per