Add required human review gate to the Agent pipeline #1

Open
woogi wants to merge 28 commits from agent-mode into main
Owner

Summary

Agent web jobs now require human review before final RFIs/reports are issued:

  • Pipeline stops after Brain consolidation, persists a review queue (outputs/<job>/review/), and parks the job in needs_review
  • Blocking items: high/critical severity, low-confidence, and sensitive-category findings; non-blocking audit sample of clean clusters
  • Review API (GET /jobs/{id}/review, POST .../review-decisions, POST .../finalize-review) + frontend review queue UI with reason codes for rejections
  • Finalizer applies decisions (rejections suppressed with reason codes, not deleted), performs bounded targeted reruns for clarifications (degrade to visible analysis_gap), drafts RFIs only for kept issues, rebuilds final conflicts/counts from kept findings
  • Two-phase email: review-required on needs_review, final report only after finalization
  • Per-decision feedback labels; aggregate metrics redact source_text/images/comments by default
  • Restart recovery from job artifacts (status derivation + job.json hydration)
  • CLI --no-review bypass; AGENT_REQUIRE_REVIEW / AGENT_REVIEW_AUDIT_SAMPLE / REVIEW_AGGREGATE_INCLUDE_TEXT env knobs
  • Classic pipeline unchanged

Plan + design spec included under docs/superpowers/.

Testing

  • 65 non-LLM tests, all passing (pytest); review logic fully covered without PDFs/network/OpenRouter
  • Full-flow coverage: gate policy, queue build, store round-trip, runner gate, API state transitions, finalization (confirm/reject/unsure/clarify), restart recovery, email flow, metrics redaction

Notes

  • Manual browser check of the review UI still pending (JS verified by review only — no JS engine in dev env)
  • Review endpoints are state-changing and unauthenticated — do not expose beyond trusted LAN without reverse-proxy auth (documented in README)
  • Spec discrepancy recorded: spec said audit sample = 5 items total; implementation queues all non-blocking findings as audit items plus 5 clean clusters (strictly more transparent; final review recommended amending the spec, not the code)
## Summary Agent web jobs now require human review before final RFIs/reports are issued: - Pipeline stops after Brain consolidation, persists a review queue (`outputs/<job>/review/`), and parks the job in `needs_review` - Blocking items: high/critical severity, low-confidence, and sensitive-category findings; non-blocking audit sample of clean clusters - Review API (`GET /jobs/{id}/review`, `POST .../review-decisions`, `POST .../finalize-review`) + frontend review queue UI with reason codes for rejections - Finalizer applies decisions (rejections suppressed with reason codes, not deleted), performs bounded targeted reruns for clarifications (degrade to visible `analysis_gap`), drafts RFIs only for kept issues, rebuilds final conflicts/counts from kept findings - Two-phase email: review-required on `needs_review`, final report only after finalization - Per-decision feedback labels; aggregate metrics redact source_text/images/comments by default - Restart recovery from job artifacts (status derivation + job.json hydration) - CLI `--no-review` bypass; `AGENT_REQUIRE_REVIEW` / `AGENT_REVIEW_AUDIT_SAMPLE` / `REVIEW_AGGREGATE_INCLUDE_TEXT` env knobs - Classic pipeline unchanged Plan + design spec included under `docs/superpowers/`. ## Testing - 65 non-LLM tests, all passing (`pytest`); review logic fully covered without PDFs/network/OpenRouter - Full-flow coverage: gate policy, queue build, store round-trip, runner gate, API state transitions, finalization (confirm/reject/unsure/clarify), restart recovery, email flow, metrics redaction ## Notes - Manual browser check of the review UI still pending (JS verified by review only — no JS engine in dev env) - Review endpoints are state-changing and unauthenticated — do not expose beyond trusted LAN without reverse-proxy auth (documented in README) - Spec discrepancy recorded: spec said audit sample = 5 items total; implementation queues all non-blocking findings as audit items plus 5 clean clusters (strictly more transparent; final review recommended amending the spec, not the code)
woogi added 3 commits 2026-07-28 12:33:11 -07:00
Wire specialist waves, Brain consolidation, and Classic-compatible reports so Agent mode can run end-to-end via OpenRouter without changing the default Classic path.

Co-authored-by: Cursor <cursoragent@cursor.com>
Build agent-mode branch image in CI under a branch tag.
Docker Release / build-and-push (push) Successful in 3m38s
Docker Release / release (push) Has been skipped
ac328d34fd
Publish agent-mode pushes as :agent-mode (and :sha-<commit>) so the experimental fork is available from the registry without overwriting the Classic :latest image.

Co-authored-by: Cursor <cursoragent@cursor.com>
Add required human review gate to the Agent pipeline.
Docker Release / build-and-push (push) Successful in 1m10s
Docker Release / release (push) Skipped
1c1d2ff21b
Agent web jobs now stop after Brain consolidation and enter needs_review
with a persisted review queue (blocking: high-severity, low-confidence,
sensitive-category findings; audit sample of clean clusters). Humans
decide confirm/reject/unsure/needs_clarification via new review API and
frontend queue; a finalizer applies decisions (rejections suppressed with
reason codes), performs bounded targeted reruns for clarifications,
drafts RFIs only for kept issues, and only then marks the job done and
sends the final email. Two-phase email (review-required, then final
report), per-decision feedback labels with redacted aggregate metrics,
restart recovery from job artifacts, and CLI --no-review bypass.
Classic pipeline unchanged. 65 non-LLM tests.
woogi added 1 commit 2026-07-28 13:34:35 -07:00
Point email links at conchecker.scoutitsystems.com and show build SHA in header.
Docker Release / build-and-push (push) Successful in 58s
Docker Release / release (push) Skipped
4ecc7c5cef
APP_BASE_URL default (config, .env.example, both compose files) is now
https://conchecker.scoutitsystems.com with no port, so review-required
and final-report email links use the public site. CI bakes the short
commit SHA into the image as APP_BUILD via a Docker build-arg; /health
returns version+build and the site header shows the build so it's easy
to confirm which image is deployed. Local runs default to 'dev'.
woogi added 1 commit 2026-07-28 14:23:45 -07:00
Add job run logs, OpenRouter model picker, and discipline grouping.
Docker Release / build-and-push (push) Successful in 55s
Docker Release / release (push) Skipped
afa1089311
- Job logs: each job's stdout/stderr is teed into outputs/<id>/job.log
  (survives restarts) and served at GET /jobs/{id}/log as text/plain, so
  full run logs can be shared for debugging and refinement.
- Model picker: GET /models proxies OpenRouter's public model list with
  per-1M-token pricing (1h cache, 502 on failure); the UI shows a model
  dropdown with costs when OpenRouter compute is selected, and the pick
  overrides vision+text models for that job (Classic and Agent modes).
- Conflicts in the report view are grouped by discipline pair
  (collapsible sections, severity-ordered within groups) instead of one
  flat severity-only list.
woogi added 1 commit 2026-08-02 07:57:46 -07:00
Merge main: dual model dropdowns + richer job logs, adapted for agent-mode.
Docker Release / build-and-push (push) Successful in 1m0s
Docker Release / release (push) Skipped
f7e1b6bb7c
- llm.py: set_model_overrides(vision, text) replaces the single job override;
  UI picks still beat per-call agent model args, but never name the hybrid
  local model (avoids main's hybrid footgun); local->cloud fallback uses the
  text pick.
- jobs.py: timestamped line-split tee (job_log.py), in-memory log + log_tail
  polls, full log on terminal states (done/error/needs_review/finalization_error),
  log-only disk recovery, error email links to the run log, and failed runs now
  append the full traceback to job.log. Keeps pipeline_mode, job.json, and the
  review gate.
- models.py: vision/text split via architecture modalities, pricing kept;
  /models returns {vision, text, defaults}; /check takes vision_model/text_model
  (replacing model); /health adds text_model. models_catalog.py dropped.
- UI: two priced dropdowns (OpenRouter compute only) + live run-log panel.
- Tests updated for dual overrides and the /models shape; new coverage for
  traceback capture and local-model immunity.
woogi added 1 commit 2026-08-05 13:17:20 -07:00
Verbose per-call LLM logging, raw request/response dumps, and end-of-log cost summary.
Docker Release / build-and-push (push) Successful in 1m13s
Docker Release / release (push) Skipped
7488cf68c5
- [LLM] line per call: stage, model, prompt size, output size, cost, parsed item counts
- LLM_RAW_DUMP: full prompt/response JSON per call under outputs/<job>/llm_raw/
- Cost block at tail of job.log (per-stage, per-model, cached vs live)
- Agent mode: reset llm cost counters per job; review finalization now teed into job.log + dumps
woogi added 1 commit 2026-08-05 13:28:12 -07:00
Agent-mode UI: remove hybrid compute option, add clickable sheet links in review queue.
Docker Release / build-and-push (push) Successful in 1m1s
Docker Release / release (push) Skipped
6f10062b93
woogi added 1 commit 2026-08-05 13:42:23 -07:00
Fix model dropdown loading, cache models for 24h, and style build tag.
Docker Release / release (push) Skipped
Docker Release / build-and-push (push) Successful in 1m1s
5305325d81
woogi added 1 commit 2026-08-05 13:50:37 -07:00
Add cache-busting meta tags and default build tag text.
Docker Release / build-and-push (push) Successful in 56s
Docker Release / release (push) Skipped
4b3b62b3fa
woogi added 1 commit 2026-08-05 13:59:52 -07:00
Add /models fetch timeout and health-derived default fallback for model dropdowns.
Docker Release / build-and-push (push) Successful in 57s
Docker Release / release (push) Skipped
32544bc2af
woogi added 1 commit 2026-08-05 14:07:35 -07:00
Guard syncPipelineOptions when hybrid radio is removed.
Docker Release / build-and-push (push) Successful in 55s
Docker Release / release (push) Skipped
5c1fccfb35
woogi added 1 commit 2026-08-07 05:35:56 -07:00
Fix sheet-extraction page loss: bare-list wrap, compact retry, reasoning cap
Docker Release / build-and-push (push) Successful in 1m1s
Docker Release / release (push) Skipped
76e0a52658
- SheetExtractorAgent accepts top-level array responses as the objects
  array instead of discarding them (recovered the failure mode behind
  16/38 failed sheets on job 475a6f184dd1)
- Second-chance compact retry per page before declaring extraction failed
- call_json: reasoning_effort param (cloud-only extra_body), finish_reason
  capture + explicit max_tokens log line, finish_reason in raw dumps,
  cache key covers reasoning_effort
- EXTRACT_MAX_TOKENS default 16384 -> 32768 (Gemini thinking tokens count
  against the cap); new EXTRACT_REASONING_EFFORT=low default for extract
- tests: 5 new fallback-ladder tests
woogi added 1 commit 2026-08-09 06:17:43 -07:00
Kill extract-wave truncation: 65k ceiling, hard thinking budget, reasoning-token telemetry
Docker Release / build-and-push (push) Successful in 57s
Docker Release / release (push) Skipped
3d7fce7bf9
Job 98194fa8d215 showed every extract call hitting the 32k cap with only
~20k chars visible despite reasoning effort=low - Gemini 2.5 Pro still
burned ~25k thinking tokens per sheet.

- EXTRACT_MAX_TOKENS default 32768 -> 65536 (model output ceiling)
- new EXTRACT_REASONING_MAX_TOKENS (default 2048): OpenRouter reasoning
  max_tokens / Gemini thinking_budget; takes precedence over effort
- log per-call reasoning token counts (usage.completion_tokens_details)
  and include thinking count in the finish_reason=length marker
woogi added 1 commit 2026-08-09 06:21:33 -07:00
Sync .env.example with new extract token defaults
Docker Release / build-and-push (push) Successful in 57s
Docker Release / release (push) Skipped
228e8bd031
woogi added 1 commit 2026-08-10 06:22:00 -07:00
Add non-technical pipeline overview: infographic + flow explainer doc
Docker Release / build-and-push (push) Successful in 59s
Docker Release / release (push) Skipped
0ea0b0e897
woogi added 10 commits 2026-08-10 10:27:30 -07:00
woogi added 2 commits 2026-08-12 12:27:22 -07:00
feat: text-layer grounding (extractor authority, guard rescue tier, verifier oracle + hi-DPI crops)
Docker Release / build-and-push (push) Successful in 1m25s
Docker Release / release (push) Skipped
570300324f
- backend/text_layer.py: PyMuPDF text-layer extraction, fuzzy evidence
  bbox matching, 300-DPI crop rendering, coverage-gap signal
- extractor (classic + agent): TEXT LAYER block appended at call sites;
  grounding guard gains text-layer rescue tier (grounding=text_layer stamp)
- verifier: {text_layer} oracle excerpt + evidence-located hi-DPI crops
  replacing full-page images (fallback preserved, I2 guard intact)
- coverage gaps: text-bearing pages with zero extraction -> failed-scope
  gap findings (agent) / log-only (classic)
- config knobs: TEXT_LAYER_ENABLED/MIN_CHARS/MAX_CHARS, VERIFY_TEXT_MAX_CHARS,
  VERIFY_HI_DPI_CROPS, VERIFY_CROP_DPI, VERIFY_CROP_MARGIN_PTS
- tests: 22 new (text_layer unit, grounding/render, runner-level flow)
Spec: docs/superpowers/specs/2026-08-12-text-layer-grounding-design.md
All checks were successful
Docker Release / build-and-push (push) Successful in 1m25s
Docker Release / release (push) Skipped
You are not authorized to merge this pull request.
This pull request can be merged automatically.
View command line instructions

Checkout

From your project repository, check out a new branch and test the changes.
git fetch -u origin agent-mode:agent-mode
git checkout agent-mode
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: woogi/Conflict_Checker#1