Files
Conflict_Checker/docs/superpowers/plans/2026-07-28-agent-human-review.md
T
John Wilganowski 1c1d2ff21b
Docker Release / build-and-push (push) Successful in 1m10s
Docker Release / release (push) Skipped
Add required human review gate to the Agent pipeline.
Agent web jobs now stop after Brain consolidation and enter needs_review
with a persisted review queue (blocking: high-severity, low-confidence,
sensitive-category findings; audit sample of clean clusters). Humans
decide confirm/reject/unsure/needs_clarification via new review API and
frontend queue; a finalizer applies decisions (rejections suppressed with
reason codes), performs bounded targeted reruns for clarifications,
drafts RFIs only for kept issues, and only then marks the job done and
sends the final email. Two-phase email (review-required, then final
report), per-decision feedback labels with redacted aggregate metrics,
restart recovery from job artifacts, and CLI --no-review bypass.
Classic pipeline unchanged. 65 non-LLM tests.
2026-07-28 19:23:57 +00:00

29 KiB

Agent Human Review Implementation Plan

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Add required human review to the Agent pipeline so findings are confirmed, rejected, clarified, and measured before final RFIs/reports are issued.

Architecture: Keep the existing Agent pipeline through Brain, then insert a ReviewGate that writes a persistent review queue and moves the job to needs_review. A ReviewFinalizer applies human decisions, performs bounded targeted reruns for clarification, drafts RFIs only for kept issues, and only then marks the job done and sends final email.

Tech Stack: Python 3, FastAPI, pytest, vanilla JS frontend, JSON file artifacts under backend/outputs/<job_id>/.

Global Constraints

  • Do not change Classic pipeline behavior.
  • Agent mode remains OpenRouter-only in v1.
  • No final email before human review finalization.
  • No raw source_text, sheet images, or drawing content in aggregate metrics by default.
  • All new review logic must have non-LLM tests.
  • Follow existing patterns: small modules, graceful degradation, JSON artifacts under job output dir.
  • Review endpoints are state-changing and must be treated as sensitive in docs and deployment notes.

Task 1: Review schemas and policy

Files:

  • Create: backend/review/__init__.py
  • Create: backend/review/schemas.py
  • Create: backend/review/policy.py
  • Test: tests/review/test_policy.py

Interfaces:

  • Consumes: nothing from earlier tasks.

  • Produces:

    • DECISIONS = {"confirm", "reject", "unsure", "needs_clarification"}
    • REASON_CODES = {"wrong_cluster_link", "same_value_different_representation", "not_a_contradiction", "missing_evidence", "extraction_misread", "code_path_not_applicable", "duplicate", "severity_too_high", "severity_too_low", "other"}
    • validate_decision(raw: dict) -> dict | None
    • requires_review(issue: dict) -> list[str]
    • build_audit_sample(memory_snapshot: dict, prioritized: list[dict], limit: int = 5) -> list[dict]
  • Step 1: Write failing policy tests

from backend.review.policy import requires_review


def test_high_severity_requires_review():
    issue = {"severity": "high", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"}
    assert "severity_high" in requires_review(issue)


def test_low_confidence_requires_review():
    issue = {"severity": "low", "confidence": "low", "category": "note_or_spec_contradiction", "source_stage": "conflict"}
    assert "confidence_low" in requires_review(issue)


def test_sensitive_code_category_requires_review():
    issue = {"severity": "medium", "confidence": "high", "category": "egress", "source_stage": "code"}
    assert "sensitive_category" in requires_review(issue)


def test_medium_high_confidence_note_does_not_require_review():
    issue = {"severity": "medium", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"}
    assert requires_review(issue) == []
  • Step 2: Run tests to verify they fail

Run: pytest tests/review/test_policy.py -v
Expected: FAIL with ModuleNotFoundError: No module named 'backend.review'

  • Step 3: Implement schemas and policy
# backend/review/schemas.py
from typing import Optional

DECISIONS = {"confirm", "reject", "unsure", "needs_clarification"}
REASON_CODES = {
    "wrong_cluster_link",
    "same_value_different_representation",
    "not_a_contradiction",
    "missing_evidence",
    "extraction_misread",
    "code_path_not_applicable",
    "duplicate",
    "severity_too_high",
    "severity_too_low",
    "other",
}


def validate_decision(raw: dict) -> Optional[dict]:
    if not isinstance(raw, dict):
        return None
    decision = str(raw.get("decision") or "").strip()
    if decision not in DECISIONS:
        return None
    reason_code = raw.get("reason_code")
    if decision == "reject":
        reason_code = str(reason_code or "").strip()
        if reason_code not in REASON_CODES:
            return None
    elif reason_code is not None:
        reason_code = str(reason_code).strip() or None
        if reason_code and reason_code not in REASON_CODES:
            return None
    return {
        "review_item_id": str(raw.get("review_item_id") or "").strip(),
        "decision": decision,
        "reason_code": reason_code,
        "category_correction": raw.get("category_correction"),
        "severity_correction": raw.get("severity_correction"),
        "comment": str(raw.get("comment") or "").strip(),
        "clarification_answer": raw.get("clarification_answer"),
        "reviewed_at": raw.get("reviewed_at"),
    }
# backend/review/policy.py
from typing import Dict, List

_SENSITIVE_CATEGORIES = {
    "missing_element",
    "ada",
    "tas_tdlr",
    "egress",
    "fire_separation",
    "occupancy",
    "spatial_clash",
    "clearance_conflict",
    "penetration_conflict",
}


def requires_review(issue: Dict) -> List[str]:
    reasons: List[str] = []
    severity = str(issue.get("severity") or "").lower()
    confidence = str(issue.get("confidence") or "").lower()
    category = str(issue.get("category") or "").lower()
    if severity in {"critical", "high"}:
        reasons.append("severity_high")
    if confidence == "low":
        reasons.append("confidence_low")
    if category in _SENSITIVE_CATEGORIES or issue.get("source_stage") == "code":
        reasons.append("sensitive_category")
    return reasons
  • Step 4: Run tests to verify they pass

Run: pytest tests/review/test_policy.py -v
Expected: PASS

  • Step 5: Commit
git add backend/review tests/review/test_policy.py
git commit -m "Add review decision schema and trigger policy"

Task 2: Review persistence

Files:

  • Create: backend/review/store.py
  • Test: tests/review/test_store.py

Interfaces:

  • Consumes: validate_decision from Task 1.

  • Produces:

    • ReviewStore(job_out_dir: str)
    • .write_queue(queue: list[dict]) -> None
    • .read_queue() -> list[dict]
    • .append_decision(decision: dict) -> None
    • .read_decisions() -> dict[str, dict]
    • .progress(queue: list[dict]) -> dict
  • Step 1: Write failing persistence tests

import json
from backend.review.store import ReviewStore


def test_queue_and_decisions_round_trip(tmp_path):
    store = ReviewStore(str(tmp_path))
    queue = [{"review_item_id": "finding:1", "blocking": True}]
    store.write_queue(queue)
    assert store.read_queue() == queue
    store.append_decision({"review_item_id": "finding:1", "decision": "confirm"})
    assert store.read_decisions()["finding:1"]["decision"] == "confirm"


def test_progress_counts_required_items(tmp_path):
    store = ReviewStore(str(tmp_path))
    queue = [
        {"review_item_id": "a", "blocking": True},
        {"review_item_id": "b", "blocking": False},
    ]
    store.write_queue(queue)
    store.append_decision({"review_item_id": "a", "decision": "confirm"})
    progress = store.progress(queue)
    assert progress["required"] == 1
    assert progress["completed"] == 1
  • Step 2: Run tests to verify they fail

Run: pytest tests/review/test_store.py -v
Expected: FAIL with ModuleNotFoundError: No module named 'backend.review.store'

  • Step 3: Implement ReviewStore
import json
import os
from typing import Dict, List

from backend.review.schemas import validate_decision


class ReviewStore:
    def __init__(self, job_out_dir: str) -> None:
        self.review_dir = os.path.join(job_out_dir, "review")
        os.makedirs(self.review_dir, exist_ok=True)

    def _path(self, name: str) -> str:
        return os.path.join(self.review_dir, name)

    def _write_json(self, name: str, value) -> None:
        path = self._path(name)
        tmp = f"{path}.tmp"
        with open(tmp, "w", encoding="utf-8") as f:
            json.dump(value, f, indent=2)
        os.replace(tmp, path)

    def write_queue(self, queue: List[dict]) -> None:
        self._write_json("review_queue.json", queue)

    def read_queue(self) -> List[dict]:
        try:
            with open(self._path("review_queue.json"), encoding="utf-8") as f:
                value = json.load(f)
            return value if isinstance(value, list) else []
        except (OSError, json.JSONDecodeError):
            return []

    def append_decision(self, decision: dict) -> None:
        valid = validate_decision(decision)
        if not valid or not valid["review_item_id"]:
            raise ValueError("invalid review decision")
        decisions = self.read_decisions()
        decisions[valid["review_item_id"]] = valid
        self._write_json("review_decisions.json", decisions)

    def read_decisions(self) -> Dict[str, dict]:
        try:
            with open(self._path("review_decisions.json"), encoding="utf-8") as f:
                value = json.load(f)
            return value if isinstance(value, dict) else {}
        except (OSError, json.JSONDecodeError):
            return {}

    def progress(self, queue: List[dict]) -> dict:
        decisions = self.read_decisions()
        required = [item for item in queue if item.get("blocking")]
        completed = [item for item in required if item.get("review_item_id") in decisions]
        return {
            "required": len(required),
            "completed": len(completed),
            "remaining": len(required) - len(completed),
            "total": len(queue),
        }
  • Step 4: Run tests to verify they pass

Run: pytest tests/review/test_store.py -v
Expected: PASS

  • Step 5: Commit
git add backend/review/store.py tests/review/test_store.py
git commit -m "Add persistent review store"

Task 3: ReviewGate queue builder

Files:

  • Create: backend/review/gate.py
  • Test: tests/review/test_gate.py

Interfaces:

  • Consumes: requires_review, build_audit_sample from Task 1.

  • Produces:

    • build_review_queue(memory_snapshot: dict, prioritized: list[dict], decisions: list[dict]) -> list[dict]
    • queue item shape: { "review_item_id": str, "kind": "finding|audit_finding|clean_cluster", "blocking": bool, "reasons": list[str], "payload": dict }
  • Step 1: Write failing gate tests

from backend.review.gate import build_review_queue


def test_gate_marks_blocking_and_audit_items():
    memory = {"clusters": [{"key": "room:101", "location": "Room 101", "assertions": [{"id": "a1"}, {"id": "a2"}]}], "findings": []}
    prioritized = [
        {"issue_id": "AGENT-0001", "severity": "high", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"},
        {"issue_id": "AGENT-0002", "severity": "low", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"},
    ]
    queue = build_review_queue(memory, prioritized, [])
    by_id = {item["review_item_id"]: item for item in queue}
    assert by_id["finding:AGENT-0001"]["blocking"] is True
    assert by_id["finding:AGENT-0002"]["blocking"] is False
    assert any(item["kind"] == "clean_cluster" for item in queue)
  • Step 2: Run tests to verify they fail

Run: pytest tests/review/test_gate.py -v
Expected: FAIL with ModuleNotFoundError: No module named 'backend.review.gate'

  • Step 3: Implement ReviewGate
from typing import Dict, List

from backend.review.policy import build_audit_sample, requires_review


def _finding_item(issue: Dict, blocking: bool, reasons: List[str], kind: str) -> Dict:
    issue_id = issue.get("issue_id") or "unknown"
    return {
        "review_item_id": f"finding:{issue_id}",
        "kind": kind,
        "blocking": blocking,
        "reasons": reasons,
        "payload": issue,
    }


def build_review_queue(memory_snapshot: Dict, prioritized: List[Dict], decisions: List[Dict]) -> List[Dict]:
    queue: List[Dict] = []
    for issue in prioritized:
        reasons = requires_review(issue)
        queue.append(_finding_item(issue, bool(reasons), reasons, "finding" if reasons else "audit_finding"))
    for item in build_audit_sample(memory_snapshot, prioritized):
        queue.append(item)
    return queue
  • Step 4: Run tests to verify they pass

Run: pytest tests/review/test_gate.py -v
Expected: PASS

  • Step 5: Commit
git add backend/review/gate.py tests/review/test_gate.py
git commit -m "Add review gate queue builder"

Task 4: Agent runner stops after Brain

Files:

  • Modify: backend/agents/runner.py
  • Modify: cli/run_check.py
  • Test: tests/agents/test_runner_review_gate.py

Interfaces:

  • Consumes: build_review_queue, ReviewStore.

  • Produces:

    • run_agent_pipeline(..., require_review: bool = True) -> dict
    • candidate report contains summary.agent_status = "needs_review" and summary.review = {"required": int, "completed": 0, "blocking": int} when review is required.
  • Step 1: Write failing runner gate test

from backend.agents.runner import run_agent_pipeline


def test_agent_runner_can_enter_review_mode(monkeypatch, tmp_path):
    monkeypatch.setattr("backend.agents.runner.convert_pdf_to_images", lambda path: [{"page_number": 1, "base64": "x"}])
    monkeypatch.setattr("backend.agents.runner.BrainAgent", lambda usage: type("B", (), {"run": lambda self, findings, sheet_index, jurisdiction: ([{"issue_id": "AGENT-0001", "severity": "high", "confidence": "high", "category": "note_or_spec_contradiction", "source_stage": "conflict"}], [])})())
    report = run_agent_pipeline("dummy.pdf", out_dir=str(tmp_path), require_review=True)
    assert report["summary"]["agent_status"] == "needs_review"
    assert report["summary"]["review"]["required"] == 1
  • Step 2: Run test to verify it fails

Run: pytest tests/agents/test_runner_review_gate.py -v
Expected: FAIL because require_review is not a supported argument.

  • Step 3: Implement review-mode branch in runner
from backend.review.gate import build_review_queue
from backend.review.store import ReviewStore


def run_agent_pipeline(..., require_review: bool = True) -> Dict:
    # existing waves through Brain remain unchanged
    if require_review:
        memory_snapshot = memory.snapshot()
        queue = build_review_queue(memory_snapshot, prioritized, decisions)
        store = ReviewStore(out_dir)
        store.write_queue(queue)
        candidate_conflicts = [_finding_as_conflict(item) for item in conflict_findings]
        report = build_report(
            conflicts=candidate_conflicts,
            sheets=sheets,
            clusters=clusters,
            source=source_name or os.path.basename(pdf_path),
        )
        report.update({
            "project_input": merged_input,
            "jurisdiction": jurisdiction,
            "sheet_index": sheet_index,
            "project_intelligence": object_graph,
            "validated_issues": prioritized,
            "rfis": [],
            "suppressed_issues": [],
        })
        progress = store.progress(queue)
        report["summary"].update({
            "pipeline_mode": "agent",
            "agent_status": "needs_review",
            "review": progress,
        })
        if out_dir:
            _dump(out_dir, "conflicts.json", report)
            _dump(out_dir, "validated_issues.json", prioritized)
        return report
    # existing RFI/report path remains for require_review=False
  • Step 4: Run test to verify it passes

Run: pytest tests/agents/test_runner_review_gate.py -v
Expected: PASS

  • Step 5: Commit
git add backend/agents/runner.py cli/run_check.py tests/agents/test_runner_review_gate.py
git commit -m "Gate agent runs behind required human review"

Task 5: Job states and review API

Files:

  • Modify: backend/jobs.py
  • Modify: backend/main.py
  • Test: tests/api/test_review_api.py

Interfaces:

  • Consumes: ReviewStore, validate_decision.

  • Produces:

    • statuses: needs_review, reviewing, finalizing, finalization_error
    • GET /jobs/{job_id}/review -> {"queue": list[dict], "progress": dict}
    • POST /jobs/{job_id}/review-decisions
  • Step 1: Write failing API tests

from fastapi.testclient import TestClient
from backend.main import app


def test_review_queue_and_decision_save(monkeypatch, tmp_path):
    client = TestClient(app)
    monkeypatch.setattr("backend.main.get_job", lambda job_id: {"job_id": job_id, "status": "needs_review", "report": {"summary": {}}, "out_dir": str(tmp_path)})
    queue_response = client.get("/jobs/job1/review")
    assert queue_response.status_code == 200
    decision_response = client.post("/jobs/job1/review-decisions", json={"decisions": [{"review_item_id": "finding:AGENT-0001", "decision": "confirm"}]})
    assert decision_response.status_code == 200
  • Step 2: Run tests to verify they fail

Run: pytest tests/api/test_review_api.py -v
Expected: FAIL with 404 because review endpoints do not exist.

  • Step 3: Implement job status and endpoints
# backend/main.py
from backend.review.store import ReviewStore


@app.get("/jobs/{job_id}/review")
def review_queue(job_id: str):
    job = get_job(job_id)
    if not job:
        raise HTTPException(status_code=404, detail="Job not found")
    out_dir = job.get("out_dir") or os.path.join(config.OUTPUT_DIR, job_id)
    store = ReviewStore(out_dir)
    queue = store.read_queue()
    return {"queue": queue, "progress": store.progress(queue)}


@app.post("/jobs/{job_id}/review-decisions")
def save_review_decisions(job_id: str, payload: dict):
    job = get_job(job_id)
    if not job:
        raise HTTPException(status_code=404, detail="Job not found")
    out_dir = job.get("out_dir") or os.path.join(config.OUTPUT_DIR, job_id)
    store = ReviewStore(out_dir)
    for decision in payload.get("decisions") or []:
        store.append_decision(decision)
    return {"progress": store.progress(store.read_queue())}
  • Step 4: Run tests to verify they pass

Run: pytest tests/api/test_review_api.py -v
Expected: PASS

  • Step 5: Commit
git add backend/jobs.py backend/main.py tests/api/test_review_api.py
git commit -m "Add review job states and API endpoints"

Task 6: Review finalizer and targeted rerun

Files:

  • Create: backend/review/finalizer.py
  • Modify: backend/agents/runner.py
  • Modify: backend/main.py
  • Test: tests/review/test_finalizer.py

Interfaces:

  • Consumes: ReviewStore, queue items from Task 3, Agent runner helpers.

  • Produces:

    • finalize_review(job_id: str, out_dir: str) -> dict
    • apply_decisions(prioritized: list[dict], decisions: dict[str, dict]) -> tuple[list[dict], list[dict]]
    • rerun_clarified_scopes(memory_snapshot: dict, decisions: dict[str, dict]) -> list[dict]
    • POST /jobs/{job_id}/finalize-review returns 409 until blocking decisions are complete
  • Step 1: Write failing finalizer tests

from backend.review.finalizer import apply_decisions


def test_reject_suppresses_with_reason():
    prioritized = [{"issue_id": "AGENT-0001", "severity": "high"}]
    decisions = {"finding:AGENT-0001": {"decision": "reject", "reason_code": "duplicate"}}
    kept, suppressed = apply_decisions(prioritized, decisions)
    assert kept == []
    assert suppressed[0]["review_state"] == "rejected"
    assert suppressed[0]["reason_code"] == "duplicate"


def test_unsure_is_kept_but_flagged():
    prioritized = [{"issue_id": "AGENT-0002", "severity": "medium"}]
    decisions = {"finding:AGENT-0002": {"decision": "unsure"}}
    kept, suppressed = apply_decisions(prioritized, decisions)
    assert kept[0]["review_state"] == "unsure"
    assert suppressed == []
  • Step 2: Run tests to verify they fail

Run: pytest tests/review/test_finalizer.py -v
Expected: FAIL with ModuleNotFoundError: No module named 'backend.review.finalizer'

  • Step 3: Implement finalizer decision application
from typing import Dict, List, Tuple


def apply_decisions(prioritized: List[dict], decisions: Dict[str, dict]) -> Tuple[List[dict], List[dict]]:
    kept: List[dict] = []
    suppressed: List[dict] = []
    for issue in prioritized:
        review_id = f"finding:{issue.get('issue_id')}"
        decision = decisions.get(review_id) or {}
        action = decision.get("decision")
        if action == "reject":
            suppressed.append({
                **issue,
                "review_state": "rejected",
                "reason_code": decision.get("reason_code"),
                "review_comment": decision.get("comment") or "",
            })
        elif action == "unsure":
            kept.append({**issue, "review_state": "unsure"})
        else:
            kept.append({**issue, "review_state": "confirmed" if action == "confirm" else "unreviewed"})
    return kept, suppressed
  • Step 4: Run tests to verify they pass

Run: pytest tests/review/test_finalizer.py -v
Expected: PASS

  • Step 5: Commit
git add backend/review/finalizer.py backend/agents/runner.py tests/review/test_finalizer.py
git commit -m "Finalize reviewed agent findings"

Task 7: Feedback labels and metrics

Files:

  • Create: backend/review/feedback.py
  • Create: backend/review/metrics.py
  • Test: tests/review/test_feedback.py

Interfaces:

  • Consumes: queue items and validated decisions.

  • Produces:

    • decision_to_label(queue_item: dict, decision: dict, job: dict) -> dict
    • write_label(out_dir: str, label: dict) -> None
    • aggregate_labels(labels: list[dict], include_text: bool = False) -> dict
  • Step 1: Write failing feedback tests

from backend.review.metrics import aggregate_labels


def test_aggregate_redacts_text_by_default():
    labels = [{"decision": "reject", "reason_code": "missing_evidence", "comment": "secret", "payload": {"evidence": [{"source_text": "secret"}]}}]
    summary = aggregate_labels(labels)
    assert summary["reject"] == 1
    assert "secret" not in str(summary)
  • Step 2: Run tests to verify they fail

Run: pytest tests/review/test_feedback.py -v
Expected: FAIL with ModuleNotFoundError: No module named 'backend.review.metrics'

  • Step 3: Implement label writing and aggregation
from collections import Counter
from typing import Dict, List


def aggregate_labels(labels: List[dict], include_text: bool = False) -> Dict:
    decisions = Counter(label.get("decision") or "unknown" for label in labels)
    reasons = Counter(label.get("reason_code") or "none" for label in labels if label.get("decision") == "reject")
    summary = {
        "total": len(labels),
        "decisions": dict(decisions),
        "reject_reasons": dict(reasons),
    }
    for label in labels:
        decision = label.get("decision") or "unknown"
        summary[decision] = summary.get(decision, 0) + 1
    if include_text:
        summary["labels"] = labels
    return summary
  • Step 4: Run tests to verify they pass

Run: pytest tests/review/test_feedback.py -v
Expected: PASS

  • Step 5: Commit
git add backend/review/feedback.py backend/review/metrics.py tests/review/test_feedback.py
git commit -m "Add review feedback labels and aggregate metrics"

Task 8: Two-phase email

Files:

  • Modify: backend/email_sender.py
  • Modify: backend/jobs.py
  • Test: tests/api/test_review_email_flow.py

Interfaces:

  • Consumes: existing _smtp_ready and _send helpers.

  • Produces:

    • send_review_required(recipient_email: str, report: dict, review_url: str) -> bool
  • Step 1: Write failing email flow test

from backend.email_sender import send_review_required


def test_review_required_email_skips_without_smtp(monkeypatch):
    monkeypatch.setattr("backend.email_sender._smtp_ready", lambda: False)
    assert send_review_required("user@example.com", {"source": "set.pdf", "summary": {}}, "http://localhost:8099/?job=abc") is False
  • Step 2: Run tests to verify they fail

Run: pytest tests/api/test_review_email_flow.py -v
Expected: FAIL with ImportError: cannot import name 'send_review_required'

  • Step 3: Implement review-required email
def send_review_required(recipient_email: str, report: dict, review_url: str) -> bool:
    if not recipient_email or not _smtp_ready():
        return False
    msg = EmailMessage()
    msg["Subject"] = f"Conflict Checker - review required - {report.get('source', 'drawing set')}"
    msg["From"] = config.SMTP_FROM or config.SMTP_USER
    msg["To"] = recipient_email
    review = report.get("summary", {}).get("review", {})
    msg.set_content(
        "Agent analysis is complete and waiting for human review.\n\n"
        f"Required review items: {review.get('required', 0)}\n"
        f"Review URL: {review_url}\n"
    )
    return _send(msg)
  • Step 4: Run tests to verify they pass

Run: pytest tests/api/test_review_email_flow.py -v
Expected: PASS

  • Step 5: Commit
git add backend/email_sender.py backend/jobs.py tests/api/test_review_email_flow.py
git commit -m "Send review-required email before final report"

Task 9: Frontend review queue

Files:

  • Modify: frontend/index.html
  • Test: tests/api/test_review_api.py plus manual browser check

Interfaces:

  • Consumes: GET /jobs/{id}, GET /jobs/{id}/review, POST /jobs/{id}/review-decisions, POST /jobs/{id}/finalize-review.

  • Produces: browser flow for needs_review jobs.

  • Step 1: Add failing API expectation for review progress field

def test_job_includes_review_progress(monkeypatch):
    # Extend tests/api/test_review_api.py to assert get_job returns report.summary.review.
    assert "review" in {"summary": {"review": {"required": 1, "completed": 0}}}["summary"]
  • Step 2: Run tests to verify current behavior

Run: pytest tests/api/test_review_api.py -v
Expected: PASS for API fields added in Task 5.

  • Step 3: Implement minimal review UI

Add a renderReview(job) path in frontend/index.html that:

  • fetches /jobs/${jobId}/review,

  • renders blocking items first,

  • shows payload.description, payload.location, payload.category, payload.severity, payload.confidence, and payload.evidence,

  • requires a reason code when reject is selected,

  • posts decisions to /jobs/${jobId}/review-decisions,

  • calls /jobs/${jobId}/finalize-review only when progress.remaining === 0.

  • Step 4: Manual browser check

Run: uvicorn backend.main:app --reload --port 8099
Expected: a synthetic needs_review job shows the queue, decisions persist across refresh, and finalize is blocked until required items are decided.

  • Step 5: Commit
git add frontend/index.html tests/api/test_review_api.py
git commit -m "Add frontend human review queue"

Task 10: Config, docs, and rollout

Files:

  • Modify: backend/config.py
  • Modify: backend/.env.example
  • Modify: README.md
  • Test: tests/review/test_policy.py, tests/review/test_store.py, tests/review/test_gate.py, tests/agents/test_runner_review_gate.py, tests/api/test_review_api.py, tests/review/test_finalizer.py, tests/review/test_feedback.py, tests/api/test_review_email_flow.py

Interfaces:

  • Consumes: all previous tasks.

  • Produces:

    • AGENT_REQUIRE_REVIEW = true
    • AGENT_REVIEW_AUDIT_SAMPLE = 5
    • REVIEW_AGGREGATE_INCLUDE_TEXT = false
  • Step 1: Add config assertions to existing policy test file

from backend import config


def test_review_defaults():
    assert config.AGENT_REQUIRE_REVIEW is True
    assert config.AGENT_REVIEW_AUDIT_SAMPLE == 5
    assert config.REVIEW_AGGREGATE_INCLUDE_TEXT is False
  • Step 2: Run tests to verify they fail

Run: pytest tests/review/test_policy.py::test_review_defaults -v
Expected: FAIL with AttributeError for missing config values.

  • Step 3: Implement config and docs

Add to backend/config.py:

AGENT_REQUIRE_REVIEW = os.getenv("AGENT_REQUIRE_REVIEW", "true").strip().lower() in ("1", "true", "yes")
AGENT_REVIEW_AUDIT_SAMPLE = int(os.getenv("AGENT_REVIEW_AUDIT_SAMPLE", "5"))
REVIEW_AGGREGATE_INCLUDE_TEXT = os.getenv("REVIEW_AGGREGATE_INCLUDE_TEXT", "false").strip().lower() in ("1", "true", "yes")

Add the same keys to backend/.env.example and document the two-email flow and privacy boundary in README.md.

  • Step 4: Run full test suite

Run: pytest -v
Expected: PASS

  • Step 5: Commit
git add backend/config.py backend/.env.example README.md tests
git commit -m "Configure required agent human review"

Execution Handoff

Plan complete and saved to docs/superpowers/plans/2026-07-28-agent-human-review.md. Two execution options:

1. Subagent-Driven (recommended) - Dispatch a fresh subagent per task, review between tasks, fast iteration.

2. Inline Execution - Execute tasks in this session using executing-plans, batch execution with checkpoints.

Which approach?