How it works
Bubo orchestrates code review across SCM access, agent execution, state, filtering, confidence evaluation, and feedback loops. Each run produces idempotent, inline, graded, and measurable feedback.
Architecture
Discover → review → filter → post → learn
The pipeline, step by step
- Discover. List open MRs/PRs per project, skipping any already reviewed at the current head SHA.
- Review. Fork a worker, check out the diff, capture AI-provenance for the governance trail, run the agent skill, and parse the findings.
- Filter. Normalize each finding’s category to the canonical taxonomy, then
apply the operator’s policy — confidence floor (optionally per-class
calibrated), the
gate/collaboratesurface mode, and dispute suppression. Optionally run independent verification lenses; a finding the majority refute is recordedREFUTEDinstead of posted. - Post. Map each surviving finding to a changed line and post it inline — or
store it as
PLANNEDwhendry_runis on. - Acknowledge. Zero findings → one change-level “all good” comment, so “reviewed and happy” is distinguishable from “never ran”.
- Persist. SQLite records reviewed SHAs and finding fingerprints, so Bubo never spams or double-posts across repeated polls.
- Grade.
--sync-outcomeslater checks which findings were resolved, replied to, disputed, deleted, or merged unresolved — the signal behind the metrics and the learning loop above.
SQLite stores reviewed SHAs and finding fingerprints for deduplication. Outcome sync feeds resolved and disputed signals into calibration.
- State closes back on discovery. SQLite remembers reviewed SHAs and finding fingerprints, so the poller skips work it has already done and never double-posts.
- Outcomes close back on precision.
--sync-outcomesgrades what humans did with each finding; those accept/dispute rates feed dispute suppression and per-class confidence calibration — so the filter gets sharper on the categories this team rejects.
Everything above is orchestration: provider access, forking, state assembly, the deterministic filter/verify/post path, and telemetry.
The review judgment comes from the review skill you configure — swap the model or the skill and the orchestration is unchanged. See Configuration reference for every lever, and Metrics & telemetry for the spans and outcome grading.