Overview

Bubo 🦉

Bubo is patient — it sits silently, sees all, and strikes only when sure. It watches your repositories and speaks only when it finds something.

Agentic AI code review — with the LLM of your choice

01Self-hosted automated review and inline posting
02Bring-your-own-LLM
03GitLab or GitHub
04Built-in metrics
05Governance, provenance, audit, and ROI metrics
06MCP for on-demand reviews

Trust

Security & compliance posture

Secrets never leak

Sensitive information is redacted before it touches the LLM, reports, logs, or the database.

Signed & attested

Release artifacts are cosign-signed keyless via Sigstore + GitHub OIDC, carry SLSA Build L3 provenance, and ship an SBOM in SPDX JSON — all independently verifiable.

Report a vuln

Disclose responsibly per SECURITY.md.

Get started

See it in action

What a review looks like

merge request · inline finding
config/settings.example+ api_token = "••••••••••••••••"
LR
LLM Reviewer · botissue · blocking · security

Hard-coded credential committed

Impact anyone with repo access can use the token until it is revoked.

Fix remove it, rotate, and inject via a CI secret.

confidence 0.98
A blocking finding, posted inline — structured impact / evidence / fix with a confidence score.
recall · learning
LR
LLM Reviewer · botstyle · dismissed

Naming nit — skipped, not re-posted

Why this category was dismissed here before; suppression is on.

Recall & learning — a class your team keeps dismissing stops getting posted.

Real production metrics

Measured RoI

Real production numbersMeasured May 28 – Aug 10, 2026Last updated August 10, 2026

Speed is a given. Bubo's edge is signal density: multi-agent orchestration and deep codebase context surface only real structural and security flaws — in minutes — saving hundreds of senior-engineer hours.

75.3k
Added LoC reviewed
across 108 instrumented runs in 5 projects
0
False positives · 0 disputes
across 191 tracked outcomes
156
Findings accepted
91 blocking · 81.7% acceptance rate
~$22.2K
Reviewer time returned
for $36.83 of recorded cost — est. (see note)

Return on investment

Loaded reviewer rate
~$90/hr
Reviewer time returned
~$22.2K
Return on recorded spend
~602×
Recorded cost / accepted finding
$0.24
Accepted findings / recorded $
4.24
Peer AI reviewers
$24–30/dev-mo

Reviewer time freed

Added LoC reviewed
75,320
Human first pass @300 LoC/hr
251.1 hrs
Bubo review wall time
4.6 hrs
Reviewer hours returned (est.)
~246.5
Reviewer-days returned (est.)
~30.8
Cost / reviewer-hour
$0.15

Outcomes & precision

Findings returned
218
Posted
194 (89.0%)
Outcomes tracked
191
Developer replied
155
Accepted / resolved
156
— of which blocking
91
Acceptance rate
81.7%
Disputed
0
False positives
0

Code reviewed

Projects
5
Added LoC measured
75,320
LoC-instrumented runs
108 of 276
Instrumentation coverage
39.1%
Blocking findings
126
Non-blocking findings
92

Cost & efficiency

Total recorded cost
$36.83
Total tokens
25,565,663
Avg tokens / run
92,629
Avg recorded cost / run
$0.13
Runs with recorded cost
218 of 276
Latency p50 / p95
128s / 308s

Aggregates

Period
May 28 – Aug 10, 2026
Total review runs
276
Successful or clean
269
Completion rate
97.5%
Failed runs
6 (2.2%)
Total run wall time
11.1 hrs

Note: these are real production numbers, not a benchmark or simulation. They are drawn from all Bubo history in the live production database across five GitLab projects, measured continuously from May 28, 2026 through August 10, 2026. Last updated August 10, 2026.

Coverage is partial and stated rather than smoothed over. Reviewer-time figures use only the 108 of 276 runs (39.1%) carrying LoC instrumentation, and count added lines only. Recorded provider cost covers 218 of 276 runs (79.0%) — 58 runs have zero or missing model pricing, so the true spend is higher than $36.83. The ~602× return charges that full recorded cost against only the reviewer-time value measurable in the instrumented subset; it is a directional provider-cost ratio, not audited all-in ROI. Host, engineering, and operational costs are not recorded. “Accepted” means resolved and neither disputed nor marked false-positive; zero recorded false positives reflects the outcome labels present in production, not independently proven precision.

Reviewer-hours and dollar figures are estimates. Careful peer review runs ~200–400 LoC/hr (300 midpoint); review effectiveness drops sharply above ~200 LoC/hr.[1][2] Reviewer time is valued at a fully-loaded ~$90/hr — the BLS May 2024 median for software developers ($133,080, ~$64/hr) at a conservative 1.4× benefits/overhead load.[3] Peer AI reviewers list at ~$24–30/developer-month (CodeRabbit, Greptile, Qodo, Graphite).[4] Bubo performs technical review — correctness, structure, and security — not subject-matter-expert or domain-specific review, and reads diff LoC at scale where a human would sample and skim mechanical changes; treat the figures as directional.

  • [1] Kemerer & Paulk, “The Impact of Design and Code Reviews on Software Quality,” IEEE Trans. Software Eng., 2009. sites.pitt.edu
  • [2] SmartBear, “Best Practices for Peer Code Review” (SmartBear/Cisco study). smartbear.com
  • [3] U.S. Bureau of Labor Statistics, “Software Developers,” Occupational Outlook Handbook, May 2024. bls.gov
  • [4] Vendor list pricing, accessed 2026: CodeRabbit, Greptile, Qodo, Graphite.
MountainOwlMountainOwl
Bubo · MIT licensed · © 2026