Pimp My IDE / Garage Dispatch
Back to garage
September 26, 2026 | pull requests / review queues / metrics

Your merge time has three waiting rooms.

GitHub's new review-stage report separates the wait for a first review, the review conversation, and the approved pull request that still has not merged.

The take. A single merge-time number tells you that work waited. Stage medians and p90s tell you where. Repair that queue instead of asking every reviewer to go faster.
Split the review clock

GitHub cut the review wait into three parts.

GitHub added a pull_request_review_times array to its daily repository-level Copilot usage reports. Each qualifying row now has median and 90th-percentile minutes for ready to first review, first to final review, and final review to merge.[1]

The split matters because each delay has a different owner. A long first wait points toward reviewer discovery or load. A long review conversation points toward change size, unclear intent, or unresolved evidence. A long merge tail points toward queues, branch policy, release windows, or ownership after approval.

Do not tune reviewer speed when the approved patch is parked behind the merge gate.

The denominator is narrower than the headline.

The report counts pull requests opened by a person and reviewed by at least one other person. Bot reviews and author reviews do not qualify. The stage count will usually be lower than the general merged pull request count. Data begins with the release and has no backfill. A day with no qualifying merges returns an empty array, not zero time.[1]

The REST documentation also limits who can fetch the report. Enterprise owners, billing managers, organization owners, and authorized roles need the relevant Copilot metrics permission and policy. The endpoint returns signed download links for daily report files.[2]

A percentile is a tail alarm, not a total.

Put the median beside the p90 for each stage. The median describes the middle qualifying pull request. The p90 exposes the slower tail. A wide gap means some work experiences a much worse queue than the middle case.

Do not add the three p90 stage values and label the result the end-to-end p90. The 90th-percentile pull request in one stage may not be the same pull request in another stage. The same caution applies to adding medians. Use the stage values to locate friction. Measure the full ready-to-merge interval separately if you need its distribution.

Metrics should start a repair, not a contest.

DORA recommends measuring one application or service in its own context. It warns against turning a metric into a goal, comparing unlike systems, and collecting data without making an improvement. Its suggested loop is practical: set a baseline, discuss the delivery friction, pick the biggest constraint, make one change, and check progress.[3]

  1. First review wait is longest. Assign review rotations, narrow ownership, or route by changed area.
  2. Review conversation is longest. Cut batch size. Put intent, screenshots, tests, and risky decisions in the opening packet.
  3. Merge tail is longest. Inspect required checks, queue policy, branch freshness, release windows, and who lands approved work.
  4. The p90 gap is wide. Sample the slow pull requests. Look for one repository, team boundary, change class, or time zone that the median hides.

Make one repair. Keep the denominator, date range, repository, median, and p90 beside the result. Then compare the same stage after enough qualifying merges have accumulated.

Interactive makeover / review queue instrument

Review queue split timer

Traditional purpose replaced: stare at one merge-time average and blame the whole review process. Better version: compare the middle case with the slow tail across three named stages, then print one bounded repair card.

Load one repository snapshot

Enter the median and p90 minutes from one report period. The defaults are an example, not production data.

Inspection line
1. READY TO FIRST REVIEW45 / 180 min
0 min480 min
0 min960 min
2. FIRST TO FINAL REVIEW90 / 360 min
0 min480 min
0 min960 min
3. FINAL REVIEW TO MERGE25 / 120 min
0 min480 min
0 min960 min
MEDIAN LINE / EXAMPLE DATAsample size required
Ready to first review
45 minreview discovery
First to final review
90 minreview conversation
Final review to merge
25 minapproved merge tail
Longest selected stageReview conversation
Stage sum, not total160 min

Inspect the review packet.

The example median spends the most time between first and final review. Check change size, stated intent, test evidence, and unresolved decisions.

Print the queue repair card

Replace the required fields with repository, date range, qualifying count, and the action owner.

Sources read, not vibes

Open the source log
  1. GitHub Changelog, "Usage metrics API adds pull request review stages", September 25, 2026. Field names, three stage definitions, median and p90 values, qualifying-review rules, no-backfill boundary, and empty-array behavior.
  2. GitHub Docs, REST API endpoints for Copilot usage metrics, API version 2026-03-10. Report endpoints, signed download links, policy requirements, permissions, and report scope.
  3. DORA, "DORA's software delivery performance metrics", updated January 5, 2026. Context-specific measurement, metric pitfalls, bottleneck selection, code-review time as a local indicator, and the baseline-change-recheck loop.

Source boundary. GitHub defines the report and its denominator. DORA supplies general measurement guidance. Pimp My IDE did not access a private report or measure a repository. The values in the split timer are illustrative. The repair suggestions are operating hypotheses, not findings.