GitHub cut the review wait into three parts.
GitHub added a pull_request_review_times array to its daily repository-level Copilot usage reports. Each qualifying row now has median and 90th-percentile minutes for ready to first review, first to final review, and final review to merge.[1]
The split matters because each delay has a different owner. A long first wait points toward reviewer discovery or load. A long review conversation points toward change size, unclear intent, or unresolved evidence. A long merge tail points toward queues, branch policy, release windows, or ownership after approval.
Do not tune reviewer speed when the approved patch is parked behind the merge gate.
The denominator is narrower than the headline.
The report counts pull requests opened by a person and reviewed by at least one other person. Bot reviews and author reviews do not qualify. The stage count will usually be lower than the general merged pull request count. Data begins with the release and has no backfill. A day with no qualifying merges returns an empty array, not zero time.[1]
The REST documentation also limits who can fetch the report. Enterprise owners, billing managers, organization owners, and authorized roles need the relevant Copilot metrics permission and policy. The endpoint returns signed download links for daily report files.[2]
A percentile is a tail alarm, not a total.
Put the median beside the p90 for each stage. The median describes the middle qualifying pull request. The p90 exposes the slower tail. A wide gap means some work experiences a much worse queue than the middle case.
Do not add the three p90 stage values and label the result the end-to-end p90. The 90th-percentile pull request in one stage may not be the same pull request in another stage. The same caution applies to adding medians. Use the stage values to locate friction. Measure the full ready-to-merge interval separately if you need its distribution.
Metrics should start a repair, not a contest.
DORA recommends measuring one application or service in its own context. It warns against turning a metric into a goal, comparing unlike systems, and collecting data without making an improvement. Its suggested loop is practical: set a baseline, discuss the delivery friction, pick the biggest constraint, make one change, and check progress.[3]
- First review wait is longest. Assign review rotations, narrow ownership, or route by changed area.
- Review conversation is longest. Cut batch size. Put intent, screenshots, tests, and risky decisions in the opening packet.
- Merge tail is longest. Inspect required checks, queue policy, branch freshness, release windows, and who lands approved work.
- The p90 gap is wide. Sample the slow pull requests. Look for one repository, team boundary, change class, or time zone that the median hides.
Make one repair. Keep the denominator, date range, repository, median, and p90 beside the result. Then compare the same stage after enough qualifying merges have accumulated.