What to know

  • The reports split review time into three stages.
  • Only qualifying human-authored and human-reviewed changes are included.
  • Missing observations must not be interpreted as zero delay.

A more detailed view of review time

GitHub expanded its repository-level Copilot usage reports on September 25 with timing for three stages: ready for review to first review, first to final review, and final review to merge. The API reports medians and 90th percentiles, allowing teams to distinguish typical behavior from longer delays.

The data covers qualifying pull requests opened by a person and reviewed by another person. Bot and author reviews do not count as timed reviews. GitHub also says the data is not backfilled and that days without qualifying merges return an empty array. Those limits are essential when comparing the new fields with broader merge totals.

Source: GitHub: Pull request review stage metrics

Different waits call for different explanations

A long wait before the first review can suggest a capacity or routing problem. A long interval between reviews may reflect substantial revision, difficult requirements or fragmented discussion. A delay after final review may involve release coordination or an additional check. The measurement helps locate the interval; it does not establish the cause on its own.

Teams should examine representative changes before deciding what to alter. A complex security repair and a routine documentation update should not necessarily be expected to move at the same speed. Combining them without context can encourage a faster-looking process that simply changes the mix of work being measured.

The higher percentile is useful for examining uneven experience. A reasonable median can coexist with a subset of changes that wait much longer. Those cases may reveal ownership gaps or specialized knowledge concentrated in one reviewer. Looking at the actual changes helps distinguish a recurring bottleneck from a small number of unusual events.

Avoid turning an incomplete measure into a target

Byte Watchr’s assessment is that the new fields are best used to investigate a process rather than rank individuals. Review time includes waiting, coordination and the characteristics of the work. Treating the number as a direct measure of personal productivity would discard much of the context needed to interpret it fairly.

Early comparisons also need enough observations. An empty day is not evidence that every change moved instantly, and a short initial period may contain too few qualifying merges to support a stable trend. Reporting the count beside each timing measure makes those limits visible and helps prevent misleading before-and-after claims.

For organizations evaluating coding assistants, the release adds a useful part of the evidence. Faster code generation can move work into review sooner without shortening the complete delivery cycle. Measuring where the resulting changes spend time gives teams a more concrete basis for deciding whether to improve generation, review capacity, test reliability or release coordination next.

Sources & further reading

  1. GitHub: Pull request review stage metrics

Factual statements are grounded in the linked material. Interpretation and illustrative examples are Byte Watchr analysis. Vendor claims are identified as claims, rather than independent testing.

The event date records the source announcement or documented operation. The coverage edition groups recent developments and is separate from the publication date. Actual publication is recorded above.

Corrections policy · About this byline