>
Developer

The four-layer architecture behind a fast huge PR view

When someone tells you they rebuilt a pull-request view so it stops stuttering on big reviews, the natural follow-up is “what did they change?” Most write-ups answer by walking the rendering pipeline top to bottom. That walk is useful, but it buries the interesting part under the mechanism. The interesting part is what the team decided to own, what they decided to delegate to the browser, and where they drew the boundaries between them. This is a decomposition of the GitHub Copilot app’s PR-view rebuild along those boundary lines, so an engineer staring at a similar janky screen can decide which of the four layers they actually need to fix.

If your team’s worst-case review is “a few hundred files and maybe 50 comments,” most of this does not apply. If your team touches migrations, generated changes, vendor bumps, or any diff that crosses a thousand files, you are the audience for this piece. The lines drawn in the GitHub rebuild are exactly the lines you will end up drawing yourself, even if you never ship a Copilot clone.

The four layers, ordered by where they fail

The rebuild splits the surface into four concerns, and they fail in roughly this order, starting from the one that breaks most often:

  • The height model that says how tall each comment row actually is.
  • The rendering loop that places the comments and the lines in the right vertical positions.
  • The data pipeline that feeds both layers without duplicating work.
  • The instrumentation layer that tells you whether any of the above is healthy.

The order is not random. The most common cause of jank in a tall-list view is the height model, not the rendering loop. The reason the height model is so often the bug is that it is the one layer where “guess and patch later” looks reasonable in the design doc and then quietly fails at runtime. Most teams hit this once, decide to live with the glitch, and then the glitch becomes the product. The rebuild’s most expensive decision was to refuse that compromise.

Layer 1: the height model you cannot fake

A line of code has a known height before you measure it (set by font metrics and line height). A comment does not. Its height depends on markdown wrap width, image load state, expandable sections, reply box visibility, reaction counts that decide whether metadata wraps onto a second line, and other small choices that are decided at render time. If you assume an average height and correct later, you have a virtualizer that is mostly right and occasionally wildly wrong. In the rebuild’s worst case, “wildly wrong” was a 600-pixel comment rendered as if it were 30 pixels tall, which means every cumulative height below it was wrong by 570 pixels by the time the user scrolled past the first one.

The proper fix is to know a comment’s height before you place it. That sounds circular, because the height depends on what is inside. The escape hatch is that the height depends on a small, enumerable set of factors (text length, wrap width, whether a quoted reply is expanded, whether the image has loaded), and each of those can be computed quickly without doing the full layout. The rebuild’s height model is, in effect, a closed-form expression over those factors rather than a guess-and-correct estimate. That is the architectural boundary: the browser still does layout for painting, but the renderer does not depend on layout output to decide where to put anything.

If you read the rebuild write-up carefully, the team’s description of this layer is closer to “we made height correct by construction” than “we made the patch faster.” The phrasing matters. Speed on its own would have been an improvement; correctness is what unblocks the rest of the system.

Layer 2: the renderer no longer owns positioning

Once the height model is correct, the rendering layer becomes much smaller. It stops needing a cumulative height table, stops needing to update positions when a single comment’s height changes, and stops needing to gamble on a buffer that might collapse if a tall comment lands in the middle. It also stops being the place where jank shows up. Most of the slowdowns teams remember can be traced back to the layer below, not the one they were looking at.

The architectural win from this is that the renderer becomes a thin presentational layer. It paints lines and comments where the height model says they go. If you ever need to change how comments look, you change a renderer. If you ever need to change how tall comments can be, you change the height model. The two layers stop stepping on each other, which is the kind of separation of concerns that pays back every time a new comment type is added (a reaction here, a suggestion there, an expanded state somewhere else). Each addition becomes a renderer change plus a height-model parameter, not a re-architecture of the whole surface.

Layer 3: the data pipeline as a first-class citizen

The least glamorous layer, and the one most teams under-invest in, is the data pipeline that feeds the surface. A pull request is not one file. It is the metadata, every file’s diff, every comment, every reply, every reaction, every edit, every suggestion, often arriving in waves as the server processes the request and as the client asks for more. A perfect renderer cannot paper over a pipeline that stalls, duplicates work it already did, or throws away results it had computed. A pipeline that works correctly lets the renderer be invisible.

In the rebuild this layer got a real owner, real metrics, and a clear contract with the rest of the system. That is the line most teams fail to draw. They treat the data fetch as a “we’ll add GraphQL later” detail and then wonder why the renderer occasionally blanks for two seconds when a second tab asks for the same PR.

Layer 4: the feedback loop you need before you can iterate

The fourth layer is a measurement loop the team built before they shipped any of the architectural changes: concrete metrics (frame time, scroll latency, comment mount time), an instrumented surface that answers those questions on every build, and an unattended harness that compares a candidate build against the worst-case PR in the corpus. None of the other layers can be improved without this layer, because there is no other honest way to know whether a change helped or hurt.

That is the most underrated part of the rebuild. Most teams land on the architectural redraw because they read about it in someone else’s post-mortem, and then ship the redraw with no way to know whether it made their specific screen faster. The team behind the rebuild did it the other way around: they had the loop first, used it to confirm the failure was real, used it to identify the height model as the dominant cause, used it to confirm the layer-by-layer redraw was moving the numbers, and used it to gate the rollout. Everything else is decoration without that.

The order to build it in, if you have to ship your own

If your team is staring at a similar slow surface and deciding whether to follow this rebuild’s lead, the path of least regret is the same path the rebuild took, in the same order:

  • Build the measurement loop first. Until you have frame time, scroll latency, and a stable comparison harness against your worst-case data, every architectural change is a guess.
  • Fix the data pipeline second. A renderer that does not stall because of bad fetches is a renderer your team can debug at all. Pipeline bugs are the kind that look like renderer bugs and waste the most time.
  • Fix the height model third. This is the expensive decision, both in design and in code. Do it once you have a measurement loop that can tell you whether your fix landed.
  • Treat the renderer as a thin presentational layer fourth. It is much cheaper to change once the layers below it are stable, and once it is cheap to change, your team can react to design feedback without rewriting positioning code.

Trade-offs

The rebuild is real cost, and pretending otherwise would be dishonest. The shape of the trade is the part most teams underestimate.

  • A closed-form height model is more code than a guess-and-patch estimate. You are betting that worst-case PRs are common enough on your team to justify the engineering.
  • A specialized pipeline is more code than a generic fetch on demand. You are betting the same thing one layer down.
  • Stable measurement is real engineering investment. Someone has to own the metrics, the harness, and the comparison infrastructure. It does not maintain itself.
  • New comment types do not get cheaper over time. Each new reaction, suggestion, or expanded state is a parameter in the height model, not a one-line config change.

If your team’s reviews are mostly small and unlikely to grow, none of this matters and a simpler approach is fine. If you have migrations, generated changes, or anything that crosses a thousand files, the trade is worth it. Most teams sit in the middle, and the right answer depends on how often the worst case shows up in the queue.

Coach’s note

If you are about to write a virtualized list with mixed-height children, do not start with the virtualizer. Start by measuring the children. The virtualizer is the easy part; the height model is where every rebuild in this category eventually gets stuck. Get the loop running first, then push the boundaries of what is known about each child before render, and you will land in the same place the Copilot team did without spending the month they spent.

Leave a comment