One edit can arrive in several shapes.
VS Code documents two visible kinds of Copilot inline suggestion. Ghost text continues code at the cursor. Next edit suggestions predict another edit and its location. The new unified system can choose current-line text, a nearby rewrite, a farther edit, or no suggestion.[1]
GitHub and Microsoft say they trained one model to cover completion, next-edit, and longer-distance edit behavior. A shared diff-patch format lets one response carry several ordered edits. The client can cache later patches and present them after the developer accepts the first change.[2]
The model proposes an edit. The editor decides when, where, and whether that edit deserves attention.
A client rule moved the dismissal result.
The team's second technical post describes an inherited client behavior. Ignored ghost text could return as a next edit after the cursor moved. The system treated those as different views. The developer saw the same unwanted suggestion twice.
In the reported A/B test, removing that behavior changed the dismissal comparison by 26 percentage points. The reported result moved from a 15.9 percent increase to a 10.1 percent decrease against the production setup. This is a first-party experiment on GitHub's product. It does not predict results in another editor or codebase.[3]
Fast can still interrupt.
The team also tested speculative decoding, cache delay, debounce timing, progressive reveal, and diff-based rendering. Their chosen setup speculated the current-line completion to reduce latency. It delayed some cached suggestions so they matched the developer's typing rhythm. Longer ghost text appeared in stages rather than as one large block.
Those choices expose the real control problem. A suggestion can be correct and still arrive too early, cover too much code, or repeat after dismissal. Test interruption cost beside acceptance. Record reverts, repeated rejects, retained code, latency, and the cases where silence was the best response.
Give people a brake.
VS Code lets users collapse next edits, disable suggestions by language, and snooze all inline suggestions in five-minute steps. A metered connection blocks new automatic requests while keeping explicit requests available.[1] These are product controls, not decoration. They let a developer reclaim attention without uninstalling the feature.
Start with one workflow and one week of evidence. Keep a short set of must-not-repeat failures. Test the same model with the actual renderer and timing rules. If rejected text returns in another shape, count that as one repeated interruption.