One developer pulled the plug for a month.
The author of "One month without AI" describes an individual experience, not a controlled study. He reports running several coding agents at once, spending longer reviewing some small tasks than he expected to spend writing them, and finding a generated test that did not cover the changed scenario.[1] After turning AI features off, he says he returned to smaller changes, direct test-driven development, and code he could defend in review.
That account cannot tell us what will happen on another team. It does name measurements worth collecting. Track review delay, context switches, changes you can explain, tests you can repair without a transcript, and time to recover after a bad suggestion.
Speed at the generation step can become debt at the explanation step. Time both.
The tools now remember and multiply.
GitHub says agentic autofix can read enabled Copilot memories and store a successful fix pattern as memory for later security alerts. GitHub also says those memories can inform code review and its cloud agent. Both agentic autofix and Copilot Memory were in public preview when the announcement was published.[2]
A useful pattern can save repeated work. A weak pattern can also travel. The product announcement does not claim that every stored fix is correct. Teams need a review route for what enters memory, where it is reused, and how it is removed.
VS Code 1.139 also made large agent-session lists faster. Microsoft reports a vendor benchmark on one development machine with about 645 sessions. The first listing fell from 1.3 seconds to 0.1 seconds, while refresh fell from 0.6 seconds to 0.15 seconds. Microsoft warns that people with fewer sessions should expect a smaller difference.[3] Faster catalogs help operators find work. They do not prove that a person understood the work inside each session.
The off switch is a test instrument.
VS Code documents chat.disableAIFeatures as the setting that hides built-in chat and inline suggestions and disables the Copilot extensions. It can be set for one workspace or for the user.[4] That makes a bounded comparison practical. You do not need a permanent vow or a team-wide policy to collect a manual baseline.
Pick one representative task. Disable AI features for that workspace. Record elapsed time, interruptions, failed tests, review corrections, and whether you can explain each changed file. Run a similar task with your normal assistant setup. Compare the whole delivery cycle, not output volume.
Put assistance on a gearbox.
- Off. Use a manual run to measure your current understanding and tool fluency.
- Pair. Ask for search, explanation, or alternatives. Keep the human on the keyboard for the final change.
- Delegate. Let the assistant produce a patch only when the review budget, test boundary, and recovery route are explicit.
The right gear can change by task. A routine rename and an unfamiliar authorization change do not need the same assistance policy. Keep the mode visible in the review card so the reviewer knows how the patch was produced.