Not everything left after handing work to AI should be cut. This piece separates the judgments people need to make from the work of explaining and restoring the same judgments again.
While AI drafts something, I can work on other things. But once the draft is done, I have to come back. I need to check whether the conditions I already set were reflected, whether facts and guesses got mixed together, and whether this round of edits changed something else.
The problem I kept running into wasn’t a lack of output. It was that the work of making the output usable for real work kept remaining.
That doesn’t mean all of the remaining work should be cut. The judgments a person has to make are different from the work of explaining and restoring the same judgments again. Telling these two apart is where this piece starts.
A day of building a website
Recently I built the TYPE website with AI. Over one day I revised the research introduction, the articles, and the service pages in turn. Looking back, what I did fell into three kinds.
First, I said again what had already been decided. The paper title in the work brief didn’t match the title in the latest manuscript. AI pointed out that the two titles differed, and I answered which one was current. The title had already been decided. That decision just hadn’t made its way into the brief.
Second, I explained again a distinction I had already made. The page for a new service I was preparing also listed our existing services. To me, they were obviously separate, but that distinction wasn’t written down anywhere. I had to say again, “This is a new service. It’s different from the existing ones.”
Third, I made a judgment for the first time: how much of the research to show on the site. At first I felt too much was exposed and cut it back. Then I felt readers didn’t have enough to judge it, and added it back. No one could have decided this in advance. It was something I had to decide by weighing the purpose of the site against the value of the research.
Splitting the remaining work in two
On the surface, all three scenes looked alike. I was answering a request to confirm something, or looking at a result and asking for changes. But they were different in kind.
The first two were restoration. A judgment I already had wasn’t kept in a form AI could apply to the next task. In the first scene, a changed decision hadn’t been reflected in the record. In the second, a distinction that existed only in my head had never been written down. Either way, I wasn’t thinking anything new; I was pulling out a judgment I already had. This kind of work can be reduced, and it’s better to reduce it.
The third was judgment. It was a new question that came up as the purpose of the site became more concrete, and the answer had to come from my own criteria. This isn’t something to reduce. If anything, moments like this should reach me clearly, instead of getting buried under restoration work.
The hard part is that the two arrive looking the same. Both come as a question like “Is this right?” or as a request for changes. So if you remove every check to cut the remaining work, the judgments disappear too; and if you check everything to protect judgment, the cost of restoration stays.
The remaining work in research, too
Related research also looks at the verification, integration, and management work left after using AI. Lee et al. analyzed 936 examples of generative AI use shared by 319 knowledge workers, and describe how, when people use AI, their critical thinking shifts toward verifying information, integrating responses, and stewarding the task.[1] Even when there is less making, the checking and connecting remain.
The remaining work isn’t always light, either. In an experiment with experienced developers, task completion time actually increased under the condition where they could use the AI tools of the time.[2][3] Getting a result quickly and finishing the work can be two different things.
Problems on the restoration side are also reported in AI development. Anthropic’s write-up on long-running agents describes work state not carrying over well enough between sessions, and agents declaring a whole task done after building only part of it. In response, they use progress logs, Git history, feature lists, and verification steps.[4] The two scenes I went through are a close problem, in that a judgment that already existed wasn’t kept in a form the next task could use.
The questions I’m studying
So what interests me is less how much of the remaining work we can cut, and more which work to keep and which to cut. These are the three questions I’m holding onto now.
- Under what conditions should a judgment already made carry straight into the next task?
- How can we recognize the moments when a person needs to make a new judgment, and convey them clearly?
- Does the cost of keeping this distinction end up larger than the cost we’re trying to cut?
I’m studying these questions as a design problem in human–AI collaboration (the research page is in Korean). The current approach is a design proposal grounded in the literature, not a method whose effect has been proven.
After we hand work to AI, what do we keep doing? Of that, which judgments should people make themselves, and what shouldn’t we have to explain again?
In the next piece, I want to look at what the instructions, records, and review practices used to carry decided judgments into the next task actually solve, and what they leave behind.
References
- Lee et al. (2025). The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers. CHI 2025. Source
A self-report survey study; it does not establish increased work time or causation. - METR (2025-07-10). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. Source
A randomized experiment with 16 experienced developers on 246 real issues. Under early-2025 tools and large, familiar repositories, completion time increased by 19%; it cannot be generalized to all work. - METR (2026-02-24). We are Changing our Developer Productivity Experiment Design. Source
The researchers’ follow-up. Alongside possible improvements, they note that participation and task-selection bias make it hard to estimate the current effect precisely. - Anthropic (2025-11-26). Effective harnesses for long-running agents. Source
A developer’s engineering write-up; not a method validated for all kinds of work.