All field notes

Evidence and evaluation

How do you tell whether AI is actually giving your team time back?

Measure the whole task, from preparation to an accepted result. Include review, corrections, and upkeep, and compare similar work at the same quality standard.

The short answer

Measure the time it takes people to reach an accepted result, including preparation, review, corrections, and upkeep. Compare similar tasks at the same quality standard. A faster first draft is promising, but it does not tell you whether the whole process gives your team time back.

I would start with a small, visible record of the work. It can be an ordinary spreadsheet. The important part is agreeing on what counts before the tool begins to feel indispensable.

Research is a reason to measure locally

In Generative AI at Work, Brynjolfsson, Li, and Raymond studied an AI assistant used by 5,172 customer support agents. The revised paper reports a 15 percent average increase in issues resolved per hour, with substantial differences across workers. That is evidence from a particular support setting, not a forecast for your nonprofit.

Other settings have produced different results. METR's early-2025 randomized study found that 16 experienced open-source developers took, on average, 19 percent longer to complete tasks in repositories they knew well when AI tools were allowed, even though they believed the tools helped them work faster.

That result has a date and a context. In its February 2026 update, METR thought developers were probably benefiting more from newer tools, based on conversations with participants. It also explained why selection effects and time measurement problems made its follow-up data unreliable for estimating the current effect.

My practical inference is to measure the work in front of us. A result from another organization can suggest questions to ask. It cannot stand in for our own baseline.

Define a finished task

Choose one recurring task and write down the finish line. For a weekly announcement, that might mean approved copy with correct facts, working links, and formatting ready for the person who publishes it. Keep that standard the same with and without AI.

Record several examples of the current process before changing it. Include typical work and difficult work. Note differences such as length, missing information, and the experience of the person completing it. Comparing a short, complete brief with a long, disorganized one tells us little about the tool.

Count all the work

For each attempt, I would record:

  • Preparation: finding material, removing information the tool should not receive, and giving instructions.
  • Production: drafting, prompting again, and any waiting that requires someone's attention.
  • Review and correction: checking facts, repairing omissions, and getting approval.
  • Upkeep: adjusting templates, repairing connections, or helping someone use the workflow.
  • Outcome: accepted, revised again, abandoned, or completed using the previous method.

Record one-time setup separately from recurring upkeep. Do not let a failed attempt disappear from the log when someone finishes the task manually.

Also distinguish staff effort from elapsed time. A draft can take longer to arrive while requiring less attention. Conversely, a quick response can demand so much supervision that it interrupts other work. Both are worth knowing.

A small example

Suppose, purely as an illustration, an announcement takes 30 minutes by the current method. With AI, gathering the input takes 6 minutes, producing the draft takes 3, and review and correction take 17. The total is 26 minutes: a possible saving of 4 minutes on that example, before any additional upkeep.

Calling that “a draft in three minutes” would leave most of the work out. Nor would one example justify multiplying four minutes across a year. Try comparable work repeatedly and keep the exceptions visible.

Look beyond speed

I would review the record with the people doing the work. Are errors being caught earlier? Does a second person now inherit the checking? Are staff spending the released time on something useful, or switching between more unfinished tasks?

Speed, quality, and the experience of work deserve separate observations. A tool might make writing less exhausting without saving minutes. That can matter, but it should be described as a reported experience rather than a measured productivity gain.

Agree on a review date and decide whether to continue, revise, or stop. Repeat the comparison when the task or tool changes materially. I want a time-saving claim to describe a result we can explain: whose time, doing which work, at what standard, and over what period.