A new software or AI tool can feel helpful without improving the whole process. Productivity assessment should measure completed, acceptable work and its consequences, including checking, correction and handover. Activity counts and personal enthusiasm provide only part of that evidence.
Define the outcome and quality threshold
Choose a task that matters to the business and specify when it is complete. A draft article is different from an approved article; a suggested code change is different from a reviewed, tested change ready for use. Apply the same quality threshold to both comparison routes.
Record errors, rework and unresolved tasks rather than counting only successes. A faster first draft can lose its advantage if another person spends more time checking unsupported claims or correcting formatting.
Measure the whole effort
Include preparation, prompting, implementation, review, testing, correction and operational handover. Distinguish active work from waiting time when relevant. Note who performs each stage and whether effort shifts to another team.
Keep the observation process proportionate. A simple agreed time log and review record can be more useful than collecting many activity signals that do not answer the question. Protect private customer and staff information during measurement.
Compare suitable tasks
Use tasks with sufficiently similar difficulty and context. Record experience, familiarity, tooling and relevant changes during the comparison. Where feasible, suitable random allocation can reduce selection bias; a small informal trial should retain its limitations.
A tool used only on easy tasks cannot fairly be compared with the normal route's hardest tasks. Likewise, repeated work may become quicker through learning. Explain the comparison method before presenting a percentage improvement.
Keep published research in context
METR's July 2025 study of early-2025 AI tools examined experienced open-source developers working on familiar repositories. Tasks were randomly assigned to allow or disallow AI use. Observed completion time and developers' perceived benefit differed.
That historical result concerns a particular setting and tool period. It is not Giraffe Digital research and does not establish the review cost of publishing, the performance of every developer or the value of today's tools. Its useful lesson here is to measure the actual process rather than assuming perceived speed equals completed-work improvement.
Separate time, capacity and cash
Time released might allow more work, faster customer responses or reduced effort. Those are different outcomes from a reduction in expenditure. Record the benefit that actually occurred and avoid counting the same improvement more than once.
Include licensing and ongoing maintenance where assessing costs. Check whether another bottleneck prevents additional throughput. A quicker drafting stage may have little effect if approval capacity remains unchanged.
Review the decision over time
Keep the tool version, configuration, task selection and review date in the record. Repeat an assessment when the process changes materially, rather than treating one result as permanent. Retain known weaknesses and the conditions under which the tool is useful.
Use a proportionate decision: adopt for suitable work, improve the workflow, narrow its use or stop the trial. A mixed result can support a more specific application instead of a sweeping claim about all software or AI.
Productivity review checklist
- Define complete work and an equal quality threshold.
- Include preparation, review and correction effort.
- Compare suitable tasks with documented limitations.
- Keep historical research distinct from your results.
- Separate capacity benefits from cash savings.
- Record the decision and conditions for another review.
Our AI evaluation guide covers answer quality. The dashboard guide explains useful outcome reporting.


