SynHy Article

How to Measure Productivity When People and AI Agents Share the Work

A practical attribution model for measuring human and AI agent work together, connecting costs, tasks, outcomes, capacity planning, and finance evidence.

AI Productivity Needs Work-Level Evidence

AI productivity is easy to claim and hard to prove. A company may know how many licenses it bought, how often employees use an assistant, or how many prompts were sent, but those signals do not show what work improved or whether the improvement was worth the cost.

The measurement problem becomes sharper when people and AI agents share the same workflow. A person may plan the task, an agent may draft or code part of it, another person may review it, and the final outcome may depend on all three contributions.

Productivity measurement therefore has to move from tool usage to work-level attribution.

Why License Counts Hide the Truth

License counts measure access, not value. Active users measure adoption, not impact. Prompt volume measures activity, not completed work. None of these numbers tells a manager whether AI reduced cycle time, improved quality, avoided rework, or shifted effort to higher-value tasks.

They can even mislead. A team may generate more AI activity because the work is confusing, because outputs require cleanup, or because staff are experimenting without a production use case. Without a link to the actual work item, the organization cannot separate useful acceleration from noisy activity.

The right unit of analysis is the task, ticket, case, customer request, document, or workflow step that the business already uses to manage work.

What Unattributed AI Work Costs

The direct cost is budget uncertainty. Finance sees vendor invoices, department leaders see anecdotes, and executives hear broad productivity claims, but no one can connect cost to delivered work with enough confidence to fund, cut, or redirect the program.

A simple measurement gap estimate is monthly AI spend multiplied by the percentage of spend not tied to a work item. If a company spends $18,000 per month on AI tools and 70 percent is unattributed, $12,600 of monthly spend is difficult to defend with operating evidence.

The estimate does not mean the spend is wasted. It means the business lacks the evidence needed to decide whether the spend should grow.

A Diagnostic for Blended Work

Choose one workflow where AI assistance is already common, then map the work record. Identify who started the task, what AI tool or agent contributed, what the person changed, which system recorded the output, and how the business judged completion.

  • Is there a baseline for human effort before AI was introduced?
  • Can AI activity be tied to a specific ticket, case, document, or customer request?
  • Does the record distinguish drafting, analysis, action, review, and rework?
  • Are quality and outcome measures attached to the same work item?

If the workflow cannot answer these questions, the ROI discussion will depend on impressions rather than evidence.

Options for Measuring Without Overbuilding

The lightest option is a manual sample. For two weeks, record AI-assisted work items, time saved or added, review effort, and outcome quality. A stronger option is tagging AI-assisted tickets or cases inside the system where the work already lives.

A third option is telemetry: connect AI tool use, human time, cost, and outcomes automatically at the work-item level. A fourth is abstention from ROI claims. If a company cannot measure a workflow yet, it can still run a pilot, but it should not present the pilot as proven productivity improvement.

The measurement approach should match decision importance. Board-level funding deserves stronger evidence than a small team experiment.

The Productivity Attribution Model

A practical model has six fields: work item, baseline effort, human effort, AI contribution, AI cost, and outcome. Baseline effort records how the task normally performed. Human effort records planning, review, and exception handling. AI contribution records drafting, coding, analysis, retrieval, or action. Outcome records cycle time, quality, revenue, risk reduction, or customer result.

The model should also mark confidence. A measured time log is stronger than a rough estimate, and a quality score based on review defects is stronger than a manager's impression. Confidence labels keep early measurements useful without pretending they are more precise than they are.

This turns AI productivity from a tool story into an operating record.

Worked Example: An Engineering Sprint

An engineering team uses AI agents to draft tests, summarize issue context, and propose code changes. Before attribution, leadership sees subscription cost and developer enthusiasm. After attribution, each Jira issue records which AI tool contributed, how much review was required, whether defects changed, and whether cycle time improved.

The result may be mixed. AI may help simple test scaffolding, add review burden on ambiguous changes, and improve documentation consistency. That mixed answer is more useful than a single productivity percentage because it tells leaders where to expand, constrain, or redesign the workflow.

The same pattern works outside engineering for claims processing, estimate preparation, document intake, scheduling, and customer follow-up.

Measures That Make ROI Defensible

Useful measures include cycle time, throughput, review time, rework rate, defect rate, customer response time, avoided manual hours, AI cost per completed work item, and percentage of work items with attributable AI contribution. Measures should be compared to a baseline rather than reported alone.

Finance needs classification evidence too. Some work may support new capability development, while other work supports maintenance, service, or routine operations. The work item should carry enough context that finance and operations can discuss cost treatment without reconstructing the task later.

ROI becomes defensible when cost, contribution, and outcome appear in the same record.

Next Step: Trace One Workflow

Pick one workflow with visible AI use and trace 25 completed work items. For each one, record baseline expectation, AI contribution, human review, outcome, quality issue, and estimated AI cost. Do not average the result until the outliers have been reviewed.

SynHy uses this kind of operating evidence because AI value is rarely uniform across a whole company. It appears in specific workflows where the task is understood, the handoffs are clear, and the measurement record is good enough to guide action.

The first trace should produce a decision: expand the pattern, revise the workflow, narrow the tool, or stop claiming value until better evidence exists.

Sources and Methodology

This article was triggered by Tempo's August 25, 2026 Business Wire announcement of Workforce Intelligence for human and agentic productivity ROI and Tempo's own explanation of workforce intelligence for engineering teams. It also references Kyndryl's 2025 Readiness Report on ROI pressure and FASB material on accounting for and disclosure of software costs.

The attribution model and measurement gap formula are SynHy original analysis. They are designed for operational decision support and should be reconciled with finance policy, accounting guidance, privacy obligations, and labor rules before formal reporting.

Source selection favored the current product signal, vendor documentation explaining the measurement category, and durable finance and readiness sources that show why AI ROI needs evidence beyond adoption metrics.