Guide

How to Measure AI Adoption Beyond Licenses and Active Users

Joring · September 6, 2026 · 12 min read

Most companies measure AI adoption the way they measured SaaS adoption: licenses, logins, active users, usage volume. Those numbers are easy to collect and easy to move, and none of them answers the question leadership is asking. Is AI changing how work gets done?

The gap is expensive. In 2024, three quarters of knowledge workers said they were already using generative AI at work, and most of them were bringing their own tools rather than waiting for a rollout.1 A year later, an MIT report on enterprise deployments concluded that 95 percent of organizations were getting zero return on their generative AI investment.2 Usage is everywhere. Evidence of better work is rare, and the way adoption is measured is a large part of why.

This guide sets out four levels of measurement, a scorecard you can copy, the mistakes that make adoption dashboards look better than the work, and a cadence for running it. The worked examples use round, illustrative numbers.

  1. 1Access

    Can employees use approved AI tools?

  2. 2Adoption

    Are employees using them?

  3. 3Proficiency

    Are they using AI effectively for valuable work?

  4. 4Impact

    Is AI improving productivity, quality or outcomes?

Figure 1. The four levels of AI adoption measurement. Each level is a prerequisite for the next.
On this page

Usage is not the same as AI adoption

Start with two people on the same sales team. Both show up as “active” in every license report.

AttributeEmployee AEmployee B
Sessions per dayFiveTwo
Typical request“Summarize this” and quick factual questionsPrepares a discovery call from CRM notes and the prospect’s last three support tickets
Company contextNone. Copies answers into email by handPastes the account history and the qualification checklist into the prompt
Repeatable workflowNoThe same three-step routine before every first call
Effect on the taskSaves a few minutes of searchingCall preparation takes 20 minutes instead of 45
Table 1. Two employees on the same sales team. Both count as “active” in every license report. Illustrative.

Each level of the framework sees these two differently. Access says both can use AI. Adoption says both do, and that A does more of it. Only proficiency separates them, and only impact tells you that B’s change is worth spreading to the rest of the team.

The four levels of AI adoption measurement

Each level answers a different question and each is a prerequisite for the next. Reporting only the first two is how organizations end up with a high adoption rate and no idea whether anything improved.

Level 1: Access. Can employees use AI?

Access is the share of the workforce that could use an approved AI tool today: a license, a seat on an enterprise plan, or an SSO entitlement to a sanctioned product. It is the easiest level to measure and the easiest to mistake for progress.

Metrics

  • Employees with AI access
  • Licensed employees
  • Access by department
  • Approved AI tools available

Where the data comes from. The identity provider and each vendor’s license admin console, joined to the HR system for department and role.

Watch out for. Access is a prerequisite for adoption, not evidence of it. A fully licensed workforce with 30 percent usage is a rollout problem, not an adoption success.

Level 2: Adoption. Are employees using AI?

Adoption is whether people with access use the tools, how often, and whether they keep coming back. This is the level most dashboards stop at, and it is where the headline number lives.

Metrics

  • Weekly active AI users
  • Monthly active AI users
  • Percentage of licensed employees using AI
  • Usage frequency
  • Usage by department
  • Usage by tool
  • Retention over time

The core formula is simple: AI adoption rate = active AI users ÷ employees with AI access. If 600 of 1,000 licensed employees used an approved tool in the last 30 days, that is a 60 percent monthly adoption rate. Measured weekly, the same organization might sit at 41 percent. Both numbers are true; the weekly one is the one to manage.

Where the data comes from. Vendor admin reports (the ChatGPT Enterprise, Copilot and Gemini consoles each expose active-user counts), browser or network telemetry for tools outside the approved list, and identity-provider sign-in logs as a floor.

Watch out for. Define “active” as at least one real interaction, not a login. Report weekly as well as monthly; the gap between the two tells you how many people are occasional visitors. And report by department and by tool, because a company-wide rate hides the sales team at 72 percent and finance at 18.

Level 3: Proficiency. Are employees using AI effectively for valuable work?

Proficiency is where most measurement programs stop short, because it cannot be read off a license report. It has four observable parts.

Workflow complexity

The same tool, two very different requests.

Basic

“Summarize this meeting.”

More sophisticated

“Review the meeting transcript, identify objections, compare them with CRM account history, identify unresolved questions, and draft follow-up actions.”

Figure 2. Two requests to the same assistant. The second carries context, spans several steps and produces something the employee can act on.

Complexity cuts both ways. In a field experiment with 758 consultants, people using GPT-4 on tasks within the model’s capability completed 12.2 percent more tasks, 25.1 percent faster, at 40 percent higher rated quality. On a task chosen to sit outside that capability, the people using AI were less likely to reach the right answer than those working without it.3 Knowing what people use AI for matters because the same tool helps or harms depending on the task.

Use-case diversity

Is AI applied to one kind of task, or across the work a role does? The categories worth tracking are research, analysis, decision support, customer preparation, planning, document creation and repetitive operational workflows. A team whose usage sits entirely in “document creation” has found one use, not adopted a capability.

Organizational context

Generic prompts get generic output. Proficient use brings the organization’s own information into the work: CRM data, internal documentation, customer information, company policies, analytics and the approved tools that hold them. The share of usage that carries this kind of context is one of the cleanest proficiency signals available.

Repeatable workflows

The clearest sign of proficiency is a task a role does the same way, with AI, every time. In sales, preparing for a discovery call. In product, analyzing customer interviews. In customer success, preparing a quarterly business review. In finance, explaining the variance between forecast and actuals. Count them per team, and count the ones with an owner.

The behaviors underneath

Published frameworks converge on the same set of behaviors: deciding what to delegate, describing the task with the right context, judging the output critically, and taking responsibility for the result. Anthropic’s AI Fluency framework, developed with Rick Dakan and Joseph Feller, names them delegation, description, discernment and diligence.4 The judging part is not automatic: a Microsoft Research survey of knowledge workers found that the more confidence people placed in the AI, the less critical-thinking effort they reported, while confidence in their own skills pushed the other way.5 Verification is a behavior to measure, not to assume.

Where the data comes from. Sampled conversation review under an explicit privacy policy, aggregate behavior analytics from tools that provide them, workflow inventories kept by team leads, and self-assessment as the weakest input.

Watch out for. Never measure proficiency by reading individuals’ conversations without consent. Every proficiency metric on the scorecard below can be produced in aggregate.

Level 4: Impact. Is AI making work better?

Usage measures whether AI is being used. Impact measures whether work is getting better, in units the business already tracks.

CategoryExample measures
TimeTask completion time, cycle time, time to first draft, research time
ThroughputAccounts researched, documents analyzed, customer requests completed, reports produced
QualityError rates, rework, completeness, consistency, manager review scores
Business outcomesSales productivity, support resolution time, engineering throughput, campaign production, customer retention, operating cost
Table 2. Impact measures, all of which live in systems the organization already runs.

Attribution is the hard part. Three rules make it tractable. Measure against a baseline taken before the rollout. Compare adopters with non-adopters in the same role over the same period, rather than the whole company against last year. And separate modeled estimates from measured results, labeling each.

Expect the gains to be uneven. In the best-known field study, giving customer support agents an AI assistant raised issues resolved per hour by 14 percent on average, with a 34 percent gain for novice and low-skilled agents and little change for experienced ones.6 That distribution, large for some people and negligible for others, is why impact has to be measured per workflow and per population, not as one company-wide number.

Do not substitute self-reported time savings for measured ones. In a 2025 randomized trial, experienced open-source developers took 19 percent longer on tasks when allowed to use AI tools, despite having forecast that the tools would make them faster.7 People are poor judges of their own speed with AI, in both directions.

Where the data comes from. Time studies for named workflows, and the CRM, ticketing, content and finance systems that already record throughput and quality.

Watch out for. An impact number without a named workflow behind it is a modeled estimate. Report it as one.

The AI adoption scorecard

Treat this scorecard as your AI adoption KPIs: one metric per question, reported on the same cadence, grouped by level so a strong number at one level cannot hide a weak one below it. Every row can be produced without reading anyone’s conversations.

Level 1: Access

Employees with approved AI access
Calculate
Employees with an active license or SSO entitlement ÷ headcount
Source
Identity provider, license admin consoles
Cadence
Monthly
Good trend
Rises to plan, then stays flat

Level 2: Adoption

Weekly active AI users
Calculate
Employees with at least one real interaction in the week ÷ employees with access
Source
Vendor admin reports, browser or network telemetry
Cadence
Weekly
Good trend
Rises, then holds; the gap to monthly active narrows
Usage frequency
Calculate
Median active days per user per month
Source
Vendor admin reports, telemetry
Cadence
Monthly
Good trend
The median climbs, not just the top decile
Usage by department and tool
Calculate
Active users ÷ headcount, per department and per tool
Source
Vendor admin reports joined to the HR system
Cadence
Monthly
Good trend
Lagging departments close the gap

Level 3: Proficiency

AI use cases by role
Calculate
Share of activity in each use-case category, per role
Source
Sampled conversation review under a privacy policy, or aggregate analytics
Cadence
Quarterly
Good trend
Shifts from lookup toward analysis, drafting and repeatable work
Repeatable AI workflows
Calculate
Documented workflows in regular use, per team
Source
Team leads, a workflow inventory
Cadence
Quarterly
Good trend
The count grows and every workflow has an owner
Workflow sophistication
Calculate
Share of usage that brings company context or spans multiple steps
Source
Sampled review or aggregate analytics
Cadence
Quarterly
Good trend
The share rises across roles, not only among early adopters

Level 4: Impact

Time saved
Calculate
Measured task time against the baseline, for named workflows. Not survey estimates.
Source
Time studies, systems of record
Cadence
Quarterly
Good trend
Measured task time for each named workflow falls against its baseline
Output and throughput
Calculate
Units completed per person, for named workflows
Source
CRM, ticketing, content and finance systems
Cadence
Quarterly
Good trend
Rises without a drop in quality
Quality
Calculate
Error rate, rework and review scores, for named workflows
Source
QA sampling, manager review
Cadence
Quarterly
Good trend
Stable or improving while throughput rises
Table 3. An AI adoption scorecard: one metric per question, grouped by level so a strong number at one level cannot hide a weak one below it.

Ten rows is enough. If a metric cannot be tied to one of the four questions, it belongs in an appendix, not on the scorecard.

How to report AI adoption: the same organization, three ways

Here is one sales organization reported three ways. Each step adds a population, a use case, and finally a change in the work itself. The figures are illustrative.

  1. Weak

    85 percent of employees used ChatGPT this month.

    One tool, one month, and no sense of what the usage was for or whether it recurred.

  2. Better

    72 percent of the sales organization uses AI weekly, primarily for account research, meeting preparation and follow-up.

    A stricter window, a specific population and the use cases behind the number.

  3. Best

    Sales standardized three AI-assisted workflows: account research, call preparation and follow-up. Weekly use rose from 41 to 72 percent in one quarter, and median call-preparation time fell from 45 minutes to 20.

    Named workflows, a measured change in adoption and a measured change in the work itself.

The goal is not more AI activity. The goal is better work, and a report that cannot show the work is not a report on adoption.

Pitfalls when measuring AI adoption

  1. Counting logins as usage. Opening a tool is not an interaction. Every vendor console can distinguish the two; make sure your definition of “active” does.
  2. Optimizing for the adoption rate. It is the easiest AI metric to move, which is exactly the problem. Make the tool the default browser tab, run a launch campaign, add a reminder to the weekly all-hands, and weekly active users climb without anyone changing how they work. An organization can reach a high weekly active rate while most of that usage is quick lookups, one-line questions and drafts that get rewritten by hand. Before celebrating the rate, be able to say what work AI is being used for, how sophisticated those workflows are, whether effective practices are spreading, where people are struggling, and which workflows create value.
  3. Trusting self-reported time saved. Survey answers about productivity are beliefs, not measurements, and the developer trial above shows how far the two can diverge.7 Use self-reports to decide where to run a time study, never as the result.
  4. Measuring one vendor’s dashboard. Most employees who use AI at work brought at least one tool themselves.1 A Copilot admin report describes Copilot, not adoption. Measure across every tool in use, including the ones nobody rolled out.
  5. Skipping the baseline. Without 30 days of pre-rollout data on access, usage and the tasks people were already doing, every later number is a level, not a change.
  6. Measuring individuals instead of aggregates. Ranking people by prompt count damages trust and rewards volume. Report by team, role and workflow, and keep individual conversations private by policy, not by promise.

Running the measurement

Measurement only pays off if it feeds a loop: see what work people do today, find where AI could improve it, learn who is already using AI well, spread what they do, and check whether the organization improved. Four steps make that loop run.

1. Baseline before the rollout

Take 30 days of data before anything changes: who has access, who uses which tools and how often, and a short survey of the tasks each role spends the most time on. For the sales team in the examples above, the baseline would be 41 percent weekly active and a median of 45 minutes to prepare each discovery call.

2. Instrument each level

Wire the scorecard to its sources: the identity provider and license consoles for access, vendor admin reports and telemetry for adoption, sampled review or aggregate analytics for proficiency, and the systems of record for impact. Agree the privacy policy for the proficiency sample before collecting a single conversation.

3. Report on a cadence

A weekly adoption pulse for the rollout team. A monthly adoption and proficiency report for department heads, by team and by tool. A quarterly impact review per named workflow for leadership, with the baseline alongside the current figure. Each audience gets the level it can act on.

4. Intervene, then re-baseline

Every intervention, whether a documented workflow, training, or a change to which tools are approved, gets its own before-and-after. Once that sales team documented the call-preparation workflow and trained on it, weekly active rose from 41 to 72 percent and median preparation time fell from 45 minutes to 20 in one quarter. The next quarter’s baseline is 72 and 20, not 41 and 45.

Questions about measuring AI adoption

What is AI adoption?

AI adoption is the degree to which employees have built AI tools into how they do their work. Narrowly it means people use the tools they have access to; usefully it also covers whether they use AI repeatedly, for meaningful tasks, with effective workflows, and whether that use changes how the work gets done.

How do you calculate AI adoption rate?

AI adoption rate is active AI users divided by employees with AI access over a fixed period. If 600 of 1,000 licensed employees used an approved tool in 30 days, the rate is 60 percent. Define “active” as a real interaction rather than a login, and report the rate by department and by tool.

What is a good AI adoption rate?

There is no universal benchmark. Adoption varies by role, tool and how long the rollout has been running, so compare against your own baseline and set targets per team relative to access. A weekly active rate that keeps rising for three months, with usage deepening at the same time, is a better signal than any single threshold.

What is the difference between AI usage and AI proficiency?

Usage is how often someone interacts with AI. Proficiency is how well: whether they bring the right context, break complex tasks down, iterate on outputs, verify results and apply AI to valuable, repeatable workflows. A heavy user can be a weak practitioner, and a light user can have fully absorbed AI into one high-value process.

Can you measure AI adoption without reading employees’ conversations?

Yes. Access and adoption come from license and usage counts. Proficiency can be measured in aggregate: the mix of use-case categories, the share of workflows that carry company context and the number of documented, repeatable workflows per team. Impact comes from task time, throughput and quality in your systems of record, none of which require reading a prompt.

How do you measure ROI from employee AI adoption?

Work backwards from specific workflows. Measure time, throughput or quality for each workflow against a baseline, attribute the change to the population that adopted it, and cost it. Treat modeled time savings as a starting signal only, validate them against cycle times and rework rates in your own systems, and report ROI by workflow and by team rather than as one company-wide figure.

Sources

  1. Microsoft and LinkedIn, “2024 Work Trend Index Annual Report: AI at Work Is Here. Now Comes the Hard Part,” May 2024. https://www.microsoft.com/en-us/worklab/work-trend-index/ai-at-work-is-here-now-comes-the-hard-part
  2. MIT NANDA, “The GenAI Divide: State of AI in Business 2025,” preliminary findings, July 2025. https://nanda.media.mit.edu/ai_report_2025.pdf
  3. Dell’Acqua et al., “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality,” Harvard Business School Working Paper 24-013, 2023. https://www.hbs.edu/faculty/Pages/item.aspx?num=64700
  4. Anthropic, with Rick Dakan and Joseph Feller, “AI Fluency: Framework and Foundations,” Anthropic Academy. https://www.anthropic.com/ai-fluency
  5. Lee, Sarkar et al., “The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers,” CHI 2025. https://dl.acm.org/doi/10.1145/3706598.3713778
  6. Brynjolfsson, Li and Raymond, “Generative AI at Work,” Quarterly Journal of Economics 140(2), 2025; NBER Working Paper 31161, 2023. https://www.nber.org/papers/w31161
  7. Becker, Rush, Barnes and Rein (METR), “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” July 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study

Related reading