AI use and occupationsFatih Kansoy ↗
Occupation observations April–May 2026Five observation windows Definitions
Research & methods

Improving AI-use measurement

Constructive proposals for regional measurement, external validation and research access.

Platform records provide evidence about activities people bring to AI systems. Occupational research also needs to know what was published, how categories were assigned, whether those categories travel across languages, and how activity relates to workers and outcomes. The contributions below distinguish implemented methods from questions the public data cannot yet settle.

Publication coverage before country rankings

We publish original occupation shares, task counts, retained mass and missing-row status. The near-perfect relationship between country volume and published task counts shows why visible breadth alone is an unreliable adoption ranking.

Further progress requires privacy-safe sample bands, reason codes for absent cells, classifier coverage and release-specific disclosure rules. These would help distinguish thin evidence from low measured activity. The present public aggregates do not recover hidden activity or supply sampling confidence intervals.

Source-preserving regional allocations

The allocation method splits each published source share across documented destination links, with weights summing to one. It preserves unmatched activity, retains the original country denominator, compares stricter link scenarios and reports marginal mapping bounds. This enables descriptive international and UK occupation shares and employment concentration ratios. Donor means remain a separate diagnostic.

Conservation is a necessary arithmetic condition, not proof of semantic validity. Regional occupational specialists can assess ambiguous links, local task descriptions and the meaning of equal splits. Such evidence should refine weights independently of the employment or wage outcomes later used to evaluate them. A crosswalk relation is not itself an allocation probability.

Validate classification across languages and work contexts

The released Multilingual occupation reference contains official ESCO labels and available descriptions for 3,010 occupations across 28 languages. It provides common identifiers for language comparisons and sample design. The evaluation protocol defines label-level scoring and a separate task-description evaluation.

Platform-classifier performance has not been evaluated here. That needs consented or licensed descriptions, independently judged occupational relevance, repeated classification under fixed settings and disagreement reporting. Labels alone cannot establish performance on real conversations. A useful regional partnership would bring occupational experts and representative local examples into that assessment, including examples with no work context or no valid category.

Build comparability across observation windows

The April–May occupation comparison applies fixed weights to common countries and occupations. It retains each month's original denominator and publication support. This is a pair of comparable constructions, not evidence of adoption growth: country user composition, classification and disclosure may still change.

Earlier weekly task snapshots remain separate. A longer diffusion series needs consistent sampling and definitions, and ideally old and new procedures run on overlapping samples to measure methodological breaks. A comparable release protocol would be more informative than joining all available dates into a line.

Join activity to outcomes with credible timing

Prepared UK usage, employment and pay panel, UK usage and recruitment-advert panel and European usage and employment panel panels supply native occupation codes, fixed activity measures, dates, populations and source warnings. They reduce the work needed to start an outcome study; they do not contain an estimated AI effect.

The historical annual outcomes precede measured activity. July 2026 is the only full advert month after the public predictor release and is affected by a source-quality notice. No clean forward-test month is available in this snapshot. A future forecast design should freeze the predictor before outcomes arrive, compare against capability measures and occupation-specific trends, evaluate on held-out outcomes, and report sensitivity to sector shocks and data breaks. Adverts are not successful hires.

Representative worker and firm surveys are still needed to measure who uses AI, how often, on which tasks and with which systems. Linking survey adoption to administrative outcomes would address workforce diffusion and economic effects more directly than a consumer conversation share.

Make exposure reconstruction testable

The independent public-input reconstruction reaches a Spearman correlation of 0.89465, below the declared 0.95 reconstruction threshold, and shares six rather than the required eight occupations with the published top ten. The allocation extension does not repair that separate failed exposure benchmark.

Reference code, privacy-safe intermediate aggregates, task-grouping decisions and independently evaluated time weights would help locate disagreement. The failed local time-weight experiment remains excluded from substantive findings; the [replacement protocol](/ai-use-and-occupations/publications/time-weight rerun protocol) is prospective. Survey task frequency cannot substitute for the fraction of working time spent on a task.

Make the evidence easier to examine and reuse

The site supplies original observations, allocations, explicit mapping assumptions, accounting checks, source flags, field definitions, accessible charts and complete downloads. It records observation dates separately from release dates. This lets researchers compare a finding under another mapping or connect a documented predictor to suitable outcomes.

Research access benefits from stable identifiers, clear component-specific reuse terms, machine-readable provenance and preserved snapshots. Uncertainty should remain separated into publication gaps, mapping ambiguity, classifier error, sampling variation and outcome identification. Resolving one does not resolve the others.