Published platform records make it possible to study activities that people bring to generative-AI systems. Turning those records into occupational statistics raises a broader set of methodological questions: how disclosure shapes the observed distribution, how classifications travel across languages and labour markets, how observation windows can be compared, and what additional evidence is required before use can be connected to workers or outcomes.
This agenda follows the frozen evidence and validation checks assessed on 9 September 2026. It states research priorities rather than completed findings.
Model the publication process
Across the observed geographies, published task breadth is closely associated with platform volume, while task and occupation facets retain different shares of published mass. An absent detailed row therefore cannot be interpreted as zero use.
A stronger statistical account would document privacy-safe sample bands, retained-mass summaries, publication reason codes and effective sample information for released cells. When exact counts cannot be disclosed, support categories or interval information can still distinguish thin evidence from low measured activity. The present data preserve rounded-zero and absent rows separately, but public aggregates cannot recover the hidden distribution or provide sampling confidence intervals.
Treat classification bridges as estimands
O*NET, ESCO, ISCO and UK SOC describe work through different structures. A semantic link may be exact, narrow or broad; an equal mean over linked occupations is a taxonomy descriptor, not a mass-preserving allocation. Future work should compare alternative bridges under a clearly stated estimand.
Useful evidence would include versioned relations, mapping rationales, unmatched destinations and sensitivity to alternative allocations. Where national usage shares are the target, weights should preserve source mass and include an explicit unmatched destination. Validation must assess the meaning of those weights as well as their arithmetic. The current methods expose equal-donor routes so that their limits can be audited.
Validate classification across languages and work contexts
Occupational classifiers may perform differently across languages, sectors and local descriptions of work. The common leading activity labels observed in the United Kingdom, Germany and Türkiye make this an empirical question; they do not by themselves establish a classifier artefact.
A multilingual evaluation should report confusion patterns, disagreement rates and sensitivity to repeated classification by language, occupation and use context. Licensed or consented task descriptions and carefully designed synthetic cases can test parts of the pipeline, but performance on an evaluation set must remain distinct from performance on platform traffic. Local occupational catalogues provide a further test of whether source categories preserve their meaning after translation.
Build comparability across observation windows
The three earlier task snapshots use one-week windows, while the later occupation observations cover calendar months and follow changed measurement procedures. Joining them into a single line risks attributing a methodological break to diffusion.
Future releases could run old and new procedures on overlapping samples, retain stable taxonomies where feasible and publish machine-readable comparability flags. A time series should separate changes in activity from changes in sampling, classification and disclosure before trends are interpreted.
Make exposure reconstruction testable
Observed-exposure measures combine capability, work-related use, automation and task-time aggregation. The independent public-input reconstruction reported here reaches a Spearman correlation of 0.89465, below the prespecified 0.95 threshold, and shares 6 rather than the required 8 occupations with the published top ten.
Versioned reference code, privacy-safe intermediate aggregates, task-grouping decisions and independently evaluated time weights would allow external researchers to locate disagreement at a particular stage. A small canonical fixture could test each transformation without exposing individual conversations. The failed local time-weight experiment is excluded from substantive findings; the validation record describes the prospective replacement protocol.
Connect use to workforce diffusion
Platform activity, workforce diffusion and economic effects require different denominators and research designs. Representative worker or firm surveys could measure who uses which systems and for what tasks. Repeated observations could then be linked, under documented classifications, to employment, pay, vacancies, productivity or task requirements.
The first useful product would be a harmonized occupation-outcome panel with stable identifiers, documented breaks and explicit population coverage. Studies of hiring or pay would additionally require credible timing, comparison groups, pre-trend analysis and controls for competing sector shocks. The current PIAAC comparison checks partial work-activity content; it does not establish AI adoption or a labour-market effect.
Quantify mapping and measurement uncertainty
Alternative classifications, disclosure thresholds and missing cells are sources of uncertainty, but they are not interchangeable with sampling error. A research programme should report them separately: sampling intervals where a probability model supports them, mapping scenarios where relationships are ambiguous, and partial-identification bounds where unpublished mass constrains an estimand.
This separation would make international comparisons more informative without forcing incompatible evidence into a single ranking. The global inventory of 121 measured geographies supplies the comparative frame; Europe provides an audited classification case; the United States shows what can be learned before a crosswalk is introduced.
Improve research access and preservation
Clear denominators, accessible tables, machine-readable provenance and stable citations are part of measurement infrastructure. Future releases should preserve source and observation dates, document revisions, provide human-readable field names alongside raw variables and deposit citable snapshots when licensing permits. A common glossary and linked bibliography support the same aim.