AI use and occupationsFatih Kansoy ↗
Occupation observations April–May 2026Five observation windows Definitions
Full publication

Measuring AI use across occupations: publication, classification and validation

Direct occupation shares, publication support and a global data inventory.

Download full PDFRead & copy citation

Fatih Kansoy
9 September 2026

Abstract

Public records of generative-AI activity offer direct evidence about use, but their interpretation as occupational statistics depends on the platform population, disclosure process and classification system. This paper constructs an auditable account of published consumer activity across 121 countries and areas in May 2026. Europe serves as a detailed classification case because the source O*NET categories must be related to ESCO, ISCO and UK SOC statistics; the United States supplies a native SOC benchmark before that crosswalk is introduced. The primary outcome is the direct published occupation share, while fixed-catalogue task coverage is treated as a secondary measure of publication breadth. Across the 121 May geographies, positive task-cell counts correlate 0.996 with global usage share, and median published mass is 42.25% in the task facet versus 79.17% in the occupation facet. European mapping alternatives materially alter occupational descriptions. In the native US comparison, direct occupation shares correlate 0.6284 with a published exposure measure, while an independent public-input exposure proxy reaches 0.89465 and fails its prespecified reconstruction gate. Comparisons with theoretical exposure scores and PIAAC work activities provide partial construct evidence rather than validation of workforce adoption or time saved. The results support a measurement framework that keeps platform activity, publication support, occupational mapping and labour-market outcomes analytically distinct.

Keywords: generative AI; occupational statistics; platform use; classification crosswalks; publication support; O*NET; ESCO; ISCO; PIAAC

1. Research question and contribution

Generative-AI systems are used for activities associated with work, but a conversation about a task is not a worker, a job or an hour of production. This distinction becomes consequential when platform records are joined to occupational statistics. The set of categories visible for a geography may depend on how much platform activity is observed and released, while a crosswalk can change an occupational distribution even when the underlying conversations remain fixed.

This paper asks which occupational descriptions of published platform use are supported by the public data, how publication and classification choices affect those descriptions, and where official employment and independent activity surveys aid interpretation. The evidence is global: 121 countries and areas have direct occupation observations in May 2026. Europe is the principal classification case because the source O*NET categories must pass through ESCO, ISCO or UK SOC to be read beside regional labour statistics. The United States provides a native benchmark because source occupations can be joined to detailed SOC employment codes without the European bridge.

The contribution has three parts. First, the study separates country usage volume, relative per-capita platform use, direct occupation shares, task coverage and task-share intensity, and quantifies their distinct publication support. Second, it exposes typed occupational mappings and alternative routes while preserving source categories and native US comparisons. Third, it places the resulting measures beside a published observed-exposure score, theoretical capability measures and PIAAC work-activity profiles, with common samples and denominators reported for each comparison.

The intended contribution is methodological: evidence about how occupational measures are constructed before those measures enter studies of diffusion or labour-market effects. The observations describe classified consumer conversations. They do not identify users' jobs, provide a representative workforce sample, measure hours automated or estimate causal changes in pay, productivity or employment.

Handa et al. (2025) establish the task-based analysis of Claude conversations underlying the Economic Index. Chatterji et al. (2025) study consumer ChatGPT use and distinguish work from non-work activity. Tomlinson et al. (2025) use Copilot interactions to study occupational applicability. These studies demonstrate the value of observed interactions while also showing why tasks represented in platform use must be distinguished from users' occupational identities.

Appel et al. (2025) describe uneven country and enterprise use and introduce an index relative to working-age population. The June 2026 Economic Index release adds monthly country usage shares, that relative per-capita index and occupation as well as task classifications. The present analysis uses those published country fields directly. A count of released task cells is examined as a publication diagnostic rather than substituted for the publisher's volume or per-capita series.

Capability-based measures address a different object. Eloundou et al. (2023) assess whether language models can reduce the time required for tasks. Felten, Raj and Seamans (2021, 2023) relate AI capabilities to occupational abilities; the ILO (2025) and OECD measures likewise describe technological exposure. Correlation with these scores can inform construct interpretation, but does not validate observed worker adoption.

Studies of realised outcomes require further design. Brynjolfsson, Li and Raymond (2025) analyse productivity in a workplace deployment. Humlum and Vestergaard (2025) connect adoption evidence to Danish labour-market outcomes, while Bick, Blandin and Deming (2024, revised 2025) measure adoption in population surveys. These studies have an identified population, measures of actual use and an outcome design—ingredients that public conversation shares do not supply.

Fan and Nguyen (2026) combine usage evidence with labour costs, illustrating why occupation and country denominators matter for valuation. Steele and Cruz (2026) compare career-related AI projections and examine disagreement across models. The present paper addresses an earlier measurement problem: how much occupational information survives publication and mapping, and how closely the resulting descriptors align with other constructs.

3. Data, scope and observation windows

The primary source is the Anthropic Economic Index release of 26 June 2026, pinned to revision 2ea58ff75e4247d26810c37f10c179edc2466cac. Its consumer country data include one-month windows from 1 April to 1 May 2026 and from 1 May to 1 June 2026, with the end date exclusive. The April window contains 114 countries and areas with direct occupation observations; May contains 121.

Three earlier country samples provide task evidence: 4–11 August 2025, 13–20 November 2025 and 5–12 February 2026. They contain 113, 116 and 117 task geographies, respectively. Sampling, classifiers and taxonomy changed, so these weekly observations are retained as separate diagnostics rather than joined to the monthly occupation evidence as a continuous series.

The full registry contains 250 entries: 248 countries and areas extracted from UNSD M49, plus Kosovo and Taiwan added explicitly as source geographies. Across the four pinned consumer releases, 180 recognised geographic identifiers have at least one country-level row and 128 have a positive published task cell in an archived window. The 581 task country-period pairs are observations, not countries. Official employment series are selected for 182 registry entries, of which 113 intersect May occupation evidence and 106 intersect April evidence.

The May roster comprises 118 UN member states, the State of Palestine, Puerto Rico and Taiwan. These categories describe statistical coverage and do not adjudicate political status. The global inventory preserves source geography labels and makes every registry entry discoverable.

The fixed task catalogue is O*NET 30.2. Employment is selected from Eurostat, ILOSTAT and UK APS at the finest eligible classification detail, with the actual year, population, age scope and ICLS definition retained. The US uses the BLS National Employment Matrix. Restricted comparator vectors, restricted mapping derivatives and respondent-level PIAAC records are not redistributed.

4. Measurement framework

Let UcmU_{cm} be geography cc's published percentage share of global consumer usage in month mm, and PcP_c its population aged 15–64 in the publisher's reference. The relative per-capita index is conceptually

Acm=Ucm/100Pc/kPk.A_{cm}=\frac{U_{cm}/100}{P_c/\sum_kP_k}.

The published value is used directly. It compares platform usage share with population share; it is not a resident adoption rate.

For a published detailed occupation row oo, the direct occupation share is

Scmo=pctcmo100.S_{cmo}=\frac{\operatorname{pct}_{cmo}}{100}.

Its denominator is consumer usage in the relevant geography and month. Let Pcm\mathcal{P}_{cm} be the published occupation rows. Retained occupation mass is

Mcm=oPcmScmo.M_{cm}=\sum_{o\in\mathcal{P}_{cm}}S_{cmo}.

This sum precedes mapping and is not renormalized. A published rounded zero remains distinct from an absent row. For mass bookkeeping an absent row adds no published mass, but actual use remains unknown.

For occupation oo, let ToT_o be its fixed catalogue tasks and scmts_{cmt} a positive published task share in fractional units (source percentage divided by 100). Catalogue task coverage is

Ccmo=tTo1 ⁣{t has a positive published row in (c,m)}To,C_{cmo}=\frac{\sum_{t\in T_o}\mathbf{1}\!\left\{t\text{ has a positive published row in }(c,m)\right\}}{|T_o|},

Let Pcmtask\mathcal{P}^{\mathrm{task}}_{cm} be the set of published task rows. Task-share intensity is

Icmo=tToPcmtaskscmtTo.I_{cmo}=\frac{\sum_{t\in T_o\cap\mathcal{P}^{\mathrm{task}}_{cm}}s_{cmt}}{|T_o|}.

Coverage records published breadth; intensity divides released task mass by catalogue size. Tasks may belong to several occupations, so neither statistic can be summed across occupations to recover national usage.

Quantity Denominator Interpretation
UU: global usage share Global consumer usage in the source window Relative geographic volume
AA: per-capita usage index Geography's share of population aged 15–64 Relative platform intensity
SS: direct occupation share Geography's consumer usage Published occupation-tagged mass
CC: task coverage Tasks in a fixed occupation catalogue Breadth of positive publication
II: task-share intensity Occupation catalogue task count Released task-share fraction per task
Mapped SS Available semantic donors Nonadditive taxonomy descriptor

5. Publication support and missingness

The May task sample contains 121 geographies. Seventy publish fewer than 100 positive task cells and 85 publish fewer than 200. Country-level overall fields are not fabricated where absent. The source's two-decimal precision also matters: each numeric row contributes an arithmetic rounding range of up to 0.005 percentage points, which is distinct from a sampling confidence interval.

The source appendix documents hourly sampling, a two-step task classifier, random selection among the highest-confidence candidate tasks and examples of classification error. An occupational tag therefore describes the activity assigned to a conversation. Even complete publication would not turn it into verified evidence of the user's job.

Reporting grades summarize positive task breadth, retained occupation mass, repeat monthly observation, employment recency, classification detail and matched employment. Applying the stated rules across the 250-entry registry yields 29 grade A, 11 grade B, 81 grade C and 129 grade D entries. These are post-review reporting rules with practical thresholds, not preregistered sample selection or validated measures of representativeness.

The 11 grade-A entries designated as the European core are Belgium, France, Germany, Italy, the Netherlands, Poland, Portugal, Spain, Switzerland, Türkiye and the United Kingdom. Some non-European geographies also receive grade A. The core defines where the paper gives European mapping results substantive emphasis; it does not define the global data universe.

Grade Registry entries Observable reporting rule
A 29 At least 200 positive tasks; occupation mass at least 90%; both months; employment from 2020 or later; at least two-digit detail; at least 80% matched employment
B 11 At least 100 positive tasks; occupation mass at least 70%; remaining conditions as for A
C 81 May occupation evidence, with at least one A/B condition unmet
D 129 No May detailed occupation facet

Table 1. Reporting support across the 250-entry registry. Source: Geography support (download).

6. Published task breadth and platform volume

Across the 121 May geographies, the Spearman correlation between positive published task-cell counts and global consumer usage share is 0.9960. A regression of log task-cell counts on log usage share gives a slope of 1.034 and R2=0.935R^2=0.935. The corresponding April association is similarly strong.

The regression includes an intercept and uses positive, nonmissing source shares and task counts. HC3 standard errors describe the fitted relationship, although their usual repeated-sampling interpretation is limited because source availability does not constitute a random sample of countries.

The near-unit elasticity demonstrates that task-cell breadth is entangled with representation on the platform and in the release. It does not separately identify conversation volume, disclosure thresholds, classifier coverage and true task composition. The association is descriptive and cannot establish a causal publication mechanism.

Window Geographies Spearman Log-log slope HC3 SE R2R^2
April 2026 114 0.992 1.121 0.051 0.937
May 2026 121 0.996 1.034 0.047 0.935

April 2026

Geographies
114
Spearman
0.992
Log-log slope
1.121
HC3 SE
0.051
R2R^2
0.937

May 2026

Geographies
121
Spearman
0.996
Log-log slope
1.034
HC3 SE
0.047
R2R^2
0.935

Table 2. Natural-log OLS with an intercept. Source: Volume regression (download).

7. Retained mass across facets

The median May sum of published task shares is 42.25% of geography-level consumer usage, compared with 79.17% in the direct occupation facet. Both medians use the same 121 geographies and are calculated before any occupation crosswalk or employment weighting. The contrast motivates the choice of direct occupation shares as the primary descriptor.

In the United Kingdom, 437 detailed occupations have positive direct shares, while 324 catalogue occupations have at least one positive task. The additional occupation detail does not remove platform selection, identify users' jobs or recover unpublished tasks. A coarser facet can retain more mass while conveying less about within-occupation task breadth.

Geography Positive tasks Task mass (%) Occupation mass (%) Positive occupations
United States 1,104 92.44 98.29 515
United Kingdom 680 84.96 97.67 437
Netherlands 353 77.18 95.64 307
Türkiye 290 71.69 94.61 281
Median of 121 63 42.25 79.17 109

United States

Positive tasks
1,104
Task mass (%)
92.44
Occupation mass (%)
98.29
Positive occupations
515

United Kingdom

Positive tasks
680
Task mass (%)
84.96
Occupation mass (%)
97.67
Positive occupations
437

Netherlands

Positive tasks
353
Task mass (%)
77.18
Occupation mass (%)
95.64
Positive occupations
307

Türkiye

Positive tasks
290
Task mass (%)
71.69
Occupation mass (%)
94.61
Positive occupations
281

Median of 121

Positive tasks
63
Task mass (%)
42.25
Occupation mass (%)
79.17
Positive occupations
109

Table 3. May 2026, original geography consumer denominators. Source: Facet comparison (download).

8. Why Europe and the United States play different roles

The source occupation categories are American O*NET categories regardless of conversation geography. European employment tables generally use ISCO; ESCO supplies a typed semantic bridge, and the United Kingdom uses SOC 2020. Europe therefore provides a tractable audit of what changes when a common source taxonomy is related to multiple official statistical systems. The exercise does not assume that European sources share identical populations or reference periods: Eurostat, ILOSTAT and APS definitions remain source-specific.

The United States answers a complementary question. O*NET children can be aggregated to native six-digit SOC codes and joined directly to BLS employment. This removes the international crosswalk from the comparison, so construct agreement can be assessed before that additional mapping is introduced.

The common leading labels also motivate a classification audit. In May, the United Kingdom, Germany and Türkiye all rank “Document Management Specialists” and “Librarians and Media Collections Specialists” first and second; the United Kingdom and Germany share 12 of their top 15 labels. These labels describe assigned conversation activities. The similarity may reflect common use, user selection, shared task definitions, classifier behaviour or several factors together; determining which requires targeted validation.

The European core also differs in the support supplied by occupation publication and official employment. “Matched” below is the share of the official employment total attached to a mapped score in the displayed route. It is not the share of workers whose tasks have been validated. The UK period is the APS reporting window, while other entries show the selected employment year.

European core Positive tasks Occupation mass (%) Employment period Detail Matched employment (%)
Belgium 226 93.83 2025 2 99.8
France 744 98.19 2025 2 97.9
Germany 626 97.76 2025 2 99.4
Italy 420 96.78 2025 2 99.5
Netherlands 353 95.64 2025 2 99.0
Poland 262 94.18 2025 2 98.8
Portugal 218 92.39 2025 2 99.7
Spain 528 97.46 2025 2 99.8
Switzerland 234 93.63 2025 2 95.6
Türkiye 290 94.61 2025 2 99.4
United Kingdom 680 97.67 2026-03 4 99.6

Belgium

Positive tasks
226
Occupation mass (%)
93.83
Employment period
2025
Detail
2
Matched employment (%)
99.8

France

Positive tasks
744
Occupation mass (%)
98.19
Employment period
2025
Detail
2
Matched employment (%)
97.9

Germany

Positive tasks
626
Occupation mass (%)
97.76
Employment period
2025
Detail
2
Matched employment (%)
99.4

Italy

Positive tasks
420
Occupation mass (%)
96.78
Employment period
2025
Detail
2
Matched employment (%)
99.5

Netherlands

Positive tasks
353
Occupation mass (%)
95.64
Employment period
2025
Detail
2
Matched employment (%)
99.0

Poland

Positive tasks
262
Occupation mass (%)
94.18
Employment period
2025
Detail
2
Matched employment (%)
98.8

Portugal

Positive tasks
218
Occupation mass (%)
92.39
Employment period
2025
Detail
2
Matched employment (%)
99.7

Spain

Positive tasks
528
Occupation mass (%)
97.46
Employment period
2025
Detail
2
Matched employment (%)
99.8

Switzerland

Positive tasks
234
Occupation mass (%)
93.63
Employment period
2025
Detail
2
Matched employment (%)
95.6

Türkiye

Positive tasks
290
Occupation mass (%)
94.61
Employment period
2025
Detail
2
Matched employment (%)
99.4

United Kingdom

Positive tasks
680
Occupation mass (%)
97.67
Employment period
2026-03
Detail
4
Matched employment (%)
99.6

Table 4. May 2026 publication support and selected employment context. Source: Geography support.

9. European mapping sensitivity

The typed ESCO–O*NET library contains 4,253 links, including 498 exact links. Exact, exact-plus-narrow and all-typed donor means remain separate; an untyped 8,627-link alternative supplies a further sensitivity route. No direct ISCO group link is expanded into invented ESCO children.

For route aa and ESCO occupation ee, let De(a)D_e^{(a)} be the available O*NET donors. The mapped descriptor is

Scme(a)=1De(a)oDe(a)Scmopublished.\overline{S}_{cme}^{(a)}=\frac{1}{|D_e^{(a)}|}\sum_{o\in D_e^{(a)}}S^{\mathrm{published}}_{cmo}.

Here ScmopublishedS^{\mathrm{published}}_{cmo} equals the source share when a row exists and zero published mass for an absent row within a supported geography-month. The donor set includes all linked O*NET catalogue occupations with a defined published-mass value, including donors with no published row. Donors outside the available catalogue are excluded. This convention measures published mass; actual use for an unpublished donor remains unknown.

ISCO four-digit values are means over represented ESCO occupations. Coarser groups are direct means over represented ISCO4 groups. These are semantic taxonomy means: they are nonadditive and cannot be divided by employment shares as though they were allocated national conversation shares.

The table compares each alternative with the all-typed route on pairwise common ISCO4 support. Correlations summarize within-geography occupational ranks across the European core. A high coefficient may coexist with missing destination groups or consequential changes in particular occupations.

Descriptor Alternative Geographies Median ρ\rho P10 / P90 Common ISCO4
Occupation share Exact 11 0.822 0.802 / 0.842 262
Occupation share Exact + narrow 11 0.817 0.804 / 0.832 280
Occupation share Untyped 11 0.892 0.889 / 0.900 412
Task coverage Exact 11 0.767 0.743 / 0.812 262
Task coverage Exact + narrow 11 0.759 0.737 / 0.811 280
Task coverage Untyped 11 0.893 0.858 / 0.916 412

Occupation share · Exact

Geographies
11
Median ρ\rho
0.822
P10 / P90
0.802 / 0.842
Common ISCO4
262

Occupation share · Exact + narrow

Geographies
11
Median ρ\rho
0.817
P10 / P90
0.804 / 0.832
Common ISCO4
280

Occupation share · Untyped

Geographies
11
Median ρ\rho
0.892
P10 / P90
0.889 / 0.900
Common ISCO4
412

Task coverage · Exact

Geographies
11
Median ρ\rho
0.767
P10 / P90
0.743 / 0.812
Common ISCO4
262

Task coverage · Exact + narrow

Geographies
11
Median ρ\rho
0.759
P10 / P90
0.737 / 0.811
Common ISCO4
280

Task coverage · Untyped

Geographies
11
Median ρ\rho
0.893
P10 / P90
0.858 / 0.916
Common ISCO4
412

Table 5. May 2026; pairwise common support. Source: Mapping correlation distribution (download).

10. Native US construct and employment comparisons

Native O*NET-SOC children are aligned to six-digit SOC codes. Task measures are averaged within code, while occupation shares are summed at the original integer hundredths of a percentage point before conversion to fractions. The resulting May descriptors are compared with the publisher's March 2026 occupation exposure file.

May US descriptor Common SOC occupations Spearman with March exposure
Published-task coverage 756 0.5869
Task-share intensity 756 0.5814
Direct occupation share 756 0.6284

Table 6. Native SOC comparison. Source: US measure validation and US validation pairs.

The moderate associations show that direct shares, task descriptors and exposure are related but not interchangeable. They differ in period and construction: the exposure measure incorporates work-related activity, global API use, capability and automation gates, and task-time aggregation. Employment weights do not enter these rank correlations.

The BLS National Employment Matrix supplies 831 detailed line items and an official 2025 total of 170,280,800 jobs, including unincorporated self-employment. Exact matching supports 772 groups and 92.53% of the official total. Unmatched BLS aggregate codes are not divided using an invented concordance.

For a native occupation with employment EoE_o, the relative concentration ratio is

Rcmo=ScmoEo/Enational.R_{cmo}=\frac{S_{cmo}}{E_o/E_{\mathrm{national}}}.

The ratio compares a group's share of classified consumer use with its share of jobs. It is not the probability that a worker uses AI. The conditional employment-weighted task coverage on matched groups is 8.977%, an average catalogue statistic rather than a national exposure rate.

11. Exposure reconstruction

The published exposure method combines work-related consumer task counts CtworkC_t^{\mathrm{work}} and global API task counts AtA_t:

Wt=Ctwork+At.W_t=C_t^{\mathrm{work}}+A_t.

For Wt>0W_t>0, and with consumer automation share qtq_t, the automation factor is

αt=0.5+0.5(Ctworkqt+AtWt).\alpha_t=0.5+0.5\left(\frac{C_t^{\mathrm{work}}q_t+A_t}{W_t}\right).

The gated task value is

rt=1{Wt100}1{βt0.5}αt,r_t=\mathbf{1}\{W_t\ge100\}\,\mathbf{1}\{\beta_t\ge0.5\}\,\alpha_t,

with rt=0r_t=0 when Wt=0W_t=0, without evaluating the zero denominator. For occupation-specific task-time weights wotw_{ot}, exposure is

Xo=tTowotrttTowot.X_o=\frac{\sum_{t\in T_o}w_{ot}r_t}{\sum_{t\in T_o}w_{ot}}.

Both sums range over the occupation's full task set. Tasks that fail either gate contribute zero to the numerator but remain in the time-weight denominator; weights are not renormalized over eligible tasks.

An independent public-input proxy reaches Spearman 0.89465 on 756 occupations and shares 6 occupations with the published top ten. It fails the prespecified gate of correlation at least 0.95 and at least 8 common top-ten occupations. Equal-weight aggregation of already-published task scores reaches 0.87154, but that check embeds publisher outcomes and is not an independent input reconstruction.

Exact reconstruction is blocked by unavailable original time fractions, work-share imputation, similar-task grouping, employment allocation and uncensored intermediates. The evidence does not identify time weights as the sole source of disagreement. Failure of the reconstruction benchmark is distinct from the construction of direct published-share descriptors.

12. Theoretical exposure comparators

All-typed ISCO4 descriptors are compared with ILO, OECD, Eloundou and Felten scores on their pairwise common occupational samples. Positive association with some capability measures is consistent with observed platform activity concentrating in tasks that language models can assist. It does not establish a time-exposure fraction or causal employment channel.

The OECD input is the published reversed normalized exposure measure, with higher values denoting greater exposure. Its negative correlations are retained rather than sign-flipped. Differences may reflect capability definitions, mapping and selected platform activity; the present data do not separately identify their contributions.

Comparator Descriptor Geographies Median ρ\rho P10 / P90 Common occupations
ILO Occupation share 11 0.630 0.617 / 0.654 406
ILO Task coverage 11 0.580 0.527 / 0.605 406
OECD Occupation share 11 -0.136 -0.180 / -0.120 411
OECD Task coverage 11 -0.106 -0.197 / -0.051 411
Eloundou GPT-4 beta Occupation share 11 0.666 0.642 / 0.705 412
Eloundou GPT-4 beta Task coverage 11 0.608 0.559 / 0.655 412
Felten language Occupation share 11 0.701 0.675 / 0.746 411
Felten language Task coverage 11 0.634 0.568 / 0.691 411

ILO · Occupation share

Geographies
11
Median ρ\rho
0.630
P10 / P90
0.617 / 0.654
Common occupations
406

ILO · Task coverage

Geographies
11
Median ρ\rho
0.580
P10 / P90
0.527 / 0.605
Common occupations
406

OECD · Occupation share

Geographies
11
Median ρ\rho
-0.136
P10 / P90
-0.180 / -0.120
Common occupations
411

OECD · Task coverage

Geographies
11
Median ρ\rho
-0.106
P10 / P90
-0.197 / -0.051
Common occupations
411

Eloundou GPT-4 beta · Occupation share

Geographies
11
Median ρ\rho
0.666
P10 / P90
0.642 / 0.705
Common occupations
412

Eloundou GPT-4 beta · Task coverage

Geographies
11
Median ρ\rho
0.608
P10 / P90
0.559 / 0.655
Common occupations
412

Felten language · Occupation share

Geographies
11
Median ρ\rho
0.701
P10 / P90
0.675 / 0.746
Common occupations
411

Felten language · Task coverage

Geographies
11
Median ρ\rho
0.634
P10 / P90
0.568 / 0.691
Common occupations
411

Table 7. May 2026, European core, all-typed mapping. Source: Comparator distribution (download).

13. PIAAC work-activity evidence

PIAAC provides independent survey evidence on the frequency of broad work activities. A frozen bridge relates survey items to seven O*NET activity domains. The primary specification compares, across occupations within a survey geography and domain, the survey share reporting an activity at least weekly with the corresponding share of catalogue tasks. All typed links are used at ISCO two-digit level.

Every finite coefficient is retained, including those flagged for weak alignment. Country means are unweighted means of coefficients rather than pooled person-level correlations. Each public activity cell requires at least 30 valid respondents, and each rank correlation must meet the protocol's common-occupation requirement. Survey weights enter activity shares; replicate-weight design intervals are not claimed for the rank summaries.

Activity domain Survey geographies Mean ρ\rho Median ρ\rho Occupations per correlation
Analysis 16 0.733 0.740 26–35
Communication 16 0.373 0.390 19–31
Computer work 16 0.479 0.470 19–31
Documentation 16 0.568 0.587 26–35
Information 16 0.472 0.464 18–31
Measurement 16 0.390 0.428 26–35
Numerical processing 16 0.290 0.283 26–35

Analysis

Survey geographies
16
Mean ρ\rho
0.733
Median ρ\rho
0.740
Occupations per correlation
26–35

Communication

Survey geographies
16
Mean ρ\rho
0.373
Median ρ\rho
0.390
Occupations per correlation
19–31

Computer work

Survey geographies
16
Mean ρ\rho
0.479
Median ρ\rho
0.470
Occupations per correlation
19–31

Documentation

Survey geographies
16
Mean ρ\rho
0.568
Median ρ\rho
0.587
Occupations per correlation
26–35

Information

Survey geographies
16
Mean ρ\rho
0.472
Median ρ\rho
0.464
Occupations per correlation
18–31

Measurement

Survey geographies
16
Mean ρ\rho
0.390
Median ρ\rho
0.428
Occupations per correlation
26–35

Numerical processing

Survey geographies
16
Mean ρ\rho
0.290
Median ρ\rho
0.283
Occupations per correlation
26–35

Table 8. ISCO2, at-least-weekly survey share versus catalogue task share. Source: PIAAC across occupation summary (download).

Detailed France and Spain results show that agreement becomes uneven at four-digit level, including near-zero and negative coefficients for some French domains. The number of common occupations varies by country and activity, so a single pooled validation coefficient would conceal relevant variation. These comparisons test partial work-content alignment. Survey frequency is not a time fraction, and catalogue composition is not observed adoption.

Activity domain France ρ\rho France occupations Spain ρ\rho Spain occupations
Analysis 0.630 33 0.500 32
Communication 0.410 27 0.275 14
Computer work 0.358 27 0.433 14
Documentation 0.500 33 0.502 32
Information -0.073 27 0.291 14
Measurement -0.044 33 0.361 32
Numerical processing -0.022 33 0.233 32

Analysis

France ρ\rho
0.630
France occupations
33
Spain ρ\rho
0.500
Spain occupations
32

Communication

France ρ\rho
0.410
France occupations
27
Spain ρ\rho
0.275
Spain occupations
14

Computer work

France ρ\rho
0.358
France occupations
27
Spain ρ\rho
0.433
Spain occupations
14

Documentation

France ρ\rho
0.500
France occupations
33
Spain ρ\rho
0.502
Spain occupations
32

Information

France ρ\rho
-0.073
France occupations
27
Spain ρ\rho
0.291
Spain occupations
14

Measurement

France ρ\rho
-0.044
France occupations
33
Spain ρ\rho
0.361
Spain occupations
32

Numerical processing

France ρ\rho
-0.022
France occupations
33
Spain ρ\rho
0.233
Spain occupations
32

Table 9. ISCO4, all-typed mapping, at-least-weekly survey share versus catalogue task share. Source: PIAAC detailed country results (download).

14. Time weights and prospective validation

A frozen local-model experiment generated implausible task-hour values: one task received two quadrillion weekly hours and 12 occupation totals exceeded 168 hours. Structural validity, exact identifiers and normalized shares cannot make those quantities credible. The weights and their exposure correlations remain in the audit record but are excluded from substantive findings. They are not clipped, selectively regenerated or chosen according to agreement with the target.

The prospective replacement protocol specifies three blinded independent draws, a 40-hour analytical allocation that includes uncovered activity, exact task identifiers, finite nonnegative entries, sum constraints and whole-draw rejection rules. A mean would be defined only for occupations with three valid draws, with missingness and between-draw variation reported.

Before a confirmatory run, the endpoint and settings, independent time-use evidence, validation sample and success criteria should be frozen and publicly registered. PIAAC frequency responses and O*NET importance ratings cannot substitute for observed time allocations. The protocol is a design for future validation, not an executed study.

15. Interpretation and implications

The evidence supports a hierarchy of empirical uses. Global usage share and the published per-capita index describe relative platform scale. Direct occupation shares describe released occupational concentration with substantially greater retained mass than task rows. Task coverage describes breadth in a fixed catalogue and is strongly associated with geography-level platform volume. These measures answer different questions and should remain separate.

For Europe, an exact semantic relation, an equal donor mean and an employment weight are not successive refinements of one quantity; each defines a different operation. Reporting alternative routes, unmatched destinations and common samples prevents arithmetic precision from being mistaken for semantic validity. The native US route shows what can be checked before the international bridge is introduced.

The UK source-share-floor diagnostic illustrates the sensitivity of catalogue breadth to small published cells. On the same ONS coding-index mapping route and APS employment support, conditional employment-weighted May task coverage is 6.0209% at baseline and 1.4994% after a 0.10 percentage-point source-share floor. The difference is sensitivity to publication thresholds, not evidence of changing adoption.

The global inventory remains analytically useful because it retains sparse and heterogeneous evidence under explicit status fields. Reporting grades constrain which comparisons receive emphasis while preserving source observations for other geographies. They should not be read as an international quality league table.

The current data can support occupation-focused descriptive research and help formulate worker or firm studies. Estimating employment effects would require outcome data, timing and a credible comparison design. Estimating work-time automation would require independently validated time allocation and work-related classification. Employment reweighting, crosswalk selection and source-share floors cannot supply these missing ingredients.

16. Reproducibility and access

Numerical tables are generated from frozen CSV outputs. The source manifest records URLs, retrieval times and hashes; the public catalogue records keys, units, row counts and file hashes. Code separates publication states, taxonomy aggregation, employment joins and validation comparisons.

The public package contains the licensed subset of data and code. Dataset-specific source and redistribution terms remain attached. The technical data note documents schema, construction and file lineage; the data catalogue provides human-readable metadata and downloads; the glossary defines recurring terms.

References

This bibliography is shared by the paper, technical data note and methodological pages.