AI use and occupationsFatih Kansoy ↗
Occupation observations April–May 2026Five observation windows Definitions
Research & methods

Methods

Definitions, denominators, international classifications and publication support.

The analysis separates four stages: the location attached to platform activity, the occupational activity assigned by the source classifier, the information retained in the public release, and the statistical classification used to describe that activity. Each stage changes what can be inferred.

Observation windows and source vintage

The primary source is the Anthropic Economic Index release of 26 June 2026, pinned to revision 2ea58ff75e4247d26810c37f10c179edc2466cac. April and May 2026 country consumer observations supply the direct occupation and overall-usage fields. The fixed task catalogue is O*NET 30.2. Source release and documentation.

A conversation is assigned an occupational activity by the publisher's classifier. The assigned activity is not evidence of the user's occupation. Country evidence concerns the documented consumer surface; the first-party API file has no country dimension and is not allocated to countries here.

Historical task observations cover one week in August 2025, November 2025 and February 2026. The April and May 2026 observations cover calendar months. Changes in sampling and classification mean that the five windows are retained as distinct observations rather than a continuous adoption series.

Country volume and relative use per capita

Let UcmU_{cm} be geography cc's published percentage share of global consumer usage in month mm, and let PcP_c be its population aged 15–64 in the publisher's reference. Conceptually, the relative per-capita index is

Acm=Ucm/100Pc/kPk.A_{cm}=\frac{U_{cm}/100}{P_c/\sum_k P_k}.

The analysis uses the publisher's released index. A value above one denotes more platform use per working-age resident than the reference average. It is not the fraction of residents who use AI.

Direct occupation shares and publication mass

The primary measure selects the detailed soc_occupation facet, hierarchy level 0 and metric pct. For a published row for occupation oo,

Scmo=pctcmo100.S_{cmo}=\frac{\operatorname{pct}_{cmo}}{100}.

Its denominator is the geography's consumer usage in that window. Changing geography does not convert an O*NET category into a local occupational classification.

Let Pcm\mathcal{P}_{cm} be the set of published occupation rows. Retained occupation mass is

Mcm=oPcmScmo.M_{cm}=\sum_{o\in\mathcal{P}_{cm}}S_{cmo}.

The sum is calculated before mapping and is not rescaled to one. The residual may contain unpublished or unclassified activity and rounding; the public tables do not identify the contribution of each source.

Source status Meaning Treatment
Published positive A source row has a positive rounded share Retain the published value
Published rounded zero A row exists and displays zero Preserve zero and its publication status
Not published No source row appears for a catalogue occupation Mark the raw share missing; actual use is unknown

For retained-mass bookkeeping, an unpublished row contributes zero published mass. This is not an imputation of zero actual use. The source's two-decimal rounding permits an arithmetic range of up to 0.005 percentage points per numeric row; that range is not a sampling interval.

Secondary task measures

Let ToT_o be the fixed set of catalogue tasks associated with occupation oo, and let scmts_{cmt} be a task's positive published share in fractional units. Catalogue task coverage is

Ccmo=tTo1 ⁣{t has a positive published row in (c,m)}To.C_{cmo}=\frac{\sum_{t\in T_o}\mathbf{1}\!\left\{t\text{ has a positive published row in }(c,m)\right\}}{|T_o|}.

Let Pcmtask\mathcal{P}^{\mathrm{task}}_{cm} be the set of published task rows. Task-share intensity is

Icmo=tToPcmtaskscmtTo.I_{cmo}=\frac{\sum_{t\in T_o\cap\mathcal{P}^{\mathrm{task}}_{cm}}s_{cmt}}{|T_o|}.

Coverage measures published breadth in a catalogue. Intensity divides released task mass by catalogue size. Because a task may belong to several occupations, neither measure can be added across occupations to recover national usage. Their denominators are neither employment nor working time.

European classification routes

The typed ESCO–O*NET crosswalk contains exact, narrow and broad relations. Alternative routes remain separate. The source library contains 4,253 typed links, including 498 exact links; link counts do not measure the share of employment with an exact correspondence. European Commission crosswalk.

For route aa, let De(a)D_e^{(a)} be the available O*NET donors for ESCO occupation ee. The ESCO descriptor is

Scme(a)=1De(a)oDe(a)Scmopublished.\overline{S}_{cme}^{(a)}=\frac{1}{|D_e^{(a)}|}\sum_{o\in D_e^{(a)}}S^{\mathrm{published}}_{cmo}.

Here ScmopublishedS^{\mathrm{published}}_{cmo} equals the source share when a row exists and zero published mass for an absent row within a supported geography-month. The donor set includes all linked O*NET catalogue occupations with a defined published-mass value, including donors with no published row. Donors outside the available catalogue are excluded. This convention measures published mass; actual use for an unpublished donor remains unknown.

ISCO four-digit descriptors are equal means of represented ESCO occupations. Coarser ISCO groups are direct means over represented four-digit groups. These values are nonadditive taxonomy descriptors. They are not national conversation shares, and dividing them by European employment shares does not produce the native concentration ratio defined below. Alternative routes are sensitivity scenarios, not confidence intervals.

Europe is the audited case because these bridges are required to place source categories beside European and UK statistics. Applying the same bridge to other labour markets would be an additional transfer assumption.

Official employment and the native US benchmark

Employment context preserves the source population, classification, year and quality flags. Eurostat and ILOSTAT commonly provide ISCO groups; UK APS uses SOC 2020. The ONS coding-index route uses lexical weights, which are not observed worker transitions.

For the United States, O*NET children aggregate directly to native six-digit SOC codes: occupation shares are summed and task descriptors are averaged. Shares are summed at their original integer rounding precision before conversion to fractions. This exact code route makes the United States a useful benchmark because it avoids the European bridge.

The BLS National Employment Matrix supplies 2025 base-year employment, including unincorporated self-employment. It counts jobs rather than workers. Its official total is 170,280,800 jobs; 772 matched detailed groups cover 92.53% of that total. The 2035 projections are not outcomes in this study. BLS definitions.

For a native group with employment EoE_o and official national total EnationalE_{\mathrm{national}}, the descriptive concentration ratio is

Rcmo=ScmoEo/Enational.R_{cmo}=\frac{S_{cmo}}{E_o/E_{\mathrm{national}}}.

A value above one means that the group's share of classified consumer use exceeds its share of national jobs. It does not estimate a worker's probability of using AI. Unmatched employment remains in the national denominator and is reported separately.

Prospective mass-preserving allocation

A future European share construction would require weights that allocate each source occupation across destinations, including an unmatched category:

S~cmg(a)=owog(a)Scmo,gG{unmatched}wog(a)=1.\widetilde{S}^{(a)}_{cmg}=\sum_o w^{(a)}_{og}S_{cmo}, \qquad \sum_{g\in\mathcal{G}\cup\{\mathrm{unmatched}\}}w^{(a)}_{og}=1.

This defines a proposed estimand, not a result in the current release. The weights require substantive justification and sensitivity analysis. Employment-based weights can mechanically affect a later usage-to-employment ratio; adding-to-one is necessary for allocation but does not validate the mapping.

Reproducibility

Frozen inputs, unique observation keys, source precision and publication states are retained in the outputs. Comparison tables report common occupational samples. Source attribution and redistribution rules accompany downloads. The technical data note records construction and file-level lineage, while validation distinguishes successful data checks from the unpassed exposure reconstruction.