Claude Occupation
EvidenceBy Fatih Kansoy
fatih.ai ↗
Research & methods

Validation

What the external comparisons establish and what remains unresolved.

Compare the measures

Validation is specific to a claim. Reproducing published arithmetic, comparing two occupational constructs and testing agreement with survey work activities are different exercises. The project reports each with its sample and interpretation.

US comparisons before a European crosswalk

Native US categories allow the country descriptors to be compared with Anthropic's March 2026 observed-exposure score before introducing ESCO or ISCO mapping. Direct occupation shares are summed to native six-digit SOC groups; the task descriptors are averaged within those groups. Source rounding ties are preserved.

May 2026 US descriptor Common occupations Spearman correlation with March exposure
Published-task coverage 756 0.5869
Task-share intensity 756 0.5814
Direct occupation share 756 0.6284

Direct occupation shares align more closely in this comparison, but the coefficients do not establish equivalence. Monthly consumer shares and the March exposure construction differ in period, sample and definition. Exposure incorporates work-related use, capability, automation and time aggregation. These comparisons cannot isolate which difference produces the gap. Anthropic's exposure study.

The exact exposure reconstruction remains unpassed

The separate independent public-input reconstruction achieves a Spearman correlation of 0.89465 across 756 occupations, with 6 occupations shared between the two top tens. The prespecified acceptance criteria were a correlation of at least 0.95 and at least 8 common top-ten occupations. The reconstruction does not pass.

Aggregating Anthropic's already published task scores with equal weights yields 0.87154. That is a downstream diagnostic which embeds publisher outputs; it is not an independent reconstruction.

Original time fractions, work-share imputation, similar-task grouping, employment allocation and uncensored intermediates are not all recoverable. The evidence does not establish that any single missing input explains the disagreement. The direct occupation dataset remains a description of published shares, rather than a claimed replication of the exposure score.

A native employment join with a defined population

The BLS National Employment Matrix contains 831 detailed line items and 170,280,800 base-year jobs in 2025. Exact native matching supports 772 detailed groups and 92.53% of the official total. The extraction is checked against the saved official source table, and unmatched categories remain explicit.

This verifies the statistical join and its denominator. It does not establish that Claude users represent the population in the table. The Matrix counts jobs and includes unincorporated self-employment; that population differs from both individual workers and the narrower OEWS universe. BLS employment definitions.

PIAAC provides a partial content check

The OECD Survey of Adult Skills measures the frequency of work activities. A frozen, explicitly interpretive bridge connects survey items with seven O*NET activity domains. The primary descriptive comparison at ISCO two-digit level ranks occupations by the survey share reporting an activity at least weekly and by the corresponding share of catalogue tasks.

Activity Survey geographies Mean rank correlation Occupations per comparison
Analysis 16 0.733 26–35
Documentation 16 0.568 26–35
Computer work 16 0.479 19–31
Information acquisition 16 0.472 18–31
Measurement 16 0.390 26–35
Communication 16 0.373 19–31
Numerical processing 16 0.290 26–35

Agreement varies across domains and becomes uneven at finer occupational detail. Weak and negative results remain in the outputs. There is no universal validation coefficient.

Each public survey activity cell requires at least 30 valid respondents. Survey weights enter the activity shares, but these rank summaries do not claim replicate-weight confidence intervals. England and Flanders retain their actual survey geography rather than being labelled as full UK and Belgian samples. OECD PIAAC database.

The comparison concerns a subset of work content. Frequency is not time allocation, and catalogue shares are not representative AI adoption. Agreement therefore cannot validate task-time weights or resolve the exposure reconstruction gap.

Mapping and theory comparisons

Mapping alternatives are compared on their pairwise common occupational support. A strong correlation on that subset can coexist with missing destinations or substantial changes for individual occupations. The alternatives are sensitivity checks, not statistical confidence bounds on an unknown national share.

The study also compares occupation descriptors with ILO, OECD, Eloundou and Felten capability measures. Score direction and paired sample sizes are retained, including negative correlations. Theoretical feasibility and observed platform use answer different questions; disagreement is evidence to examine rather than a reason to reverse a source score or select a more favourable sample.

Excluded time-weight experiment

A local model experiment generated task-hour estimates that failed basic plausibility checks. The weights and their correlations are excluded from substantive conclusions. Normalizing or clipping invalid quantities would not validate them.

A prospective replacement protocol requires three blinded independent draws, exact task identifiers, nonnegative finite allocations, explicit total-time constraints and whole-draw rejection rules. Independent time-use evidence and acceptance criteria must be specified before a confirmatory run. That study has not been executed; PIAAC frequencies and O*NET importance scores do not substitute for observed time allocations.

Reading the evidence together

The checks support reproducible descriptions of published occupation activity, visible classification sensitivity and limited comparisons with independent work-activity information. They also identify the additional evidence required for regional exposure estimates and labour-market impact studies.

The full paper reports the detailed tables and literature. The data catalogue provides the comparison files and their underlying support. The measurement agenda connects the unresolved questions to concrete research proposals.