Claude Occupation
EvidenceBy Fatih Kansoy
fatih.ai ↗
Research & methods

Improving AI-use measurement

Constructive proposals for regional measurement, external validation and research access.

The Economic Index has made real platform activity available for economic research. Its public occupation shares, geographic indicators, methodology and downloadable data support questions that capability assessments alone cannot answer. This project builds on that contribution by examining what is needed to make the evidence more useful across regional statistical systems.

The agenda below follows the frozen data and independent checks documented here, assessed on 9 September 2026. Proposals are distinct from completed results. They do not assume access to private conversations or an existing partnership.

Make publication support a first-class data product

What the evidence shows. Published task breadth is strongly associated with country platform volume, and the retained mass differs considerably between task and occupation facets. An absent detailed row cannot be assumed to mean no use.

What would help. Publish privacy-safe sample bands, retained-mass summaries, publication reason codes and information about the effective sample behind released cells. Where a direct estimate is not appropriate, a documented support category can still help researchers distinguish thin evidence from genuinely low activity.

What this project contributes. The explorer reports the observation period, publication status and retained mass alongside occupation shares. The data retain rounded-zero and absent rows separately. These diagnostics make publication support visible; they do not recover the hidden distribution or provide sampling confidence intervals. The relevant source definitions are in the June release.

Connect occupational classifications without hiding assumptions

What the evidence shows. O*NET is a useful common starting point, but European and UK statistics use different classifications. Semantic matches can be exact, narrow or broad. An average over linked occupation shares is a descriptor; it is not a mass-preserving allocation.

What would help. Maintain versioned bridges with relation types, mapping rationales, unmatched destinations and alternative allocations. Evaluate the links with regional occupational experts and actual descriptions of local work. Publish the result of disagreement rather than forcing every source occupation into one destination.

What this project contributes. Audited mapping routes and sensitivity files expose those choices. The proposed next method would preserve source mass under explicit allocation weights and carry mapping uncertainty into regional comparisons. Validation would need to establish the adequacy of the allocations, not merely that their arithmetic adds up. Current crosswalk methods.

Test classification across languages and work contexts

What is already documented. Anthropic describes its classification approach and provides examples of classification error. Its independent-research programme also discusses how prompt wording and the difference between public test conversations and platform traffic affect classification. Independent research programme.

What would help. Report validation by language, occupation and use context, with confusion patterns, disagreement rates and repeated-classification sensitivity. A shared multilingual evaluation set could use licensed or consented examples of tasks and carefully designed synthetic cases; agreement on such cases would still need to be distinguished from performance on actual platform traffic.

A regional collaboration proposal. Occupational-classification specialists, labour economists and language experts could jointly review where local job descriptions diverge from US catalogue tasks. The output would be an evaluation protocol and documented correction set. No such partner agreement or completed multilingual validation is claimed here.

Create bridges between observation windows

What the evidence shows. Earlier task snapshots and later monthly occupation observations differ in their windows and measurement procedures. A line joining them can appear to measure diffusion while also reflecting changes in observation.

What would help. Run overlapping samples under both old and new classification procedures, retain stable taxonomies where feasible, and distribute machine-readable comparability flags. Separate genuine changes in use from changes caused by the measurement system before interpreting a trend.

What this project contributes. Historical weekly diagnostics and calendar-month evidence remain separate. Source dates, release dates and publication support are visible. The site does not present a continuous regional adoption series where the inputs cannot sustain one.

Make exposure reconstruction independently testable

What is already published. Anthropic's observed-exposure framework combines capability, work-related use, automation and task-time aggregation. A methodological formula is available. Exposure research and appendix.

What the independent check finds. Our public-input reconstruction reaches a rank correlation of 0.895, below the prespecified 0.95 threshold. Several original intermediates cannot be recovered; missing time weights alone do not explain the entire gap.

What would help. Supply versioned reference code, privacy-safe intermediate aggregates, task grouping and allocation decisions, and independently evaluated time weights. A small canonical test fixture would let external researchers verify each transformation without access to individual conversations.

What this project contributes. The achieved checks and the unpassed gate are reported separately. The failed local time-weight experiment is excluded from substantive results; its replacement remains a prospective protocol. Validation record.

Connect use, diffusion and impact through independent evidence

Conversation data describe activity on a product. Workforce diffusion requires a population of workers or firms and an appropriate sampling design. Economic impact additionally requires outcomes, timing and a credible comparison. These are related but separate research tasks.

A regional research programme could connect occupational descriptions to repeated worker and firm surveys, national employment and earnings statistics, and task requirements in job advertisements. The most useful first deliverable would be a harmonized occupation-outcome panel with auditable classification links and clearly documented breaks. Tests of hiring or pay responses would follow a stated empirical design, including pre-trends and competing sector shocks.

Potential collaborators include national statistical institutes, public employment services, universities, occupational standards bodies and licensed job-data providers. These are proposed roles, not an announced consortium. The present PIAAC comparison checks partial work-activity content; it does not establish AI adoption or a labour-market effect.

Broaden access to the evidence

Clear denominators, accessible tables and reproducible views are part of measurement infrastructure. They make it easier for a policymaker to understand a result, a journalist to check its scope and a researcher to reuse the underlying observations.

This site provides country-specific entry points, source-aware charts, full-text research, methods with equations and downloadable data. Further useful work includes teaching examples, multilingual explanations reviewed by domain specialists, documented contribution routes and an archival identifier after a suitable deposit is completed.

Anthropic's August 2026 pilot demonstrates one route to independent aggregate-data research and invites expressions of interest for future studies. Access is not guaranteed, and the pilot does not give researchers raw conversations. A collaboration could extend the questions studied here while preserving independent analysis and privacy. Programme announcement.