# Multilingual occupation evaluation protocol

Status: evaluation design and reference preparation; no model or platform-classifier accuracy result has been produced.

The reference file contains official ESCO v1.1.2 occupation identifiers, their ISCO4 parent, and available labels/descriptions in 28 languages. Common URIs create paired language examples. A label translation is not a human-validated description of work in a particular country.

## Label-retrieval benchmark

Freeze a model endpoint, version, prompt, retrieval corpus, language set, seed and settings before scoring. Reserve an occupation-disjoint held-out set so translations of a test occupation never appear in training or worked examples. Report how many source occupations and language pairs remain after missing descriptions are excluded. Stratify by ISCO major group and report both macro averages and the full language/group table.

Predict an ISCO4 code or abstain. Evaluate exact ISCO4 accuracy, ISCO2 accuracy, abstention rate and paired cross-language disagreement. Report all examples in the denominator, including abstentions. If repeated draws are used, retain every draw and report variability rather than selecting the best. The scorer accepts one frozen draw at a time; run separately for each draw.

Do not present lookup of an unchanged catalogue label in the same catalogue as a test of conversational classification. Separate exact label retrieval, paraphrased labels and occupation descriptions; disclose whether reference descriptions entered retrieval. Independently inspect errors and ambiguous gold labels.

## Local task-description evaluation

Obtain licensed or consented descriptions across target languages, regions and occupations. Include non-work, ambiguous, multi-activity and out-of-catalogue cases. Two independent regional occupational reviewers assign acceptable activities and document disagreement before model predictions are revealed. Adjudicate disagreements and report agreement; a multi-label task requires a multi-label scorer rather than forcing a unique gold occupation.

Where possible, repeat the same underlying task in several languages and distinguish translation effects from local differences in work content. Keep matched examples together when splitting data and constructing uncertainty intervals. Freeze the sample and success criteria before a confirmatory run. Representative platform inference requires access to an appropriate sample and knowledge of the production classifier; this reference benchmark alone cannot establish it.

## Scoring file contract

The accompanying research-code package contains `score_multilingual.py`. Input predictions have `esco_uri,language,predicted_isco4`. Exactly one row per frozen example and draw is required. Join gold labels from the official reference; reject unknown identifiers, invalid four-digit outputs and duplicate examples. Blank predictions are abstentions and count as incorrect in unconditional accuracy. Outputs report per-language sample size, exact and ISCO2 accuracy, abstention, and within-occupation cross-language disagreement. Confidence intervals or representative-population claims require a sampling design beyond this deterministic score.
