# Protocol for a task-time measurement study

Status: prospective design, frozen locally on 9 September 2026; not externally preregistered and not executed. This protocol does not retrospectively validate the existing experiment. A model endpoint, independent human time-use evidence and a dated public registration must be specified before a confirmatory run.

## Decision on the existing estimates

Exclude the failed local-model weights and their correlations from substantive conclusions. One task received two quadrillion weekly hours and twelve occupation totals exceeded 168 hours. Retain the original responses, numeric exports, quality flags and downstream arithmetic in an audit archive. Do not clip the extreme value, selectively retry occupations, select weights on benchmark fit, or infer from the reported 0.901 proxy correlation that time weights are unimportant. Normalization checks test arithmetic, not measurement validity.

## Question and target quantity

Can a fixed elicitation procedure approximate independently measured within-occupation task-time allocations? The unit is one occupation-task in the pinned O*NET 30.2 catalogue. The target is a compositional allocation of a standardized 40-hour analytical workweek, not the actual working hours of every worker. Because O*NET activities overlap, the prompt must explicitly allocate each hour once across the supplied tasks and an additional uncovered-activity category. The uncovered share remains visible; task shares are not silently rescaled to exclude it. Evidence is needed to show that such a decomposition is meaningful.

## Design to freeze before generation

1. Name the endpoint, immutable model revision or digest, provider, parameters, seeds where supported, full prompt, occupation/task universe, and the allowed budget. Do not select a model using Anthropic's exposure targets. A larger model may be a useful candidate; size alone is not a validity criterion.
2. Obtain an independent human task-time dataset with documented occupation matching and population. Freeze a development set and a held-out validation set before any tuning. Prevent the held-out time observations and Anthropic occupation scores from entering prompts, retrieval or model selection. Training contamination cannot be ruled out merely by omitting a target from a prompt.
3. Generate exactly three independent draws per occupation from the same frozen specification. Each draw receives the same occupational information in a new session and no previous responses. Record all three, including failures. For technical transport failures only, retry the identical request at most twice, retaining each failed attempt. A content failure is not a transport failure and receives no selective replacement.
4. Freeze the parse rules and plausibility checks: exact task IDs, no duplicates or omitted tasks, finite nonnegative allocations, every allocation at most 40 hours, sum including uncovered activity equal to 40 within 0.01 hours. Reject the whole draw if any check fails. Do not clip or renormalize a rejected draw into compliance. Report extreme concentration, entropy and disagreement as diagnostics, not automatic deletion rules chosen after inspection.
5. The confirmatory three-draw average is defined only when all three draws pass. Otherwise mark the occupation unavailable in the primary result and report its missingness. Average the unrounded shares arithmetically. Report draw dispersion alongside the mean; this is model instability, not a survey confidence interval.

## Validation and analysis

Before accessing held-out outcomes, register the minimum matched occupations, human-data reliability requirements, missingness tolerance and numerical success thresholds justified by the available human instrument. They cannot responsibly be chosen now without seeing the measurement design of that instrument. Use held-out mean absolute share error, rank association, compositional distance and agreement in concentrated tasks; compare against equal-task and simple task-frequency baselines on identical support. Report results both by occupation and pooled with explicit weights. Do not treat O*NET importance or frequency ratings, or PIAAC activity-frequency responses, as observed task-time fractions.

Only after independent time-use validation should the weights enter a fixed public-input US exposure sensitivity. Report the same-sample improvement and deterioration, absolute error, rank correlation and top-ten overlap. No tuning on those outcomes. Such a sensitivity still cannot reproduce unavailable work filters, grouping, employment allocation or unrounded internal inputs. Country transfer requires separate national evidence and remains outside this protocol's validation claim.

Release the registration timestamp, model/settings hashes, all successful and failed draw metadata, permitted numeric outputs, missingness, validation code and uncertainty diagnostics. Separate arithmetic reproducibility, empirical time-use validation and exposure reconstruction as three distinct acceptance decisions.
