The Compact Institute

Measuring the AI economy

A shared data commons for the government, labs, and the public

Anushree Chaudhuri

Who gets to measure the AI economy?

The measurement gap

In November 2025, economists at Stanford's Digital Economy Lab reported that employment among 22- to 25-year-olds in the occupations most exposed to AI had fallen about 16 percent relative to older workers in the same occupations. Young software developers were down close to 20 percent from a late-2022 peak. The Bureau of Labor Statistics’ reported unemployment rates would not have revealed this shift; the Stanford researchers relied on private ADP payroll records data covering roughly 25 million workers each month to make this finding possible.

The Current Population Survey (CPS), which produces the official unemployment rate, contains only a few dozen young software developers in a typical month; estimates for a group that small can move by 20 percentage points from sampling noise alone. Because the CPS only surveys 60,000 households per month, the federal system can produce a reliable national unemployment rate while completely missing changes within a particular occupation and age group.

Our national accounts and the unemployment survey were built in response to the Great Depression, when household surveys and periodic reporting were among the best methods available. However, firms adopt AI task by task, and usage can shift between monthly releases. Frontier labs record which tasks people use their models for and how usage differs across occupations in near real-time. They can also estimate whether a model is assisting with a task or doing more of the work itself. Anthropic publishes part of this through its Economic Index, and OpenAI has used aggregated ChatGPT data in its own work on occupational exposure. These releases are useful, but the companies decide what to measure and what to publish. Researchers and public agencies cannot independently test the impact of policy interventions or build on the underlying data.

What other systems show

Examples from other countries show that better measurement does not require one centralized government database. Estonia's X-Road system leaves individual data records with the separate agencies that collect them and moves queries between those agencies over an encrypted exchange. India's Account Aggregator framework uses a regulated intermediary that relays encrypted financial data without being able to read it. The Nordic register-based systems operate on high levels of public trust, when records share stable identifiers and statistical offices have a legal right to link them. The Eurostat code requires that the authority producing the statistic has enough independence that political or financial interests cannot change the definition of a metric if results become inconvenient.

A measurement commons

A new measurement commons for the AI economy could borrow from these systems. Raw payroll data, government records, and model-usage logs would remain at their sources: payroll companies, statistical agencies, and AI labs, respectively. A shared standard, probably anchored to the existing Department of Labor’s O*NET task structure, would define occupation, task, exposure, and realized usage. An independent, government-chartered third party would run agreed computations and publish the results. Secure multiparty computation or trusted execution environments could join data without requiring either side to hand over its raw records; differential privacy could limit what a published statistic can reveal about any worker, task, or prompt. The public would receive real-time aggregate indicators, while vetted researchers could work with more detailed data inside a secure environment. The UN Privacy Enhancing Technologies Lab has already tested this general approach with national statistical offices that needed to reconcile confidential trade data.

Participation in this system requires careful consideration of incentives. A first mover lab participating in a pilot could bear the reputational cost of a negative result while competitors use the findings to their benefit without contributing data. The exchange therefore has to be reciprocal: participating labs would receive privacy-protected benchmarks by industry, firm size, occupation, and task, without access to another lab's raw data. They could use those results to decide where adoption is stalling and to give enterprise customers independently audited evidence about how a product is used and which labor outcomes follow. Federal procurement could make this easier to mandate: covered AI contracts could require recurring data contributions under a fixed term and publication schedule. Any later public benefit or market protection for frontier developers could carry the same requirement.

A first pilot

Junior-hiring subsidies, wage insurance, and taxes on AI-related rents need timely, occupation-level evidence. States and cities also need a credible counterfactual for testing local policy interventions. A first pilot could test whether changes in model usage within a set of occupations predict changes in hiring or wages over the following months. One lab would be enough to test the computation system, but public reporting should either include at least two labs or avoid identifying the contributor. The pilot agreement would set a minimum term, a fixed release schedule, and rules for publishing analyses specified before any party can withdraw.

Long version

Summary

The federal government's labor statistics can estimate aggregate national unemployment rates reliably, but are too small and slow to detect many occupation-level shifts in real-time. On the other hand, frontier AI companies can track how their models are used at the task level. Because the public cannot inspect the underlying data, a few firms determine what is measured and released. This essay proposes a measurement commons that joins lab, payroll, and government data without exposing raw records.

Drawing on several national systems, I propose a shared measurement commons. Raw data would remain with the originating agency, payroll company, or lab. An independent, government-chartered nonprofit would run agreed computations and publish privacy-protected results, using a common exposure-and-usage standard anchored to O*NET. Public users, vetted researchers, government agencies, and contributing firms would receive different levels of access. Contributing firms would receive private, privacy-protected benchmarks as an incentive to participate. Continued contribution could also be made a condition of covered government contracts or other public benefits, under fixed reporting and publication rules.

Introduction

In November 2025, economists at Stanford's Digital Economy Lab reported that employment of workers aged 22 to 25 in the occupations most exposed to AI had fallen about 16 percent relative to older workers in the same occupations, with young software developers down close to 20 percent from a late-2022 peak (Brynjolfsson, Chandar, and Chen 2025). They found this using records from the global payroll provider company ADP, with a dataset that covers roughly 25 million workers each month. Normally these changes are the responsibility of the Bureau of Labor Statistics to report. However, our current federal statistical system could not have produced the same result. The Current Population Survey (CPS), which generates the official unemployment rate, contains only a few dozen young software developers in a typical month, far too few to measure a change of this size. Its monthly estimates for such occupations can vary by 20 percent or more from sampling noise alone.

Around the same time, the UK's Office for National Statistics reported its main labor figures as "official statistics in development" rather than full official statistics, because response rates to the Labour Force Survey had dropped below 70 percent and the data was no longer reliable enough to release normally (Economics Observatory 2024). Survey response rates have been falling across rich countries for years, and the U.S. National Academies concluded in 2024 that analysts trying to track AI's labor effects were "flying blind" because federal statistics are too slow and aggregated (NASEM 2024).

The instruments the United States uses to measure its economy were built in the mid-twentieth century and are losing accuracy and relevance at a time when they matter most. The parts of the economy that AI is changing the fastest are also parts that these instruments measure worst. At the same time, the frontier AI companies record how their models are used at the level of individual tasks, close to real time, and can estimate when a model substitutes for human work. Governments and the public do not have access to the underlying data. This creates an information asymmetry that makes it easy for a small number of actors to misreport or even manipulate data on AI labor impacts, increasing risks of extreme power concentration and scenarios like gradual disempowerment.

The rest of the essay has two parts. The first compares the U.S. survey system with the Nordic register model, Estonia's X-Road, India's Account Aggregator framework, Brazil's Cadastro Único, and the rules used by Eurostat. The second proposes a measurement commons, a shared occupation-and-task standard, incentives for lab participation, and a first pilot. Readers interested mainly in implementation can skip to "Modelling a measurement commons" or "A first pilot."

History of the U.S. economic measurement system

The current U.S. national income accounts and the monthly unemployment survey were both products of the Great Depression era. Before these accounts existed, the federal government tracked the economy through partial indicators and proxies like freight car loadings and industrial production indices; the Commerce Department later described Hoover and Roosevelt as having fought the Depression without any comprehensive measure of national output (BEA 2000). Responding to this experience, the Department of Commerce commissioned Simon Kuznets to build the accounts we use today. His 1934 report to the Senate produced the first official estimates of national income. Kuznets warned in that report that national income should not be read as a measure of welfare, a caution that later use of GDP mostly ignored.

The unemployment side of accounting followed the same path. The U.S. had no direct national count of the unemployed before the 1930s, and during the early Depression there were competing estimates from indirect methods circulated widely (Census Bureau). The Works Progress Administration launched a monthly household sample survey of unemployment in 1940using probability sampling, which was a new method at the time. That survey moved to the Census Bureau in 1942 and became the Current Population Survey (CPS) in 1948, with analysis passing to the Bureau of Labor Statistics (BLS) in 1959.

The oldest continuous federal statistical function that remains in use today is in tracking agricultural output. Agricultural statistics began with a $1,000 appropriation in 1839 and a Division of Statistics in 1863, set up so that ordinary farmers had a neutral source of crop and price information and were not at an informational disadvantage against speculators (USDA NASS). After leaks of cotton estimates were used to trade ahead of the market in 1905, the Department built a physical lockup procedure for sensitive reports. This data required permission to access in sealed rooms with a rule that not even the Secretary could access the numbers before release.

The pattern in how our national income accounts, unemployment rates, and agricultural output are tracked shows that it is moments of economic crisis and necessity that generate improvements and redesigns of our measurement capacity, using the best methods available at the time. In creating our accounts system, the federal government also treated manipulation of official statistics as a primary risk and designed against it, often at the cost of more efficient data collection and linking strategies.

Comparing international measurement systems

Most countries fall into several distinct models for how they measure a population and their economy. The procedure by which records about a person or a firm are joined across different government databases is the key differentiator across these systems. I outline some of these systems across the following contexts: the U.S. survey-based approach, the Nordic countries’ register model, Estonia’s decentralized datalake, India’s consent-based identification sharing, and trade-offs in China on surveillance and accuracy. The design choices involved (often more a product of history or crisis than intention) determine the speed, granularity, and efficiency of a country’s statistics, as well as the privacy and political risks absorbed by the institutions governing the system.

The United States: surveys conducted by siloed agencies

Unlike many other countries, the United States does not have a single national statistical office. Official statistics come from a decentralized network of 16 recognized statistical agencies and units, 13 of them principal agencies (these directly answer to the President or Congress), each inside a different cabinet department with its own budget, legal authority, and confidentiality rules (statspolicy.gov). OMB's Chief Statistician and the Interagency Council on Statistical Policy coordinate them but do not have full operational control, and agencies cannot freely cross-access or merge their data.

The main labor measurement instrument, the CPS, is a survey of about 60,000 households a month run by the Census Bureau for BLS (Census methodology). It uses a 4-8-4 rotating panel: a household is interviewed for four months, rested for eight, and interviewed for four more. The unemployment rate is an estimate from this sample. Thus, although a national unemployment rate figure may be roughly accurate, an estimate for a single occupation by age and by location can depend on just a handful of households in the panel and have large sampling errors. This is one of the reasons the AI effect on young developers was not obvious in official data even though it was easily uncovered by analyzing payroll.

Because the United States has no universal personal identifier or population register system, census and survey-based approaches are the only way to collect economic accounting data. The Social Security number is not a clean identifying key, since not everyone has one, it is recorded inconsistently, and its use is legally restricted. There are also two statutes that block off the most important data: Title 13 makes Census data confidential and usable only for statistical purposes, and Title 26 section 6103 keeps federal tax data confidential and lets the IRS share it with Census only for narrow statistical purposes (CRS IF12957). Measuring AI’s economic effects would require linking firm-level economic activity to tax records, which is therefore close to impossible without new authority, agency rules, or legislation.

The Census Bureau's linkage hub at Center for Administrative Records Research and Applications (CARRA) offers a workaround. Incoming records are matched probabilistically against a reference file built from Social Security Administration data, and each matched person is assigned a Protected Identification Key, “an anonymous identifier as unique as a SSN” that then links files internally without exposing the actual identifier (Census CARRA 2014). Access for outside researchers runs through the Standard Application Process, and the National Secure Data Service, mandated by the 2022 CHIPS and Science Act, is now piloting a shared-services model for secure linkage (NSDS). To patch the survey gaps, statistical agencies and researchers increasingly lean on private data: ADP payroll records, Lightcast job postings, and LinkedIn profiles, which are fast and granular but not representative or linked to government files (Pew 2025).

Fig. 1Figure 1. A comparison of survey-based versus register-based measurement. While countries with a census-based system like the U.S. use a household survey to make population-level estimates, a register-based system has full population data directly.

The Nordic register model

On the other hand, the Nordic countries are able to produce annual statistics, including a full equivalent to a census, by linking administrative registers the state already maintains. The United Nations Economic Commission for Europe (UNECE) describes the model as treating all statistical registers as one system rather than separate files (UNECE). The system works using three base registers: a population register keyed to a personal identity number, a business register keyed to a business identity number, and a register of addresses, buildings, and dwellings keyed to a dwelling identifier. Activity registers for employment, income and tax, education, and pensions link to these base registers through the same keys. The personal identity number allows the agency to link each person in a dwelling and building through an address code, link each person to an employer through the job record's business and establishment identifier, and attach tax, education, and welfare records to the right person. The census is just the annual join of these registers on their keys and thus requires no field survey to produce.

Fig. 2Figure 2. In the Nordic register model, a universal personal ID links the base registers and the census is recomputed from them each year.

Central population registers were established across the Nordic countries between 1964 and 1969. Denmark ran the first fully register-based census in 1981, Finland in 1990, Sweden and Norway in 2011. There is a substantial cost savings to this approach: Statistics Norway reported its 2011 census cost about NOK 3 per person, roughly a tenth of the 2001 census, while a traditional census costs far more per head (for comparison, the 2020 U.S. census cost 43 USD per person according to Government Accountability Office estimates).

The legal basis for the Nordic system is a national statistics act that grants the statistical office the right to access administrative records at the unit level and link them (all linked microdata remain confidential). This system only works because it has three components that exist together: universal identifier, a statutory linkage right, and enough public trust to maintain both. The latter precondition of public trust is required for a model like this to work, but it is also the hardest to design if it doesn’t exist already.

Estonia and X-Road

Estonia built a digital version of register-based statistics on a national data-exchange layer called X-Road. X-Road is not a central database in the way the Nordic register-based census is; data stays in each registry and is queried point to point on demand, so there is no central data lake (x-road.global). Each organization connects through a Security Server, which signs, encrypts, time-stamps, and logs every message. A Central Server maintains the registry of members and the trust policy but does not participate in the live data transfer paths. When Statistics Estonia needs data, its Security Server opens a mutually authenticated, encrypted channel to the provider's Security Server, the provider's system queries its own database, so that only the result is returned. Records can be joined across registries because every registry uses the same identifier, the isikukood, an 11-digit personal identification code. For the 2021 census, Statistics Estonia drew on nearly 30 registers this way (rahvaloendus.ee). The system processes about 2.2 billion transactions a year across roughly 52,000 organizations. An overarching design rule in Estonia’s system is “once only,” meaning a citizen provides each fact to the state only one time and it is reused across all databases, reducing attrition and data fatigue.

Fig. 3Figure 3. Estonia's X-Road: records join across registries by the isikukood, with no central data lake.

X-Road demonstrates that a data-sharing system does not require pooling data in one place. It can leave data with its originating source and securely transfer only queries and results, protected by cryptographic controls and a catalogue of data connectors.

India: Aadhaar, India Stack, and consent-based sharing

India built population-scale digital infrastructure in about a decade, organized in layers known collectively as India Stack. The identity layer is called “Aadhaar,” a 12-digit number backed by a central biometric repository run by UIDAI, which exposes an authentication service that returns only yes or no and a consent-based electronic KYC service. A virtual ID, a revocable 16-digit alias, lets people authenticate without exposing the Aadhaar number itself. The payments layer is called UPI (Unified Payments Interface), run by the national payments corporation, which now supports more than 10 billion transactions a month (indiastack.org) and has made the vision of a largely cashless economy possible (BBC).

The most innovative aspect is the data layer, the Account Aggregator framework, built under a design called the Data Empowerment and Protection Architecture. An account aggregator is a regulated intermediary that lets a person share financial data stored by one institution with another, under explicit consent, without the aggregator accessing the data. The roles involved are the user, the account aggregator, the financial information provider that stores the data, and the financial information user that consumes it (ReBIT spec). The user links accounts and approves a request through the aggregator's app; the aggregator generates a digitally signed consent artifact; the data user submits a request with an ephemeral public key; and the data provider encrypts the financial payload end to end so that only the data user can read it. The aggregator relays ciphertext and never accesses the decryption key.

Fig. 4Figure 4. India's Account Aggregator: the aggregator only relays the ciphertext and cannot read the data that is transferred.

This is a working example of a consent-managed, data-blind intermediary at national scale. The framework crossed 100 million consents by August 2024, with about 155 data providers and 475 data users (Sahamati). It is regulated by the central bank with its API standard published by a dedicated body, and a separate non-profit runs the ecosystem registry. However, India also shows the cost of getting identity wrong. Aadhaar's biometric authentication fails for several percent of attempts, and researchers have linked those failures to denied rations and, in the most serious accounts, to deaths (Drèze, Khera et al.).

Brazil: the Cadastro Único

Brazil's Cadastro Único is a single registry of low-income households that helps determine who is eligible for dozens of social programs. The unit of registration is the household, headed by a responsible adult, with individual records nested underneath (gov.br/MDS). Each person historically received a Social Identification Number generated by the public bank, Caixa Econômica Federal, which ran the database and made benefit payments. Since a 2023 law, the tax identifier, the CPF, has become the primary key for every person in the registry.

Collection of the underlying data is decentralized, while putting it into operation is centralized. Households enroll in person at municipal social-assistance centers, where an enumerator records household composition, per-capita and total income, housing conditions, and education for each member. The national database validates entries on the spot for duplication and document consistency, and runs periodic cross-checks against the tax registry and the social-security and formal-employment database; the modernized system now blocks a registration whose data diverge from the tax authority until the person regularizes it (Agência Gov 2025). Two recurring processes keep it up-to-date: 1) an investigation is triggered when a family's data conflict with other federal databases, and 2) a review is triggered when a family has not updated its record for more than two years.

The registry covers about 94.5 million people and helps inform more than 60 federal programs as of 2025 (World Bank 2025), having grown from a 2006 base of about 11 million families (World Bank case study). It is a good example of a registry that is both a measurement instrument and an operational system, as the same database tracks both poverty levels and the delivery of social services. It also shows how a registry can resist gaming, as eligibility is anchored to the tax identifier and re-verified against employment records on a fixed cycle.

China: capacity rising, transparency falling

China runs a powerful statistical system incorporating several data categories, but with little transparency into how its layers are computed. Its National Bureau of Statistics has converged toward international standards and centralized provincial GDP accounting, which eliminated the long-standing gap between summed provincial output and the national figure (CSIS). It also uses satellite imagery, administrative records, digital-payment data, and machine learning for nowcasting.

Labor statistics are organized around the hukou household-registration system, which historically left the roughly 300-million rural-migrant workforce partly invisible—China only began publishing a surveyed urban unemployment rate in 2018. However, these metrics are subject to manipulation based on the agency in charge of defining the values: in 2023 the statistical office suspended its youth unemployment series after it passed 21 percent, then resumed it months later under a method that excluded students (Atlantic Council), changing the percentage. Thus, China is a cautionary case that surveillance capacity does necessarily not produce trustworthy or consistent labor measurement, and suggests that whoever controls the definition of accounting measures may also control the result without appropriate protections in place.

Comparison table

SystemKey identifierCore instrumentRecord joined byGovernance risk it illustrates
United Statesno universal identifier; PIK is a partial workaroundsurvey of about 60,000 households and separate administrative dataprobabilistic matching; tax data remains separatesmall occupation and age cells are too noisy
Nordicpersonal identity numberpopulation, business, and dwelling registersdeterministic join on shared keysrequires a universal ID and public trust
Estoniaisikukoodabout 30 registries queried over X-Roadon-demand query, no central data pooldistributed control requires a mandatory catalogue
IndiaAadhaar and a signed consent artifactAadhaar identity and Account Aggregatorconsent-managed and end-to-end encryptedauthentication failures can exclude people
BrazilNIS, transitioning to CPFCadastro Único family registrycross-check against tax and employment recordsoperational use can exclude poor households
Chinahukou-linked registrationNBS surveys, administrative records, and commercial datacentralized state accountingdefinitions can exclude migrant workers

Fig. 6There are cautionary lessons to be learned from failed attempts in other countries to modernize a measurement system. For example, Kenya tried to build a national digital ID and had it halted repeatedly by its courts for rolling out ahead of data-protection and inclusion safeguards (Future of Privacy Forum). Argentina's statistical agency understated inflation by roughly half for years after 2007 and drew the first formal IMF censure ever issued for bad data. Ultimately, privately produced price indices were used to fill the gap (Reuters/Investing). In both cases, the timing of the modernization did not align with developing appropriate safeguards against misuse, and in Argentina’s case, led to capture of the system by private actors. The European Statistics Code of Practice has components that address both of these failures: it requires statistical authorities to operate free from political or external pressure, and it coordinates a system of independent national offices that produce comparable statistics without a central authority owning the raw data (Eurostat 2025).

Frontier lab data

The AI companies have privately assembled high-frequency, task-level data that current public systems lack. Anthropic analyzes how its models are used through Clio, which summarizes and clusters millions of conversations into aggregated usage statistics without human reviewers reading raw conversations (Tamkin et al. 2024), and publishes part of this as its Economic Index, broken down by occupation and task and by whether a task is being automated or done collaboratively (Anthropic Economic Index). OpenAI built its occupational map of AI's labor effects partly by validating against aggregated, anonymized ChatGPT usage data from its own systems (OpenAI 2026). Dario Amodei has stated that Anthropic has run its index for over a year and that governments have access to data types the company lacks, while the company has usage data the government lacks (Amodei, Policy on the AI Exponential).

This asymmetry is likely to persist. Foundation models have high fixed costs and low marginal costs, which gives them large economies of scale and pushes the market toward concentration, so a small number of firms control the usage data and that is unlikely to change soon (Korinek and Vipra 2025). As these models become general infrastructure used across the economy, Korinek and Vipra argue they begin to resemble utilities, which is one basis for regulating access to those records.

Current exposure studies that began with Eloundou et al. (2023) measure whether a model could cut the time to do a task by at least half, requiring extensive labeling to be accurate. Human raters and GPT-4 score the detailed work tasks in the public O*NET occupational database for whether a model could cut the time to do a task by at least half, with no employment or usage data in the estimate (Eloundou et al. 2023). They state that exposure does not distinguish work AI augments from work it displaces and make no prediction about adoption. Their estimate that around 80 percent of workers have at least 10 percent of tasks exposed is an upper bound on technical possibility, but the research does not translate to a forecast of job loss. The Stanford payroll study is the first large-scale sign that exposure is turning into displacement for some workers. It runs on ADP's administrative payroll records, a private dataset covering about 25 million workers a month that the researchers reached through a corporate partnership, and even there the authors caution that other factors may explain part of what they find (Brynjolfsson, Chandar, and Chen 2025). At the macro level the effects look modest so far. Acemoglu adds no new data and instead runs published task-exposure and cost-savings estimates through a task-based growth model, which puts total factor productivity gains at no more than about two-thirds of a percent over ten years, and he notes that AI can raise measured GDP through low-value activity or energy use without raising welfare (Acemoglu 2024). Card and DiNardo, working decades earlier from Current Population Survey wage data, are a warning against the reasoning AI commentary often uses, since without an independent measure of how much a technology changes the demand for skills, attributing a wage shift to that technology becomes untestable (Card and DiNardo 2002). And much of AI's value reaches people as free or bundled tools that GDP does not capture at all, which is why Brynjolfsson and Collis ran their own online choice experiments, asking people what they would take to give up a service. The median US consumer valued search engines at around $17,000 a year, almost none of which appears in the accounts (Brynjolfsson, Collis, et al.).

The current proposals for closing the gap treat data sharing as a transfer between two parties. Amodei argues that labs should get access to government microdata and that the government should expand its own statistics. Leicht and Ball argue the reverse, that government should receive the labs' usage data, by analogy to the privileged information-sharing that already exists between firms and the national security apparatus for cyber risk, and they note that AI developers who have a commercial interest in the narrative on labor impacts should not be the only ones who can access this data (Leicht and Ball 2026). However, neither argues for access by external researchers, by the general public, or by the states and municipalities that could actually run economic experiments.

Modelling a measurement commons

Building out a shared measurement commons would need to create a standing system that measures the AI economy in close to real time, places the data under independent governance, and lets government, qualified researchers, and the public access data beyond what individual firms voluntarily choose to disclose. By drawing from existing models, we can assemble a system with well-tested components: 1) leave data at its source and move only results, as X-Road does; 2) join records through agreed keys and computation rather than a central pool; mediate access through a data-blind consent layer, as India's account aggregators do; and 3) create governing rules for oversight, verifiability, and independence, as Eurostat and the USDA lockup do. Figure 5 shows a potential architecture.

Fig. 5Figure 5. A possible measurement commons. Raw data remains at each source; only verified, privacy-protected results are shared.

A common concern is that raw usage logs and confidential government records are too sensitive to share. Testing privacy-enhancing computation reduces that risk by letting parties compute a joint result on data that none of them can access with all of the attributors tied to it. The clearest existing framework for this is structured transparency, which separates any data-sharing arrangement into five components: input privacy, so a party can process data hidden from it; output privacy, so the published result cannot be reverse-engineered to an individual; verification of inputs, so a party can confirm the data is genuine; verification of outputs, so a party can confirm the computation ran correctly; and governance of the data pipeline and actors, so each stakeholder has guarantees about how the data is used (Trask et al. 2020).

All the working tools from existing models can map onto these components. Secure multiparty computation provides input privacy: several parties jointly compute a function over their combined inputs while each keeps its own input secret, so a lab and a statistical agency can compute a joint labor statistic without either handing over its records. Differential privacy provides output privacy: calibrated noise bounds what any published number can reveal about a single person or prompt, which is the technique the US Census Bureau adopted for the 2020 census. Trusted execution environments and federated computation offer alternative routes to input privacy, keeping raw data in place and exporting only results. Zero-knowledge proofs support the verification components.

These tools also make the system's independence auditable. This helps avoid the information-capture that occurred in China and Argentina, where a party with a stake in the measurement output also had control over the collection, analysis, or definition of the data. The system should also only publish outputs and aggregates, not raw data or model capabilities. Structured transparency cannot protect against altering data if the source data is copied; thus output verification lets an independent auditor confirm that a published statistic came from unaltered data without any party having access to raw records.

The UN's Privacy Enhancing Technologies Lab ran a pilot in which the US Census Bureau and three other national statistical offices computed shared statistics on trade data without exposing their inputs (UN PET Lab), so similar pilots to one this system would require have been done before. The pilot, run through the UN statistical system, took on a long-standing problem in international trade data. When two countries report the same trade flow, the exporter's figure and the importer's figure often disagree, and reconciling them normally means each office revealing confidential business-level records it is not allowed to share. The statistical offices of the United States, Canada, the United Kingdom, the Netherlands, and Italy instead computed the combined trade totals jointly, so that each office received the reconciled figures and the gaps between country pairs while the underlying import and export records stayed inside each office (UN PET Lab, International Trade). The lab ran two versions of the same computation. One used secure multiparty computation across a peer-to-peer network of nodes, built with OpenMined, and added differential privacy to the published figures. The other kept each office's data inside a hardware enclave, built with Oblivious, and returned signed reports of the result. The two runs showed that more than one privacy-enhancing architecture can answer the same question, with different tradeoffs in how much control each office keeps over its data.

Building a shared exposure and usage standard

A measurement system needs agreed-upon definitions such that its numbers can be compared and trusted by all users. Today the relevant measures are incompatible. Academic exposure scores, the labs’ internal automation-and-augmentation classifications, and realized-usage data each use different units. The Stanford payroll paper combined three of these: ADP payroll data, the Eloundou exposure taxonomy, and Anthropic’s usage-based classification of which tasks are automated. A shared standard would define a common occupational and task vocabulary, anchored to the existing O*NET task structure that both the exposure studies and the labs already use, and a common way to express realized usage: the share of a task's time where a model is doing the work. With that standard, the gap between what models can do and what they are used for could be tracked as a public indicator rather than reconstructed paper by paper. This standard is also required to enable privacy-preserving computation, since secure computation across a lab's usage data and an agency's employment data requires both sides to agree on the categories being joined.

The National Academies has already proposed an institutional form for a system like this: an independent, non-profit, government-chartered entity that operates the infrastructure and expertise for public-private data sharing and integrated analysis, using secure multiparty computation to preserve privacy (NASEM 2024). This entity would operate the secure-computation layer, maintain the shared standard, run the independent verification, and publish outputs through tiered access: aggregate real-time indicators for the public, secure-enclave microdata access for vetted researchers through the existing Standard Application Process, and operational indicators for government to use in policy and evaluation. The consent and access logic could follow a process similar to India’s data-blind intermediary model, where a custodian entity manages access and computation but cannot itself read the underlying data.

The Opportunity Atlas is a narrower precedent for this access model. Economists at Opportunity Insights and Census researchers used anonymized IRS and Census records to publish neighborhood-level mobility estimates while the underlying records remained restricted. Disclosure controls limit what estimates from these neighborhood-level numbers can reveal about any specific household or individual by adding small amounts of random noise to each record.

Legal scaffolding to enable this in the U.S. also partly exists already. The Foundations for Evidence-Based Policymaking Act and its confidentiality provisions created a presumption that statistical agencies can access federal data for statistical purposes, and the National Secure Data Service is already testing privacy-preserving linkage as a shared service (NSDS). Two limitations of the current legislation are that 1) federal tax data sits behind Title 26 and cannot be linked without specific new authority—this is the single largest legal obstacle. And 2) private data cannot be legally compelled; it can only be obtained by agreement, which means good and rigorously-tested incentive design is key to encourage labs to participate.

The independence requirement should follow the Eurostat code, with the entity insulated from both political co-option and the commercial interests of its data contributors. Release should follow the USDA lockup process, so that no single party can access or alter a statistic before its simultaneous public release. Purpose limitation and consent should follow the Indian and Brazilian models, with data used only for the agreed measurement and with safeguards against being repurposed for surveillance. There then needs to be a mandated verification layer that makes all of this easily auditable through legally enforceable and checkable properties, though this could start through voluntary agreements.

Why labs would participate

First-mover advantage alone is too weak of an incentive to support a voluntary data-sharing system. A participating lab could become the sole target of criticism if the estimates show negative labor effects, while competitors use the public results without contributing any of their own data. A lab could also leave once the findings become inconvenient. A pilot therefore needs a benefit for participants and terms that require continued participation even if there is an unfavorable result.

Participating labs could receive a private benchmark generated from the joint computation after the public aggregate release. It could show how the lab's usage differs by industry, firm size, occupation, and task, and which usage patterns precede changes in hiring, wages, or vacancies. The benchmark would disclose no competitor's raw data and would use the same minimum-cell and privacy rules as the public series. A lab could use it to decide where adoption is stalling and to give enterprise customers independently audited evidence about how its tools are used and which labor outcomes follow.

Government procurement contracts provide a stronger condition for continued participation. Federal contracts already use reporting clauses, with noncompliance reflected in contractor performance records. A covered AI contract could require recurring contributions to the measurement commons for the term of the contract, subject to procurement-law review. If the government later gives frontier developers other selective benefits or market protections, continued data contribution could be attached to those benefits as well.

The pilot agreement would set a minimum participation period, a fixed reporting calendar, and pre-specified analyses. Withdrawal would apply only to future data after a notice period; it would not cancel publication of analyses already authorized. The custodian would admit any lab on the same terms. If contracts and voluntary agreements still produced too little coverage, legislation could impose access obligations on foundation models treated as economic infrastructure.

A first pilot

The secure computation can first be tested with one lab and a statistical agency. Any public labor-effects release should either include at least two labs or suppress lab-identifying results. The agency could test whether rising model use within an occupation is followed by changes in hiring or wages. Each lab would produce a monthly table at the O*NET occupation and task level, with usage intensity and the share of activity classified as automation rather than assistance.

The agency would join those cells to employment and wage series using occupation and time, plus broad geography where possible. The agreed calculation would run through secure multiparty computation or a trusted execution environment. Minimum cell sizes and differential privacy would apply before release.

For each major capability release τ, the pilot would estimate:

ln y₍ₒₜ₎ = αₒ + λₜ + Σₖ βₖ (Aₒ × 1[t − τ = k]) + Σₖ γₖ (Eₒ × 1[t − τ = k]) + ε₍ₒₜ₎   (1)

Here, y₍ₒₜ₎ is employment in occupation o and month t, Aₒ is the occupation's realized automation share, and Eₒ is its technical exposure. The βₖ coefficients compare occupations where technical exposure has turned into actual use with similarly exposed occupations where it has not. Coefficients before the release test whether those groups were already on different trajectories (the parallel trends assumption).

A second specification would test the early-career pattern within occupations:

ln y₍ₒₐₜ₎ = α₍ₒₐ₎ + λ₍ₒₜ₎ + μ₍ₐₜ₎ + δ (Aₒ × Youngₐ × Postₜ) + ε₍ₒₐₜ₎   (2)

The occupation-by-time effects absorb demand shifts that affect everyone in an occupation versus just younger workers. The estimate then comes from the gap between early-career and older workers in the same occupation, on the premise that entry-level hiring adjusts first. Wage and vacancy variations of the first equation could use the same design.

The lab would need to agree on the O*NET and automation definition, then produce a recurring aggregate extract. It would also need to let an independent auditor inspect the joint computation. The statistical partner supplies outcome series and publishes the result on a fixed schedule. An independent custodian would document the code and privacy budget, including exclusions and release rules, so neither party could revise an awkward result before publication.

Before the pilot begins, the parties would need to sign an 18-month data-contribution agreement covering a baseline period, at least one major capability release, and the follow-up period. The agreement would specify the monthly extract and every planned public release. A participant could withdraw from later rounds after a notice period, but it could not block publication of an analysis already authorized.

Participating labs would receive their benchmark when the public aggregate is released. It would compare the lab's own usage distribution with pooled reference distributions and report only cells that meet the disclosure thresholds. Public and participant outputs would be released at the same time, and the benchmark would contain no other lab's records.

The pilot would produce a maintained indicator of the gap between technical exposure, realized use, and recorded labor outcomes. If the pre-release trends fail, the result is still descriptively useful. If they pass, the same setup could test wage and vacancy effects or evaluate a local hiring subsidy against a current counterfactual.

Closing

Good measurement instruments are preconditions for most proposed policy responses to mitigate AI's adverse economic effects. A junior-employment subsidy, as well as other interventions like wage insurance, an automation tax, and retention incentives, depend on labor-market data granular enough that employers cannot manipulate it and timely enough that it can inform quick evidence-based iterations of policy design. If the tax revenue generated from human labor decreases, proposals to raise revenue from AI labor generally target corporate profits or rents but offer no method to measure AI-specific rents (AGI Social Contract 2026). Thus, a system that can identify and measure where AI value is being created is an important precondition for taxing it, or for designing an alternative policy mechanism. In addition to improving our understanding of labor market effects, the most useful capability a system like this enables is being able to conduct policy experiments and broadly-accessible evaluation. For example, this would allow us to run an intervention in some states or cities and then measure the difference against a real-time counterfactual.

The New Deal generation built the national accounts and unemployment measures we know today as a response and recovery to the Great Depression. Much of the capacity to measure the impacts of the AI economy already exists and is running inside the frontier labs, but building the measurement system required to create actionable policy insights from that data cannot once again wait for a Depression-level or greater crisis. A shared measurement commons would build the system that lets government, researchers, and the public act on that data, using privacy-preserving computation, structured transparency, and the independent custodian methods described here. Increasing market and power concentration will make that harder to implement the later it is done.

Figures