The spread in AI productivity estimates — a randomised trial's measured 19% slowdown at one pole, task-exposure projections of 1.5 percentage points a year at the other — is not noise to be averaged away but a structural artefact of what each study design can see, so the right number depends on the decision: measured trials matched to your task mix for firm investment, population-scale quasi-experiments for policy, and national accounts anchored to the low-to-middle of the projection range for macro forecasting while firm adoption sits near a fifth.
The most rigorous single trial of AI-assisted work yet run — a randomised controlled trial — found that experienced open-source developers took 19% longer with AI tools than without (95% CI: 2%–39% longer). At the other pole, a task-exposure projection from Goldman Sachs holds that generative AI could add roughly 1.5 percentage points a year to US labour productivity growth over a decade after widespread adoption. Both numbers come from credible institutions doing careful work. Neither is wrong in its own terms.
METR RCT found allowing AI increased completion time by 19%, CI +2% to +39%arxiv.org
Goldman projects ~1.5pp/yr US labour productivity uplift post-adoptionwww.goldmansachs.com
The thesis of this report is that the conflict is structural: each study design can only see what it holds still. A controlled trial on assigned tasks cannot see organisational adaptation; a national accounts series cannot isolate AI from everything else in the economy; a projection is not an observation at all. These differences are not defects to be fixed before a "true number" emerges — they are permanent features of measuring a general-purpose technology mid-diffusion, and they will persist for years.
The practical conclusion, developed in Section 7, is that weighting is decision-context-specific. There is no single correct estimate, but there are demonstrably wrong ways to use each class of evidence — and the wrong ways are currently the popular ways.
The literature sorts into five design classes, and most public confusion comes from quoting numbers across class boundaries as if they measured the same thing.
**(a) RCTs and lab trials** randomise access to AI within a defined task set: METR's developer trial, Peng et al.'s Copilot experiment, Noy & Zhang's writing tasks, Dell'Acqua et al.'s consultant and P&G experiments, Cui et al.'s three firm-run trials, Choi et al.'s legal tasks, and Microsoft's M365 Copilot licence-allocation experiment. They hold the task and tool constant and deliver clean causal identification of task-level speed and quality — but they sample recruited or volunteer workers, run for hours to months, and cannot detect whole-job effects, reallocation or spillovers.
Cui et al. ran three firm RCTs as encouragement designs within ordinary business operationspapers.ssrn.com
**(b) Quasi-experimental field deployments** exploit staggered rollouts or policy variation with difference-in-differences identification. Brynjolfsson, Li and Raymond's call-centre study is the canonical example — and it is quasi-experimental, not an RCT, a distinction press coverage routinely blurs. Humlum and Vestergaard's Danish study — circulated as "Large Language Models, Small Labor Market Effects" and retitled "Still Waters, Rapid Currents" in its March 2026 revision — links representative surveys of roughly 25,000 workers to administrative payroll records. These designs see real workflows over months to years, at the cost of resting on parallel-trends assumptions rather than randomisation.
The study analyses the staggered introduction of an AI assistant across 5,172 support agents using DiD, not randomisationarxiv.org
Humlum & Vestergaard link ~25,000 workers across 11 occupations to Danish admin payroll datawww.nber.org
**(c) Self-reports, surveys and internal benchmarks**: GitHub's developer surveys, the St. Louis Fed's Real-Time Population Survey, Google's and Klarna's corporate claims, MIT NANDA's enterprise-pilot review. Cheap, broad, fast — and, as Section 5 shows, systematically measuring perception rather than output.
St. Louis Fed RPS is a nationally representative survey eliciting self-reported counterfactual time savingswww.stlouisfed.org
**(d) Aggregate task-exposure projections**: Goldman Sachs, McKinsey, Acemoglu, Aghion & Bunel, the OECD, the Penn Wharton Budget Model, Anthropic's Economic Index. These are not measurements of anything. They map occupations to exposed tasks, apply assumed cost savings and adoption paths, and extrapolate. Quoting Goldman's 1.5pp or Acemoglu's ceiling as if someone had observed them is the second routine misclassification.
Acemoglu's published estimate is a model-derived ceiling: TFP gains over ten years of less than 0.53%academic.oup.com
**(e) National accounts and TFP measures** — the BLS labour productivity and total factor productivity series — capture everything that actually happens in the economy and attribute none of it. BLS itself notes it "implicitly captures AI use through its capital measure of software used in production"; no dedicated AI category exists.
BLS captures AI only implicitly via software capital; no dedicated AI categorywww.bls.gov
Each class answers a different question. The measurement gap is what happens when the answers are read as competing estimates of the same quantity.
Five mechanisms move estimates in predictable directions by design class.
**Task scope.** RCTs measure assigned, self-contained tasks. Noy & Zhang's professionals completed one-off incentivised writing tasks averaging 27 minutes in the control group (a working-paper figure) — and the trial's roughly 40% time saving (37% in the working-paper specification; CI not reported; p<0.001) says little about a job in which such tasks are a small slice. Humlum and Vestergaard offer exactly this reconciliation: RCTs measure narrow, cleanly scoped tasks, while whole-job data captures the dilution of AI assistance across everything else a worker does. Narrow task scope inflates; whole-job scope deflates.
Control group averaged 27 minutes; treatment cut time by 10 minuteswww.science.org
Authors reconcile their small effects with 15–50% RCT effects via task-scope dilutionbfi.uchicago.edu
**Worker selection.** METR deliberately recruited experienced maintainers averaging five years on their own codebases — the population where AI has least to add and the most implicit standards to violate. Brynjolfsson, Li and Raymond found the mirror image: a 15% average gain concentrated among novices — roughly 34%, as stated in the 2023 working paper — with minimal impact on the most experienced agents (CIs not reported in reviewed sources). Selection can also corrupt a design over time: METR's follow-up suffered developers refusing to work without AI, which the organisation says "likely biases downwards our estimate".
16 developers, 246 tasks, ~5 years' average prior experience with the repositoryarxiv.org
15% average (QJE 2025); ~34% for novices per the 2023 draft; minimal effect for experienced workerswww.nber.org
METR reports growing refusal to participate without AI, biasing estimates downmetr.org
**Time horizon.** Short trials catch workers on the learning curve — METR's developers typically had only a few dozen hours with Cursor, and the authors flag possible learning effects appearing only after several hundred hours. But longer horizons cut both ways: Humlum's two-year window may precede the workflow redesign through which time savings become output, while quarterly macro data arrives too early in diffusion to show anything.
Developers used Cursor only a few dozen hours; strong learning effects may appear latermetr.org
**Spillovers.** Task-level gains do not sum to economy-level gains, and no micro design can see general-equilibrium effects — reallocation, new task creation, bottlenecks. This is precisely where the projections diverge: Goldman assumes roughly 25% of tasks eventually automated and includes labour reallocation and new-task creation; Acemoglu restricts to the 4.6% profitably automatable near term and excludes both — a parameter choice Goldman says cuts his estimate "by 5½ times relative to ours".
Goldman attributes the gap to automatable-share and reallocation assumptionswww.aei.org
**Organisational adaptation.** Firms must rebuild workflows before tools show up in output — the intangible-capital argument behind Brynjolfsson, Rock and Syverson's productivity J-curve, in which TFP is first understated during the investment build-out and later overstated. Microsoft's own licence-randomised experiment found workers did not shift responsibilities without broad institutional effort. Pre-adaptation measurement suppresses; post-adaptation self-congratulation inflates.
J-curve: TFP underestimated during GPT diffusion, overestimated laterwww.nber.org
Few substantive changes in work patterns; larger shifts require institutional effortarxiv.org
On this analysis, the sign and size of each headline number are largely predictable from where its design sits on these five dimensions — which is why the gap will not close with better data collection alone.
Source: Primary studies as catalogued in the report's evidence ledger; confidence intervals and sample sizes in the companion table (Studies published 2023–2026; data collected 2020–2025)
The chart shows point estimates only — the bar format cannot draw error bars — so the confidence intervals live in the companion table below, and several are wide enough to change the story. Aggregate projections are deliberately excluded: they are denominated in percentage points a year of economy-wide productivity growth, not per-task percentage effects, and plotting them on the same axis would commit the category error this report warns against. They appear in Section 6.
| Study | Design | Effect | 95% CI / uncertainty | Notes |
|---|---|---|---|---|
| METR 2025 | RCT | 19% slower | CI 2%–39% slower | 16 devs, 246 tasks |
| METR follow-up 2026 | RCT (selection-confounded) | 18% faster (returning), 4% faster (new) | returning: 38% faster to 9% slower; new: 15% faster to 9% slower | METR: likely a lower bound; "very weak evidence" |
| Peng et al. 2023 | RCT/lab | 55.8% faster | CI 21%–89% (per GitHub's write-up) | one task; N=95 per GitHub blog |
| Noy & Zhang 2023 | RCT | 40% less time; +18% quality | CI not reported; p<0.001; 0.8/0.4 SD | 453 professionals, incentivised writing tasks |
| Cui et al. 2025 | RCT ×3 firms | +26.1% tasks completed | SE 10.3% → derived 95% CI ≈ +6% to +46% | 4,867 developers; derived CI = 26.08 ± 1.96×10.3 |
| Dell'Acqua et al. 2023 | RCT | +25.1% speed, +12.2% tasks (in-frontier); −19pp correctness (out-frontier) | CI not reported in reviewed sources | 758 BCG consultants |
| Choi et al. 2024 | RCT | ~22% avg time saved (12–32% by task) | CI not reported | 60 law students |
| Dell'Acqua et al. 2025 | RCT | +16.4% individual speed; +0.37 SD quality | CI not reported in reviewed sources | 776 P&G professionals |
| Brynjolfsson et al. 2025 (QJE) | Quasi-exp DiD | +15% issues/hr avg; +34% novices | CI not reported in reviewed sources | 5,172 agents; 2023 draft said 14%/N=5,179 — use QJE figure |
| Humlum & Vestergaard 2025 | Quasi-exp DiD | ~0% earnings/hours | CIs rule out >2% | 25,000 workers; self-reported savings 2.8% |
| Microsoft M365 Copilot RCT | RCT (license allocation) | small/insignificant output changes | no summary CI | 56 firms; excluded from chart (no point estimate) |
Read top to bottom, the randomised evidence on narrow or novel tasks clusters between roughly +15% and +56%: Peng et al.'s lab trial found developers 55.8% faster on a single HTTP-server task (95% CI 21%–89%, per GitHub's own write-up of the experiment — the arXiv abstract states no CI); Noy & Zhang's RCT found 40% less time and 18% higher quality on writing tasks (CI not reported; p<0.001; the Science-published figures supersede the SSRN draft's 444 subjects and 37%). Cui et al.'s pooled firm RCTs — published in Management Science in 2026 but run on GPT-3.5-era Copilot around 2022–23 — found 26.1% more completed tasks, though the derived 95% CI (26.08 ± 1.96×10.3) spans roughly +6% to +46%. Dell'Acqua et al.'s BCG experiment found +25.1% speed inside the AI capability frontier but 19 percentage points lower correctness on a task outside it — the frontier's location being invisible to the workers themselves.
Copilot treatment group 55.8% fasterarxiv.org
N=95, CI 21%–89%, P=.0017github.blog
40% time reduction, 18% quality gain, N=453www.science.org
26.08% increase, SE 10.3%, 4,867 developerspapers.ssrn.com
+25.1% speed in-frontier; −19pp correctness out-frontierpapers.ssrn.com
Against this cluster stands the one RCT on experienced developers doing their own real work in mature codebases: METR's 19% slowdown (CI 2%–39% slower). Its February 2026 follow-up — 57 developers, 143 repositories, over 800 tasks — now estimates returning original developers 18% faster (95% CI: 38% faster to 9% slower) and newly recruited developers 4% faster (95% CI: 15% faster to 9% slower). But METR itself calls the follow-up "only very weak evidence", because 30–50% of developers said they were withholding tasks they did not want to do without AI, and refusals to participate at all rose — selection that confounds the estimate even as METR argues it is likely a lower bound on the true effect. A small illustration of how easily figures drift: METR's own blog summary misstates the original result as 20%; the paper says 19%.
Follow-up estimates, CIs, selection caveats, "very weak evidence"metr.org
The quasi-experimental rows bracket the field. Brynjolfsson, Li and Raymond's call-centre deployment found +15% issues resolved per hour on average (CI not reported in reviewed sources) — the peer-reviewed QJE figure, revised from 14% and N=5,179 in the 2023 draft still widely cited. Humlum and Vestergaard's economy-scale Danish study found precise nulls on earnings and recorded hours, with confidence intervals ruling out effects larger than 2% two years after ChatGPT's launch — nulls that hold for intensive users, early adopters and heavily invested workplaces.
QJE version reports 15%, N=5,172arxiv.org
Precise null on earnings and hours, CIs rule out >2%www.nber.org
The chart's bottom rows are self-reports, plotted for contrast rather than credence: generative-AI users in the St. Louis Fed's nationally representative survey reported saving 5.4% of their work hours as of November 2024 (1.4% averaged across all workers, users and non-users alike) — about 2.2 hours in a 40-hour week, a self-reported counterfactual with no measured comparator — and Danish workers 2.8%. Both sit an order of magnitude below the trial effects, for reasons Section 5 takes up.
Users report 5.4% of work hours saved, ~2.2 hrs/weekwww.stlouisfed.org
Two fragility notes belong on this ledger. First, one of the most spectacular claimed results — an MIT student paper reporting 44% more materials discovered and 39% more patents from an AI tool — was disavowed by MIT in May 2025 over data-integrity concerns and withdrawn from arXiv; it should carry no weight, and its prior circulation is a caution about how eagerly dramatic numbers travel. Second, note how many rows read "CI not reported": for several headline effects — Dell'Acqua's consulting and P&G experiments, Choi et al.'s legal trial — reviewed sources supply point estimates without interval estimates, which limits how hard any of these numbers can be leaned on.
MIT withdrew support and requested arXiv removal of Toner-Rodgers paperwww.aei.org
Where the same subjects supply both a forecast and a measured outcome, the pattern is one-directional. METR's developers forecast a 24% speedup before the trial, still believed AI had made them 20% faster after it — and were measured 19% slower. External experts did no better: economists forecast a 39% speedup and ML experts 38%, against the same measured slowdown.
Forecast 24%, post-study belief 20%, measured 19% slowdownarxiv.org
Economics experts forecast 39% shorter, ML experts 38%arxiv.org
The pattern extends beyond one trial. Microsoft's early M365 Copilot users reported saving about 14 minutes a day — roughly 3% of an eight-hour (480-minute) day, a derived figure — while the licence-randomised experiment across 56 firms found only small, statistically insignificant changes in time in Word and documents authored. Danish workers self-reported time savings of 2.8% of work hours — about an hour a week — against administrative payroll data showing precise nulls, with the paper estimating that only 3–7% of the reported productivity gain passed through to pay.
Users reported ~14 minutes/day savedwww.microsoft.com
Small and insignificant changes in Word time and documents authoredarxiv.org
2.8% self-reported savings; 3–7% pass-through; null earnings/hourswww.nber.org
The gap is not universal, and honesty requires the counter-example: for GitHub Copilot, perception and lab measurement point the same way — over 90% of surveyed developers perceive faster task completion, and the controlled experiment measured 55.8% faster. But those are separate studies on overlapping populations, not a paired forecast-and-outcome design; in every case found where forecast and measurement exist for the same subjects, perception exceeded measurement.
>90% perceive faster completion in GitHub surveysgithub.blog
The corporate record sharpens the lesson. Klarna's claim that its AI assistant did the work of 700 full-time agents came with no published methodology; fourteen months later its CEO conceded the cost-driven approach had produced "lower quality" service and the firm began rehiring humans — even as it raised the self-reported equivalence to 853 agents. Google's much-quoted "75% of new code is AI-generated" is a volume metric with no published methodology, not a productivity claim; the company's separate estimate of a roughly 10% engineering-velocity gain is an internal self-report.
700-agent claim, Feb 2024 press releasewww.klarna.com
CEO "lower quality" admission and 853-agent updatewww.customerexperiencedive.com
75% code-share statementwww.fastcompany.com
~10% engineering velocity, internally measuredwww.aol.com
The consequence follows directly: self-report bias runs consistently positive where it can be checked, so unaudited self-reports and internal benchmarks — including your own staff's — should be discounted heavily, and the St. Louis Fed's survey authors themselves warn that reported time savings may overstate productivity if freed time is not redeployed.
Authors caution reported savings could overestimate productivity gainswww.stlouisfed.org
Source: Observed: BLS Productivity and Costs (USDL-26-0785, latest vintage), AI-era years 2023–2025. Trend = 2015–19 BLS average (≈1.3%, derived). Implied paths derived: trend plus each estimate's stated annual contribution at full adoption (Released June 4, 2026; annual data through 2025)
Does anything in the national accounts corroborate the micro effects? The honest answer runs both ways. US nonfarm business labour productivity grew 2.1% in 2023, 3.0% in 2024 and 2.1% in 2025, against a 2015–19 average of roughly 1.3% a year (derived from BLS annual figures) — so the popular claim that there is "nothing in the macro data" is too strong. But the composition undercuts the AI-efficiency reading: private nonfarm business TFP — the closer proxy for efficiency gains — decelerated to 0.8% in 2025 from 1.5% in 2024, and BLS notes that fall "accounted for most of the labor productivity decline" in that series. (Nonfarm business and private nonfarm business are adjacent but distinct series — 2.1% versus 2.2% for 2025 labour productivity.)
Annual labour productivity 2015–2025, 2025 = 2.1%www.bls.gov
TFP +0.8% in 2025 vs +1.5% in 2024www.bls.gov
What is unambiguous in the data is AI spending, not AI efficiency. St. Louis Fed analysts calculate that information-processing-equipment investment contributed 0.90 percentage points to Q1 2025 real GDP growth — more than two standard deviations above its long-run average and above dot-com-era peaks. Indeed's economists read the 2025 pattern — high labour productivity, decelerating TFP — as capital deepening doing the work, an interpretation rather than a BLS causal claim. San Francisco Fed President Mary Daly put the state of play plainly in February 2026:
IPE investment contributed 0.90pp to Q1 2025 GDP growthwww.stlouisfed.org
TFP deceleration implies capital spending, not efficiency, drove 2025 productivitywww.hiringlab.org
the macroeconomic literature does not find economy-wide productivity gains from AI while microeconomic studies tend to find some gain.
Mary C. Daly, President, Federal Reserve Bank of San Francisco, February 17, 2026
Daly, "The AI Moment?", Feb 17 2026www.frbsf.org
The strongest argument against reading this as disconfirmation is the J-curve: Brynjolfsson, Rock and Syverson argue that during a general-purpose technology's build-out, unmeasured intangible investment makes TFP look worse than it is, with the flattery coming later. The IT precedent supports patience in both directions: business-sector TFP growth ran 1.8% a year from 1995 to 2004, then fell to 0.4% from 2004 to 2018 — gains arrived years after the technology, then faded. Absence of TFP evidence, two to three years into diffusion, is not evidence of absence.
Productivity J-curve mechanismwww.nber.org
TFP 1.8%/yr 1995Q4–2004Q4, 0.4% 2004Q4–2018Q1www.frbsf.org
Two measurement caveats bound what any quarter can prove. BLS revisions are large relative to the effects being hunted: 2023 TFP was revised from 0.7% to 1.3% between March and December 2024, and 2023 labour productivity moved from 1.2% as first reported to roughly 2.0–2.1% in the current vintage (a comparison derived across BLS release vintages).
2023 TFP revised 0.7% to 1.3%www.bls.gov
2023 labour productivity 1.2% first-reported vs current vintage, derived across release vintageswww.bls.gov
Which claims fail the arithmetic, then? On this evidence, none is refuted — firm adoption stood at about 18% at end-2025 by the Fed Board's measure and 19.8% of firms (32% employment-weighted) in the Census BTOS as of May 2026, far from the "widespread adoption" the big projections condition on. But the observed record so far is fully consistent with the low-to-mid projections — Acemoglu's ~0.05pp a year of TFP (the published paper's headline prediction of "less than 0.53%" over ten years, alongside the unadjusted 0.66% ceiling it also reports — the NBER-working-paper figure most press coverage cites — itself down from 0.71% in the first draft), the OECD's 0.25–0.6pp TFP band, the Dallas Fed's 0.3pp scenario — without requiring the high ones. Goldman's 1.5pp path, and Aghion & Bunel's 0.8–1.3pp historical-analogy range, remain untested, not confirmed.
Adoption ~18% of firms end-2025www.federalreserve.gov
BTOS national AI use 19.8%, 32% employment-weighted, May 2026www.census.gov
Dallas Fed: 0.3pp/yr "a more reasonable scenario"www.dallasfed.org
OECD US estimate 0.25–0.6pp/yr TFPwww.oecd.org
Aghion & Bunel GPT-analogy: +0.8 to +1.3pp/yr over next decadewww.frbsf.org
The taxonomy and the mechanisms yield decision-context-specific weighting priors, summarised in the matrix and defended below.
| Decision context | Weight most | Weight least | The trap to expect |
|---|---|---|---|
| Firm investment | RCTs / quasi-experiments matched to your task mix and worker mix | Vendor benchmarks; your own staff's self-reports | Time saved ≠ output gained without workflow redesign (the Humlum pass-through problem) |
| Policy | Population-scale quasi-experiments (Denmark-style); national accounts | Aggregate projections read as forecasts | Binding constraint is organisational adaptation and diffusion, not model capability |
| Macro forecasting | National accounts + adoption data; projection range as a fan | Any single point projection | Labour productivity conflates capital deepening with efficiency — watch TFP; expect the J-curve's later overstatement phase |
**Firm-level investment.** Weight controlled evidence closest to your task mix and worker mix: a call centre staffed with novices should expect something nearer Brynjolfsson's +34% for new agents; a team of senior engineers maintaining a mature codebase should take METR's −19% seriously as a live possibility, not an anomaly. Expect the pass-through problem: Denmark's workers really did save 2.8% of their hours, and it showed up in neither output proxies nor pay — time saved becomes output gained only with workflow redesign, which Microsoft's experiment suggests requires institutional effort rather than local enthusiasm. Discount vendor benchmarks — Microsoft's own report concedes that the tasks in its lab sub-studies, where Copilot users completed tasks in 26–73% of the time taken by non-users, were selected to favour Copilot — and your staff's self-assessments, for the Section 5 reasons. Pilot with measured outcomes — cycle time, error rates, resolved issues — not surveys.
Responsibility shifts require broad institutional effortsarxiv.org
Microsoft caveat that lab tasks overstate real-world impactwww.microsoft.com
**Policy.** Weight population-scale quasi-experimental evidence and national accounts, because policy cares about wages, hours and displacement, which only those designs observe. The Danish result — precise nulls two years in, despite employer investment — is the best current estimate of near-term labour-market impact in a high-adoption rich economy. Treat the projection literature as scenario bounds, not forecasts: its thirty-fold spread (0.05–1.5pp a year) is parameter disagreement, not sampling error, and carries no confidence intervals at all. The evidence across METR's learning-curve caveat, Microsoft's institutional-effort finding and MIT NANDA's 95%-of-pilots-show-no-P&L-impact result (that report's own figure, from a non-peer-reviewed methodology of 300+ initiative reviews, 52 interviews and 153 surveys) points the same way: the binding constraint is organisational adaptation and diffusion, not model capability.
95% of organisations studied saw no measurable P&L impactmlq.ai
**Macro forecasting.** Anchor on national accounts plus adoption data, and hold the projection range as a fan rather than picking a winner. On this evidence the sensible prior weights the low-to-middle of the fan while adoption sits near a fifth of firms, shifting weight upward only as adoption, TFP and quasi-experimental results move together. Watch TFP, not labour productivity alone — 2025 showed how capital deepening can hold the headline number up while the efficiency signal weakens — and remember the J-curve cuts both ways: today's possible understatement is tomorrow's likely overstatement.
The design-flaw arithmetic worth internalising: narrow task scope, volunteer or novice-heavy selection, self-report instruments, vendor-selected benchmarks and short novel tasks all inflate estimates; short pre-adaptation horizons, experienced workers on mature codebases, unamortised learning-curve costs and invisible spillovers all suppress them. A number's likely bias is readable from its study design before its magnitude is.
Finally, the re-opening dynamic. No single number can settle the question during the diffusion period, because the structural gap is a property of the designs, not of the current data vintage. The checkpoints: quarterly BLS productivity releases (Q2 2026 lands on August 6, 2026), the annual TFP release, Census BTOS adoption readings (noting the November 2025 question redesign that breaks comparability with earlier vintages), and the next experimental waves — METR is redesigning its study precisely because selection effects have degraded the old design. Each new quarter should be checked against the fan in Chart 3: sustained TFP acceleration alongside rising adoption would move weight toward the Aghion & Bunel and Goldman paths; continued TFP deceleration through 30–40% adoption would start to falsify them. Until then, the only defensible answer to "what is AI's productivity effect?" is: which decision are you making?
Q2 2026 data due; Q1 2026 +0.3% q/q annualised, +2.8% y/ywww.bls.gov
BTOS question redesign Nov 2025 breaks series comparabilitywww.census.gov
Produced by AI. A cluster of agents wrote this report from the cited sources. Check it against those sources before you act on it.