The Short Answer
AI productivity gains may not scale evenly; the research tests whether the “AI productivity dividend” diminishes—i.e., whether returns flatten as firms increase exposure or use intensity. The focus is on the *shape* (concave vs. linear), not just the average impact.
For practitioners, this means ROI expectations from “more AI” may face diminishing marginal payoff. Rollouts should be planned and measured around task fit, workflow capacity, and whether additional exposure continues to translate into measurable productivity.
A key caveat is methodological: using the wrong pre-trend tests or clustering choices can produce patterns that aren’t real. The study emphasizes validated measurement by testing the pipeline on synthetic data where the true diminishing pattern is known.
On this page
- Introduction: why “average AI gains” might mislead everyone
- Why This Matters (and where you’ll feel it first)
- What the researchers are trying to measure: the “shape” of AI’s payoff, not just the average
- The measurement strategy: exposure-first difference-in-differences built for curvature
- The part you’ll want to remember: their framework is validated on a synthetic world with known concavity
- Beyond curvature: separating intensity, adoption, duration, and production-side mechanisms
- Key Takeaways
AI’s Productivity Returns: Do They Flatten as Firms Use It?
Introduction: why “average AI gains” might mislead everyone
There’s a common story about AI productivity: early users saw big gains, so the whole economy will too. But this new research (based on the original paper) challenges a key assumption behind many forecasts—that those gains scale evenly. The big question isn’t just whether AI helps. It’s whether the AI productivity dividend diminishes as firms get more exposed, adopt later, or use it more intensively.
This paper takes a very specific approach to that question. Instead of focusing on the average impact of AI, it tries to identify the shape of the return—whether it’s linear, keeps climbing, or flattens (concave). The authors build and validate a measurement framework using linked firm–worker data concepts, anchored around the release of ChatGPT in November 2022. And they go one step further: before touching confidential-style data structures, they validate the whole pipeline on a synthetic panel where the “true” diminishing-returns pattern is known.
Why This Matters (and where you’ll feel it first)
This is significant right now because policy and investment decisions are currently being made with imperfect beliefs about how AI benefits scale. If you assume linear returns, you’ll forecast larger productivity growth from “catching up” firms—especially those with lower initial exposure. But if returns are concave, then the marginal payoff of pushing AI deeper into the economy can shrink fast, and the “productivity dividend” can look a lot smaller than expected.
Here’s a scenario you can actually picture: imagine a mid-sized Canadian manufacturer that’s considering rolling out generative AI across engineering, scheduling, and customer support. The firm might hear optimistic claims from early adopters. But the real payoff may depend on where the firm’s workers sit in the task landscape—how much of their payroll is already in AI-friendly roles, whether complementary tools and workflows exist, and whether the firm has enough operational capacity to translate AI output into measurable productivity. This research is directly about that scaling problem: does each additional unit of exposure or use keep paying off the same way, or does it get harder to extract value?
It also builds on a growing AI research base where many studies show benefits at the task level (writing, consulting, customer support, etc.), and firm-level studies often find positive productivity associations. What’s different here is the emphasis on curvature: the idea that productivity gains might taper off due to saturation, bottlenecks, or selection effects. The paper explicitly argues that getting the shape right is crucial for potential output estimates and output-gap thinking—an issue that’s not just academic when central banks, governments, and large employers are making planning assumptions.
And finally, it surfaces a practical methodological warning: if you test the wrong kind of “pre-trends” with the wrong clustering strategy, you can end up finding patterns that aren’t real. That matters because the temptation in applied economics is to run a familiar test and trust it. This paper shows that familiarity can be dangerous.
What the researchers are trying to measure: the “shape” of AI’s payoff, not just the average
The paper’s core idea is simple to state but hard to measure: AI returns may diminish, meaning deeper exposure or more intensive use yields smaller marginal improvements. In the language of the paper, the central object is whether the productivity gain is concave in exposure/use intensity.
The three “margins” where diminishing returns can show up
The authors frame diminishing returns along three economic dimensions:
- Intensity margin: Do productivity gains flatten as firms are more exposed to AI-capable tasks or use AI more heavily?
- Adoption margin: Do marginal adopters (firms that adopt because their exposure is higher) get smaller benefits than earlier, more entrenched adopters?
- Duration margin: Do productivity benefits plateau over time after adoption?
These correspond to different mechanisms:
- Concavity in intensity can reflect task saturation or bottlenecks in complementary inputs (think: training, workflows, process redesign).
- Declining gains at the adoption margin can reflect positive selection—early adopters may be better positioned to benefit.
- A plateau over time distinguishes a one-off level shift from sustained growth, echoing the idea behind the “productivity J-curve” literature.
Why “average gains from early adopters” can be wrong
A key reason forecasts differ by an order of magnitude is that many projections implicitly extrapolate from early adopters and treat effects as if they scale linearly. But this paper stresses that average effects don’t identify curvature. Two scenarios can produce the same average gain but radically different scaling for later adopters.
Here’s the comparison the paper implicitly forces you to confront:
| What you assume about AI returns | What you’ll predict for “catch-up” firms | What the world might actually do |
|---|---|---|
| Linear scaling (constant marginal returns) | Big productivity payoff as more firms adopt and deepen use | Returns could flatten due to saturation, complementarities, or selection |
| Concave scaling (diminishing marginal returns) | Smaller marginal payoff at higher exposure/use intensity | Benefits rise early, then taper off |
| Increasing returns (convex) | Late adopters eventually “catch up” and surpass early ones | Less consistent with saturation and task bottleneck logic |
The paper’s framework is designed to detect which shape fits the data-generating process—without just trusting averages.
The measurement strategy: exposure-first difference-in-differences built for curvature
To study the shape of returns, you need a design that can credibly compare productivity changes across firms with different AI exposure levels—while addressing the fact that adoption and usage aren’t random.
Step 1: build a pre-determined exposure score (before generative AI is widely available)
The authors use firm exposure constructed from pre-ChatGPT occupational composition. Specifically, a firm’s exposure is a payroll-weighted average of occupation-level exposure scores for generative AI, using the workforce mix from 2019–2021 (before the shock becomes widely available).
They define exposure as:
E_f = sum_s (payroll_share_{f,o}) * exposure_score_o
They also use two exposure indices:
- The GPT-exposure measure (from Eloundou et al., 2024)
- The AIOE index (from Felten et al., 2021)
These are strongly related but not identical (correlation ρ = 0.848), and the framework allows robustness checks with either.
Step 2: estimate changes around ChatGPT’s release using continuous-treatment DiD
Instead of treating adoption as binary, the paper uses a continuous treatment difference-in-differences approach. The timeline is anchored around the ChatGPT period:
- Pre-period: 2019–2022
- Post-period: 2023–2025
The main design uses a binned event-study style first (quintiles of exposure), then moves to dose–response estimation using long differences.
Step 3: run pre-specified concavity tests (the paper’s “curvature detector”)
The authors include three different one-sided tests against linearity, each aimed at detecting concavity in the dose–response of productivity to exposure.
They don’t just run one regression and hope. They specify multiple tests in advance, then check whether the evidence points toward diminishing returns.
A key result from the validation (details below) is that the primary slope-difference test rejects linearity with:
- slope difference: −0.010
- standard error: 0.003
- one-sided p-value: 0.002
That’s the paper’s central “diminishing returns” signal in their simulation.
The part you’ll want to remember: their framework is validated on a synthetic world with known concavity
This is where the paper earns trust. Many empirical papers tell you what their method does, but not whether it can recover the curvature you care about—especially when treatment is endogenous and exposure is measured with noise.
The synthetic panel: 20,000 enterprises, realistic data structure, embedded concave effects
The authors validate every estimator on a calibrated synthetic dataset that imitates the structure of Statistics Canada’s linked microdata ecosystem. The synthetic panel includes:
- 20,000 enterprises
- 40 industries
- 10 provinces
- annual observations 2015–2025
- employer–worker-style structure mimicking Canadian linked data files
Most importantly, the simulation embeds a known concave effect of generative AI use/exposure on productivity (with a phase-in over about three years).
The validation results: does the method recover the “true” dose–response curve shape?
In the synthetic data, the framework:
- tracks the true dose–response curve closely
- detects concavity with high power
- avoids manufacturing diminishing returns when they shouldn’t exist
A few validation numbers the paper highlights:
- Concavity test rejection (primary test) is strong: slope difference −0.010 with SE 0.003
- Confidence interval coverage of the main estimate: 0.97
- Power of the concavity test: 0.93
- In the event-study, pre-2023 coefficients are small and statistically indistinguishable from zero, consistent with parallel trends in the simulation setup.
A practical pitfall they discover: pre-trend tests can over-reject with industry-level clustering
One of the most “real-world” lessons comes from a methodological trap: if you cluster pre-trend tests at roughly 40 industries, you can reject a true null too often.
Their simulation finds the industry-clustered linear pre-trend test rejects the true null about 20% of the time at a nominal 5% level. That’s over-sizing—meaning you may think you see differential trends when you don’t.
The paper points toward safer inference choices, like clustering at the enterprise level for the main joint test and using wild cluster bootstrap methods when appropriate.
Beyond curvature: separating intensity, adoption, duration, and production-side mechanisms
Detecting concavity is the headline. But the paper also tries to answer why and where the diminishing returns show up, using multiple complementary modules.
Intensity (use depth): do heavy users keep getting the same returns?
The paper compares productivity effects between:
- low-intensity adopters (one or two AI applications)
- high-intensity adopters (three or more)
Using a doubly robust difference-in-differences method (built for selection concerns), the validation finds both groups get similar long-difference gains in the simulation. That pattern is consistent with saturation at the intensive margin (i.e., diminishing returns show up as flattening rather than a sharp collapse).
Adoption margin: do marginal adopters benefit less?
They decompose the exposure effect into:
- an adoption gradient (how adoption rates change across exposure quintiles)
- a marginal-return component for firms induced to adopt by higher exposure
In the synthetic CSBC-style survey portion, adoption rises steeply at low exposure and flattens at higher exposure. The decomposition indicates that marginal returns decline—again pointing to diminishing returns at the adoption margin (though the authors note that the decomposition gets noisy where adoption saturates).
Duration margin: do effects plateau after adoption?
For earlier adopters (2017–2021), the paper estimates dynamic treatment effects using a staggered-adoption framework. In the validation:
- there’s a dip in the adoption year (consistent with reorganization costs)
- the effect rises afterward
- later on, the trajectory slows into a plateau
This is exactly the kind of pattern you would miss if you only looked at a single “post” coefficient.
Production-function evidence: diminishing returns show up in TFP and value added
Finally, they embed AI exposure into a production-function framework where productivity follows a law of motion that includes AI exposure. Concavity tests run on outcomes like:
- TFP
- log value added
- log employment
In the simulation, concavity is strong for TFP and value added, and employment effects are small and not significantly concave. That matters because it suggests the measured diminishing returns aren’t merely an artifact of labor adjustment.
Key Takeaways
- The main question is curvature, not averages. This research is built to detect whether AI productivity returns diminish (concave dose–response), not just whether AI helps.
- In their validated synthetic panel (20,000 enterprises), the main concavity test rejects linearity strongly: slope difference
−0.010(SE0.003, one-sided p =0.002), with high confidence interval coverage (0.97) and power (0.93). - Diminishing returns can show up on multiple margins:
- intensity flattening (saturation with deeper use)
- adoption effects declining for marginal adopters
- duration plateauing after early gains
- The authors also identify a practical warning: pre-trend testing can over-reject when clustered at about 40 industries, so inference choices (and clustering strategy) matter a lot.
- The framework is designed for application to Canadian linked microdata inside secure environments (using Statistics Canada-style data architecture), and it aims to inform how AI might affect potential output and the persistence of the productivity gap.
If you want, I can also turn this into a “what it means for AI strategy” guide for executives (how to think about exposure, intensity, adoption sequencing, and complementarity constraints) based on the same framework.
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results: