Difference-GMM and system-GMM
- Evaluation
- Evaluation
- Evaluation Methods
- Data Management
- Performance Monitoring and Evaluation Framework (PMEF)
- Monitoring and evaluation framework
Difference‑GMM and system‑GMM are dynamic panel data methods designed to analyse situations where current farm outcomes (e.g. income, productivity or environmental proxies) depend heavily on past outcomes and where policy support and performance may influence each other over time. These methods help obtain more credible estimates of how CAP support relates to farm‑level results than simpler correlational approaches, especially when long multi‑year panels are available.
Page contents
Basics
In a nutshell
Difference-GMM (Diff-GMM) and system-GMM (SYS-GMM) are dynamic panel data methods used when today's outcomes depend strongly on yesterday's outcomes, a pattern called persistence (for example, current farm income depending on past income), and when policy variables and outcomes may influence each other over time. They extend standard panel regressions by explicitly modelling this persistence and by using internal instruments (lagged values of the variables in the dataset) to reduce bias from ‘chicken‑and‑egg’ problems between support and performance.
The generalised method of moments (Hansen, 1982) is an estimation technique used to identify unknown parameters in models without requiring full knowledge of the underlying data distribution. It works by using ’moment conditions’, which are mathematical equations that include parameters and data and should give an expected value of zero if the parameters are correct. It then finds the parameter values that bring the sample averages of these conditions as close to zero as possible. The GMM estimates model parameters by asking that certain features of the data, such as averages or variances, line up with what the model predicts. It is useful when we are unsure about the exact shape of the data distribution, because it only relies on conditions for these selected features rather than on a full distributional assumption. When we have more such conditions than parameters to estimate, GMM combines all of them in an optimal way, giving more weight to the conditions that are measured more precisely, so that the resulting estimates are as accurate as possible.
Diff-GMM / SYS-GMM work with naturally occurring, non-random data (i.e. the data come from real‑world administrative or survey sources (e.g. FADN, IACS), where farms self‑select into CAP measures or are selected by eligibility rules – not by a random draw). These methods improve credibility but do not reach the causal certainty of randomized experiments. They are designed to extract more credible information on policy-outcome relationships from long farm‑level datasets.
The rational of the model is that the change in outcome y for farm i this year depends on:
- How much y changed last year and two years ago.
- How much the policy support indicators (ISI) changed this year and last year.
- How much other factors (X), such as farm characteristics or external conditions. Changed.
- Plus a remaining unexplained part (the error term).
In more formal terms, the model is:
Δy_(i,t)=β_1 Δy_(i,t-1)+β_2 Δy_(i,t-2)+∑_k δ_k ΔISI_(k,i,t)+∑_k δ_k ΔISI_(k,i,t)+∑_j γ_j ΔX_(j,i,t)+ϵ_(i,t)
where y = income, ISI = policy indicator, X = control variables, Δ = change over time (first difference), i = individual farm, t = time period (typically years) and ϵ_(i,t) = error term.
In particular, Δy_(i,t) is the difference between the current year and the past (first difference) of the dependent variable y:
Δy_(i,t)=y_(i,t) - y_(i,t_(-1))
Δy_(i,t-1) represents the lagged first difference:
Δy_(i,t_(-1) )=y_(i,1) - y_(i,t_(-2))
The coefficient β_1 is the ‘autoregressive’ coefficient that indicates the ‘persistency’ of the dependent variable.
Similarly ∑_k ΔISI_(k,i,t) includes the first differences of all ISI variables considered, and ΔX_(j,i,t) represents the first differences of the control variables adopted to reduce the confounding effect or selection bias.
The coefficients δ_k and γ_j represent the coefficients for the respective ISI indicators and control variables (these variables typically capture the difference between the current and previous year, but in particular cases, additional time lags can be included).
Diff‑GMM and SYS‑GMM allow for ‘chicken‑and‑egg’ situations, where more performant farms may get more CAP support and more support may improve results. To handle this, both methods use the farm’s own history:
- Difference‑GMM looks at changes over time within each farm (how income changes when support changes). By focusing on changes, it automatically removes all farm characteristics that do not change over time (such as soil quality or location). It then uses past levels of income and support as internal instruments to separate policy effects from feedback (=reverse causality, where farm performance influences support) and noise (=unobserved shocks from weather, market, etc., influencing both support and income, creating bias). In other words, because support and outcomes influence each other, the method uses older historical information from the same farm to help disentangle true policy effects from situations where performance affects support, or where random shocks distort the data.
- System‑GMM combines two types of equations:
- The ‘equation in changes’, where the model uses changes over time (Δ income, Δ support, etc.). This is what Diff‑GMM uses exclusively.
- Looks at how changes in income relate to changes in support.
-
The ‘equation in levels’, where the model uses absolute values of the variables (income, support, etc.) instead of only their changes.
- Looks at how income itself relates to support itself and other variables.
This usually gives more precise and less biased results when variables like income and support are very persistent over time (i.e. they barely change over time).
- The ‘equation in changes’, where the model uses changes over time (Δ income, Δ support, etc.). This is what Diff‑GMM uses exclusively.
Both Diff-GMM and SYS-GMM rely on important assumptions:
- The past values used as instruments are good predictors of current variables, but are not correlated with unexpected shocks (random disturbances) in the current period
- The error terms do not exhibit patterns over time (no serial correlation), specifically, there should be no correlation between the error in one period and the error in another period (after taking first differences).
When these conditions hold reasonably well, Diff‑GMM and SYS‑GMM can provide more credible estimates of how CAP support is related to farm outcomes than simple OLS or fixed‑effects models, especially with long farm‑level panels.
Pros and cons
| Advantages | Disadvantages |
|---|---|
| These methods explicitly model the fact that current farm outcomes depend on past outcomes and that CAP support and performance may influence each other over time (the ’chicken‑and‑egg’ problem). This makes them more appropriate than simple OLS or standard fixed‑effects models when income, productivity, or other indicators are highly persistent. | Diff‑GMM and SYS‑GMM are technically demanding. Poor choices about instruments, lag structure, or specification can easily generate biased or invalid estimates, especially if used as a ’black box’ without strong econometric expertise. |
| By differencing and using internal instruments, Diff‑GMM and SYS‑GMM remove all time‑invariant farm characteristics (e.g. soil quality, altitude, long‑run managerial ability), even when these are not observed in the data. | These methods tend to generate many instruments, particularly with longer panels. Too many or weak instruments can invalidate standard diagnostic tests (e.g. Hansen test) and lead to over‑fitting (where the model starts to fit the noise instead of the true signal because it is using too many instruments), giving a false sense of precision. |
| When long farm‑level panels are available (e.g. many years of FADN/FSDN data) and the necessary diagnostics are satisfied, these methods can deliver more reliable estimates of the association between CAP support and outcomes than simpler correlational approaches. | The reliability of these estimates depends on two conditions: past values must be good predictors of current outcomes without being affected by today's random events, and unexpected shocks should not create patterns that carry over from one year to the next. When these conditions are violated, the results can be misleading. |
| They can accommodate multiple policy variables, lagged effects and dynamic adjustment processes (e.g. gradual impact of investment support on productivity), which are common in CAP contexts. | They require reasonably long, well‑measured panels with limited missing data. Measurement error, short time series or high attrition (farms entering and exiting the panel) can seriously undermine performance. |
| Diff‑GMM and SYS‑GMM do not fully solve all causal inference problems; time‑varying unobserved factors that affect both support and outcomes can still bias results, so findings should be interpreted as ’improved but still assumption‑dependent’ rather than as definitive causal evidence |
When to use?
Diff‑GMM and SYS-GMM are used when we follow the same farms over time and when results today are strongly linked to the past results (for example, farm income today may be strongly influenced by farm income in the past, as higher income in the past may have allowed increased investments, resulting in increased profitability and thus higher income today).
SYS-GMM are generally preferred over Diff-GMM for CAP’s persistent income/support data.
Preconditions
To use Diff-GMM and SYS-GMM credibly in CAP evaluation, national evaluators need:
- Long panel data, especially the same farms observed (in general, for more than five years, better, more than 7-10). GMM requires longer T (time periods) and higher N (farms) than Fixed Effects, plus rigorous diagnostic testing.
- Highly persistent outcomes. For example, variables like farm income, productivity, where current values strongly depend on past values.
- To consider endogeneity, in particular ‘chicken-and-egg’ situations where CAP support and outcomes influence each other (e.g. profitable farms get different support)
- Advanced econometric expertise to perform dynamic panel estimation, instrument selection and interpret the battery of diagnostic tests used to assess the validity and quality of the model
- Detailed policy timing knowledge that is particularly relevant to justify lagged values as valid instruments and interpret short/long-run effects correctly.
Diff‑GMM and SYS‑GMM are not suitable for short panels (less than five years), small samples, non-persistent outcomes or when simpler fixed effects suffice.
When to use this technique in the context of the CAP Strategic Plan assessment
These methods are most appropriate when:
- a longitudinal micro‑dataset is available (typically five or more years of farm‑level data, ideally with many farms and several policy reforms or changes over time);
- the main outcome (e.g. farm income, productivity, technical efficiency) is highly persistent from year to year, so that including lagged outcomes is substantively important; and
- there are serious concerns about endogeneity (for example, more profitable farms may receive different levels of support, or adjust their behaviour in ways that affect both future support and outcomes).
- different ISI measures (e.g., direct payments, agri-environmental schemes, investment subsidies) must be disentangled, especially when they are implemented together.
- the analysis must accommodate lagged effects and dynamic adjustment processes, for example, when investment subsidies affect productivity gradually over several years, or when farms take time to adapt to new environmental requirements.
For national CAP evaluation, this means Diff‑GMM and SYS‑GMM are candidates when the evaluator has a sufficiently long FADN‑type panel and wishes to go beyond simple correlation to obtain more credible medium‑ and long‑run income or productivity effects of support.
Dynamic GMM methods have been used to:
- Estimate how direct payments and other CAP instruments affect farm productivity and technical change over time, controlling for farm‑specific unobserved characteristics and outcome persistence.
- Analyse the income transfer efficiency of different CAP support instruments, distinguishing short‑run and longer‑run effects on farm income.
- Study dynamic relationships between agricultural policy, trade performance or environmental outcomes (for example, exports, carbon emissions or technical efficiency trajectories) using multi‑year panel data.
Step-by-step
Step 1 – Start with a panel in which the same farms (or regions) are observed over several years. The dataset should contain:
- key outcomes (e.g. farm income, productivity, or environmental proxies),
- CAP support variables
- other factors that change over time (for example, herd size, crop mix, input prices, or simple weather indicators).
Step 2 – Specify a dynamic model by setting up an equation where the outcome today depends on:
- the outcome in previous years (to capture persistence),
- current CAP support and other time‑varying controls.
Farm‑specific and year‑specific effects are included so that permanent differences between farms and common shocks by year are absorbed rather than estimated directly.
Step 3 – Estimate with Diff‑GMM or SYS‑GMM:
- Diff‑GMM works with changes over time (differences) to remove fixed farm effects and uses past levels of variables as internal instruments.
- SYS‑GMM combines equations in differences and in levels, using both past levels and past changes as instruments, which is useful when income and support are very persistent.
The interest is in the coefficients on CAP support and other policy variables, not in the fixed effects themselves. In panel‑data models (like fixedeffects, Diff‑GMM, SYS‑GMM), fixed effects are included to control for all stable characteristics of each farm, i.e. things that do not change over time, such as soil type, altitude, long‑term management style, etc. These fixed effects are useful because they remove bias from constant differences between farms, but they are not policy‑relevant and cannot be interpreted.
Step 4 – After estimation, it is essential to verify that:
- the instruments are strong and valid;
- there is no serial correlation in the errors,
- the number of instruments is kept under control.
If needed, the specification (=step 2) can be refined (e.g. changing lags, simplifying the instrument set).
Step 5 – Results should be read as relationships that account for:
- dynamic effects (current’s outcome depending on past’s),
- unobserved, time‑invariant farm characteristics,
- and some of the “chicken‑and‑egg” feedback between support and performance.
Example:
Consider a model estimating the effect of CAP support on farm income. Suppose the coefficient for agri-environmental support (AES), denoted as ΔISI-AES (the year-to-year change in policy support), is 0.15. This means that a EUR 1 000 increase in AES is associated with a EUR 150 increase in farm income, after accounting for:
- the dynamic effect of past income on current income,
- all unobserved, time-invariant farm characteristics,
- part of the ‘chicken-and-egg’ feedback whereby higher-performing farms receive more support.
Because the GMM estimator explicitly deals with these sources of bias, this estimate is more credible than a simple correlation and more reliable than simple before-and-after comparisons, which would ignore persistence, selection patterns and other confounding factors.
Evaluators should still discuss possible remaining biases (for example, unobserved shocks that change over time) and, where possible, compare findings with simpler models and other causal approaches.
Main takeaway points
- Diff‑GMM and SYS‑GMM use longitudinal farm data and the farm’s own history to handle persistence and feedback between CAP support and outcomes.
- They are beneficial when outcome (e.g, income, productivity, technical efficiency) and policy support are both highly persistent over time, and when simple OLS or fixed-effects models are likely to be biased due to endogeneity (the ‘chicken-and-egg’ problem where support and outcomes influence each other simultaneously) or when lagged dependent variables are included as regressors.
- Their credibility depends strongly on good data, careful instrument choice, and thorough diagnostic testing, so they require experienced quantitative analysts.
Learning from practice
Farm Income and Transfer Efficiency
Biagini, L., Antonioli, F., Severini, S., (2020), The Role of the Common Agricultural Policy in Enhancing Farm Income: A Dynamic Panel Analysis Accounting for Farm Size in Italy, Journal of Agricultural Economics, 71(3), pp. 652-675.
This study uses SYSS-GMM on Italian FADN micro-data (2008–2014) to estimate the income transfer efficiency of CAP measures (decoupled payments, agri-environmental payments, on-farm investment subsidies, coupled payments). The dynamic panel approach accounts for endogeneity, simultaneity bias, and omitted variables. Results show that decoupled direct payments provide the highest contribution to farm income, followed by agri-environmental payments. Coupled payments have no significant impact. Large farms benefit from greater transfer efficiency than small farms.
Ciliberti, S., Frascarelli, A., (2019), The income effect of CAP subsidies: implications of distributional leakages for transfer efficiency in Italy, Bio-based and Applied Economics.
This paper applies an Arellano-Bond (Diff-GMM) linear dynamic panel-data estimation on Italian FADN data (2008-2014) to evaluate the income distributional effects of CAP subsidies, distinguishing between the single payment scheme, coupled payments, and Pillar II aids. While all major subsidy types show a statistically detectable income effect, the authors find that coupled payments are subject to more pronounced distributional leakages through factor markets than decoupled support, resulting in a lower net income gain per euro transferred. These findings are complementary with Biagini, Antonioli and Severini (2020): the significant coupled payment coefficient captures a gross income effect prior to leakage adjustment, whereas the absence of a significant efficiency effect in Biagini et al. (2020) reflects a stricter welfare-economic criterion applied through a more robust estimator.
Total factor productivity (TFP)
Biagini et al. (2023), The impact of CAP subsidies on the productivity of cereal farms in the EU, Food Policy.
This study applies a three-step estimation strategy culminating in a SYS-GMM estimator on FADN data from France, Germany, Italy, Poland, Spain, and the UK (2008-2018). It assesses the relationship between different types of CAP subsidies (coupled direct payments, decoupled direct payments, agri-environmental scheme payments, less-favoured area payments) and farm total factor productivity (TFP). Results confirm that most CAP subsidies have a negative-to-insignificant effect on TFP, except for agri-environmental subsidies, which can increase farms' productivity. The study compares results across countries and among farms with different productivity levels.
Mary, S., (2013), Assessing the Impacts of Pillar 1 and 2 Subsidies on TFP in French Crop Farms, Journal of Agricultural Economics, 64(1), pp. 133-144.
Mary applied a GMM technique to French FADN crop farm data, estimating a Cobb-Douglas production function and then using GMM regressions to link farm subsidies to TFP. This two-step approach addresses the fact that inputs are likely correlated with productivity shocks. Results show that set-aside, Less Favoured Areas payments, and livestock payments have a statistically negative effect on productivity.
Rizov, M., Pokrivcak, J., Ciaian, P., (2013), CAP Subsidies and Productivity of the EU Farms, Journal of Agricultural Economics, 64(3), pp. 537-557.
This influential paper investigates the impact of CAP subsidies on farm TFP across EU-15 countries using FADN data and a structural semi-parametric estimation algorithm combined with GMM regressions. The main finding is that subsidies impacted negatively on farm productivity before the decoupling reform; after decoupling, the effect became more nuanced and in several countries turned positive.
Further reading
- EU CAP Network (2024): Assessing the effectiveness and efficiency of CAP income support instruments
- Hansen, L. P., (1982). Large sample properties of generalized method of moments estimators. Econometrica: Journal of the econometric society, 1029-1054.