Institutional Changes and Social Preferences: Experimental Evidence from a Developing Economy

Authors
Affiliation

Andrew S. Griffen

Faculty of Economics, University of Tokyo

Yasuyuki Sawada

Faculty of Economics, University of Tokyo

Abstract

This paper estimates Fehr-Schmidt (FS) social preferences in rural Burkina Faso using data from a large-scale RCT of a school-based management institution combined with lab-in-the-field experiments on the public goods (PG) and hypothetical dictator (HD) games. We combine the structural model with the RCT data to try to understand the causal channels behind the experimental results. The FS model fits the data well and can rationalize both the differential impacts and contribution levels observed between the PG and HD games. The FS model suggests that the randomized introduction of the institution operates primarily through reductions in “envy” rather than decreases in “guilt” preferences and can explan the increased contributions in both lab-in-the-field experiments. The results imply complementary between democratic institutions and prosocial behavior. The villagers also appear to have non-W.E.I.R.D. social preferences that look quite different from the published (and globally non-representative) literature on structural FS estimates. Implications for the literature and policy are discussed.

Introduction

The goal of this paper is to understand the causal channels through which a RCT of school-based management (SBM) called the COGES program (Sawada et al. 2022) affected social preferences in rural Burkina-Faso1. School-based management is the name given to a range of school decentralization policies intended to increase accountability of schools in developing countries (Casey et al. 2012). In our RCT, we randomly assigned some villages the opportunity to elect community members to a COGES via community wide elections. The COGES members were then empowered to change policies within the schools using feedback from the broader community. Such decentralization policies aim to delegate more local control with the goal of improving school performance. The randomization of elections and institutions is rare and gives us an opportunity to learn how formal institutions interact with informal norms. In a companion paper, we indeed found large impacts of COGES on children’s human capital outcomes (Sawada et al. 2024) consistent with COGES’ intended goal to improve schools. We had also hypothesized that these new institutions might have broader impacts in the community, specifically through a social capital channel. To test this hypothesis, we measured villagers’ behavior in two lab-in-the-field experiments: the public goods game and the hypothetical dictator game. We found positive impacts on contributions in both games. This and other indirect evidence led us to interpret COGES as having increased social capital within the villages (Sawada et al. 2022).

An alternative interpretation of the experimental results is that COGES program may have directly affected villagers’ preferences. An influential literature on social preferences has been developed to rationalize non-Nash play in economic games as well as various form of prosociality observed in many human societies. A leading example is the model by Fehr and Schmidt (1999) (hereafter FS), which posits two forms of disutilty from inequity: disadvantageous inequality (“envy”) and advantageous inequality (“guilt”). Although it has been argued that such FS preferences are not able to rationalize some experimental results (Andreoni and Bernheim 2009; Charness and Rabin 2002; Daruvala 2010) and it is difficult to distinguish between various forms of prosocial preferences (e.g., Bolton and Ockenfels (2000)), an advantage of the FS model is it provides a parsimonious and tractable deviation from neoclassical preferences that can be easily taken to the data (Cooper and Kagel 2016). Structural estimation in behavioral economics is useful not only to disentangle and estimate mechanisms (DellaVigna 2018) but also provides parameter estimates that can be compared across studies (Heckman 2000). In fact, combining RCTs and structural estimation addresses exactly this tension between internal and external validity (Todd and Wolpin 2023). Although previous research has used RCTs to validate (Todd and Wolpin 2006) and to estimate (Attanasio et al. 2012) structural models, we take the view that our randomized introduction of elections and SBM institutions as potentially having changed the underlying preferences, which would be expressed in changes in play in the lab-in-the-field experiments. The formation of prosocial behavior (Bowles and Gintis 1998; Gintis 2003) as well as interactions between formal institutions and informal norms (North 1990; Ostrom 1990; Alesina and Giuliano 2015) have long been of interest to researchers, particularly in societies transitioning to more Western and modern institutions (Ensminger 1996). Our results are consistent with results finding crowding-in of prosocial behavior (Dell et al. 2018; Martinez-Bravo et al. 2017) and provides direct causal evidence on such channels.

Our estimated structural FS model reveals several findings. The first main result is that the RCT impact operated primarily through reductions in envy preferences. Intuitively this appears through increases in the right tail of contributions in the treatment group. However, the model counterfactuals allow a more precise quantitative result as we can turn off various parts of the guilt, envy, and belief parameters to decompose the RCT impact into changes in guilt and envy as well as changes in equilibrium beliefs. The program itself explains 25.2% of the impact in the PG game and the entirety of the impact in the dictator game. Changes in beliefs, which affect play only in the public goods game, explains a further 66.8% of the impact in the PG game. Our interpretation of the result is that the introduction of democratic institutions makes villagers more tolerant of disadvantageous inequality. In the public goods game specifically this moves the villagers toward a more efficient allocation, which suggests important complementarities between democratic institutions and efficiency. The villagers seem to already have high levels of the guilt parameter which causes almost no low levels of contributions. One interpretation is that such preferences reflect the presence of informal insurance arrangements that avoid low payoffs. Their evolved social preferences seem designed to deal with risk but that more formal institutions can help to realize additional efficiency gains. The estimated model also fits the distribution of contributions well and perhaps surprisingly the same RCT parameters can explain both different levels of contribution between PG and HD games as well as differential experimental impacts in the two games. These two results are a broader test of the FS model itself.

The second main finding is that the estimated envy and guilt parameters for the villagers in our data look quite different from Fehr-Schmidt’s proposed calibration of their model (Fehr and Schmidt 1999). In fact, our estimated parameters for disadvantageous and advantageous inequality lay outside the range of estimates from a meta-analysis of the 43 papers estimating FS preferences in the existing literature (Nunnari and Pozzi 2024). The vast majority of these papers are estimated on Western (86%) and student populations (74%). We interpret our result as evidence for non-W.E.I.R.D. preferences among our study population (Henrich et al. 2001; Henrich 2020). Like other psychological traits, FS preferences also perhaps look quite different in traditional societies, which has implications not only for the published literature but also policies that might be implemented in these communities.

RCT

In this section, we briefly describe the RCT that we have previously evaluated in Sawada et al. (2022) and Sawada et al. (2024). The COGES project was a joint endeavour by the government of Burkina Faso and JICA that aimed to increase the quality of and access to primary schools by giving villagers more control over decisions made within their local schools2. The mechanism was through the democratic election of COGES committee members and interchange between COGES and community members to determine policies (an “action plan”) that were important for the school. Common action plans include building female toilets, providing meals to students, and providing housing to teachers.

COGES is one example of school-based management, which is a part of a larger trend in development economics of community driven development (Fearon et al. 2009; Casey et al. 2012; Nguyen and Rieger 2017) in which decisions are delegated to local actors, sometimes through democratic mechanisms. Although school-based management is increasingly pushed by the development community, the track record is actually mixed as discussed by Sawada et al. (2022). Due to limitations in the number of schools that could be served in the first-year, COGES was set-up as a randomized roll-out in which approximately half of the schools were treated in the first-year and the remaining schools were treated from the second-year3. Conveniently for researchers, a randomized roll-out is often more politically viable when there are capacity constraints to scaling a program. The schools were located in villages within Ganzourgou province and randomization was done within strata defined by educational districts and school types (public schools, private Islamic schools, and private Catholic schools). From the second year onwards, both treatment and control schools had village-wide elections and then implemented action plans. Unfortunately this makes it impossible to assess the long-term cumulative impacts of the COGES program because the duration of RCT was for a single year.

Lab-in-the-field experiments

In both treatment and control villages, we conducted two lab-in-the-field experiments: the public goods game and the hypothetical dictator game. The experiments were conducted at both baseline and endline during which there were two rounds of the public goods game and a single survey question about the hypothetical dictator game at each time period. As explained in Sawada et al. (2022), the timing of the lab-in-the-field experiments was slightly unusual and affects the interpretation of the experimental impacts. So the first set of lab-in-the-field experiments were not conducted precisely at baseline but rather immediately after the first-year elections in the treatment villages. We call this difference between treated and control schools an “election effect”. The second set of lab-in-the-field experiments were conducted at the end of the first year after the second-year elections had been conducted in both treatment and control schools. We call this difference between treatment and control schools an “implementation effect” because it captures the effect of implementing the action plan during the first year4. In general, we only find experimental implementation effects so that will be our focus. We next introduce some general notation that will be used to describe the games and the FS model.

Game notation

Games \(g \in \{pg,hd\}\) are either public goods (\(pg\)) or hypothetical dictator (\(hd\)). In each game \(g\) there a finite number of players indexed by \(i = 1,...,n^g\). Let the contribution of player \(i\) in game \(g\) be \(c_i^g\). For the PG game, we also need to keep track of the vector of contributions for the \(n^{pg}\) players \(c^{pg} = \{c_1^{pg},...,c_n^{pg}\}\) as well as the vector of contributions without the \(i\)th player \(c_{-i}^{pg} = \{c_1^{pg},...,c_{i-1}^{pg},c_{i+1}^{pg},...,c_n^{pg}\}\). We use analogous notation for monetary payoffs with the payoff for player \(i\) \(x_i^g\), the vector of payoffs \(x^g\), and without the \(i\)th player \(x_{-i}^g\). In the implementation of the lab-in-the-field experiments, the endowment was set to \(e = 500\) FCFA in each game, which is approximately the wage for one day of labor in rural Burkina Faso. For each game, the choices were discrete and restricted to the set \(c_i^g \in \C = \{0,100,200,300,400,500\}\). For the public goods game there were 4 players while for the hypothetical dictator game there were 2 “players” (the dictator and the hypothetical player). So we have \(n^{pg} = 4\) and \(n^{hd} = 2\). At both baseline and endline, the players played the PG game twice and the hypothetical dictator game once so for a typical player we will have 6 observations of play. However, because of attrition and sample replenishment, participants sometimes played more or fewer times.

Public goods game

In our public goods game, each player \(i\) made an anonymous contribution \(c_i^{pg}\) from their endowment and the contributions were then summed, doubled, and divided equally among the \(n^{pg}\) group members. In addition to their share of the summed contributions, each player kept any amount they did not contribute. Given the players’ contribution vector \(c^{pg}\), player \(i\)’s monetary payoff is:

\[ x_i^{pg}(c^{pg}) = 500 - c_i^{pg} + \tfrac{2}{4}\sum_{j=1}^{4} c_{j}^{pg}. \] The player keeps \(500 - c_i^{pg}\) with certainty while the second term depends on the (unknown) behavior of the other players. We write the payoff \(x_i^{pg}(c^{pg})\) as a function of the entire vector of contributions. Notice that because \(\tfrac{d x_i^{pg}}{d c_i^{pg}} = - 1 + \tfrac{2}{4} < 0\) each player’s monetary payoff is a decreasing function of their contribution regardless of the contributions of the other players. Under the assumption that \(i\)’s utility only depends on their own payoff \(c_i^{pg} = 0 \,\, \forall i\) is a dominant strategy Nash equilibrium.

Hypothetical dictator game

The second game that participants played was a hypothetical dictator game. Again from an initial endowment of 500, player \(i\) was asked to contribute a hypothetical amount \(c_i^{hd}\) to another randomly chosen player \(j\) in the public goods game group. Given the contribution \(c_i^{hd}\), the corresponding payoffs are \(x_i^{hd} = 500 - c_i^{hd}\) and \(x_j^{hd} = c_i^{hd}\). Compared to the public goods game there is no payoff uncertainty so the payoff vector can be written as \(x^{hd}(c_i^{hd})\) because it depends only on player \(i\)’s choice.

Fehr-Schmidt preferences

With FS preferences, player \(i\)’s utility depends not only their own payoff \(x_i\) but also on how \(x_i\) compares to the other players’ payoffs \(x_{-i}\) according to5: \[ u_i(x^g(c^g)) = x_i^g - \alpha_i \tfrac{1}{n^g-1}\sum_{j\neq i} \text{max}\{x_j^g - x_i^g,0\} - \beta_i \tfrac{1}{n^g-1}\sum_{j\neq i}\text{max}\{x_i^g - x_j^g,0\} \] Disadvantageous inequality \(x_j - x_i > 0\) reduces utility according to the degree of the player \(i\)’s “envy” parameter \(\alpha_i\). Player \(i\)’s utility can also be reduced through advantageous inequality when \(x_i - x_j > 0\). Player \(i\) feels “guilt” according to the value of \(\beta_i\). In principle, these parameters can also be negative so that the player would enjoy others having more than them (negative \(\alpha_i\)) or enjoy having more than other players (negative \(\beta_i\)). Fehr and Schmidt (1999) argue negative guilt is unlikely but negative envy is plausibly related to potlach type behavior.

To take the model to the data, we make the following modifications: \[ u_i(x^g(c^g)) = (x_{i}^g - \alpha_{i}\tfrac{1}{n^g-1}\sum_{j\neq i}\text{max} \{x_j^g - x_i^g,0 \} - \beta_{i}\tfrac{1}{n^g-1}\sum_{j\neq i}\text{max} \{x_i^g - x_j^g, 0 \})/\lambda + \epsilon_{ic} \]

First, we add type-I extreme value shocks \(\epsilon_{ic}\) that are assumed to be independent across individuals and choices. Second, the parameter \(\lambda\) sets the overall scale of the utility over monetary payoffs relative to the shocks à la McKelvey and Palfrey (1995). We assume the individual heterogeneity is distributed according to:

\[ \begin{pmatrix} \alpha_{i} \\ \beta_{i} \end{pmatrix} \sim \N\left( \begin{pmatrix} \alpha(s_i) \\ \beta(s_i) \end{pmatrix}, \begin{bmatrix} \sigma_{\alpha}^2 & \rho\sigma_{\alpha}\sigma_{\beta} \\ \rho\sigma_{\alpha}\sigma_{\beta} & \sigma_{\beta}^2 \end{bmatrix} \right) \] where \(\alpha(s_i)\) and \(\beta(s_i)\) are indices of player \(i\)’s observable characteristics \(s_i\): \[ \begin{aligned} \alpha(s_i) = & \, \alpha_0 + \alpha_1 \text{COGES}_i \mathbb{1}_{\{t = 1\}} + \alpha_2 \text{COGES}_i \mathbb{1}_{\{t = 2\}} + \alpha_3 \mathbb{1}_{\{t = 2\}} + \alpha_4 \mathbb{1}_{\{r = 2\}} + \\ & \alpha_5 \text{age}_i + \alpha_6 \text{educ}_i + \alpha_7 \text{male}_i + \alpha_8 \text{muslim}_i + \alpha_9 \text{public}_i \end{aligned} \] where \(\text{COGES}_i\) is an indicator for treatment, time is \(t \in {1,2}\) denotes time (baseline or endline), and \(r \in {1,2}\) for round (for the public goods game). The remainding variables cover the individual observables of gender, education, and age as well as indicators for the type of school. The index \(\beta(s_i)\) is defined analogously. These indices are where randomized exposure to the COGES program enters, which can shift the means of the alpha-beta distribution by exposure to the experiment to uncover the experimental mechanisms. Particular focus will be on \(\alpha_0\) and \(\beta_0\) that capture intrinsic levels of envy and guilt and the parameters that control exposure to implementation effect of the COGES RCT \(\alpha_2\) and \(\beta_2\). As a placebo effect, we should also expect not to find any impacts of \(\alpha_1\) and \(\beta_1\).

We assume both \(s_i\) and the distributions themselves are all known by other players but that the realized values of \(\alpha_i\) and \(\beta_i\) are \(i\)’s private information. The \(\alpha_i\) and \(\beta_i\) are time and round invariant which will induce persistence across games. The individual heterogeneity creates a mixture so the model resembles a flexible mixed-logit (McFadden and Train 2000). In principle the identification of heterogeneity in a mixed logit only requires a single decision for each individual (Hess and Train 2011) although it is a common intuition that multiple observations helps distinguish within from between individual correlation of choices. Note as well that the model with heterogeneity nests the homogenous model (\(\sigma_{\alpha} = \sigma_{\beta} = \rho = 0\)) as a special case.

Beliefs

In the public goods game, the players face uncertainty over the others players’ choices. Consequently, the players need to form beliefs about the choices of the other players. So we next introduce notation for beliefs. Let the belief over the other players contributions \(c_{it}\) be given by: \[ \pi(c_{-i}|s) = \prod_{j\neq i} \pi(c_{j}|s_j). \] The product follows from the independence of the shocks across players given the state space. To implement the procedure, we first estimate \(\pi(c_{j}|s_j)\) and then construct \(\pi(c_{-i}|s)\) for each player \(i\) in each game they play with state space \(s\). This also introduces some parsimony into the model in that the individual contribution probabilities depend on \(s_j\) and not \(s\). Given this probability distribution over \(c_{-i}\) , the expected utility for player \(i\) of making choice \(c_{i}\) is \[ \displaystyle\sum \pi(c_{-i}|s) u_i(x(c_i, c_{-i})) = v_i(c_{i}). \] Note that this object depends only on the choice \(c_i\) so that given beliefs, the expected payoffs, envy, and guilt can all be estimated directly in the data and the game can be recast into a standard discrete choice framework (Rust 1996).

Estimation

The estimation is done in two steps. First, we estimate beliefs for the public goods game using an unrestricted multinomial logit. The estimated beliefs are then used to construct the expected payoff, envy, and guilt terms for the public goods game which are inputs into the structural model. The structural model is then estimated using simulated maximum likelihood. To account for the estimated terms, we bootstrap the standard errors.

Belief parameters

To estimate the beliefs about the play of other players in the public goods game group, we follow Rust (1996) and estimate the following multinomial logit models for public goods game contribution choice6. \[ \begin{aligned} \pi(c_j^{pg} | s_j) = \frac{exp(\mu_{c}(s_j))}{\sum\limits_{c \in \C}exp(\mu_{c}(s_j))}, \end{aligned} \] where \(\mu_{c}(s_j)\) is an index of the observable characterstics of the players. We impose a normalization that \(\mu_{0}(s_j) = 0\). Given the estimated parameters \(\hat{\mu}_{c}(s_j)\) we can constuct \(\hat{\pi}(c_j^{pg}|s_j)\) and then \(\hat{\pi}(c_{-i}|s) = \prod_{j\neq i} \hat{\pi}(c_{j}|s_j)\) which is used to integrate out the choices for the other players.

Envy and guilt terms

Using the predicted probabilities from the multinomial logit, we can then compute the expected payoff, envy, and guilt terms to add to the structural model7. The expected envy term for player \(i\) given choice \(c_i\) is \[ \begin{aligned}\label{expected-envy} \sum_{c_{-i} \in \C_{-i}}\hat{\pi}(c_{-i}|s)\tfrac{1}{n^g-1}\sum_{j\neq i}\textup{max}\{x_j(c_i, c_{-i}) - x_i(c_i, c_{-i}),0\}, \end{aligned} \] and the expected guilt term is \[ \begin{aligned}\label{expected-guilt} \sum_{c_{-i} \in \C_{-i}}\hat{\pi}(c_{-i}|s)\tfrac{1}{n^g-1}\sum_{j\neq i}\textup{max}\{x_i(c_i, c_{-i})- x_j(c_i, c_{-i}),0\}. \end{aligned}. \]

Likelihood function

Let \(\K_i\) be the sequence of games played by player \(i\) with element \(k\) and let \(d_{ic}^{k} = 1\) if player \(i\) made choice \(c\) in game \(k\) with \(\sum_{c \in \C} d_{ic}^{k} = 1\). We can write the probability of the observed sequence of choices for player \(i\) as:

\[ p_i = \int \underset{k \in \K_i}{\prod} \underset{c \in \C}{\prod} Pr(d_{ic}^{k} = 1 | \alpha_i, \beta_i, \hat{\pi}_{i}^{k}, \lambda)^{d_{ic}^{k}} f(\alpha_i, \beta_i | \alpha(s_i), \beta(s_i), \sigma_\alpha, \sigma_\beta, \rho) d\alpha_i d\beta_i \] where the the individual specific parameters, beliefs, and \(\lambda\) paramater determine the probability of the sequence of choices. The choice probabilities are assumed to be independent conditional on these values, which follows from the independence assumption on the type-I extreme value errors. We add \(\hat{\pi}_{i}^{k}\) to the conditioning to denote that the estimated beliefs are used to compute expected payoffs, guilt, and envy, which then affect the structural choice probabilities. The integration is then done over a bivariate normal mixing distribution. To implement the procedure, we need to numerically integrate so we replace the integral with a draw \((\alpha_i^r, \beta_i^r)\) from the bivariate normal given the current guess of the parameters \((\alpha(s_i), \beta(s_i), \sigma_\alpha, \sigma_\beta, \rho)\): \[ \begin{aligned} & p_i = \frac{1}{R}\sum_{r=1}^{R} \underset{k \in \K_i}{\prod} \underset{c \in \C}{\prod} Pr(d_{ic}^{k} = 1 | \alpha_i^r, \beta_i^r, \lambda)^{d_{ic}^{k}}. \end{aligned} \]

The likelihood is then the log sum of the individual specific likelihoods \(\mathcal{L} = \sum_{i} log(p_i)\), which we can use to find maximum likelihood estimates of the structural parameters. Initial calibration of a simplified version of the model suggested small values of \(\alpha_0\) and large values of \(\beta_0\) as well as reasonably large values of the \(\lambda\) to capture within-individual variation in contributions resulted in a good initial fit. In addition, to our hand calibrated values, we also tried simulating annealing to find other possible starting values as well as starting the model at the mean estimates reported in Nunnari and Pozzi (2024). These produced similar final parameter estimates. Apart from the intercepts, the observed heterogeneity parameters were all initialized to zero and the unobserved heterogeneity estimates were set to small values. Using analytic gradients was critical for numerical stability of the estimation given that the FS model has many flat spots that numerical gradients struggle with. To account for the estimated \(\hat{\pi}_{i}^{k}\) in the likelihood function, we bootstrapped the standard errors by repeatedly re-drawing the estimation sample, re-estimating the beliefs, and then re-estimating the structural parameters from the same initial starting value.

Data

Table 1 displays summary statistics from the data. The average contribution in the public goods game was 321 FCFA and 279 FCFA in the hypothetical dictator game. Clearly the data will reject a model without some form of other-regarding or social preferences. These totals represent 64.2% and 55.8% of the 500 FCFA endowment respectively and are on the high end of the range of contributions observed in these type of field experiments.

The average years of education is only 2.48 years and a majority (67.8%) of participants have 0 years of education. The average age is 40.5 years, most are male (57.6%), and most have children in public schools (68.2%).

Table 2 replicates the experimental impacts of Sawada et al. (2022) where the main finding was that COGES affected contributions at endline.8,

There were several interesting features of the data and experimental results. First, the absolute level of contributions is large in both games. Participants contributed between 53.2% and 66.6% of the 500 FCFA endowments. In meta-analyses of the literature show that the average contribution is 37.7% of the endowment in the public goods game (Zelmer 2003) and 28% in the dictator game (Engel 2011). So the data from our participants are definitely on the high end of observed contributions.

Second, COGES had large impacts on both public goods game and dictator game contributions. As explained in Sawada et al. (2022), the “baseline” data represent impact of the COGES elections and the endline impacts represent impacts of implementation of the COGES9. A new reported finding is that COGES also had impacts on the hypothetical dictator game contributions. At endline the raw impact of 31.0 represents a 9.31 percentage point (pp) increase relative to the control group. For the public goods game, Ledyard (1995) reports a range of 40-60% of the endowment so 9.31 pp increase is large and moves approximately half-way through the range of observed contributions. The impact on the hypothetical dicator game is 16.3 which is 52.6% of the impact on the public goods game. So clearly the experiment had differential impacts on these two measures.

In figure 1 above, we can also look at the distribution of contributions by treatment status. A very striking feature of the distribution is there are almost no contributions of zero in both games. This is very unusual compared to existing studies. In most public goods games experiments a fair amount of free-riding is common. For example, Andreoni (1988) reports that 34.3% of participants free-ride10. In the seminal paper on dictator games, Forsythe et al. (1994) report between 10-20% of zero contributions11.

Figure 1 also foreshadows the identification of the model. Zero contributions suggests there should be a large value of guilt because if individuals contribute 0, then they will almost certainly end up with a higher payoff than the other player(s), particularly in the hypothetical dictator game because of the lack of payoff uncertainty. At the same time some participants contribute everything. This indicates that there should be some individuals with very low or even negative values of envy. Contributing everything will guarantee that other players end up with a higher payoff, which indicates that some players are not bothered by this or even get positive utility through having negative envy.

To give a sense of the model behavior estimates, we can examine the behavior of a very stylized version of the model, \[ c_i^*(\alpha, \beta) = \underset{c_i}{\text{argmax}} \,\, x_i - \alpha \tfrac{1}{n^g-1}\sum_{j\neq i}\text{max} \{x_j - x_i,0 \} - \beta\tfrac{1}{n^g-1}\sum_{j\neq i}\text{max} \{x_i- x_j,0 \} \] which shows how the optimal choice varies for different values of \(\alpha\) and \(\beta\). For the public goods game, we take the average observed play in the data to form beliefs and take expectations over \(c_{-i}\).

Figure 2 reveals several interesting patterns related to the identification. First, negative values of \(\alpha\) are required to generate contribution of the entire endowment, especially in the hypothetical dictator game. The intuition is that you must have positive utility (negative \(\alpha\)) in order to generate behavior where you know others will have larger payoff than you. In the public goods game, the required parameters for such behavior are not as extreme because even for less negative values of \(\alpha\) you may still contribute all of your endowment if you think it is likely others are contributing their full endowments. This where uncertainty can amplify the contributions in the public goods game, which is indeed what we see in the data. But overall envy parameters should be small and negative for some of the individuals. The second feature is that guilt must be large. In the dictator game, in order to avoid 0 contributions, for most values of \(\alpha\) suggests a guilt parameter \(\beta\) larger than 0.5.

Although it seems that for a given parameter distribution the average contribution of the public goods game would be larger, it is unclear whether the same deep parameter distribution could explain the distribution of contribution in both games. So that is one test of the model. There is also within variation in contribution for the same players, which is handled by the Type-I extreme value shocks and the scaling parameter. So we need to estimate the full version of the model to check these issues.

In figure 3 above, we can look at the joint distribution of contributions by examining a jittered scatter plot. The first panel displays the public goods game in round 2 versus round 1. The second panel pools the public goods game and looks at the public goods games vs. hypothetical dictator games contributions. The positive correlation is evident, which shows the consistency of choices by players and is suggestive of a distribution of envy and guilt in the population. The correlation between the public goods game contribution in round 1 and round 2 is 0.60. You can also see the increase in round 2 relative to round 1 as there is more weight towards the upper left quadrant. There is a postive but smaller correlation of 0.36 for the second panel. You can also see the larger average contribution of the public goods game.

So overall the challenge for taking the FS model to our data will be for it to explain the observed distributions of the contributions, the higher mean contribution for the public goods games, the high within-person correlation across games, and the larger impact estimate of the public goods compared to the hypothetical dictator game. A key difference between the two games is that the public goods game has uncertainty over the other players’ actions. As is standard for micro data and structural models, the specification of heterogeneity will be essential for fitting the model to the data (Heckman 2001), which has been shown specifically to be important for estimating FS models (Bellemare 2023).

Envy and guilt histograms

For the public goods game, variation in \(s\) across players and groups will induce variation in \(\hat{p}(c_{-i}|s)\) which will vary the expected envy and guilt terms across individuals given the same choice of \(c_i\). There will also be within-group variation because each individual player will face a different set of players. Of course, how the individuals value the guilt and envy terms relative to their expected monetary payoffs will also vary according to the parameters of the structural model through their individual specific \(\alpha_{i}\) and \(\beta_{i}\) parameters. Figure 4 displays these expected envy and guilt frequency histograms for the public goods game using beliefs estimated in the data. The top and bottom panels refer to envy and guilt respectively and contributions go from 0 to 500 moving left to right.

The simplest case is for a contribution \(c_i = 0\), which results in an envy term of 0 because the payoff is weakly greater than all group members. Subsequent increases in contributions result in greater average envy as increased contributions increase the likelihood of a lower payoff than other group members. However there is still considerable variation in expected envy across groups. Expected guilt displays a reversed pattern with 0 guilt for a contribution of \(c_i = 500\) as the payoff is weakly less than all other group members. Decreasing contributions increases the expected guilt but again there is substantial variation across individuals because of variation in group configurations induced by variation in \(s\) both within and across groups.

In figure 5 above, we can repeat the exercise for the hypothetical dictator game. Because there is no payoff uncertainty in the hypothetical dictator game the mass points are degenerate. The envy term is 0 for contributions \(c_i \in \{0,100,200\}\) because \(i\)’s own payoff \(500 - c_i\) will exceed the payoff \(c_i\) for the hypothetical receiver. Similarly the guilt term is 0 for contributions \(c_i \in \{400,500\}\) because the receiver’s payoff \(x_j = c_i\) will exceed \(i\)’s own payoff \(x_i = 500 - c_i\). For other contributions, guilt and envy move in opposite directions as increasing contribution increases envy and decreases guilt.

Parameters and model fit

The structural model parameters are shown in Table 4 above. There are two main findings of interest. The first is how the implementation of the COGES program affected the \(\hat{\alpha}_2\) and \(\hat{\beta}_2\) parameters. The magnitude of the estimates suggests that most of the COGES impact operates through decreases in the envy parameter \(\hat{\alpha}_2 =\) -0.030 rather than changes in the guilt parameter \(\hat{\beta}_2 =\) -0.002. Not surprisingly the election effects \(\hat{\beta}_1 =\) 0.001. and \(\hat{\alpha}_1 =\) -0.016 appear small because no such election effect exists in the reduced form. The relevant rows are highlighted in green.

However, it is difficult to compare the magnitude of the estimates without simulating the model, which will be the subject of the next section. But the fact that implementation effects seems to operate primarily through changes in envy rather guilt is interesting. The reason appears to be that, although reductions in envy preferences would also increase contributions, they would have resulted in a different pattern of contributions in the hypothetical dictator game (more right tail contributions), which are not as prominent. However, because of the uncertainty in the public goods game, increases in guilt can increase the right tail contributions in the public goods game, which is what is observed in the data. So this results in a better fit of the contribution distributions by treatment and control across the two games.

The second finding is the values of \(\hat{\alpha}_0\) and \(\hat{\beta}_0\) are quite different from the existing literature. This relates to our non-W.E.I.R.D. interpretation. Our estimate of -0.22 is negative and much smaller than the average calibrated value of 0.35 reported by Fehr and Schmidt (1999) or the average value of 0.85 from a meta-analysis of the literature (Nunnari and Pozzi 2024). Similarly our estimated value of \(\beta\) is 0.75, which is again much different than the value of 0.32 reported by both Fehr and Schmidt (1999) and Nunnari and Pozzi (2024).

From the observed contribution distributions in the Burkina-Faso data, we see that there are almost no contributions of zero, which indicates that levels of guilt need to be extremely high. Contributing zero would guarantee a higher level of payoff, which our participants completely avoid, so the estimation sets \(\hat{\beta}_0\) high to fit the data. At the other end, the mass of individuals who contribute their entire endowment of 500 ensure that others will have a larger payoff than them. To rationalize such behavior there need to be individuals who actually derive utility from having a lower payoff. This explains the negative envy that we estimate.

We think that The large difference that estimate is suggestive of a possible W.E.I.R.D. bias in the existing literature (Henrich et al. 2010b; Henrich 2020). In the paper surveyed by Nunnari and Pozzi (2024), out of 43 published / working papers estimating FS preferences 32 used college students (74%) and 37 were from Western populations (86%). The study populations came from China, Denmark, France, Germany, Italy, Netherlands, Spain, Sweden, Switzerland, Turkey, UK, and US with none from Africa or from traditional societies: this is the first paper. This is consistent with the finding that Henrich et al. (2010b) 96% of “standard subjects” come from societies representing 12% of the world population and they argue that these societies are psychologically peculiar: impersonal, individualistic, weaker kinship ties, etc. (Henrich 2020)

Other parameters of interest show that envy and guilt are both increasing in age and education. Being male is associated with higher envy but no change in guilt, which may have some evolutionary explanation related to status (Ridley 1994). There is also substantially more between than within individual variation in envy and guilt, which is needed to explanation the persistence of contributions and suggests that preferences are stable.

Model fit is shown in Figure 6. Qualitatively the model fit is good. It replicates the shape of the distribution with few low contributions in both the public goods and hypothetical dictator games, a peak at contributions between 200 to 300, and then another peak at contributions of 500. In addition, such structural models are often quite parsimonious relative to unrestricted statistical models (Keane and Wolpin 1997). For example, this model does not have choice specific intercepts, which could be used to trivially improve the model fit but which would lack a clear intrepretation.

In table 5, we can also compare the impact estimate computed in data simulated from our model to the impact estmates from the RCT. Because there were only statistically significant impact estimates at endline (“implementation effects”), we focus on the those impacts. Qualitatively the impacts are close and we alo fail to reject the null that the impact estimates are the same.

RCT Decomposition

The decomposition of the RCT impacts are shown in table 6. The first row shows the impact estimates from the estimated model. We then repeatedly turn off different parameters and re-compute the implementation effects. The goal is to understand how the RCT affected the parameters of intesest in the FS model and quantitatively how those can be used to understand the RCT impact.

The counterfactual \(\hat{\alpha}_2 = 0\) removes the impact estimates due to changes in envy. The impact in the public goods game falls to 21.7. So we conclude that changes in the envy parameter explain 25.2% of the impact of the RCT. Decreases in envy also explain basically the entirety of the impacts on the dictator game with the impact falling almost to zero: -2.18. The impacts through changes in the guilt parameter (\(\hat{\beta}_2 = 0\)) is quite small with the impacts in the PG 29.8 and HD 11.1 both basically unchanged. This is also confirmed when examining \(\hat{\alpha}_2 = \hat{\beta}_2 = 0\) as the impacts are similar to \(\hat{\alpha}_2 = 0\) suggesting not much interaction between the two channels.

The counterfactual \(\hat{\mu}_{COGES} = 0\) imagines if the players’ preferences have changed but they do not believe that other players’ preferences have changed as expressed in their beliefs. This counterfactual obviously only affects the PG game. The impact estimate falls to only 9.62. The idea of this counterfactual is to isolate effects in preferences only independently have changes through thinking that others’ preferences are different. For example, suppose (as our estimates imply) that the RCT reduces envy. So in order to maximize your utility you would likely contribute more because you no longer feel envious that other have more than you. But if you also perceive that others’ envy has also decreased then they likely will also contribute more, which would then nullify the increased utility you get from others’ having more than you (e.g., if you had negative utility). To the extent that you are aware of this (through your beliefs), you would contribute even more. It is through this mechanism that it seems the same change in preference parameter can explain the differential impacts we see in the PG and HD games. As a final exercise, we set all possible channels to zero (\(\hat{\alpha}_2 = \hat{\beta}_2 = \hat{\mu}_{COGES} = 0\)) the model predicts close to a zero impact 1.44 as expected.

Conclusion

This paper estimates the effect that the experimental introduction of a school-based management program called COGES had on social preferences in rural Burkina Faso. The program can be viewed as the randomized introduction of both democratic elections and decentralized institutions related to the schools. The RCT showed positive impacts on both human and social capital, which highlights the importance of institutions on such outcomes. Estimating a model of social preference allows us to map the lab-in-the-field experiments into structrual parameters that can be compared to existing studies. We found that the COGES program primarily operated through decreases in envy which suggests important complementarity between informal and formal institutions that can increase efficiency. We also found preference estimates that differ markedly from the existing literature with much smaller envy and much larger guilt parameters. This is plausibly a W.E.I.R.D. effect (Henrich et al. 2010a) given the convenience samples used in much of the existing literature (Nunnari and Pozzi 2024). An interesting direction for future research would be to understand the aspects of the economic and social lives of the villagers that give rise to such preferences (Henrich et al. 2001). The FS model itself, despite its parsimony, provides a quite good qualitative fit to the data and has no problem generating most of the major empirical features of the data and RCT: avoidance of zero contribution, larger average PG contributions, larger experimental PG impacts, and a mass of contribution of the entire endowment. This an interesting finding as well that such a parsimonious model can still fit data from such a different context.

References

Alesina, Alberto, and Paola Giuliano. 2015. “Culture and Institutions.” Journal of Economic Literature 53 (4): 898–944.
Andreoni, James. 1988. “Why Free Ride?: Strategies and Learning in Public Goods Experiments.” Journal of Public Economics 37 (3): 291–304.
Andreoni, James, and B Douglas Bernheim. 2009. “Social Image and the 50–50 Norm: A Theoretical and Experimental Analysis of Audience Effects.” Econometrica 77 (5): 1607–36.
Attanasio, Orazio P, Costas Meghir, and Ana Santiago. 2012. “Education Choices in Mexico: Using a Structural Model and a Randomized Experiment to Evaluate Progresa.” The Review of Economic Studies 79 (1): 37–66.
Bajari, Patrick, Han Hong, John Krainer, and Denis Nekipelov. 2010. “Estimating Static Models of Strategic Interactions.” Journal of Business & Economic Statistics 28 (4): 469–82.
Bajari, Patrick, Han Hong, and Stephen P Ryan. 2010. “Identification and Estimation of a Discrete Game of Complete Information.” Econometrica 78 (5): 1529–68.
Bellemare, Charles. 2023. Estimation of Structural Models Using Experimental Data from the Lab and the Field. Cambridge University Press.
Bolton, Gary E, and Axel Ockenfels. 2000. “ERC: A Theory of Equity, Reciprocity, and Competition.” American Economic Review 91 (1): 166–93.
Bowles, Samuel, and Herbert Gintis. 1998. “The Moral Economy of Communities: Structured Populations and the Evolution of Pro-Social Norms.” Evolution and Human Behavior 19 (1): 3–25.
Casey, Katherine, Rachel Glennerster, and Edward Miguel. 2012. “Reshaping Institutions: Evidence on Aid Impacts Using a Preanalysis Plan.” The Quarterly Journal of Economics 127 (4): 1755–812.
Charness, Gary, and Matthew Rabin. 2002. “Understanding Social Preferences with Simple Tests.” The Quarterly Journal of Economics 117 (3): 817–69.
Cooper, David J, and John H Kagel. 2016. “Other-Regarding Preferences.” The Handbook of Experimental Economics 2: 217.
Daruvala, Dinky. 2010. “Would the Right Social Preference Model Please Stand Up!” Journal of Economic Behavior & Organization 73 (2): 199–208.
Dell, Melissa, Nathan Lane, and Pablo Querubin. 2018. “The Historical State, Local Collective Action, and Economic Development in Vietnam.” Econometrica 86 (6): 2083–121.
DellaVigna, Stefano. 2018. “Structural Behavioral Economics.” In Handbook of Behavioral Economics: Applications and Foundations 1, vol. 1. Elsevier.
Engel, Christoph. 2011. “Dictator Games: A Meta Study.” Experimental Economics 14 (4): 583–610.
Ensminger, Jean. 1996. Making a Market: The Institutional Transformation of an African Society. Cambridge University Press.
Fearon, James D, Macartan Humphreys, and Jeremy M Weinstein. 2009. “Can Development Aid Contribute to Social Cohesion After Civil War? Evidence from a Field Experiment in Post-Conflict Liberia.” American Economic Review 99 (2): 287–91.
Fehr, Ernst, and Simon Gächter. 2000. “Cooperation and Punishment in Public Goods Experiments.” American Economic Review 90 (4): 980–94.
Fehr, Ernst, and Klaus M Schmidt. 1999. “A Theory of Fairness, Competition, and Cooperation.” The Quarterly Journal of Economics 114 (3): 817–68.
Forsythe, Robert, Joel L Horowitz, Nathan E Savin, and Martin Sefton. 1994. “Fairness in Simple Bargaining Experiments.” Games and Economic Behavior 6 (3): 347–69.
Gintis, Herbert. 2003. “Solving the Puzzle of Prosociality.” Rationality and Society 15 (2): 155–87.
Heckman, James J. 2000. “Causal Parameters and Policy Analysis in Economics: A Twentieth Century Retrospective.” The Quarterly Journal of Economics 115 (1): 45–97.
Heckman, James J. 2001. “Micro Data, Heterogeneity, and the Evaluation of Public Policy: Nobel Lecture.” Journal of Political Economy 109 (4): 673–748.
Henrich, Joseph. 2020. The WEIRDest People in the World: How the West Became Psychologically Peculiar and Particularly Prosperous. Farrar, Straus; Giroux.
Henrich, Joseph, Robert Boyd, Samuel Bowles, et al. 2001. “In Search of Homo Economicus: Behavioral Experiments in 15 Small-Scale Societies.” American Economic Review 91 (2): 73–78.
Henrich, Joseph, Steven J Heine, and Ara Norenzayan. 2010a. “Most People Are Not WEIRD.” Nature 466 (7302): 29–29.
Henrich, Joseph, Steven J Heine, and Ara Norenzayan. 2010b. “The Weirdest People in the World?” Behavioral and Brain Sciences 33 (2-3): 61–83.
Hess, Stephane, and Kenneth E Train. 2011. “Recovery of Inter-and Intra-Personal Heterogeneity Using Mixed Logit Models.” Transportation Research Part B: Methodological 45 (7): 973–90.
Keane, Michael P, and Kenneth I Wolpin. 1997. “The Career Decisions of Young Men.” Journal of Political Economy 105 (3): 473–522.
Ledyard, John O. 1995. “Public Goods: A Survey of Experimental Research.” In Handbook of Experimental Economics, edited by J. Kagel and A. Roth.
Martinez-Bravo, Monica, Gerard Padró i Miquel, Nancy Qian, Yiqing Xu, and Yang Yao. 2017. “Making Democracy Work: Formal Institutions and Culture in Rural China.” Unpublished Manuscript.
McFadden, Daniel, and Kenneth Train. 2000. “Mixed MNL Models for Discrete Response.” Journal of Applied Econometrics, 447–70.
McKelvey, Richard D, and Thomas R Palfrey. 1995. “Quantal Response Equilibria for Normal Form Games.” Games and Economic Behavior 10 (1): 6–38.
Nguyen, Tu Chi, and Matthias Rieger. 2017. “Community-Driven Development and Social Capital: Evidence from Morocco.” World Development 91: 28–52.
North, Douglass C. 1990. Institutions, Institutional Change and Economic Performance. Cambridge university press.
Nunnari, Salvatore, and Massimiliano Pozzi. 2024. Meta-Analysis of Distributional Preferences.
Ostrom, Elinor. 1990. Governing the Commons: The Evolution of Institutions for Collective Action. Cambridge university press.
Ridley, Matt. 1994. The Red Queen: Sex and the Evolution of Human Nature. Penguin UK.
Rust, John. 1996. “Estimation of Dynamic Structural Models, Problems and Prospects: Discrete Decision Processes.” Advances in Econometrics: Sixth World Congress 2: 119–70.
Sawada, Yasuyuki, Takeshi Aida, Andrew S Griffen, et al. 2024. Democracy and Human Capital: Experimental Evidence from Burkina Faso.
Sawada, Yasuyuki, Takeshi Aida, Andrew S Griffen, Eiji Kozuka, Haruko Noguchi, and Yasuyuki Todo. 2022. “Democratic Institutions and Social Capital: Experimental Evidence on School-Based Management from a Developing Country.” Journal of Economic Behavior & Organization 198: 267–79.
Todd, Petra E, and Kenneth I Wolpin. 2006. “Assessing the Impact of a School Subsidy Program in Mexico: Using a Social Experiment to Validate a Dynamic Behavioral Model of Child Schooling and Fertility.” American Economic Review 96 (5): 1384–417.
Todd, Petra E, and Kenneth I Wolpin. 2023. “The Best of Both Worlds: Combining Randomized Controlled Trials with Structural Modeling.” Journal of Economic Literature 61 (1): 41–85.
Zelmer, Jennifer. 2003. “Linear Public Goods Experiments: A Meta-Analysis.” Experimental Economics 6 (3): 299–310.

Footnotes

  1. The COGES acronym comes from the French name for the project: COmites de Gestion dans des EcoleS primaires.↩︎

  2. JICA refers to the Japan International Cooperation Agency, which is Japan’s foreign development and aid agency.↩︎

  3. For details about the randomization and issues about crossovers and no-shows see Sawada et al. (2022).↩︎

  4. If the election effect fades outs, then this second-year impact also identifies a difference between implementation and election effects. However, generally we find no election effects in the first-year so this not likely to be a concern. If there are sleeper effects of the election, then the endline is a compounded effect of the election and the implementation of the action plan.↩︎

  5. This is a slight abuse of notation to write a more general version because for the HD game we should have \(u_i(x^{hd}(c_i^{hd}))\).↩︎

  6. See also Bajari, Hong, Krainer, et al. (2010) and Bajari, Hong, and Ryan (2010). As an alternative estimation strategy, we tried estimating the equilibrium by having an inner loop that searches for the equilibrium beliefs by specifying an initial set of beliefs, simulating from the model, and updating beliefs until convergence followed by an outer loop that searched over the parameter space. This is obviously much more time consuming, which is the benefit of the Rust (1996) procedure. When searching for the equilibrium, we can also only use aggregate data because it would be too computationally difficult to iterate separate beliefs to convergence for different subgroups or individuals. So such a procedure is difficult to include much heterogeneity across games.↩︎

  7. We refer to these as envy and guilt terms because we interpret actual envy and guilt as each term multiplied by its respective \(\alpha_{i}\) or \(\beta_{i}\), which captures how players value the envy or guilt terms. For example, two players may have quite different preferences over the same amount of payoff inequality.↩︎

  8. The current sample is slightly different from Sawada et al. (2022) because we needed complete data on observables on all group members who played the public goods game in order to estimate the beliefs. However, the magnitude of the RCT impacts is unchanged.↩︎

  9. Due to the timing of the data collection, the COGES elections had already occured↩︎

  10. This is for a “partner” design in which players face the same (anonoymous) players in each round and with no opportunity for punishment. Fehr and Gächter (2000) report approximately 50% free-riding in the final rounds of a repeated public goods game (without punishment). Both these papers find some convergence to the Nash equilibrium with repeated play. In contrast we actually find increases in contributions across both round and time. In comparison, averaging across treatment and control, our data has only 1.44% of free-riders. Similarly in the hypothetical dictator game, we find only 0.75% of participants contributed zero. In a meta-analyses of dictator games, Engel (2011) reports 36% of participants gave zero.↩︎

  11. They also find larger fraction of zero contributors when the game is incentivized (i.e., not hypothetical) so our lower percentage of zero contributions could be explained by being hypothetical. However, our public goods game was incentized and we still see much lower percentage of zero contributions, which suggests that is not driving the result.↩︎