Section 2

Data and Methods

We collected biographical information on the Nobel laureates from 1901–2023 from a range of historical sources, covering 739 winners in chemistry, physics, physiology, and economics.1 We exclude the prizes for Peace and for Literature, as accomplishments in these domains are more subjective and can be intentionally awarded to individuals who faced challenging childhood circumstances.

For each winner, we collected demographic information and parents' occupations.2 We supplemented the data by contacting living winners directly to request information about their parents. In total, we were able to identify fathers' occupations for 714 out of the 739 laureates, but mothers for only 181; our analysis therefore focuses on father occupation as a measure of childhood socioeconomic status.3 Appendix B.1 details how we assembled this biographical record.

Appendix B.1 · Biographical Data

We hand-collected biographical information for every Nobel laureate in physics, chemistry, physiology, and economics. For each laureate, we reviewed multiple sources to identify their birth dates, birth locations, and the occupations of their parents. The availability of parent occupational information was heterogeneous. Many winners described their parents' work in their acceptance speeches or biographies. Parent occupations were also recorded in documents like obituaries, biographies, scientific directories and encyclopedia entries. When we could not find information on father occupation, we reached out to living laureates, a handful of whom responded with information about their parents. We obtained data on father occupations for 714 out of 739 laureates (96.2%), but only 181 mother occupations (25.3%). Mothers' occupations are also generally less informative about laureates' childhood socioeconomic status, because so many women in this era were out of the labor force.

When laureates moved before the age of six, we assigned them to their location at age six. Results were unchanged if we used other ages as anchors for their home country. When laureates were born to expatriates (i.e. parents living temporarily in countries other than their home country), we assigned them to their parents' home countries, as the home country occupational distributions are likely to be more relevant for determining parental rank. We similarly assigned them to home countries when parents were serving in a foreign service. For example, Ronald Ross (physiology, 1902) was the son of a British general serving in India; at age 8, he moved to the Isle of Wight for school. Our assumption is that his parents' socioeconomic status is better characterized by the average status of an army general in the United Kingdom than that of an army general in India. There are only a handful of cases like these, so they do not have a substantial bearing on our results.

We construct cohorts of winners by rounding each winner's birth year to the nearest decade (e.g., a winner born in 1893 belongs to the 1890 cohort, whereas a winner born in 1898 belongs to the 1900 cohort). There are six winners born between 1835–1843, all of whom are initially assigned to the 1840 cohort. However, due to the lack of availability of census data during these years, we then map these observations to 1850 instead, the closest available decade.

Using father occupation as a measure of childhood socioeconomic status is a common approach in the historical intergenerational mobility literature . Father occupation is a good predictor of other measures of socioeconomic status, and is often the only such variable available in historical sources. We use IPUMS data from the 1850–1970 U.S. population censuses to calculate occupational earnings and education levels covering the birth years of the majority of winners.45 Appendix B.2 describes the construction of these occupational earnings and education ranks in full.

Appendix B.2 · Occupational Earnings & Education from the U.S. Census

To estimate the socioeconomic status of laureates' parents, we aimed to identify the average income and education ranks of people with the same occupation as each laureate's father in the U.S. Census. We used data from the 1850–1970 U.S. population censuses, extracted from IPUMS and proceeded as follows.

We restricted the census sample to working age (18–65 years old) men, and dropped all individuals who were not in the labor force. We also dropped census observations for which the occupational category variable takes on the following values: "students", "disabled individuals with no reported occupation", "other non-occupation", "n/a (blank)". Excluding the unemployed could marginally bias our ranks downward, but the bias should be small, given that (i) the historic U.S. unemployment rate has typically fluctuated between only 5–8%; and (ii) the unemployed are distributed across many different occupations, limiting the degree of bias.

We matched each laureate's father occupation to the nearest harmonized occupational category (occ1950, created by IPUMS). When fathers held multiple occupations, we assigned the one described as the primary occupation during the laureate's childhood years. Fathers described only as aristocrats or large estate holders, with no more specific occupation reported, were assigned both income and education ranks of 100. shows the list of occupational categories used in the sample.

There is inevitably some measurement error in classifying business owners, as the historical record was often unclear on the scale of the business; as such, a business owner could represent a relatively upper or middle class family. We classified business owners with code 290, which is "Managers, officials, and proprietors, not otherwise classified." This puts business owners, on average, at the 86th education percentile and the 93rd income percentile. Farm owners exist as an occ1950 category, but similarly cover a wide range of socioeconomic ranks. To ensure that our assumptions about these occupational ranks were not driving any of our results, we verified that our primary results are similar when we exclude business owners and farmers (Appendix Tables and ).

Individual and household earnings were available beginning in 1950, and years of education in 1940, so we used these values for all earlier years. We excluded negative or zero earnings values. Literacy is recorded in earlier census years, but is considered less reliable by historians, as collection methodology varied substantially across enumerators . For census years after 1940/50, we applied a three-census moving average to smooth out fluctuations in occupational income and education from sampling variation and interpolated missing values. These adjustments have little bearing on our results, as they only affect census years after 1940/50, and most Nobel laureates were born earlier.

We also estimated the population share of each occupation in each decade, which is essential for calculating ranks. To calculate ranks, we assigned individuals an earnings level based on their occupation, and then ordered individuals within each decadal census according to those earnings levels. We then assigned the mean earnings rank within each occupation group, and collapsed the data to occupation-decade level. These earnings ranks were then matched to the Nobel laureates database at the occupation-decade level. We performed a similar exercise for education ranks.

Since occupational incomes were not observed before 1950, there is no change in the relative ordering of occupations before 1950. However, we observe the share of people with each occupation in all census years, so the ranks of occupations do change continuously over time. For instance, trade workers tended to have relatively high education ranks in 1850, but had lower ranks by 1950, due to the growing number of more educated white collar workers.

We calculated the mean total individual income and education ranks for working-age men in each occupation category in each decade, and assigned these ranks to laureates' birth decades and fathers' occupations.6 These ranks reflect the average income or education rank that would have been held by an adult male in the given occupation in the birth year of the laureate. Income and educational ranks are highly but imperfectly correlated; some example outliers are members of the clergy and (to a lesser extent) teachers (high education but low income ranks), and tradesmen (high income but low education ranks).7

Occupational rank data are not available for all countries in which laureates were born, so our estimates are based on the U.S. occupational rank distribution. We adjusted for the different occupational distributions in less-developed countries by matching country-decade pairs to the U.S. decade with the most similar level of development. For example, in 1960, a professor in India would have a higher relative socioeconomic rank in India than a professor would have in the U.S. To account for these differences, we classify the Indian professor using the U.S. rank distribution from 1850, the year in the U.S. data with the most similar per capita GDP to India in 1960 — thus assigning them a higher socioeconomic rank within their country.8 Appendix B.3 sets out the country-level GDP adjustments behind these national and global rank measures.

Appendix B.3 · Country-Level GDP per Capita & Population

We used country of childhood in two ways in our analysis. For our primary analysis, we used country of childhood to adjust the relative occupational distributions, to reflect the fact that an individual with only some schooling (say, literacy) would occupy a higher education rank in a poor country than in a rich country. Second, for the section on income differences across countries only, we used national GDP per capita to adjust occupational earnings for living standards across countries. This second adjustment reflects the fact that an individual at the 90th percentile of the earnings distribution in India has a much lower global income rank than an individual at the 90th percentile in the United States.

We used historical estimates of GDP per capita and population for 217 countries from 1830–1970, from , with adjustments described below.

Adjusting for occupation distributions. The distribution of occupational population shares affects the resulting socioeconomic rank of each occupation, even if the ordering of occupations is unchanged. For example, the percentile rank of a sawyer in the U.S. is much higher in 1850, a time when roughly 44% of the population were farmers, compared to 1950, when that share is down to only 9%. Similarly, countries will have different occupational distributions at different levels of economic development. However, we only had data on the historical occupational distribution of the United States.

We therefore used the U.S. historical occupational distribution to proxy for the occupational distribution in other countries at different points in time. We assigned each country the U.S. occupational distribution from the decade when its development level (measured by GDP per capita) was nearest to that of the U.S., using minimum absolute distance.

Consider an example of how we would calculate the earnings rank of an Indian tailor in 1950. First, we identified the U.S. decade in our sample data with the nearest GDP per capita level to that of India in 1950 — this was 1850. In 1850 in the U.S., a tailor had an earnings rank of 84; we assign this rank to the Indian tailor for the primary analysis. By contrast, in the U.S. in 1950, a tailor would have an earnings rank of only 55.

Calculating global income ranks. Our primary analysis focuses on occupational ranks within countries. Since most winners are from the West, this tells us how effective the West is at mobilizing its own talent. In the final part of the results, we instead consider global income ranks, which tell us how effective human society is at mobilizing global talent. Consider again the Indian tailor in 1950 above. The tailor occupied the 84th earnings rank in India, but their global earnings rank was much lower, since India in 1950 was a very poor country.

We calculate a global earnings rank as follows. First, for each country-decade-occupation group, we scale the occupational income by the GDP per capita ratio of each country with the United States in the same decade. This gives us a real income estimate for each country-decade-occupation group. Second, we create a synthetic global population based on each country's occupation shares (adjusted for development level, as above) and their population. An individual's rank in this synthetic population is that individual's global income rank.

Let us continue with the example of the Indian tailor above. We have already estimated that the tailor is at the 84th earnings percentile in India in 1950. In 1950, U.S. GDP per capita was approximately 16 times higher than Indian GDP; we therefore multiply the Indian tailor's predicted income by 0.06. After conducting this calculation for every country-occupation-decade, the Indian tailor has a global income rank of 46. This reflects the fact that the tailor is relatively well off in India, but India is relatively poor globally. Appendix Figure B1 summarizes the example.

For individuals from poor countries, global ranks are systematically lower than national ranks, and vice versa. A U.S. tailor in 1950 would be at the 97th percentile globally. Since most of the Nobel laureates are from rich countries, the average global earnings rank is therefore higher than the average national earnings rank.

We describe these as national rank measures, as they describe the relative status of laureates' fathers within their own countries. To address income differences across countries, we use the following procedure to rank each country-decade-occupation in the global income distribution.

For each country-decade pair, we create a synthetic population of individuals, reflecting the occupation distribution in the U.S. decade with the nearest per capita GDP, as above. We assign each person an occupational income equal to the 1950 U.S. census income for that occupation, multiplied by the proportional GDP difference between that country and the United States, in the given decade. We then rerank all individuals in the synthetic global population for each decade, creating a new global income rank for each country-decade-occupation group.

For example, a tailor in 1950 has a U.S. occupational income rank of 55 — their position in the U.S. national income distribution. According to our global measure, a U.S. tailor in 1950 would be at the 97th percentile globally, reflecting high U.S. incomes compared with the rest of the world. Conversely, we would classify an Indian tailor in 1950 at the 46th global percentile, reflecting India's relative poverty in 1950.

References (5)
  1. Long, J., Ferrie, J. (2013). Intergenerational Occupational Mobility in Great Britain and the United States since 1850. American Economic Review. View source →
  2. Song, X., Massey, C., Rolf, K., Ferrie, J., Rothbaum, J., Xie, Y. (2020). Long-term decline in intergenerational mobility in the United States since the 1850s. Proceedings of the National Academy of Sciences. View source →
  3. Long, J., Ferrie, J. (2007). The Path to Convergence: Intergenerational Occupational Mobility in Britain and the US in Three Eras. The Economic Journal. View source →
  4. Collins, W., Margo, R. (2006). Historical Perspectives on Racial Differences in Schooling in the United States. Handbook of the Economics of Education. View source →
  5. Fariss, C. J., Anders, T., Markowitz, J. N., Barnum, M. (2022). New Estimates of Over 500 Years of Historic GDP and Population Data. Journal of Conflict Resolution. View source →

Report a bug

0 / 2000