Extras

Replication

Download replication package (.zip)

Python code + the underlying datasets · Code MIT, data CC BY 4.0. See “What's inside” below.

Ask an AI about this paper

Paste this into the AI assistant of your choice. It links a plain-text Markdown copy of the paper, the PDF, and the replication download, and asks the model to help you understand the paper and run new analyses on the data.

I want to understand and work with an academic paper: "Access to Opportunity in the Sciences: Evidence from the Nobel Laureates" by the Development Data Lab.

Here are the resources:
- Paper (Markdown, full text — easiest to read): https://devdatalab.github.io/nobel-prizes-digipub/paper.md
- Paper (PDF): https://devdatalab.github.io/nobel-prizes-digipub/paper/nobel_prizes.pdf
- Replication package (code + data, zip): https://devdatalab.github.io/nobel-prizes-digipub/downloads/nobel-laureates-replication.zip

The replication zip contains: data/ (CSV datasets, documented in data/README.md, with every column listed in data/codebook.csv) and code/ (one self-contained Python script per result, documented in code/README.md and AGENTS.md). README.md maps each figure and table to its script. Run code/run_all.py --check to reproduce every result. The main dataset is data/nobelists_analysis.csv — one row per laureate with the father's occupation, socioeconomic percentile ranks (med_inctot_rank, global_med_inctot_rank, educ_rank), and birth decade (cohort).

Please help me:
1. Summarise the paper's central finding and how it is measured.
2. Explain how the replication package is organised and how to run a result.
3. Propose and, where you can, carry out a new analysis on the dataset (for example, a breakdown I could compute from nobelists_analysis.csv).

Start by reading the Markdown paper (or the PDF) and the README files from the zip, then ask me what I want to explore.

What's inside

data/ holds the source datasets (read-only inputs). code/ holds the analysis — one self-contained Python script per result, each reading from data/ and writing JSON into its own output/ folder. code/run_all.py --check re-runs every script and confirms it reproduces the shipped results. The README.md and AGENTS.md files at the top of the package (and in data/ and code/) orient both humans and AI agents.

data/            input datasets (CSV) and codebook.csv
code/            one analysis script per result
  run_all.py     run everything (--check to verify, --only NAME for a subset)
  manifest.py    the script list and the exhibits each reproduces
  examples/      a worked example to copy for your own analysis
README.md        start here: quick start and where each result comes from
AGENTS.md        instructions for AI coding assistants
CITATION.cff     how to cite the paper
LICENSE.md       MIT (code), CC BY 4.0 (data)

Data files

FileRowsDescription
codebook.csv331Data dictionary — every column of every file below, with its type, coverage, an example value and (for key columns) a description.
nobelists_analysis.csv735Main analysis dataset — one row per laureate with father occupation, socioeconomic ranks, cohort, and derived columns. The primary input for almost every result.
nobelists_clean.csv739Pre-analysis cleaned roster, before rank imputation and derived columns.
laureates_grounded_search.csv735LLM grounded-search records used to verify/fill father names and occupations, with source URLs and reasoning.
occ1950_sci_field_classification.csv113Maps IPUMS occ1950 codes to a parent scientific field (parent-field-match analysis).
ipums_parent_field_census_shared.csv22,176IPUMS census population shares by occupation × decade × country — the over-representation baseline.
occ_rank_timeseries.csv3,939US income- and education-rank series per occupation group, 1850–1970 (input to Table A4).
occ_shares_occ1950_table.csv269Per-occ1950 laureate vs census shares and over-representation ratios.
occ_shares_25group_table.csv24The occupation-shares table rolled up to 25 groups.
occ_shares_10group_table.csv10The occupation-shares table rolled up to 10 broad groups.

Code files

Each script lives in its own folder, code/<name>/, and reproduces the exhibits listed:

ScriptReproducesDescription
rank_histograms/Figure 1Distribution of fathers' national-income, education, and global-income ranks in 5-percentile buckets, with gender splits.
main_binscatter_ranks/Figure 2, Figure A2Binned means of father rank by prize year and by birth cohort, with fitted trend lines.
parent_field_match/Figure 3, Table 3, Figures A4, A5Permutation test of whether a laureate works in their father's field, plus match-share graphs by field.
main_spec/Table 1, Table A7Headline OLS — father rank (and bottom-80/90 shares) regressed on prize year.
heterogeneity_spec/Table 2The main result split by gender, bloc, region, and prize category, with equality F-tests.
most_common_occupations/Figure A1The 10 most common father occupations — laureate vs census share and over-representation ratio.
other_rank_functions/Figure A3Robustness cuts — median rank by year and bottom-80/90 shares by year.
father_occ_by_category/Table A1Father-occupation frequencies split by prize category.
occ_shares/Tables A2, A3Occupation-share tables (10- and 25-group), the occ1950 crosswalk, and the father-occupation frequency list.
occ_representation_change/Table A4Change in occupation representation between census baselines.
main_robust/Tables A5, A6Robustness variants of the headline regression.
gender_spec_robust/Table A8The male–female gap in father rank re-estimated on four exclusion samples.
regional_analysis/Table A10Region/bloc variables, binned means by region, region OLS, and top-percentile shares (US vs rest of world).
gender_analysis/SupplementaryDistributional tests (Wilcoxon/KS), gender-split trends, KDE, and a gender × post-2000 interaction.
birth_distribution/SupplementaryLaureate counts by birth decade.
winners_cat_gender/SupplementaryLaureate counts by prize category and gender.

Table A13 is the data file occ1950_sci_field_classification.csv itself. Not reproduced by code: Table A9 (gender gap, alternate codings; not yet ported), Table A11 (a reference list of countries per region — each laureate's region is the bloc column of nobelists_analysis.csv), and the advisor-tree comparison, whose advisor-linkage data is not included.

Two known differences from the printed tables: the permutation test (Figure 3, Table 3, Figure A4) uses a fixed NumPy seed, so it is exactly repeatable but its simulated p-values can differ slightly from the original Stata run; and Table A4's census shares are approximate (see the note in occ_representation_change.py).

Citing

Novosad, Paul, Sam Asher, Catriona Farquharson, and Eni Iljazi (2026). "Access to Opportunity in the Sciences: Evidence from the Nobel Laureates." Working paper.

Report a bug

0 / 2000