Extras
Replication
Python code + the underlying datasets · Code MIT, data CC BY 4.0. See “What's inside” below.
Ask an AI about this paper
Paste this into the AI assistant of your choice. It links a plain-text Markdown copy of the paper, the PDF, and the replication download, and asks the model to help you understand the paper and run new analyses on the data.
I want to understand and work with an academic paper: "Access to Opportunity in the Sciences: Evidence from the Nobel Laureates" by the Development Data Lab.
Here are the resources:
- Paper (Markdown, full text — easiest to read): https://devdatalab.github.io/nobel-prizes-digipub/paper.md
- Paper (PDF): https://devdatalab.github.io/nobel-prizes-digipub/paper/nobel_prizes.pdf
- Replication package (code + data, zip): https://devdatalab.github.io/nobel-prizes-digipub/downloads/nobel-laureates-replication.zip
The replication zip contains: data/ (CSV datasets, documented in data/README.md, with every column listed in data/codebook.csv) and code/ (one self-contained Python script per result, documented in code/README.md and AGENTS.md). README.md maps each figure and table to its script. Run code/run_all.py --check to reproduce every result. The main dataset is data/nobelists_analysis.csv — one row per laureate with the father's occupation, socioeconomic percentile ranks (med_inctot_rank, global_med_inctot_rank, educ_rank), and birth decade (cohort).
Please help me:
1. Summarise the paper's central finding and how it is measured.
2. Explain how the replication package is organised and how to run a result.
3. Propose and, where you can, carry out a new analysis on the dataset (for example, a breakdown I could compute from nobelists_analysis.csv).
Start by reading the Markdown paper (or the PDF) and the README files from the zip, then ask me what I want to explore.What's inside
data/ holds the source datasets (read-only inputs). code/ holds
the analysis — one self-contained Python script per result, each reading from data/ and writing JSON into its own output/ folder. code/run_all.py --check re-runs every script and confirms it reproduces the
shipped results. The README.md and AGENTS.md files at the top of the package (and in data/ and code/) orient both humans and AI agents.
data/ input datasets (CSV) and codebook.csv
code/ one analysis script per result
run_all.py run everything (--check to verify, --only NAME for a subset)
manifest.py the script list and the exhibits each reproduces
examples/ a worked example to copy for your own analysis
README.md start here: quick start and where each result comes from
AGENTS.md instructions for AI coding assistants
CITATION.cff how to cite the paper
LICENSE.md MIT (code), CC BY 4.0 (data) Data files
| File | Rows | Description |
|---|---|---|
codebook.csv | 331 | Data dictionary — every column of every file below, with its type, coverage, an example value and (for key columns) a description. |
nobelists_analysis.csv | 735 | Main analysis dataset — one row per laureate with father occupation, socioeconomic ranks, cohort, and derived columns. The primary input for almost every result. |
nobelists_clean.csv | 739 | Pre-analysis cleaned roster, before rank imputation and derived columns. |
laureates_grounded_search.csv | 735 | LLM grounded-search records used to verify/fill father names and occupations, with source URLs and reasoning. |
occ1950_sci_field_classification.csv | 113 | Maps IPUMS occ1950 codes to a parent scientific field (parent-field-match analysis). |
ipums_parent_field_census_shared.csv | 22,176 | IPUMS census population shares by occupation × decade × country — the over-representation baseline. |
occ_rank_timeseries.csv | 3,939 | US income- and education-rank series per occupation group, 1850–1970 (input to Table A4). |
occ_shares_occ1950_table.csv | 269 | Per-occ1950 laureate vs census shares and over-representation ratios. |
occ_shares_25group_table.csv | 24 | The occupation-shares table rolled up to 25 groups. |
occ_shares_10group_table.csv | 10 | The occupation-shares table rolled up to 10 broad groups. |
Code files
Each script lives in its own folder, code/<name>/, and reproduces the exhibits listed:
| Script | Reproduces | Description |
|---|---|---|
rank_histograms/ | Figure 1 | Distribution of fathers' national-income, education, and global-income ranks in 5-percentile buckets, with gender splits. |
main_binscatter_ranks/ | Figure 2, Figure A2 | Binned means of father rank by prize year and by birth cohort, with fitted trend lines. |
parent_field_match/ | Figure 3, Table 3, Figures A4, A5 | Permutation test of whether a laureate works in their father's field, plus match-share graphs by field. |
main_spec/ | Table 1, Table A7 | Headline OLS — father rank (and bottom-80/90 shares) regressed on prize year. |
heterogeneity_spec/ | Table 2 | The main result split by gender, bloc, region, and prize category, with equality F-tests. |
most_common_occupations/ | Figure A1 | The 10 most common father occupations — laureate vs census share and over-representation ratio. |
other_rank_functions/ | Figure A3 | Robustness cuts — median rank by year and bottom-80/90 shares by year. |
father_occ_by_category/ | Table A1 | Father-occupation frequencies split by prize category. |
occ_shares/ | Tables A2, A3 | Occupation-share tables (10- and 25-group), the occ1950 crosswalk, and the father-occupation frequency list. |
occ_representation_change/ | Table A4 | Change in occupation representation between census baselines. |
main_robust/ | Tables A5, A6 | Robustness variants of the headline regression. |
gender_spec_robust/ | Table A8 | The male–female gap in father rank re-estimated on four exclusion samples. |
regional_analysis/ | Table A10 | Region/bloc variables, binned means by region, region OLS, and top-percentile shares (US vs rest of world). |
hetero_trends/ | Table A12 | Subgroup-specific time-trend regressions by gender, category, and region. |
gender_analysis/ | Supplementary | Distributional tests (Wilcoxon/KS), gender-split trends, KDE, and a gender × post-2000 interaction. |
birth_distribution/ | Supplementary | Laureate counts by birth decade. |
winners_cat_gender/ | Supplementary | Laureate counts by prize category and gender. |
Table A13 is the data file occ1950_sci_field_classification.csv itself. Not
reproduced by code: Table A9 (gender gap, alternate codings; not yet ported), Table A11
(a reference list of countries per region — each laureate's region is the bloc column of nobelists_analysis.csv), and the advisor-tree
comparison, whose advisor-linkage data is not included.
Two known differences from the printed tables: the permutation test (Figure 3, Table 3,
Figure A4) uses a fixed NumPy seed, so it is exactly repeatable but its simulated
p-values can differ slightly from the original Stata run; and Table A4's census shares
are approximate (see the note in occ_representation_change.py).
Citing
Novosad, Paul, Sam Asher, Catriona Farquharson, and Eni Iljazi (2026). "Access to Opportunity in the Sciences: Evidence from the Nobel Laureates." Working paper.