Social Media and Scientific Careers · data explorer
These are the raw averages behind my estimates: publications and citations per year for US research academics, grouped by the year they opened a Twitter account and the year they finished their PhD. Each comparison group sits next to them, so you can see the selection problem before any estimator touches it.
Each panel plots one group of joiners against the comparison groups you select. Solid lines are Twitter joiners; dashed lines are people who had not joined. Color marks the comparison: the dashed green line is the matched control group for the solid green line. With the x axis in years since entry, the vertical rule at zero is the year the joiners opened their accounts.
The "All, pooled" chip at the end of the entry-year and PhD-year rows adds a panel that pools every entry year or every PhD year; select it alone to see only that panel. In years since entry, a pooled line's composition shifts at long horizons, because late cohorts have fewer post-entry years in the data and drop out first. When a panel pools several entry years or PhD-year bins, the joiner lines are plain pooled averages. The never-adopter and not-yet-joiner lines are first averaged within each entry year and PhD-year bin and then weighted by the number of joiners in that cell, so they share the joiners' composition. Matched controls already share it by construction, with one adjustment: some joiners found three controls and others only one, and joiners who found more controls tend to differ from those who found fewer. Each control is therefore weighted by one over the number of controls in its cell, so every matched pair of joiner and controls counts once on each side. Without that weight, the control line sits well below the joiner line before entry for reasons of bookkeeping alone. Hover over a panel to see each value and the number of person-years behind it; points with fewer than 10 person-years are not drawn.
Every panel groups people by the year they finished their PhD (or other terminal degree: MD, JD, MFA and the like), which I collected by hand from CVs and faculty pages. Only people with a collected degree year appear, so the panels cover roughly a quarter of joiners and a third of never adopters. In the matched samples, joiner and controls are placed by the joiner's PhD year, so a matched pair always appears in the same panel.
Values are plain means with no winsorizing, so a few very highly cited people can move a thin cell. The panel runs from 2001, the first year OpenAlex tracks citing works, to 2020.
Matched controls take their joiner's value, so a matched pair always stays together. Never adopters have a field but no recorded gender or ethnicity and no entry year to define a pre-entry window, so they appear only in the field split. Not-yet joiners are left out of the pre-entry split, since their own pre-entry window lies years after the cohort they are compared with.