· 5 min read

My collaboration network for 2010 to 2020 (+ other plots)

In what has become a bit of an annual tradition, here is my collaboration network for 2010 to 2020. This year was rough. Of the two first-author papers published this year, one was pre-pandemic. I think it’s fair to say this wasn’t the level of productivity I was expecting of myself. Hopefully, a few projects still in the pipeline will come out early next year.

All that said, I’m thankful for a strong network of kind collaborators who picked up my slack when necessary, checked in on me even when we didn’t have an active project, and understood when childcare issues caused last minute Zoom cancellations.

You’ll have plenty of time to work with famous, smart, and/or fun people — 2020 was a good reminder of the importance of working with kind people.

An animation stepping one year at a time from 2010 to 2020. The main panel is the coauthor network: blue nodes are papers, sized by citation count on a legend from 5 to 100, and red nodes are collaborators. Each year the papers and coauthors added that year light up while everything earlier fades to the background, so the network grows outward from a few small clusters into one dominant component plus three or four that stay separate. Five smaller panels advance in step: new manuscripts per year against a long-run average line, new collaborators per year against its own average, citations accumulating per paper as papers age, cumulative citations across all papers, and new citations each year. The citation panels climb throughout and steepen after about 2015, ending near 1,400 cumulative citations; new manuscripts per year peaks at seven in 2015.

The first time I made this plot, I noted how many components I had and how disjointed the collaboration networks were. Since then, there’s now (1) a dominant connected component, (2) my NYU component (top middle) that will likely always be disconnected, (3) my Health Policy and Management cluster (top left), which has a reasonable chance of connecting with the rest of the group now that Sara is at Stanford, and (4) a single paper with Alex (middle right), which will almost certainly join the rest of the group at some point. It’s also interesting to note the trajectories of papers (in terms of citations) in the lower left. A couple papers seem to get some traction, but for the most part, my papers tend to add citations at around 5-10 cites per year.

Below is a plot of collaborations (circles) over time (x-axis) by collaborator (y-axis). I’ve worked fairly consistently with two people, Nancy and Jarvis, for six years, which is pretty wild. Most of my collaborations are bursty with a rush of papers and then long dormant periods but a handful are pretty regular with ~1 paper per year. Most of my collaborators are one-time collaborators.

Collaborations over time, one row per collaborator with no names shown. The x-axis runs 2010 to 2020 and circle size is the number of collaborations that year, from 1 to 5. A grey line joins each collaborator's first collaboration to their last. Most rows are a single small circle, a one-time collaborator, and vertical stacks of circles mark one paper with many coauthors, the tallest in 2020. The longest rows span 2014 to 2020, about six years; the rest are short bursts separated by long gaps.

Another thing we can look at is which of my collaborators also collaborate together (conditional on me being on the paper). Below, I show the top ten (in terms of number of collaborations) collaborators with a horizontal bar chart for the number of times we have worked together. The lower right plot shows dots and lines of intersecting collaborators along with how often this subset of collaborators appears in my collaborations (vertical bars).

An UpSet plot of my ten most frequent collaborators, shown by name on the chart. Horizontal bars on the left give each person's total collaborations, sorted shortest at the top to longest at the bottom. Vertical bars above show how many papers each combination of those people appears on: 6, then 5, 5, 3, 3, 2, 2, and four combinations with 1. Connected dots below each bar mark which people are in that combination. The tallest bar, six papers, is a group of four; one of the fives is a single person who appears in no other combination.

For example, there are six papers with Jason, Pam, Nancy, and Jarvis and additional two with the same group minus Jason. Sara is the outlier here with a 5 collaborations — none of which involve another top ten collaborator.

Lastly, for kicks I wanted to see who my “most efficient” collaborator is. That is, conditional on more than one project together, who has the highest average number of citations per project?

A scatter plot with number of papers together on the x-axis, 1 to about 11, and total citations on the y-axis, 0 to about 370. Circle size is years since the last collaboration, larger meaning more recent, on a legend of 0, 3, 7 and 10. Colour is citations per paper among people with more than one collaboration, running dark purple below 40 through magenta and red to yellow above 120; people with a single collaboration are black. Most points sit at one or two papers with few citations. Two overlapping yellow circles at two papers and about 310 citations are the highest citations per paper on the chart. The two points furthest right, at ten and eleven papers, reach about 340 and 370 total citations.

The answer (two yellow dots in upper left) is Nishant and Rafa at about 150 citations per project. (The upper right is Jarvis followed by Nancy.)

Code is here. Note there are five files and you need to change my_id on line 21 of 01_pull_data.R to your Google Scholar ID.