Back

Science-wide mapping and ranking of institutions based on affiliated authors

Ioannidis, J.; Baas, J.; Boverhof, R.; Voyant, C.

2025-12-19 scientific communication and education
10.64898/2025.12.16.694665 bioRxiv
Show abstract

Standardized information on the size of institutions and on their concentration of high-impact scientists is missing. Here, we generate institution-level data on 6,979 institutions with >=50 authors who have had >=15 Scopus-indexed items. We present data on the number and proportion of top-cited authors (according to a composite citation indicator) for each institution. The primary analysis considers 1,214,702 authors (145,704 top-cited in career-long impact) who started publishing in 1980 or later, have published >=40 Scopus-indexed items and >=5 single-, first- or last-authored items. There is modest overlap among top institutions based on volume and based on number of top-cited authors. However, there is minimal overlap when institutions with the highest proportions of top-cited authors are considered; these include mostly research institutes and tech organizations and a minority of universities. Institutions with few eligible authors have large uncertainty. We propose a percentile ranking of the 2,380 largest institutions that adjusts the number of top-cited authors for the proportion of those who have high (>95th percentile) self-citation rates and high rates of publications in Scopus-discontinued titles and further penalizes for authorships of retracted papers not due to publisher/journal errors. Institutions with the highest self-citation, discontinuation, and retraction penalties cluster in specific countries. Among countries with >=15 institutions among the 2,380, Saudi Arabia, China, Malaysia, Iran, India and Indonesia have the lowest median percentile rankings of their institutions (3.9th-21st percentile) due to high penalties. The publicly available datasets that we provide allow institution-level research assessments balancing high impact against retractions and gaming practices.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.