Extracting knowledge of NCI research directions from funding data using language processing.

e13547 Background: In fiscal year (FY) 2019, 42% of the $6 billion NCI budget went towards nearly 5,000 research project grants, of which about 60% are R01 type. Given the enormity of allocated resources, there is a need for the scientific community to have a more rigorous understanding of the cancer research landscape. While the NCI Budget Fact Book publishes statistics based on pre-designated codings, it is unclear if this method yields the best representation of fields within oncology. Open questions include: how many distinguishable areas of cancer research are being funded? Are there differences in growth rate, publication rate and geographic distribution among those areas? Addressing these questions in a systematic manner is a well-suited problem for unsupervised machine learning. Methods: We analyzed 55,362 R-type grants from FY 2010-2021 (up to 1/31/21) from NIH ExPORTER. Preprocessing was done on ‘Project Terms’ to weight their importance using TF*IDF vectorization and principal component analysis. We used minibatch K-means clustering repeated 100 times, with the best iteration chosen by Calinski-Harabasz clustering quality score. Over 100 repetitions, the Adjusted Rand Index was 0.9907±0.0037 (mean ±standard deviation), indicating robustness to initial conditions of K-means. For publication rate analysis, FY 2020 and 2021 grants were excluded, and 2021 grants were excluded for trajectory analysis. Optimal cluster number was determined based on a combination of inertia, Calinski-Halabasz, and silhouette scores. Results: We found the optimal number of 24 clusters to best represent separation of the R-type grant research directions. These 24 clusters clearly represent topics such as immunotherapy, cohort risk-factor studies, and imaging. Notable trends include increased funding of immunotherapy and targeted inhibitor clusters averaging +$9.9M and +$9.2M growth per year respectively over 10 years, and decreased funding of pathway regulation and genetics clusters averaging -$7.8M and -$5.7M per year. These examples suggest a broader trend that funding is shifting from basic to translational science. Further analysis shows that the targeted inhibitor cluster is most geographically skewed, with 30% of grants going to institutions in just three cities. The number of journal articles (per grant) also shows a bias, with development/training and genetics clusters having publication rates of 32.6 and 20.8 and randomized control trials having a publication rate of 6.3. The average publication rate of NCI was 14.8. Conclusions: Using a novel framework for unsupervised clustering of NCI grant key-phrases, we can organize research projects more holistically than keyword searching, and more efficiently than manual categorization. Our model identifies growing and shrinking areas of research and points out biases in location and publication rate across these areas.

Paper

Full text

PDF

Extracting knowledge of NCI research directions from funding data using language processing.

Semantic Scholar · Medicine · 2021

Abstract

e13547 Background: In fiscal year (FY) 2019, 42% of the $6 billion NCI budget went towards nearly 5,000 research project grants, of which about 60% are R01 type. Given the enormity of allocated resources, there is a need for the scientific community to have a more rigorous understanding of the cancer research landscape. While the NCI Budget Fact Book publishes statistics based on pre-designated codings, it is unclear if this method yields the best representation of fields within oncology. Open questions include: how many distinguishable areas of cancer research are being funded? Are there differences in growth rate, publication rate and geographic distribution among those areas? Addressing these questions in a systematic manner is a well-suited problem for unsupervised machine learning. Methods: We analyzed 55,362 R-type grants from FY 2010-2021 (up to 1/31/21) from NIH ExPORTER. Preprocessing was done on ‘Project Terms’ to weight their importance using TF*IDF vectorization and principal component analysis. We used minibatch K-means clustering repeated 100 times, with the best iteration chosen by Calinski-Harabasz clustering quality score. Over 100 repetitions, the Adjusted Rand Index was 0.9907±0.0037 (mean ±standard deviation), indicating robustness to initial conditions of K-means. For publication rate analysis, FY 2020 and 2021 grants were excluded, and 2021 grants were excluded for trajectory analysis. Optimal cluster number was determined based on a combination of inertia, Calinski-Halabasz, and silhouette scores. Results: We found the optimal number of 24 clusters to best represent separation of the R-type grant research directions. These 24 clusters clearly represent topics such as immunotherapy, cohort risk-factor studies, and imaging. Notable trends include increased funding of immunotherapy and targeted inhibitor clusters averaging +$9.9M and +$9.2M growth per year respectively over 10 years, and decreased funding of pathway regulation and genetics clusters averaging -$7.8M and -$5.7M per year. These examples suggest a broader trend that funding is shifting from basic to translational science. Further analysis shows that the targeted inhibitor cluster is most geographically skewed, with 30% of grants going to institutions in just three cities. The number of journal articles (per grant) also shows a bias, with development/training and genetics clusters having publication rates of 32.6 and 20.8 and randomized control trials having a publication rate of 6.3. The average publication rate of NCI was 14.8. Conclusions: Using a novel framework for unsupervised clustering of NCI grant key-phrases, we can organize research projects more holistically than keyword searching, and more efficiently than manual categorization. Our model identifies growing and shrinking areas of research and points out biases in location and publication rate across these areas.

Similar papers

© 2026 NYSGPT2525 LLC