Mapping Metadata Inequities: Regional Disparities in Crossref Scholarly Records

Author
Affiliation

Luis Montilla

Crossref

Published

May 28, 2025

Abstract

The Crossref REST API provides the global community with access to over 165 million scholarly metadata records, serving as a foundational resource for discovery, citation, integrity assessment and provenance tracking (Hendricks et al. 2020). However, disparities in the completeness of this metadata can pose significant challenges to equitable access and utilization of scholarly knowledge. The diversity of barriers to metadata enrichment is complex and multifaceted, including but not limited to lack of institutional infrastructure, automation challenges, lack of awareness and/or training, language barriers or low availability of staff and resources to perform these tasks. Crossref schemas allow members to deposit rich metadata beyond basic bibliographic elements, namely abstracts, list of references, author identification via ORCIDs, institutional affiliation and also identification via ROR IDs, funding information, including funder and award IDs, journal update policies via the Crossmark service, license information, all of which contribute to the realize an open interconnected scholarly knowledge network that we aspire to as part of the Research Nexus. The regional differences in terms of journals included in the Crossref overall data have been previously described (Asubiaro & Onaolapo, 2023). Here, instead, we will explore regional disparities patterns in metadata completeness, focusing on highly relevant metadata fields such as references, author affiliations, ORCIDs, and funding information

Extended poster

I presented a short spatial analysis of Crossref member metrics in the Workshop on Open Citations & Open Scholarly Metadata 2025.

A group of people

Citation

BibTeX citation:
@online{montilla2025,
  author = {Montilla, Luis},
  title = {Mapping {Metadata} {Inequities:} {Regional} {Disparities} in
    {Crossref} {Scholarly} {Records}},
  date = {2025-05-28},
  url = {https://www.luismmontilla.com/events/bologna2025/},
  langid = {en},
  abstract = {The Crossref REST API provides the global community with
    access to over 165 million scholarly metadata records, serving as a
    foundational resource for discovery, citation, integrity assessment
    and provenance tracking (Hendricks et al. 2020). However,
    disparities in the completeness of this metadata can pose
    significant challenges to equitable access and utilization of
    scholarly knowledge. The diversity of barriers to metadata
    enrichment is complex and multifaceted, including but not limited to
    lack of institutional infrastructure, automation challenges, lack of
    awareness and/or training, language barriers or low availability of
    staff and resources to perform these tasks. Crossref schemas allow
    members to deposit rich metadata beyond basic bibliographic
    elements, namely abstracts, list of references, author
    identification via ORCIDs, institutional affiliation and also
    identification via ROR IDs, funding information, including funder
    and award IDs, journal update policies via the Crossmark service,
    license information, all of which contribute to the realize an open
    interconnected scholarly knowledge network that we aspire to as part
    of the Research Nexus. The regional differences in terms of journals
    included in the Crossref overall data have been previously described
    (Asubiaro \& Onaolapo, 2023). Here, instead, we will explore
    regional disparities patterns in metadata completeness, focusing on
    highly relevant metadata fields such as references, author
    affiliations, ORCIDs, and funding information}
}
For attribution, please cite this work as:
Montilla, Luis. 2025. “Mapping Metadata Inequities: Regional Disparities in Crossref Scholarly Records.” Workshop on Open Citations and Open Scholarly Metadata 2025, May 28. https://www.luismmontilla.com/events/bologna2025/.