In our previous instalments of the blog series about matching (see part 1 and part 2), we explained what metadata matching is, why it is important and described its basic terminology. In this entry, we will discuss a few common beliefs about metadata matching that are often encountered when interacting with users, developers, integrators, and other stakeholders. Spoiler alert: we are calling them myths because these beliefs are not true! Read on to learn why.
We’ve just released an update to our participation report, which provides a view for our members into how they are each working towards best practices in open metadata. Prompted by some of the signatories and organizers of the Barcelona Declaration, which Crossref supports, and with the help of our friends at CWTS Leiden, we have fast-tracked the work to include an updated set of metadata best practices in participation reports for our members.
It’s been a while, here’s a metadata update and request for feedback In Spring 2023 we sent out a survey to our community with a goal of assessing what our priorities for metadata development should be - what projects are our community ready to support? Where is the greatest need? What are the roadblocks?
The intention was to help prioritize our metadata development work. There’s a lot we want to do, a lot our community needs from us, but we really want to make sure we’re focusing on the projects that will have the most immediate impact for now.
In the first half of this year we’ve been talking to our community about post-publication changes and Crossmark. When a piece of research is published it isn’t the end of the journey—it is read, reused, and sometimes modified. That’s why we run Crossmark, as a way to provide notifications of important changes to research made after publication. Readers can see if the research they are looking at has updates by clicking the Crossmark logo.
Crossref allows citation linking using Digital Object Identifiers (DOIs) between research produced by different organizations (without the need for individual agreements between them). This ensures that citation links are persistent - that they work over long periods of time. However, there is no purely technical solution to the problem of broken links on the web; Crossref members have to keep these links updated, along with rich metadata that everyone in the scholarly ecosystem relies on.
At Crossref, every metadata record that our members register for their content needs to have a unique DOI attached to it, both as a container for that record and as a locator for others to use. A DOI does not signify any value or accuracy of the thing it locates; the value lies in the record’s metadata which gives context about the object (such as contributors, funding bodies, abstract/summary) and enables connections with other entities (such as people (e.g. ORCID) or organizations (e.g. ROR)).
DOIs include 3 parts:
Of these three parts of the DOI, members (or their service providers) create the last part, the suffix. Because DOIs must be unique and persistent, members need a reasonable way to create and manage their suffixes, which should be opaque.
Here, we share the rules, guidelines and some examples to help you decide how to approach your suffixes. You can also go straight to our suffix generator.
Rules are shared by all DOI registration agencies.
Each DOI must be unique
Only use approved characters: DOI suffixes can be any alphanumeric string that includes combinations the following approved characters:
Letters of the Roman alphabet, A-Z (see below on case insensitivity)
Numbers, 0-9
-._;()/ (hyphen, period, underscore, semicolon, parentheses, forward slash). Note that the non-breaking hyphen (U+2011), figure dash (U+2012), en dash (U+2013), and em dash (U+2014) are not approved characters. The only approved hyphen is the hyphen-minus (U+002D).
Suffixes are case insensitive, so 10.1006/abc is the same in the system as 10.1006/ABC. Note that using lowercase is better for accessibility.
Guidelines for creating a DOI suffix
In part because there are few rules, it can be helpful to have some guidance in how to approach suffixes. This advice applies to DOIs at all levels, whether at journal or book level (a title-level DOI), or volume, issue, article, or chapter level.
The most important part of creating your DOIs is to understand that because DOIs are unique, persistent and ‘dumb’, once they are created, they will always work. There is never a need to delete or update existing DOIs.
Best practices for DOI suffixes:
Suffixes are best when they include short strings that are easily displayed and typed but are ‘dumb’ - meaning, the suffixes contains no readable information, including metadata.
Use a random approach: this ensures the DOIs are opaque or ‘dumb’ and minimizes attempts at interpretation or prediction (more on opaque suffixes below). Try our suggested DOI registration workflow, including our suffix generator. Any random generator will also work. Note, if you’re using the Crossref XML plugin for OJS,you don’t need to create your suffixes as the plugin will generate them for you automatically.
Keep suffixes short. This makes them easier to read and to re-type. Remember, DOIs will appear online and in print.
Best practice DOI example:10.3390/s18020479 This example appears to be opaque because it includes no obvious information.
Avoid the following in DOI suffixes:
The function of suffixes is technical in nature so they are most problematic when they are treated as information to be read, interpreted and/or predicted. Remember, DOIs are persistent and not subject to correction or deletion.
While it may be tempting, using a pattern, such as a sequence, can cause problems. Services and tools that use DOIs may, for example, try to predict future DOIs that are not registered and may never be (more on opaque suffixes below).
Don’t include information like journal title (or initials), page number or date. This kind of information should be included in the metadata but can cause problems when included in suffixes for 2 main reasons:
Information in the suffix that conflicts with information in the metadata is confusing.
Information like journal title (or initials) may change or be found to be incorrect, as with dates, but DOIs are persistent, cannot be deleted and are not subject to correction. See more on opaque suffixes below.
Example problematic DOI suffix:10.5555/2014-04-01 This example is not opaque because it includes a date, which should be included in the metadata instead of in the suffix.
Proceed with caution in DOI suffixes:
Determining how to create suffixes and manage the over time can be a challenge. We recognize that some systems have requirements that don’t follow this advice and that human readability is helpful in managing DOIs.
If you must use a suffix with meaning, internal system identifiers can work, with careful management. Because things like ISBNs are themselves metadata, we don’t recommend using them in suffixes.
Just remember that while you and readers may recognize an ISBN, for example, the DOI system itself doesn’t and DOIs are not subject to correction or deletion.
No matter your approach, it’s worth taking some time to understand the emphasis on opaque suffixes.
Once a DOI has been registered with us, it should always be used for the same content. Even if the content moves to a new website or a new owner, the same DOI should continue to be used. Though the DOI never changes, its associated metadata is kept up-to-date by the relevant Crossref member.
What if your content already has a DOI?
Sometimes members may acquire a journal that already has DOIs registered for some articles. It’s important to keep and continue to use the DOIs that have already been registered and not change them - DOIs need to be persistent.
It doesn’t matter if the prefix on the existing DOI is different from the prefix belonging to the acquiring member. As content can move between members, the owner of a DOI is not necessarily the same as the owner of the prefix. Read more about transferring responsibility for DOIs.
The importance of opaque identifiers
What are opaque suffixes & why they are important
Suffixes are ‘dumb numbers.’ They are essentially meaningless on their own and meant to be that way–opaque. One good reason for that is because when something is meaningless, it doesn’t need to be corrected.
DOIs should not include information that can be understood, interpreted or predicted, especially information that may change. Page numbers and dates are examples of information that shouldn’t be included in suffixes. It is particularly problematic if the suffix includes information that conflicts with the metadata associated with the DOI.
We’ve referred to creating ‘suffix patterns’ in the past but information that includes or implies a pattern is also problematic. A sequence of numbers, for example, lends itself to the assumption that future DOIs can be predicted.
Scraping for DOIs - or what appear to be DOIs–is common, as is the likelihood that what is–or appears to be–a pattern will be treated as such. Just as the timing of DOI registration is important, in order to avoid unregistered DOIs, their construction is critical to avoiding interpretation.
More information on creating DOIs
Here are a few other resources that discuss creating DOIs and the importance of using opaque suffixes.