Skip to main content

Reference Guide

Why Scientific Databases Disagree

Why two records for the same substance can differ in formula, mass or structure: substance versus compound records, standardisation, and provenance.

[Research Use Only] This guide is for controlled laboratory research workflows only. It is not for human or veterinary use and does not provide applied-use guidance.

Disagreement is usually structural, not sloppy

Look up the same substance in two places and the entries may differ — a different molecular formula, a mass that does not match to the second decimal, a structure drawn with stereochemistry one record omits. The instinct is to decide which source is wrong. Usually neither is. Chemical databases make different, deliberate choices about what a record represents, and two records can describe the same material faithfully while disagreeing on the page. Knowing which choices those are converts an apparent contradiction into a readable difference.

Substance and compound are not the same layer

PubChem holds structural information in two separate databases, and the distinction between them explains more confusion than any other single fact in chemical informatics. Substance records are versioned sample descriptions from individual contributors, held without normalisation processing — essentially as provided and interpreted by whoever deposited them. Compound records are derived from those substances by automated standardisation, which verifies whether structures are chemically sensible, recognises equivalent chemicals between depositors, and generates a preferred chemical representation.

  • A Substance record (SID) is what a depositor supplied, preserved as supplied
  • A Compound record (CID) is what standardisation produced from such deposits
  • The mapping is many-to-one: many SIDs can standardise to a single CID
  • A difference between two SIDs may be a difference between two depositors, not two substances

What standardisation actually does

Standardisation is not tidying. It is a defined pipeline of verification and normalisation steps, and it is the reason a compound record can be more consistent than any individual deposit behind it — and also the reason a compound record may not look exactly like the deposit you started from.

  1. DepositA contributor submits a substance record. It is stored as provided and interpreted, without normalisation.
  2. VerificationAutomated checks establish whether the structure is chemically sensible — rooted in physical reality.
  3. NormalisationRepresentation is regularised so equivalent structures from different depositors converge.
  4. Compound recordA preferred chemical representation is generated, and the substances that standardise to it are mapped to it.
Because the substance layer is preserved unchanged, both views remain available: what each depositor said, and what standardisation made of it. Comparing across the two layers without noticing which is which is a reliable source of false contradictions.

Legitimate reasons two records differ

Each of the following produces a genuine, defensible difference between records. None of them means a database has made an error.

CauseWhat differs on the pageWhy it is not an error
Salt, solvate or hydrate formFormula and mass include the counter-ion or solventThe records describe different materials — a salt is not the free form
Parent versus supplied formOne record strips associated components, the other retains themOne answers what the molecule is; the other what the material is
StereochemistryOne record specifies configuration, another leaves it undefinedStandardisation can add stereo annotation where deposits lacked it
Isotopic specificationLabelled and unlabelled forms appear as distinct entriesIsotopic composition is part of structural identity
Protonation and chargeNeutral, protonated or deprotonated forms are drawn differentlyNormalisation rules differ between systems and are applied deliberately
Average versus monoisotopic massTwo different numbers, both labelled massThey are different quantities, not competing estimates of one
Update timingOne record reflects a revision the other has not yet takenCuration is continuous; records are snapshots
Deposited versus curated dataA depositor value differs from a curated oneThe two layers have different provenance and different guarantees
Deprecated or duplicate recordsTwo entries exist where one is supersededHistoric identifiers are retained so old citations keep resolving
SynonymsA name maps to a related but non-identical structureNames are not identifiers, and rarely carry the precision assumed of them
Ten mechanisms, none of them a mistake. The practical consequence is that comparing two records means comparing what they are records of, before comparing their contents.

Mass is not one number

Two mass figures for the same molecule are among the most common apparent contradictions, and they are almost always both correct. Average mass is computed from the natural isotopic abundance of each element; monoisotopic mass uses the mass of the single most abundant isotope. They answer different questions and are used by different techniques. The underlying atomic weights are themselves not fixed constants: CIAAW recommends standard atomic weights applicable to normal materials, and for elements with well-documented natural variation in isotopic abundance the recommended value is an interval rather than a single number — argon, for instance, is given as [39.792, 39.963]. Values are revised as evaluation continues.

  • Average mass reflects natural isotopic abundance across the element
  • Monoisotopic mass uses the most abundant isotope of each element
  • For some elements the standard atomic weight is an interval, not a point value
  • Substantial deviations can occur where isotopic fractionation has taken place
  • Recommended values are periodically revised, so two sources may reflect different evaluations

Repetition is not independent confirmation

This is the principle worth carrying beyond chemistry. Because many depositors can contribute records that standardise to the same compound, a value appearing on several substance records is not, by itself, several authorities agreeing. The paper describing PubChem standardisation gives a concrete illustration: for guanine, 153 entries in Substance standardise and map to the single structure at CID 764. Counting those as 153 confirmations would be a serious misreading — they are 153 deposits, some of which may share an origin, aggregated onto one record precisely because they were recognised as the same chemical.

Comparing two records properly

A short procedure removes most false disagreements before they start.

  • Establish what each record represents — a deposited substance, or a standardised compound
  • Check whether one includes a counter-ion, solvent or water of hydration the other does not
  • Compare stereochemistry explicitly, including whether it is specified at all
  • Confirm both mass figures are the same kind of mass
  • Check dates, and whether one record has been superseded
  • Compare structures or standard identifiers rather than names
  • Trace provenance before treating agreement between records as corroboration

Research Checklist

  • Confirm batch identity and records.
  • Document all preparation inputs.
  • Keep use within controlled laboratory workflows.
  • Do not infer applied-use suitability from guide content.

Frequently Asked Questions

What is the practical difference between an SID and a CID?

A Substance record holds what a particular depositor supplied, kept without normalisation processing. A Compound record is produced by automated standardisation of such deposits and represents a preferred chemical representation. Many substance records can map to one compound record, which is why an SID tells you what somebody said and a CID tells you what standardisation concluded.

Why do two records show different molecular formulae for the same substance?

Most often because they are not describing the same material. A salt, solvate or hydrate carries additional components in its formula, and a record of the parent structure will not. Stereochemistry, isotopic specification and protonation state can also differ between records deliberately.

Which mass should be used?

It depends on the technique and the question. Average mass reflects natural isotopic abundance; monoisotopic mass uses the most abundant isotope of each element. Neither is the correct one in general — but comparing one against the other and finding a discrepancy is a comparison error rather than a database error.

If several databases agree, is the value confirmed?

Only if they arrived at it independently. Records frequently propagate from shared sources, and aggregation is a design feature rather than an accident — 153 substance entries standardise onto the single guanine compound record. Agreement is worth something when provenance is distinct, and considerably less when it is not.

Sources

The technical statements in this guide are drawn from the following. Where a definition is contested or a figure depends on method, the guide says so rather than picking one.

  1. PubChem chemical structure standardizationJournal of Cheminformatics, via PMC, National Library of MedicineUsed for the separation of the Substance and Compound databases, substance records being held without normalisation processing, the purpose of standardisation, the verification and normalisation stages, the many-to-one mapping, the treatment of stereochemistry and isotopic specification, differences in protonation normalisation, and the guanine CID 764 example with 153 mapped substance entries.
  2. PubChem Substance and Compound databasesNucleic Acids ResearchUsed for the structure and purpose of the two databases and the role of depositor-contributed records.
  3. Standard atomic weightsIUPAC Commission on Isotopic Abundances and Atomic Weights (CIAAW)Used for standard atomic weights being recommended values applicable to normal materials, interval values for elements with documented natural variation, the argon example, deviations arising from isotopic fractionation, and periodic revision.