Reference Guide
Why Scientific Databases Disagree
Why two records for the same substance can differ in formula, mass or structure: substance versus compound records, standardisation, and provenance.
[Research Use Only] This guide is for controlled laboratory research workflows only. It is not for human or veterinary use and does not provide applied-use guidance.
Disagreement is usually structural, not sloppy
Look up the same substance in two places and the entries may differ — a different molecular formula, a mass that does not match to the second decimal, a structure drawn with stereochemistry one record omits. The instinct is to decide which source is wrong. Usually neither is. Chemical databases make different, deliberate choices about what a record represents, and two records can describe the same material faithfully while disagreeing on the page. Knowing which choices those are converts an apparent contradiction into a readable difference.
Substance and compound are not the same layer
PubChem holds structural information in two separate databases, and the distinction between them explains more confusion than any other single fact in chemical informatics. Substance records are versioned sample descriptions from individual contributors, held without normalisation processing — essentially as provided and interpreted by whoever deposited them. Compound records are derived from those substances by automated standardisation, which verifies whether structures are chemically sensible, recognises equivalent chemicals between depositors, and generates a preferred chemical representation.
- A Substance record (SID) is what a depositor supplied, preserved as supplied
- A Compound record (CID) is what standardisation produced from such deposits
- The mapping is many-to-one: many SIDs can standardise to a single CID
- A difference between two SIDs may be a difference between two depositors, not two substances
What standardisation actually does
Standardisation is not tidying. It is a defined pipeline of verification and normalisation steps, and it is the reason a compound record can be more consistent than any individual deposit behind it — and also the reason a compound record may not look exactly like the deposit you started from.
- DepositA contributor submits a substance record. It is stored as provided and interpreted, without normalisation.
- VerificationAutomated checks establish whether the structure is chemically sensible — rooted in physical reality.
- NormalisationRepresentation is regularised so equivalent structures from different depositors converge.
- Compound recordA preferred chemical representation is generated, and the substances that standardise to it are mapped to it.
Legitimate reasons two records differ
Each of the following produces a genuine, defensible difference between records. None of them means a database has made an error.
| Cause | What differs on the page | Why it is not an error |
|---|---|---|
| Salt, solvate or hydrate form | Formula and mass include the counter-ion or solvent | The records describe different materials — a salt is not the free form |
| Parent versus supplied form | One record strips associated components, the other retains them | One answers what the molecule is; the other what the material is |
| Stereochemistry | One record specifies configuration, another leaves it undefined | Standardisation can add stereo annotation where deposits lacked it |
| Isotopic specification | Labelled and unlabelled forms appear as distinct entries | Isotopic composition is part of structural identity |
| Protonation and charge | Neutral, protonated or deprotonated forms are drawn differently | Normalisation rules differ between systems and are applied deliberately |
| Average versus monoisotopic mass | Two different numbers, both labelled mass | They are different quantities, not competing estimates of one |
| Update timing | One record reflects a revision the other has not yet taken | Curation is continuous; records are snapshots |
| Deposited versus curated data | A depositor value differs from a curated one | The two layers have different provenance and different guarantees |
| Deprecated or duplicate records | Two entries exist where one is superseded | Historic identifiers are retained so old citations keep resolving |
| Synonyms | A name maps to a related but non-identical structure | Names are not identifiers, and rarely carry the precision assumed of them |
Mass is not one number
Two mass figures for the same molecule are among the most common apparent contradictions, and they are almost always both correct. Average mass is computed from the natural isotopic abundance of each element; monoisotopic mass uses the mass of the single most abundant isotope. They answer different questions and are used by different techniques. The underlying atomic weights are themselves not fixed constants: CIAAW recommends standard atomic weights applicable to normal materials, and for elements with well-documented natural variation in isotopic abundance the recommended value is an interval rather than a single number — argon, for instance, is given as [39.792, 39.963]. Values are revised as evaluation continues.
- Average mass reflects natural isotopic abundance across the element
- Monoisotopic mass uses the most abundant isotope of each element
- For some elements the standard atomic weight is an interval, not a point value
- Substantial deviations can occur where isotopic fractionation has taken place
- Recommended values are periodically revised, so two sources may reflect different evaluations
Repetition is not independent confirmation
This is the principle worth carrying beyond chemistry. Because many depositors can contribute records that standardise to the same compound, a value appearing on several substance records is not, by itself, several authorities agreeing. The paper describing PubChem standardisation gives a concrete illustration: for guanine, 153 entries in Substance standardise and map to the single structure at CID 764. Counting those as 153 confirmations would be a serious misreading — they are 153 deposits, some of which may share an origin, aggregated onto one record precisely because they were recognised as the same chemical.
Comparing two records properly
A short procedure removes most false disagreements before they start.
- Establish what each record represents — a deposited substance, or a standardised compound
- Check whether one includes a counter-ion, solvent or water of hydration the other does not
- Compare stereochemistry explicitly, including whether it is specified at all
- Confirm both mass figures are the same kind of mass
- Check dates, and whether one record has been superseded
- Compare structures or standard identifiers rather than names
- Trace provenance before treating agreement between records as corroboration
Research Checklist
- Confirm batch identity and records.
- Document all preparation inputs.
- Keep use within controlled laboratory workflows.
- Do not infer applied-use suitability from guide content.
Frequently Asked Questions
What is the practical difference between an SID and a CID?
A Substance record holds what a particular depositor supplied, kept without normalisation processing. A Compound record is produced by automated standardisation of such deposits and represents a preferred chemical representation. Many substance records can map to one compound record, which is why an SID tells you what somebody said and a CID tells you what standardisation concluded.
Why do two records show different molecular formulae for the same substance?
Most often because they are not describing the same material. A salt, solvate or hydrate carries additional components in its formula, and a record of the parent structure will not. Stereochemistry, isotopic specification and protonation state can also differ between records deliberately.
Which mass should be used?
It depends on the technique and the question. Average mass reflects natural isotopic abundance; monoisotopic mass uses the most abundant isotope of each element. Neither is the correct one in general — but comparing one against the other and finding a discrepancy is a comparison error rather than a database error.
If several databases agree, is the value confirmed?
Only if they arrived at it independently. Records frequently propagate from shared sources, and aggregation is a design feature rather than an accident — 153 substance entries standardise onto the single guanine compound record. Agreement is worth something when provenance is distinct, and considerably less when it is not.
Sources
The technical statements in this guide are drawn from the following. Where a definition is contested or a figure depends on method, the guide says so rather than picking one.
- PubChem chemical structure standardizationJournal of Cheminformatics, via PMC, National Library of MedicineUsed for the separation of the Substance and Compound databases, substance records being held without normalisation processing, the purpose of standardisation, the verification and normalisation stages, the many-to-one mapping, the treatment of stereochemistry and isotopic specification, differences in protonation normalisation, and the guanine CID 764 example with 153 mapped substance entries.
- PubChem Substance and Compound databasesNucleic Acids ResearchUsed for the structure and purpose of the two databases and the role of depositor-contributed records.
- Standard atomic weightsIUPAC Commission on Isotopic Abundances and Atomic Weights (CIAAW)Used for standard atomic weights being recommended values applicable to normal materials, interval values for elements with documented natural variation, the argon example, deviations arising from isotopic fractionation, and periodic revision.