The scientific and regulatory information about terpenes is, in aggregate, rich. In practice it is nearly unusable without heavy engineering, because it is fragmented across sources that were never designed to work together. This section names the specific failure modes. Each of them corresponds to a design decision in Terpedia described in the next section.
Fragmentation¶
A team that wants a complete picture of a single terpene must consult, at a minimum:
Chemical registries for identity and structure: PubChem, ChEBI, ChEMBL, the FDA’s UNII substance registry, CAS.
Natural-product databases for occurrence and classification: COCONUT, LOTUS, TeroKit, KNApSAcK, SuperNatural, Dr. Duke’s phytochemical database.
Biochemical databases for enzymes and reactions: Rhea, BRENDA, UniProt.
Composition databases for essential oils and cannabis: EssoilDB, CannabisDatabase.ca, commercial and state laboratory result sets.
Assay and target data: PubChem BioAssay, ChEMBL, and terpene–protein interaction sets.
Regulatory lists: FEMA GRAS, FDA SCOGS and food-additive listings, EU flavouring and cosmetic-allergen annexes, state cannabis rules.
Literature and patents: PubMed, patent full-text indexes, and grey literature such as monographs and ethnobotanical compendia.
Ontologies for biology and disease: MeSH, MONDO, NCBI Taxonomy, Cellosaurus.
Each has its own identifier scheme, update cadence, file format, and license. None of them is terpene-focused. Several of them disagree with one another about names, stereochemistry, and even which compounds count as terpenoids.
Identity¶
Names do not identify molecules. The Terpedia record for limonene carries 58 aliases, including (+)-limonene, (R)-(+)-p-mentha-1,8-diene, d-limonene, dipentene, and the FEMA and CAS designations. Certificates of analysis, ingredient declarations, supplier specifications, and papers each use whichever name their authors were trained on. Stereochemistry is often dropped: “limonene” on a COA might be the (R)-enantiomer, the (S)-enantiomer, or an unresolved mixture, and only enantioselective chromatography can tell. Salts, hydrates, glycosides, and esters of the same core appear as separate entries in some databases and are folded together in others.
The only robust identity key is a structure-based one: the standard InChI and its hashed InChIKey. Terpedia unions its sources on exact standard InChI and reports coverage in those terms Terpedia, LLC, 2026. Anything less produces counts and joins that silently mix compounds.
Counting¶
Because identity is unstable, counts are unstable. As How many terpenes are there? described, published totals for “known terpenes” span an order of magnitude, and even a single well-curated database can be counted several ways. In Terpedia’s own counting report, the same TeroKit source yields 145,354 confirmed identities under its explicit terpene/terpenoid categories and a further 23,061 in categories labeled “Others” and “Steroids” that are deliberately excluded as ambiguous Terpedia, LLC, 2026. A general natural-product table is not a terpene table just because some of its entries are terpenes. Any vendor who quotes a single large “compounds in our database” figure without a scope is asking you to trust an unstated counting rule.
Measurement variability¶
Laboratory terpene profiles are the most commercially important terpene data and the least reliable. Terpenes are volatile, oxidize on standing, and are lost to sample preparation; laboratories differ in extraction solvent, instrument (GC-MS versus GC-FID), calibration standards, reporting thresholds, and the list of analytes they quantify. Inter-laboratory comparisons in the cannabis sector routinely show divergent results on split samples, and proficiency-testing programs remain immature relative to established food and pharmaceutical testing.
Terpedia’s own research quantifies how bad this is. In a provenance audit of a commercial cannabis terpene archive, the identity of the source laboratory database was more strongly associated with a sample’s assigned terpene profile class (normalized mutual information 0.324) than the cultivar’s commercial name was (normalized mutual information 0.112), with two laboratories’ profiles mapping almost entirely to a single class McShan & Trapp, 2026. In plain terms: knowing which lab produced a profile told you more about the profile than knowing which cultivar it claimed to describe. This finding led Terpedia to downgrade its own earlier seven-class “terpotype” classifier from a biological claim to a provisional archive partition McShan & Trapp, 2026, an example of the platform applying its evidence rules to itself.
The practical consequences for industry are direct. Cultivar names are not a chemistry. Reference profiles must be built from many samples and must carry laboratory identity as a covariate. And a product’s terpene claims can only be as good as the measurement pipeline behind them.
Claims inflation¶
Effect claims about terpenes propagate through the market by repetition. A receptor-binding result in a cell assay becomes “activates CB2”; that becomes “anti-inflammatory”; that becomes label copy. Each step drops a qualifier. Terpedia’s claims investigation catalogs this drift explicitly: it treats every promotional effect statement as a hypothesis, joins it to whatever receptor or assay evidence exists, and reports two separate fields, whether the compound has any linked mechanism evidence and whether the effect itself is supported by appropriately scoped studies Terpedia, LLC, 2026. A compound can have the first without the second. Most do.
The failure mode is not dishonesty; it is the absence of a data model that distinguishes evidence types. When “associated with CB2” and “clinically shown to reduce anxiety” are stored in the same free-text field, the distinction is gone by the time the content team sees it.
Provenance and licensing¶
Public databases carry heterogeneous licenses. Wikidata is CC0. COCONUT and LOTUS are open. CannabisDatabase.ca is CC BY-NC 4.0, which restricts commercial reuse. BRENDA is license-scoped. Some regulatory and monograph sources are redistributable only as documents. A pipeline that merges these into one undifferentiated table has destroyed the information a legal team needs to approve commercial use. It has also destroyed the information a scientist needs to judge reliability: a source-curated association from a community database and a reviewed UniProt entry are not the same grade of evidence.
Provenance also has a time dimension. Sources refresh; records change; a number that was correct in June may be wrong in September. Without release identifiers, retrieval timestamps, and content hashes on every row, there is no way to reproduce a result or to know whether a discrepancy is a data error or a data update.
The cost of doing it yourself¶
Every organization that works seriously with terpenes has, at some point, built an internal spreadsheet or database that tries to solve the problems above. These efforts share a life cycle: they are built by one motivated scientist, they encode that person’s naming conventions, they fall out of date when sources refresh, and they cannot be defended when a regulator or a customer asks for the source of a specific value. The work is not differentiating; it is the same reconciliation done again in every company.
Commercial chemistry platforms exist but are priced and designed for pharmaceutical discovery, are not terpene-focused, do not carry cannabis or essential-oil composition data, and do not link to product-level certificates of analysis. The gap between “general chemistry database” and “usable terpene knowledge for a product team” is exactly the gap Terpedia fills.
- Terpedia, LLC. (2026). How many terpenes are there? A source-audited census of candidate terpene and terpenoid identities. Working draft, 4 September 2026. https://github.com/Terpedia/census
- Terpedia, LLC. (2026). Terpedia terpene/terpenoid counts. Entity counting report, retrieved 4 September 2026 from the Terpedia Knowledge API. https://github.com/Terpedia/kb
- McShan, D. C., & Trapp, S. (2026). Provenance audit of a commercial cannabis terpene archive: database identity versus commercial name. Companion manuscript to the Terpotype classifier, 4 September 2026. https://github.com/Terpedia/strain/blob/main/manuscript/cannabis_archive_audit_article.md
- McShan, D. C., & Trapp, S. (2026). A name-independent terpene-profile classification of commercial Cannabis sativa: the Terpedia Terpotype Classifier (TTC-7). https://github.com/Terpedia/strain/blob/main/manuscript/terpotype_classification_article.md
- Terpedia, LLC. (2026). Terpedia claims investigation: promotional terpene effects recast as hypotheses. https://github.com/Terpedia/claims