Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Abstract

Terpenes are the largest family of natural products and the working vocabulary of the cannabis, hemp, flavor, fragrance, essential-oil, functional-food, and natural-product-pharmaceutical industries. The information needed to make terpene decisions is scattered across dozens of databases that do not share identifiers, names, licenses, or standards of evidence, and laboratory terpene profiles vary more by laboratory than by product. Terpedia is a provenance-first knowledge platform that reconciles more than thirty sources on structural chemical identity, attaches evidence type and license to every statement, and exposes the result through a browsable encyclopedia, a documented API, domain agents, and a product-and-certificate catalog. This paper explains the chemistry, the commercial landscape and its evidence, the data problem, the platform, and concrete industry workflows, with runnable examples executed against the live Terpedia Knowledge API.

Keywords:terpenesterpenoidsknowledge graphprovenancecertificate of analysiscannabisflavor and fragranceessential oils

Terpenes are the largest family of natural products and the working vocabulary of several industries at once. They are the volatile signature of a cannabis cultivar, the top notes of a fragrance, the citrus in a beverage, the active principle in a chest rub, and the scaffold of two of the most important drugs of the last century. Product decisions worth billions of dollars a year turn on which terpenes are present, at what levels, and what can truthfully be said about them.

The information needed to make those decisions is scattered across dozens of public chemistry, biology, and literature databases that do not share identifiers, naming conventions, licenses, or standards of evidence. A single molecule can carry fifty synonyms. Cultivar names carry almost no chemical information. Laboratory terpene profiles vary more by which lab produced them than by what they claim to measure. Promotional “effects” propagate through the market faster than the evidence behind them. Anyone who has tried to build a terpene-aware product, quality program, or content library knows the cost: weeks of manual reconciliation, brittle spreadsheets, and claims that cannot be defended when a regulator or a customer asks for the source.

Terpedia is a knowledge platform built to remove that cost. It ingests more than thirty public and licensed sources into a single provenance-preserving knowledge base, resolves names to chemical identities, links compounds to organisms, proteins, products, measurements, and literature, and exposes the result through a browsable encyclopedia (kb.terpedia.com), a documented API, a suite of domain agents (chat.terpedia.com), a product-and-certificate-of-analysis catalog (Terproduct), and a research program that publishes its methods.

Terpedia’s differentiator is not that it has the most records. It is that every record carries its source, its version, its license, and its evidence type, and that the platform is engineered to keep claims smaller than the data. A database record is not a validated natural product. A receptor association is not efficacy. A docking score is not affinity. A cultivar name is not a chemistry. Those distinctions are enforced in the data model and surface in every API response, which is what makes the output usable in regulated settings.

What this paper covers

  1. The chemistry. What terpenes are, how plants make them, why there are so many, and which physical properties matter for products.

  2. The business landscape. Where terpenes create value in cannabis and hemp, flavor and fragrance, essential oils and aromatherapy, functional foods, pharmaceuticals, and industrial biotechnology, and what the evidence actually supports.

  3. The data problem. Fragmentation, identity, counting, measurement variability, claims inflation, and licensing.

  4. The platform. Terpedia’s architecture, sources, scale, access surfaces, partnerships, and research portfolio.

  5. Live examples. Runnable queries against the public Terpedia Knowledge API, executed when this document was built and re-runnable in the browser.

  6. Use cases. Concrete workflows for formulation, quality, regulatory, content, sourcing, IP, and data science teams.

  7. Evidence rules and limitations. What Terpedia will and will not assert, and where coverage is still incomplete.

  8. Getting started. Integration paths and the roadmap.

Key numbers

All counts are dated, scoped observations from Terpedia’s own counting reports, not marketing totals; the scope and identity key for each is given in Scale, stated carefully.

Measure

Value (September 2026)

Candidate terpene/terpenoid identities in the two-source (COCONUT + TeroKit) census union

268,924, of which 225,905 have an exact PubChem CID match

Source-declared terpenoid identities from COCONUT

199,234 unique identities

Confirmed terpene/terpenoid identities under the TeroKit classification

145,354 unique identities

LOTUS occurrence triples (compound reported in organism, each with a DOI)

674,422 triples over 227,316 structures, 37,468 organisms, 91,379 papers

Normalized cannabis laboratory terpene measurements

409,655 non-null measurements from 43,018 source rows

CannabisDatabase.ca records and relations

12,039 source records; 54,251 relation records

Public sources in the knowledge base catalog

More than 30 (see Appendix A: Source inventory)

Who should read this

This paper is written for the people who have to make terpene decisions stick: R&D and formulation leads, quality and laboratory managers, regulatory and compliance officers, product and content teams, sourcing managers, and the data and AI groups that support them. It assumes technical literacy but not a chemistry degree. Sections 1 and 2 can be read as a primer; Sections 3 to 7 are the case for Terpedia; the live-examples section is for anyone who wants to see the data rather than read about it.