A knowledge platform earns trust by being clear about what it will not say and where it is incomplete. This section states Terpedia’s evidence rules and its known limitations as of September 2026.
The rules¶
Terpedia applies the same rules to its data model, its API, its agents, and its own publications.
A database record is not a validated natural product. A compound’s presence in COCONUT, TeroKit, or any other source, including Terpedia’s own union, means a source classified it as a terpenoid. It does not mean the structure has been independently verified, isolated, or shown to occur in nature. Counts are reported as candidate identities with their identity key.
An occurrence is not a concentration. A found_in_taxon statement means
a cited work reported detecting the compound in that organism. It says
nothing about level, tissue, or commercial viability.
A receptor association is not efficacy. A source-curated compound–protein link, a binding constant, or an assay hit describes a molecular interaction under specified conditions. It is not evidence of a physiological effect, a therapeutic benefit, or human relevance. Terpedia’s claims register tracks mechanism evidence and effect support as separate fields for exactly this reason Terpedia, LLC, 2026.
A docking score is not affinity. Computational predictions are ranked hypotheses for experimental testing. They are labeled as such and are never merged into the same field as measured interactions.
A graph path is not proof of in-vivo production. Linking a genome, an enzyme, and a product in a pathway graph shows that a route is plausible, not that an organism uses it.
A cultivar name is not a chemistry. Names are attributes of measured records, not sources of chemical information; where Terpedia’s own research found that laboratory identity predicted profile shape better than cultivar name, it revised its own classifier accordingly McShan & Trapp, 2026.
A claim is a hypothesis until scoped evidence supports it. Promotional effect statements are catalogued as hypotheses with an explicit uncertainty boundary and, where possible, a falsifiable test.
Every response carries its boundary. The context block returned by the
knowledge API opens with the instruction not to extend claims beyond the
cited evidence. Downstream language models are constrained by the data
layer, not only by their prompts.
Known limitations¶
Terpedia’s coverage reports are public, and the following gaps are real.
Resolver coverage is incomplete. The open resolver currently returns no entry for some compounds that are clearly in scope, including thujone and artemisinin, because the resolved-entity layer has been populated from specific source adapters (CannabisDatabase.ca, the terpene registry, EssoilDB, curated core) and has not yet been extended across the full BigQuery census union. The live-examples section shows a miss deliberately. Coverage is versioned and expanding; the SQL layer already holds the structures.
Some labels are unresolved. Occurrence statements sourced from Wikidata sometimes display a bare Q-identifier where the taxon label has not yet been fetched. The identifier is correct and resolvable; the presentation is incomplete.
Source encoding defects propagate. At the time of writing, some CannabisDatabase.ca descriptions carry a character-encoding fault (Greek letters rendered as two-character sequences). This originates in ingestion and is being corrected; it is mentioned here because a platform that promises provenance should also disclose its own defects.
Ingestion quality varies by source. The coverage matrix records, for example, 39 errors in the latest EssoilDB run and a large number of reference and row errors in the SAIR interaction set. Validated tables are marked validated; partial and raw-only sources are marked as such. Consumers should read the status column, not just the source name.
Several adapters are planned, not live. A release-aware UniProt adapter, the remaining PubChem RDF shards, and the Fuseki promotion of MeSH and MONDO are in progress. LOTUS structure serving is validated; the full occurrence metadata promotion awaits a complete upstream archive.
Licensing constrains reuse. CannabisDatabase.ca content is CC BY-NC 4.0 and BRENDA is license-scoped. Commercial users must filter or license accordingly. Terpedia keeps the license on every record so that this is possible, and will not present license-restricted data as if it were open.
Counts are observations, not properties. Every figure in this paper is a dated observation from a specific scope. When a source refreshes, the count is re-run and appended, not overwritten. Readers quoting Terpedia numbers should quote the date and scope with them.
Terpedia is not a medical or legal authority. The platform does not diagnose, recommend treatment, pre-clear claims, or certify compliance. Its agents are configured to say so.
Why this section is in a marketing document¶
Buyers of scientific infrastructure have learned to discount vendor claims. The most persuasive thing Terpedia can show a quality director or a regulatory lead is not a large number; it is a platform that already behaves the way their auditors will want it to. Every limitation above is visible in the product today, and every rule above is enforced in the data rather than promised in a slide.
- Terpedia, LLC. (2026). Terpedia claims investigation: promotional terpene effects recast as hypotheses. https://github.com/Terpedia/claims
- McShan, D. C., & Trapp, S. (2026). Provenance audit of a commercial cannabis terpene archive: database identity versus commercial name. Companion manuscript to the Terpotype classifier, 4 September 2026. https://github.com/Terpedia/strain/blob/main/manuscript/cannabis_archive_audit_article.md