How the numbers work

Every ranking on this site is one line of arithmetic, printed on the page that uses it. This is that line, and the reasoning behind it.

The problem this site exists to fix

Consider two strains and one effect:

Strain A energetic 64% of 11 reports Strain B energetic 52% of 8,431 reports

Sorted by percentage, Strain A wins. It should not. Seven people out of eleven is what noise looks like -- 64% there is not a measurement of Strain A, it is barely a measurement at all. Sort by effect on any strain site and watch the top of the list fill with strains nobody has heard of, each with a handful of reviews.

What we do instead

Each rate is pulled toward the database-wide average for that effect, by an amount that depends on how little evidence stands behind it:

adjusted = (k + m * p0) / (n + m) k reports of this effect for this strain n total reports for this strain p0 average rate for this effect across the whole database m 50, the amount of evidence needed to move halfway off p0

With a database average of 30%, the two strains above become 36% and 52%, and the order is right. Strain B's estimate barely moved, because 8,431 reports are plenty. Strain A's fell most of the way back to average, because eleven reports are not.

Two things this buys, for free

A strain nobody has reviewed scores as average, not as zero. Put n = 0 into that formula and it returns exactly p0. This matters most for the effects people sort to avoid: in a naive model, a strain with no data has "0% anxiety" and wins the low-anxiety sort outright. Absence of evidence must never read as good news, and here it cannot.

There is no special case. One report and ten thousand go through the same line, so there is no threshold to get wrong.

Turning that into an order

Your weights and the adjusted rates make the score:

score = SUM over effects of weight * (adjusted_rate - average_rate)

Each term measures a strain against a typical strain rather than against zero. A score of 0 means average on everything you asked about; positive means better than typical for your weights. Because an effect with no data returns exactly the average, its term is exactly zero -- so a strain can never climb your ranking by being under-researched.

There is no 0-100 match percentage, because there is no ceiling to measure against: the maximum depends on weights you picked seconds ago. The table ranks, shows each term, and lets you check it.

Confidence

Every pooled rate carries a tier, derived from the data rather than assigned:

Two independent sources are required for the top tier because sample size alone cannot catch a methodology problem. Forty thousand reports from one site with one checkbox layout are forty thousand reports of that checkbox layout.

What this does not fix

Bias. Shrinking estimates fixes variance, not slant. If a source's reviewers skew toward people who liked a strain enough to write about it, more reports buy a more precise estimate of a biased quantity. Nothing in the arithmetic detects that, and this site does not pretend otherwise.

Double counting. Sources are pooled by summing their counts, which weights each by its sample size. If two sources syndicate the same underlying reviews, that counts them twice, and we cannot see it from the outside. So the number of distinct sources is shown beside every pooled figure, and each source's own numbers stay visible on the strain page.

Correlated effects. Energetic and sleepy are close to opposites. Weighting both is expressing roughly one preference twice, and it counts twice. Correcting for that properly would need a covariance matrix and would make the score impossible to print. Printing the score is worth more.

The name on the jar. This is the largest limitation and no statistic touches it. Strain names are not standardised, not certified, and not enforced. Two jars labelled with the same name, from different growers, can differ more from each other than from a third strain entirely. The chemistry section shows ranges across samples rather than a single figure for exactly this reason, and a range is the honest shape of that data.

Ranking on chemistry

Laboratory measurements are not percentages of people, so they get their own version of the same idea. Each strain's median is pulled toward the typical strain by how little evidence supports it:

shrunk = typical + (this strain's median - typical) * n / (n + m)

Here m is not chosen. It is calculated, separately for each ingredient, as the ratio of how much samples of one strain scatter to how much strains differ from each other. That ratio is the honest answer to "how many lab results before I believe this strain is really different from average".

What it returned is worth stating, because it was not what we expected. For THC it is 0.73; for limonene 0.66; for caryophyllene 1.02. All far below the figure the same calculation gives for self-reported effects. In plain terms: a laboratory measuring the same strain twice agrees far more than two people describing it do, so chemistry needs comparatively little correction, and a handful of lab results genuinely does pin a strain down. The correction still matters at the extreme, which is where it was always needed: three samples no longer outrank five hundred.

Why every ingredient is scored in standard deviations

THC runs around 18%, limonene around 0.2%. A weighted total over raw percentages would be a THC ranking with a rounding error attached, and a request for more limonene would never visibly change the order. So each contribution is expressed as how far the strain sits from typical, measured in the spread between strains:

score = SUM of weight * (shrunk - typical) / spread

A term then reads as "this strain sits 0.8 standard deviations above a typical strain on limonene", which is comparable across ingredients and printable on the page. An ingredient never measured for a strain returns exactly the typical value, so its term is exactly zero and no strain can climb a ranking by being under-tested.

The published range is the middle 80%, not the extremes

Raw minimum and maximum are useless here, and the reason is instructive. Blue Dream's 1,793 results run from 0.4% to 32.7% THC — but six of those samples are under 5%, which is not flower anyone sold as Blue Dream, and one is over 30%. Seven bad rows out of 1,793 would define both ends of the published figure. The 10th to 90th percentile for the same strain is 14.4% to 22.2%. Raw extremes are kept as a way to spot bad data, never as the range shown to a reader.

What the research says about strain names

The largest limitation on this site is not statistical, and no amount of arithmetic touches it: a strain name is not a regulated, certified or standardised thing. Varietal names cannot be trademarked in the United States, federal trademark protection is unavailable for cannabis goods anyway, and plant variety protection reaches only hemp. There is no registry and no authority.

The published research is consistent about what follows from that:

This site shows ranges rather than single numbers because of that last point, and marks its classification labels as "how it is marketed" because of the first. None of it makes strain names useless — they are how people shop, and comparing them is the point of this tool. It does mean a number here describes what was measured in samples carrying that name, which is a weaker and more honest claim than it first appears.

Where the numbers come from

Every effect rate, every lab figure and every claim about lineage is stored with a source, a URL, a sample size and the date it was verified. Anything missing one of those fails the build and does not appear. The status page reports how much of the database is thinly evidenced, because a site that hides its own weak spots is asking to be taken on faith.