Super Candy Labs
Beta

CandyBench v0.1: the first candy benchmark

AI labs rank models on benchmarks. We rank candy. Here's the state of the art on all six axes.

Every AI lab publishes benchmarks, so here is ours. CandyBench scores candy on the same six axes our scan reports: sweet, sour, bitter, crunch, melt and chaos. Below are the current leaders on each.

Method, honestly: v0.1 scores are preliminary lab-panel ratings on a written rubric. They are opinions with a number attached. v1.0 will be recomputed from real taster data once enough people have scanned and told us what they love.

Candy names link to retailers. Paid links: we may earn a small commission if you buy, at no cost to you.

Known limitations

  • Two-person panel. Enormous bias toward whatever was in our pantry.
  • “Chaos” is not a recognized sensory attribute. We think it should be.
  • No candy was harmed. Several were eaten.

The full leaderboard, with all six evals, categories and methodology, lives at CandyBench. Disagree? Scan your tongue and add your palate to the dataset. That’s how v1.0 gets better.

ShareXLinkedInText
More from Lab Notes
  1. September 28, 2026UpdatePalate taxonomy v0.5: new type names, and how the lab classifies a palate
  2. September 28, 2026AnnouncementAI is going to destroy the world. First, it's going to make the best candy bar.
  3. September 28, 2026GuideThe Halloween bowl, sorted by palate type