Pulsar AI benchmark

The held-out set is the network’s golden calibration tasks — questions with a known answer, graded the moment they are answered. We publish the raw score, whatever it is; no cherry-picking, no invented numbers.

AI agreement with human consensus

Human accuracy on held-out goldens (30d)

0

labels resolved by consensus

Not enough graded data yet — this benchmark fills in as the network answers calibration tasks.

Method: golden tasks carry a known answer and are graded at submit time (is_gold_pass); the AI’s suggested answer is compared to the human-consensus result on the same tasks. Scores update continuously.

What this page answers

  • AI against human consensus
  • measured model accuracy
  • AI quality benchmark
  • calibration tasks and score
  • what is the Pulsar AI benchmark
  • how the Pulsar AI benchmark works
  • the Pulsar AI benchmark explained
  • the Pulsar AI benchmark guide
  • how to start with the Pulsar AI benchmark
  • the Pulsar AI benchmark reviews
  • is the Pulsar AI benchmark worth it
  • the Pulsar AI benchmark pricing
  • the Pulsar AI benchmark alternatives
  • the Pulsar AI benchmark for beginners
  • why the Pulsar AI benchmark
  • the Pulsar AI benchmark — what you get