—
AI agreement with human consensus
—
Human accuracy on held-out goldens (30d)
0
labels resolved by consensus
Not enough graded data yet — this benchmark fills in as the network answers calibration tasks.
Method: golden tasks carry a known answer and are graded at submit time (is_gold_pass); the AI’s suggested answer is compared to the human-consensus result on the same tasks. Scores update continuously.
What this page answers
- AI against human consensus
- measured model accuracy
- AI quality benchmark
- calibration tasks and score
- what is the Pulsar AI benchmark
- how the Pulsar AI benchmark works
- the Pulsar AI benchmark explained
- the Pulsar AI benchmark guide
- how to start with the Pulsar AI benchmark
- the Pulsar AI benchmark reviews
- is the Pulsar AI benchmark worth it
- the Pulsar AI benchmark pricing
- the Pulsar AI benchmark alternatives
- the Pulsar AI benchmark for beginners
- why the Pulsar AI benchmark
- the Pulsar AI benchmark — what you get