Smiling Buddha Open Research
Reproducible baselines

Physical AI benchmarks.

Standardised evaluation suites for physical AI systems. Each benchmark ships with data, protocol, reference implementation, and a public leaderboard — as a complete bundle only.

Benchmark suites · In preparation

The first Smiling Buddha benchmark suite is being defined by founding contributors. Each benchmark ships with data, evaluation protocol, reference implementation, and public leaderboard — released as a complete bundle only when the protocol has passed community review. Apply as founding contributor →

Suites in preparation

Five benchmark suites

SB-Perception-01

Long-horizon object tracking in cluttered environments

Track objects through occlusion, motion blur, and lighting changes across continuous captures.

In developmentFounding contributors call open
SB-Manip-03

Precision insertion under partial observability

Fork-and-pallet-jack style insertion tasks under camera occlusion. Physical fidelity, tolerance, and cycle time.

In developmentFounding contributors call open
SB-Nav-02

Dynamic-environment path planning

Navigate through a busy warehouse simulated distribution centre. Safety, efficiency, and stop-events.

In developmentFounding contributors call open
SB-Cleaning-01

Floor-cleaning coverage under obstacle drift

Achieve target coverage while obstacles change positions between passes. Coverage, redundancy, shine-index outcome.

In developmentFounding contributors call open
SB-Safety-01

Edge-case classification under distribution shift

Classify novel edge cases correctly across held-out distribution shifts. Precision, recall, false-safe rate.

In developmentFounding contributors call open
How they'll work

How the benchmarks will work

Once the protocol has passed founding-contributor review and the reference bundle is complete, submission will follow a three-step protocol.

01

Download the benchmark bundle

Each benchmark will ship with data, evaluation code, reference baseline, and reproducibility protocol. All MIT-licensed.

02

Run the reference protocol

Follow the documented protocol so results are comparable to every other submission. No hyperparameter tuning post-hoc.

03

Submit results + code

Push to the benchmark's repository. Community-reviewed; verified submissions appear on the public leaderboard.