Standardised evaluation suites for physical AI systems. Each benchmark ships with data, protocol, reference implementation, and a public leaderboard — as a complete bundle only.
Track objects through occlusion, motion blur, and lighting changes across continuous captures.
Fork-and-pallet-jack style insertion tasks under camera occlusion. Physical fidelity, tolerance, and cycle time.
Navigate through a busy warehouse simulated distribution centre. Safety, efficiency, and stop-events.
Achieve target coverage while obstacles change positions between passes. Coverage, redundancy, shine-index outcome.
Classify novel edge cases correctly across held-out distribution shifts. Precision, recall, false-safe rate.
Once the protocol has passed founding-contributor review and the reference bundle is complete, submission will follow a three-step protocol.
Each benchmark will ship with data, evaluation code, reference baseline, and reproducibility protocol. All MIT-licensed.
Follow the documented protocol so results are comparable to every other submission. No hyperparameter tuning post-hoc.
Push to the benchmark's repository. Community-reviewed; verified submissions appear on the public leaderboard.