01
Benchmarking
NEXBENCH 2.1 is now public
Builders can measure whether an autonomous Web3 agent is both capable and safe across execution, swaps, bridging, DeFi, research, security, portfolio analysis, and governance.
- A public leaderboard, agent comparison view, submission flow, and methodology guide now share the same hash-sealed run manifests.
- The zero-runtime-dependency CLI can scaffold an adapter, run the public suite, verify results, mint a valid manifest, and submit it for review.
- Deterministic forks, frozen research corpora, safety-zeroed scoring, and programmatic graders keep results reproducible and harder to game.