Repository Details
Shared by
HelloGitHub Rating
0 ratings
Free•Apache-2.0
Claim
Discuss
Collect
Share
3.6k
Stars
No
Chinese
Python
Language
Yes
Active
265
Contributors
592
Issues
Yes
Organization
None
Latest
1k
Forks
Apache-2.0
License
More
This is an open-source framework for evaluating AI agents and large models. You can use a single command to spin up agents like Claude Code and Codex CLI in Docker and run benchmarks with evaluation sets such as SWE-Bench. If local speed is slow, you can connect to cloud platforms like Daytona to run multiple environments simultaneously.
Comments
Rating:
No comments yet