下拉刷新
Repository Details
Shared bynavbar_avatar
repo_avatar
HelloGitHub Rating
0 ratings
AI Agent Evaluation Framework
FreeApache-2.0
Claim
Collect
Share
3.6k
Stars
No
Chinese
Python
Language
Yes
Active
265
Contributors
592
Issues
Yes
Organization
None
Latest
1k
Forks
Apache-2.0
License
More
This is an open-source framework for evaluating AI agents and large models. You can use a single command to spin up agents like Claude Code and Codex CLI in Docker and run benchmarks with evaluation sets such as SWE-Bench. If local speed is slow, you can connect to cloud platforms like Daytona to run multiple environments simultaneously.
Included in:
Vol.124
Tags:
AI Agent
AI
Python

Comments

Rating:
No comments yet