ZJU-REAL/ageval
Agent eval on one running base. Swap the agent under test with plugins; run the same dataset anywhere.
GitHub repository with 71 stars and 0 forks.
Language: Python
Topics: agents, benchmark, cli, evals, evaluation, plugins, python, reproducibility