facebookresearch/swe-sweep
How many bugs can LMs find & fix in large codebases?
GitHub repository with 34 stars and 3 forks.
Language: Python
Topics: ai, ai-agents, benchmark, benchmarking, harbor, harbor-framework, llm, llm-benchmark, llm-benchmarking, llm-benchmarks