dbeley/nixos-benchmark
An all-in-one nix shell to benchmark your system and compare your results.
GitHub repository with 13 stars and 1 forks.
Language: Python
Topics: benchmark, benchmarking, nix, nixos
An all-in-one nix shell to benchmark your system and compare your results.
GitHub repository with 13 stars and 1 forks.
Language: Python
Topics: benchmark, benchmarking, nix, nixos
2026-06-05: 13 stars and 1 forks.
🔍 The hardest search benchmark in the wild — vague, multi-turn, proactive. 200 long-horizon tasks with persona-driven progressive disclosure, scored by verifiable schema-free knowledge-graph evaluation. No vibes, just triplet F1.
GitHub repository with 780 stars and 9 forks.
Trending score: 1.88; stars gained: +102; forks gained: +0.
Language: Python
Topics: agentic-ai, benchmark, llm, proactive-agent, search, search-agent
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
GitHub repository with 1,273 stars and 328 forks.
Trending score: 0.92; stars gained: +7; forks gained: +1.
Language: Python
Topics: benchmark, llm, ai, language-model-agent, conversational-agents
SpatialBench: Is Your Spatial Foundation Model an All-Round Player?
GitHub repository with 93 stars and 0 forks.
Trending score: 0.74; stars gained: +5; forks gained: +0.
Language: Python
Topics: 3d-reconstruction, benchmark, spatial-foundation-model
Generate degraded speech datasets for noise-robust ASR benchmarking
GitHub repository with 15 stars and 0 forks.
Trending score: 0.50; stars gained: +2; forks gained: +0.
Language: Python
Topics: asr, audio, audiomentations, benchmark, cli, dataset-generation
Benchmark for the quality of LLM-generated test suites — anti-fragility, rigor, mocking discipline, reuse — scored against human baselines, not coverage. Python, JS/TS, Go.
GitHub repository with 18 stars and 1 forks.
Trending score: 0.49; stars gained: +2; forks gained: +0.
Language: Python
Topics: benchmark, claude, code-quality, llm, mocha, pytest
Local-first CLI for benchmarking LLMs on real hardware — quality, speed, reliability, and a real multi-turn agent loop.
GitHub repository with 32 stars and 6 forks.
Trending score: 0.46; stars gained: +2; forks gained: +0.
Language: Python
Topics: agent, benchmark, cli, evaluation, llm, local-llm
The agent that grows with you
GitHub repository with 182,705 stars and 31,323 forks.
Trending score: 5.95; stars gained: +1,867; forks gained: +361.
Language: Python
Topics: ai, ai-agent, ai-agents, anthropic, chatgpt, claude
Academic Research Skills for Claude Code: research → write → review → revise → finalize
GitHub repository with 27,643 stars and 2,276 forks.
Trending score: 5.52; stars gained: +1,079; forks gained: +89.
Language: Python
Topics: academic-pipeline, academic-writing, ai-research, claude, claude-code, literature-review
Learn it. Build it. Ship it for others.
GitHub repository with 28,771 stars and 4,705 forks.
Trending score: 5.32; stars gained: +1,261; forks gained: +238.
Language: Python
Topics: agents, ai, ai-agents, ai-engineering, computer-vision, course
An opinionated list of Python frameworks, libraries, tools, and resources
GitHub repository with 301,435 stars and 28,046 forks.
Trending score: 4.60; stars gained: +518; forks gained: +24.
Language: Python
Topics: awesome, collections, python, python-frameworks, python-libraries, python-tools
Use claude-code for free in the terminal, VSCode extension or discord like OpenClaw (voice supported)
GitHub repository with 32,540 stars and 4,942 forks.
Trending score: 4.56; stars gained: +467; forks gained: +82.
Language: Python
The agent engineering platform.
GitHub repository with 138,601 stars and 22,962 forks.
Trending score: 4.53; stars gained: +171; forks gained: +31.
Language: Python
Topics: ai, anthropic, gemini, langchain, llm, openai
MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research · 浏览器里运行的安卓模拟器 · Browser-hosted Android Simulator · Verifiable Evaluation · Scalable Online RL Training
GitHub repository with 527 stars and 83 forks.
Trending score: 3.00; stars gained: +33; forks gained: +4.
Language: TypeScript
Topics: agent, agents, ai, android, automation, benchmark
🔍 The hardest search benchmark in the wild — vague, multi-turn, proactive. 200 long-horizon tasks with persona-driven progressive disclosure, scored by verifiable schema-free knowledge-graph evaluation. No vibes, just triplet F1.
GitHub repository with 780 stars and 9 forks.
Trending score: 1.88; stars gained: +102; forks gained: +0.
Language: Python
Topics: agentic-ai, benchmark, llm, proactive-agent, search, search-agent
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
GitHub repository with 1,273 stars and 328 forks.
Trending score: 0.92; stars gained: +7; forks gained: +1.
Language: Python
Topics: benchmark, llm, ai, language-model-agent, conversational-agents
Cinebench Advanced Edition Portable with extended test profiles, command-line runner, and comparison charts—full benchmark toolkit unlocked.
GitHub repository with 26 stars and 0 forks.
Trending score: 0.84; stars gained: +6; forks gained: +0.
Topics: advanced-edition, benchmark, cinebench, cpu, gpu, hardware
SpatialBench: Is Your Spatial Foundation Model an All-Round Player?
GitHub repository with 93 stars and 0 forks.
Trending score: 0.74; stars gained: +5; forks gained: +0.
Language: Python
Topics: 3d-reconstruction, benchmark, spatial-foundation-model
AIDA64 Extreme hardware diagnostics benchmarks and sensor monitoring for PCs.
GitHub repository with 29 stars and 0 forks.
Trending score: 0.59; stars gained: +3; forks gained: +0.
Topics: aida64, benchmark, cpu, diagnostics, gpu, hardware