dbeley/nixos-benchmark

An all-in-one nix shell to benchmark your system and compare your results.

GitHub repository with 13 stars and 1 forks.

Language: Python

Topics: benchmark, benchmarking, nix, nixos

Open provider repository

Latest metric snapshot

2026-06-05: 13 stars and 1 forks.

Similar repositories

1. VibeBench/VibeSearchBench

🔍 The hardest search benchmark in the wild — vague, multi-turn, proactive. 200 long-horizon tasks with persona-driven progressive disclosure, scored by verifiable schema-free knowledge-graph evaluation. No vibes, just triplet F1.

GitHub repository with 780 stars and 9 forks.

Trending score: 1.88; stars gained: +102; forks gained: +0.

Language: Python

Topics: agentic-ai, benchmark, llm, proactive-agent, search, search-agent
2. sierra-research/tau2-bench

τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

GitHub repository with 1,273 stars and 328 forks.

Trending score: 0.92; stars gained: +7; forks gained: +1.

Language: Python

Topics: benchmark, llm, ai, language-model-agent, conversational-agents
3. Ropedia/SpatialBench

SpatialBench: Is Your Spatial Foundation Model an All-Round Player?

GitHub repository with 93 stars and 0 forks.

Trending score: 0.74; stars gained: +5; forks gained: +0.

Language: Python

Topics: 3d-reconstruction, benchmark, spatial-foundation-model
4. karamouche/noisekit

Generate degraded speech datasets for noise-robust ASR benchmarking

GitHub repository with 15 stars and 0 forks.

Trending score: 0.50; stars gained: +2; forks gained: +0.

Language: Python

Topics: asr, audio, audiomentations, benchmark, cli, dataset-generation
5. rollinsio/beyond-test-coverage

Benchmark for the quality of LLM-generated test suites — anti-fragility, rigor, mocking discipline, reuse — scored against human baselines, not coverage. Python, JS/TS, Go.

GitHub repository with 18 stars and 1 forks.

Trending score: 0.49; stars gained: +2; forks gained: +0.

Language: Python

Topics: benchmark, claude, code-quality, llm, mocha, pytest
6. outsourc-e/bench-loop

Local-first CLI for benchmarking LLMs on real hardware — quality, speed, reliability, and a real multi-turn agent loop.

GitHub repository with 32 stars and 6 forks.

Trending score: 0.46; stars gained: +2; forks gained: +0.

Language: Python

Topics: agent, benchmark, cli, evaluation, llm, local-llm

Trending in Python

1. NousResearch/hermes-agent

The agent that grows with you

GitHub repository with 182,705 stars and 31,323 forks.

Trending score: 5.95; stars gained: +1,867; forks gained: +361.

Language: Python

Topics: ai, ai-agent, ai-agents, anthropic, chatgpt, claude
2. Imbad0202/academic-research-skills

Academic Research Skills for Claude Code: research → write → review → revise → finalize

GitHub repository with 27,643 stars and 2,276 forks.

Trending score: 5.52; stars gained: +1,079; forks gained: +89.

Language: Python

Topics: academic-pipeline, academic-writing, ai-research, claude, claude-code, literature-review
3. rohitg00/ai-engineering-from-scratch

Learn it. Build it. Ship it for others.

GitHub repository with 28,771 stars and 4,705 forks.

Trending score: 5.32; stars gained: +1,261; forks gained: +238.

Language: Python

Topics: agents, ai, ai-agents, ai-engineering, computer-vision, course
4. vinta/awesome-python

An opinionated list of Python frameworks, libraries, tools, and resources

GitHub repository with 301,435 stars and 28,046 forks.

Trending score: 4.60; stars gained: +518; forks gained: +24.

Language: Python

Topics: awesome, collections, python, python-frameworks, python-libraries, python-tools
5. Alishahryar1/free-claude-code

Use claude-code for free in the terminal, VSCode extension or discord like OpenClaw (voice supported)

GitHub repository with 32,540 stars and 4,942 forks.

Trending score: 4.56; stars gained: +467; forks gained: +82.

Language: Python
6. langchain-ai/langchain

The agent engineering platform.

GitHub repository with 138,601 stars and 22,962 forks.

Trending score: 4.53; stars gained: +171; forks gained: +31.

Language: Python

Topics: ai, anthropic, gemini, langchain, llm, openai

dbeley/nixos-benchmark

Latest metric snapshot

Similar repositories

1. VibeBench/VibeSearchBench

2. sierra-research/tau2-bench

3. Ropedia/SpatialBench

4. karamouche/noisekit

5. rollinsio/beyond-test-coverage

6. outsourc-e/bench-loop

Trending in Python

1. NousResearch/hermes-agent

2. Imbad0202/academic-research-skills

3. rohitg00/ai-engineering-from-scratch

4. vinta/awesome-python

5. Alishahryar1/free-claude-code

6. langchain-ai/langchain

Trending topic: benchmark

1. Purewhiter/mobilegym

2. VibeBench/VibeSearchBench

3. sierra-research/tau2-bench

4. ZeptoSeniorMoat/Cinebench-Advanced-Edition-Portable

5. Ropedia/SpatialBench

6. LimeLevelLearn/AIDA64-Extreme-Cracked-2026