whilehq/whileai-sdk
Scientific RL and SFT post-training for AI agents: build evals that can fail, simulate situations, judge checked against people, pass@1 with a 95% interval, train with GRPO/SFT/DPO, prove every gain on a held-out set. uv add whileai
GitHub repository with 15 stars and 1 forks.
Language: Python
Topics: agent-evals, agents-md, ai-agents, claude-code, dpo, evaluation, grpo, llm-as-a-judge, llm-evals, pass-at-k