#llm-evaluation
Open source repositories tagged with #llm-evaluation, ranked by health score.
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.
A full-stack AI Red Teaming platform securing AI ecosystems via Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.
Creating realistic RL environments requires collaboration between researchers, engineers, and domain experts across many dimensions: artifacts, environments tools, dynamism of the environment, reproducibility, and more. There is no open source framework for building these environments effectively. Until now.
Dart framework for stateful AI agents: tool use, skills, sub-agent delegation, planning, streaming, evals, and multi-provider LLM support.
Agentic RL 中文零基础教程(33 章):从概念到 GRPO 实战,含 TRL 最小可跑示例;26–33 章附一套可运行的三方判别模型实证工程(encoder vs LLM-LoRA vs 规则基线)。第 25 章讲清 Jev / TypeSafe System One 与 RL 的能力边界 | Chinese Agentic RL tutorial (33 chapters) + a reproducible discriminative-model benchmark