Control center / Discover

Discover

Browse by category instead of drowning in search results. Every category opens a submenu of subcategories with live counts, and checking one narrows the grid to exactly that slice.

Verified against GitHub · scores pending
15 results for “Evals” across 2 categories
Sort by:Type:Trust:Maturity:15 shown

Token & Cost Optimization

1

TensorZero

tensorzero12kToolstable

Open source LLM gateway and optimization framework unifying inference, observability and evals.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
12k
Downloads
Updated
2026-06-11
Skill installs
Trust pendingMomentum not matched

LLM Ops & Observability

14

AutoEvals

braintrustdata1.0kRepositorystable

Open source tool for evaluating AI model outputs using established best practices.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
1.0k
Downloads
Updated
2026-07-29
Skill installs
Trust pendingMomentum not matched

OpenAI Evals

openai19kRepositoryexperimental

Framework for creating and running evaluations on large language models and systems.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
19k
Downloads
2.5k
Updated
2026-04-14
Skill installs
Trust pendingMomentum not matched

Opik

comet-ml22kToolstable

Open-source LLM evaluation, tracing, and monitoring by Comet.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
22k
Downloads
3527k
Updated
2026-09-08
Skill installs
Trust pendingMomentum not matched

Phoenix

Arize-ai11kToolstable

Open-source LLM tracing, evaluation, and observability platform.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
11k
Downloads
2215k
Updated
2026-09-08
Skill installs
Trust pendingMomentum not matched

Weave

wandb1.1kToolexperimental

Toolkit to track, evaluate, and monitor LLM applications by W&B.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
1.1k
Downloads
947k
Updated
2026-09-07
Skill installs
Trust pendingMomentum not matched

Promptfoo

promptfoo25kToolexperimental

Test, evaluate, and red-team LLM prompts and apps from the CLI.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
25k
Downloads
Updated
2026-09-08
Skill installs
Trust pendingMomentum not matched

Evidently

evidentlyai7.9kToolexperimental

Open source framework to evaluate and monitor machine learning and LLM systems.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
7.9k
Downloads
Updated
2026-08-31
Skill installs
Trust pendingMomentum not matched

DeepEval

confident-ai18kToolstable

Open source evaluation framework for unit testing large language model outputs.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
18k
Downloads
6221k
Updated
2026-09-06
Skill installs
Trust pendingMomentum not matched

LM Evaluation Harness

EleutherAI14kRepositoryexperimental

Unified framework for evaluating language models across hundreds of benchmark tasks.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
14k
Downloads
1512k
Updated
2026-09-01
Skill installs
Trust pendingMomentum not matched

TruLens

truera3.5kRepositorystable

Library for evaluating and tracking the performance of LLM applications with feedback functions.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
3.5k
Downloads
97k
Updated
2026-09-04
Skill installs
Trust pendingMomentum not matched

Prompt Flow

microsoft11kToolactive

Toolkit for building, evaluating and deploying prompt based LLM application workflows.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
11k
Downloads
79k
Updated
2026-08-26
Skill installs
Trust pendingMomentum not matched

Giskard

Giskard-AI5.8kRepositorystable

Open source testing framework to detect vulnerabilities and evaluate quality of LLM and ML models.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
5.8k
Downloads
25k
Updated
2026-09-07
Skill installs
Trust pendingMomentum not matched

Ragas

explodinggradients16kToolslowing

Framework for evaluating retrieval augmented generation pipelines with reference free metrics.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
16k
Downloads
1582k
Updated
2026-02-24
Skill installs
Trust pendingMomentum not matched

PromptTools

hegelai3.1kToolslowing

Open source tools for testing and experimenting with LLM prompts and vector databases.

Best match because its name, category, capabilities, or owner matches “Evals”. Adoption and trust break close ties.

Stars
3.1k
Downloads
82
Updated
2026-02-11
Skill installs
Trust pendingMomentum not matched
Discover · SkillPilot