📊
Evals & Observability
146 tools
Evaluation harnesses, LLM tracing, monitoring, and prompt/agent observability. Dashboards, judges, replay tools, and regression suites for LLM apps.
B
BenchLocal
BenchLocal
Local SEO benchmarking and reporting platform
1 story
Claude Console
Anthropic
Develop with Claude
1 story
AA-Briefcase
Artificial Analysis
Agentic Knowledge Work Benchmark
0 stories
AEGIS
Applied Theory LLC
AI Agent Compliance, Governance & Control
0 stories
Agent Arena
Arena Intelligence, Inc.
AI Agent Performance Leaderboard
0 stories
Agent Bricks
Databricks
Production AI agents
0 stories
Agent Sandbox
Kubernetes SIG Apps
Kubernetes SIG Apps project for agent sandboxes.
0 stories
Agent Sessions
jazzyalex
Session management for AI agents
0 stories
A
Agent View
Agent View
Agent View software product
0 stories
agent-trace
Open Source
Open-source agent tracing
0 stories
Agentation
dip Corporation
Software product from dip Corporation
0 stories
A
AgentRank
AgentRank
AgentRank
0 stories
AgentsView
Kenn Software LLC
A local-first desktop and web app for browsing, searching, and analyzing your past AI coding sessions.
0 stories
AGNTCY
Outshift by Cisco
Building the Internet of Agents (IoA)
0 stories
AI Gateway
Vercel
The AI Gateway for developers
0 stories
A
AI Usage
AI Usage
AI Usage
0 stories
AI21 Maestro
AI21 Labs
AI workflow orchestration platform
0 stories
aiewf-eval
Daily
A framework for evaluating multi-turn LLM conversations with support for text, realtime audio, and speech-to-speech models.
0 stories
Andon Labs
Andon Labs
AI research lab for agent systems
0 stories
Antithesis
Antithesis
Bug-free systems, unlimited velocity
0 stories
ARC-AGI-3
ARC Prize Foundation
The first interactive reasoning benchmark designed to measure human-like intelligence in AI agents.
0 stories
Arena
Arena Intelligence, Inc.
Crowdsourced AI Model Evaluation Platform
0 stories
Artificial Analysis
Artificial Analysis, Inc.
AI model benchmarks and rankings
0 stories
ASI-Evolve
Generative Artificial Intelligence Research Lab (GAIR)
Let AI Do the Research, You Keep the Insight
0 stories
ASSERT
Microsoft
Adaptive Spec-driven Scoring for Evaluation and Regression Testing
0 stories
Attention Head Visualiser
HeyNEO
Mapping what each head in GPT-2 actually does
0 stories
AutoHypothesis
arteemg
Official repository
0 stories
AutomationBench-AA
Artificial Analysis
Agentic SaaS Workflow Benchmark
0 stories
Baidu Qianfan
Baidu, Inc.
An enterprise-grade, all-in-one large model platform centered on Agent development
0 stories
Benchmarks.AI
Time Doctor
AI benchmarking software
0 stories
Better Agents
LangWatch
Build reliable, testable, production-grade AI agents with Better Agents CLI - the reliability layer for agent development
0 stories
Block-Sparse Featurizers
Goodfire
Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds
0 stories
Braintrust
Braintrust Data, Inc.
Build and monitor reliable AI applications.
0 stories
BridgeBench V3
BridgeMind
The World's #1 Vibe Coding Benchmark
0 stories
BrowseComp-Plus
OpenAI
A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
0 stories
ccusage
Independent
Claude Code usage analyzer
0 stories
Chrome DevTools for agents
Google
Agent-oriented Chrome DevTools
0 stories
Claude Counter
Solo Software
Track Claude usage at a glance
0 stories
Claude Token Count API
Anthropic
Count the number of tokens in a Message.
0 stories
C
ClawdMeter
Independent
Independent software product
0 stories
Code Arena
Arena Intelligence, Inc.
Build & Test with AI Coding Models
0 stories
Context Arena
Context Arena
Unverified software product
0 stories
Context.ai
Explore Interfaces Inc.
AI analytics for conversation data
0 stories
Coval
Coval
AI agent testing platform
0 stories
Cua-Bench
Cua AI, Inc.
Make your agents better at computers
0 stories
CyberBench
Vals AI
Can AI agents find security bugs and patch them?
0 stories
Daybreak
Daybreak
Daybreak
0 stories
Deep Agents Deploy
LangChain
Deploy deep agents with LangChain
0 stories
DFlash
Z Lab
Block Diffusion for Flash Speculative Decoding
0 stories
Dify
LangGenius, Inc.
The Platform for Production-Ready Agentic Workflows.
0 stories
Dogfood
Dogfood
Dogfood software product
0 stories
DSPy
Stanford NLP Group
Program, don’t prompt, your LLMs.
0 stories
EdgeBench
ByteDance Seed
An ultra-long-horizon benchmark built to measure learning from environments
0 stories
EnterpriseRAG-Bench
Onyx
Enterprise RAG benchmark
0 stories
Entire
Entire
Entire
0 stories
eot-bench
LiveKit
The open benchmark for end-of-turn detection
0 stories
Evals
OpenAI
A framework for evaluating language models and LLM-powered systems.
0 stories
Fullstack Code Arena
Arena
Build, Deploy, and Evaluate with Fullstack Code Arena
0 stories
Future AGI
Future AGI
AI Agents hallucinate, fix it faster.
0 stories
G
GeneBench-Pro
OpenAI
A research-level benchmark measuring how AI agents navigate ambiguity and make consequential judgments in computational biology.
0 stories
GEPA
gepa-ai
Optimize Anything with LLMs
0 stories
Gepa-Viz
Modaic
Visualization software product
0 stories
GitHub Repo Signals
Gooseworks
Extract and score leads from GitHub repositories using stars, forks, issues, PRs, comments, and contributions.
0 stories
Google ADK
Google
Agent development kit for building AI agents.
0 stories
Google AI Edge Gallery
Google LLC
Explore and run AI models on-device.
0 stories
GuideLLM
Red Hat
SLO-aware Benchmarking and Evaluation Platform for Optimizing Real-World LLM Inference
0 stories
G
Gym-Anything
Carnegie Mellon University
Gym-style research software toolkit
0 stories
Hermes Agent Control Room
Hermes
Agent control room
0 stories
howtoeval
Unverified vendor
Unverified software product
0 stories
HydraDB
AGI Context, Inc.
The Graph AI Runs On.
0 stories
Interfere
Interfere, Inc.
Interfere
0 stories
KernelBench
Scaling Intelligence Lab at Stanford University
Can LLMs Write Efficient GPU Kernels?
0 stories
KiloBench
Kilo Code Inc.
AI Coding Model Benchmark Results
0 stories
Klaimee
Klaimee
Klaimee
0 stories
LangSmith
LangChain
Ship great agents faster with LangSmith
0 stories
LangSmith Engine
LangChain
LangSmith platform engine
0 stories
LangSmith Fleet
LangChain
Agents for the whole company
0 stories
LangSmith Sandboxes
LangChain
Isolated code execution for LangSmith
0 stories
Latitude
Latitude Data S.L.
Make your AI agents self-healing
0 stories
LLM Checker
Pavelevich
Intelligent Ollama Model Selector
0 stories
LLM Council
Evolo Pty Ltd
When the Stakes Are Too High for One AI Answer
0 stories
LongSeeker
PolarSeeker
Elastic Context Orchestration for Long-Horizon Search Agents
0 stories
Lucent
Lucent AI, Inc.
AI software product from Lucent AI, Inc.
0 stories
Lumetric
Lumetric
You ask. It works.
0 stories
Mastra
Kepler Software Inc.
Open-source AI agent framework for TypeScript
0 stories
MCPMark Verified
EVAL SYS
Stabilized, version-pinned subset of MCPMark's standard tasks for reliable and reproducible MCP evaluation.
0 stories
Medmarks
Medmarks
Medmarks
0 stories
Meta-Harness
Stanford IRIS Lab
Open-source research harness
0 stories
micro1
micro1
Data lab to train frontier models & evaluate agents
0 stories
Microsoft Agent 365
Microsoft
Control plane for AI agents
0 stories
minigepa
Paolo Anzani
Mini GEPA implementation
0 stories
MirrorCode
Epoch AI
What's the largest software project AI can complete on its own?
0 stories
Mistral Studio
Mistral AI
Your AI production platform.
0 stories
M
ModelClock
ModelClock
ModelClock
0 stories
Morpheus
Skyfall AI
A Big-World Benchmark for Continual Reinforcement Learning
0 stories
M
MPMWorlds
Cornell University
Material-Point-Method Simulations for Inferring and Extrapolating Physical Dynamics
0 stories
Neo
Neo Research Inc.
Autonomous AI Agent to build AI Agents, evaluate AI Agents, Optimize LLM Prompts, Train AI Models and Finetune LLMs
0 stories
Observability
TrueFoundry
Observability for TrueFoundry applications and infrastructure.
0 stories
OctoTools
Stanford University
An Agentic Framework with Extensible Tools for Complex Reasoning
0 stories
Open Instruct
Ai2
AllenAI's post-training codebase
0 stories
OpenAI Agents SDK
OpenAI
Build sandbox, text, and voice agents with a small set of primitives.
0 stories
OpenInspect
Cole Murray
An open framework for background coding agents.
0 stories
Opik
Comet
LLM observability and evaluation platform
0 stories
Overmind
Overmind Technology Inc.
Software platform
0 stories
Paper Assistant Tool
Google
Gemini-backed automated feedback for scientific manuscripts
0 stories
Parallel Monitor API
Parallel Web Systems
Web monitoring API
0 stories
Pareto
Pareto, Inc.
Expert data for AGI
0 stories
Phoenix
Arize AI
Open-source AI observability and evaluation.
0 stories
P
PhoenixScore
Unknown vendor
Unverified software product
0 stories
Plurai
Plurai
Plurai
0 stories
Promptfoo
Promptfoo
AI/LLM evaluation and red-teaming platform
0 stories
Raindrop Workshop
Raindrop
Debug your AI agent locally
0 stories
Ramp Sheets
Ramp
Spreadsheet-oriented finance workflows
0 stories
Recursive
Recursive
Recursive
0 stories
Rerun
Rerun
Visualize and debug multimodal data
0 stories
RoPoLL
Amazon Web Services
Robust Panel of LLM-as-Judge
0 stories
Sakana AI Recursive Self-Improvement (RSI) Lab
Sakana AI
A dedicated research group within Sakana AI tasked with redesigning the AI development process itself with AI.
0 stories
Seer
Seer
Stop guessing on retrieval quality
0 stories
Seer Agent
Sentry
Ask any question about your application and Seer Agent finds the right telemetry to answer it.
0 stories
sentrux
Sentrux
The sensor that helps AI agents close the feedback loop.
0 stories
Sentry MCP
Sentry
Sentry MCP plugs Sentry's API directly into your LLM.
0 stories
Silico
Goodfire
Build AI models the way you write software
0 stories
SkillOpt
Microsoft
Microsoft software product
0 stories
Smithery
Clavia, Inc.
Connect agents to services in minutes
0 stories
SophontAI
Sophont
Open Medical Superintelligence
0 stories
Stop Slop
Unverified
A skill for removing AI tells from prose.
0 stories
SWE-Bench Pro
Scale AI
Evaluating challenging long-horizon software engineering tasks in public open source repositories
0 stories
SWE-check
Cognition
10x Faster Bug Detection
0 stories
Tessl
Tessl AI Limited
The AI-native software development platform
0 stories
THUNDERDOME
Thunderdome
THUNDERDOME
0 stories
Tinker
Thinking Machines Lab
Tinker is a training API for researchers
0 stories
tokenjuice
Vincent Koc
Lean output compaction for terminal-heavy agent workflows.
0 stories
Tokenscope
HduSy
Claude Code token cost, in your menu bar
0 stories
TokenSpeed
LightSeek Foundation
Speed-of-light LLM inference
0 stories
Traversal
Traversal
Traversal
0 stories
Trifle
Trifle, Inc.
Time-Series Metrics Made Simple
0 stories
Vals AI
Vals AI
Benchmark Generative AI for Enterprise Applications.
0 stories
Vellum
Vocify, Inc.
Your own, personal intelligence
0 stories
Vertex AI
Google Cloud
Build, deploy, and scale AI applications with Vertex AI.
0 stories
vitest-evals
Sentry
Vitest-based evals tool
0 stories
W&B LEET
Weights & Biases
Explore and compare local W&B runs from the terminal with the LEET TUI.
0 stories
WANDR
Perplexity AI
A benchmark for high-volume, evidence-heavy knowledge work
0 stories
Watchmen
Watchmen B.V.
Zero Trust, run as a service.
0 stories
Weights & Biases
Weights and Biases, LLC
The AI developer platform
0 stories
WorldModelGym
Reka AI
A decision-based fidelity benchmark for world models
0 stories
ZenMux
AI Force Singapore Pte. Ltd.
Unified API for all models, intelligent routing, and AI Model Insurance.
0 stories