📊
Evals & Observability
159 tools
Evaluation harnesses, LLM tracing, monitoring, and prompt/agent observability. Dashboards, judges, replay tools, and regression suites for LLM apps.
Arena
Arena Intelligence, Inc.
Introducing Agent Mode: Agentic AI is now measured in the Arena.
1 story
B
BenchLocal
stevibe
Test LLMs on real tasks. Compare models side-by-side.
1 story
Claude Console
Anthropic
Build on the Claude Platform
1 story
AA-Briefcase
Artificial Analysis
Agentic Knowledge Work Benchmark
0 stories
AEGIS
Applied Theory LLC
AI Agent Compliance, Governance & Control
0 stories
Agent Arena
Arena Intelligence, Inc.
Benchmark AI models on real agent work
0 stories
Agent Bricks
Databricks
AI agents for your enterprise
0 stories
Agent Development Kit (ADK)
Google
Build powerful AI agents with flexibility and control.
0 stories
Agent Sandbox
Kubernetes SIG Apps
A Kubernetes-native solution for running AI agents in isolated environments.
0 stories
Agent Sessions
jazzyalex
Live Per-Session Quota Burn for Codex and Claude on macOS
0 stories
A
Agent View
Agent View
Property viewings reimagined.
0 stories
agent-trace
Open Source
Open-source agent tracing
0 stories
Agentation
Dip
The missing feedback layer for AI coding agents.
0 stories
A
AgentRank
AgentRank
Every MCP server & agent tool, ranked
0 stories
AgentsView
Kenn Software LLC
A local-first desktop and web app for browsing, searching, and analyzing your past AI coding sessions.
0 stories
AGNTCY
Outshift by Cisco
The Internet of Agents
0 stories
AI Feedback
Amplitude
Listen to users at scale with AI
0 stories
AI Gateway
Vercel
One API for all AI models.
0 stories
AI21 Maestro
AI21 Labs
Optimization framework for real-world AI agents
0 stories
aiewf-eval
Daily
A long-context eval
0 stories
Amida Technology Platform
Amida Technology Solutions, Inc.
A Low-Code AI Test Automation Platform for Enterprise Applications
0 stories
Andon
Andon Labs
Platform for building, running, and evaluating any Safe Autonomous Organization.
0 stories
Antithesis
Antithesis Operations LLC
Autonomous Testing Platform
0 stories
ARC-AGI-3
ARC Prize Foundation
A benchmark for measuring intelligence in interactive environments
0 stories
Arize Phoenix
Arize AI
Open-source AI observability and evaluation.
0 stories
Artificial Analysis
Artificial Analysis, Inc.
Independent AI model and API provider benchmarking
0 stories
ASI-Evolve
Generative Artificial Intelligence Research Lab
Let AI Do the Research, You Keep the Insight
0 stories
ASSERT
Microsoft
Adaptive Spec-driven Scoring for Evaluation and Regression Testing
0 stories
Attention Head Visualiser
HeyNEO
Mapping What Each Head in GPT-2 Actually Does
0 stories
AutoHypothesis
Artem G
Open-source framework for agentic quantitative finance research.
0 stories
AutomationBench
Zapier
AI benchmark leaderboard
0 stories
AutomationBench-AA
Artificial Analysis
Independent leaderboard for Zapier's AutomationBench
0 stories
Baidu AI Cloud Qianfan
Baidu, Inc.
AI-native application development platform
0 stories
Benchmarks AI
Time Doctor
Compare User Productivity to Real Peers
0 stories
Better Agents
LangWatch
Build reliable, testable, production-grade AI agents with Better Agents CLI - the reliability layer for agent development
0 stories
Block-Sparse Featurizers
Goodfire
A family of methods for decomposing activations into multidimensional concept blocks.
0 stories
Braintrust
Braintrust Data, Inc.
The AI observability platform.
0 stories
BridgeBench V3
BridgeMind
The World's #1 Vibe Coding Benchmark
0 stories
BrowseComp-Plus
OpenAI
A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
0 stories
ccusage
ryoppippi
A CLI tool for analyzing Claude Code usage from local JSONL files.
0 stories
Chrome DevTools MCP
Google
Chrome DevTools for coding agents
0 stories
Claude Counter
Solo Software
A minimal browser extension that shows token count, cache timer, and usage bars on claude.ai.
0 stories
Clawdmeter
Hermann Haraldsson
ESP32 desk dashboard that shows Claude Code usage
0 stories
Code Arena
Arena Intelligence, Inc.
Build & Test with AI Coding Models
0 stories
Context Arena
Context Arena
Unverified software product
0 stories
Context.ai
Explore Interfaces Inc.
The analytics platform for LLM products.
0 stories
Coval
Coval
The simulation and evaluation platform for AI agents.
0 stories
Cua-Bench
Cua AI, Inc.
A benchmark for computer-use agents on professional software
0 stories
Cuey
Cuey
De-risk AI answers without leaving ChatGPT, Claude, Gemini, and other AI tools.
0 stories
CyberBench
Vals AI
Benchmarking AI models on real-world cybersecurity vulnerabilities.
0 stories
Daybreak
OpenAI
Frontier AI for cyber defenders.
0 stories
DeepAgents Deploy
LangChain
An open alternative to Claude Managed Agents
0 stories
DFlash
Z Lab
Diffusion-based speculative decoding for faster LLM inference
0 stories
Dify
LangGenius, Inc.
The platform for agentic workflow development.
0 stories
Dogfood
Dogfood
Dogfood your product, the efficient way
0 stories
DSPy
Stanford NLP Group
Programming—not prompting—language models.
0 stories
EdgeBench
ByteDance Seed
Measuring real-world environment learning and discovering a new scaling law
0 stories
Enterprise Worlds
Vibrant Labs
Executable environments for training and evaluating AI agents on realistic enterprise workflows
0 stories
EnterpriseRAG-Bench
Onyx
A RAG benchmark for the real world.
0 stories
Entire
Entire Inc.
Developer platform for humans and agents
0 stories
eot-bench
LiveKit
A benchmark for end-of-turn voice detection
0 stories
FrontierMath: Open Problems
Epoch AI
A collection of unsolved mathematics problems that have resisted serious attempts by professional mathematicians.
0 stories
Fullstack Code Arena
LMArena
Build, Deploy, and Evaluate with Fullstack Code Arena
0 stories
Future AGI
Future AGI
AI Agents hallucinate, fix it faster.
0 stories
Gemini Enterprise
Google Cloud
The AI platform for work.
0 stories
GeneBench-Pro
OpenAI
A research-level benchmark for a harder kind of AI progress.
0 stories
GEPA
Sky Computing Lab
Reflective prompt evolution for AI systems.
0 stories
GEPA-Viz
Modaic
Make prompt optimization visible and cheaper to inspect.
0 stories
GitHub Repository Signals
Gooseworks
Extract and score leads from GitHub repositories.
0 stories
Google AI Edge Gallery
Google LLC
Run AI models locally on your Android device.
0 stories
GuideLLM
Red Hat
Benchmark LLM inference performance
0 stories
G
Gym-Anything
Carnegie Mellon University
Verified desktop-app agent environments
0 stories
Hermes Agent Control Room
Hermes
Control Room-first template for managing Hermes agents from one VPS agent to specialist teams and orchestrated workflows
0 stories
How to Eval AI Agents
Raindrop AI
The 2026 Guide
0 stories
HydraDB
AGI Context, Inc.
The Graph AI Runs On.
0 stories
Inspect Petri
Meridian Labs
Automated multi-turn agent for alignment auditing of language models.
0 stories
Interfere
Interfere, Inc.
Ship software that never breaks
0 stories
ITSMBench
Vibrant Labs
Enterprise-agent evaluation benchmark for IT service management
0 stories
KernelBench
Scaling Intelligence Lab
Can LLMs Write Efficient GPU Kernels?
0 stories
KiloBench
Kilo Code Inc.
Because your benchmark score doesn't pay the API bills.
0 stories
Klaimee
Klaimee
Insure your agent. Get covered, fast.
0 stories
LangSmith
LangChain
The platform for agent engineering.
0 stories
LangSmith Engine
LangChain
Turns agent traces into issues, evals, fixes, and memory updates.
0 stories
LangSmith Fleet
LangChain
Collaborative agent workspace for sharing and running capable agents where teams work.
0 stories
LangSmith Sandboxes
LangChain
Give agents a secure computer
0 stories
Latitude
Latitude Data S.L.
Open-source agent monitoring.
0 stories
LLM Checker
Pavelevich
Intelligent Ollama Model Selector
0 stories
LLM Council
Evolo Pty Ltd
Parallel models, rankings, and synthesis for transparent deep research.
0 stories
LongSeeker
PolarSeeker
Elastic Context Orchestration for Long-Horizon Search Agents
0 stories
Lucent
Lucent AI, Inc.
Turn audience intelligence into landing pages that convert.
0 stories
Lumetric
Lumetric
You ask. It works.
0 stories
Mastra
Kepler Software Inc.
The TypeScript AI agent framework.
0 stories
MCPMark Verified
EVAL SYS
MCPMark: Stress-Testing Comprehensive MCP Benchmark
0 stories
Medmarks
Sophont AI
Open medical LLM evaluation suite
0 stories
Meta-Harness
Stanford IRIS Lab
Searches for better scaffolding around a fixed model
0 stories
micro1 Intelligence Platform
micro1
Human intelligence platform for AI training, evaluation, and data workflows at scale.
0 stories
Microsoft Agent 365
Microsoft
The control plane for AI agents
0 stories
Mini GEPA
Paolo Anzani
Mini GEPA implementation
0 stories
MirrorCode
Epoch AI
A long-horizon SWE benchmark for autonomous coding runs.
0 stories
Mistral AI Studio
Mistral AI
Build your frontier with Studio.
0 stories
ModelClock
stevibe
Probe LLM knowledge cutoffs with dated software release versions.
0 stories
Monitor API
Parallel Web Systems, Inc.
Automated web monitoring for news, products, research, filings, events, corporate development
0 stories
Morpheus
Skyfall AI
A Big-World Benchmark for Continual Reinforcement Learning
0 stories
M
MPMWorlds
Cornell University
Material-Point-Method Simulations for Inferring and Extrapolating Physical Dynamics
0 stories
Neo
Neo Research Inc.
Fully autonomous AI engineering agent
0 stories
Numbat
Perplexity AI
Visibility into AI agent activity on endpoints, with on-device detection, optional pre-action blocking, and forensic reconstruction.
0 stories
OctoTools
Stanford University
An Agentic Framework with Extensible Tools for Complex Reasoning
0 stories
Open Instruct
Allen Institute for AI
A fully open-source framework for instruction tuning and preference tuning of LMs.
0 stories
OpenAI Agents SDK
OpenAI
A lightweight, powerful framework for multi-agent workflows.
0 stories
OpenAI Evals
OpenAI
Evals is a framework for evaluating OpenAI models and an open-source registry of benchmarks.
0 stories
OpenInspect
CM Engineering
Open-source background agent system
0 stories
Opik
Comet
Open-source LLM evaluation and observability platform
0 stories
Overmind
Overmind Technology Inc.
Understand your infrastructure.
0 stories
Paper Assistant
Google
AI-assisted scientific paper review
0 stories
Pareto AI
Pareto, Inc.
AI-powered operations for your business
0 stories
PhoenixScore
Hyperbrowser
Score your tweet with the open-source X algorithm.
0 stories
Plurai
Plurai Inc.
Vibe-training for real-time, tailored agent evals and guardrails.
0 stories
Prime Agent
Prime Intellect
Autonomous AI research agent
0 stories
Promptfoo
Promptfoo, Inc.
AI security testing and evaluation
0 stories
Qwen-Scope
Alibaba Cloud
Decoding Intelligence, Unleashing Potential
0 stories
Raindrop Workshop
Raindrop
Local trace inspection for AI-agent runs
0 stories
Ramp Sheets
Ramp
The #1 AI spreadsheet editor, built for financial modeling and analysis
0 stories
Recovery Lab
Orreco
An interactive visualization of Cloudflare Agents' chat recovery system.
0 stories
Recursive
Recursive Superintelligence, Inc.
Recursive self-improving superintelligence to automate knowledge discovery.
0 stories
Rerun
Rerun Technologies AB
The multimodal data stack for physical AI.
0 stories
RoPoLL
Amazon Web Services
Robust Panel of LLM Judges
0 stories
Sakana AI Recursive Self-Improvement (RSI) Lab
Sakana AI
A dedicated research group within Sakana AI, tasked with redesigning the AI development process itself with AI.
0 stories
Seer
Seer
AI Reliability Platform
0 stories
Seer
Sentry
AI debugging agent
0 stories
sentrux
Sentrux
The sensor that helps AI agents close the feedback loop.
0 stories
Sentry MCP Server
Sentry
Give your AI assistant access to Sentry
0 stories
Silico
Goodfire
Automated interpretability and RL experimentation platform
0 stories
SkillOpt
Microsoft
Optimize agent skills against measured task batches
0 stories
Smithery
Clavia, Inc.
The MCP Registry
0 stories
SophontAI
Sophont
Open Medical Superintelligence
0 stories
Stop Slop
Unverified
A skill for removing AI tells from prose.
0 stories
SWE-bench Pro
Scale AI
A more challenging benchmark for coding agents
0 stories
SWE-check
Cognition
10x Faster Bug Detection
0 stories
Tessl
Tessl AI Limited
AI-native development platform
0 stories
Thunderdome
Thunderdome
Open Source Agile Planning Poker app
0 stories
Tinker
Thinking Machines Lab
A research API for fine-tuning language models
0 stories
Token counting
Anthropic
Count tokens before sending a request
0 stories
TokenJuice
Vincent Koc
Lean output compaction for terminal-heavy agent workflows.
0 stories
TokenScope
HduSy
Claude Code token cost, in your menu bar
0 stories
TokenSpeed
LightSeek Foundation
A speed-of-light LLM inference engine.
0 stories
Traversal
Traversal, Inc.
The AI SRE for complex systems
0 stories
Trifle
Trifle, Inc.
Track what matters. Skip the infrastructure.
0 stories
Usage.ai
Usage AI
Insurance for your cloud commitments.
0 stories
Vals AI
Vals AI
Benchmark Generative AI for Enterprise Applications.
0 stories
Vellum
Vocify, Inc.
The AI product development platform.
0 stories
Vercel Observability
Vercel
Monitor, debug, and understand your applications on Vercel.
0 stories
VisionAgent
LandingAI
An open-source Large Multimodal Model (LMM) powered agent for computer vision.
0 stories
Vitest Evals
Sentry
Evaluation tooling for Vitest
0 stories
W&B LEET
Weights & Biases
Explore and compare local W&B runs from the terminal with the LEET (Lightweight Experiment Exploration Tool) TUI.
0 stories
WANDR
Perplexity AI
Benchmark for deep and wide research
0 stories
Watchmen
Watchmen B.V.
Zero Trust, run as a service.
0 stories
Weights & Biases
Weights & Biases
The AI developer platform.
0 stories
WorldModelGym
Reka AI
a decision-based fidelity benchmark for world models
0 stories
ZenMux
AI Force Singapore Pte. Ltd.
One API for all AI models.
0 stories