📊
Evals & Observability
149 tools
Evaluation harnesses, LLM tracing, monitoring, and prompt/agent observability. Dashboards, judges, replay tools, and regression suites for LLM apps.
Vals AI
Vals AI
Benchmark Generative AI for Enterprise Applications.
5 stories
AA-Briefcase
Artificial Analysis
Agentic Knowledge Work Benchmark
3 stories
Artificial Analysis
Artificial Analysis, Inc.
Independent analysis of AI
3 stories
DFlash
Z Lab
Block Diffusion for Flash Speculative Decoding
3 stories
LangSmith
LangChain
The Agent Engineering Platform
3 stories
Weights & Biases
Weights and Biases, LLC
The AI developer platform
3 stories
AI Gateway
Vercel
The AI Gateway for developers
2 stories
ITSMBench
Vibrant Labs
IT service management environment for testing whether agents can carry out policy-governed work inside a system of record.
2 stories
Mastra
Kepler Software Inc.
TypeScript AI Framework for Agents and Apps
2 stories
OpenAI Agents SDK
OpenAI
Build sandbox, text, and voice agents with a small set of primitives.
2 stories
Silico
Goodfire
Build AI models the way you write software
2 stories
SkillOpt
Microsoft
Text-space optimization for frozen agents
2 stories
Agent Arena
Arena Intelligence, Inc.
AI Agent Performance Leaderboard
1 story
Arena
Arena Intelligence, Inc.
Crowdsourced AI Model Evaluation Platform
1 story
AutomationBench-AA
Artificial Analysis
Agentic SaaS Workflow Benchmark
1 story
Baidu Qianfan
Baidu, Inc.
An enterprise-grade, all-in-one large model platform centered on Agent development
1 story
Block-Sparse Featurizers
Goodfire
Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds
1 story
Braintrust
Braintrust Data, Inc.
The active observability platform for agents
1 story
Chrome DevTools for agents
Google
Help your agent build, debug, and verify your code correctly.
1 story
Claude Console
Anthropic
Develop with Claude
1 story
Code Arena
Arena Intelligence, Inc.
Build & Test with AI Coding Models
1 story
Context.ai
Explore Interfaces Inc.
Build, deploy, and improve AI agents.
1 story
Cua-Bench
Cua AI, Inc.
Make your agents better at computers
1 story
Daybreak
OpenAI
Frontier AI for defenders.
1 story
Enterprise Worlds
Vibrant Labs
Executable enterprise environments for measuring operational agents
1 story
G
GeneBench-Pro
OpenAI
A research-level benchmark measuring how AI agents navigate ambiguity and make consequential judgments in computational biology.
1 story
G
Gym-Anything
Carnegie Mellon University
Turn Any Software into an Agent Environment
1 story
LangSmith Engine
LangChain
Your proactive agent engineer
1 story
LangSmith Sandboxes
LangChain
Give agents a secure computer
1 story
Latitude
Latitude Data S.L.
Make your AI agents self-healing
1 story
Medmarks
Sophont
Open medical AI benchmarks
1 story
MirrorCode
Epoch AI
What's the largest software project AI can complete on its own?
1 story
Morpheus
Skyfall AI
A Big-World Benchmark for Continual Reinforcement Learning
1 story
Numbat
Perplexity AI
Visibility into AI agent activity on endpoints, with on-device detection, optional pre-action blocking, and forensic reconstruction.
1 story
OctoTools
Stanford University
An Agentic Framework with Extensible Tools for Complex Reasoning
1 story
Plurai
Plurai Inc.
The real world trust platform for AI agents
1 story
Raindrop Workshop
Raindrop
The local debugger your agent is missing.
1 story
Ramp Sheets
Ramp
The #1 AI spreadsheet editor, built for financial modeling and analysis
1 story
Sentry MCP
Sentry
Sentry MCP plugs Sentry's API directly into your LLM.
1 story
SWE-Bench Pro
Scale AI
Evaluating challenging long-horizon software engineering tasks in public open source repositories
1 story
Tinker
Thinking Machines Lab
Tinker is a training API for researchers
1 story
tokenjuice
Vincent Koc
Lean output compaction for terminal-heavy agent workflows.
1 story
TokenSpeed
LightSeek Foundation
Speed-of-light LLM inference
1 story
WANDR
Perplexity AI
A benchmark for high-volume, evidence-heavy knowledge work
1 story
AEGIS
Applied Theory LLC
AI Agent Compliance, Governance & Control
0 stories
Agent Bricks
Databricks
Production AI agents
0 stories
Agent Development Kit (ADK)
Google
Build production agents, not prototypes.
0 stories
Agent Sandbox
Kubernetes SIG Apps
Enables easy management of isolated, stateful, singleton workloads, ideal for use cases like AI agent runtimes.
0 stories
Agent Sessions
jazzyalex
Live Per-Session Quota Burn for Codex and Claude on macOS
0 stories
A
Agent View
Agent View
Property viewings reimagined.
0 stories
agent-trace
Open Source
Open-source agent tracing
0 stories
Agentation
Dip
Visual feedback for AI coding agents
0 stories
A
AgentRank
AgentRank
Google PageRank for AI agents.
0 stories
AgentsView
Kenn Software LLC
A local-first desktop and web app for browsing, searching, and analyzing your past AI coding sessions.
0 stories
AGNTCY
Outshift by Cisco
Building the Internet of Agents (IoA)
0 stories
A
AI Usage
AI Usage
AI Usage
0 stories
AI21 Maestro
AI21 Labs
Optimization framework for real-world AI agents
0 stories
aiewf-eval
Daily
A framework for evaluating multi-turn LLM conversations with support for text, realtime audio, and speech-to-speech models.
0 stories
Andon
Andon Labs
Platform for building, running, and evaluating Safe Autonomous Organizations
0 stories
Antithesis
Antithesis
Bug-free systems, unlimited velocity
0 stories
ARC-AGI-3
ARC Prize Foundation
The first interactive reasoning benchmark designed to measure human-like intelligence in AI agents.
0 stories
ASI-Evolve
Generative Artificial Intelligence Research Lab (GAIR)
Let AI Do the Research, You Keep the Insight
0 stories
ASSERT
Microsoft
Adaptive Spec-driven Scoring for Evaluation and Regression Testing
0 stories
Attention Head Visualiser
HeyNEO
Mapping what each head in GPT-2 actually does
0 stories
AutoHypothesis
arteemg
an open-source agentic framework for quantitative finance research
0 stories
B
BenchLocal
BenchLocal
Test LLMs on real tasks. Compare models side-by-side.
0 stories
Benchmarks.AI
Time Doctor
Directory of AI Benchmarks
0 stories
Better Agents
LangWatch
Build reliable, testable, production-grade AI agents with Better Agents CLI - the reliability layer for agent development
0 stories
BridgeBench V3
BridgeMind
The World's #1 Vibe Coding Benchmark
0 stories
BrowseComp-Plus
OpenAI
A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
0 stories
ccusage
Independent
Coding (Agent) CLI Usage Analysis
0 stories
Claude Counter
Solo Software
Shows ~token count, cache timer, and native session/weekly usage bars on claude.ai.
0 stories
Claude Token Count API
Anthropic
Count the number of tokens in a Message.
0 stories
Clawdmeter
Independent
ESP32 desk dashboard that shows Claude Code usage
0 stories
Context Arena
Context Arena
Unverified software product
0 stories
Coval
Coval
Voice AI Testing & Evaluation Platform
0 stories
CyberBench
Vals AI
Can AI agents find security bugs and patch them?
0 stories
Deep Agents Deploy
LangChain
Deploy a model-agnostic, open source agent to production with a single command.
0 stories
Dify
LangGenius, Inc.
The Platform for Production-Ready Agentic Workflows.
0 stories
Dogfood
Dogfood
Dogfood software product
0 stories
DSPy
Stanford NLP Group
Program, don’t prompt, your LLMs.
0 stories
EdgeBench
ByteDance Seed
An ultra-long-horizon benchmark built to measure learning from environments
0 stories
EnterpriseRAG-Bench
Onyx
A RAG benchmark for the real world.
0 stories
Entire
Entire Inc.
Developer platform for humans and agents
0 stories
eot-bench
LiveKit
The open benchmark for end-of-turn detection
0 stories
Fullstack Code Arena
Arena
Build, Deploy, and Evaluate with Fullstack Code Arena
0 stories
Future AGI
Future AGI
AI Agents hallucinate, fix it faster.
0 stories
Gemini Enterprise Agent Platform
Google Cloud
Innovate, build, and deploy enterprise ready agents
0 stories
GEPA
gepa-ai
Optimize prompts, code, and more with AI-powered Reflective Optimization
0 stories
gepa-viz
Modaic
Live visualization for GEPA prompt-optimization runs.
0 stories
GitHub Repo Signals
Gooseworks
Extract and score leads from GitHub repositories using stars, forks, issues, PRs, comments, and contributions.
0 stories
Google AI Edge Gallery
Google LLC
Explore, Experience, and Evaluate the Future of On-Device Generative AI with Google AI Edge.
0 stories
GuideLLM
Red Hat
SLO-aware Benchmarking and Evaluation Platform for Optimizing Real-World LLM Inference
0 stories
Hermes Agent Control Room
Hermes
Control Room-first template for managing Hermes agents from one VPS agent to specialist teams and orchestrated workflows
0 stories
howtoeval
Unverified vendor
Unverified software product
0 stories
HydraDB
AGI Context, Inc.
The Graph AI Runs On.
0 stories
Interfere
Interfere, Inc.
Production visibility for teams that ship fast
0 stories
KernelBench
Scaling Intelligence Lab at Stanford University
Can LLMs Write Efficient GPU Kernels?
0 stories
KiloBench
Kilo Code Inc.
AI Coding Model Benchmark Results
0 stories
Klaimee
Klaimee
AI Agent Insurance, Certification & Guarantee
0 stories
LangSmith Fleet
LangChain
Agents for the whole company
0 stories
LLM Checker
Pavelevich
Intelligent Ollama Model Selector
0 stories
LLM Council
Evolo Pty Ltd
When the Stakes Are Too High for One AI Answer
0 stories
LongSeeker
PolarSeeker
Elastic Context Orchestration for Long-Horizon Search Agents
0 stories
Lucent
Lucent AI, Inc.
See what's really happening in your product
0 stories
Lumetric
Lumetric
You ask. It works.
0 stories
MCPMark Verified
EVAL SYS
Stabilized, version-pinned subset of MCPMark's standard tasks for reliable and reproducible MCP evaluation.
0 stories
Meta-Harness
Stanford IRIS Lab
End-to-End Optimization of Model Harnesses
0 stories
micro1
micro1
Data lab to train frontier models & evaluate agents
0 stories
Microsoft Agent 365
Microsoft
The Control Plane for Agents
0 stories
minigepa
Paolo Anzani
Mini GEPA implementation
0 stories
Mistral Studio
Mistral AI
Your AI production platform.
0 stories
ModelClock
stevibe
Probe LLM knowledge cutoffs with dated software release versions.
0 stories
M
MPMWorlds
Cornell University
Material-Point-Method Simulations for Inferring and Extrapolating Physical Dynamics
0 stories
Neo
Neo Research Inc.
Autonomous AI Agent to build AI Agents, evaluate AI Agents, Optimize LLM Prompts, Train AI Models and Finetune LLMs
0 stories
Observability
TrueFoundry
Observability for TrueFoundry applications and infrastructure.
0 stories
Open Instruct
Ai2
AllenAI's post-training codebase
0 stories
OpenAI Evals
OpenAI
Evaluations test model outputs to ensure they meet style and content criteria that you specify.
0 stories
OpenInspect
Cole Murray
Self-hostable background coding agents for software teams
0 stories
Opik
Comet
AI Observability & Evals For the Agentic Era
0 stories
Overmind
Overmind Technology Inc.
Prevent your next outage
0 stories
Paper Assistant Tool
Google
Gemini-backed automated feedback for scientific manuscripts
0 stories
Parallel Monitor API
Parallel Web Systems Inc.
Automated web monitoring for news, products, research, filings, events, corporate development
0 stories
Pareto
Pareto, Inc.
Expert data for AGI
0 stories
Phoenix
Arize AI
The open-source platform for agent development and evaluation
0 stories
PhoenixScore
Hyperbrowser
Score your tweet with the open-source X algorithm.
0 stories
Promptfoo
Promptfoo
Build Secure AI Applications
0 stories
Recursive
Recursive
Recursive self-improving superintelligence to automate knowledge discovery.
0 stories
Rerun
Rerun
The Data Layer for Physical AI
0 stories
RoPoLL
Amazon Web Services
Robust Panel of LLM-as-Judge
0 stories
Sakana AI Recursive Self-Improvement (RSI) Lab
Sakana AI
A dedicated research group within Sakana AI tasked with redesigning the AI development process itself with AI.
0 stories
Seer
Seer
Stop guessing on retrieval quality
0 stories
Seer Agent
Sentry
Seer knows everything. Ask it anything.
0 stories
sentrux
Sentrux
The sensor that helps AI agents close the feedback loop.
0 stories
Smithery
Clavia, Inc.
Connect agents to services in minutes
0 stories
SophontAI
Sophont
Open Medical Superintelligence
0 stories
Stop Slop
Unverified
A skill for removing AI tells from prose.
0 stories
SWE-check
Cognition
10x Faster Bug Detection
0 stories
Tessl
Tessl AI Limited
Agent Enablement Platform
0 stories
Thunderdome
Thunderdome
Open Source Agile Planning Poker app
0 stories
Tokenscope
HduSy
Claude Code token cost, in your menu bar
0 stories
Traversal
Traversal, Inc.
The AI SRE for complex systems
0 stories
Trifle
Trifle, Inc.
Time-Series Metrics Made Simple
0 stories
Vellum
Vocify, Inc.
Your own, personal intelligence
0 stories
vitest-evals
Sentry
Harness-backed AI evaluation tests on top of Vitest.
0 stories
W&B LEET
Weights & Biases
Explore and compare local W&B runs from the terminal with the LEET TUI.
0 stories
Watchmen
Watchmen B.V.
Zero Trust, run as a service.
0 stories
WorldModelGym
Reka AI
A decision-based fidelity benchmark for world models
0 stories
ZenMux
AI Force Singapore Pte. Ltd.
Unified API for 100+ AI Models
0 stories