aiewf-eval
A long-context eval
Open-source multi-turn benchmark framework for evaluating LLM conversations, measuring latency, tool use, instruction following, and knowledge grounding for text and speech-to-speech models used in voice-agent scenarios.

Recent stories
0 linked stories
No linked stories yet.