One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
-
Updated
Sep 19, 2026 - Python
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
Versatile Evaluation of Speech and Audio
A lightweight, self-hostable tool for conducting subjective listening tests in a browser.
An objective reconstruction evaluation toolkit for audio tokenizers, neural codecs, audio VAEs, and vocoders
Benten is an audio evaluation and analytics platform built specifically for Voice AI agents. It connects directly to voice AI platforms to discover your agents, process call audio, and measure conversation quality, response latency, and speech dynamics.
Extract grounded evidence from video files to enable automated review and visual understanding for coding agents.
Transparent, reproducible, modular evaluation for AI-generated music
Supplementary materials (pipeline, data, R analysis) for a study of latency and anchor-choice confounds in objective binaural quality metrics (BAM-Q, BINAQUAL), evaluated across 11 immersive audio format variants and 3 professionally produced recordings.
Professional portfolio for AI Evaluation, Audio Analysis, and UX Research. Specialized in Human-in-the-Loop (HITL) data integrity and design-focused research.
A reproducible benchmark for evaluating general-purpose agents on DJ mixing and real Mixxx software operation.
A standalone tool for evaluating Automatic Speech Recognition (ASR) models, particularly optimized for medical/clinical speech recognition, using Word Error Rate (WER) metric
To associate your repository with the audio-evaluation topic, visit your repo's landing page and select "manage topics."