Multimodal AI Evaluation Framework
Engineered a capability-aware benchmark dashboard unifying text, vision, audio, and agent tasks into a single FastAPI backend. Optimized engine execution using lmms-eval, faster-whisper, and inspect-ai subprocesses, implemented real-time SSE streaming, and normalized diverse engine outputs into a unified JSON schema.