Multimodal AI Evaluation Framework
Engineered a capability-aware benchmark dashboard unifying text, vision, audio, and agent tasks through a FastAPI backend. Optimized lmms-eval, faster-whisper, and inspect-ai execution using subprocesses, implemented real-time SSE streaming for evaluation logs, and normalized diverse engine outputs into a unified JSON schema.