2609.02231

Total: 1

#1 PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment [PDF1] [Copy] [Kimi] [REL]

Authors: Fan Yuxuan, Huang Miaojun, Zhang Haimei, Wu Jingshen, Liu Hao

Interview assessment requires per-criterion judgments grounded in behavioral evidence, yet surging applicant volumes have made human-only evaluation costly and inconsistent, while existing AI approaches yield opaque scores without traceable rationale. We introduce PhoenixNest-Video, an evidence-grounded multimodal agent framework for automated video interview assessment. It builds a semantic video graph as structured working memory, performs rubric-conditioned retrieval with cross-modal verification across visual, audio, and textual streams, and produces per-criterion scores anchored to the candidate's materials. A Scorer trained via Rubrics-based Reinforcement Learning with dual rewards for rubric alignment and score-level differentiation internalizes the discriminative structure of multi-level rubrics. PhoenixNest-Video attains 91.50\% grade-level accuracy on VInterview-2025, outperforming substantially larger proprietary models. A compact, rubric-grounded agent therefore scores candidates in closer agreement with an expert panel than direct prompting of much larger models, and exposes the evidence behind each score for human review.

Subject: Artificial Intelligence

Publish: 2026-09-02 07:40:43 UTC