Learning Video LLM with Streaming Speech Transcription at Scale (CVPR 2025)
Joya Chen
chenjoya
AI & ML interests
Video LLM
Recent Activity
authored a paper about 22 hours ago
Rethinking Expressivity and Efficiency in Test-Time Training updated a dataset 23 days ago
chenjoya/Live-WhisperX-526K upvoted a paper about 2 months ago
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning