WangJ. x 3
bookRéférences 3
Références · 2026-01-20
Vx2text: End-to-end learning of video-based text generation from multimodal inputs
We present Vx2Text, a framework for text generation from multimodal inputs consisting of video plus ...
Références · 2023-11-06
A cry for help: Early detection of brain injury in newborns
… base model (encoder) is a 76M-parameter convolutional neural network that has demonstra...