WangJ. x 3

bookRéférences 3

Références · 2026-01-20

Vx2text: End-to-end learning of video-based text generation from multimodal inputs

We present Vx2Text, a framework for text generation from multimodal inputs consisting of video plus ...

WangJ.LinX.BertasiusG.ChangS.F.

Références · 2023-11-06

A cry for help: Early detection of brain injury in newborns

… base model (encoder) is a 76M-parameter convolutional neural network that has demonstra...

OnuC.C.LatremouilleS.GorinA.WangJ.

Références

Emu: Enhancing image generation models using photogenic needles in a haystack

Training text-to-image models with web scale image-text pairs enables the generation of a wide range...

WangJ.DaiX.HouJ.MaC.Y.TsaiS.WangR.

Mots-clés associés