Penguin-VL is a compact vision-language model family built to study how far multimodal efficiency can be pushed by redesigning the vision encoder, rather than only scaling data or model size.
Abstract: Video encoding draws high research interest, due to the enormous demand for video traffic and real-time encoding for transmission. In video encoding standards such as HEVC (High-Efficiency ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results