MOSS-VL-Instruct-0708

An 11B multimodal vision-language model for image and video understanding. Attach an image or video (or none) and ask anything, turn by turn.

Radio
MultimodalTextbox
Examples