An 11B multimodal vision-language model for image and video understanding. Attach an image or video (or none) and ask anything, turn by turn.
Model Card | GitHub