visual-question-answering

Amazon: Nova Pro 1.0

Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of tasks. As of December 2024, it achieves state-of-the- ...

Amazon 292.97K context $0.8/M input tokens $3.2/M output tokens $0.001/M image tokens

Amazon: Nova Lite 1.0

Text image 2 text

# New

Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output. Amazon Nova Lite can handle real-time cu ...

Amazon 292.97K context $0.06/M input tokens $0.24/M output tokens

FREE

Meta: Llama 3.2 11B Vision Instruct (free)

Text image 2 text

# Free

Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It excels in tasks such as image captioning and visual question answ ...

Meta Llama 128K context $0 input tokens $0 output tokens $0.079/K image tokens

Meta: Llama 3.2 11B Vision Instruct

Text image 2 text

Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It excels in tasks such as image captioning and visual question answ ...

Meta Llama 128K context $0.055/M input tokens $0.055/M output tokens $0.079/K image tokens

Meta: Llama 3.2 90B Vision Instruct

Text image 2 text

The Llama 90B Vision model is a top-tier, 90-billion-parameter multimodal model designed for the most challenging visual reasoning and language tasks. It offers unparalleled accuracy in image caption ...

Meta Llama 128K context $0.35/M input tokens $0.4/M output tokens $0.506/K image tokens

Visual question answering

Amazon: Nova Pro 1.0

Amazon: Nova Lite 1.0

Meta: Llama 3.2 11B Vision Instruct (free)

Meta: Llama 3.2 11B Vision Instruct

Meta: Llama 3.2 90B Vision Instruct