text-image-2-text

MiniMax: MiniMax-01

MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, with 45.9 billion parameters activated per inference, and can ha ...

Rifx.Online 976.75K context $0.2/M input tokens $1.1/M output tokens

OpenAI: o1

Text image 2 text

The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason using ...

OpenAI 195.31K context $15/M input tokens $60/M output tokens $0.022/M image tokens

FREE

Google: Gemini 2.0 Flash Thinking Experimental (free)

Text image 2 text

# Free

Gemini 2.0 Flash Thinking Mode is an experimental model that's trained to generate the "thinking process" the model goes through as part of its response. As a result, Thinking Mode is capable of stro ...

Google 39.06K context $0 input tokens $0 output tokens

xAI: Grok 2 Vision 1212

Text image 2 text

Grok 2 Vision 1212 advances image-based AI with stronger visual comprehension, refined instruction-following, and multilingual support. From object recognition to style analysis, it empowers develope ...

X AI 32K context $2/M input tokens $10/M output tokens $0.004/M image tokens

70% OFF

nova-lite

Text image 2 text

# Discount

Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output. Amazon Nova Lite can handle real-time c ...

Amazon 292.97K context $0.06/M input tokens $0.24/M output tokens

70% OFF

nova-pro

Text image 2 text

# Discount

Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of tasks. As of December 2024, it achieves state-of-the ...

Amazon 292.97K context $0.8/M input tokens $3.2/M output tokens $0.001/M image tokens

gemini-exp-1206

Text image 2 text

Experimental release (December 6, 2024) of Gemini. ...

Google 8K context $4/M input tokens $16/M output tokens

40% OFF

Gemini 1.5 Pro

Text image 2 text

# Discount

Google's latest multimodal model, supporting image and video in text or chat prompts. Optimized for language tasks including:Code generation Text generation Text editing Problem solving...

Google 1.91M context $2.5/M input tokens $10/M output tokens $0.003/M image tokens

Amazon: Nova Pro 1.0

Text image 2 text

# New

Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of tasks. As of December 2024, it achieves state-of-the- ...

Amazon 292.97K context $0.8/M input tokens $3.2/M output tokens $0.001/M image tokens

Amazon: Nova Lite 1.0

Text image 2 text

# New

Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output. Amazon Nova Lite can handle real-time cu ...

Amazon 292.97K context $0.06/M input tokens $0.24/M output tokens

40% OFF

Claude-3-Haiku-20240307

Text image 2 text

# Discount

Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See the launch announcement and benchmark results [here](https ...

Anthropic 195.31K context $0.5/M input tokens $2.5/M output tokens $0.4/K image tokens

40% OFF

Gemini Flash 1.5

Text image 2 text

# Discount

Gemini 1.5 Flash is a foundation model that performs well at a variety of multimodal tasks such as visual understanding, classification, summarization, and creating content from image, audio and vid ...

Google 976.56K context $0.15/M input tokens $0.6/M output tokens $0.04/K image tokens

GPT-4o mini

Text image 2 text

GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more afford ...

OpenAI 125K context $0.15/M input tokens $0.6/M output tokens $0.007/M image tokens

40% OFF

gpt-4o

Text image 2 text

# Discount

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of GPT-4 Turbo while being twi ...

OpenAI 125K context $2.5/M input tokens $10/M output tokens $0.004/M image tokens

40% OFF

GPT-4o mini

Text image 2 text

# Discount # 40%Off # Discount

GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more afford ...

OpenAI 125K context $0.15/M input tokens $0.6/M output tokens $0.007/M image tokens

40% OFF

Claude 3.5 Sonnet-20240620

Text image 2 text

# Discount # 40%Off # Discount

Claude 3.5 Sonnet delivers better-than-Opus capabilities, faster-than-Sonnet speeds, at the same Sonnet prices. Sonnet is particularly good at:Coding: Autonomously writes, edits, and runs code w...

Anthropic 195.31K context $3/M input tokens $15/M output tokens $0.005/M image tokens

Mistral: Pixtral Large 2411

Text image 2 text

Pixtral Large is a 124B open-weights multimodal model built on top of Mistral Large 2. The model is able to understand documents, charts and natural images. The mode ...

MistralAI 125K context $2/M input tokens $6/M output tokens $0.003/M image tokens

Google: Gemini Pro Vision 1.0

Text image 2 text

Google's flagship multimodal model, supporting image and video in text or chat prompts for a text or code response. See the benchmarks and prompting guidelines from [Deepmind](https://deepmind.googl ...

Google 16K context $0.5/M input tokens $1.5/M output tokens $0.003/M image tokens

Google: Gemini Pro 1.5

Text image 2 text

Google's latest multimodal model, supporting image and video in text or chat prompts. Optimized for language tasks including:Code generation Text generation Text editing Problem solving...

Google 1.91M context $1.25/M input tokens $5/M output tokens $0.003/M image tokens

Google: Gemini Flash 1.5

Text image 2 text

Gemini 1.5 Flash is a foundation model that performs well at a variety of multimodal tasks such as visual understanding, classification, summarization, and creating content from image, audio and vide ...

Google 976.56K context $0.075/M input tokens $0.3/M output tokens $0.04/K image tokens

Mistral: Pixtral 12B

Text image 2 text

The first image to text model from Mistral AI. Its weight was launched via torrent per their tradition: https://x.com/mistralai/status/1833758285167722836 ...

MistralAI 4K context $0.1/M input tokens $0.1/M output tokens $0.144/K image tokens

OpenAI: ChatGPT-4o

Text image 2 text

Dynamic model continuously updated to the current version of GPT-4o in ChatGPT. Intended for research and evaluation. Note: This model is currently experimental and not suitable fo ...

OpenAI 125K context $5/M input tokens $15/M output tokens $0.007/M image tokens

Anthropic: Claude 3.5 Sonnet (2024-06-20)

Text image 2 text

Claude 3.5 Sonnet delivers better-than-Opus capabilities, faster-than-Sonnet speeds, at the same Sonnet prices. Sonnet is particularly good at:Coding: Autonomously writes, edits, and runs code wi...

Anthropic 195.31K context $3/M input tokens $15/M output tokens $0.005/M image tokens

FREE

Google: Gemini Pro 1.5 Experimental

Text image 2 text

# Free

Google's latest multimodal model, supporting image and video in text or chat prompts. Optimized for language tasks including:Code generation Text generation Text editing Problem solving...

Google 1.91M context $0 input tokens $0 output tokens $0.003/M image tokens

Anthropic: Claude 3 Opus

Text image 2 text

Claude 3 Opus is Anthropic's most powerful model for highly complex tasks. It boasts top-level performance, intelligence, fluency, and understanding. See the launch announcement and benchmark result ...

Anthropic 195.31K context $15/M input tokens $75/M output tokens $0.024/M image tokens

Anthropic: Claude 3 Sonnet

Text image 2 text

Claude 3 Sonnet is an ideal balance of intelligence and speed for enterprise workloads. Maximum utility at a lower price, dependable, balanced for scaled deployments. See the launch announcement and ...

Anthropic 195.31K context $3/M input tokens $15/M output tokens $0.005/M image tokens

Anthropic: Claude 3 Haiku

Text image 2 text

Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See the launch announcement and benchmark results [here](https: ...

Anthropic 195.31K context $0.25/M input tokens $1.25/M output tokens $0.4/K image tokens

Anthropic: Claude 3.5 Sonnet

Text image 2 text

Claude 3.5 Sonnet delivers better-than-Opus capabilities, faster-than-Sonnet speeds, at the same Sonnet prices. Sonnet is particularly good at:Coding: Autonomously writes, edits, and runs code wi...

Anthropic 195.31K context $3/M input tokens $15/M output tokens $0.005/M image tokens

Qwen2-VL 7B Instruct

Text image 2 text

Qwen2 VL 7B is a multimodal LLM from the Qwen Team with the following key enhancements:SoTA understanding of images of various resolution & ratio: Qwen2-VL achieves state-of-the-art performance o...

Qwen 32K context $0.1/M input tokens $0.1/M output tokens $0.144/K image tokens

FREE

Meta: Llama 3.2 11B Vision Instruct (free)

Text image 2 text

# Free

Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It excels in tasks such as image captioning and visual question answ ...

Meta Llama 128K context $0 input tokens $0 output tokens $0.079/K image tokens

Meta: Llama 3.2 11B Vision Instruct

Text image 2 text

Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It excels in tasks such as image captioning and visual question answ ...

Meta Llama 128K context $0.055/M input tokens $0.055/M output tokens $0.079/K image tokens

GPT-4o

Text image 2 text

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of GPT-4 Turbo while being twi ...

OpenAI 125K context $2.5/M input tokens $10/M output tokens $0.004/M image tokens

OpenAI: GPT-4o

Text image 2 text

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of GPT-4 Turbo while being twi ...

OpenAI 125K context $2.5/M input tokens $10/M output tokens $0.004/M image tokens

OpenAI: GPT-4o-mini

Text image 2 text

GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more afford ...

OpenAI 125K context $0.15/M input tokens $0.6/M output tokens $0.007/M image tokens

Google: Gemini 1.5 Flash-8B

Text image 2 text

Gemini 1.5 Flash-8B is optimized for speed and efficiency, offering enhanced performance in small prompt tasks like chat, transcription, and translation. With reduced latency, it is highly effective ...

Google 976.56K context $0.037/M input tokens $0.15/M output tokens

Qwen2-VL 72B Instruct

Text image 2 text

Qwen2 VL 72B is a multimodal LLM from the Qwen Team with the following key enhancements:SoTA understanding of images of various resolution & ratio: Qwen2-VL achieves state-of-the-art performance...

Qwen 32K context $0.4/M input tokens $0.4/M output tokens $0.578/K image tokens

Meta: Llama 3.2 90B Vision Instruct

Text image 2 text

The Llama 90B Vision model is a top-tier, 90-billion-parameter multimodal model designed for the most challenging visual reasoning and language tasks. It offers unparalleled accuracy in image caption ...

Meta Llama 128K context $0.35/M input tokens $0.4/M output tokens $0.506/K image tokens

Google: Gemini Experimental 1121 (free)

Text image 2 text

Experimental release (November 21st, 2024) of Gemini. ...

Rifx.Online 8K context $0 input tokens $0 output tokens

Google: LearnLM 1.5 Pro Experimental (free)

Text image 2 text

An experimental version of Gemini 1.5 Pro from Google. ...

Rifx.Online 8K context $0 input tokens $0 output tokens

Google: Gemini 1.5 Flash-8B

Text image 2 text

Gemini 1.5 Flash-8B is optimized for speed and efficiency, offering enhanced performance in small prompt tasks like chat, transcription, and translation. With reduced latency, it is hig ...

Google 976.56K context $0.037/M input tokens $0.15/M output tokens

Meta: Llama 3.2 11B Vision Instruct

Text image 2 text

Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It excels in tasks such as image captioning and visual ...

Meta llama 128K context $0.055/M input tokens $0.055/M output tokens $0.079/K image tokens

Meta: Llama 3.2 90B Vision Instruct

Text image 2 text

The Llama 90B Vision model is a top-tier, 90-billion-parameter multimodal model designed for the most challenging visual reasoning and language tasks. It offers unparalleled accuracy in ...

Meta llama 128K context $0.35/M input tokens $0.4/M output tokens $0.506/K image tokens

Meta: Llama 3.2 90B Vision Instruct (free)

Text image 2 text

The Llama 90B Vision model is a top-tier, 90-billion-parameter multimodal model designed for the most challenging visual reasoning and language tasks. It offers unparalleled accuracy in ...

Rifx.Online 4K context $0 input tokens $0 output tokens

Qwen2-VL 72B Instruct

Text image 2 text

Qwen2 VL 72B is a multimodal LLM from the Qwen Team with the following key enhancements:SoTA understanding of images of various resolution & ratio: Qwen2-VL achieves state-of-the-ar...

Qwen 32K context $0.4/M input tokens $0.4/M output tokens $0.578/K image tokens

Mistral: Pixtral 12B

Text image 2 text

The first image to text model from Mistral AI. Its weight was launched via torrent per their tradition: https://x.com/mistralai/status/1833758285167722836 ...

Mistralai 4K context $0.1/M input tokens $0.1/M output tokens $0.144/K image tokens

Google: Gemini Flash 8B 1.5 Experimental

Text image 2 text

Gemini 1.5 Flash 8B Experimental is an experimental, 8B parameter version of the Gemini 1.5 Flash model. Usage of Gemini is subject to Google's [Gemini Term ...

Google 976.56K context $0 input tokens $0 output tokens

Qwen2-VL 7B Instruct

Text image 2 text

Qwen2 VL 7B is a multimodal LLM from the Qwen Team with the following key enhancements:SoTA understanding of images of various resolution & ratio: Qwen2-VL achieves state-of-the-art...

Qwen 32K context $0.1/M input tokens $0.1/M output tokens $0.144/K image tokens

OpenAI: ChatGPT-4o

Text image 2 text

Dynamic model continuously updated to the current version of GPT-4o in ChatGPT. Intended for research and evaluation. Note: This model is currently experimental and n ...

Openai 125K context $5/M input tokens $15/M output tokens $0.007/M image tokens

Anthropic: Claude 3.5 Sonnet (2024-06-20)

Text image 2 text

Claude 3.5 Sonnet delivers better-than-Opus capabilities, faster-than-Sonnet speeds, at the same Sonnet prices. Sonnet is particularly good at:Coding: Autonomously writes, edits, an...

Anthropic 195.31K context $3/M input tokens $15/M output tokens $0.005/M image tokens

Google: Gemini Flash 1.5

Text image 2 text

Gemini 1.5 Flash is a foundation model that performs well at a variety of multimodal tasks such as visual understanding, classification, summarization, and creating content from image, ...

Google 976.56K context $0.075/M input tokens $0.3/M output tokens $0.04/K image tokens

Google: Gemini Pro 1.5

Text image 2 text

Google's latest multimodal model, supporting image and video in text or chat prompts. Optimized for language tasks including:Code generation Text generation Text editing Prob...

Google 1.91M context $1.25/M input tokens $5/M output tokens $0.003/M image tokens

Anthropic: Claude 3 Haiku

Text image 2 text

Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See the launch announcement and benchmark results ...

Anthropic 195.31K context $0.25/M input tokens $1.25/M output tokens $0.4/K image tokens

Anthropic: Claude 3 Opus

Text image 2 text

Claude 3 Opus is Anthropic's most powerful model for highly complex tasks. It boasts top-level performance, intelligence, fluency, and understanding. See the launch announcement and be ...

Anthropic 195.31K context $15/M input tokens $75/M output tokens $0.024/M image tokens

Anthropic: Claude 3 Sonnet

Text image 2 text

Claude 3 Sonnet is an ideal balance of intelligence and speed for enterprise workloads. Maximum utility at a lower price, dependable, balanced for scaled deployments. See the launch an ...

Anthropic 195.31K context $3/M input tokens $15/M output tokens $0.005/M image tokens

Google: Gemini Pro Vision 1.0

Text image 2 text

Google's flagship multimodal model, supporting image and video in text or chat prompts for a text or code response. See the benchmarks and prompting guidelines from [Deepmind](https:// ...

Google 16K context $0.5/M input tokens $1.5/M output tokens $0.003/M image tokens