Multimodal inputs
Send text and image content in the same user message.
Multimodal input allows a supported model to process text, images, and other content together. This page uses a text-and-image chat request. It does not imply that every model supports images, audio, or video.
Choose a model
Before sending a request, confirm that:
- The model supports image input.
- The model supports the Chat Completions interface.
- Image format, count, size, and resolution meet its limits.
- The current API key can access the model.
Text and image example
{
"model": "YOUR_MULTIMODAL_MODEL_ID",
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "Describe the main content of this image."
},
{
"type": "image_url",
"image_url": {
"url": "https://example.com/image.jpg"
}
}
]
}
]
}This is an example compatible structure. Support for URLs, Base64 data, or other input methods depends on the selected model.
Prepare an image
| Check | Description |
|---|---|
| Accessibility | The service processing the request must be able to access the image URL |
| File format | Use a MIME type explicitly supported by the model |
| File size | Stay within per-image and total request limits |
| Image count | Stay within the number allowed in one request |
| Data permission | Make sure that you are allowed to submit and process the image |
A local file path or a webpage that requires a browser login usually cannot be used as a remote image input. Do not make a private file public only to call the API.
Control usage
Images may use a different metering method from ordinary text. Remove unnecessary images and repeated history, then review media usage and cost according to the model documentation.
Troubleshooting
- The image cannot be read: confirm that the URL returns an image and is reachable from the service.
- The format is unsupported: convert the file to a supported format instead of changing only its extension.
- The request is too large: compress, resize, or remove images.
- The response structure differs: follow the interface documentation for the selected model.