Image Analysis with GPT-4.1-mini and Structured Outputs
In this tutorial, we will learn how to analyze images using a multimodal AI model and return the results directly as structured data. We will use GPT-4.1-mini with Azure OpenAI and Structured Outputs.
Instead of using a separate service for every task, we can send an image together with our instructions to the model and extract information such as a description, tags, detected objects and OCR text.
The important point is that the AI response should not simply be arbitrary text. Our application needs a defined data structure. For this purpose, we use JSON Schema with Pydantic.
Prerequisites
- Python must be installed on your computer.
- An Azure OpenAI deployment with GPT-4.1-mini must be available.
- Basic knowledge of Python and FastAPI.
- Basic knowledge of REST APIs and JSON.
- An Azure OpenAI client must be configured in the backend.
Step 1: Define the Pydantic Schema
First, we define the structure that our application expects from the AI using Pydantic:

This class defines both the contract for our application and the schema for the model's structured response. This means that we do not need to manually create the JSON Schema as a dictionary.
Step 2: Send the Image and Prompt to the Model
The image is read by the backend and then passed to the model as a Base64 Data URL.
In addition, we send a prompt describing which information should be extracted from the image. For example:

The user only sees the upload functionality. The actual instructions used for the image analysis remain in the backend.
Step 3: Use Structured Outputs
Now comes the important part. Instead of simply asking the model to return JSON, we use Structured Outputs:

With text_format=AnalysisResult, we use our Pydantic class as the schema for the response.
The model therefore receives a defined structure, and our backend can use the result directly as an AnalysisResult:

We no longer need to remove Markdown code fences or manually parse the returned text into JSON.
Step 4: Return the Result to the Frontend
Since our FastAPI endpoint also uses AnalysisResult, we can return the result directly:

The frontend then receives a predictable structure:

Why Structured Outputs?
A simple prompt such as “respond only with JSON” can work, but it is not an ideal contract between an AI model and an application.
With Structured Outputs, we explicitly define the expected structure.
This provides several benefits:
- The backend receives a predictable data structure.
- The Pydantic model can be used for both the API and the AI response.
- Less manual JSON parsing is required.
- The frontend can work directly with a defined structure.
- The prompt does not need to describe the complete JSON format.
Structured Outputs as a Contract
While building the application, I noticed that the model could sometimes return information that was not actually present in the image, such as an incorrect price or object. This is a common AI hallucination. Structured Outputs does not prevent hallucinations, but it provides a reliable contract between the AI and the application.
By defining the expected fields and types with Pydantic, the backend knows exactly what structure to expect. To reduce hallucinations, I also added explicit instructions to the prompt, such as “Do not invent prices or links. If the information cannot be reliably identified, return an empty value.”
This combination makes the application more predictable: Structured Outputs protects the structure, while clear prompting and validation help protect the data.
OCR or Multimodal AI?
For pure text recognition, Azure provides specialized OCR and Document Intelligence solutions. These are particularly useful when the main goal is to reliably extract text, layouts, tables or positions.
In our case, however, we do not only want to recognize text. We want to analyze the entire image and extract different types of information from it.
That makes a multimodal model such as GPT-4.1-mini an interesting alternative:
Image + instructions → model → structured data
This allows us to combine image understanding and structured data extraction in a single step.
Conclusion
With Azure OpenAI and GPT-4.1-mini, we can relatively easily build an application that analyzes images and converts the results directly into a structure that our software can consume.
The key point is not only the model we use, but also the contract between the AI and the application.
With Pydantic + Structured Outputs, a free-form AI response becomes a structured API response.
This allows AI to do more than just generate text. It can act as an additional processing layer between an image and our application.
I hope this short tutorial was helpful.
Now let’s code!