Use Images with Modern Machine Vision
There are many companies with vast amount of images that are either include poorly or not at all metadata. The way to find a picture is based on luck or remembering the long long file path where the picture is located. You can imagine that with just 1000 pictures, it’s really hard to keep up what you’ve got or how to find a picture. Or in the case someone else would need need to find a picture.
Using machine vision for image data
Not so long ago OpenAI introduced us something called GPT-Vision. This is a full featured machine vision solution to see what’s in your pictures or videos. We all remember Dall-E and Midjourney for create images with stable diffusion, but this goes other way around. In the past this was the luxury of highly trained machine learning models. Models where trained to recognise predefined items in pictures. Now with gpt-vision we have the capability to label what is included in the image. We can also ask to create meaningful descriptions. The semantic meaning and lwhat is happening in the picture, is something that helps us human’s a lot. When we search pictures, we want to find that picture that includes the scenery or items we required. Instead of asking rain, people, joggins shoes, we like to ask a person running in rain. Or instead of measurement, wall, window glass, we want to ask how to measure windows glass size. To achieve this we need to extract the meaningful data from the images and be able to execute semantic searches agains our dataset. This can be again achieved with the retrieval augmented framework as described in the earlier posts.
Metadata for images
We created a predefined model on Buildgrid platform to handle image data and to create a image search solution. You just upload your images through our platform and we’ll take it from there. We will securely store the pictures in Azure blog storage, handle format conversion and populate metadata for the images. After this you have a image client where you can search images like you would do a google search. It’s super fast as the data is indexed when processed and you ca easily access the images you require. We will also handle the indexing and updates. You probable have files with same names, different file types and in different places. To avoid file name collisions and corrupted data, we will manage this for you. You can just dump all the pictures needed into the system and see the full index of what’s included in your model. Then you can remove or upload more. We will check if the picture already existist or are they new and also remove the ones you ask.
Image labeling
Another useful application is image labeling. Your use-case is that you would like to find all images that are invoices or all images that are invoices from certain vendor. This is also achievable. With the image model you can label, fetch and enrich. This method is really helpful and can be used to automate handling of scanned content to transform it to machine readable json or XML data.
Final thoughts
I think image extraction will be a game changer in many business processes. There are many use-cases where humans or pre-trained models handles this data. Its really vulnerable for small differences in documents. And fixing it is expensive. It’s easier to implement a modern gpt-vision solution to translate and get the wanted information from the image data.