Skip to content
RAG Data AI

Talk with Your Data

by Fredrik

The basic use-case for RAG (retrieval augmented generation) is to upload data, convert to embeddings and provide large language models the right context. This approach ensures accurante answers and will increase your productivity by ten fold. In many cases is this called “Talk with your data” or Q&A solution with AI. The framework behind is RAG, as described in the blog Harnessing the Power of RAG: Streamlining Business Communication with Advanced AI . You can check more details there.

Website data

One example could be to add a website or websites to your vector database. The vector database is where we store the embeddings that are number representation of the textual data. RAG solutions usually offer a simple web scraping feature.You can add a base url or sitemap, define the link depth and refresh interval. The system will crawl the given pages for their textual data and add this data to the vector store. And voila! Then we have succesfully created a context for our RAG model and can start using it. One great feature here is also that you will get a citation to the original document. In this case if would be a web site URL where you could check the origin if needed. This is one of the big advantages in RAG models.You don’t have to trust a black-box answer.You can actually check the exact wording used in the original source document. Remember! Crawling pages might be prohibited. You always need to check the ToC of the page before doing it. Check also the chunks that are created. If the crawling function is poorly developed, it includes javascript and style data that will mess up your vector storage. It will slow down searches and consume tokens to produce the answer. OpenAI models are so sophisticated that they will not probably add the scripts and style parts to the final answers. You are still adding unwanted noise that could lower the accurancy of the responses.

Files and storages

Another popular approach is uploading different files or connecting a file storage system to the model. This helps you a lot when you have hundreds or thousands of pages information. And you should be able to find the right answer or generate content that is in line with the given instructions. Files can be in any readable text format, if could be PDF, word, txt, excel, power point etc documents. But please keep in mind the old truth. The output is as good as the input. If you throw in bad quality data, that will not only affect your responses, but also the speed of the solution. Remember! The document readers are good in reading in text based data. But when you have graphs, images, tables, it gets more complicated. Please ensure first that the solution is capable to handle these formats also. E.g. your PDF could be a scanned document and thus actually and image. Then you need OCR features to extract the text from the documents. So it needs to checking and data preparation before you hit upload to your model!

Other datasources?

Now when you’ve got really exited, you’re thinking what else can I do with this! That’s really good thinking and there are really a lot of possibilities. I will cover those more in the coming blogs and tell you there in more details. But in general you can add all kind of data sources and data types. The data needs just to be convert to text and then to embeddings. After that the RAG pipeline is effective and offers you the needed results. If you would have operative business systems, then the right approach is a bit different. That requires functions and API’s to ensure real-time data fetching to create context on fly. Don’t worry, we’ve got this covered via integration partnerships.

Conclusion

Using your own text-based data sources is a great start for your RAG world. It enables you to create virtual assistants, chatbots for customer service. Also internal services and helpful advisers to advise on domain specific topics. No need to ask that one person who’s super busy an annoyed with you asking obvious stuff all the time 😊