Episode Details
Back to EpisodesUnlocking Unstructured Data for Enterprise AI Success
Description
An estimated 80% of all existing digital data isn’t neatly organized in rows and columns, but that unstructured data can unlock significant value and fuel enterprise AI applications, explains data solutions specialist Kailin Hart.
Get tech leader insights to move faster and smarter.
Get more stories by subscribing to The Forecast.
Video transcript:
Kaitlin Hart: I’m Caitlin Hart, director of strategic partnerships at Pryon. I help build together ecosystems and technologies with our partner to build full value driven solutions for our customers.
Ken Kaplan: Now tell me how you got into this.
Kaitlin Hart: Well, I’m a data nerd. I’ve been a data nerd in all kinds of different capacities of the market for almost 20 years. I have spent a lot of time in machine learning before it was popular. And then most recently I’ve moved into unstructured data transformation at Pryo.
Ken Kaplan: And what pulled you to the other side?
Kaitlin Hart: To the other side of data? Well, so I spent so much time in structured data trying to help customers solve really critical problems there, but it’s only 20% of all data is actually structured. And so the types of problems that they would often talk to me about, data discoverability and being able to build more full applications, we actually couldn’t serve in structured data. And so when you think about 80% of all data is unstructured, it’s a huge opportunity for us to drive the needle, deliver more value. And that’s really what we’re doing here at Prime.
Ken Kaplan: And what is some of the unstructured data examples that people wanted to tap into and weren’t able to?
Kaitlin Hart: It’s PDFs, it’s Word docs, it’s websites. It could be research documents. It could be legal briefs. It could be really any … There’s about … I’m trying to think about dozens of different types of unstructured data. Even getting out to multimodal type of content like video data, audio. It’s a very large mass of data. And the difference between unstructured and structured data, structure is very organized, right? So it’s columns and rows and it’s very lightweight and size. Unstructured data because it varies so much in the type of content it actually is. So then does the size of that data. So it makes it very complex to actually transform it at scale.
Ken Kaplan: How are people coming to Priyan and what do they want and how do you help them?
Kaitlin Hart: Yeah, people who are focused on security, they’re really driven to us because we take that as a day one priority. We’ve been engineered from day one not to train on customer data to make sure that their IP is protected at scale. And that’s how we have customers like Nvidia and Leidos and some of the market leaders who take their data super seriously.
Ken Kaplan: And take us under the hood a little bit of … Describe how it works.
Kaitlin Hart: Yeah. So well, I’m a data nerd, so I don’t want to go too deep or I might lose everybody. But when you think about transformation, ETL, structured data, it’s a similar process. It’s a lot more complex because again, that data’s a lot more complex. So there’s a series of different types of models, machine learning based models that we’ve developed ourselves that handle the vectorization, the semantic understanding, query routing to question mapping. All that stuff is actually really complex, especially again at scale and to do that out of the box without training on a customer data. So we do all of that and at the very end we have the LLM, which does the smoothing and consolidation of the answer and we’re agno