Episode Details

Back to Episodes

Own Your Token Machine

Published 2 weeks ago
Description

In this video interview, Liqid Founder and CTO Sumit Puri argues the path to affordable inference runs through composable AI infrastructure that dynamically taps into pools of GPUs, CPUs and DRAM while optimizing usage to manage costs. He says Liqid allows a single server to access a pool of 30 AMD GPUs to meet performance-hungy AI needs.

Get tech leader insights to move faster and smarter.

Get more stories by subscribing to The Forecast.

Video transcript:

Sumit Puri, Founder & CTO, Liqid: Since we spoke to you guys last, one of the major trends that we’re seeing is the enterprises are finally past the initial exploration phase. Now they’re in the adoption and deployment phase of their journey of AI. And I think a lot of them are trying to figure out how exactly they are going to do this journey. And the first crossroads that they’re up against is, am I going to take all of my AI capability and move it into the cloud? Or am I going to own the infrastructure on-prem or in a colo that’s required to do AI? And what it all comes down to is tokenomics. At the end of the day, these companies have a limited amount of dollars and a limited amount of power in many cases that they can deploy. And what they’re trying to figure out is how do I get the most amount of tokens out of that very, very scarce resource?

And so those are the discussions that customers are having now. And the cost of these tokens is front in mind for them. If we notice what’s happening in the industry, large organizations are now saying, “Hey, we’re going to limit the amount of tokens that you as an engineer inside of our organization can consume because the cost of that is becoming very high.” And so the one great way to address that increasing cost is to own your own token machine. And so that’s one of the conversations that we are having with our customers is one way to reduce the cost of those tokens is bring the infrastructure, bring the GPUs on-prem or in a colo so it’s a one-time expense and you can consume all the tokens that you want out of that investment that you make. It is about the model that you are looking to deploy and you’re going to, in any environment, you’re going to have a variety of models.

[Related: Rise of AI Agents Forges IT Industry Partnerships]

You’re going to have big models, you’re going to have small models, you’re going to have medium sized models. The model actually, the size of the model will dictate the type of infrastructure that you need. If I’m running a very, very large model, I’m going to need a large quantity of GPUs in order to run that model. And so you have to figure out if I’m going to build a very big system, how do I get those large quantity of GPUs into play? One way to do it is independent scaling of resources. We come in and let customers say, “I don’t want to scale my compute as I’m scaling my GPUs. Allow me to just scale GPUs.” As an example, today we’re at the AMD show. The announcement that we’re making today is our ability to take a single server and scale up to 30 AMD GPUs to that single machine.

And the benefit of that is we can now start to run these very large frontier models on a single server infrastructure. These models are so big. Previously you had to run them in a cluster. Now we come in and say we can take those models that were previously clustered and consolidate them down to a single server, a single instance, back to that tokenomic story. That’s how we reduce the cost of deploying these models. Models will change over time. And so some cases you might want a very small model that will require, let’s say, a single GPU. Sometimes you will wa

Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us