Penguin Solutions, Inc. Ordinary Shares Rosenblatt's 6th Annual Technology Summit: The Age of AI (Part II)
Review the key takeaways and the transcript of this earnings call.
Transcript
Preview the first fifteen paragraphs, organized by speaker.
Good afternoon, and welcome to the Rosenblatt Securities Sixth Annual Age of AI Scaling Tech Conference. My name is Sajal Dogra. I am one of the semiconductor analysts here at Rosenblatt. It is my pleasure to introduce Mark Adams, Penguin Solutions' Senior Vice President of Global Marketing. Mark has extensive leadership experience in the design and implementation of complex enterprise infrastructure for cloud computing, AI, big data storage, and high-performance computing. We have a buy rating on Penguin with an $80 12-month price target. Penguin is a leading provider of MemoryAI and AI computing solutions for enterprise sovereign AI and Neo Cloud. Penguin is built on 25 years of engineering expertise, bringing together differentiated software, advanced MemoryAI, ComputeAI systems, and industry-leading partner solutions in a full-stack AI factory platform. I will kick off the fireside chat with a few questions, and we will take questions from the audience.
To ask a question, click on the quote bubble graphic on the top right-hand corner, and I will read the question to Mark. Please keep in mind that Penguin is in their earnings quiet period. The purpose of this chat is to understand Penguin's market strategy. Mark, thank you for joining us again this year.
Hey, it is great to be here, Sajal. Thanks so much for having me. Looking forward to the conversation.
Likewise. Mark, back when I used to design chips at Qualcomm in San Diego where you live, Beowulf cluster was this first high-performance compute engineering marvel at NASA that Penguin then commercialized through ClusterWareAI, taking compute and using software to make it work like a scalable, high-performance system. With that historical context, Penguin has undergone a fairly dramatic transformation. The company is now transitioning into an AI factory platform. Can you discuss the steps with us that has brought Penguin to this stage?
Sure. It's a great question, and it's interesting you mentioned NASA and the heritage of where this goes back. Penguin actually was inducted just a couple of years ago into the Space Technology Hall of Fame, going back all the way to that work with NASA. The origins of this were really around as supercomputing was evolving, was this idea of taking scalable nodes and then being able to assemble those nodes into very high-performance compute clusters that could take large problems, break them down into smaller problems, work on them in scalable parallel ways, and then to put the results back together.
What's interesting is that while Penguin has gone through a pretty significant transformation, primarily from being a more hardware and maybe research organization-centric, high-performance computing solution provider to where we are today, operating at the intersection of really large memory solutions, which we'll talk about a little bit later, and AI infrastructure. The core infrastructure and the core architectural model that really underlies what's going on today in this explosion of AI really is actually quite consistent and evolved from that original model. What we bring to this is this 25 years of heritage and skill set that we've been honing for this, and that we have been now in the right place at the right time, where the thing we've been training for for over two decades has evolved as being this skill set that almost every enterprise now is looking to take advantage of.
That's awesome. If I try to understand Penguin versus competitors, there's these open agnostic end-to-end solutions versus proprietary vendor lock systems. Penguin just became an NVIDIA AI factory partner. How does that work, and will ClusterWareAI and OriginAI support AMD racks at a similar feature level, or does it need changes in software and tooling?
That's another good question. I think philosophically, one of the things Penguin's always done, has tried to operate with a mindset of living with a foot in the world of our customer, and really thinking about the workload first, the customer first, and then really looking at the palette of technologies that are available, and trying to think about how can we best pick from everything that's out there to assemble for the customer's objective, the best possible solution. Rather than just no matter what problem someone shows up with, taking the same architecture and the same solution and jamming it down, we like to think that we take a much more thoughtful approach of combining the best of the best. Now, with that much said, NVIDIA's done tremendous work.
As you mentioned, we actually for some time have been a DGX-ready managed services partner, which means that we've demonstrated the top level of experience in managing and operationalizing large AI clusters. They announced a new program that is an AI factory specialized partner, specific to producing these large AI factories, and we were an initial inductee into that. We've got tremendous depth specific to NVIDIA clusters. In fact, we manage these alongside our customers and have accumulated over 4 billion hours of runtime experience of managing GPU clusters at scale. With that much said, we're also completely open. Just as you know in our heritage, we were delivering both Intel-based CPU architectures and AMD-based CPU architectures.
In today's world, where there's an explosion of silicon options coming available for customers, we're equally happy to bring to a customer that needs it an NVIDIA cluster, an AMD cluster, or other clusters. One example that I think we talked about in a recent earnings call that we did was the delivery of a cluster to Sandia National Laboratories, a system called Spectra. Which was based on a Kneron silicon advanced accelerator, which was very nicely mapped to the workload that the customer wanted to run. That's another example, even beyond NVIDIA or AMD, where there's other silicon becoming available. I think agnostic is maybe too generic a word. I think we're opinionated in many regards, but we think that we're opinionated as an advocate for the customer. We'll look at what they're trying to do and recommend what we think would work best.
You mentioned these billions of GPU runtime experience in massive clusters. How does that inform the value in ClusterWareAI that can easily be recreated by a competitor? Is it a learning flywheel, or is it just purely a moat based on all the engineering knowledge and best practices you've had?
Yeah, it's been an interesting journey as we've gone along. As we implemented these large GPU systems really starting in 2016, 2017, before really this whole AI mania that's gripped the world had even fully crystallized. We were outputting these systems together using that HPC architecture, but now with the incorporation of GPUs. As you start to do that at that scale, you start to learn not just about how these systems run, but you start to learn about how they fail. Ideally, you start to learn about how they start to signal the potential for failures of individual components before those components actually fail.
What we've really done is taken the learning that as each of those failures would occur, we would go back after the fact and say, "Was there anything in the logs we could have seen that would have given us a clue that that GPU was going to fail or that that network adapter was going to fail?" Let's continue to compile all of that learning. Then really what's happened now with our core software platform called ClusterWareAI that we use to deploy and manage these systems, is that we've distilled all of that learning, just as if you had a top engineer who was your best data center engineer who just had a knack for keeping your systems running well.
We've tried to take that learning and using analytics techniques, distill all that learning into the software. So now while we do have top services engineers who work in the data center alongside our customers managing these systems, you've got this silent assistant that's running with ClusterWareAI that's doing things that no person could do, and that it can watch millions of data points at every minute of the day as a cluster runs and being able to start to spot, you know what? There's going to be a network issue over there. Why don't we gracefully take that node out of the cluster before it fails and maybe causes disruption to a user job and put it on the side where a technician can repair that and gracefully put it back in.
The end result of this whole thing we're talking about here is higher operational throughput for our customers, fewer disruptions, more uptime as they run these very large systems. By the way, we're continuing to learn every day. I don't want to even make it sound like we're done. As new GPUs become available, new system elements become available, we're constantly taking that learning and feeding it back into the engine.
FULL TRANSCRIPT
Continue the full translated transcript in StockNow.
Access every statement, the English original, and speaker-by-speaker history with StockNow Pro.
View the full transcript with ProCall participants
2 people spoke on this call — only 1 are shown here.
PARTICIPANT LIST
View participant details in StockNow.
Log in to see executives and analysts, their roles, and complete speaking history.
Log in to view all participantsKeep exploring
