GSI Technology CG 46th Annual Growth Conference
Review the key takeaways and the transcript of this earnings call.
Transcript
Preview the first fifteen paragraphs, organized by speaker.
Hi, everyone. Thanks for joining the session. I am Kingsley Crane, a technology analyst here at Canaccord. Really pleased to have the GSI Technology team with us here today. DDA, thanks for joining. Thanks for having me.
Let's kick it off. You just reported your June quarter. You have described this as a pivotal point. What does GSI look like today versus 18 months ago, and just at a high level?
Yeah. We certainly progressed the company quite a long ways. 18 months ago, we were just coming out with Gemini-II, which is really our first commercialized part for the AI market. Fast-forward to now, since then, the Gemini-II is in production now as far as the hardware is concerned. We have gotten some third-party validation from folks like Cornell. Cornell University actually had a board of ours and did a RAG comparison with an NVIDIA GPU, and it found at comparable performance, we were 98% less power, so it was certainly a huge movement in the market space. Since then, we have also engaged in a few POCs. We have one drone surveillance POC that is being funded by the DoD.
We also have recently won a phase 1 POC out of a municipality in Taiwan for Smart City, which is nice because it leveraged a lot of what we're doing for the POC for the drone, so it's not a complete new lift. Since then, we've also gone quite a ways with developing an AI SDK, or AI-assisted SDK, I should say. This is important because it's going to help us enable the ecosystem. If you look at our model now, we are writing our own applications for all these POCs, which is fine to really showcase the technology, but to really scale, we're going to have to have tools that the customers can use. We are quite a ways with that SDK. We'll have our alpha version out in this fall. Let me see what else have we done since then.
We've won a couple SBIRs, one with the U.S. Army, for ruggedized edge node. This could be a nice product for us and for the Army. Essentially, it's going to be a server that can do object detection, or it can do SAR imagery, and it can be done, again, at the edge. So it's a ruggedized server they can put at the back of a Humvee or something. Let me see, 18 months. We've also started our next-generation device, the Plato. Plato is going to be. It leverages some of the technology, obviously, from Gemini-II, but it's going to adjust a different market.
The way that Gemini-II is that it's. We have a small bandwidth coming from memory into our chip because really, the intent of the chip is to download a model or a database one time, and then once it's in our chip, we run it really, really fast. So our internal bandwidth is extreme, so it's perfect for search applications or HPC kind of applications. Plato, on the other hand, is going to be developed more for LLMs at the edge. We're going to open up the pipe so that we can get data into the part faster, and then we'll scale down the internal bandwidth to match that. That'll actually be used for vision language models, large language models at the edge.
It's going to be extreme because it's going to have almost data center performance at a 2-10-watt power budget. So it's going to be a powerhouse for edge applications.
Really helpful overview. Investors are trying to understand the AI supply chain and target the bottleneck. Can you just help us understand why SRAM is so important for AI, how durable that business is, and what is a cyclical, AI super cycle this time?
Sure. I'm going to answer that in a couple of different ways. First of all, it's important to our company because right now it's the cash cow. We're just starting the AI story, and so we need something to offset the bills and we've been doing SRAMs now for 30 years, shipped over 140 million devices. So we're certainly a leader in that market. It's really helped offset the bills or the cost there. As far as the market itself, the SRAM isn't directly in the AI, it's not in a data center, but it really helps the infrastructure. What I mean by that is one of our largest customers we've talked about is Cadence. They make emulation systems. These are systems that if you're an IC manufacturer or designer, I should say, you do a design and in the old days, you just do a simulation.
Then, is the part working? Do I have any major bugs? It wasn't until you went to first silicon, got the actual chip, you looked at it and said, "Wow, it's dead on arrival. Part doesn't work." That's a real problem nowadays because you have to create a mask set in order to get that chip. The mask sets cost upwards of $30 million a mask set, so you don't want to be wasting a lot of $30 million mask sets. Now guys like Cadence make these emulation systems where they actually emulate the design in software, and these are large, expensive systems, and our highest-end parts go into those systems.
We're really helping the front-end design. Another one of our larger customers is KYEC. They actually are part of the manufacturing process. In this particular case, they do burn-in. Burn-in is basically a way to electrically really challenge a part to make sure there's what's called no infant mortality rate. In other words, you don't want a part to be put in a system, go out to the field, and then fail because it's a weak chip. So they do burn-in to try and get rid of those. KYEC is doing the back end for all the latest GPUs that are out there, and they use one of our high-end parts. So we're indirectly supporting the AI with our SRAM division.
Can you help us understand why compute-in-memory is so critical from a performance and cost perspective?
There is this thing called von Neumann model. I do not want to get too much into detail on it, but essentially, if you look at the way a GPU and a CPU work, they have their processing elements. When they are being asked to do something, do a calculation or some kind of process, they have to go outside the chip to fetch data from memory, bring it back, use it, and once they have used it, they have to write it back to memory. There is this constant data transfer, data flow back and forth. It takes a tremendous amount of power, which I am sure you have all heard, data centers, what is their biggest problem? Power. With the APU, we have done something much different, and this is all patent protected because we know we have something unique is with the Gemini-II family, as I mentioned, we have a small pipeline going into the chip.
We bring in data one time, and once it is there, we run it very fast. The way we do that is we have, first of all, a large memory inside, but the process or the calculation is actually done in the memory bit line, in the memory itself. Our bit processors are actually coupled with our memory. We are not going out fetching data, we are not bringing it back. We do not have this constant transfer, and so the performance is high, but the power is extremely low. As I mentioned, this Cornell paper, 98% less power, and that is because of that CIM architecture.
For Gemini-II, can you help us get a sense of that path towards commercialization, just moving from proof of concept into product, and then what do you think success looks like in a year from now?
FULL TRANSCRIPT
Continue the full translated transcript in StockNow.
Access every statement, the English original, and speaker-by-speaker history with StockNow Pro.
View the full transcript with ProCall participants
2 people spoke on this call — only 1 are shown here.
PARTICIPANT LIST
View participant details in StockNow.
Log in to see executives and analysts, their roles, and complete speaking history.
Log in to view all participantsKeep exploring
