Penguin Solutions just gave its ClusterWareAI platform a substantial facelift, rolling out an AI-powered operations agent and automated GPU management tools designed to make running massive GPU clusters less of a nightmare. The update, announced on June 25, turns what was already a full-stack AI factory operating system into something that can talk back to you in plain English about what your hardware is doing.

The company, which trades on NASDAQ under the ticker PENG, has deployed nearly 100,000 GPUs and has accumulated over four billion hours of GPU runtime experience.

What the update actually does

The centerpiece of the ClusterWareAI upgrade is what Penguin Solutions calls an AI Factory Operations Agent. Think of it as a conversational layer sitting on top of your GPU cluster, letting operators ask questions about performance in natural language instead of digging through dashboards and log files.

The second major addition is automated remediation for Kubernetes-based inference workloads. When something breaks in a GPU cluster running inference tasks, downtime translates directly into lost revenue and wasted compute. Automated remediation means the system can detect problems and fix them without waiting for a human to intervene.