Michael Wu is GM and President of Phison Technology Inc. (USA), a leading provider of NAND controllers and NAND storage solutions.gettyAgentic AI is moving from concept to implementation. Developers are building systems that can plan, use tools and carry context across multiple steps rather than respond to a single prompt. NVIDIA is expanding the ecosystem around open frameworks like NemoClaw, while OpenAI is advancing model capabilities. Gartner expects 33% of enterprise software applications to include agentic AI by 2028, up from less than 1% in 2024, underscoring how quickly autonomy is moving into mainstream software environment. ​Raw compute power once defined the limits of local systems. Increasingly, memory is becoming the constraint that determines what can run. ​Why Local Agentic AI Is Gaining Traction ​Running AI locally is no longer a niche requirement. For many organizations, it’s becoming a practical consideration tied to data, cost control and responsiveness. ​Keeping inference on-premises can protect sensitive data, reduce reliance on external services, lower repeated-query costs and improve responsiveness for always-on experiences. ​Local deployment gives developers and device makers more control over models, data and runtime environments. It can reduce the variability associated with shared cloud infrastructure and support environments where connectivity is limited or offline operation is required. ​As agentic systems take on more responsibility, these advantages become more relevant as systems interact with proprietary data. Cloud deployment, however, remains the practical choice for many advanced workloads, particularly where larger models or elastic scale are required. ​The Hidden Constraint Behind Agentic Workloads ​Agentic AI raises the bar for what systems must handle. Rather than simply generating responses, these systems maintain context, track state across multiple steps and coordinate actions between tools and data sources. ​This increases memory demand in ways that traditional chatbot use cases don’t. ​A short interaction may require only a limited context window. An agent tracking an evolving task may need to reference prior steps and incorporate new inputs. When multiple agents operate together, each with distinct roles, the memory footprint expands further. ​At Phison, software engineers use agents for code review. As the agents inspect changes, retrieve files and invoke tools, active context grows. When fast memory fills, the system may need to shorten context, recompute cached state, use a smaller model or move work to the cloud, making complex reviews slower and less consistent. This is where many local systems begin to encounter practical limits. ​GPUs have finite VRAM, and system memory is also bounded. Once those limits are reached, developers are forced into trade-offs, like reducing model size, shortening context windows, applying compression techniques or lowering model precision. ​These adjustments enable execution, but they can affect consistency and reliability, particularly in complex, multistep workflows. Over time, the gap between theoretical capability and real-world performance becomes visible. ​When Hardware Runs Out Of Room ​Even with modern hardware, memory constraints can appear quickly. Large models may require tens of gigabytes just to load, before accounting for context, caching or runtime overhead. ​As a result, many local deployments default to smaller models, not always because they’re sufficient, but because they’re the only viable option with available memory. ​For agentic workloads, this creates a structural mismatch. The tasks benefit from more capable models, longer active context or multiple concurrent sessions, while the hardware environment constrains what can realistically run. ​Developers and system designers are left navigating trade-offs that include smaller or more heavily quantized models, shorter context windows, lower concurrency, more recomputation or larger and more expensive hardware. ​None of these options scale easily, particularly for mainstream device environments. ​Rethinking How Memory Is Used ​Addressing this challenge requires rethinking how memory is utilized across the AI stack, rather than relying solely on additional high-speed capacity. ​Not all data needs to reside in the fastest memory at all times. Frequently accessed data benefits from proximity to compute, while less active data can be managed in higher-capacity, lower-speed tiers. ​This hierarchy already exists in traditional computing. What is changing is how it’s applied to AI workloads, particularly those involving large models and long-running processes. ​By aligning data placement with usage patterns, systems can extend their effective memory footprint without requiring proportional increases in costly resources. ​This doesn’t eliminate performance trade-offs, but it can make more demanding workloads feasible within constrained environments. ​For agentic AI, that flexibility supports longer context retention, more stable state management and more consistent execution across multistep workflows. ​What This Means For The Next Generation Of Devices A common mistake is evaluating AI-ready hardware by peak compute alone, without considering the model, active context, concurrency and runtime overhead the system must support. ​As agentic AI evolves, device makers must decide between prioritizing maximum hardware capability, often at higher cost, or delivering meaningful performance within more scalable environments. ​Memory is central to that balance. ​When evaluating AI-ready hardware, technology leaders should ask whether a system can support the models, context windows and persistent workloads they expect to run over time. NPUs can improve efficiency, but memory capacity and utilization often determine how much of that capability can be used in practice. ​Compute performance continues to improve across CPUs and GPUs. Without sufficient memory to support larger models and persistent workloads, however, those gains may not translate into usable capability. ​This is particularly relevant for laptops and edge systems, where power, cost and physical constraints are more tightly managed than in data centers. ​A more flexible approach to memory can help bridge that gap, enabling systems to advanced workloads without disproportionate increases in hardware complexity. ​Looking Ahead Agentic AI reflects a broader shift toward systems that can operate across steps, interact with tools and adapt over time. ​An underlying layer determines what runs, where it runs and how reliably it performs. Memory sits at the center of that layer. ​As local AI evolves, attention will increasingly move beyond compute toward how effectively systems manage and extend memory resources. ​That shift may be less visible, but it’ll play a meaningful role in determining how broadly these systems can be deployed in practice. ​In many cases, it’ll help determine whether agentic AI becomes a practical capability in real-world environments. ​Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?