WebLLM: The Rise of AI That Runs Directly in Your Browser

For the last few years, the dominant architecture for generative AI has been straightforward:

Your application → Cloud API → Large Language Model → Response

Every time you interact with an AI application, your prompt or data is typically sent to a remote inference service.

But a different architecture is emerging: