WebLLM: The Rise of AI That Runs Directly in Your Browser
For the last few years, the dominant architecture for generative AI has been straightforward:
Your application → Cloud API → Large Language Model → Response
Every time you interact with an AI application, your prompt or data is typically sent to a remote inference service.
But a different architecture is emerging:






