vLLM runs on 500K GPUs as co-founder Simon Mo raises $150M for Inferact, making the case for open-weight models in production AI infrastructure.

Low-Memory LLM Inference: Meet AirLLM As open-source Large Language Models (LLMs)...

vLLM runs on 500K GPUs as co-founder Simon Mo raises $150M for Inferact, making the case for open-weight models in production AI infrastructure.